Recovery and Incident Manager
Location: Remote
Compensation: Salary
Reviewed: Tue, Jul 21, 2026
This job expires in: 29 days
Job Summary
Leading incident management and operational resilience initiatives, the full-time salaried Recovery and Incident Manager with AI-Ops will collaborate with cross-functional teams to enhance system reliability and automate incident response in a hybrid work environment.
Key responsibilities
- Monitor and analyze major incident responses, serving as a senior escalation point for operational incidents
- Implement Site Reliability Engineering practices to improve platform reliability and develop proactive strategies
- Automate incident detection and remediation processes using AIOps solutions and optimize operational workflows
Required qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field (or equivalent experience)
- 8+ years of experience in IT Operations, Site Reliability Engineering, or related environments
- 5+ years of experience leading incident management or operational transformation initiatives
- Strong experience with AIOps, ITSM, ITIL, and cloud platforms (AWS, Azure, or GCP)
- Hands-on experience with monitoring platforms like Dynatrace and ServiceNow for workflow automation
COMPLETE JOB DESCRIPTION
The job description is available to subscribers. Subscribe today to get the full benefits of a premium membership with Virtual Vocations. We offer the largest remote database online...