Senior AI Infrastructure Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Mon, Jul 20, 2026
This job expires in: 22 days
Job Summary
To support the management of expansive AI ecosystems, the full-time Senior AI Infrastructure Engineer will lead technical operations, service reliability, and platform engineering efforts in a remote environment, focusing on NVIDIA GPU infrastructure and Kubernetes operations.
Key responsibilities
- Lead the investigation and resolution of complex infrastructure and platform-related incidents, acting as a senior escalation point during critical events
- Drive improvements in platform reliability and operational processes while identifying opportunities for automation
- Mentor AI Infrastructure & Platform Operations Engineers and develop operational standards and best practices
Required qualifications
- 7+ years of experience in infrastructure operations, site reliability engineering, or related technical roles
- Expert-level Linux administration and troubleshooting skills
- Strong experience operating Kubernetes in production environments
- Proven experience leading technical investigations and managing complex incidents
- Strong understanding of observability, monitoring, and service reliability practices
COMPLETE JOB DESCRIPTION
The job description is available to subscribers. Subscribe today to get the full benefits of a premium membership with Virtual Vocations. We offer the largest remote database online...