Remote Jobs Sign In

Senior AI Infrastructure Engineer

Location: Remote
Compensation: To Be Discussed
Reviewed: Mon, Jul 20, 2026
This job expires in: 22 days

Job Summary

To support the management of expansive AI ecosystems, the full-time Senior AI Infrastructure Engineer will lead technical operations, service reliability, and platform engineering efforts in a remote environment, focusing on NVIDIA GPU infrastructure and Kubernetes operations.

Key responsibilities
  • Lead the investigation and resolution of complex infrastructure and platform-related incidents, acting as a senior escalation point during critical events
  • Drive improvements in platform reliability and operational processes while identifying opportunities for automation
  • Mentor AI Infrastructure & Platform Operations Engineers and develop operational standards and best practices
Required qualifications
  • 7+ years of experience in infrastructure operations, site reliability engineering, or related technical roles
  • Expert-level Linux administration and troubleshooting skills
  • Strong experience operating Kubernetes in production environments
  • Proven experience leading technical investigations and managing complex incidents
  • Strong understanding of observability, monitoring, and service reliability practices

COMPLETE JOB DESCRIPTION

The job description is available to subscribers. Subscribe today to get the full benefits of a premium membership with Virtual Vocations. We offer the largest remote database online...