Senior Cloud Infrastructure Engineer
Job Summary
Passionate about cloud infrastructure operations, the full-time Senior Cloud Infrastructure Engineer will guide NVIDIA Cloud Partners in advancing operational capabilities for large-scale NVIDIA accelerated infrastructure, focusing on Day 2 operations, infrastructure health, and observability, while working remotely or onsite in Santa Clara or Seattle.
Key responsibilities
- Lead NCP Day 2 operational readiness efforts, collaborating with partners to establish systems and procedures for managing NVIDIA accelerated infrastructure
- Develop and implement continuous infrastructure validation methods for GPU, CPU, storage, and network health across AI clusters
- Establish observability and operational telemetry, implementing monitoring, alerting, and dashboards for compute and AI workloads
Required qualifications
- BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience
- 8+ years of experience in infrastructure engineering, Site Reliability Engineering, or similar roles in large-scale production environments
- Strong experience with Linux-based distributed systems and cloud infrastructure in production
- Deep understanding of Kubernetes, containers, and the operational lifecycle of large multi-node environments
- Experience with automation for infrastructure lifecycle management and failure detection
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 47,803 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee