Software Engineer, AI Infrastructure
Location: Remote
Compensation: Salary
Reviewed: Tue, Sep 29, 2026
This job expires in: 30 days
Job Summary
Focused on bringing up and optimizing distributed training and inference workloads, the full-time Software Engineer, AI Infrastructure will benchmark, analyze, and debug large-scale AI systems across multiple GPU platforms in a remote or onsite environment.
Key responsibilities
- Bring up, validate, and debug large-scale AI clusters and end-to-end workloads
- Benchmark AI pre-training, post-training, and inference workloads using NVIDIA AI software stacks
- Perform root-cause analysis of failures in distributed environments and contribute to resilience tooling
Required qualifications
- Bachelor's or Master's in Computer Science or a related technical field (or equivalent experience)
- Experience developing software for AI, HPC, or systems-level applications
- Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution
- Background with debugging and scaling distributed systems
- Strong programming skills in Python and C/C++
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 39,986 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee