Network SRE Expert
Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Aug 27, 2026
This job expires in: 30 days
Job Summary
To support the development of an AI-operated GPU cloud, the full-time Network SRE Expert will manage InfiniBand and RoCEv2 networks, ensuring optimal performance and reliability for GPU clusters while working remotely from San Jose, CA or Austin, TX.
Key responsibilities
- Oversee the deployment and management of InfiniBand fabrics and RoCEv2 networks for GPU clusters, ensuring efficient data flow and performance tuning
- Implement monitoring and diagnostics using UFM and other tools to identify and resolve network issues, including link degradation and congestion
- Collaborate with platform teams to enhance AIOps capabilities by integrating telemetry data and developing predictive maintenance workflows
Required qualifications
- 5+ years of experience in data center networking, with a minimum of 3 years focused on InfiniBand or RoCE fabrics
- Hands-on experience with Nvidia/Mellanox InfiniBand switches and RoCEv2 deployment, including relevant configurations
- Strong understanding of IB subnet management, QoS, and performance monitoring tools
- Proficiency in diagnosing network issues using tools such as ibdiagnet and perfquery
- Familiarity with NCCL and its implications for GPU communication and network topology
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 47,980 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee