Remote Jobs Sign In

Network SRE Expert

Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Aug 27, 2026
This job expires in: 30 days

Job Summary

To support the development of an AI-operated GPU cloud, the full-time Network SRE Expert will manage InfiniBand and RoCEv2 networks, ensuring optimal performance and reliability for GPU clusters while working remotely from San Jose, CA or Austin, TX.

Key responsibilities
  • Oversee the deployment and management of InfiniBand fabrics and RoCEv2 networks for GPU clusters, ensuring efficient data flow and performance tuning
  • Implement monitoring and diagnostics using UFM and other tools to identify and resolve network issues, including link degradation and congestion
  • Collaborate with platform teams to enhance AIOps capabilities by integrating telemetry data and developing predictive maintenance workflows
Required qualifications
  • 5+ years of experience in data center networking, with a minimum of 3 years focused on InfiniBand or RoCE fabrics
  • Hands-on experience with Nvidia/Mellanox InfiniBand switches and RoCEv2 deployment, including relevant configurations
  • Strong understanding of IB subnet management, QoS, and performance monitoring tools
  • Proficiency in diagnosing network issues using tools such as ibdiagnet and perfquery
  • Familiarity with NCCL and its implications for GPU communication and network topology

Complete Job Description

The complete job description is available to members. Premium membership includes:

Full access to 47,980 remote jobs from human-vetted companies, updated daily

Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job

Twice-monthly live group coaching and the full Remote Career Center

20% member discount on Career Services

Backed by a 30-day money-back guarantee