Senior Lab Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Tue, Sep 01, 2026
This job expires in: 30 days
Job Summary
Owning the operational reliability of VAST clusters in a remote full-time capacity, the Senior Lab Reliability Engineer will manage health monitoring, issue resolution, and automation strategies while serving as the primary technical escalation point for complex cluster issues.
Key responsibilities
- Own the operational reliability of VAST clusters in the lab environment, including proactive health monitoring and upgrade planning
- Serve as the primary technical escalation point for complex cluster issues, collaborating with engineering for deeper investigations
- Shape the automation and tooling strategy for the lab environment, including infrastructure-as-code practices and internal tooling development
Required qualifications
- 4+ years of professional experience in systems engineering, storage engineering, or related roles
- Deep hands-on experience with enterprise storage systems at an operational or reliability level
- Strong Linux systems administration skills, including networking and storage management
- Proficient in scripting/programming (Python, Bash, or similar) for automation and tooling
- Experience with Docker and Kubernetes in operational environments
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 45,346 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee