Cloud Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Mon, Oct 05, 2026
This job expires in: 30 days
Job Summary
To ensure operational readiness and reliability for production systems on a federal cloud platform, the full-time Cloud Site Reliability Engineer (SRE) will design automated remediation processes, define service-level objectives, and enforce infrastructure-as-code practices, working both onsite in Arlington, VA and remotely.
Key responsibilities
- Design and implement automated systems for self-healing operations to enhance reliability
- Define data-driven service-level objectives and error budgets for critical services
- Enforce infrastructure-as-code standards and drive a containerization-first approach using Kubernetes
Required qualifications
- Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent practical experience)
- 5+ years of SRE experience with demonstrated technical leadership
- Expertise in AWS and GovCloud, along with deep Kubernetes and Terraform knowledge at production scale
- Strong software engineering background in Python and/or Go
- Proven experience with observability platforms and incident management
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 38,620 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee