Site Reliability Engineer
Location: Remote
Compensation: Salary
Reviewed: Mon, Jul 27, 2026
This job expires in: 29 days
Job Summary
Joining a remote team, the full-time Site Reliability Engineer will take ownership of the operational health, reliability, and availability of a FedRAMP High cloud platform, proactively identifying and resolving issues while driving continuous improvement in collaboration with engineering teams.
Key responsibilities
- Monitor the health, availability, performance, and security of production services
- Investigate, troubleshoot, and resolve complex production incidents as the L3 escalation point
- Develop and maintain operational runbooks, dashboards, and standard operating procedures
Required qualifications
- 3-6 years of experience in Site Reliability Engineering, Production Engineering, or a senior production support role
- Strong understanding of Linux operating systems and networking fundamentals
- Experience troubleshooting distributed applications running in Kubernetes
- Proficiency in scripting or programming languages such as Python, Bash, or Go
- Experience with monitoring and observability platforms like Prometheus or Datadog
COMPLETE JOB DESCRIPTION
The job description is available to subscribers. Subscribe today to get the full benefits of a premium membership with Virtual Vocations. We offer the largest remote database online...