Site Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Fri, Jul 24, 2026
This job expires in: 29 days
Job Summary
Joining the Platform & Production Reliability team, the full-time Application Site Reliability Engineer will manage the reliability and performance of .NET/C# services on Windows, focusing on operational excellence and incident response in a remote environment.
Key responsibilities
- Participate in the on-call rotation for production trading systems and lead incident response during service disruptions
- Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues
- Build and maintain Grafana dashboards and Prometheus alerts to monitor application and infrastructure health
Required qualifications
- 3-5 years of experience debugging and supporting .NET/C# applications in production
- Strong PowerShell scripting skills and experience with Python or Bash
- Hands-on experience with Grafana, Prometheus, and AWS
- Familiarity with CI/CD pipelines and deployment strategies
- Experience troubleshooting Aurora PostgreSQL or other relational databases
COMPLETE JOB DESCRIPTION
The job description is available to subscribers. Subscribe today to get the full benefits of a premium membership with Virtual Vocations. We offer the largest remote database online...