Platform Reliability Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Mon, Sep 28, 2026
This job expires in: 30 days
Job Summary
To support enterprise AI platforms, the contract Platform Reliability Engineer will build observability, implement OpenTelemetry instrumentation, and manage cloud cost monitoring while working remotely.
Key responsibilities
- Design and implement OpenTelemetry-based instrumentation and centralized logging across AI agents and services
- Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability dashboards for platform components
- Configure proactive alerts for service degradation, execution failures, and resource utilization, supporting incident management and operational runbooks
Required qualifications
- 7+ years of experience in platform engineering, SRE, DevOps, or production infrastructure operations
- Hands-on experience with OpenTelemetry SDKs and telemetry pipelines
- Strong experience in implementing distributed tracing, metrics, logging, and alerting
- Experience monitoring AWS infrastructure and services, including CloudWatch and EKS
- Proficiency in Python or another scripting language, with experience in Terraform or equivalent IaC tools
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 38,646 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee