Staff AI Telemetry Engineer
Location: Remote
Compensation: To Be Discussed
Reviewed: Thu, Aug 27, 2026
This job expires in: 27 days
Job Summary
To architect the "nervous system" of the AI-native NeoCloud platform, the full-time Staff AI Observability & Telemetry Engineer will build a high-fidelity telemetry infrastructure that captures and analyzes millions of hardware and software signals per second, enabling real-time performance insights for GPU workloads and network fabric.
Key Responsibilities
- Architect and scale a high-cardinality telemetry infrastructure using time-series databases capable of handling massive ingestion rates
- Integrate complex hardware-level exporters into the Kubernetes observability stack for a unified view of the cluster
- Develop automated dashboards and alerting pipelines to proactively manage degraded hardware impacting customer jobs
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field
- 6+ years of software or site reliability engineering experience, with expertise in the Prometheus/OpenTelemetry ecosystem
- Advanced proficiency in Go and experience writing custom Kubernetes metric exporters
- Hands-on experience with kernel-level tracing tools (eBPF, BCC) and performance tuning of Linux systems
- Strong familiarity with AI hardware metrics and high-performance network telemetry
Complete Job Description
The complete job description is available to members. Premium membership includes:
Full access to 44,387 remote jobs from human-vetted companies, updated daily
Resume Builder - AI-powered tool to craft, enhance, and tailor your resume to a specific job
Twice-monthly live group coaching and the full Remote Career Center
20% member discount on Career Services
Backed by a 30-day money-back guarantee