What are the responsibilities and job description for the Staff Observability Platform Engineer position at Programming.com?
π Staff Observability Platform Engineer | AI/GPU Infrastructure | Remote/Hybrid
π Hybrid β Seattle, Houston, or New York
Have you personally operated an observability backend at real production scale β not just consumed dashboards built by another team? We're looking for someone who can quantify that scale and explain the engineering decisions behind it.
What you'll do:
πΉ Design, build, and operate large-scale metrics, logging, and tracing platforms
πΉ Own observability backend architecture in distributed Kubernetes and AI/GPU environments
πΉ Operate and scale Mimir, Thanos, VictoriaMetrics, Cortex, Loki, or Elasticsearch at production scale
πΉ Design Prometheus-based architectures (remote write, high availability, retention, global querying)
πΉ Build OpenTelemetry Collector pipelines (receivers, processors, exporters, sampling)
πΉ Identify and remediate high-cardinality metrics
πΉ Establish observability standards across engineering teams
What we're looking for:
β Personal ownership (not just consumption) of a metrics/logs backend at meaningful scale
β Ability to quantify that scale: active series, ingestion rate, TB/day, retention, cluster size
β Real experience making cardinality, retention, ingestion, and storage-cost trade-offs
β Strong hands-on production Kubernetes experience (multi-cluster, networking, autoscaling)
β Strong production code-reading/review ability (Go and/or Python)
β Understanding of retries, timeouts, backpressure, circuit breaking, and failure handling
Nice to have:
π Large-scale GPU fleets, NVIDIA DCGM, InfiniBand/RoCE/NVLink, Slurm/HPC
π Custom Prometheus exporters, custom OpenTelemetry components
π AI/ML infrastructure or distributed training platforms
- π Hybrid β Seattle, Houston, or New York