Demo

Staff Observability Platform Engineer (AI / GPU Infrastructure)

Recruitment.ai
York, NY Contractor
POSTED ON 8/22/2026
AVAILABLE BEFORE 9/20/2026

Job Title: Staff Observability Platform Engineer (AI / GPU Infrastructure)

 Location: Seattle, WA / Houston, TX / New York, NY (Onsite)
 Employment Type: Contract


Position Overview

We are looking for a highly experienced Staff Observability Platform Engineer to design, build, and operate large-scale observability platforms supporting AI/GPU infrastructure and Kubernetes environments.

This is not a traditional monitoring or dashboard role. You''ll own the observability backend platform, ensuring it remains scalable, reliable, and cost-efficient as telemetry volumes grow.


Key Responsibilities

  • Design, build, and scale enterprise metricsloggingtracing, and telemetry platforms.
  • Architect and operate distributed observability backends using PrometheusMimirThanosVictoriaMetricsCortexLokiElasticsearch, or similar technologies.
  • Build and optimize OpenTelemetry Collector pipelines including routing, filtering, sampling, and exporters.
  • Optimize cardinalityingestionretentionstoragequery performance, and infrastructure cost.
  • Manage large-scale Kubernetes observability across multi-cluster environments.
  • Troubleshoot production issues involving metricslogstraces, networking, storage, and distributed systems.
  • Write and review production-quality code and establish observability standards across engineering teams.
  • Partner closely with Platform, Infrastructure, Security, and Application Engineering teams.

Required Qualifications

  • 8 years of experience in ObservabilityPlatform EngineeringSRE, or Infrastructure Engineering.
  • Hands-on experience operating production-scale observability platforms such as MimirThanosVictoriaMetricsCortexLoki, or Elasticsearch.
  • Strong expertise with PrometheusOpenTelemetry, and telemetry pipeline design.
  • Experience managing large-scale production environments with measurable metrics (ingestion rates, active time series, storage, retention, cluster size, etc.).
  • Strong understanding of cardinalityretention strategiesstorage architecturesamplingquery optimization, and cost management.
  • Deep hands-on experience with Kubernetes, distributed systems, networking, service discovery, autoscaling, and reliability engineering.
  • Strong programming skills in Go and/orn Python.
  • Solid understanding of production engineering concepts including retries, backpressure, buffering, circuit breaking, graceful degradation, scalability, and failure handling.

Preferred Qualifications

  • Experience with GPU infrastructureAI/ML platforms, or HPC environments.
  • Knowledge of NVIDIA DCGMInfiniBandRoCE/RDMANVLinkNVSwitchNCCL, or Slurm.
  • Experience building custom Prometheus ExportersOpenTelemetry Collectors, or large multi-cluster Kubernetes observability platforms.

Ideal Candidate

We''re looking for someone who has personally owned and operated observability backends at production scale, understands the trade-offs between cardinality, retention, storage, performance, and cost, and can build scalable observability platforms for modern AI/GPU infrastructure. Experience limited to dashboards, alerts, or consuming monitoring tools without backend ownership will not be sufficient for this role.

Hourly Wage Estimation for Staff Observability Platform Engineer (AI / GPU Infrastructure) in York, NY
$73.00 to $86.00
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Staff Observability Platform Engineer (AI / GPU Infrastructure)?

Sign up to receive alerts about other jobs on the Staff Observability Platform Engineer (AI / GPU Infrastructure) career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$101,387 - $124,118
Income Estimation: 
$119,030 - $151,900
Income Estimation: 
$131,953 - $159,624
Income Estimation: 
$169,825 - $204,021
Income Estimation: 
$166,631 - $195,636
Income Estimation: 
$162,237 - $199,353
Income Estimation: 
$181,083 - $218,117
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Recruitment.ai

  • Recruitment.ai San Jose, CA
  • Project Overview Building new platform for secured environment Commercial and government deployments AWS-based: Kubernetes, S3, EKS, SQS, SNS Some infrastr... more
  • 11 Days Ago

  • Recruitment.ai Seattle, WA
  •  8 years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.  You’ve operated observability infrastructure at serio... more
  • 12 Days Ago

  • Recruitment.ai Houston, TX
  • Staff Observability Platform Engineer (SRE) Locations: Seattle, WA (Hybrid), Houston, TX (Hybrid), New York, NY (Hybrid) What You’ll Do Design, build, and ... more
  • 13 Days Ago


Not the job you're looking for? Here are some other Staff Observability Platform Engineer (AI / GPU Infrastructure) jobs in the York, NY area that may be a better fit.

  • CoreWeave York, NY
  • Job Description. Job Description. CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, to... more
  • 27 Days Ago

  • LangChain York, NY
  • About LangChain At LangChain, our mission is to make intelligent agents ubiquitous. We provide the agent engineering platform and open source frameworks de... more
  • 10 Days Ago

AI Assistant is available now!

Feel free to start your new journey!