Demo

Staff Observability Platform Engineer (SRE)

Recruitment.ai
Houston, TX Contractor
POSTED ON 8/20/2026
AVAILABLE BEFORE 9/18/2026
Staff Observability Platform Engineer (SRE)
Locations: Seattle, WA (Hybrid), Houston, TX (Hybrid), New York, NY (Hybrid)

What You’ll Do
  • Design, build, and evolve observability platforms across metrics, logs, traces, alerting, and telemetry pipelines.
  • Lead the implementation of scalable observability solutions that support Nscale’s growing GPU and AI infrastructure.
  • Partner with SRE, infrastructure, platform, and AI/ML teams to ensure observability is embedded throughout the software and infrastructure lifecycle.
  • Drive improvements in monitoring coverage, alert quality, service health visibility, and incident response effectiveness.
  • Develop standards, frameworks, and reusable patterns that simplify observability adoption across engineering teams.
  • Identify reliability risks and operational blind spots, helping teams proactively address them before they impact customers.
  • Contribute to architectural decisions around telemetry collection, storage, retention, cardinality management, and performance optimization.
  • Lead technical initiatives and projects that improve platform scalability, reliability, and operational efficiency.
  • Mentor engineers and provide technical guidance through design reviews, code reviews, and knowledge sharing.
  • Participate in incident investigations and postmortems, translating operational learnings into durable platform improvements.
  • Evaluate new observability technologies and practices, balancing innovation with operational simplicity and long-term maintainability.
About You
  • Experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines.
  • Strong experience building and operating observability platforms in cloud-native, distributed environments.
  • Deep hands-on experience with several of the following technologies: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic, or similar platforms.
  • Strong software engineering skills with proficiency in Go, Python, or equivalent languages.
  • Experience operating and troubleshooting Kubernetes-based platforms at scale.
  • Strong understanding of monitoring, logging, tracing, telemetry pipelines, and modern observability practices.
  • Experience designing systems with scalability, reliability, performance, and operational simplicity in mind.
  • Proficiency with Infrastructure-as-Code tools such as Terraform, Ansible, or equivalent.
  • Ability to lead technical initiatives and influence engineering decisions across multiple teams.
  • Excellent communication skills with the ability to explain technical tradeoffs and align stakeholders around pragmatic solutions.
Preferred
  • Experience operating observability systems in GPU, AI/ML, HPC, or large-scale compute environments.
  • Familiarity with Slurm, Kubernetes GPU scheduling, or AI infrastructure platforms.
  • Experience with high-volume telemetry pipelines and streaming technologies such as Kafka, Vector, or Fluent Bit.
  • Knowledge of observability challenges related to model training, inference workloads, GPU utilization, and distributed AI systems.
  • Experience mentoring engineers and helping grow technical capability across teams.

Hourly Wage Estimation for Staff Observability Platform Engineer (SRE) in Houston, TX
$37.00 to $43.00
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Staff Observability Platform Engineer (SRE)?

Sign up to receive alerts about other jobs on the Staff Observability Platform Engineer (SRE) career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$97,257 - $120,701
Income Estimation: 
$123,167 - $152,295
Income Estimation: 
$76,670 - $90,826
Income Estimation: 
$91,609 - $118,978
Income Estimation: 
$92,877 - $110,401
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Recruitment.ai

  • Recruitment.ai York, NY
  • Job Title: Staff Observability Platform Engineer (AI / GPU Infrastructure) Location: Seattle, WA / Houston, TX / New York, NY (Onsite) Employment Type: Con... more
  • 11 Days Ago

  • Recruitment.ai San Jose, CA
  • Project Overview Building new platform for secured environment Commercial and government deployments AWS-based: Kubernetes, S3, EKS, SQS, SNS Some infrastr... more
  • 11 Days Ago

  • Recruitment.ai Seattle, WA
  •  8 years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.  You’ve operated observability infrastructure at serio... more
  • 12 Days Ago


Not the job you're looking for? Here are some other Staff Observability Platform Engineer (SRE) jobs in the Houston, TX area that may be a better fit.

  • Persona AI Houston, TX
  • Job Title: Staff Platform Engineer - Developer Infrastructure Department: Software Employment Type: Full-Time Location: Houston, TX Travel: 10% Who We Are ... more
  • 2 Days Ago

  • Affirm Houston, TX
  • Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or comp... more
  • 6 Days Ago

AI Assistant is available now!

Feel free to start your new journey!