Demo

Senior SRE (AI)

Altimetrik
Mountain View, CA Full Time
POSTED ON 8/1/2026
AVAILABLE BEFORE 8/30/2026

AI-First SRE (Senior)


Key Responsibilities

  • Build and extend O11Y pipelines: metrics, distributed tracing, logging, SLO/SLI definitions and dashboards consumed platform-wide.
  • Instrument AI systems in production: token usage and cost, latency-per-inference, tool-call success rates, response-quality signals, agent loop detection, runaway-cost alerting.
  • Extend tracing across agentic flows — planner → executor → retrieval → tool calls — spanning GenOS, AI Gateway, and MCP Gateway.
  • Define SLOs where "available" includes response quality and tool-call success, not just HTTP 200s.
  • On-call and incident command for AI platform surfaces: model/provider degradation, Bedrock/Gemini failover, semantic-cache issues, prompt-injection events — blameless postmortems with tracked remediations.
  • Own the SRE side of the progressive-delivery seam: canary analysis, automated rollback decisioning, chaos/resilience testing, blast-radius controls — DevOps builds the pipeline; you define the gates.
  • Build AIOps agents in the Traffic Agent pattern: intelligent triage, anomaly detection, auto-remediation with hard guardrails.
  • Drive UX Availability (R30) and Foundations Capabilities (R30); land network, O11Y, and cloud cost savings (tracked YTD KPIs).


Must-Have Qualifications

  • 7 years SRE/production engineering on Tier-0/Tier-1, high-traffic systems.
  • Strong Go or Python — this is a build role: tooling, automation, instrumentation.
  • Has built observability stacks, not just consumed them: Prometheus/Grafana, OpenTelemetry, or equivalent at scale, including cardinality and cost control.
  • Production LLM/ML monitoring: Langfuse, Arize, WhyLabs, or homegrown — token/cost tracking, drift and quality metrics.
  • Working fluency in AI-system failure modes: nondeterminism, provider limits and outages, context-window overflow, agent loops, cache poisoning.
  • Kubernetes AWS operational depth — debugs across cluster, mesh, and gateway layers.
  • Structured incident-command and postmortem experience.


Nice-to-Have

  • AIOps / LLM-applied-to-ops: auto-triage, incident summarization, remediation agents.
  • Chaos engineering (Litmus, Gremlin, or homegrown); eBPF or deep network debugging.
  • FinOps / cost engineering; fintech or regulated-industry reliability experience.

Salary.com Estimation for Senior SRE (AI) in Mountain View, CA
$160,159 to $198,088
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Senior SRE (AI)?

Sign up to receive alerts about other jobs on the Senior SRE (AI) career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$105,489 - $131,507
Income Estimation: 
$128,913 - $157,494
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Altimetrik

  • Altimetrik Mountain View, CA
  • GenAI GenOS Engineer (Senior) About the Team GenOS is a Generative AI Operating System — the platform every GenAI experience is built, deployed, and govern... more
  • 2 Days Ago

  • Altimetrik Mountain View, CA
  • AI-First DevOps Engineer (Senior) About the Team The Development Efficiencies pillar owns the paved roads, PDLC velocity, and self-serve tooling behind an ... more
  • 2 Days Ago

  • Altimetrik Pittsburgh, PA
  • Role: Oracle HR Help Desk Implementation Specialist Position Summary: We are seeking an experienced Oracle HR Help Desk Implementation Specialist to lead a... more
  • 10 Days Ago


Not the job you're looking for? Here are some other Senior SRE (AI) jobs in the Mountain View, CA area that may be a better fit.

  • Apple Cupertino, CA
  • As a Sr. Site Reliability Engineer at Apple, you will be responsible for driving the reliability, scalability, and observability of our cloud platform. You... more
  • 27 Days Ago

  • AMD San Jose, CA
  • WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data... more
  • 15 Days Ago

AI Assistant is available now!

Feel free to start your new journey!