What are the responsibilities and job description for the Senior SRE (AI) position at Altimetrik?
AI-First SRE (Senior)
Key Responsibilities
- Build and extend O11Y pipelines: metrics, distributed tracing, logging, SLO/SLI definitions and dashboards consumed platform-wide.
- Instrument AI systems in production: token usage and cost, latency-per-inference, tool-call success rates, response-quality signals, agent loop detection, runaway-cost alerting.
- Extend tracing across agentic flows — planner → executor → retrieval → tool calls — spanning GenOS, AI Gateway, and MCP Gateway.
- Define SLOs where "available" includes response quality and tool-call success, not just HTTP 200s.
- On-call and incident command for AI platform surfaces: model/provider degradation, Bedrock/Gemini failover, semantic-cache issues, prompt-injection events — blameless postmortems with tracked remediations.
- Own the SRE side of the progressive-delivery seam: canary analysis, automated rollback decisioning, chaos/resilience testing, blast-radius controls — DevOps builds the pipeline; you define the gates.
- Build AIOps agents in the Traffic Agent pattern: intelligent triage, anomaly detection, auto-remediation with hard guardrails.
- Drive UX Availability (R30) and Foundations Capabilities (R30); land network, O11Y, and cloud cost savings (tracked YTD KPIs).
Must-Have Qualifications
- 7 years SRE/production engineering on Tier-0/Tier-1, high-traffic systems.
- Strong Go or Python — this is a build role: tooling, automation, instrumentation.
- Has built observability stacks, not just consumed them: Prometheus/Grafana, OpenTelemetry, or equivalent at scale, including cardinality and cost control.
- Production LLM/ML monitoring: Langfuse, Arize, WhyLabs, or homegrown — token/cost tracking, drift and quality metrics.
- Working fluency in AI-system failure modes: nondeterminism, provider limits and outages, context-window overflow, agent loops, cache poisoning.
- Kubernetes AWS operational depth — debugs across cluster, mesh, and gateway layers.
- Structured incident-command and postmortem experience.
Nice-to-Have
- AIOps / LLM-applied-to-ops: auto-triage, incident summarization, remediation agents.
- Chaos engineering (Litmus, Gremlin, or homegrown); eBPF or deep network debugging.
- FinOps / cost engineering; fintech or regulated-industry reliability experience.