What are the responsibilities and job description for the Site Reliability Engineer position at Evlo AI?
About The Role
The Site Reliability Engineer owns the reliability, availability, and operational maturity of production services running across AWS. The role spans Kubernetes, Terraform, CI/CD, observability, incident response, and the automation required to operate distributed systems at scale.
The team is building dependable platform capabilities for engineering teams that ship frequently and serve demanding workloads. This role matters because it turns operational risk into measurable engineering improvements through resilient architecture, clear service-level objectives, and disciplined automation.
Key Responsibilities
The Site Reliability Engineer owns the reliability, availability, and operational maturity of production services running across AWS. The role spans Kubernetes, Terraform, CI/CD, observability, incident response, and the automation required to operate distributed systems at scale.
The team is building dependable platform capabilities for engineering teams that ship frequently and serve demanding workloads. This role matters because it turns operational risk into measurable engineering improvements through resilient architecture, clear service-level objectives, and disciplined automation.
Key Responsibilities
- Design and operate highly available AWS infrastructure using Kubernetes, Terraform, Helm, and infrastructure-as-code best practices
- Define and enforce service-level objectives, error budgets, and reliability standards across production services
- Build and maintain observability systems using Prometheus, Grafana, OpenTelemetry, and centralized logging platforms
- Automate deployment, scaling, backup, recovery, and routine operational workflows through Python, Go, or Bash
- Lead incident response, coordinate remediation during high-severity events, and produce blameless post-incident reviews with actionable follow-up work
- Improve CI/CD pipelines with deployment safety controls, automated testing, progressive delivery, and reliable rollback procedures
- Partner with software engineers to identify performance bottlenecks, eliminate recurring operational toil, and strengthen failure-mode testing
- 3–8 years of experience in site reliability engineering, DevOps, platform engineering, or production infrastructure roles
- Hands-on experience operating Kubernetes workloads in AWS, including networking, IAM, autoscaling, storage, and production troubleshooting
- Strong proficiency with Terraform or an equivalent infrastructure-as-code framework, plus practical experience managing cloud resources through version control
- Experience building observability and alerting systems with tools such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent platforms
- Proficiency in Python, Go, or Bash for operational automation, tooling, and service integration; familiarity with Linux internals and networking fundamentals
- Bachelor’s degree in computer science, engineering, or a related technical field, or equivalent professional experience
- Bonus: Experience with AWS certifications, service mesh technologies, GitOps, chaos engineering, distributed databases, or compliance requirements for production infrastructure