What are the responsibilities and job description for the Site Reliability Engineer position at Evlo AI?
About The Role
The role focuses on building, scaling, and maintaining the core infrastructure and platform services that power highly available production environments. The engineer will collaborate closely with software delivery teams to ensure optimal reliability, observability, and performance across distributed, multi-cloud architectures.
This position requires a deep understanding of container orchestration, infrastructure as code, and continuous deployment methodologies. The engineer will actively work to eliminate manual operations through automation, drive architecture reviews, and systematically improve system resilience.
Key Responsibilities
The role focuses on building, scaling, and maintaining the core infrastructure and platform services that power highly available production environments. The engineer will collaborate closely with software delivery teams to ensure optimal reliability, observability, and performance across distributed, multi-cloud architectures.
This position requires a deep understanding of container orchestration, infrastructure as code, and continuous deployment methodologies. The engineer will actively work to eliminate manual operations through automation, drive architecture reviews, and systematically improve system resilience.
Key Responsibilities
- Design, deploy, and manage multi-region Kubernetes clusters on public cloud infrastructure (AWS or GCP)
- Build and optimize infrastructure as code (IaC) templates using Terraform to maintain repeatable, secure, and compliant cloud environments
- Develop and maintain comprehensive observability pipelines using Prometheus, Grafana, and OpenTelemetry to proactively detect and alert on system anomalies
- Architect and secure robust CI/CD pipelines utilizing GitLab CI, GitHub Actions, or ArgoCD for automated, zero-downtime deployments
- Participate in a blameless on-call rotation, leading incident mitigation, conducting root-cause analyses, and implementing long-term engineering fixes to prevent regressions
- Collaborate with backend engineering teams on capacity planning, database tuning, and system architecture to ensure seamless horizontal scaling
- 3–7 years of experience in site reliability engineering, DevOps, or systems engineering managing production-grade, high-traffic cloud environments
- Strong hands-on expertise with containerized environments, specifically Docker and production-scale Kubernetes cluster management
- Demonstrated experience writing production-grade Infrastructure as Code using Terraform or Pulumi
- Proficiency in at least one programming or scripting language, such as Python, Go, or Bash, for automating operational tasks
- Solid understanding of Linux internals, networking concepts (DNS, TCP/IP, Load Balancing), and secure cloud architecture patterns
- Bachelor's degree in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience
- Bonus: Experience with GitOps deployment workflows (ArgoCD, Flux), service mesh architectures (Istio, Linkerd), or managing large-scale distributed databases