What are the responsibilities and job description for the Site Reliability Engineer position at Evlo AI?
About The Role
The role focuses on maintaining high availability, low latency, and robust security across large-scale distributed systems supporting millions of active users.
The team works closely with software engineers to automate infrastructure provisioning, enhance observability, and drive incident response protocols.
Key Responsibilities
The role focuses on maintaining high availability, low latency, and robust security across large-scale distributed systems supporting millions of active users.
The team works closely with software engineers to automate infrastructure provisioning, enhance observability, and drive incident response protocols.
Key Responsibilities
- Design and manage cloud infrastructure on AWS using Terraform and Kubernetes, ensuring high availability and zero-downtime deployments
- Build and maintain comprehensive monitoring, logging, and alerting systems using Prometheus, Grafana, and Datadog
- Automate deployment pipelines and release processes with CI/CD tools such as GitHub Actions and ArgoCD
- Lead incident response, conduct thorough post-mortem analyses, and implement preventive measures to eliminate recurring failure modes
- Optimize infrastructure costs, resource utilization, and database performance across multiple production environments
- Participate in an on-call rotation and collaborate with development teams to ensure production readiness for new services
- 3–6 years of experience in Site Reliability Engineering, DevOps, or systems administration within cloud-native environments
- Strong proficiency in Linux systems administration, networking fundamentals (TCP/IP, DNS, TLS), and containerization technologies (Docker, Kubernetes)
- Hands-on experience with Infrastructure as Code (Terraform, CloudFormation) and configuration management tools
- Proficiency in at least one scripting or programming language: Python, Go, or Bash
- Solid understanding of distributed systems architecture, microservices patterns, and resilience engineering principles
- Bonus: Bachelor's degree in Computer Science, certified Kubernetes administrator (CKA), or experience with service mesh technologies like Istio