What are the responsibilities and job description for the Site Reliability Engineer position at Evlo AI?
About The Role
The role is responsible for the availability, latency, performance, efficiency, and capacity management of a high-throughput global cloud infrastructure. This SRE position acts as a bridge between software development and systems engineering, designing automation-first platforms to scale systems sustainably.
You will collaborate closely with product development teams to improve service reliability, build robust observability pipelines, and drive the adoption of infrastructure-as-code best practices across the entire engineering organization.
Key Responsibilities
The role is responsible for the availability, latency, performance, efficiency, and capacity management of a high-throughput global cloud infrastructure. This SRE position acts as a bridge between software development and systems engineering, designing automation-first platforms to scale systems sustainably.
You will collaborate closely with product development teams to improve service reliability, build robust observability pipelines, and drive the adoption of infrastructure-as-code best practices across the entire engineering organization.
Key Responsibilities
- Design, provision, and maintain multi-region Kubernetes clusters on AWS using Terraform to ensure high availability and disaster recovery readiness
- Build and optimize continuous integration and continuous deployment (CI/CD) pipelines using GitHub Actions or GitLab CI to automate software delivery safely
- Establish comprehensive observability across microservices by configuring Prometheus, Grafana, Datadog, and OpenTelemetry for proactive alerting and performance monitoring
- Lead incident response operations, conduct blameless post-mortems, and build automated remediation systems to reduce Mean Time to Resolution (MTTR)
- Configure and manage scalable data stores and caching layers, including PostgreSQL, Redis, and Kafka, focusing on performance tuning and replication policies
- Develop internal CLI tools and automation scripts in Go or Python to eliminate manual operational toil and streamline developer workflows
- 3 to 6 years of experience in SRE, DevOps, or systems engineering roles managing production environments at scale
- Strong proficiency in writing infrastructure as code (IaC) with Terraform and managing containerized workloads using Kubernetes
- Solid scripting and programming experience in Go, Python, or Bash for automation and tool development
- Deep understanding of networking protocols, security standards (TLS, IAM, VPC), and Linux systems internals
- Bachelor's degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience
- Bonus: Experience with service meshes (Istio, Linkerd), GitOps paradigms (ArgoCD), and cloud-native databases at scale