What are the responsibilities and job description for the DevOps Engineer position at Evlo AI?
About The Role
The DevOps Engineer builds and operates the infrastructure that runs production services at scale, with a focus on automation, availability, deployment safety, and measurable system performance. The role spans cloud infrastructure, Kubernetes, CI/CD, observability, and incident response across development and production environments.
You will partner with software engineers and SREs to improve delivery velocity without compromising reliability. The team needs an engineer who can turn operational requirements into reusable platforms, eliminate manual work through infrastructure as code, and lead practical improvements to production resilience.
Key Responsibilities
The DevOps Engineer builds and operates the infrastructure that runs production services at scale, with a focus on automation, availability, deployment safety, and measurable system performance. The role spans cloud infrastructure, Kubernetes, CI/CD, observability, and incident response across development and production environments.
You will partner with software engineers and SREs to improve delivery velocity without compromising reliability. The team needs an engineer who can turn operational requirements into reusable platforms, eliminate manual work through infrastructure as code, and lead practical improvements to production resilience.
Key Responsibilities
- Design, provision, and maintain AWS infrastructure using Terraform, including VPCs, IAM, EKS, RDS, S3, and related production services
- Build and improve CI/CD pipelines with GitHub Actions, GitLab CI, or equivalent tools for automated testing, security checks, deployments, and rollback
- Operate Kubernetes workloads across development and production environments, including deployments, Helm charts, ingress, autoscaling, secrets, and resource management
- Implement observability with Prometheus, Grafana, OpenTelemetry, and centralized logging to track service health, latency, capacity, and error rates
- Automate operational workflows with Python, Go, or Bash, reducing toil and improving consistency across infrastructure and application teams
- Participate in incident response, troubleshoot complex production failures, and produce clear post-incident actions that improve system reliability
- Define and enforce infrastructure, security, and operational standards through code reviews, documentation, runbooks, and platform tooling
- 3–8 years of experience in DevOps, site reliability engineering, platform engineering, or production infrastructure
- Hands-on experience operating workloads in AWS or another major cloud provider, with a strong understanding of networking, IAM, compute, storage, and managed databases
- Production experience with Kubernetes and container technologies, including Docker, Helm, service discovery, ingress, and autoscaling
- Proficiency with Terraform or another infrastructure-as-code tool and experience managing infrastructure through version-controlled workflows
- Practical experience building CI/CD pipelines and implementing automated testing, deployment strategies, secrets management, and rollback procedures
- Strong troubleshooting and communication skills, with experience participating in on-call rotations and resolving production incidents
- Bachelor’s degree in computer science, engineering, information technology, or a related field, or equivalent professional experience; Bonus: experience with Go, Argo CD, Istio, PostgreSQL, security automation, SLOs, and cost optimization