Demo

Site Reliability Engineer

Evlo AI
Atlanta, GA Full Time
POSTED ON 7/26/2026
AVAILABLE BEFORE 8/24/2026
About The Role

The role focuses on building, scaling, and maintaining the core infrastructure and platform services that power highly available production environments. The engineer will collaborate closely with software delivery teams to ensure optimal reliability, observability, and performance across distributed, multi-cloud architectures.

This position requires a deep understanding of container orchestration, infrastructure as code, and continuous deployment methodologies. The engineer will actively work to eliminate manual operations through automation, drive architecture reviews, and systematically improve system resilience.

Key Responsibilities

  • Design, deploy, and manage multi-region Kubernetes clusters on public cloud infrastructure (AWS or GCP)
  • Build and optimize infrastructure as code (IaC) templates using Terraform to maintain repeatable, secure, and compliant cloud environments
  • Develop and maintain comprehensive observability pipelines using Prometheus, Grafana, and OpenTelemetry to proactively detect and alert on system anomalies
  • Architect and secure robust CI/CD pipelines utilizing GitLab CI, GitHub Actions, or ArgoCD for automated, zero-downtime deployments
  • Participate in a blameless on-call rotation, leading incident mitigation, conducting root-cause analyses, and implementing long-term engineering fixes to prevent regressions
  • Collaborate with backend engineering teams on capacity planning, database tuning, and system architecture to ensure seamless horizontal scaling

What We Are Looking For

  • 3–7 years of experience in site reliability engineering, DevOps, or systems engineering managing production-grade, high-traffic cloud environments
  • Strong hands-on expertise with containerized environments, specifically Docker and production-scale Kubernetes cluster management
  • Demonstrated experience writing production-grade Infrastructure as Code using Terraform or Pulumi
  • Proficiency in at least one programming or scripting language, such as Python, Go, or Bash, for automating operational tasks
  • Solid understanding of Linux internals, networking concepts (DNS, TCP/IP, Load Balancing), and secure cloud architecture patterns
  • Bachelor's degree in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience
  • Bonus: Experience with GitOps deployment workflows (ArgoCD, Flux), service mesh architectures (Istio, Linkerd), or managing large-scale distributed databases

Salary.com Estimation for Site Reliability Engineer in Atlanta, GA
$86,716 to $108,052
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Site Reliability Engineer?

Sign up to receive alerts about other jobs on the Site Reliability Engineer career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$92,877 - $110,401
Income Estimation: 
$120,933 - $155,034
Income Estimation: 
$114,618 - $136,401
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Evlo AI

  • Evlo AI Washington, DC
  • About The Role The role is responsible for the availability, latency, performance, efficiency, and capacity management of a high-throughput cloud-native pl... more
  • 1 Day Ago

  • Evlo AI Boston, MA
  • About The Role The Customer Success Manager owns the post-sale customer lifecycle, driving adoption, retention, and growth across a portfolio of enterprise... more
  • 1 Day Ago

  • Evlo AI Atlanta, GA
  • About The Role The role drives the vision, strategy, and execution for core platform APIs and developer-facing infrastructure, translating complex technica... more
  • 1 Day Ago

  • Evlo AI Raleigh, NC
  • About The Role The role serves as the central operational hub for two high-performing Product and Engineering executives, managing complex schedules, cross... more
  • 1 Day Ago


Not the job you're looking for? Here are some other Site Reliability Engineer jobs in the Atlanta, GA area that may be a better fit.

  • qgenda Atlanta, GA
  • Who We Are QGenda is redefining healthcare workforce management everywhere care is delivered. We're on a mission to empower the healthcare industry to bett... more
  • 1 Day Ago

  • Morgan Stanley Alpharetta, GA
  • In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to re... more
  • 2 Days Ago

AI Assistant is available now!

Feel free to start your new journey!