Demo

Cloud Site Reliability Engineer (SRE)

Everforth ECS
Arlington, VA Full Time
POSTED ON 9/19/2026
AVAILABLE BEFORE 1/16/2027
Everforth ECS is seeking a Cloud Site Reliability Engineer (SRE) to work in our Arlington, VA office/remotely.
Our Philosophy
We believe the job of an SRE is to engineer the cloud to run itself. That means writing software and automation that lets systems detect and recover from failure on their own, rather than relying on someone to notice an alert and manually fix it. When something breaks, self-healing comes first, deep root-cause debugging happens after service is restored, not instead of it. We're looking for someone who automates the operational task by default, not documents the runbook for doing it by hand.

About the Role
This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You'll define what “reliable enough” looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation.

Responsibilities
Self-Healing Operations
  • Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention
  • Shift the team's posture from “is it running, how do we fix it” to “how do we make it fix itself”
  • Automate service restoration first; investigate root cause after
Uptime Goals & Reliability
  • Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them
  • Use live metrics to decide what's “reliable enough” and where to invest next
Infrastructure
  • Enforce infrastructure-as-code and configuration-as-code, no manual tech change
  • Own Terraform standards and reusable modules adopted across programs
  • Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling)
  • Set CI/CD and pipeline-as-code standards, including progressive delivery
Observability & Incidents
  • Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct
  • Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation
Collaboration & Leadership
  • Partner with development and contractor teams leads to embed reliability and automation across the software
  • Mentor engineers toward this same automation-first philosophy
  • Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation

Salary Range: $130,000 - $180,000

General Description of Benefit

Requirements:
  • Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent practical experience)
  • 5 years of SRE experience (or equivalent), with demonstrated technical leadership
  • 10 years of general work experience
  • Track record building self-healing/auto-remediating systems, not just dashboards
  • Jenkins experience
  • Expert AWS knowledge, GovCloud experience strongly preferred
  • Deep Kubernetes and Terraform expertise at production scale
  • Strong software engineering background (Python and/or Go)
  • Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc.)
  • Proven incident command and postmortem experience
  • Strong communication skills across technical and federal leadership audiences
  • Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable)
  • US Citizenship

Req Benefits:
Benefits - Everforth ECS

Benefits:

Health Insurance, Vacation & Paid Time Off, 401K Plan

Salary : $130,000 - $180,000

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Cloud Site Reliability Engineer (SRE)?

Sign up to receive alerts about other jobs on the Cloud Site Reliability Engineer (SRE) career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$114,618 - $136,401
Income Estimation: 
$144,264 - $191,312
Income Estimation: 
$140,435 - $166,410
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Everforth ECS

  • Everforth ECS Arlington, VA
  • Everforth ECS is seeking an AI Solutions Engineer to work in our Arlington, VA office/remotely . Our Philosophy We believe AI only creates value when it's ... more
  • 1 Day Ago

  • Everforth ECS Arnold, MO
  • ECS is searching for a GEOINT Analyst - Exploitation Specialist in support of the US Government in St. Louis, MO . We are seeking an experienced analyst to... more
  • 2 Days Ago

  • Everforth ECS Dahlgren, VA
  • Everforth ECS seeks an Intermediate Financial Analyst to work Hybrid remote/onsite, with a minimum of 3 business days onsite at the Dahlgren, VA customer s... more
  • 2 Days Ago

  • Everforth ECS Dahlgren, VA
  • Everforth ECS seeks a Basic Cost Analyst to support the Missile Defense Agency in Dahlgren, VA . The successful candidate will: Collaborate with team membe... more
  • 2 Days Ago


Not the job you're looking for? Here are some other Cloud Site Reliability Engineer (SRE) jobs in the Arlington, VA area that may be a better fit.

  • Tines Washington, DC
  • Founded in 2018 with co-headquarters in Dublin and Boston, Tines powers some of the world's most important workflows. Our intelligent workflow platform app... more
  • 1 Month Ago

  • C-Serv Arlington, VA
  • Site Reliability Engineer II DC Metro area, offices in Reston · Hybrid · 24/7 FedRAMP Operations · Rotational Shift · Initial Contract till March 27. KEY R... more
  • 1 Day Ago

AI Assistant is available now!

Feel free to start your new journey!