Demo

Site Reliability Engineer

Gridiron IT
Arlington, VA Contractor
POSTED ON 7/25/2026
AVAILABLE BEFORE 8/23/2026

DevOps SRE

Location: SCIF ~75%

Remote work: Yes

Clearance: TS/SCI


Role Overview

DevOps Site Reliability Engineer (SRE) essential to ensuring the continuous availability, performance, and health of an enterprise platform. This role focuses on maintaining reliable and stable software deployments in multi-cloud environments, establishing comprehensive monitoring and alerting systems, and providing rapid-reaction troubleshooting. You will bridge the gap between development and operations to uphold the mandated 99.9% uptime requirement for a critical defense enterprise platform.

Desired Qualifications

- Strong background in Site Reliability Engineering principles and practices.

- Hands-on experience with Kubernetes deployments and orchestration.

- Proven experience managing and troubleshooting multi-cloud environments (GCP, Azure, AWS).

- Expertise in setting up and managing monitoring, logging, and automated alerting systems.

- Proficiency in scripting and automation (e.g., Python, Bash) for system stabilization and diagnostic tasks.

What we are looking for in a candidate

- A strong interest in maintaining highly reliable software deployments under stringent uptime requirements.

- Ability to remain calm and systematic during high-pressure downtime incidents.

- Excellent diagnostic and troubleshooting skills with an eye for rapid resolution.

- A "build-to-manage" mindset, focusing on automating repetitive tasks to reduce operational toil.

Key Responsibilities

- Continuous Monitoring: Provide 24/7 infrastructure health monitoring and automated telemetry tracking to maintain the mandated 99.9% core platform availability.

- Incident Response: Serve on-call to respond to major incidents or platform downtime within 1 hour of notification, executing rapid platform downtime response and system stabilization maneuvers.

- Automation: Develop and maintain a library of scripts and internal tools to streamline and automate repetitive diagnostic tasks, health checks, and data-gathering procedures.

- Dashboarding & Telemetry: Create and maintain automated dashboards for uptime, incident status, API latency, and other critical support metrics to ensure platform health visibility.

- Vendor Coordination: Lead direct engineering-level coordination with Cloud Service Providers (CSPs) during outages or underlying infrastructure issues.

- Reliability Engineering: Continuously evaluate, monitor, and provide recommended improvements to logging, system metrics, and architecture to ensure the platform remains rapidly scalable and highly available.


Clearance

Applicants selected will be subject to a security investigation and may need to meet eligibility requirements for access to classified information. TS/SCI required.


Compensation and Benefits

Salary Range $120,000 - $210,000/YR (Compensation is determined by various factors, including but not limited to location, work experience, skills, education, certifications, seniority, and business needs. This range may be modified in the future.)


Benefits: Gridiron offers a comprehensive benefits package including medical, dental, vision insurance, HSA, FSA, 401(k), disability & ADD insurance, life and pet insurance to eligible employees. Full-time and part-time employees working at least 30 hours per week on a regular basis are eligible to participate in Gridiron’s benefits programs.


Gridiron IT Solutions is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status or disability status.


Gridiron IT is a Women Owned Small Business (WOSB) headquartered in the Washington, D.C. area that supports our clients' missions throughout the United States. Gridiron IT specializes in providing comprehensive IT services tailored to meet the needs of federal agencies. Our capabilities include IT Infrastructure & Cloud Services, Cyber Security, Software Integration & Development, Data Solution & AI, and Enterprise Applications. These capabilities are backed by Gridiron IT's experienced workforce and our commitment to ensuring we meet and exceed our clients' expectations.


Salary : $120,000 - $210,000

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Site Reliability Engineer?

Sign up to receive alerts about other jobs on the Site Reliability Engineer career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$92,877 - $110,401
Income Estimation: 
$120,933 - $155,034
Income Estimation: 
$114,618 - $136,401
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Gridiron IT

  • Gridiron IT Lexington, MA
  • Location: Lexington, MA or Rome, NY Work Type: Full-Time / Hybrid Remote Work: 50% - Candidate must be onsite 3x a week at the designated AFB (Lexington or... more
  • 11 Days Ago


Not the job you're looking for? Here are some other Site Reliability Engineer jobs in the Arlington, VA area that may be a better fit.

  • Tiger Analytics Inc. Washington, DC
  • Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our productio... more
  • 1 Day Ago

  • SpaceX Washington, DC
  • SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today... more
  • 1 Day Ago

AI Assistant is available now!

Feel free to start your new journey!