Demo

Senior System Reliability Engineer

On-Demand Group
Eagan, MN Contractor
POSTED ON 9/27/2026
AVAILABLE BEFORE 10/25/2026

Senior Site Reliability Engineer (SRE)

Job Type: Contract

Pay Range: $80–$100/hour


Position Overview

We are seeking an experienced Senior Site Reliability Engineer (SRE) to support highly available, cloud-native applications within a Digital Commerce environment.

This role will focus heavily on monitoring, observability, reliability, and performance of Kubernetes-based applications and services running within Microsoft Azure and Azure Kubernetes Service (AKS).


The Senior SRE will partner closely with DevOps, Platform Engineering, Production Support, Infrastructure, Network, Security, Architecture, and Development teams to improve the reliability, scalability, performance, and security of distributed applications.

The ideal candidate brings strong hands-on experience with Azure, AKS/Kubernetes, observability and application performance monitoring, along with the troubleshooting skills necessary to identify and resolve complex production issues.


Key Responsibilities

  • Monitor and improve the reliability, availability, latency, and performance of Kubernetes-based applications and services running on Azure Kubernetes Service (AKS).
  • Plan, design, deploy, and operate Site Reliability Engineering capabilities for cloud-based products and services.
  • Design and maintain monitoring and observability frameworks using tools such as Dynatrace, Azure Monitor, and Application Insights.
  • Develop monitoring and alerting that proactively identifies symptoms and performance degradation rather than simply reporting outages.
  • Analyze telemetry, logs, metrics, and monitoring data to identify application and infrastructure bottlenecks.
  • Recognize and address substandard application or infrastructure performance based on established KPIs.
  • Troubleshoot complex issues across distributed, cloud-native systems.
  • Continuously automate and improve capabilities to increase reliability, scalability, performance, and security.
  • Partner with DevOps and Platform Engineering teams to integrate monitoring and observability capabilities into applications, infrastructure, and automated pipelines.
  • Work closely with Infrastructure, Network, Security, Architecture, and Development teams to build and maintain highly available Azure environments.
  • Support incident response and help identify root causes of production reliability and performance issues.
  • Design and configure proactive alerting mechanisms and thresholds to enable rapid identification and resolution of issues.
  • Contribute to the design and implementation of CI/CD pipelines and automated operational processes.
  • Document processes, technical designs, operational procedures, and monitoring standards.
  • Identify cross-team operational risks and drive issues toward resolution through engineering, troubleshooting, and operational improvements.
  • Participate in regulatory and compliance activities as needed.


Required Qualifications

  • 5 years of experience in Software Engineering, Systems Engineering, Operations Engineering, DevOps, SRE, or related technical roles.
  • 2 years of hands-on Site Reliability Engineering, DevOps, or similar cloud-native engineering experience.
  • Strong experience supporting cloud-native applications hosted within Microsoft Azure.
  • Hands-on experience supporting and monitoring Kubernetes / Azure Kubernetes Service (AKS) environments.
  • Experience monitoring application availability, uptime, latency, infrastructure, and performance across large distributed systems.
  • Strong knowledge of observability and application performance monitoring concepts.
  • Experience troubleshooting complex cloud, application, infrastructure, and system-related issues.
  • Strong debugging and problem-solving skills within distributed environments.
  • Experience designing or implementing CI/CD pipelines.
  • Experience with version control systems such as Git.
  • Working knowledge across systems, networking, security, databases, storage, and cloud infrastructure.
  • Experience collaborating across DevOps, Platform Engineering, Production Support, Infrastructure, Architecture, and Development teams.
  • Strong written and verbal communication skills with the ability to communicate technical monitoring and reliability insights to both technical and non-technical stakeholders.


Preferred Qualifications

  • Deep experience monitoring Kubernetes / AKS environments, containerized applications, services, and workloads.
  • Hands-on experience with Dynatrace.
  • Experience with Azure Monitor and Application Insights.
  • Experience designing and implementing enterprise monitoring and observability frameworks.
  • Experience configuring proactive, symptom-based alerting and thresholds.
  • Experience analyzing telemetry and monitoring data to identify performance bottlenecks.
  • Strong understanding of application performance monitoring within distributed and microservices-based environments.
  • Experience improving reliability through automation and SRE practices.
  • Experience supporting high-volume, customer-facing web or digital commerce applications.
  • Experience participating in incident response, root cause analysis, and continuous operational improvement.


What We're Looking For

The strongest candidate will bring a combination of Site Reliability Engineering, Azure, Kubernetes/AKS, and observability expertise.


This is not simply a traditional DevOps or cloud infrastructure role. We are looking for someone who understands how to use monitoring and telemetry to determine what is happening across complex distributed applications, proactively identify performance or reliability issues, and work across engineering teams to resolve the underlying problems.

Candidates with hands-on experience using Dynatrace, Azure Monitor, Application Insights, and Kubernetes/AKS observability will be particularly relevant.


The projected hourly range for this position is $80–$100.

On-Demand Group (ODG) provides employee benefits which includes healthcare, dental, and vision insurance. ODG is an equal opportunity employer that does not discriminate on the basis of race, color, religion, gender, sexual orientation, age, national origin, disability, or any other characteristic protected by law.

Salary : $80 - $100

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Senior System Reliability Engineer?

Sign up to receive alerts about other jobs on the Senior System Reliability Engineer career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$114,618 - $136,401
Income Estimation: 
$144,264 - $191,312
Income Estimation: 
$140,435 - $166,410
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at On-Demand Group

  • On-Demand Group Litchfield, MN
  • Job Title: Senior Manager of IT & Business Systems Job Type: Permanent Job Location: 100% Onsite Overview The Senior Manager of Information Technology & In... more
  • 2 Days Ago


Not the job you're looking for? Here are some other Senior System Reliability Engineer jobs in the Eagan, MN area that may be a better fit.

  • Allied Reliability Washington, WA
  • ROCKWOOL is seeking a Sr. Electrical Engineer to join our Group Technology team! This position requires travel of 60%, both internationally and nationally,... more
  • 28 Days Ago

  • System One Columbia, SC
  • Job Title: Senior AWS Site Reliability Engineer (SRE) Location: Birmingham, Alabama Type: Contract To Hire Work Model: Onsite – onsite Hours: 40.0 Security... more
  • 8 Days Ago

AI Assistant is available now!

Feel free to start your new journey!