What are the responsibilities and job description for the Senior System Reliability Engineer position at On-Demand Group?
Senior Site Reliability Engineer (SRE)
Job Type: Contract
Pay Range: $80–$100/hour
Position Overview
We are seeking an experienced Senior Site Reliability Engineer (SRE) to support highly available, cloud-native applications within a Digital Commerce environment.
This role will focus heavily on monitoring, observability, reliability, and performance of Kubernetes-based applications and services running within Microsoft Azure and Azure Kubernetes Service (AKS).
The Senior SRE will partner closely with DevOps, Platform Engineering, Production Support, Infrastructure, Network, Security, Architecture, and Development teams to improve the reliability, scalability, performance, and security of distributed applications.
The ideal candidate brings strong hands-on experience with Azure, AKS/Kubernetes, observability and application performance monitoring, along with the troubleshooting skills necessary to identify and resolve complex production issues.
Key Responsibilities
- Monitor and improve the reliability, availability, latency, and performance of Kubernetes-based applications and services running on Azure Kubernetes Service (AKS).
- Plan, design, deploy, and operate Site Reliability Engineering capabilities for cloud-based products and services.
- Design and maintain monitoring and observability frameworks using tools such as Dynatrace, Azure Monitor, and Application Insights.
- Develop monitoring and alerting that proactively identifies symptoms and performance degradation rather than simply reporting outages.
- Analyze telemetry, logs, metrics, and monitoring data to identify application and infrastructure bottlenecks.
- Recognize and address substandard application or infrastructure performance based on established KPIs.
- Troubleshoot complex issues across distributed, cloud-native systems.
- Continuously automate and improve capabilities to increase reliability, scalability, performance, and security.
- Partner with DevOps and Platform Engineering teams to integrate monitoring and observability capabilities into applications, infrastructure, and automated pipelines.
- Work closely with Infrastructure, Network, Security, Architecture, and Development teams to build and maintain highly available Azure environments.
- Support incident response and help identify root causes of production reliability and performance issues.
- Design and configure proactive alerting mechanisms and thresholds to enable rapid identification and resolution of issues.
- Contribute to the design and implementation of CI/CD pipelines and automated operational processes.
- Document processes, technical designs, operational procedures, and monitoring standards.
- Identify cross-team operational risks and drive issues toward resolution through engineering, troubleshooting, and operational improvements.
- Participate in regulatory and compliance activities as needed.
Required Qualifications
- 5 years of experience in Software Engineering, Systems Engineering, Operations Engineering, DevOps, SRE, or related technical roles.
- 2 years of hands-on Site Reliability Engineering, DevOps, or similar cloud-native engineering experience.
- Strong experience supporting cloud-native applications hosted within Microsoft Azure.
- Hands-on experience supporting and monitoring Kubernetes / Azure Kubernetes Service (AKS) environments.
- Experience monitoring application availability, uptime, latency, infrastructure, and performance across large distributed systems.
- Strong knowledge of observability and application performance monitoring concepts.
- Experience troubleshooting complex cloud, application, infrastructure, and system-related issues.
- Strong debugging and problem-solving skills within distributed environments.
- Experience designing or implementing CI/CD pipelines.
- Experience with version control systems such as Git.
- Working knowledge across systems, networking, security, databases, storage, and cloud infrastructure.
- Experience collaborating across DevOps, Platform Engineering, Production Support, Infrastructure, Architecture, and Development teams.
- Strong written and verbal communication skills with the ability to communicate technical monitoring and reliability insights to both technical and non-technical stakeholders.
Preferred Qualifications
- Deep experience monitoring Kubernetes / AKS environments, containerized applications, services, and workloads.
- Hands-on experience with Dynatrace.
- Experience with Azure Monitor and Application Insights.
- Experience designing and implementing enterprise monitoring and observability frameworks.
- Experience configuring proactive, symptom-based alerting and thresholds.
- Experience analyzing telemetry and monitoring data to identify performance bottlenecks.
- Strong understanding of application performance monitoring within distributed and microservices-based environments.
- Experience improving reliability through automation and SRE practices.
- Experience supporting high-volume, customer-facing web or digital commerce applications.
- Experience participating in incident response, root cause analysis, and continuous operational improvement.
What We're Looking For
The strongest candidate will bring a combination of Site Reliability Engineering, Azure, Kubernetes/AKS, and observability expertise.
This is not simply a traditional DevOps or cloud infrastructure role. We are looking for someone who understands how to use monitoring and telemetry to determine what is happening across complex distributed applications, proactively identify performance or reliability issues, and work across engineering teams to resolve the underlying problems.
Candidates with hands-on experience using Dynatrace, Azure Monitor, Application Insights, and Kubernetes/AKS observability will be particularly relevant.
The projected hourly range for this position is $80–$100.
On-Demand Group (ODG) provides employee benefits which includes healthcare, dental, and vision insurance. ODG is an equal opportunity employer that does not discriminate on the basis of race, color, religion, gender, sexual orientation, age, national origin, disability, or any other characteristic protected by law.
Salary : $80 - $100