What are the responsibilities and job description for the SRE (Site Reliability Engineering) Lead position at Pacer Group?
Job Title: SRE (Site Reliability Engineering) Lead
Location: Woonsocket, RI
Work Arrangement: Hybrid
Employment Type: Contract
Duration: 12 months
Domain: Healthcare / Retail Tech
Pay Rate: $41.00/hr W2 | $50.00/hr C2C
Application Deadline: September 15, 2026
SKILLS REQUIRED
Primary (Must-Have):
• 8 years of Senior Software Engineering experience in SRE, DevOps, or Platform Engineering for distributed systems
• On-call Incident Commander (IC) experience leading P1/P2 incidents and tuning production time-series anomaly detection models
• Strong production programming proficiency in Python, React, and Java for operational tooling
• Hands-on design of SLIs, SLOs, and error budgets, plus deep observability experience (Prometheus, Grafana, OpenTelemetry, Loki/Splunk/Elasticsearch)
• Cloud platform expertise in GCP, Rancher K3s, advanced Kubernetes operations, and data pipeline observability (Airflow, Tidal)
Secondary (Good to Have):
• Experience owning Production Readiness Reviews, fault injection/chaos engineering, or holding TIC (Technical Incident Commander) certification
• Telemetry data analytics using SQL, Google BigQuery, and PostgreSQL
• Hands-on experience with streaming platforms (Kafka), service mesh (Istio, Envoy), and IaC (Terraform, Ansible)
POSITION OVERVIEW
Operating within the Platform Reliability and Infrastructure organization, this role leads Site Reliability Engineering initiatives for large-scale distributed production systems. Reporting to the Senior Manager of Infrastructure & Reliability, the SRE Lead will solve complex observability, incident response, and system availability challenges to ensure high availability, reduced blast radius, and continuous performance across business-critical customer and patient-facing applications.
ROLES & RESPONSIBILITIES
• Serve as an on-call Incident Commander (IC) for P1/P2 incidents, providing structured leadership updates and driving swift restoration of mission-critical services.
• Tune, validate, and manage time-series anomaly detection models within a production observability context to drive proactive incident mitigation.
• Design, implement, and maintain SLIs, SLOs, error budgets, and full-stack observability frameworks using Prometheus, Grafana, OpenTelemetry, and log aggregation suites.
• Write, deploy, and maintain production-quality operational tooling using Python, React, and Java across GCP, Rancher K3s, and Kubernetes environments.
• Diagnose and resolve complex workflow orchestration failures, batch processing issues, and scheduler performance bottlenecks within Airflow and Tidal data pipelines.
BENEFITS
Medical | Dental | Vision | 401(k)
EEOC Compliance:
We are an equal opportunity employer, and all qualified applicants will receive consideration for employment.
DISCLAIMER
AI Usage Policy: Pacer Group uses AI to assist in screening applications. Final hiring decisions are made by human recruiters based on qualifications and experience.
Salary : $41 - $50