What are the responsibilities and job description for the Site Reliability Engineer position at Scale.jobs?
About The Role
The role is responsible for the reliability, scalability, and performance of production systems running on distributed cloud infrastructure. This SRE team sits at the intersection of software engineering and systems engineering, ensuring that core platform services remain resilient under heavy load.
The engineer will focus on building automated tooling, self-healing infrastructure, and robust observability pipelines to eliminate manual operational work. This is a critical technical position where infrastructure is treated as a software problem, directly impacting the availability of customer-facing services.
Key Responsibilities
The role is responsible for the reliability, scalability, and performance of production systems running on distributed cloud infrastructure. This SRE team sits at the intersection of software engineering and systems engineering, ensuring that core platform services remain resilient under heavy load.
The engineer will focus on building automated tooling, self-healing infrastructure, and robust observability pipelines to eliminate manual operational work. This is a critical technical position where infrastructure is treated as a software problem, directly impacting the availability of customer-facing services.
Key Responsibilities
- Design, provision, and maintain secure, multi-region Kubernetes clusters on AWS or GCP using Terraform for Infrastructure as Code (IaC)
- Build and maintain robust CI/CD deployment pipelines using GitHub Actions, GitLab CI, or ArgoCD to ensure seamless, zero-downtime releases
- Develop, configure, and maintain comprehensive observability dashboards and alerting systems using Prometheus, Grafana, and Datadog
- Participate in a blameless on-call rotation, leading incident response, conducting root-cause analysis (RCA), and implementing preventative automation
- Optimize cloud spend, system performance, and network latency by conducting capacity planning and performance bottleneck analysis
- Write custom tooling and automation scripts in Go or Python to replace manual operations and standard runbooks
- 3–6 years of experience in a Site Reliability Engineering, DevOps, or systems engineering role managing high-throughput production environments
- Strong proficiency in containerization and orchestration technologies, specifically Docker and Kubernetes (EKS, GKE, or self-managed)
- Deep familiarity with Infrastructure as Code (IaC) principles and tools, with advanced knowledge of Terraform or Pulumi
- Solid software engineering skills in at least one backend language, preferably Go, Python, or Ruby, for writing automation tools and API integrations
- Hands-on experience configuring observability stacks and defining Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
- Bonus: Experience with service mesh technologies (Istio, Linkerd), cloud security compliance (SOC2, ISO27001), or managing multi-tenant SaaS databases at scale