What are the responsibilities and job description for the Site Reliability Engineer position at Scale.jobs?
About The Role
The role focuses on building, maintaining, and scaling the infrastructure that powers high-throughput production systems. The team ensures application reliability, uptime, and performance by designing self-healing systems and reducing operational toil through automation.
You will collaborate closely with software engineering teams to architect resilient microservices, manage multi-region cloud environments, and establish robust monitoring and incident response frameworks.
Key Responsibilities
The role focuses on building, maintaining, and scaling the infrastructure that powers high-throughput production systems. The team ensures application reliability, uptime, and performance by designing self-healing systems and reducing operational toil through automation.
You will collaborate closely with software engineering teams to architect resilient microservices, manage multi-region cloud environments, and establish robust monitoring and incident response frameworks.
Key Responsibilities
- Design and implement Infrastructure as Code (IaC) using Terraform to manage AWS resources across multi-account environments
- Build and maintain Kubernetes clusters, managing container orchestration, ingress controllers, and service meshes at scale
- Develop automated CI/CD pipelines using GitHub Actions or GitLab CI to ensure safe, repeatable, and fast deployments
- Establish comprehensive observability dashboards and alerting rules using Prometheus, Grafana, and Datadog to minimize Mean Time to Resolution (MTTR)
- Participate in an on-call rotation, conducting blameless post-mortems and implementing long-term engineering fixes to prevent incident recurrence
- Optimize cloud infrastructure spend, system latency, and resource utilization through continuous profiling and capacity planning
- 4 years of professional experience in Site Reliability Engineering, DevOps, or Systems Engineering roles
- Deep expertise with AWS services (EC2, EKS, RDS, VPC, IAM) and containerization technologies like Docker and Kubernetes
- Strong scripting and automation skills in Python, Go, or Bash for tool development and infrastructure automation
- Production experience managing infrastructure as code with Terraform and configuration management with Ansible
- Solid understanding of networking concepts (DNS, TCP/IP, load balancing, SSL/TLS) and Linux system administration internals
- Bonus: Experience with service meshes like Istio, database administration at scale, or holding AWS Solutions Architect / CKA certifications