What are the responsibilities and job description for the Site Reliability Engineer position at Harrison Clarke?
We're partnering with a well-funded Series A AI infrastructure startup building a cloud-native platform designed to support highly available, distributed systems at scale.
As an early engineering hire, you'll play a key role in improving the reliability, scalability, and operational maturity of the platform. You'll work closely with software engineers to automate infrastructure, strengthen observability, and ensure production systems remain resilient as the company grows.
Key Responsibilities
- Operate and scale Kubernetes environments across AWS, Azure, or GCP.
- Build and maintain infrastructure using Terraform, Helm, and GitOps practices.
- Enhance platform reliability through monitoring, alerting, logging, and observability improvements.
- Automate deployments, scaling, recovery processes, and day-to-day operational tasks.
- Troubleshoot complex production issues across infrastructure, networking, and distributed systems.
- Improve production readiness, resilience, security, and overall platform performance.
- Partner with engineering teams to embed operational best practices into the development lifecycle.
- Support incident response, capacity planning, and disaster recovery initiatives.
Requirements
- 5-10 years' experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure.
- Strong experience running Kubernetes workloads in production.
- Hands-on experience with AWS, Azure, or GCP.
- Proven experience with Infrastructure as Code using Terraform.
- Familiarity with Helm, GitOps, Argo CD, or similar deployment tooling.
- Solid understanding of Linux, networking, DNS, load balancing, and cloud security.
- Experience building and maintaining CI/CD pipelines.
- Strong knowledge of observability tooling such as Prometheus, Grafana, and OpenTelemetry.
- Experience supporting distributed systems, including technologies such as Kafka, Redis, PostgreSQL, or similar.
- Proficiency in Go, Python, Bash, or another scripting/programming language.
- Excellent troubleshooting skills across application, infrastructure, and network layers.