Demo

Site Reliability Engineer

Harrison Clarke
Palo Alto, CA Full Time
POSTED ON 7/21/2026
AVAILABLE BEFORE 8/18/2026

We're partnering with a well-funded Series A AI infrastructure startup building a cloud-native platform designed to support highly available, distributed systems at scale.


As an early engineering hire, you'll play a key role in improving the reliability, scalability, and operational maturity of the platform. You'll work closely with software engineers to automate infrastructure, strengthen observability, and ensure production systems remain resilient as the company grows.


Key Responsibilities


  • Operate and scale Kubernetes environments across AWS, Azure, or GCP.
  • Build and maintain infrastructure using Terraform, Helm, and GitOps practices.
  • Enhance platform reliability through monitoring, alerting, logging, and observability improvements.
  • Automate deployments, scaling, recovery processes, and day-to-day operational tasks.
  • Troubleshoot complex production issues across infrastructure, networking, and distributed systems.
  • Improve production readiness, resilience, security, and overall platform performance.
  • Partner with engineering teams to embed operational best practices into the development lifecycle.
  • Support incident response, capacity planning, and disaster recovery initiatives.


Requirements


  • 5-10 years' experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure.
  • Strong experience running Kubernetes workloads in production.
  • Hands-on experience with AWS, Azure, or GCP.
  • Proven experience with Infrastructure as Code using Terraform.
  • Familiarity with Helm, GitOps, Argo CD, or similar deployment tooling.
  • Solid understanding of Linux, networking, DNS, load balancing, and cloud security.
  • Experience building and maintaining CI/CD pipelines.
  • Strong knowledge of observability tooling such as Prometheus, Grafana, and OpenTelemetry.
  • Experience supporting distributed systems, including technologies such as Kafka, Redis, PostgreSQL, or similar.
  • Proficiency in Go, Python, Bash, or another scripting/programming language.
  • Excellent troubleshooting skills across application, infrastructure, and network layers.



Salary.com Estimation for Site Reliability Engineer in Palo Alto, CA
$91,966 to $120,104
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Site Reliability Engineer?

Sign up to receive alerts about other jobs on the Site Reliability Engineer career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$114,618 - $136,401
Income Estimation: 
$144,264 - $191,312
Income Estimation: 
$140,435 - $166,410
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Harrison Clarke

  • Harrison Clarke Palo Alto, CA
  • AI Agent Engineer | Palo Alto | Seed-Stage, VC-Backed We are working with a well-funded, early-stage AI startup based in Palo Alto that is building next-ge... more
  • 9 Days Ago


Not the job you're looking for? Here are some other Site Reliability Engineer jobs in the Palo Alto, CA area that may be a better fit.

  • Candidate Experience site Sunnyvale, CA
  • We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and trou... more
  • 29 Days Ago

  • Candidate Experience site Sunnyvale, CA
  • Join Fortinet, a cybersecurity pioneer with over two decades of excellence, as we continue to shape the future of cybersecurity and redefine the intersecti... more
  • 1 Month Ago

AI Assistant is available now!

Feel free to start your new journey!