Demo

Senior Site Reliability Engineer - Control Plane

Lambda
San Francisco, CA Full Time
POSTED ON 4/30/2025
AVAILABLE BEFORE 6/6/2025
In 2012, Lambda started with a crew of AI engineers publishing research at top machine-learning conferences. We began as an AI company built by AI engineers. That hasn't changed. Today, we're on a mission to be the world's top AI computing platform. We equip engineers with the tools to deploy AI that is fast, secure, affordable, and built to scale. Whether they need powerhouse GPU hardware on-site or the flexibility of cloud-based solutions, we've got the horsepower to make it happen. Lambda’s AI Cloud has been adopted by the world’s leading companies and research institutions including Anyscale, Rakuten, The AI Institute, and multiple enterprises with over a trillion dollars of market capitalization. Our goal is to make computation as effortless and ubiquitous as electricity.

If you'd like to build the world's best deep learning cloud, join us.

  • Note: This position requires presence in our San Francisco office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

What You’ll Do

  • Design and implement cloud-native architectures that deliver the "four nines" (99.99%) of reliability while balancing performance and cost efficiency
  • Develop comprehensive monitoring and alerting systems with actionable dashboards that provide real-time visibility into system health
  • Implement SLIs, SLOs, and SLAs across services and maintain error budgets to guide development priorities
  • Automate deployments using tools like Argo and Terraform
  • Create robust incident management processes, escalation paths, and documentation
  • Architect fault-tolerant systems with graceful degradation capabilities to handle component failures
  • Design and implement disaster recovery solutions with regular testing procedures
  • Lead post-incident reviews that focus on systemic improvements rather than individual blame
  • Champion reliability best practices and system design principles
  • Build automated, auditable, and compliant processes to improve efficiency and productivity

You

  • 5 years of experience in Site Reliability Engineering or DevOps roles
  • Strong understanding of cloud platforms (AWS, GCP, Azure) and their core services
  • Experience designing and implementing monitoring and observability solutions at scale
  • Proven track record managing production incidents and driving root cause analysis
  • Proficiency with Infrastructure as Code tools and CI/CD pipeline implementation
  • Strong understanding of network architecture, load balancing, and content delivery
  • Expertise in performance tuning and system optimization techniques
  • Experience with container orchestration platforms like Kubernetes
  • Knowledge of database administration and optimization strategies
  • Solid coding skills in at least one language (Python, Go, Bash) for automation

Nice to Have

  • Experience with high-throughput, low-latency systems
  • Knowledge of security best practices and implementing defense-in-depth strategies
  • Experience with multi-region and distributed systems and solving consistency/availability challenges
  • Background in chaos engineering or similar reliability testing methodologies
  • Understanding of compliance frameworks (SOC 2, ISO 27001, etc.)
  • Familiarity with message queuing systems and event-driven architectures
  • Background working with specialized computing hardware

Salary Range Information

Based on market data and other factors, the annual salary range for this position is $245,000 - $385,000. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, ~350 employees (2024) and growing fast
  • We offer generous cash & equity compensation
  • Our investors include Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, US Innovative Technology, Gradient Ventures, Mercato Partners, SVB, 1517, Crescent Cove.
  • We are experiencing extremely high demand for our systems, with quarter over quarter, year over year profitability
  • Our research papers have been accepted into top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
  • Health, dental, and vision coverage for you and your dependents
  • Commuter/Work from home stipends for select roles
  • 401k Plan with 2% company match (USA employees)
  • Flexible Paid Time Off Plan that we all actually use

A Final Note

You do not need to match all of the listed expectations to apply for this position. We are committed to building a team with a variety of backgrounds, experiences, and skills.

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Compensation Range: $245K - $385K

Salary : $245,000 - $385,000

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Senior Site Reliability Engineer - Control Plane?

Sign up to receive alerts about other jobs on the Senior Site Reliability Engineer - Control Plane career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$114,618 - $136,401
Income Estimation: 
$144,264 - $191,312
Income Estimation: 
$140,435 - $166,410
Income Estimation: 
$114,618 - $136,401
Income Estimation: 
$144,264 - $191,312
Income Estimation: 
$140,435 - $166,410
Income Estimation: 
$140,435 - $166,410
Income Estimation: 
$151,875 - $212,356
Income Estimation: 
$169,957 - $202,398
Income Estimation: 
$95,852 - $118,073
Income Estimation: 
$100,690 - $126,032
Income Estimation: 
$120,143 - $165,703
Income Estimation: 
$92,877 - $110,401
Income Estimation: 
$120,933 - $155,034
Income Estimation: 
$114,618 - $136,401
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Lambda

Lambda
Hired Organization Address San Francisco, CA Full Time
In 2012, Lambda started with a crew of AI engineers publishing research at top machine-learning conferences. We began as...
Lambda
Hired Organization Address San Jose, CA Full Time
In 2012, Lambda started with a crew of AI engineers publishing research at top machine-learning conferences. We began as...
Lambda
Hired Organization Address Dallas, TX Full Time
In 2012, Lambda started with a crew of AI engineers publishing research at top machine-learning conferences. We began as...
Lambda
Hired Organization Address San Francisco, CA Full Time
In 2012, Lambda started with a crew of AI engineers publishing research at top machine-learning conferences. We began as...

Not the job you're looking for? Here are some other Senior Site Reliability Engineer - Control Plane jobs in the San Francisco, CA area that may be a better fit.

Senior Software Engineer - Control Plane

Snowflake Computing, San Mateo, CA

Senior Site Reliability Engineer

Tekfen Ventures, San Francisco, CA

AI Assistant is available now!

Feel free to start your new journey!