Demo

Machine Learning Operations Lead

Together AI
San Francisco, CA Full Time
POSTED ON 11/18/2025
AVAILABLE BEFORE 12/17/2025
About The Role

Together AI is building the AI Inference & Model Shaping Platform that brings the most advanced generative AI models to the world. Our platform powers multi-tenant serverless workloads and dedicated endpoints, enabling developers, enterprises, and researchers to harness the latest LLMs, multimodal models, image, audio, video, and reasoning models at scale.

We are looking for an exceptional MLOps Engineering Lead to partner closely with our cross-functional engineering, infrastructure, research, and sales teams to ensure excellence of our ML API offerings. Your primary focus will be on delivering world-class inference and fine-tuning in our public APIs and customer deployments by building automation and operations processes.

This role is ideal for a highly motivated and technically adept individual who excels in fast-paced, dynamic environments. You will be in charge of designing and scaling our ML processes & tooling at production scale – optimizing operations to ensure availability and reliability for our services, across differing tenants and user loads, and in a multi-cluster deployment. You will serve as a passionate advocate for internal and external customers, providing feedback to the wider engineering and infrastructure teams to improve our systems and core business metrics. If you thrive in a collaborative, problem-solving environment and are driven to deliver operational excellence, we encourage you to apply for this exciting opportunity.

Responsibilities

  • Own availability and performance SLAs for production inference and fine-tuning services across serverless and dedicated deployments
  • Own & improve testing, deployment, configuration management, and monitoring practices for multi-cluster ML infrastructure – partnering closely with Infra SREs
  • Build self-serve tooling and automation to reduce operational toil and enable internal users (MLOps, customer experience) and self-serve offerings
  • Define and enforce configuration best practices for inference engines (vLLM, tvLLM, Pulsar) to prevent runtime issues
  • Lead incident response, conduct postmortems, and drive reliability improvements
  • Hire, mentor, and grow an MLOps engineering team
  • Partner with infrastructure and ML engineering teams to improve system reliability and cost efficiency

Requirements

  • 5 years operating production ML inference or training systems at scale
  • 2 years leading engineering teams, with experience building teams from scratch
  • Deep expertise with Kubernetes, multi-cluster orchestration, and ML serving frameworks
  • Strong track record owning production SLAs (e.g. availability, TTFT, TPS)
  • Experience with LLM inference serving systems (vLLM, TRT-LLM, or similar)
  • Ability to influence cross-functional teams and make deployment/architecture decisions

Nice to Have

  • Experience building internal developer platforms or self-serve tooling
  • Background in cost optimization for GPU infrastructure
  • Contributions to open-source ML infrastructure projects

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $160,000 - $280,000 equity benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy

Salary : $160,000 - $280,000

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Machine Learning Operations Lead?

Sign up to receive alerts about other jobs on the Machine Learning Operations Lead career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$77,900 - $95,589
Income Estimation: 
$101,387 - $124,118
Income Estimation: 
$77,900 - $95,589
Income Estimation: 
$101,387 - $124,118
Income Estimation: 
$101,387 - $124,118
Income Estimation: 
$119,030 - $151,900
Income Estimation: 
$179,606 - $233,815
Income Estimation: 
$211,413 - $298,244
Income Estimation: 
$119,030 - $151,900
Income Estimation: 
$149,493 - $192,976
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Together AI

Together AI
Hired Organization Address San Francisco, CA Contractor
As a Sales Development Engineer Representative at Together.ai, your territory is the world. You will have the opportunit...
Together AI
Hired Organization Address San Francisco, CA Full Time
Role Together AI is looking for an ML Engineer who will develop systems and APIs that enable our customers to perform in...
Together AI
Hired Organization Address San Francisco, CA Full Time
About The Role Together.ai is at the forefront of AI infrastructure development, creating robust platforms and framework...
Together AI
Hired Organization Address San Francisco, CA Full Time
This role focuses on enabling custom models and dedicated inference on Together. We are responsible for optimizing autos...

Not the job you're looking for? Here are some other Machine Learning Operations Lead jobs in the San Francisco, CA area that may be a better fit.

Machine Learning Operations Lead

togetherai, San Francisco, CA

Machine Learning Operations Engineer

Recruiting From Scratch, San Francisco, CA

AI Assistant is available now!

Feel free to start your new journey!