Demo

Engineering Manager (AI Inference)

Perplexity AI
San Francisco, CA Full Time
POSTED ON 5/24/2026
AVAILABLE BEFORE 6/23/2026
About the Role

We are looking for an Inference Engineering Manager to lead our AI Inference team. This is a unique opportunity to build and scale the infrastructure that powers Perplexity's products and APIs, serving millions of users with state-of-the-art AI capabilities.

You will own the technical direction and execution of our inference systems while building and leading a world-class team of inference engineers. Our current stack includes Python, PyTorch, Rust, C , and Kubernetes. You will help architect and scale the large-scale deployment of machine learning models behind Perplexity's Comet, Sonar, Search, Deep Research products.
Why Perplexity?
  • Build SOTA systems that are the fastest in the industry with cutting-edge technology
  • High-impact work on a smaller team with significant ownership and autonomy
  • Opportunity to build 0-to-1 infrastructure from scratch rather than maintaining legacy systems
  • Work on the full spectrum: reducing cost, scaling traffic, and pushing the boundaries of inference
  • Direct influence on technical roadmap and team culture at a rapidly growing company
Responsibilities
  • Lead and grow a high-performing team of AI inference engineers
  • Develop APIs for AI inference used by both internal and external customers
  • Architect and scale our inference infrastructure for reliability and efficiency
  • Benchmark and eliminate bottlenecks throughout our inference stack
  • Drive large sparse/MoE model inference at rack scale, including sharding strategies for massive models
  • Push the frontier with building inference systems to support sparse attention, disaggregated pre-fill/decoding serving, etc.
  • Improve the reliability and observability of our systems and lead incident response
  • Own technical decisions around batching, throughput, latency, and GPU utilization
  • Partner with ML research teams on model optimization and deployment
  • Recruit, mentor, and develop engineering talent
  • Establish team processes, engineering standards, and operational excellence
Qualifications
  • 5 years of engineering experience with 2 years in a technical leadership or management role
  • Deep experience with ML systems and inference frameworks (PyTorch, TensorFlow, ONNX, TensorRT, vLLM)
  • Strong understanding of LLM architecture: Multi-Head Attention, Multi/Grouped-Query Attention, and common layers
  • Experience with inference optimizations: batching, quantization, kernel fusion, FlashAttention
  • Familiarity with GPU characteristics, roofline models, and performance analysis
  • Experience deploying reliable, distributed, real-time systems at scale
  • Track record of building and leading high-performing engineering teams
  • Experience with parallelism strategies: tensor parallelism, pipeline parallelism, expert parallelism
  • Strong technical communication and cross-functional collaboration skills
Nice to Have
  • Experience with CUDA, Triton, or custom kernel development
  • Background in training infrastructure and RL workloads
  • Experience with Kubernetes and container orchestration at scale
  • Published work or contributions to inference optimization research

Salary : $300,000 - $485,000

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Engineering Manager (AI Inference)?

Sign up to receive alerts about other jobs on the Engineering Manager (AI Inference) career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$151,448 - $188,145
Income Estimation: 
$203,425 - $249,816
Income Estimation: 
$213,375 - $267,876
Income Estimation: 
$190,687 - $235,769
Income Estimation: 
$151,448 - $188,145
Income Estimation: 
$203,425 - $249,816
Income Estimation: 
$213,375 - $267,876
Income Estimation: 
$190,687 - $235,769
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Perplexity AI

  • Perplexity AI York, NY
  • Perplexity is seeking a highly experienced and hands-on Cloud Security Engineer to join our dynamic security team, revolutionizing the way people search an... more
  • 1 Day Ago

  • Perplexity AI York, NY
  • Perplexity is looking for experienced Backend Engineers to build the foundational systems behind our core products. As we move toward agentic products that... more
  • 1 Day Ago

  • Perplexity AI York, NY
  • Perplexity is looking for an Applied AI Engineer to design, build, and iterate on cutting-edge agents powering our core experience in Perplexity Computer. ... more
  • 1 Day Ago

  • Perplexity AI York, NY
  • Perplexity is redefining how people search, reason, and interact with information. Our API team sits at the core of this vision, designing and operating th... more
  • 1 Day Ago


Not the job you're looking for? Here are some other Engineering Manager (AI Inference) jobs in the San Francisco, CA area that may be a better fit.

  • Beacon Engineering Resources San Francisco, CA
  • Construction Project Manager Overview We are seeking a Construction Project Manager to support the planning and execution of construction projects. This ro... more
  • 17 Days Ago

  • Chime San Francisco, CA
  • About The Role Chime serves millions of members on a mission to help them unlock financial progress. Growth Engineering builds the product surfaces and sys... more
  • 21 Days Ago

AI Assistant is available now!

Feel free to start your new journey!