Demo

Senior LLMOps / MLOps Engineer

Pacific Consultancy Services
Santa Clara, CA Contractor
POSTED ON 7/29/2026
AVAILABLE BEFORE 8/28/2026

Senior LLMOps / MLOps Engineer

Santa Clara, CA (Onsite)

Contract
UST Global // Applied Materials (AMAT)

Job Summary

We are seeking an experienced Senior LLMOps / MLOps Engineer to design, build, and optimize enterprise-scale AI platforms supporting Large Language Models (LLMs) in production environments. This role requires deep expertise in deploying, hosting, serving, and optimizing open-source LLMs while building highly scalable, GPU-accelerated inference platforms. The ideal candidate is a hands-on engineer with strong experience in MLOps, cloud-native AI infrastructure, Kubernetes, model serving frameworks, and production-grade GenAI applications.

The successful candidate will work closely with AI researchers, data scientists, and platform engineering teams to operationalize LLMs, improve inference performance, optimize GPU utilization, and build reliable AI infrastructure for enterprise workloads.

 

Key Responsibilities

  • Design, develop, deploy, and maintain scalable LLMOps and MLOps platforms for production AI workloads.
  • Build and optimize inference pipelines for open-source Large Language Models including Llama, Mistral, Gemma, and Qwen.
  • Deploy and manage high-performance model serving platforms using vLLM, SGLang, Triton Inference Server, Ray Serve, TGI, Azure ML, or Databricks Model Serving.
  • Optimize GPU utilization, inference latency, throughput, memory usage, and infrastructure cost.
  • Implement advanced inference optimization techniques including KV Cache, PagedAttention, Continuous Batching, Dynamic Batching, and Quantization.
  • Build scalable containerized AI workloads using Kubernetes and Docker.
  • Develop automation for model deployment, versioning, monitoring, rollback, and lifecycle management.
  • Integrate MLflow for experiment tracking, model registry, and deployment pipelines.
  • Deploy and maintain Retrieval-Augmented Generation (RAG) solutions with vector databases.
  • Monitor production AI systems for performance, availability, scalability, and reliability.
  • Troubleshoot inference bottlenecks, model serving issues, and distributed deployment challenges.
  • Collaborate with cross-functional teams to operationalize Generative AI applications in enterprise environments.
  • Implement AI governance, observability, responsible AI practices, and production monitoring.
  • Build reusable deployment frameworks and automation pipelines supporting enterprise AI initiatives.
  • Document architecture, deployment standards, operational procedures, and best practices.

 

Required Qualifications

  • Bachelor''s or Master''s degree in Computer Science, Artificial Intelligence, Data Science, Software Engineering, or a related technical discipline.
  • 5–7 years of experience in MLOps, LLMOps, AI Platform Engineering, or Machine Learning Infrastructure.
  • Strong programming expertise in Python with solid software engineering fundamentals.
  • Hands-on experience deploying and serving open-source Large Language Models.
  • Strong experience with model serving frameworks including:
    • vLLM
    • SGLang
    • Triton Inference Server
    • Ray Serve
    • Hugging Face TGI
    • Azure ML Model Serving
    • Databricks Model Serving
  • Experience building scalable AI platforms using Kubernetes and Docker.
  • Experience with Azure ML, Databricks, and MLflow.
  • Strong understanding of:
    • LLM Inferencing
    • Model Hosting
    • GPU Optimization
    • Quantization
    • KV Cache
    • PagedAttention
    • Continuous Batching
    • Dynamic Batching
    • Vector Databases
    • Retrieval-Augmented Generation (RAG)
  • Proven experience deploying, scaling, monitoring, troubleshooting, and optimizing production-grade LLM applications.
  • Familiarity with AI Observability, Governance, Model Monitoring, and Responsible AI practices.

 

Preferred Qualifications

  • Experience fine-tuning Large Language Models using:
    • PEFT
    • LoRA
    • QLoRA
    • Supervised Fine-Tuning (SFT)
    • Continued Pre-Training (CPT)
  • Experience with:
    • Azure AI Foundry
    • Azure OpenAI
    • Hugging Face Ecosystem
    • DeepSpeed
    • PEFT
  • Knowledge of distributed training, multi-GPU clusters, and large-scale AI infrastructure.
  • Experience working with Agentic AI frameworks such as:
    • LangGraph
    • AutoGen
    • CrewAI
  • Understanding of simulation platforms, digital twins, scientific computing, or engineering simulation workflows.

 

Technical Skills

Programming

  • Python
  • Shell Scripting

LLMOps / MLOps

  • MLflow
  • Azure ML
  • Databricks
  • Model Registry
  • Model Deployment
  • Experiment Tracking
  • CI/CD for ML

Model Serving

  • vLLM
  • SGLang
  • Triton Inference Server
  • Ray Serve
  • Hugging Face TGI
  • Azure ML Serving
  • Databricks Model Serving

Large Language Models

  • Llama
  • Mistral
  • Gemma
  • Qwen

AI Infrastructure

  • Kubernetes
  • Docker
  • GPU Clusters
  • Distributed Computing

LLM Optimization

  • Quantization
  • KV Cache
  • PagedAttention
  • Continuous Batching
  • Dynamic Batching
  • GPU Optimization

Generative AI

  • RAG
  • Vector Databases
  • Embeddings
  • Prompt Engineering

Observability & Governance

  • AI Monitoring
  • Responsible AI
  • AI Governance
  • Model Performance Monitoring

Cloud Platforms

  • Microsoft Azure
  • Azure ML
  • Databricks

 

Preferred Certifications

  • Microsoft Certified: Azure AI Engineer Associate
  • Microsoft Certified: Azure Data Scientist Associate
  • Microsoft Certified: Azure Solutions Architect Expert
  • Databricks Certified Machine Learning Professional
  • Databricks Certified Data Engineer Professional
  • Kubernetes Certified Application Developer (CKAD)
  • Certified Kubernetes Administrator (CKA)

 

Ideal Candidate

The ideal candidate is a hands-on Senior LLMOps/MLOps Engineer with extensive experience building production-scale AI infrastructure, deploying open-source Large Language Models, optimizing GPU inference performance, and implementing scalable, enterprise-grade AI platforms. They should possess strong expertise in model serving, cloud-native AI systems, Kubernetes, Azure ML, Databricks, and modern LLM optimization techniques while collaborating effectively across AI, platform engineering, and DevOps teams to deliver high-performance Generative AI solutions.

Hourly Wage Estimation for Senior LLMOps / MLOps Engineer in Santa Clara, CA
$76.00 to $96.00
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Senior LLMOps / MLOps Engineer?

Sign up to receive alerts about other jobs on the Senior LLMOps / MLOps Engineer career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$117,024 - $149,811
Income Estimation: 
$137,568 - $176,908
Income Estimation: 
$119,030 - $151,900
Income Estimation: 
$149,493 - $192,976
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Pacific Consultancy Services

  • Pacific Consultancy Services St. Louis, MI
  • ML Engineer I – Data Scientist (eCommerce Search) St. Louis, MI, USA (Onsite – 5 Days) Contract UST Global // Sigma-Aldrich Job Summary We are seeking an M... more
  • 1 Day Ago

  • Pacific Consultancy Services Cincinnati, OH
  • Job Title: MES Site IT Coordinator / Production Support Lead (Werum PAS-X) Location: Cincinatti OH (Onsite) Contract Hiring Manager Notes Mandatory areas –... more
  • 1 Day Ago

  • Pacific Consultancy Services Issaquah, WA
  • Lead II Software Testing (Automation QE with React) Issaquah, WA (Hybrid – 3 Days) Contract UST Global // Costco Mandatory Areas Must Have Skills – QE Auto... more
  • 2 Days Ago

  • Pacific Consultancy Services Atlanta, GA
  • Role :- Sr Salesforce Developer Work Location: Atlanta, GA Onsite Key Responsibilities: Collaborate with business stakeholders to gather, analyze, and docu... more
  • 3 Days Ago


Not the job you're looking for? Here are some other Senior LLMOps / MLOps Engineer jobs in the Santa Clara, CA area that may be a better fit.

  • Cyber 1 Armor Milpitas, CA
  • Senior MLOps / LLMOps Engineer Location : Milpitas 4 days onsite contracts We are looking for a Senior MLOps / LLMOps Engineer to help standardize and enha... more
  • 2 Days Ago

  • American IT Systems santclara, CA
  • Try to submit locals within 50 miles -3 resumes Senior LLMOps / MLOps Engineer Location: Santa Clara, CA 5days onsite Summary We are looking for a highly s... more
  • 2 Days Ago

AI Assistant is available now!

Feel free to start your new journey!