What are the responsibilities and job description for the Senior LLMOps / MLOps Engineer position at Pacific Consultancy Services?
Senior LLMOps / MLOps Engineer
Santa Clara, CA (Onsite)
Contract
UST Global // Applied Materials (AMAT)
Job Summary
We are seeking an experienced Senior LLMOps / MLOps Engineer to design, build, and optimize enterprise-scale AI platforms supporting Large Language Models (LLMs) in production environments. This role requires deep expertise in deploying, hosting, serving, and optimizing open-source LLMs while building highly scalable, GPU-accelerated inference platforms. The ideal candidate is a hands-on engineer with strong experience in MLOps, cloud-native AI infrastructure, Kubernetes, model serving frameworks, and production-grade GenAI applications.
The successful candidate will work closely with AI researchers, data scientists, and platform engineering teams to operationalize LLMs, improve inference performance, optimize GPU utilization, and build reliable AI infrastructure for enterprise workloads.
Key Responsibilities
- Design, develop, deploy, and maintain scalable LLMOps and MLOps platforms for production AI workloads.
- Build and optimize inference pipelines for open-source Large Language Models including Llama, Mistral, Gemma, and Qwen.
- Deploy and manage high-performance model serving platforms using vLLM, SGLang, Triton Inference Server, Ray Serve, TGI, Azure ML, or Databricks Model Serving.
- Optimize GPU utilization, inference latency, throughput, memory usage, and infrastructure cost.
- Implement advanced inference optimization techniques including KV Cache, PagedAttention, Continuous Batching, Dynamic Batching, and Quantization.
- Build scalable containerized AI workloads using Kubernetes and Docker.
- Develop automation for model deployment, versioning, monitoring, rollback, and lifecycle management.
- Integrate MLflow for experiment tracking, model registry, and deployment pipelines.
- Deploy and maintain Retrieval-Augmented Generation (RAG) solutions with vector databases.
- Monitor production AI systems for performance, availability, scalability, and reliability.
- Troubleshoot inference bottlenecks, model serving issues, and distributed deployment challenges.
- Collaborate with cross-functional teams to operationalize Generative AI applications in enterprise environments.
- Implement AI governance, observability, responsible AI practices, and production monitoring.
- Build reusable deployment frameworks and automation pipelines supporting enterprise AI initiatives.
- Document architecture, deployment standards, operational procedures, and best practices.
Required Qualifications
- Bachelor''s or Master''s degree in Computer Science, Artificial Intelligence, Data Science, Software Engineering, or a related technical discipline.
- 5–7 years of experience in MLOps, LLMOps, AI Platform Engineering, or Machine Learning Infrastructure.
- Strong programming expertise in Python with solid software engineering fundamentals.
- Hands-on experience deploying and serving open-source Large Language Models.
- Strong experience with model serving frameworks including:
- vLLM
- SGLang
- Triton Inference Server
- Ray Serve
- Hugging Face TGI
- Azure ML Model Serving
- Databricks Model Serving
- Experience building scalable AI platforms using Kubernetes and Docker.
- Experience with Azure ML, Databricks, and MLflow.
- Strong understanding of:
- LLM Inferencing
- Model Hosting
- GPU Optimization
- Quantization
- KV Cache
- PagedAttention
- Continuous Batching
- Dynamic Batching
- Vector Databases
- Retrieval-Augmented Generation (RAG)
- Proven experience deploying, scaling, monitoring, troubleshooting, and optimizing production-grade LLM applications.
- Familiarity with AI Observability, Governance, Model Monitoring, and Responsible AI practices.
Preferred Qualifications
- Experience fine-tuning Large Language Models using:
- PEFT
- LoRA
- QLoRA
- Supervised Fine-Tuning (SFT)
- Continued Pre-Training (CPT)
- Experience with:
- Azure AI Foundry
- Azure OpenAI
- Hugging Face Ecosystem
- DeepSpeed
- PEFT
- Knowledge of distributed training, multi-GPU clusters, and large-scale AI infrastructure.
- Experience working with Agentic AI frameworks such as:
- LangGraph
- AutoGen
- CrewAI
- Understanding of simulation platforms, digital twins, scientific computing, or engineering simulation workflows.
Technical Skills
Programming
- Python
- Shell Scripting
LLMOps / MLOps
- MLflow
- Azure ML
- Databricks
- Model Registry
- Model Deployment
- Experiment Tracking
- CI/CD for ML
Model Serving
- vLLM
- SGLang
- Triton Inference Server
- Ray Serve
- Hugging Face TGI
- Azure ML Serving
- Databricks Model Serving
Large Language Models
- Llama
- Mistral
- Gemma
- Qwen
AI Infrastructure
- Kubernetes
- Docker
- GPU Clusters
- Distributed Computing
LLM Optimization
- Quantization
- KV Cache
- PagedAttention
- Continuous Batching
- Dynamic Batching
- GPU Optimization
Generative AI
- RAG
- Vector Databases
- Embeddings
- Prompt Engineering
Observability & Governance
- AI Monitoring
- Responsible AI
- AI Governance
- Model Performance Monitoring
Cloud Platforms
- Microsoft Azure
- Azure ML
- Databricks
Preferred Certifications
- Microsoft Certified: Azure AI Engineer Associate
- Microsoft Certified: Azure Data Scientist Associate
- Microsoft Certified: Azure Solutions Architect Expert
- Databricks Certified Machine Learning Professional
- Databricks Certified Data Engineer Professional
- Kubernetes Certified Application Developer (CKAD)
- Certified Kubernetes Administrator (CKA)
Ideal Candidate
The ideal candidate is a hands-on Senior LLMOps/MLOps Engineer with extensive experience building production-scale AI infrastructure, deploying open-source Large Language Models, optimizing GPU inference performance, and implementing scalable, enterprise-grade AI platforms. They should possess strong expertise in model serving, cloud-native AI systems, Kubernetes, Azure ML, Databricks, and modern LLM optimization techniques while collaborating effectively across AI, platform engineering, and DevOps teams to deliver high-performance Generative AI solutions.