What are the responsibilities and job description for the ML Infrastructure Engineer position at Top Gen AI Jobs?
Home/Jobs/ML Infrastructure Engineer
ML Infrastructure Engineer
DocVerify
United State
3 years
Today
$21.7K–32.5K/yr
Full-time
Remote
Skills Required
Python
Kubernetes
TensorRT
Triton
GCP
Cloud Run
GKE
Vertex AI
Docker
AWS
TorchServe
TF Serving
Go
Rust
C
Description
Engineering role focused on building ML infrastructure for a forensic engine that runs multiple vision models per request. The goal is to support scalable training, deployment, and low-latency serving with full auditability and zero-downtime releases.
Role: ML Infrastructure Engineer
Location: Remote (EU / US)
Experience
CI/CD, model versioning, canary deployments, A/B testing, batching, quantization, hardware-aware compilation, monitoring, alerting, drift detection, throughput, GPU, TPU, MLflow, Kubeflow, ONNX Runtime
Prepare for this role
Recommended resources to build the skills for this position. Sponsored.
Top 15 Together AI Interview Questions
Zenaique
A Together AI-focused interview question set from Zenaique.
Top 25 xAI Interview Questions
Zenaique
An xAI-focused interview question set from Zenaique.
Top 25 LLM Evaluation Interview Questions
Zenaique
Curated LLM evaluation questions covering benchmarks, evals, and human review.
More Python jobs
AI / ML Engineer
Accenture
Hyderabad
Today
AI/ML Fullstack Engineer – GenAI, Python & React
S Talent HUB
Bengaluru
Today
AI Systems Engineer
Michael Page
Bengaluru
Today
Agentic AI Solutions_Specialist
Accenture
Mumbai
Today
AI Solution Engineer
Sapiens
Bengaluru
Today
GenAI Developer
Sapiens
Bengaluru
Today
ML Infrastructure Engineer
DocVerify
United State
3 years
Today
$21.7K–32.5K/yr
Full-time
Remote
Skills Required
Python
Kubernetes
TensorRT
Triton
GCP
Cloud Run
GKE
Vertex AI
Docker
AWS
TorchServe
TF Serving
Go
Rust
C
Description
Engineering role focused on building ML infrastructure for a forensic engine that runs multiple vision models per request. The goal is to support scalable training, deployment, and low-latency serving with full auditability and zero-downtime releases.
Role: ML Infrastructure Engineer
Location: Remote (EU / US)
Experience
- 3 years in ML infrastructure, MLOps, or backend systems engineering
- Strong experience with Kubernetes, Docker, and cloud-native deployments on GCP or AWS
- Hands-on experience with model serving frameworks such as Triton, TorchServe, or TF Serving
- Proficiency in Python and at least one systems language: Go, Rust, or C
- Experience with CI/CD pipelines for ML: MLflow, Kubeflow, or custom
- Understanding of GPU inference optimization: TensorRT, ONNX Runtime
- Own the systems that train, version, deploy, and serve models at scale
- Build and maintain model serving infrastructure on GCP using Cloud Run, GKE, or Vertex AI
- Design the CI/CD pipeline for model training, evaluation, and promotion to production
- Implement model versioning, canary deployments, and A/B testing frameworks
- Optimize inference latency through batching, quantization, and hardware-aware compilation
- Build monitoring and alerting for model performance, drift detection, and throughput
- Manage GPU and TPU resource allocation and cost optimization across training and serving
- Ensure every API call returns in under 200ms while maintaining full auditability and zero-downtime deployments
- Experience with Vertex AI or SageMaker in production
- Background in real-time serving systems with strict latency SLAs under 200ms p99
- Familiarity with cost modeling for GPU workloads
- Experience with distributed training across multiple GPUs or nodes
- Terraform or Pulumi for infrastructure as code
CI/CD, model versioning, canary deployments, A/B testing, batching, quantization, hardware-aware compilation, monitoring, alerting, drift detection, throughput, GPU, TPU, MLflow, Kubeflow, ONNX Runtime
Prepare for this role
Recommended resources to build the skills for this position. Sponsored.
Top 15 Together AI Interview Questions
Zenaique
A Together AI-focused interview question set from Zenaique.
Top 25 xAI Interview Questions
Zenaique
An xAI-focused interview question set from Zenaique.
Top 25 LLM Evaluation Interview Questions
Zenaique
Curated LLM evaluation questions covering benchmarks, evals, and human review.
More Python jobs
AI / ML Engineer
Accenture
Hyderabad
Today
AI/ML Fullstack Engineer – GenAI, Python & React
S Talent HUB
Bengaluru
Today
AI Systems Engineer
Michael Page
Bengaluru
Today
Agentic AI Solutions_Specialist
Accenture
Mumbai
Today
AI Solution Engineer
Sapiens
Bengaluru
Today
GenAI Developer
Sapiens
Bengaluru
Today
Salary : $21,700 - $32,500