Demo

Principal ML Ops Engineer

Jobgether
Italy, TX Full Time
POSTED ON 7/19/2026
AVAILABLE BEFORE 9/18/2026

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal ML Ops Engineer based in Italy.

This is a highly technical leadership opportunity to build and scale the infrastructure powering next-generation AI applications.
The role focuses on designing reliable, efficient, and production-grade machine learning platforms capable of supporting large-scale AI workloads.
You will take ownership of model serving systems, GPU-powered infrastructure, deployment pipelines, and operational excellence.
Working closely with infrastructure, platform, and AI teams, you will help shape the foundations of a modern AI-native cloud ecosystem.
This position offers the chance to solve complex distributed systems challenges while improving performance, scalability, and cost efficiency.
You will play a key role in defining engineering standards and building critical ML infrastructure from the ground up.

\n


Accountabilities:
  • Design, build, and operate production-grade ML inference infrastructure using modern model serving frameworks such as vLLM, TGI, Triton, or equivalent solutions.
  • Develop scalable deployment pipelines supporting reliable model releases through strategies such as blue/green deployments and canary rollouts.
  • Build and maintain auto-scaling systems, multi-model serving architectures, and intelligent request routing mechanisms.
  • Optimize GPU utilization, memory efficiency, network performance, and model artifact storage to improve system reliability and cost effectiveness.
  • Implement observability solutions to monitor inference latency, throughput, GPU usage, operational health, and infrastructure costs.
  • Manage model registries, CI/CD workflows, and automation processes to enable reproducible and efficient model deployments.
  • Own the complete lifecycle of ML systems, from development and deployment through production operations and ongoing support.
  • Establish engineering best practices and contribute to platform architecture decisions in a fast-moving, remote-first environment.
  • Collaborate with infrastructure, platform, and applied AI teams to deliver scalable and reliable AI systems.

Requirements:

  • 4 years of experience in ML Ops, Platform Engineering, SRE, or similar infrastructure-focused roles supporting machine learning systems.
  • Strong hands-on experience with production model serving frameworks such as vLLM, TGI, Triton, or comparable technologies.
  • Proven experience operating GPU-based workloads and managing containerized environments in production.
  • Strong understanding of MLOps practices, including model registries, experiment tracking, automated deployment pipelines, and lifecycle management.
  • Proficiency in Python and infrastructure-as-code tools such as Terraform, Helm, or similar technologies.
  • Solid knowledge of distributed systems, performance optimization, scalability, and reliability engineering principles.
  • Experience using AI coding assistants to accelerate software development, troubleshooting, and debugging workflows.
  • Ability to work independently with strong ownership and accountability in a remote-first environment.
  • Experience with ML platforms such as Kubeflow, MLflow, or KubeAI is a plus.
  • Knowledge of GPU scheduling, CUDA/ROCm optimization, multi-tenant inference systems, and infrastructure cost optimization is advantageous.
  • Previous experience building greenfield infrastructure projects or working in early-stage technology environments is highly valued.

Benefits:

  • Opportunity to own and shape critical ML infrastructure for a rapidly scaling AI-focused technology platform.
  • Fully remote working environment with flexibility to work from Romania.
  • Chance to build foundational systems from the ground up rather than maintaining legacy infrastructure.
  • Exposure to cutting-edge technologies across distributed systems, GPU computing, and large-scale AI model serving.
  • High level of ownership and influence over technical decisions and engineering practices.
  • Opportunity to work with experienced professionals solving complex AI infrastructure challenges.
  • Dynamic startup environment with strong growth opportunities and meaningful technical impact.


\n

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

 Why Apply Through Jobgether? 

 

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

 

 

#LI-CL1

Salary.com Estimation for Principal ML Ops Engineer in Italy, TX
$123,000 to $157,932
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Principal ML Ops Engineer?

Sign up to receive alerts about other jobs on the Principal ML Ops Engineer career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$119,030 - $151,900
Income Estimation: 
$149,493 - $192,976
Income Estimation: 
$119,030 - $151,900
Income Estimation: 
$149,493 - $192,976
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Jobgether

  • Jobgether Brazil, IN
  • This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an Senior Game Publishing A... more
  • 12 Days Ago

  • Jobgether Brazil, IN
  • This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Data Analytics and Visual... more
  • 12 Days Ago

  • Jobgether Brazil, IN
  • This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a DESENVOLVEDOR JAVA - perf... more
  • 12 Days Ago

  • Jobgether Brazil, IN
  • This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Full Stack Enginee... more
  • 13 Days Ago


Not the job you're looking for? Here are some other Principal ML Ops Engineer jobs in the Italy, TX area that may be a better fit.

  • Southern Glazer's Wine & Spirits Dallas, TX
  • Principal AI/ML Ops Platform Engineer Job ID: 41851 Location: Dallas, TX, US Overview Employer: Southern Glazer’s Wine and Spirits LLC Job Title: Principal... more
  • 15 Days Ago

  • Southern Glazers Dallas, TX
  • Overview Employer: Southern Glazer’s Wine and Spirits LLC Job Title: Principal AI/ML Operations Platform Engineer Locations: 2300 SW 145th Avenue, Miramar,... more
  • 15 Days Ago

AI Assistant is available now!

Feel free to start your new journey!