Demo

Principal AI and ML Infra Software Engineer, GPU Clusters

NVIDIA AI
Redmond, WA Full Time
POSTED ON 5/11/2026
AVAILABLE BEFORE 5/27/2026
Job Requisition ID

JR2017185

Job Category

Engineering

Time Type

Full time

We are seeking a Principal AI and ML Infra Software Engineer, GPU Clusters at NVIDIA to join our Hardware Infrastructure team. As an Engineer, you will have a pivotal role in enhancing efficiency for our researchers by implementing progressions throughout the entire stack. Your main task will revolve around collaborating closely with customers to pinpoint and address infrastructure deficiencies, facilitating groundbreaking AI and ML research on GPU Clusters. Together, we can craft potent, effective, and scalable solutions as we mold the future of AI/ML technology!

What You Will Be Doing

  • Engage closely with our AI and ML research teams to discern their infrastructure requirements and barriers, converting those insights into actionable improvements.
  • Proactively identify researcher efficiency bottlenecks and lead initiatives to systematically improve it. Drive the direction and long-term roadmaps for such initiatives.
  • Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization.
  • Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results.
  • Work closely with a variety of teams, such as researchers, data engineers, and DevOps professionals, to develop a cohesive AI/ML infrastructure ecosystem.
  • Keep up to date with the most recent developments in AI/ML technologies, frameworks, and successful strategies, and advocate for their integration within the organization.

What We Need To See

  • BS or similar background in Computer Science or related area (or equivalent experience).
  • 15 years of demonstrated expertise in AI/ML and HPC tasks and systems.
  • Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure as well as in-depth knowledge of accelerated computing (e.g., GPU, custom silicon), storage (e.g., Lustre, GPFS, BeeGFS), scheduling & orchestration (e.g., Slurm, Kubernetes, LSF), high-speed networking (e.g., Infiniband, RoCE, Amazon EFA), and containers technologies (Docker, Enroot).
  • Capability in supervising and improving substantial distributed training operations using PyTorch (DDP, FSDP), NeMo, or JAX. Moreover, an in-depth understanding of AI/ML workflows, involving data processing, model training, and inference pipelines.
  • Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with cloud computing platforms (e.g., AWS, GCP, Azure) in addition to experience with parallel computing frameworks and paradigms.
  • Dedication to ongoing learning and staying updated on new technologies and innovative methods in the AI/ML infrastructure sector.
  • Excellent communication and collaboration skills, with the ability to work effectively with teams and individuals of different backgrounds.

NVIDIA offers competitive salaries and a comprehensive benefits package. Our engineering teams are growing rapidly due to outstanding expansion. If you're a passionate and independent engineer with a love for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until May 1, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Salary.com Estimation for Principal AI and ML Infra Software Engineer, GPU Clusters in Redmond, WA
$88,130 to $107,355
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Principal AI and ML Infra Software Engineer, GPU Clusters?

Sign up to receive alerts about other jobs on the Principal AI and ML Infra Software Engineer, GPU Clusters career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$92,775 - $114,342
Income Estimation: 
$115,086 - $141,907
Income Estimation: 
$77,657 - $95,021
Income Estimation: 
$97,257 - $120,701
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at NVIDIA AI

  • NVIDIA AI Santa Clara, CA
  • Job Requisition ID JR2011694 Job Category Engineering Time Type Full time At NVIDIA, we’re tapping into the unlimited potential of AI to define the next er... more
  • 4 Days Ago

  • NVIDIA AI Westford, MA
  • Job Requisition ID JR2011761 Job Category Engineering Time Type Full time NVIDIA has been transforming computer graphics, PC gaming, and accelerated comput... more
  • 5 Days Ago

  • NVIDIA AI Santa Clara, CA
  • Job Requisition ID JR2012302 Job Category Engineering Time Type Full time NVIDIA has been transforming computer graphics, PC gaming, and accelerated comput... more
  • 5 Days Ago

  • NVIDIA AI Santa Clara, CA
  • Job Requisition ID JR2007865 Job Category Engineering Time Type Full time We are now looking for a Senior Performance Verification Engineer! As a member of... more
  • 5 Days Ago


Not the job you're looking for? Here are some other Principal AI and ML Infra Software Engineer, GPU Clusters jobs in the Redmond, WA area that may be a better fit.

  • NVIDIA AI Redmond, WA
  • Job Requisition ID JR2015886 Job Category Engineering Time Type Full time NVIDIA is searching for a highly motivated, creative engineer to join the GPU Sof... more
  • 16 Days Ago

  • Microsoft AI Redmond, WA
  • Overview Microsoft AI is looking for a Principal Software Engineer - AI Ads , to shape the future of online advertising in Mountain View, CA or Redmond, WA... more
  • 23 Days Ago

AI Assistant is available now!

Feel free to start your new journey!