Demo

Principal Technical Program Manager (TPM) - AI Infrastructure Operations

Nscale
Seattle, WA Full Time
POSTED ON 5/22/2026
AVAILABLE BEFORE 6/19/2026
About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

Role Overview

As a Technical Program Manager (TPM) for AI Infrastructure Operations, you will be the operational backbone of our high-scale, high-performance AI and High-Performance Computing (HPC) environment. You will be responsible for driving complex, cross-functional programs that ensure the stability, availability, and growth of our cutting-edge GPU fleet and Infiniband network fabrics. This role requires a blend of deep technical understanding, rigorous program management, and a relentless focus on delivering against key operational metrics (SLAs, Uptime, Availability). You will bridge the gap between engineering execution and strategic business goals, directly impacting our ability to serve customer workloads at scale.

Key Responsibilities

  • Program Leadership: Own the planning, execution, and delivery of strategic operational programs, including new data center AI infrastructure build-outs, large-scale fleet software/firmware rollouts, and the implementation of new operational tooling (in partnership with SRE).
  • Metrics and Reporting: Establish, track, and drive accountability against critical infrastructure KPIs, specifically focusing on Availability (Target 97.5%) and Uptime (Target 99%). Develop clear dashboards and communication rhythms to provide leadership with real-time visibility into operational health, program status, and risk.
  • Process Engineering: Analyze and optimize operational workflows across Fleet Operations, Network Operations, and SRE teams. Drive the standardization of incident management, change management, and postmortem processes to reduce toil and improve Mean Time to Recovery (MTTR).
  • Cross-Functional Coordination: Serve as the primary liaison between engineering teams (Hardware, Compute Platform, Network), Data Center Operations, and external vendors (GPU, Network hardware). Proactively identify and resolve dependencies, risks, and roadblocks.
  • Capacity and Readiness: Partner with Data Science/Operation Programs to translate capacity planning models into actionable infrastructure delivery and readiness roadmaps. Ensure that new hardware (GPUs, NICs, switches) is successfully integrated into the operational control plane and meets go-live criteria.
  • Risk Management: Proactively identify technical, schedule, and resource risks related to AI infrastructure scaling and stability. Develop mitigation strategies and communicate impacts clearly to stakeholders.

Required Qualifications

  • Experience: 5 years of experience in a Technical Program Management role, successfully driving large-scale, complex infrastructure or software engineering programs.
  • Technical Domain Knowledge: Strong foundational understanding of data center infrastructure, distributed systems, Linux, and networking concepts.
  • Program Management Rigor: Proven expertise in modern program management methodologies (Agile, Scrum, PMP certification preferred). Exceptional organizational, communication, and presentation skills.
  • Metrics-Driven Approach: Demonstrable experience in defining, tracking, and improving system performance based on operational metrics (e.g., Uptime, Availability, MTTR, SLOs/SLIs).
  • Execution in Ambiguity: Ability to thrive in a fast-paced, high-growth environment, managing multiple priorities and adapting to evolving technical requirements.

Preferred Qualifications

  • Direct experience managing programs related to data center infrastructure build-outs and hardware commissioning processes.
  • Specific domain knowledge of AI/HPC infrastructure, including NVIDIA GPUs, InfiniBand/RDMA networks, and the challenges of tightly-coupled systems.
  • Experience in a hyperscale or public cloud environment supporting 24/7 mission-critical services.
  • Familiarity with SRE principles, automation tooling, and continuous integration/continuous deployment (CI/CD) pipelines for infrastructure.
  • A Bachelor's or Master's degree in a technical field (Computer Science, Engineering, etc.) or equivalent practical experience.

What We Can Offer You

At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.

  • Highly competitive package (base equity) with reviews every 12 months.
  • Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
  • Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
  • Human-First Flexibility: We treat you as humans first. Our flexible workplace trusts Nscalers to deliver, giving you the autonomy to shape your day around life's moments.

Join our thriving remote-first team. Geography is no barrier to impact or connection. We build seamless virtual collaboration, empowering you, wherever you work.

Equal Opportunities Statement

We strongly encourage applications from people of colour, the LGBTQ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there's anything we can do to accommodate your specific situation, please let us know.

The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Salary.com Estimation for Principal Technical Program Manager (TPM) - AI Infrastructure Operations in Seattle, WA
$170,683 to $212,761
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Principal Technical Program Manager (TPM) - AI Infrastructure Operations?

Sign up to receive alerts about other jobs on the Principal Technical Program Manager (TPM) - AI Infrastructure Operations career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$138,649 - $191,575
Income Estimation: 
$182,502 - $249,036
Income Estimation: 
$207,946 - $249,343
Income Estimation: 
$175,165 - $219,883
Income Estimation: 
$182,642 - $260,237
Income Estimation: 
$161,406 - $211,884
Income Estimation: 
$188,022 - $236,092
Income Estimation: 
$205,940 - $255,928
Income Estimation: 
$199,907 - $266,531
Income Estimation: 
$195,700 - $270,403
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Nscale

  • Nscale Seattle, WA
  • Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nsca... more
  • 2 Days Ago

  • Nscale Point Pleasant, WV
  • About Nscale At Nscale, we are building the infrastructure for the AI revolution. Nscale is developing cutting-edge, sovereign generative AI solutions powe... more
  • 3 Days Ago

  • Nscale Point Pleasant, WV
  • About Nscale At Nscale, we are building the infrastructure for the AI revolution. Nscale is developing cutting-edge, sovereign generative AI solutions powe... more
  • 3 Days Ago

  • Nscale Seattle, WA
  • Role Overview We are seeking a customer-facing Solution Architect to lead and deepen our engagements with Nscale’s most strategic customers. This role sits... more
  • 3 Days Ago


Not the job you're looking for? Here are some other Principal Technical Program Manager (TPM) - AI Infrastructure Operations jobs in the Seattle, WA area that may be a better fit.

  • Microsoft AI Redmond, WA
  • Overview At Microsoft AI, we are on a mission to train the world’s most capable AI frontier models, pushing the boundaries of scale, performance, and produ... more
  • 4 Days Ago

  • Microsoft AI Redmond, WA
  • Overview Our team—affectionately known as the Deviation Detectives—plays a critical role in protecting the health and integrity of the Bing Ads marketplace... more
  • 2 Days Ago

AI Assistant is available now!

Feel free to start your new journey!