Demo

Principal Systems Engineer

Nscale
Seattle, WA Full Time
POSTED ON 7/31/2026
AVAILABLE BEFORE 9/15/2026
Principal Systems Engineer – GPU Supercluster Bringup

About Us

We are building AI infrastructure for frontier-scale workloads. Our platform is designed for high-density, high-performance GPU clusters that push the limits of power, networking, and distributed compute.

As a startup, we move fast, operate with ownership, and expect technical leaders to define standards—not just follow them.

The Role

We are hiring a Principal Deployment Engineer to architect and lead the bringup of large-scale GPU clusters (hundreds to thousands of GPUs). This is a technical leadership role responsible for defining how we deploy, validate, and scale AI superclusters across sites.

You will own the full lifecycle of deployment—from rack design and fabric architecture to cluster validation frameworks and production readiness standards. You will set the bar for performance, reliability, and operational excellence.

This role combines deep hands-on expertise with system-level thinking and cross-functional leadership.

What You’ll Do

End-to-End Supercluster Bringup Ownership

  • Define the technical standards for node, rack, and full-cluster bringup.
  • Lead large-scale GPU cluster deployments (multi-rack, multi-pod environments).
  • Architect high-performance network fabrics (IB, RoCE, Ethernet) optimized for AI workloads.
  • Establish cluster-level acceptance criteria and validation frameworks.

Performance & Fabric Architecture

  • Tune and validate NCCL, RDMA, GPUDirect, and collective operations at scale.
  • Identify and eliminate performance bottlenecks across hardware, topology, and firmware layers.
  • Drive congestion control and fabric optimization strategies.
  • Define performance benchmarking methodology for AI training workloads.

Deployment Strategy & Scalability

  • Design repeatable deployment models for multi-site expansion.
  • Build automation frameworks for provisioning and cluster validation.
  • Establish deployment SLAs, quality gates, and operational readiness standards.
  • Reduce time-to-capacity while increasing reliability.

Technical Leadership

  • Serve as the escalation point for complex bringup and performance issues.
  • Mentor senior engineers and shape infrastructure best practices.
  • Influence hardware selection, rack topology, and data center design decisions.
  • Partner with executive leadership on infrastructure scaling strategy.

Required

What We’re Looking For

  • 10 years of experience in large-scale infrastructure or HPC environments.
  • Proven experience bringing up large GPU clusters (hundreds GPUs).
  • Deep expertise in high-speed networking (InfiniBand, RoCE, Ethernet fabrics).
  • Strong understanding of server architecture (PCIe, NUMA, memory hierarchy).
  • Experience debugging performance issues across compute and network layers.
  • Strong automation and systems-level thinking.

Strongly Preferred

  • Experience scaling AI training clusters for frontier models.
  • Experience with liquid cooling or ultra-high-density deployments.
  • Knowledge of distributed storage systems (Lustre, Ceph, NVMe-oF).
  • Experience defining infrastructure standards in a fast-growing organization.

What Success Looks Like

  • Superclusters are brought online quickly, predictably, and at peak performance.
  • Deployment processes scale from first cluster to multi-site expansion.
  • Infrastructure becomes a competitive advantage.
  • You define the technical blueprint for how we scale AI infrastructure.

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range: $175,000 USD - $225,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Salary : $175,000

If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Principal Systems Engineer?

Sign up to receive alerts about other jobs on the Principal Systems Engineer career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$163,289 - $195,234
Income Estimation: 
$136,356 - $178,393
Income Estimation: 
$117,033 - $148,289
Income Estimation: 
$178,619 - $225,190
Income Estimation: 
$132,903 - $169,021
Income Estimation: 
$144,671 - $184,917
Income Estimation: 
$136,361 - $179,761
Income Estimation: 
$86,891 - $130,303
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Nscale

  • Nscale Seattle, WA
  • About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise cu... more
  • 1 Day Ago

  • Nscale Bellevue, WA
  • About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise cu... more
  • 1 Day Ago

  • Nscale Seattle, WA
  • About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise cu... more
  • 1 Day Ago

  • Nscale Bellevue, WA
  • About Nscale Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise cu... more
  • 1 Day Ago


Not the job you're looking for? Here are some other Principal Systems Engineer jobs in the Seattle, WA area that may be a better fit.

  • Zeno Power Seattle, WA
  • Company Overview Zeno Power is the leading developer of nuclear batteries – compact power systems that provide reliable, clean energy in frontier environme... more
  • 2 Months Ago

  • Hubble Network Seattle, WA
  • Hubble Network was founded with the intention of delivering on the promise of what Internet-of-Things (IoT) was supposed to be. We're building a global Blu... more
  • 6 Days Ago

AI Assistant is available now!

Feel free to start your new journey!