What are the responsibilities and job description for the Cluster Operations Engineer — MLOps & Orchestration position at Comet Cloud?
Cluster Operations Engineer — MLOps & Orchestration (preferably Boston/Dallas; hybrid or remote)
After we deploy a GPU cluster, someone has to make it a great place to actually train models. That's this role.
We build dedicated NVIDIA clusters — GB300/GB200 NVL72, HGX-class B300 nodes, Quantum-3 XDR Infiniband and more for enterprise customers who get a named engineer, not a ticket queue. You'd be that engineer for the platform layer: Kubernetes and Slurm orchestration, the NVIDIA software stack (DCGM, Base Command Manager, NCCL), control plane health, and the direct customer relationship that turns their workload requirements into cluster configuration.
You're a fit if you:
- Have run Kubernetes and/or Slurm in production and actually enjoy scheduler configuration
- Have a expertise in training models, workload orchestration, optimizing core infrastructure, and resolving performance issues
- Know the NVIDIA toolchain, or know one part deeply and want the rest
- Can monitor and operate control plane infrastructure without drama
- Like sitting with a customer's ML team as much as sitting in a terminal
- Write runbooks people actually use
Why here: Lean team, no process layer between you and the machines, and the newest NVIDIA hardware in production. Your customers know your name.
Boston metro or Dallas | Hybrid or remote | Periodic site travel
Comet Compute is an NVIDIA Inception member building to the NVIDIA Cloud Partner reference architecture. Equal opportunity employer.