What are the responsibilities and job description for the Lead GPU Cluster Solution Architect position at Axe Compute?
ABOUT THE ROLE
Axe Compute is seeking a Lead Architect, GPU Cluster Solutions to design GPU cluster configurations for prospective and signed engagements, translating client requirements and NVIDIA Reference Architecture into a buildable, supportable design, spanning compute, storage, networking, software and spares strategy for support.
ROLE AT A GLANCE
- Mandate: Own end-to-end technical design of GPU cluster deployments, from client requirements through NVIDIA Reference Architecture compliance and sparing strategy.
- Scope: Cluster design, network architecture (InfiniBand/RoCE/Ethernet), sparing and spares planning, connectivity design (internet/VPN/firewall, dedicated circuits), design adjustments for site and hardware constraints.
- Key Outcomes: Designs that meet client SLAs and NVIDIA Reference Architecture standards, sparing plans that protect uptime commitments, designs that account for real-world site and hardware lead-time constraints.
WHAT YOU WILL OWN
Cluster Design & Reference Architecture
- Design GPU cluster configurations (compute, storage, networking) against NVIDIA Reference Architecture for each signed engagement.
- Translate client technical requirements into a complete bill of design, including all necessary compute, storage, and networking components.
Network Architecture
- Design network topology and fabric selection, including InfiniBand, RoCE, and Ethernet options, appropriate to each client's workload and performance requirements.
- Incorporate internet, VPN, and firewall connectivity requirements into cluster designs.
- Design dedicated point-to-point network requirements where needed, including protected optical circuits and similar dedicated connectivity.
Sparing & Availability Strategy
- Formulate and own the hot/cold sparing plan for each deployment to meet contracted SLA commitments.
- Adjust sparing and design assumptions based on data center power/cooling parameters and hardware lead-time constraints.
Design Adaptation & Site Constraints
- Adjust cluster designs to fit site-specific power, cooling, and space constraints identified by the Data Center Procurement and Operations Director.
- Work with Supply Chain to align design decisions with realistic hardware delivery timing.
Cross-Functional Collaboration
- Partner with the VP, Deployments and Deployment Program Manager to ensure designs translate cleanly into buildable, trackable project plans.
- Support acceptance test design and criteria definition, ensuring test procedures validate the as-designed architecture.
REQUIRED QUALIFICATIONS
- 7 years in solutions architecture, network engineering, or systems engineering supporting GPU, HPC, or large-scale compute infrastructure.
- Deep working knowledge of NVIDIA Reference Architecture (HGX, NVL72) and GPU cluster design principles.
- Hands-on experience with InfiniBand, RoCE, and high-speed Ethernet fabric design.
- Experience designing sparing/spares strategies for mission-critical infrastructure.
- Experience incorporating firewall, VPN, and dedicated circuit (e.g., protected optical) requirements into network designs.
- Experience with high-speed shared storage solutions (e.g., Weka, Vast, DDN).
PREFERRED QUALIFICATIONS
- Experience designing clusters for large enterprise clients or neoclouds, not just internal infrastructure.
- Familiarity with NVIDIA NCP program requirements and certification processes.
- Experience with capacity or sparing modeling tools.
Salary : $140,000 - $240,000