What are the responsibilities and job description for the Infrastructure Operations Technical Lead position at Thor?
US Infrastructure Operations Technical Lead | AI & HPC Platform | Remote (US East Coast)
Thor is exclusively retained on a search for a US Infrastructure Operations Technical Lead for a fast-scaling, UK-headquartered AI and HPC compute platform building out its US operations function.
This is a hands-on player-manager role for an infrastructure leader who wants to stay technical while building and leading a team, not a purely managerial seat.
**The perfect applicant comes with hands-on Nvidia GPU product experience**
What you'll do:
- Lead a growing US Infrastructure Operations team (currently 3 engineers), owning both people leadership and hands-on technical execution
- Stay deep in the stack: Linux performance and troubleshooting, networking (TCP/IP, routing, DNS, latency debugging), storage systems, and bare-metal infrastructure
- Own incident leadership for the US region and participate in on-call rotation, leading from the front on major incidents
- Coordinate closely with a UK-based Infrastructure Operations counterpart during overlapping morning hours
- Drive automation and Infrastructure as Code (Terraform, Ansible) to reduce operational toil
- Build observability practices (metrics, logs, traces, alerting) as a first-class discipline
- Hire, mentor, and grow the team as US operations scale
What you'll bring:
- 8 years in infrastructure engineering, SRE, or large-scale production operations
- 2 years leading engineers directly
- Deep Linux expertise and strong networking fundamentals down to the packet level
- Bare-metal operations experience (Redfish, IPMI, hardware lifecycle)
- Strong scripting (Python or Bash) and automation/config management experience
- Comfort in an incident-heavy, 24x7, on-call environment
Nice to have: HPC, AI/ML, or GPU compute infrastructure exposure; InfiniBand or other low-latency networking; distributed storage (Lustre, WEKA); ITIL or PMP certification.
On offer: Highly competitive base plus company share participation, private medical insurance, dedicated learning time, and the opportunity to help shape how a rapidly scaling AI infrastructure platform operates in the US.
This search is being run confidentially. If your background fits, reach out directly and I'll share full details under NDA.