What are the responsibilities and job description for the NOC Engineer position at Axe Compute?
ABOUT THE ROLE
Axe Compute is seeking NOC Engineers who do more than watch a screen. This role owns customer experience for live GPU cluster deployments, using automation, AI tools, and continuous process improvement to drive SLA performance, not just monitor it. We're looking for people who build tools, write and improve their own runbooks and SOPs, and treat every shift as an opportunity to make the NOC better than they found it.
ROLE AT A GLANCE
- Mandate: Monitor live GPU clusters, drive incident response, and continuously improve the tools, processes, and SOPs the NOC runs on.
- Scope: 24/7 shift-based monitoring and incident response, tool-building and automation, runbook/SOP creation and maintenance, third-party coordination, SLA metric improvement.
- Key Outcomes: SLA compliance, a measurably improving NOC (fewer repeat incidents, faster resolution, better tooling), strong customer experience outcomes.
WHAT YOU WILL OWN
Monitoring & Incident Response
- Monitor cluster health, power, cooling, and network status across live deployments, triaging and resolving incidents within SLA.
- Escalate to Tier 3 (OEM or data center operator) only when an issue genuinely requires vendor-level support.
Tooling, Automation & AI Workflows
- Build and improve internal tools and automation to reduce manual, repetitive work, using AI-enabled workflows wherever they make the team faster or more accurate.
- Identify opportunities to automate recurring monitoring, triage, or reporting tasks rather than performing them manually shift after shift.
Runbooks, SOPs & Process Improvement
- Create and maintain runbooks and standard operating procedures, updating them based on real incidents rather than leaving them static.
- Contribute to after-action reviews and turn recurring issues into permanent process or tooling fixes.
Customer Experience Ownership
- Own the customer experience of every incident, communicating clearly and proactively rather than treating tickets as a checklist.
- Work directly with third parties (data center operators, OEMs, network providers) as needed to resolve client-impacting issues.
SLA & Continuous Improvement
- Track and actively work to improve SLA metrics, using data and software tooling to identify where performance is slipping before it becomes a client-facing problem.
REQUIRED QUALIFICATIONS
- 1 years in a NOC, network operations, or infrastructure monitoring role. This role can be junior, but not passive.
- Demonstrated interest or experience in automation, scripting, or tool-building, not just following existing procedures.
- Comfort working rotating shifts, including nights/weekends, as part of a 24/7 coverage model.
- Familiarity with monitoring tools (e.g., Datadog, Grafana, PagerDuty) and a genuine curiosity about AI-enabled operations tools.
PREFERRED QUALIFICATIONS
- Experience monitoring GPU or high-performance computing infrastructure.
- Scripting ability (Python, Bash, or similar) to build or modify automation.
- Additional languages beyond English are a plus for coordinating with clients and global vendors.
- Multiple shift options available to accommodate different timezones and flexible working hours.
Salary : $75,000 - $140,000