What are the responsibilities and job description for the Major Incident Management & NOC Lead position at Bi3 Careers?
Position Overview
We are seeking an experienced Major Incident Management & NOC Lead to lead enterprise NOC operations and manage critical P1/P2 incidents. This person will serve as the Incident Commander during major outages, coordinate technical teams and vendors, communicate with leadership, and drive incidents through resolution and root cause analysis.
The ideal candidate has a strong infrastructure/network operations background along with proven NOC and major incident leadership.
Key Responsibilities
- Lead NOC operations, monitoring, escalation and operational processes
- Serve as Incident Commander for P1/P2 incidents and major outages
- Coordinate network, infrastructure, application, security, cloud and vendor teams
- Provide clear incident updates to business and executive leadership
- Lead post-incident reviews, root cause analysis and corrective actions
- Manage vendor escalations and SLA performance
- Maintain SOPs, runbooks and escalation procedures
- Drive improvements in monitoring, alerting and service reliability
Required Experience
- 10 years in IT Operations, NOC and/or Major Incident Management
- Previous experience leading a NOC or enterprise operations team
- Strong P1/P2 incident management and outage leadership experience, including post-incident reviews, root cause analysis and problem management
- Strong knowledge of ITIL Incident, Problem and Change Management
- Experience with ServiceNow or a similar ITSM platform
- Experience with SLAs, SOPs, escalation procedures and vendor management
- Strong hands-on technical foundation in infrastructure and network operations, with the ability to lead network, Linux/Windows and operations teams
- Strong infrastructure knowledge including Windows/Linux, DNS, DHCP, TCP/IP, routing, load balancers and firewalls
- Experience with AWS and/or Azure
- Experience with monitoring tools such as Splunk, Dynatrace, Datadog, New Relic, AppDynamics or similar
- Strong leadership, problem-solving and communication skills
- Ability to remain calm and make decisions during critical outages
Preferred
- Bachelor’s degree or equivalent experience
- ITIL certification
- PowerShell, Python or Bash experience
- Experience in a large, high-availability enterprise environment