What are the responsibilities and job description for the Network Operations Center Manager position at Optomi?
Optomi is seeking a NOC Manager to join their client in downtown Fort Worth. This role is direct hire and required onsite in Fort Worth, TX 5 days a week.
**Relocation assistance available**
This role drives operational excellence by ensuring effective observability practices, proactive detection of service-impacting events, rapid incident response, and continuous service improvement.
Key Responsibilities:
- Lead, mentor, and develop a team of level one Observability Analysts responsible for 24x7 monitoring, event triage, incident response, and operational escalation.
- Establish and execute operational strategies for observability, monitoring, alert management, incident management, and service reliability.
- Own operational KPIs including MTTR, alert quality, incident trends, escalation effectiveness, service availability, and customer impact metrics.
- Collaborate with Infrastructure, Engineering, Network Operations, Security, and support teams to improve system reliability and operational performance.
- Define monitoring standards, alert governance policies, dashboard strategies, and operational best practices.
- Oversee identification and remediation of monitoring gaps, alert fatigue, false positives, missed detections, and escalation inefficiencies.
- Lead major incident management activities, post-incident reviews, operational readiness assessments, and root cause analysis initiatives.
- Drive operational automation, event correlation, and tooling enhancements to improve efficiency and reduce manual effort.
- Manage operational reporting, executive dashboards, service reviews, and reliability improvement initiatives.
- Ensure runbooks, operational procedures, escalation paths, knowledge articles, and incident documentation remain current and effective.
- Serve as an escalation point during critical incidents and provide leadership throughout high-severity operational events.
- Manage staffing, workforce planning, scheduling, performance management, and career development for IOC and Observability personnel.
- Foster a culture of operational excellence, accountability, collaboration, and continuous improvement.
Required Qualifications:
- 8 years of experience in NOC, Observability, Infrastructure Operations, Production Operations, Site Reliability Engineering, or related technical operations environments.
- 3 years of leadership experience managing technical operations, monitoring, observability, or support teams in a 24x7 environment.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent professional experience.
- Strong understanding of incident management, observability concepts, service reliability, operational risk management, and continuous improvement methodologies.
- Experience supporting large-scale enterprise infrastructure, cloud environments, networking, data center operations, or distributed systems.
- Proven success developing monitoring programs, alerting strategies, operational dashboards, and service health reporting.
- Experience driving improvements in operational KPIs including MTTD, MTTR, service availability, and incident reduction.
- Strong knowledge of observability and monitoring platforms such as Splunk, Datadog, Grafana, Prometheus, New Relic, LogicMonitor, SolarWinds, or similar technologies.
- Experience working with ITSM platforms such as ServiceNow and established incident management processes.
- Working knowledge of Linux, Windows, networking, cloud infrastructure, automation, and infrastructure operations.
- Experience leading root cause analysis efforts, operational reviews, and cross-functional incident response activities.
- Strong leadership, communication, stakeholder management, and team development skills.
Salary : $125,000 - $150,000