What are the responsibilities and job description for the Sr AIOPS Engineer - Incident Manager position at Reliable Software Resources?
Position: Sr AIOPS Engineer - Incident Manager / SRE
Location: Hybrid Onsite (Multiple Locations)
Contract: 6 months
Qualifications
- Bachelor's degree in Computer Science, Information Engineering, or related field (or equivalent experience).
- 8 years Of experience in IT Operations, Site Reliability Engineering, Infrastructure Operations, Network Operations, or Production Support environments.
- 5 years of experience leading incident management, transformation, or reliability engineering
- Strong experience with:
- Site Reliability Engineering (SRE)
- IT Service Management (ITSM)
- ITIL Framework
- Incident, Problem, Change, and Event Management
- Network Operations Center (NOC)
- Infrastructure Operations
- Service Desk
- Application Support
- Cloud Platforms (AWS Azure, or Google Cloud Platform)
- DevOps Practices and Toolchains
- Hands on experience with Dynatrace, monitoring platforms, and observability solutions.
- Experience using ServiceNow for ticketing, workflow automation, and service management.
- Strong understanding Of infrastructure, networking, cloud architecture, and enterprise application ecosystems
- Proven experience conducting root cause analysis and implementing preventive controls.
- Experience leading enterprise AIOps implementations.
- Experience building AI-Powered operational agents and intelligent automation solutions.