Demo

Journey-Centric Lead Site Reliability Engineer

Ideate Technologies LLC
Phoenix, AZ Contractor
POSTED ON 8/27/2026
AVAILABLE BEFORE 9/25/2026

We are seeking a highly experienced Journey-Centric Lead Site Reliability Engineer (SRE) with Banking and Financial Services (BFS) domain experience to drive end-to-end reliability, observability, automation, and operational excellence across critical customer and business journeys. This role will lead the design and implementation of modern SRE practices, unified observability platforms, self-healing capabilities, AI-driven operations, and workflow automation to ensure highly resilient, scalable, and intelligent digital services.

The ideal candidate combines deep SRE expertise with strong platform engineering, cloud operations, networking, observability, and AI Ops experience, with a particular focus on AWS AgentCore-powered operational intelligence and autonomous operations.

 

Key Responsibilities

SRE & Reliability Engineering

·        Define and implement enterprise-scale SRE best practices across critical applications and digital journeys.

·        Establish reliability frameworks, operational standards, and governance models.

·        Drive proactive reliability engineering initiatives to improve system availability, resilience, and performance.

·        Lead incident management, postmortem analysis, root cause investigations, and reliability reviews.

Unified Observability & Monitoring

·        Design and implement a unified observability strategy encompassing metrics, logs, traces, events, and user experience telemetry.

·        Build comprehensive observability dashboards for business and technology stakeholders.

·        Implement distributed tracing and end-to-end monitoring across complex microservices ecosystems.

·        Define observability standards and instrumentation frameworks across engineering teams.

Log Analytics & Trace Correlation

·        Enable unified logs, metrics, and trace correlation capabilities for rapid issue detection and troubleshooting.

·        Deploy intelligent correlation engines for root cause analysis.

·        Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through observability-driven insights.

·        Establish service dependency mapping and journey-centric operational visibility.

Self-Healing & Autonomous Operations

·        Design and implement self-healing capabilities using event-driven automation and AI-assisted remediation.

·        Develop automated recovery processes for common failure scenarios.

·        Create autonomous operational workflows that minimize manual intervention.

·        Integrate predictive alerting and automated response mechanisms.

Automation & Workflow Engineering

·        Build scalable operational automation frameworks.

·        Develop infrastructure, application, observability, and operational workflows using Infrastructure as Code (IaC), Monitoring as Code (MaC), and Observability as Code (OaC).

·        Automate deployments, monitoring, remediation, and operational runbooks.

·        Reduce operational toil through intelligent engineering solutions.

SLO, SLA & Error Budget Management

·        Define and govern measurable Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Error Budgets.

·        Partner with engineering and business teams to align reliability targets with customer expectations.

·        Establish service maturity metrics and reliability scorecards.

·        Drive data-driven operational decision-making through reliability KPIs.

Network & Platform Reliability

·        Apply deep understanding of:

o   Understanding of network layer to troubleshoot critical bandwidth/latency issues

o   No need to pass all these:

o   TCP/IP

o   DNS

o   Load Balancing

o   CDN

o   API Gateway architectures

o   Service Mesh technologies

o   VPC and cloud networking

·        Troubleshoot complex network performance and availability issues.

·        Ensure end-to-end reliability across cloud and hybrid environments.

AI Ops & AWS AgentCore

·        Implement and operationalize AI Ops platforms and autonomous operations capabilities.

·        Leverage AWS AgentCore to build intelligent operational agents for functions like:

o   Incident response

o   Root cause analysis

o   Predictive remediation

o   Capacity forecasting

o   Automated operational workflows

·        Drive adoption of GenAI-powered operational intelligence across the enterprise.

·        Integrate AI-assisted observability, automation, and service management solutions.

Hourly Wage Estimation for Journey-Centric Lead Site Reliability Engineer in Phoenix, AZ
$37.00 to $44.00
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Journey-Centric Lead Site Reliability Engineer?

Sign up to receive alerts about other jobs on the Journey-Centric Lead Site Reliability Engineer career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$76,670 - $90,826
Income Estimation: 
$91,609 - $118,978
Income Estimation: 
$92,877 - $110,401
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at Ideate Technologies LLC

  • Ideate Technologies LLC Jersey, NJ
  • Job Description Must Have Technical/Functional Skills · 10 years of experience COBOL, JCL, DB2, VSAM, CICS, ChangeMan, File Aid/File Manager, DFSORT, Syncs... more
  • Just Posted


Not the job you're looking for? Here are some other Journey-Centric Lead Site Reliability Engineer jobs in the Phoenix, AZ area that may be a better fit.

  • Intraedge Phoenix, AZ
  • Senior / Lead Site Reliability Engineer, Principal Engineer Long term contract Phoenix, AZ (Hybrid-3 days onsite) Direct client- Immediate client interview... more
  • 5 Days Ago

  • Nexthink Phoenix, AZ
  • Nexthink is the leader in digital employee experience management software. The company provides IT leaders with unprecedented insight allowing them to see,... more
  • 2 Days Ago

AI Assistant is available now!

Feel free to start your new journey!