What are the responsibilities and job description for the Senior Site Reliability Engineer position at Wells Fargo?
Title: Senior Site Reliability Engineer
Location: Des Moines, IA
Alternate Location: Minneapolis, MN / Irving, TX
Duration: 18 months
Work Engagement: W2
Work Schedule: Hybrid 3 days in office / 2 days remote
Benefits on offer for this contract position: Health Insurance, Life insurance, 401K and Voluntary Benefits
Summary:
We are seeking a Senior Site Reliability Engineer (SRE) to join a growing team responsible for the reliability, observability, automation, and operational excellence of critical enterprise applications and AI-powered platforms. This role is ideal for an experienced SRE who thrives in complex production environments and is passionate about driving platform stability, automation, self-healing capabilities, and continuous improvement.
The successful candidate will bring deep expertise in production support, incident management, observability, and automation while helping the team evolve its reliability engineering practices. Experience supporting AI/ML and LLM-based systems, including emerging Agentic AI solutions, is highly desired.
Experience within banking, financial services, or other highly regulated industries is strongly preferred.
Key Responsibilities:
Lead and support large-scale production environments using SRE and ITIL best practices.
Manage incident, problem, and change management processes to ensure service reliability and operational excellence.
Monitor, troubleshoot, and optimize business-critical applications and platforms.
Design and implement observability solutions to improve system visibility and performance.
Automate operational processes and develop self-healing solutions to reduce manual intervention.
Support application deployments and continuous delivery pipelines.
Drive root cause analysis and remediation efforts for complex production issues.
Partner with development, infrastructure, and platform engineering teams to improve system resilience and scalability.
Support AI/ML and LLM-based applications in production environments while ensuring reliability, performance, and operational governance.
Contribute to platform modernization and operational maturity initiatives.
Required Qualifications:
Applicants must be authorized to work for ANY employer in the U.S. This position is not eligible for visa sponsorship.
8+ years of experience in Site Reliability Engineering, Production Support, Platform Engineering, or related disciplines.
Proven experience leading large-scale production support environments utilizing ITIL practices.
Experience supporting AI/ML, Large Language Model (LLM), and Agentic AI platforms.
Deep understanding of reliability engineering, observability, incident response, operational risk management, and resilient system design.
Strong experience across on-premises, hybrid, and public cloud environments.
Hands-on experience with CI/CD tools including:
Jenkins
Artifactory
UDeploy
Terraform
Strong automation experience using:
Ansible
Experience with observability and monitoring platforms such as:
AppDynamics
Splunk
Grafana
BigPanda
Application Insights
Experience supporting large-scale:
Java applications
.NET applications
Oracle databases
Microsoft SQL Server databases
Strong Unix/Linux administration and troubleshooting skills, including L2/L3 production support.
Experience performing incident management, change management, problem management, and root cause analysis.