What are the responsibilities and job description for the Senior Software Engineer - Application Reliability position at MetLife, Inc?
Description and Requirements
When you join MetLife’s Global Technology team, you’ll be part of a forward-thinking group dedicated to shaping the future of digital solutions for customers worldwide. You’ll develop, maintain and support technology applications and delivery, leveraging AI, automation, and contemporary ways of working to enhance experiences and drive business outcomes. Your work will simplify complex processes, improve tech resiliency, and ensure high-performing, seamless solutions that power life’s most important moments. In this dynamic environment, you’ll collaborate with talented peers across teams and functions, expanding your skills in impactful ways. Ready to push boundaries and set new industry standards? Join us and help drive the future of technology forward.
At MetLife, we seek to make a meaningful impact on the lives of our customers and our communities. Global Technology & Operations group (GTO) is a diverse team of Agile practitioners comprised of engineers, developers, and technology leaders with the freedom to create innovative solutions that address core business challenges. This role is for a software engineer who will design, build, enhance, and support software applications and platforms for US Digital applications. You’ll work in a collaborative, agile environment and be hands-on with full stack engineering, API and integration development, cloud and DevOps practices, production reliability, observability, and AI-enabled engineering tools.
- Design, develop, test, deploy, and maintain full stack software applications across UI, API, data, and integration layers.
- Deliver high-quality, secure, scalable, and maintainable code using modern software engineering practices, including code reviews, automated testing, version control, and CI/CD pipelines.
- Analyze business and technical requirements, translate them into engineering solutions, and partner with product and platform teams to deliver application enhancements.
- Troubleshoot, debug, and resolve complex software defects across applications, API, data, batch, authentication, and integration components.
- Apply site reliability engineering practices to improve application availability, performance, resiliency, observability, monitoring, alerting, incident response, automation, and problem management.
- Participate in production support and incident response activities as an engineering owner, including root cause analysis, corrective actions, and preventive improvements.
- Use AI-enabled engineering and AIOps capabilities to accelerate development, identify incident patterns, analyze trends, automate repetitive tasks, support root cause analysis, and improve operational efficiency.
- Contribute to the integration of AI capabilities into applications, including GenAI APIs, LLM-based tools, agents, and intelligent self-service or support features.
- Collaborate with cross-functional teams to improve system design, reduce technical debt, strengthen operational readiness, and deliver reliable customer-facing digital experiences.
- Perform related duties as assigned or requested.
- Bachelor’s degree in computer science, Information Systems, Engineering, or a related field (or equivalent experience), with 7 years of experience building and supporting full-stack applications, APIs, integrations, batch processes, data flows, and customer-facing digital platforms.
- Proficiency in one or more programming languages, including Java, Python, C#, JavaScript/TypeScript, or similar technologies. Strong knowledge of software engineering practices, including object-oriented design, secure coding, automated testing, code reviews, source control, CI/CD, cloud platforms, DevOps methodologies, and/or ITSM tools such as ServiceNow.
- Experience applying Site Reliability Engineering (SRE) principles and capabilities, including observability, monitoring, alerting, performance analysis, incident response, root cause analysis, reliability improvement, and automation to enhance application stability and operational effectiveness.
- Experience with cloud platforms, DevOps practices, container technologies, and application monitoring/observability tools (ex: Splunk, Grafana, Ealstic, Azure DevOps, Docker, Kubernetes, MongoDB, CosmosDB & AppDynamics etc.,)
- Experience with AI-enabled engineering, AI SRE, or AIOps platforms, including building dashboards, agents, automation tools, AI-assisted development practices (LLM), prompt engineering.
- Hands-on experience with ServiceNow ticket management, building SRE level logs & operational dashboards, and production support reporting.
- Demonstrated ability to engineer solutions for complex applications, infrastructure, API, batch, data, authentication, and integration challenges in highly available production environments.
- Strong knowledge of production reliability practices, including incident management, triage facilitation, root cause analysis, post-incident reviews, and corrective action tracking.
- Ability to analyze technical and operational data from multiple sources, identify patterns, draw conclusions, and recommend engineering improvements.
- Strong understanding of service level objectives, service level agreements, customer-facing metrics, and production health indicators.
The expected salary range for this position is $90,000 - $120,000. This role may also be eligible for annual short-term incentive compensation and stock-based long-term incentives. All incentives and benefits are subject to the applicable plan terms.
Salary : $90,000 - $120,000