Demo

Principal Site Reliability Engineer (Cloud, Observability & Automation)

DTCC Candidate Experience Site
Jersey, NJ Full Time
POSTED ON 8/17/2026
AVAILABLE BEFORE 10/17/2026

Are you ready to make an impact at DTCC?

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.

Pay and Benefits: 

  • Competitive compensation, including base pay and annual incentive
  • Comprehensive health and life insurance and well-being benefits, based on location
  • Pension / Retirement benefits
  • Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
  • DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).

The Impact You Will Have in This Role

The Enterprise Application Support (EAS) team supports critical applications across the ITP and ECS business lines, ensuring the reliability, scalability, and performance of enterprise platforms.

As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across mission-critical systems. You will lead reliability initiatives, champion observability and automation, drive major incident response, and partner with engineering, infrastructure, and security teams to build resilient, highly available applications.

This is a hands-on technical leadership role focused on improving system performance, reducing operational risk, accelerating recovery, and advancing SRE best practices through modern cloud, observability, automation, and AI-powered technologies.

Your Primary Responsibilities:

  • Drive reliability, scalability, resiliency, and operational excellence across critical enterprise applications. 
  • Design and implement observability solutions using Splunk, Grafana, Dynatrace, ITSI, and related monitoring platforms. 
  • Define and manage SLIs, SLOs, dashboards, alerts, and operational KPIs. 
  • Lead major incident response, root cause analysis, and continuous service improvement initiatives. 
  • Build automation, self-healing capabilities, and AI-assisted operational solutions using Python, Java, Amazon Q, Kiro, and related technologies. 
  • Partner with development, infrastructure, cloud, security, and application teams to embed SRE best practices throughout the software development lifecycle. 
  • Drive operational readiness, capacity planning, performance optimization, disaster recovery, and resiliency initiatives. 
  • Identify operational risks and deliver strategic reliability improvements across the technology ecosystem. 
  • Collaborate with technical and business stakeholders to improve service reliability and operational outcomes.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience. 
  • 8 years of experience in Site Reliability Engineering, Production Engineering, DevOps, Application Support Engineering, or related disciplines.

Talent Needed for Success

  • Strong hands-on experience with AWS and cloud-native architectures. 
  • Proficiency in Python, Java, Go, or similar programming languages. 
  • Strong Linux/Unix systems administration and troubleshooting experience. 
  • Expertise in observability and monitoring platforms including Splunk, Grafana, Dynatrace, and ITSI. 
  • Experience leading major incident management and root cause investigations in complex production environments. 
  • Strong understanding of distributed systems, resiliency engineering, performance tuning, automation, and operational excellence. 
  • Excellent communication and stakeholder management skills with the ability to influence technical and business partners.

Preferred Qualifications

  • Experience with AI-assisted engineering tools such as Amazon Q, Kiro, or similar technologies. 
  • Experience designing and measuring SLOs, SLIs, and operational KPIs. 
  • Experience supporting large-scale enterprise applications in financial services or other highly regulated environments.

The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations. We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

Salary.com Estimation for Principal Site Reliability Engineer (Cloud, Observability & Automation) in Jersey, NJ
$158,791 to $188,043
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Principal Site Reliability Engineer (Cloud, Observability & Automation)?

Sign up to receive alerts about other jobs on the Principal Site Reliability Engineer (Cloud, Observability & Automation) career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$76,670 - $90,826
Income Estimation: 
$91,609 - $118,978
Income Estimation: 
$92,877 - $110,401
Income Estimation: 
$140,435 - $166,410
Income Estimation: 
$151,875 - $212,356
Income Estimation: 
$169,957 - $202,398
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Job openings at DTCC Candidate Experience Site

  • DTCC Candidate Experience Site Coppell, TX
  • Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment... more
  • 1 Day Ago

  • DTCC Candidate Experience Site Jersey, NJ
  • Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment... more
  • 2 Days Ago

  • DTCC Candidate Experience Site Jersey, NJ
  • Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment... more
  • 2 Days Ago

  • DTCC Candidate Experience Site Jersey, NJ
  • Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment... more
  • 2 Days Ago


Not the job you're looking for? Here are some other Principal Site Reliability Engineer (Cloud, Observability & Automation) jobs in the Jersey, NJ area that may be a better fit.

  • IPolarity LLC Whippany, NJ
  • Job title: Site Reliability Engineer (SRE) Bill rate: $52/hr W2 Client address: 2900 W Plano Pkwy Plano, TX 75075 - Role is hybrid (3 days/wk) Years of exp... more
  • 26 Days Ago

  • Bank of America Jersey, NJ
  • Job Description: At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do thi... more
  • 2 Days Ago

AI Assistant is available now!

Feel free to start your new journey!