Demo

Senior Site Reliability Engineer- Palo Alto, the US

Kody
Palo Alto, CA Remote Full Time
POSTED ON 8/23/2026
AVAILABLE BEFORE 9/22/2026

Senior Site Reliability Engineer (Payments Infrastructure)
Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America.

Responsibilities

  • Participate in a follow-the-sun production on-call rotation as a primary incident responder.
  • Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
  • Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes.
  • Drive reliability improvements through automation, observability, capacity planning, performance optimization, and post-incident reviews.
  • Partner with engineering teams to improve resilience, security, and operational maturity in PCI-DSS-regulated environments.
  • Lead incident management during SEV1/SEV2 events and improve response effectiveness and MTTR.
  • Cross-Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignment.
  • 5 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems.
  • Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms.
  • Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization.
  • Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements.
  • Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence.



Leadership & Operational Excellence

  • Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer impact.
  • Possesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision-making under pressure.
  • Applies a systematic and methodical approach to troubleshooting, root-cause analysis, and incident resolution in complex distributed environments.
  • Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.
  • Continuously drives engineering excellence through iterative improvement, automation, standardization, and elimination of operational toil.
  • Proven ability to lead cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events.
  • Champions a culture of operational readiness, continuous learning, post-incident improvement, and blameless accountability.
  • Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by design.

  • Competitive packages aligned with California market standards
  • Lead a dynamic and innovative team in a very rapidly growing company
  • Collaborative, inclusive environment where your contributions are recognized and valued

Salary.com Estimation for Senior Site Reliability Engineer- Palo Alto, the US in Palo Alto, CA
$141,780 to $166,306
If your compensation planning software is too rigid to deploy winning incentive strategies, it’s time to find an adaptable solution. Compensation Planning
Enhance your organization's compensation strategy with salary data sets that HR and team managers can use to pay your staff right. Surveys & Data Sets

What is the career path for a Senior Site Reliability Engineer- Palo Alto, the US?

Sign up to receive alerts about other jobs on the Senior Site Reliability Engineer- Palo Alto, the US career path by checking the boxes next to the positions that interest you.
Income Estimation: 
$114,618 - $136,401
Income Estimation: 
$144,264 - $191,312
Income Estimation: 
$140,435 - $166,410
Income Estimation: 
$114,618 - $136,401
Income Estimation: 
$144,264 - $191,312
Income Estimation: 
$140,435 - $166,410
Employees: Get a Salary Increase
View Core, Job Family, and Industry Job Skills and Competency Data for more than 15,000 Job Titles Skills Library

Not the job you're looking for? Here are some other Senior Site Reliability Engineer- Palo Alto, the US jobs in the Palo Alto, CA area that may be a better fit.

  • Amada Senior Care Palo Alto Palo Alto, CA
  • Join the Amada Senior Care Team and Make a Difference Every Day! At Amada Senior Care , we believe caregiving is more than a job—it's an opportunity to imp... more
  • 23 Days Ago

  • Candidate Experience site Sunnyvale, CA
  • Join Fortinet, a cybersecurity pioneer with over two decades of excellence, as we continue to shape the future of cybersecurity and redefine the intersecti... more
  • 2 Months Ago

AI Assistant is available now!

Feel free to start your new journey!