Skip to content
← Back to job listings

Senior Site Reliability Engineer

Kody · Palo Alto, United States

External listingfull-timeabout 1 month ago

About The Role

Join Kody, a global payment platform, as a Senior Site Reliability Engineer. In this role, you will be responsible for ensuring the reliability, availability, scalability, and operational excellence of our payment processing systems. You will own production observability, incident response, service-level management, and cloud infrastructure reliability. Additionally, you will participate in a follow-the-sun production on-call rotation, define and maintain SLOs and SLIs, drive reliability improvements, and partner with engineering teams to improve resilience and security. This position offers remote work flexibility, a competitive salary, and various benefits.

  • Assurer la fiabilité, la disponibilité, l'évolutivité et l'excellence opérationnelle de la plateforme de paiement mondiale de l'entreprise.
  • Participer à la rotation d'appel de production et agir en tant que principal répondant aux incidents, en diagnostiquant, en triant et en coordonnant la résolution des incidents de production.
  • Définir et maintenir les SLO, SLI, budgets d'erreur, normes d'alerte et processus de préparation opérationnelle, et conduire les améliorations de fiabilité par l'automatisation et l'observabilité.
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission-critical production systems
  • Strong hands-on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms
  • Strong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellence
  • Proven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirements
  • Deep understanding of distributed systems, cloud-native architectures, high availability, disaster recovery, capacity planning, and performance optimization
  • Possesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision-making under pressure
  • Applies a systematic and methodical approach to troubleshooting, root-cause analysis, and incident resolution in complex distributed environments
  • Demonstrates strong ownership and accountability, taking end-to-end responsibility for service reliability and customer impact
  • Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements
  • Proven ability to lead cross-functional incident response efforts, coordinate stakeholders, and communicate effectively during high-severity production events
  • Champions a culture of operational readiness, continuous learning, post-incident improvement, and blameless accountability
  • Demonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by design

This is an external listing. JobSpring does not represent or verify the employer. Report this listing