Skip to content
← Back to job listings

Senior Site Reliability Engineer (CCIP)

Chainlink Labs · United States

External listingfull-timeabout 1 month ago

About The Role

Join Chainlink, a leading provider of decentralized oracle networks. As a Senior Site Reliability Engineer on the CCIP Platform team, you will ensure the reliability, scalability, and operational excellence of the systems powering Chainlink's Cross-Chain Interoperability Protocol (CCIP). You will influence reliability practices across the platform, establish operational standards, and drive adoption of meaningful SLOs, SLIs, and error budgets. This role requires deep expertise in Site Reliability Engineering, production engineering, or a similar role operating large-scale distributed systems.

  • Assurer la fiabilité, l'évolutivité et l'excellence opérationnelle des systèmes alimentant le protocole d'interopérabilité inter-chaînes de Chainlink (CCIP).
  • Influencer les pratiques de fiabilité à travers la plateforme et aider à établir des normes opérationnelles qui évoluent avec l'entreprise.
  • Améliorer la sécurité des déploiements et augmenter la vitesse de livraison en faisant progresser les pratiques d'ingénierie de production.
  • Applied OpenTelemetry to improve observability across distributed systems
  • Deep expertise defining, implementing, and driving adoption of SLOs, SLIs, and error budgets across engineering organizations
  • Experience improving the reliability, scalability, and operability of production infrastructure
  • Demonstrated experience in Site Reliability Engineering, Production Engineering, or a similar role operating large-scale distributed systems
  • Built and operated production Kubernetes environments supporting critical services
  • Demonstrated technical leadership influencing reliability practices across engineering teams
  • Experience performing capacity planning and performance tuning for high-throughput distributed services
  • Previous experience working on Web3 infrastructure or within a crypto-native engineering organization
  • Applied chaos engineering or fault-injection techniques to improve production resilience
  • Partnered with software engineering teams to conduct production-readiness reviews before service launches
  • Experience leading on-call operations, including defining rotations, escalation policies, and improving alert quality

This is an external listing. JobSpring does not represent or verify the employer. Report this listing