Skip to content
← Back to job listings

Principal Site Reliability Engineer (AWS, Azure, Terraform, Kubernetes)

Fourth · London, United Kingdom

External listingfull-time11 days ago

About The Role

Join Fourth, a rapidly growing company in the field of automated, highly reliable, and zero-downtime infrastructure pipelines. As a Principal Site Reliability Engineer, you will work with architecture, product owners, and development teams to define the practical needs and requirements for products being migrated to the cloud. You will design and implement scalable, highly available, and resilient infrastructure solutions, innovate and enhance tooling for monitoring and automation, and actively participate in incident management. You will also assist in the career development of colleagues and encourage best practices.

  • Participer activement à la gestion des incidents, en conduisant une récupération rapide, une analyse des causes profondes et une amélioration continue.
  • Travailler en collaboration avec l'architecture, les propriétaires de produits et les équipes de développement pour définir les besoins pratiques et les exigences des produits migrés vers le cloud.
  • Concevoir et mettre en œuvre des solutions d'infrastructure évolutives, hautement disponibles et résilientes, conformément aux normes convenues.
  • You will exhibit both strategic and tactical sensibilities. You are adept at mapping the bigger picture into smaller, yet valuable and achievable chunks. You have excellent written and verbal communication skills, allowing you to work effectively with our worldwide development teams and SRE community to select the right patterns and practices. You understand the importance of standardisation of technology and practices and have experience of implementing these in a consistent and low-maintenance fashion. You believe useful and relevant documentation is an essential output of your work
  • Relevant and recent experience with our main tech stack: Terraform, Configuration Management (Chef ideally, but will consider Ansible or Puppet), Kubernetes (cloud based), Docker (Kubernetes or AWS ECS Fargate)
  • Extensive cloud experience (ideally Azure and AWS)
  • Outstanding communication skills, with the ability to convey complex technical concepts to non-technical stakeholders
  • Bachelor’s degree in computer science, Engineering, or a related field
  • Programming, scripting skills and OOP principles in general
  • Strong analytical and problem-solving skills
  • Proven track record of designing, implementing, and managing large-scale, highly available, and scalable infrastructure
  • Experience of working in Agile or Kanban
  • Ability to work in a fast-paced, dynamic environment, managing multiple priorities
  • Demonstrable experience in Site Reliability Engineering, Platform Engineering, DevOps, or a similar role, with experience mentoring others
  • A natural and positive team player

This is an external listing. JobSpring does not represent or verify the employer. Report this listing