Skip to content
← Back to job listings

Senior Site Reliability Engineer

Illumio · Sunnyvale, United States

External listingfull-time3 months ago

About The Role

Join our team as a Senior Site Reliability Engineer (SRE) and play a key role in ensuring the reliability, scalability, and performance of our cloud-based systems and applications. You will monitor system performance, lead incident response efforts, implement security best practices, and drive continuous improvement initiatives. The ideal candidate will have hands-on experience in supporting and managing AWS and Azure infrastructure, a passion for automation, and a track record of driving reliability and performance in cloud-based environments.

  • Monitor system performance, application health, and infrastructure metrics, and implement proactive measures to optimize performance and availability.
  • Lead incident response and resolution efforts, conducting root cause analysis, implementing corrective actions, and documenting post-incident reviews.
  • Drive continuous improvement initiatives to enhance reliability, scalability, and efficiency of infrastructure and services, leveraging automation and emerging technologies.
  • The ideal candidate will have hands-on experience in supporting, and managing AWS and Azure infrastructure, along with a passion for automation, continuous improvement, and collaboration with cross-functional teams
  • If you are passionate about AWS and/or Azure cloud platform and have a track record of driving reliability, scalability, and performance in cloud-based environments, we'd love to hear from you. Apply now to be a part of our talented team!
  • Hands-on experience in designing, deploying, and managing AWS and/or Azure infrastructure, including compute, storage, networking, and security services
  • Bachelor’s degree in computer science, Engineering, or related field; or equivalent work experience
  • Strong understanding of CI/CD principles and experience with tools such as Azure DevOps, Jenkins, or GitLab CI/CD
  • Proficiency in scripting and programming languages such as PowerShell, Python, or Go for automation and infrastructure management tasks
  • Excellent analytical, problem-solving, and communication skills, with the ability to collaborate effectively with cross-functional teams
  • 5+ years of experience working as a Site Reliability Engineer (SRE) or similar role, with a focus on AWS and/or Azure cloud platform
  • AWS or Azure certifications such as AWS/Azure Solutions Architect, Azure DevOps Engineer, or Azure Security Engineer are preferred
  • Experience with containerization technologies (e.g., Docker, Kubernetes) and microservices architecture in AWS and Azure environments is a plus

This is an external listing. JobSpring does not represent or verify the employer. Report this listing