← Back to job listings
SA
Principal Site Reliability Engineer
Saviynt · Vancouver, Canada
About The Role
Join Saviynt, a leading provider of identity governance solutions. As a Principal Site Reliability Engineer, you will work on a mission-critical SaaS platform used by global enterprises. You will solve complex reliability challenges at scale, influence architecture and engineering culture, and have opportunities for competitive compensation, benefits, and growth. Your responsibilities will include designing and building core platform components, managing Kubernetes platforms, developing internal tools, and collaborating with product development teams.
- Concevoir, construire et maintenir des services d'infrastructure partagés et des plateformes sur lesquelles les équipes de produits et d'applications dépendront.
- Architecturer, mettre en œuvre et gérer des plateformes Kubernetes hautement disponibles et évolutives en tant que service pour les consommateurs internes.
- Développer et maintenir des pipelines CI/CD robustes (par exemple, GitLab CI et ArgoCD) en tant que service, fournissant des workflows de déploiement standardisés et automatisés.
- This role requires compliance with Saviynt’s information security and privacy policies, including annual security training
- Proficiency in establishing and utilizing comprehensive Observability and Monitoring platforms (e.g., Prometheus, Grafana, ELK stack, Datadog) for shared infrastructure
- Strong programming skills in Go (Golang) and Python, with experience building robust, maintainable backend services and automation
- 9+ years of experience in an Infrastructure Development, Platform Engineering, or Site Reliability Engineering role, with a strong focus on building tools and services for other engineers
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience or equivalent military experience required
- Strong experience with RESTful API design principles and building well-documented, consumable APIs
- Hands-on experience with Relational Databases (e.g., MySQL, PostgresSQL), ideally in managing them as a service
- Knowledge of Service Mesh concepts and practical experience with solutions like Istio in a platform context
- Familiarity with Multi-Region Cloud Environments and strategies for building globally distributed and highly available platform
- A strong customer-centric mindset, treating internal development teams as your primary customers
- Extensive hands-on experience with at least one major Cloud Provider (AWS, GCP, or Azure); multi-cloud experience is a strong plus, especially in building abstractions over them
- Proven experience designing and implementing Event-Driven Architecture and message queuing systems (e.g., Kafka, RMQ, NATS) as shared services
- Demonstrable experience designing and operating Distributed Systems, with an understanding of patterns for creating reliable, shared components
- Deep expertise with Kubernetes in production environments, particularly in providing it as a platform(i.e single tenant and multi-tenant deployment architectures)
- Excellent communication skills and the ability to clearly articulate complex technical concepts to both technical and non-technical audiences
- Solid understanding and practical experience with CI/CD pipeline tools (especially GitLab CI) and experience establishing automated delivery processes for other teams
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring