← Back to job listings
RE
Senior Site Reliability Engineer (Remote Build)
Remote · Canada
About The Role
Join Remote, a leading global HR and payroll platform, as a Senior Site Reliability Engineer. In this remote position, you will be responsible for the operational excellence and infrastructure strategy of Remote Build's platform. You will work closely with the Remote Build leadership group, product managers, engineers, and customer success to ensure scalability and reliability from day one. Your tasks will include designing and maintaining infrastructure-as-code patterns, building monitoring and alerting systems, ensuring security and compliance, optimizing performance and costs, and improving developer experience.
- Design, implement, and maintain infrastructure-as-code patterns using Terraform and Kubernetes that support both standard connectors and custom builds.
- Build and maintain comprehensive monitoring, logging, and alerting systems. Lead incident response efforts, conduct post-mortems, and drive continuous improvement in system reliability.
- Work with our Security team to embed security into every layer of Build infrastructure. Ensure we meet compliance requirements across 100+ jurisdictions without creating friction for developers or customers.
- Scripting and systems knowledge: strong bash scripting. Comfortable debugging system-level issues, reading logs, and understanding Linux kernel basics
- Great communication: you explain complex infrastructure decisions clearly to both engineers and non-technical stakeholders. You write clear runbooks and documentation
- Kubernetes and AWS: deep, hands-on experience running Kubernetes in production. Solid AWS fundamentals across compute, networking, storage, and managed services
- Infrastructure-as-code: Proficiency with Terraform or similar IaC tools. You write code to define infrastructure; you don't click buttons in the console
- Senior-level SRE experience: demonstrated experience in a Site Reliability Engineering, DevOps Engineering, or SysOps role. You have stood up and operated production systems at scale
- CI/CD and deployment automation: real experience setting up and operating GitLab, GitHub Actions, Jenkins, or similar. You understand deployment strategies, rollback mechanisms, and safety nets
- Experience working with or scaling multi-tenant platforms
- Observability stack depth (Datadog, Prometheus, ELK, Grafana, or similar)
- Container registry and artifact management (ECR, Docker Hub, etc.)
- Experience in consultancy settings
- Experience with 1+ backend programming language (Elixir, Python, Go, Java, Node.js, etc.)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring