Senior Site Reliability Engineer
HiveWatch · El Segundo, United States
About The Role
Join HiveWatch as a Senior Staff Site Reliability Engineer. In this role, you will architect and maintain mission-critical edge infrastructure, ensuring exceptional performance, reliability, and observability across our distributed environment. You will provide technical leadership to our growing engineering team and report directly to our VP of Engineering. Responsibilities include owning the reliability of mission-critical systems, debugging complex production issues, building automation and tooling, and maintaining CI/CD pipelines. You will also provide technical leadership and mentorship to foster engineering excellence and reliability culture.
- Architect and maintain mission-critical edge infrastructure that connects the SaaS platform to customer systems, ensuring exceptional performance, reliability, and observability.
- Own the reliability of mission-critical systems, including production monitoring, alerting, and capacity planning, and debug and resolve complex production issues.
- Provide technical leadership and mentorship to foster engineering excellence and reliability culture, and participate in establishing SRE practices and culture.
- Hands-on experience with relational databases and SQL performance optimization
- 5+ years of SRE, DevOps, or production operations experience
- Experience with monitoring and observability tools (Prometheus, Grafana, DataDog, or equivalent)
- Proficiency in at least one object oriented programming language in our tech stack (Java, Kotlin, Python)
- 7+ years of software engineering experience with strong coding skills in production environments
- Experience with Infrastructure as Code (Terraform, CloudFormation, or similar)
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience
- Expertise with cloud platforms (AWS preferred) and containerized applications (Docker, Kubernetes)
- Strong debugging skills across distributed systems and microservices architectures
- Expertise in AWS architecture and services
- Experience in physical security, IoT, or edge computing environments
- Expertise with advanced AWS services (Kinesis, Lambda, EKS, RDS)
- Experience with Terraform and Terragrunt specifically
- Background in high-availability, multi-tenant SaaS environments
- Experience establishing SRE practices and culture from the ground up
- Knowledge of security best practices and compliance requirements
- Experience mentoring and developing junior engineers
- Track record of leading incident response and post-mortem processes
- Experience with edge computing and distributed system architectures
- Previous experience in a startup or high-growth environment (50-200 employees)
- Experience with our tech stack: Kotlin, Rust,TypeScript, Python
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring