Skip to content
← Back to job listings

Staff Site Reliability Engineer

soundhound · Toronto, CAN

Software DevelopmentRemoteExternal listingfull-time3 months ago

About The Role

### The Opportunity

We’re looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability, and performance of our infrastructure, with a deep focus on Google Cloud Platform (GCP). You will architect and maintain high-availability systems, automate operational tasks, and ensure our services can handle the demands of millions of voice AI interactions.

### What You'll Do

  • Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
  • Architect and automate CI/CD pipelines to ensure rapid, reliable deployments.
  • Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues.
  • Partner with engineering teams to optimize performance, cost, and reliability of backend services.
  • Drive incident response, post-mortem analysis, and long-term remediation efforts.
  • Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
  • Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
  • Lead department wide compliance (PCI, SOC) initiatives.

### What You'll Bring

  • 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
  • Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
  • Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Deep experience with Kubernetes, container orchestration, and service mesh architectures.
  • Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
  • Experience designing and managing high-throughput, distributed systems.
  • Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs.
  • Excellent communication skills and a demonstrated ability to mentor engineers.

###
### Preferred Qualifications

  • Experience working in a high-velocity, customer-focused environment.
  • Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript).
  • Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
  • Experience implementing security and compliance best practices in the cloud.

### Workplace & Compensation

This role is available throughout Canada.

Compensation includes salary, equity, comprehensive healthcare, paid time off, and other benefits. Our recruiting team will provide a specific salary range based on location and years of experience.

#LI-MQ1 #LI-REMOTE

This is an external listing. JobSpring does not represent or verify the employer. Report this listing