Skip to content
← Back to job listings

Site Reliability Engineer

Forward Networks · Santa Clara, CA, United States

External listingfull-time14 days ago

About The Role

Join Forward, a fast-growing early-stage company, as a Site Reliability Engineer (SRE). In this role, you will build the reliability engineering function from the ground up, defining SLOs, SLIs, error budgets, and frameworks for the engineering organization. You will drive the reliability and operational excellence of our SaaS platform, build and maintain observability infrastructure, lead incident response, and partner with engineering teams to embed reliability thinking into the software development lifecycle. This is a foundational hire with a path to leadership.

  • Definir y liderar las prácticas de ingeniería de confiabilidad desde cero, incluyendo SLOs, SLIs, presupuestos de error y los marcos que la organización de ingeniería utilizará.
  • Construir y mantener la infraestructura de observabilidad, incluyendo registro, métricas, trazado y alertas, para que el equipo siempre sepa lo que está sucediendo antes que los clientes.
  • Colaborar con equipos de ingeniería para incorporar el pensamiento de confiabilidad en el ciclo de vida del desarrollo de software, incluyendo planificación de capacidad, pruebas de carga, ingeniería de caos y revisiones de preparación para producción.
  • If you thrive in environments where you're handed a problem rather than a playbook this role is for you
  • Hands-on experience with Kubernetes and container orchestration in production environments
  • Ability to communicate clearly with both engineering teams and non-technical stakeholders — you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them
  • Experience with cloud platforms — AWS, GCP, or Azure — including infrastructure as code (Terraform, Ansible, or equivalent)
  • Strong scripting and automation skills in Python, Bash, or similar
  • Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements
  • Deep proficiency with observability tooling — Prometheus, Grafana, Datadog, Splunk, or similar
  • 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
  • Proven experience building or significantly maturing an SRE function — not just operating within one someone else built
  • Strong fundamentals in networking — TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus
  • Experience supporting enterprise or federal government customers with high availability requirements
  • Experience in a foundational or early SRE hire capacity at a growth stage company

This is an external listing. JobSpring does not represent or verify the employer. Report this listing