Skip to content
← Back to job listings

Senior Site Reliability Engineer

Salesforce · San Francisco, United States

External listingfull-time14 days ago

About The Role

Join Salesforce as a Senior Site Reliability Engineer in San Francisco. In this role, you will be a technical leader driving operational resilience through automation, observability, and AI-powered platforms. You will ensure high availability, performance, and reliability of our services while continuously improving operational efficiency. Collaborate with development teams, lead incident management, and design systems that prevent issues.

  • Conduire la réponse coordonnée aux incidents en tant que Commandant d'incident, en favorisant une récupération rapide et en garantissant des améliorations durables grâce à des post-mortems sans blâme.
  • Appliquer des pratiques d'ingénierie logicielle - automatisation, surveillance, systèmes auto-réparateurs - pour éliminer le travail pénible et améliorer l'efficacité opérationnelle.
  • Collaborer avec les équipes de développement et d'ingénierie dès le début du cycle de vie pour concevoir, construire et exploiter des systèmes fiables par défaut.
  • Track record of mentoring and technically coaching other engineers
  • Ability to work in a 24/7 global operations model, managing multiple priorities under time-sensitive conditions
  • Advanced prompt engineering skills and the ability to write precise, structured prompts and cultivate the system context that makes AI outputs reliable, secure, and production-ready
  • Experience applying AI/ML to operations — including anomaly detection, predictive analysis, LLM-based automation, and prompt engineering to build intelligent operational agents and workflows
  • Experience using AI tools (e.g., Claude Code, GitHub Copilot, Codex, Cursor, etc.) in development workflows
  • Solid background in incident management, including on-call participation, root cause analysis, and postmortem practices
  • Familiarity with large-scale internet service architectures (DNS, HTTP, Load Balancing, caching, etc.)
  • Production experience building and operating observability platforms (Grafana, Prometheus, ELK, Splunk, Datadog, or similar)
  • A demonstrated, genuine AI-first approach to engineering. Using AI to move faster, build fluency across the stack, and contribute well beyond your core specialty
  • Hands-on expertise with containerized architectures (Docker, Kubernetes) and orchestration platforms
  • Strong knowledge of distributed systems and Linux/Unix internals, with experience tuning performance and troubleshooting at scale
  • A related technical degree required
  • Excellent communication skills with demonstrated ability to lead during high-pressure incidents, present technical designs to leadership, and mentor junior engineers
  • Proven proficiency in Python and Go (GoLang) with strong software engineering practices (testing, code review, CI/CD)
  • 5+ years of experience in systems engineering and software engineering for large-scale, internet-facing services
  • Growth mindset with curiosity to explore new technologies and drive continuous improvement
  • Hands-on experience with workflow/orchestration engines (Temporal, Airflow, Argo Workflows, or similar) for building durable automation pipelines
  • Strong understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, blameless culture, and capacity planning
  • Experience with AI agent frameworks, MCP (Model Context Protocol), or building LLM-powered operational tools
  • Contributions to open-source reliability/observability tooling
  • AWS/GCP professional-level certifications
  • Prior experience in SRE organizations supporting multi-cloud or hyperscale environments
  • Experience with chaos engineering and game day exercises
  • Python and Go proficiency for systems-level tooling

This is an external listing. JobSpring does not represent or verify the employer. Report this listing