Skip to content
← Back to job listings

Staff Site Reliability Engineer (Splunk)

Okta · Washington, United States

External listingfull-time2 months ago

About The Role

Join our team as a Staff Site Reliability Engineer (Observability) and take ownership of our Splunk ecosystem. You will design, build, and maintain scalable observability infrastructure, optimize log data collection and processing, participate in incident response, and automate the deployment of observability agents and collectors. This role requires a minimum of 5 years of experience in an SRE, DevOps, or Systems Engineering role, with a focus on high-availability systems and expertise in Splunk. Enjoy benefits such as work from home opportunities, health and wellness programs, financial benefits, and time off.

  • Evoluer et gérer l'écosystème Splunk de l'entreprise, en optimisant la collecte, le traitement et le stockage des données de journal.
  • Automatiser le déploiement des agents et des collecteurs à travers des systèmes distribués complexes en utilisant Terraform et des compétences en programmation.
  • Participer aux rotations d'appel et diriger les revues post-incident pour améliorer systématiquement les processus.
  • SRE Mindset: Minimum 5+ years of experience in an SRE, DevOps, or Systems Engineering role with a focus on high-availability systems
  • Log Management: Minimum 5+ Experience scaling and managing Splunk Cloud at scale (1000+ SVCs), including Workload Management (WLM) and HEC optimization. Visualization: Expertise in creating intuitive, actionable Splunk dashboards that correlate data across multiple sources
  • Programming Proficiency: Strong coding skills in SPL, Go for building internal tools and automating workflows
  • Problem Solving: A data-driven approach to debugging complex, cross-service performance bottlenecks
  • Distributed Systems: Deep understanding of Linux internals, networking (TCP/IP, DNS, Load Balancing), and container orchestration (Kubernetes/EKS)
  • Telemetry Standards: Hands-on experience with OpenTelemetry (OTel), Vector, or similar frameworks for instrumenting applications
  • Charge-back app: Experience in implementing Splunk charge-back app for usage reporting
  • Cloud Platforms: Experience managing observability native tools within AWS or GCP

This is an external listing. JobSpring does not represent or verify the employer. Report this listing