← Back to job listings
SE
Staff Infrastructure Engineer (Observability)
SentinelOne · New-York, United States
About The Role
Join SentinelOne as a Staff Infrastructure Engineer (Observability) and play a pivotal role in shaping the future of our critical systems. You will design, implement, and optimize solutions that underpin our global platform, empowering engineering teams across the organization. This high-impact role involves end-to-end ownership of critical infrastructure, mentoring talented engineers, and accelerating software delivery. You will also act as the primary Subject Matter Expert for our core observability stack and partner strategically with diverse engineering teams to define platform requirements.
- Architect and implement robust, scalable telemetry platforms that empower SentinelOne engineers to deploy and monitor features with speed, safety, and reliability.
- Act as the primary Subject Matter Expert (SME) and administrator for our core observability stack, including Grafana, Prometheus, Thanos/Mimir/Cortex, and OpenTelemetry (OTEL) pipelines.
- Drive exemplary operational efficiency for critical observability services across AWS and GCP, meticulously balancing unwavering system reliability with smart cloud cost-optimization.
- We’re looking for people who are relentlessly curious and committed to continuous learning
- Those who thrive here actively seek out new solutions, experiment thoughtfully, and apply what they learn to drive better, faster, smarter outcomes
- We are seeking a candidate who is driven by a deep passion for observability and technical leadership
- US Citizenship and the ability to work in a government-regulated environment
- 8+ years experience in architecting, scaling, and managing enterprise-grade observability stacks utilizing Prometheus, Grafana, Thanos (or Mimir/Cortex), and OpenTelemetry (OTEL)
- Demonstrated ability to lead complex technical designs, mentor other engineers, and collaborate cross-functionally with product and application teams
- 8+ years experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or a related systems-focused field
- Experience design-engineering cloud-native infrastructure within major cloud providers (AWS or GCP) and managing production Kubernetes environments (EKS, GKE)
- Advanced proficiency with IaC and automation tools, specifically Terraform and Ansible, to manage immutable infrastructure
- Experience maintaining and optimizing high-throughput, large-scale distributed systems with a focus on cost-efficiency, scalability, and disaster recovery
- Experience working with high-security compliance frameworks, specifically FedRAMP or other sovereign cloud requirements
- 8+ years production-level programming experience in GoLang (highly desirable) or another mainstream language (e.g., Python, Java) with a strong willingness to adopt GoLang
- Familiarity with the unique operational challenges of on-premises, hybrid, or air-gapped Kubernetes deployments
- Experience designing advanced CI/CD pipelines (e.g., GitHub Actions) and implementing sophisticated deployment strategies (canary, blue-green, rolling updates)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring