Skip to content
← Back to job listings

Senior DevOps Engineer/Site Reliability Engineer

Stellar Cyber · United States

External listingfull-timeabout 2 months ago

About The Role

Join our globally distributed engineering organization as a Senior DevOps/Site Reliability Engineer. In this hands-on senior-level role, you will focus on building, operating, and scaling reliable cloud-native infrastructure and distributed data platforms. You will work closely with platform, development, and operations teams to drive automation, operational excellence, and reliability best practices for mission-critical systems. The ideal candidate will have strong expertise in Kubernetes, cloud infrastructure, observability, automation, CI/CD, incident management, and infrastructure reliability.

  • Administer and maintain Kubernetes clusters and containerized workloads, manage cloud infrastructure across OCI, AWS, GCP, or Azure environments.
  • Develop and maintain CI/CD pipelines for reliable application deployments, implement and manage Infrastructure as Code (IaC) using Terraform and Helm.
  • Drive observability initiatives including monitoring, logging, tracing, and alerting improvements, monitor, troubleshoot, and resolve production incidents.
  • Experience with CI/CD tools and deployment automation
  • Understanding of AI-driven operational tooling and automated remediation concepts
  • Excellent communication, collaboration, and problem-solving skills
  • Strong expertise with Kubernetes, Docker, and container orchestration
  • Hands-on experience managing production cloud environments
  • Familiarity with data platform technologies such as Kafka, Spark, Elasticsearch, Redis, or MongoDB
  • Resides on the East Coast
  • Knowledge of incident management and reliability engineering practices
  • 5+ years of experience in DevOps, SRE, or Platform Engineering roles
  • Experience supporting high-availability production systems and on-call operations
  • Strong programming and scripting skills in Python, Bash, or Go
  • Experience with observability platforms including Prometheus, Grafana, Loki, Alertmanager, and Elastic Stack
  • Advanced troubleshooting skills in Linux systems, networking, and distributed systems
  • Strong Infrastructure as Code experience with Terraform and Helm

This is an external listing. JobSpring does not represent or verify the employer. Report this listing