Skip to content
← Back to job listings

Senior Staff Site Reliability Engineer

Archer · San Jose, United States

External listingfull-time2 months ago

About The Role

Join our team as a Senior Staff Site Reliability Engineer (SRE) and take on a critical role in ensuring the reliability, scalability, performance, and security of our core systems and services. You will design, implement, and maintain robust infrastructure and automation solutions, develop observability strategies, optimize data pipelines, and drive continuous improvement of our CI/CD pipelines. You will also collaborate with development teams, troubleshoot complex production issues, and participate in on-call rotations.

  • Implement and maintain the infrastructure and pipeline required for an internal LLM-powered chat service, potentially leveraging platforms like OpenRouter or similar alternatives.
  • Drive the continuous improvement of our CI/CD pipelines, promoting best practices for automated testing, deployment, and release management.
  • Collaborate with development teams to ensure reliability is built into the software development lifecycle from inception.
  • Expert-level knowledge of cloud platforms (AWS preferred), including infrastructure-as-code principles
  • Proven track record in designing and implementing robust data pipelines (e.g., Kafka, Airflow, Spark)
  • Extensive experience with observability tools and practices, including Prometheus, Grafana, ELK stack, or similar
  • Excellent problem-solving, analytical, and communication skills
  • Proficiency in Docker for containerization and orchestration
  • Deep expertise in Amazon EKS, including cluster provisioning, management, and troubleshooting
  • Comprehensive understanding of security best practices for cloud environments, applications, and data
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience
  • 12+ years of experience in Site Reliability Engineering, DevOps, or a similar role with a strong focus on operational excellence
  • Ability to work independently and as part of a highly collaborative team
  • Solid understanding of networking concepts, distributed systems, and operating systems
  • Advanced scripting and programming skills in Python, Bash, and PowerShell
  • Strong background in CI/CD methodologies and tools (e.g., Jenkins, GitLab CI, ArgoCD)
  • Successful candidates must be able to demonstrate U.S. citizenship, permanent residency, or status as a protected individual to satisfy ITAR, contractual, and/or regulatory requirements
  • Certifications in AWS, Kubernetes, or other relevant technologies
  • Experience with other Kubernetes distributions or cloud providers
  • Familiarity with compliance frameworks (e.g., SOC 2, HIPAA, GDPR)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing