← Back to job listings
KL
Senior/Staff Site Reliability Engineer (Consumer Apps)
Klover · Chicago, United States
About The Role
Join Attain as a Senior/Staff Site Reliability Engineer, where you will be instrumental in building and maintaining the infrastructure that powers our systems. You will focus on automation, using AI agents to eliminate manual toil, and work closely with nearly every engineering team to ensure peak efficiency and scalability. Your responsibilities will include writing Terraform modules, developing Helm charts, monitoring databases, and improving our CI/CD pipeline.
- Automate manual processes and eliminate toil by leveraging AI agents and modern tooling.
- Collaborate with engineering teams to ensure systems are operating efficiently and are prepared for future growth.
- Design, implement, and maintain infrastructure resources using infrastructure-as-code tools like Terraform.
- You have a strong desire to automate things
- You have a willingness to learn and teach in a fast-paced, collaborative environment
- You reach for automation before you reach for a runbook, and a manual process is something you want to delete, not document
- You like to get your hands dirty and tinker with/stress test new technologies
- You treat AI agents as power tools and have real opinions about how to drive them — especially when to stop trusting them
- You are comfortable wearing many hats
- You readily provide constructive feedback, and also proactively seek feedback to improve yourself
- Strong computer science and software engineering fundamentals
- Experience with infrastructure-as-code tools such as Terraform
- Experience with SOC2 and PCI Compliance processes and requirements
- Experience working with the containerization technologies Docker, Kubernetes, and Istio or a similar service mesh technology
- A track record of replacing manual operations with durable automation
- Experience with pub sub technologies such as AWS SNS and Google Pub/Sub
- Experience with SQL database technologies such as MySQL, Google BigQuery, and Google Spanner
- Demonstrated fluency directing AI coding agents (e.g. Claude Code, Cursor, or similar) to build, operate, and debug real infrastructure; and robust and experienced judgment on verification of their work
- Experience with serverless computing technologies such as AWS Lambda and Google Cloud Functions/Google Cloud Run
- Experience with stream technologies such as Kafka and Amazon Kinesis
- 6+ years of experience building and maintaining large-scale cloud-native infrastructure (AWS and/or GCP)
- Experience with observability tools such as Datadog, Prometheus, and Grafana
- We encourage you to apply, even if your experience doesn’t match every detail on the job description. If we don’t see something that immediately fits, we will keep your resume on file for future opportunities
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring