Senior Software Engineer (Infrastructure Engineering)
CoreWeave · Sunnyvale, United States
About The Role
Join CoreWeave, a leading cloud provider specializing in GPU-accelerated workloads. As a Senior Software Engineer in Infrastructure Engineering, you will play a crucial role in the development, deployment, and monitoring of services that manage our bare-metal infrastructure. You will lead incident response efforts, build a strategy for operational support and reliability, and collaborate with cross-functional teams to improve platform reliability. This role requires 7+ years of experience in cloud operations or site reliability engineering, proficiency in Go or Python, and familiarity with incident management practices.
- Lead incident response efforts by identifying and resolving service disruptions quickly, while coaching other junior team members through resolution.
- Build a strategy around making core services perform at their best at scale, including improvements to the services for robustness as well as supportability in production.
- Collaborate with engineers across teams to improve platform reliability, resilience improvements, and disaster recovery.
- 7+ years of experience in cloud operations, site reliability engineering (SRE), or related technical roles
- Previous experience deploying containerized applications using Kubernetes
- Prior experience with Prometheus / Grafana
- Proficiency with Go or Python
- Familiarity with incident management practices and frameworks (e.g., ITIL, SRE best practices)
- Understanding of cloud platforms (e.g., Kubernetes, AWS, GCP) and basic knowledge of cloud infrastructure
- Strong analytical and problem-solving abilities
- Served on an on-call rotation supporting production services
- Excellent documentation skills and attention to detail
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring