← Back to job listings
CO
Staff Software Engineer
CoreWeave · Sunnyvale, United States
About The Role
Join our Platform & Infrastructure Engineering team as a Staff Software Engineer. You will take end-to-end ownership of the reliability and performance of our Kubernetes-based data platform, architecting and operating highly available, multi-region systems. Your work will involve scaling infrastructure, optimizing deployment pipelines, and strengthening our security posture. You should have 10+ years of experience in Platform Engineering, Infrastructure Engineering, or building highly scalable distributed systems.
- End-to-end ownership of the reliability and performance of Kubernetes-based data platform.
- Architecting and operating highly available, multi-region systems designed to meet rigorous uptime and latency standards.
- Scaling infrastructure, optimizing deployment pipelines, and strengthening security posture while managing scalable systems.
- Cloud-Native Security: Hands-on experience applying security best practices in cloud-native settings, including secrets management, network policies, and vulnerability scanning
- Observability: Strong experience building and owning full-stack observability solutions, including metrics, logging, and distributed tracing using tools such as Prometheus, Grafana, and OpenTelemetry
- Performance & Capacity: Strong command of system performance tuning, capacity planning, and resource optimization in distributed environments
- Production Ownership: Demonstrated experience owning mission-critical systems with high availability requirements (≥99.99% uptime), including incident response, SLI/SLO/SLA definition, error budget management, and blameless postmortems
- Multi-Region Architecture: Practical experience designing and operating geo-replicated, active-active, multi-region systems — with a solid grasp of traffic routing, failover strategies, and data consistency tradeoffs
- Infrastructure as Code: Proficiency with IaC tooling such as Helm, Terraform, or Pulumi, and experience with automated environment provisioning
- Data Platform Experience: Familiarity operating data platforms or data-intensive workloads, including distributed processing and streaming frameworks such as Spark, Airflow, Kafka, or Flink
- Kubernetes & Containerization: Deep expertise in Kubernetes cluster design, day-to-day operations, and production troubleshooting across containerized service environments
- CI/CD Systems: Proven track record building and operating robust CI/CD pipelines using tools such as Argo CD and GitHub Actions
- Distributed Systems: Hands-on experience developing large-scale distributed systems, databases, and backend APIs
- You bring 10+ years of hands-on experience in Platform Engineering, Infrastructure Engineering, or building highly scalable distributed systems — with a strong foundation in software design, development, and algorithmic problem-solving
- Internal Developer Platforms: Background in designing and building internal developer platforms or self-service infrastructure tooling that empowers engineering teams to move faster and operate independently
- Regulated Environments: Experience working within compliance-driven environments and a working knowledge of regulatory frameworks such as GDPR, SOC 2, HIPAA, or SOX
- You're an expert in diagnosing and solving complex distributed systems problems
- You're curious about how to continuously improve system resilience, security, and operations
- We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk
- You love building highly reliable systems that operate at scale
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring