Skip to content
← Back to job listings

Senior Software Engineer

CoreWeave · Sunnyvale, United States

External listingfull-time13 days ago

About The Role

Join our Platform & Infrastructure Engineering team as a Senior Software Engineer. You will be responsible for the reliability and performance of our Kubernetes-based data platform, architecting and operating highly available, multi-region systems. Your work will involve scaling infrastructure, optimizing deployment pipelines, and strengthening our security posture. You should have 7+ years of experience in Platform Engineering, Infrastructure Engineering, or building highly scalable distributed systems, with a strong foundation in software design, development, and algorithmic problem-solving.

  • Assumer la responsabilité de la fiabilité et des performances de la plateforme de données basée sur Kubernetes, en architecturant et en exploitant des systèmes hautement disponibles.
  • Participer à l'automatisation, à l'observabilité approfondie et à la résilience des systèmes, en soutenant des exigences de disponibilité strictes.
  • Concevoir et mettre en œuvre des pipelines CI/CD robustes, en utilisant des outils tels qu'Argo CD et GitHub Actions, et en appliquant les meilleures pratiques de sécurité.
  • You love building highly reliable systems that operate at scale
  • Data Platform Experience: Familiarity operating data platforms or data-intensive workloads, including distributed processing and streaming frameworks such as Spark, Airflow, Kafka, or Flink
  • CI/CD Systems: Proven track record building and operating robust CI/CD pipelines using tools such as Argo CD and GitHub Actions
  • Distributed Systems: Hands-on experience developing large-scale distributed systems, databases, and backend APIs
  • Cloud-Native Security: Hands-on experience applying security best practices in cloud-native settings, including secrets management, network policies, and vulnerability scanning
  • We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk
  • You bring 7+ years of hands-on experience in Platform Engineering, Infrastructure Engineering, or building highly scalable distributed systems — with a strong foundation in software design, development, and algorithmic problem-solving
  • Observability: Strong experience building and owning full-stack observability solutions, including metrics, logging, and distributed tracing using tools such as Prometheus, Grafana, and OpenTelemetry
  • You're an expert in diagnosing and solving complex distributed systems problems
  • Production Ownership: Demonstrated experience owning mission-critical systems with high availability requirements (≥99.99% uptime), including incident response, SLI/SLO/SLA definition, error budget management, and blameless postmortems
  • Infrastructure as Code: Proficiency with IaC tooling such as Helm, Terraform, or Pulumi, and experience with automated environment provisioning
  • Wondering if you're a good fit?
  • Multi-Region Architecture: Practical experience designing and operating geo-replicated, active-active, multi-region systems — with a solid grasp of traffic routing, failover strategies, and data consistency tradeoffs
  • Performance & Capacity: Strong command of system performance tuning, capacity planning, and resource optimization in distributed environments
  • You're curious about how to continuously improve system resilience, security, and operations
  • Kubernetes & Containerization: Deep expertise in Kubernetes cluster design, day-to-day operations, and production troubleshooting across containerized service environments
  • Regulated Environments: Experience working within compliance-driven environments and a working knowledge of regulatory frameworks such as GDPR, SOC 2, HIPAA, or SOX
  • Internal Developer Platforms: Background in designing and building internal developer platforms or self-service infrastructure tooling that empowers engineering teams to move faster and operate independently

This is an external listing. JobSpring does not represent or verify the employer. Report this listing