Skip to content
← Back to job listings

Staff Inference Software Engineer

CoreWeave · Sunnyvale, United States

External listingfull-time3 months ago

About The Role

Join CoreWeave's Inference team as a Staff Inference Software Engineer. In this role, you will lead cross-team design initiatives, optimize inference performance, and improve system reliability at scale. You will work deeply in distributed systems and Kubernetes-based infrastructure, focusing on scheduling, batching, and memory optimization. This position requires hands-on technical leadership and the ability to influence engineering direction across the organization.

  • Act as a technical leader driving architecture, performance, and reliability across multiple services and teams.
  • Lead cross-team design initiatives, optimizing inference performance (latency, throughput, and GPU utilization).
  • Design and operate low-latency, high-throughput systems with strict P95/P99 latency requirements.
  • This role requires hands-on technical leadership and the ability to influence engineering direction across the organization
  • Familiarity with mixed precision (BF16, FP8) and streaming inference workloads
  • Experience improving system performance using metrics-driven approaches (e.g., latency, throughput, utilization)
  • Proven experience leading cross-team technical initiatives impacting multiple services or organizations
  • Experience designing and operating low-latency, high-throughput systems with strict P95/P99 latency requirements
  • Strong programming skills in Go, Python, or C++
  • Strong understanding of distributed systems, networking, and performance optimization
  • 8–12+ years of experience building and operating large-scale distributed systems or cloud platforms
  • Deep expertise in Kubernetes at production scale, including orchestration, scheduling, and service design
  • Hands-on experience with inference systems, including batching or micro-batching strategies, caching, and memory optimization
  • You love to design and optimize high-performance distributed systems at scale
  • Exposure to large-scale AI/ML infrastructure or hyperscale cloud environments
  • You’re curious about AI inference, GPU systems, and emerging performance techniques
  • You’re an expert in building reliable, low-latency infrastructure and driving system-wide improvements
  • We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk
  • Experience with inference frameworks such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe
  • Experience leading multi-team or org-level technical initiatives
  • Experience with GPU systems and performance optimization (CUDA, NCCL, RDMA, NUMA, GPU interconnects)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing