Skip to content
← Back to job listings

Senior Software Engineer (GPU Kernel Authoring & Optimization)

CoreWeave · Sunnyvale, United States

External listingfull-time16 days ago

About The Role

Join CoreWeave, the leading AI-cloud for high-performance GPU infrastructure. As a Senior Software Engineer, you will focus on kernel authoring and optimization, writing, profiling, and tuning GPU kernels for large-scale model serving. You will lead designs, raise engineering standards, and deliver measurable improvements to latency, throughput, and reliability across our inference stack.

  • Author, profile, and optimize CUDA kernels for large-scale model serving, focusing on maximizing throughput and minimizing latency.
  • Lead design reviews, drive architecture within the team, and mentor junior engineers to elevate coding/testing standards.
  • Implement and maintain benchmarking workflows for end-to-end MLPerf Inference (and Training) runs, ensuring reproducibility and documentation.
  • Strong coding in C++ and Python; comfortable reading and writing low-level, performance-sensitive code
  • Deep understanding of GPU architecture and performance: tensor cores, warp/occupancy tuning, the memory hierarchy and bandwidth, NVLink/PCIe, and profiling with Nsight Compute/Systems
  • Familiarity with model-serving stacks (vLLM, TensorRT-LLM, llm-d, SGLang) and the kernels that dominate their inference cost
  • Strong communicator comfortable collaborating with cross-functional teams and external partners
  • Hands-on CUDA experience is required—you have written and optimized custom kernels and are fluent with the CUDA programming and memory model
  • 5+ years of experience building high-performance computing, GPU/accelerator software, or performance-critical systems
  • Contributions to OSS projects such as vLLM, SGLang, PyTorch, Triton, or CUTLASS
  • NCCL and collective-communication performance
  • HIP / ROCm and AMD GPU experience
  • Experience with SUNK (Slurm on Kubernetes) / Slurm for scheduling large GPU jobs
  • CuTe DSL for Python-based kernel authoring on NVIDIA GPUs
  • Experience with alternative accelerators such as Google TPUs and Meta's MTIA
  • Experience running MLPerf submissions or similar large-scale audited benchmarks
  • Triton or Mojo for authoring custom GPU kernels — highly desired
  • Familiarity with kernel-authoring DSLs and nano-compilers such as KNYFE and its Block DSL
  • JAX and its Pallas kernel language for authoring kernels on GPU/TPU
  • Experience with Kubernetes at production scale
  • We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match

This is an external listing. JobSpring does not represent or verify the employer. Report this listing