← Back to job listings
CO
Senior Software Engineer (GPU Kernel Authoring & Optimization)
CoreWeave · Sunnyvale, United States
About The Role
Join CoreWeave, the leading AI-cloud for high-performance GPU infrastructure. As a Senior Software Engineer, you will focus on kernel authoring and optimization, writing, profiling, and tuning GPU kernels for large-scale model serving. You will lead designs, raise engineering standards, and deliver measurable improvements to latency, throughput, and reliability across our inference stack.
- Author, profile, and optimize CUDA kernels for large-scale model serving, focusing on maximizing throughput and minimizing latency.
- Lead design reviews, drive architecture within the team, and mentor junior engineers to elevate coding/testing standards.
- Implement and maintain benchmarking workflows for end-to-end MLPerf Inference (and Training) runs, ensuring reproducibility and documentation.
- Strong coding in C++ and Python; comfortable reading and writing low-level, performance-sensitive code
- Deep understanding of GPU architecture and performance: tensor cores, warp/occupancy tuning, the memory hierarchy and bandwidth, NVLink/PCIe, and profiling with Nsight Compute/Systems
- Familiarity with model-serving stacks (vLLM, TensorRT-LLM, llm-d, SGLang) and the kernels that dominate their inference cost
- Strong communicator comfortable collaborating with cross-functional teams and external partners
- Hands-on CUDA experience is required—you have written and optimized custom kernels and are fluent with the CUDA programming and memory model
- 5+ years of experience building high-performance computing, GPU/accelerator software, or performance-critical systems
- Contributions to OSS projects such as vLLM, SGLang, PyTorch, Triton, or CUTLASS
- NCCL and collective-communication performance
- HIP / ROCm and AMD GPU experience
- Experience with SUNK (Slurm on Kubernetes) / Slurm for scheduling large GPU jobs
- CuTe DSL for Python-based kernel authoring on NVIDIA GPUs
- Experience with alternative accelerators such as Google TPUs and Meta's MTIA
- Experience running MLPerf submissions or similar large-scale audited benchmarks
- Triton or Mojo for authoring custom GPU kernels — highly desired
- Familiarity with kernel-authoring DSLs and nano-compilers such as KNYFE and its Block DSL
- JAX and its Pallas kernel language for authoring kernels on GPU/TPU
- Experience with Kubernetes at production scale
- We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring