Skip to content
← Back to job listings

Senior Staff LLM Inference Engineer

d-Matrix · United States

External listingfull-timeabout 1 month ago

About The Role

Join D-Matrix as a Principal LLM Inference Engineer, where you will be responsible for transforming novel research ideas into deployed, optimized systems. You will work across the entire inference stack, from kernel-level optimization to distributed orchestration and high-level serving APIs. Your role will involve identifying and prototyping emerging LLM inference use cases, building proof-of-concept systems, developing custom kernels, and contributing to distributed inference systems. You will collaborate closely with hardware architects and partner with product and business development to translate POCs into customer-facing demonstrations.

  • Concevoir et prototyper des cas d'utilisation émergents pour l'inférence LLM adaptés aux déploiements matériels hétérogènes.
  • Développer et affiner des noyaux personnalisés et des optimisations au niveau des opérateurs pour maximiser le débit et minimiser la latence.
  • Contribuer aux systèmes d'inférence distribués : parallélisme tensoriel/pipeline, pré-remplissage/décodage désagrégé.
  • Are results-oriented with a strong bias toward action; you own problems end-to-end from prototype to optimization to handoff
  • Have deep intuition for modern generative AI architectures and how to squeeze performance out of them at inference time
  • Value clear communication and thrive in a small, high-ownership team environment
  • Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can extend or replace them when needed
  • Enjoy pathfinding new use cases — exploring heterogeneous deployment topologies and building early-stage POCs that prove out new ideas
  • Are energized by working at the intersection of novel hardware and frontier models, and want your work to directly influence how next-generation AI silicon is used
  • Master’s or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience
  • Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level
  • Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience
  • Hands-on experience optimizing LLM inference — attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4)
  • Strong proficiency in Python and C/C++
  • Experience with heterogeneous compute deployments — scheduling inference workloads across dissimilar hardware (accelerators, CPUs, GPUs)
  • Familiarity with GPU kernel programming (CUDA/Triton) and performance profiling tools
  • Familiarity with custom silicon or ASIC-based inference (beyond GPU-only environments)
  • Experience with distributed inference: tensor parallelism, pipeline parallelism, disaggregated serving
  • Experience with production inference serving at scale (latency SLOs, continuous batching, multi-model serving)
  • Contributions to open-source inference or ML systems projects
  • Familiarity with speculative decoding, mixture-of-experts routing, or long-context serving techniques
  • Working familiarity with the material in the JAX Scaling Book or equivalent systems-level understanding of modern LLM training and inference

This is an external listing. JobSpring does not represent or verify the employer. Report this listing