Software Engineer (Training/Inference, C++)
xAI · Palo Alto, United States
About The Role
Join our team as a Software Engineer focused on training and inference. You will design and optimize large-scale model serving systems, ensuring lightning speed and perfect reliability for millions of users. Your responsibilities will include architecting and implementing scalable distributed infrastructure, optimizing latency and throughput, building reliable serving systems, and developing custom tools for issue tracing and resolution. You will also create robust CI/CD infrastructure and accelerate research on next-generation systems. Enjoy comprehensive health insurance, flexible vacation, visa sponsorship, and a 401(k) plan.
- Architect and implement scalable distributed infrastructure for model serving, including load balancing, auto-scaling, batch scheduling, and global KV cache.
- Optimize latency and throughput of model inference under real production workloads, and build reliable, high-concurrency serving systems that serve billions of users.
- Benchmark, fine-tune, and accelerate inference engines, including low-level GPU kernel work and code generation, and develop custom tools to trace, replay, and fix issues.
- Experience with large-scale, high-concurrent production serving
- Low-level inference optimizations: GPU kernels, code generation
- Experience with GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.)
- Experience with testing, benchmarking, and reliability of inference services
- Deep low-level systems programming (C/C++ or Rust)
- Experience designing and implementing CI/CD infrastructure for inference
- Algorithmic inference optimizations: quantization, speculative decoding, distillation, low-precision numerics
- Strong background in system optimizations: batching, caching, load balancing, parallelism
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring