Skip to content
← Back to job listings

Applied AI Inference Engineer

Crusoe Energy Systems · Sunnyvale, United States

External listingfull-time5 days ago

About The Role

Join our team as an Applied AI Inference Engineer, where you will focus on optimizing large language models for production use. You will own the inference stack end to end, working on profiling, optimization, and deployment. This hands-on engineering role involves coding, profiling, and low-level optimization, as well as customer-facing interactions and elements of product and technical solutions work. You will have the opportunity to make a real impact on customer deployments and contribute to the advancement of AI technology.

  • Profiling and optimizing the inference stack for large language models, focusing on performance improvements.
  • Collaborating with customer engineering teams to tailor deployments to their specific models and constraints.
  • Taking ownership of the delivery process from experimentation to production, ensuring performance goals are met.
  • Familiarity with methods for optimizing LLMs for high throughput / low latency inference
  • Strong communication skills, particularly when explaining hard technical topics to customers and teammates
  • Clear interest and hands-on experience with large language models
  • A firm grasp of how GPUs are built and how they behave
  • Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level
  • A working knowledge of AI/ML pipelines and the full path of developing and deploying ML models
  • Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python
  • A Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field
  • A track record of making software systems run faster, especially for large language models
  • Experience with CUDA or comparable technologies
  • A strong command of software engineering fundamentals, with a record of building and shipping AI/ML inference systems
  • Prior work building or tuning AI/ML projects, particularly in a customer-facing setting
  • Experience with Docker and Kubernetes

This is an external listing. JobSpring does not represent or verify the employer. Report this listing