← Back to job listings
SS
AI Systems Performance Engineer (New Graduate)
SambaNova Systems · San Jose, United States
About The Role
Join SambaNova, a leading AI technology company, as an AI Systems Performance Engineer. In this role, you will work hands-on with advanced AI models and optimize their performance on SambaNova's reconfigurable dataflow platform. You will collaborate with engineers across various teams and contribute to the development of high-performance AI applications. This is an ideal opportunity for new graduates passionate about AI and its execution on real hardware.
- Collaborate with engineers across model, compiler, runtime, and hardware teams to bring up and optimize state-of-the-art foundation models on SambaNova's platform.
- Analyze and profile model execution to identify performance bottlenecks and optimize AI workloads for throughput, latency, memory efficiency, and scalability.
- Develop tools, benchmarks, and performance analysis methodologies for large-scale AI inference, and investigate new model architectures.
- This is an ideal opportunity for a new graduate who is passionate about understanding how AI models execute on real hardware and wants to help build the next generation of high-performance AI systems
- Bachelor's or Master's degree in computer science, electrical engineering, computer engineering, or a related technical field (e.g., applied mathematics, physics, or statistics), completed or expected before the start date
- Strong programming skills in Python, C++, or a similar programming language
- Familiarity with deep learning and at least one major ML framework, such as PyTorch, TensorFlow, or JAX
- Solid foundations in algorithms, data structures, computer architecture, operating systems, or parallel computing
- Ability and enthusiasm to learn across machine learning, software systems, and hardware
- We value strong technical fundamentals, curiosity, and the ability to learn quickly. Prior production experience with large-scale AI systems is not required
- Strong analytical and problem-solving skills, with an interest in understanding and optimizing system performance
- Coursework, research, internship, or project experience in machine learning systems, computer architecture, compilers, distributed systems, or high-performance computing
- Hands-on experience with LLMs, multimodal models, or transformer architectures
- Familiarity with model inference, KV cache, batching, quantization, or distributed execution
- Experience with GPU or accelerator programming using CUDA, Triton, OpenCL, or similar technologies
- Familiarity with frameworks such as vLLM, DeepSpeed, Megatron, or TensorRT
- Understanding of memory hierarchy, caching, parallelism, or scheduling
- Experience profiling and optimizing the performance of software or ML workloads
- Research publications, open-source contributions, programming competitions, or technically challenging personal projects are a plus
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring