Staff AI Inference and Acceleration Engineer
Figure · San Jose, United States
About The Role
Join Figure's Platform Software team as a Staff AI Inference & Acceleration Engineer. You will be the technical authority on AI workload mapping, optimization, and execution across the robot's compute hardware. Your responsibilities will include owning the on-board inference architecture, partitioning inference workloads, defining a system-level compute budget, evaluating next-generation acceleration hardware, optimizing inference toolchains, and profiling inference pipelines. You should have at least 8 years of industry experience in hardware acceleration, ML systems, or compute architecture, and a strong understanding of AI/ML inference and computer architecture.
- Posséder l'architecture d'inférence embarquée et optimiser l'exécution des charges de travail d'IA sur le matériel de calcul du robot.
- Définir et maintenir un budget de calcul au niveau système pour toutes les tâches d'inférence exécutées sur le robot.
- Collaborer étroitement avec l'équipe AI/ML pour définir les contraintes d'architecture du modèle qui sont compatibles avec le matériel.
- At least 8 years of industry experience in hardware acceleration, ML systems, or compute architecture
- Deep understanding of AI/ML inference — model formats (ONNX, TFLite, etc.), inference runtimes, and deployment pipelines
- Strong understanding of computer architecture — memory hierarchies, data movement, and heterogeneous compute
- M.S. or Ph.D. in Computer Engineering, Electrical Engineering, Computer Science, or a related field — or equivalent industry experience
- Experience profiling and benchmarking inference workloads across CPU, GPU, NPU, DSP
- Hands-on experience optimizing models for edge or embedded hardware using quantization, pruning, and operator-level tuning
- Solid software engineering skills in C++ and Python
- Strong cross-functional communication skills — able to work effectively across hardware, software, and AI/ML teams
- Familiarity with low-level toolchains and compilation frameworks (e.g. TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm)
- Knowledge of real-time operating constraints and their impact on inference scheduling
- Track record of co-designing model architectures with ML teams to meet hardware constraints
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring