Skip to content
← Back to job listings

大模型推理性能优化工程师

iluvatar · 建邺区, 江苏, 中国; 闵行区, 上海市, 中国

External listingfull-time14 days ago

About The Role

岗位职责: ● 优化大语言模型推理在自研 GPU 上的性能( 基于 vLLM / SGLang 等); ● 分析 GLM / Qwen / DeepSeek 等模型结构,定位计算与访存瓶颈; ● 优化 Attention / MoE 等关键路径; ● 负责量化加速(FP8 / INT8 / AWQ / GPTQ )与端到端调优; ● 编写与优化 CUDA / Triton Kernel(GEMM / Attention / Norm / Rope 等); ● 与芯片 / 编译器团队协同做性能分析与后端优化; 任职要求: ● 熟练 C/C++/Python; ● 深入理解 GPU 架构与内存体系; ● 熟悉 vLLM/SGLang 等大模型推理框架的机制原理; ● 熟练 CUDA / Triton,能独立优化复杂 Kernel; ● 熟悉 Nsight 等性能分析工具; ● 1年以上 GPU/AI 性能优化经验; 加分项: ● 有大模型推理落地优化经验; ● 有kernel/通信或编译器优化经验; ● 有开源框架性能贡献。

This is an external listing. JobSpring does not represent or verify the employer. Report this listing