Machine Learning Infrastructure Engineer
Character.ai · United States
About The Role
Join our team as a Machine Learning Infrastructure Engineer, where you will provide crucial infrastructure support for our ML research and product development. You will build tools to diagnose cluster issues and hardware failures, monitor deployments, manage experiments, and maximize GPU allocation and utilization. This role requires a minimum of 4 years of experience in supporting infrastructure within an ML environment, experience with GPUs, and familiarity with cloud platforms and ML frameworks. Enjoy a generous benefits package, including a 401(K) contribution, top-notch health coverage, 4 weeks of PTO, paid leave for new parents, daily in-office catering, and a monthly wellness stipend.
- Provide infrastructure support to machine learning research and product development, ensuring optimal performance and reliability.
- Build and maintain tooling to diagnose cluster issues and hardware failures, enhancing the overall efficiency of the ML infrastructure.
- Monitor deployments, manage experiments, and support research efforts, maximizing GPU allocation and utilization for both serving and training.
- We’re looking for seasoned ML Infrastructure engineers with experience designing, building and maintaining training and serving infrastructure for ML research
- Experience working with GPUs
- Experience in developing tools used to diagnose ML infrastructure problems and failures
- 4+ years of experience supporting the infrastructure within an ML environment
- Experience with GPU kernel development
- Experience with supporting large language model training
- Experience with cloud platforms (e.g., Compute Engine, Kubernetes, Cloud Storage)
- Experience with large GPU clusters and high-performance computing/networking
- Experience with ML frameworks like Pytorch/TensorFlow/JAX
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring