← Back to job listings
TA
Platform Engineer (Model Shaping)
Together AI · San Francisco, United States
About The Role
Join Together as a Platform Engineer in the Model Shaping team. You will design and build the infrastructure for model customization and evaluation, collaborate with other engineering teams, and contribute to reliability improvements for the platform. You will also create and improve internal tooling for deployment, continuous integration, and observability. This role requires experience in building infrastructure or backend components of production services, strong software engineering skills in Python or Go, and cloud environment administration experience.
- Design and build Together’s systems and infrastructure for model customization, including user-facing features and internal improvements.
- Collaborate with other engineers and researchers in the team to improve the infrastructure based on the needs of projects they work on.
- Contribute to reliability improvements for the platform, participating in an on-call rotation and improving processes for incident response.
- Experienced with infrastructure automation tools (Terraform, Ansible), monitoring/observability stacks (Prometheus, Grafana), and CI/CD pipelines (GitHub Actions, ArgoCD)
- 3+ years of experience in building infrastructure or backend components of production services
- Skilled with analyzing non-trivial issues of complex software systems and documenting your findings
- Strong software engineering background in Python or Go
- Strong communication skills, willing to document systems and processes and collaborate with peers of varying technical expertise
- Have cloud environment (e.g., AWS/GCP/Azure) administration experience, preferably with a hybrid bare-metal/cloud environment
- Comfortable with the fundamentals of Linux environments and modern container/orchestration stacks (e.g., Docker and Kubernetes)
- Developing large-scale production systems with high reliability requirements
- Pipeline orchestration frameworks (e.g., Kubeflow, Argo Workflows, Flyte)
- Deployment of services for AI training or inference
- Managing GPU workloads on HPC clusters, ideally with hands-on experience in operating NVIDIA’s networking stack (e.g., NCCL, Mellanox firmware, GPUDirect RDMA)
- Maintaining or contributing to open-source projects
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring