← Back to job listings
GR
Infrastructure and MLOps Engineer
Graphcore · London, United Kingdom
About The Role
Join our Software Infrastructure team as an Infrastructure and MLOps Engineer. You will play a crucial role in scaling and managing our infrastructure, developing essential tools and services that empower our software team. Your contributions will enhance the build, test, deployment, and productization processes of our Machine Learning Software components. You will work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems.
- Develop and maintain tools and services to support AI research and engineering teams, enhancing the build, test, deployment, and productization processes of Machine Learning Software components.
- Deploy and maintain services using Kubernetes and Docker, and manage Cloud Infrastructure using tools such as Terraform.
- Collaborate with the Software Infrastructure team to manage the CI platform and services, build engineering, component integration, and packaging and release systems.
- Understanding of CI/CD principles
- Experience managing or developing in Linux environments
- Knowledge of Python
- Familiarity with cloud services (e.g. AWS)
- Managing ML accelerator hardware (e.g. DCGM)
- Maintaining machine learning applications
- Experience using Kubernetes (k8s)
- Experience of one of the following:
- Deploying ML orchestration tools (e.g. NV Ray, KFP, SkyPilot)
- Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)
- Experience with GitHub Actions
- Experience with modern observability tooling (e.g. Prometheus)
- Knowledge of Go/Java/C++ (or similar language)
- Experience with Grafana
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring