Skip to content
← Back to job listings

Infrastructure and MLOps Engineer

Graphcore · London, United Kingdom

External listingfull-time13 days ago

About The Role

Join our Software Infrastructure team as an Infrastructure and MLOps Engineer. You will play a crucial role in scaling and managing our infrastructure, developing essential tools and services that empower our software team. Your contributions will enhance the build, test, deployment, and productization processes of our Machine Learning Software components. You will work with our High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems.

  • Develop and maintain tools and services to support AI research and engineering teams, enhancing the build, test, deployment, and productization processes of Machine Learning Software components.
  • Deploy and maintain services using Kubernetes and Docker, and manage Cloud Infrastructure using tools such as Terraform.
  • Collaborate with the Software Infrastructure team to manage the CI platform and services, build engineering, component integration, and packaging and release systems.
  • Understanding of CI/CD principles
  • Experience managing or developing in Linux environments
  • Knowledge of Python
  • Familiarity with cloud services (e.g. AWS)
  • Managing ML accelerator hardware (e.g. DCGM)
  • Maintaining machine learning applications
  • Experience using Kubernetes (k8s)
  • Experience of one of the following:
  • Deploying ML orchestration tools (e.g. NV Ray, KFP, SkyPilot)
  • Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)
  • Experience with GitHub Actions
  • Experience with modern observability tooling (e.g. Prometheus)
  • Knowledge of Go/Java/C++ (or similar language)
  • Experience with Grafana

This is an external listing. JobSpring does not represent or verify the employer. Report this listing