Skip to content
← Back to job listings

Infrastructure and MLOps Engineer

Graphcore · Bristol, United Kingdom

External listingfull-time18 days ago

About The Role

Join our Software Infrastructure team as an Infrastructure and MLOps Engineer. You will play a crucial role in scaling and managing our infrastructure, developing essential tools and services that empower our software team. Your contributions will enhance the build, test, deployment, and productization processes of our Machine Learning Software components. You will work with High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems.

  • Develop and maintain tools and services to support AI research and engineering teams, enhancing the build, test, deployment, and productization processes of Machine Learning Software components.
  • Deploy and maintain services with Kubernetes and Docker, manage Cloud Infrastructure using tools such as Terraform, and oversee ML accelerator hardware.
  • Collaborate with the Software Infrastructure team to manage the CI platform and services, build engineering, component integration, and packaging and release systems.
  • Knowledge of Python
  • Familiarity with cloud services (e.g. AWS)
  • Deploying ML orchestration tools (e.g. NV Ray, KFP, SkyPilot)
  • Managing ML accelerator hardware (e.g. DCGM)
  • Experience managing or developing in Linux environments
  • Maintaining machine learning applications
  • Experience of one of the following:
  • Understanding of CI/CD principles
  • Experience using Kubernetes (k8s)
  • Experience with Grafana
  • Knowledge of Go/Java/C++ (or similar language)
  • Experience with GitHub Actions
  • Experience with modern observability tooling (e.g. Prometheus)
  • Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing