← Back to job listings
GR
Infrastructure and MLOps Engineer
Graphcore · Bristol, United Kingdom
About The Role
Join our Software Infrastructure team as an Infrastructure and MLOps Engineer. You will play a crucial role in scaling and managing our infrastructure, developing essential tools and services that empower our software team. Your contributions will enhance the build, test, deployment, and productization processes of our Machine Learning Software components. You will work with High-Performance Computing (HPC) AI platforms and gain invaluable experience in distributed systems.
- Develop and maintain tools and services to support AI research and engineering teams, enhancing the build, test, deployment, and productization processes of Machine Learning Software components.
- Deploy and maintain services with Kubernetes and Docker, manage Cloud Infrastructure using tools such as Terraform, and oversee ML accelerator hardware.
- Collaborate with the Software Infrastructure team to manage the CI platform and services, build engineering, component integration, and packaging and release systems.
- Knowledge of Python
- Familiarity with cloud services (e.g. AWS)
- Deploying ML orchestration tools (e.g. NV Ray, KFP, SkyPilot)
- Managing ML accelerator hardware (e.g. DCGM)
- Experience managing or developing in Linux environments
- Maintaining machine learning applications
- Experience of one of the following:
- Understanding of CI/CD principles
- Experience using Kubernetes (k8s)
- Experience with Grafana
- Knowledge of Go/Java/C++ (or similar language)
- Experience with GitHub Actions
- Experience with modern observability tooling (e.g. Prometheus)
- Experience with Infrastructure as Code (IaC) tools (e.g. Terraform/OpenTofu)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring