Skip to content
← Back to job listings

Principal Cloud Engineer

Graphcore · Bristol, United Kingdom

External listingfull-time26 days ago

About The Role

Join Graphcore, a leading AI technology company, as a Principal Cloud Engineer. In this hands-on technical role, you will work closely with various teams to develop and deploy cloud services on cutting-edge AI systems. You will be responsible for significant technical initiatives, mentoring junior engineers, and ensuring the peak performance of our AI systems. Enjoy a flexible work-life balance, private medical insurance, a pension plan, and opportunities for career progression and personal development.

  • Contribuer au développement et au déploiement de services cloud, en travaillant en étroite collaboration avec les équipes de développement logiciel et d'exploitation des centres de données.
  • Assurer l'intégration, la validation, l'optimisation et le développement de solutions d'IA haute performance, y compris les systèmes d'IA internes et les serveurs de haute performance.
  • Mentorer une petite équipe d'ingénieurs moins expérimentés et être responsable d'initiatives techniques significatives.
  • Expert-level, proven Linux system administration (Ubuntu, RHEL and variants)
  • Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki, Mimir, OpenTelemetry, Fluentd ,Kafka
  • Experience specifying, scoping, estimating and detailing work plans in an AGILE and SCRUM framework, including priorities, risks, issues, impacts and constraints
  • Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar
  • Experience with a version control system (preferably Git) and using it to manage system configuration or automation
  • Experience with container deployment and management tools (e.g. docker, podman, apptainer)
  • Bachelor's degree or equivalent practical experience in a relevant subject
  • Hands-on experience deploying services into public or private clouds using Infrastructure-as-Code (IAC)
  • A solid understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems), virtual networks, block storage, resource management and monitoring
  • Expert with IAC automation tools (e.g. Terraform/OpenTofu, Ansible, Packer)
  • Expert-level, proven Linux scripting ability (bash and python required)
  • Excellent communication and presentation skills, and experience dealing with end-users of IT or cloud services
  • Experience managing or operating on-premises or private-cloud environments
  • Solid infrastructure or IT experience with a proven track record of delivering technical output as an individual contributor
  • An ability to work independently and lead others on critical infrastructure without oversight, and with a focus on end-user availability
  • Experience with OpenStack deployments or the technologies they rely on (e.g. Ceph, Open vSwitch, KVM, QEMU )
  • Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions
  • Strong skillset and experience in end-to-end deployment automation and CI of containerised services. Complete automation of pipelines for build, test, deploy, manage, alert, destroy, rebuild
  • Experience with managing production Kubernetes clusters and workloads
  • Experience with managed switch configuration (e.g. EOS, SONiC, DNOS)
  • Experience with workload queue management systems (SLURM, LSF, Kueue)
  • Programming experience with Python3 utilising classes and inheritance
  • Programming experience with Go

This is an external listing. JobSpring does not represent or verify the employer. Report this listing