Skip to content
← Back to job listings

Staff Cloud Engineer

Graphcore · London, United Kingdom

External listingfull-time2 months ago

About The Role

Join Graphcore as a Staff Cloud Engineer and be part of our Cloud Platform Team. In this hands-on technical role, you will develop and deploy cloud services, work closely with various teams, and contribute to the integration, validation, and optimization of our high-performance AI solutions. You will also have the opportunity to work with cutting-edge AI systems and make a significant impact in the field of AI technology.

  • Contribuer au développement et au déploiement de services cloud, en travaillant en étroite collaboration avec les équipes de développement logiciel et d'exploitation des centres de données.
  • Participer à l'intégration, à la validation, à l'optimisation et au développement de solutions d'IA haute performance, y compris les systèmes d'IA internes et les serveurs de haute performance.
  • Développer et exploiter des services destinés aux utilisateurs finaux sur nos clouds, en transformant les exigences des utilisateurs finaux et des produits en services déployés.
  • Strong proven Linux scripting ability (bash and python required)
  • Strong proven Linux system administration (Ubuntu, RHEL and variants)
  • Solid infrastructure or IT experience with a proven track record of delivering technical output as an individual contributor
  • A solid understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems), virtual networks, block storage, resource management and monitoring
  • Experience specifying, scoping, estimating and detailing work plans in an AGILE and SCRUM framework, including priorities, risks, issues, impacts and constraints
  • Experience with IAC automation tools (e.g. Terraform/OpenTofu, Ansible, Packer)
  • Hands-on experience deploying services into public or private clouds using Infrastructure-as-Code (IAC)
  • Experience with container deployment and management tools (e.g. docker, podman, apptainer)
  • Experience managing or operating on-premises or private-cloud environments
  • Good communication and presentation skills, and experience dealing with end-users of IT or cloud services
  • Experience with a version control system (preferably Git) and using it to manage system configuration or automation
  • Bachelor's degree or equivalent practical experience in a relevant subject
  • Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki, Mimir, OpenTelemetry, Fluentd ,Kafka
  • An ability to work independently on critical infrastructure without oversight, and with a focus on end-user availability
  • Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar
  • Programming experience with Python3 utilising classes and inheritance
  • Programming experience with Go
  • Experience with managed switch configuration (e.g. EOS, SONiC, DNOS)
  • Experience with managing production Kubernetes clusters and workloads
  • Experience with workload queue management systems (SLURM, LSF, Kueue)
  • Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions
  • Strong skillset and experience in end-to-end deployment automation and CI of containerised services. Complete automation of pipelines for build, test, deploy, manage, alert, destroy, rebuild
  • Experience with OpenStack deployments or the technologies they rely on (e.g. Ceph, Open vSwitch, KVM, QEMU )

This is an external listing. JobSpring does not represent or verify the employer. Report this listing