← Back to job listings
GR
Staff Cloud Engineer
Graphcore · London, United Kingdom
About The Role
Join Graphcore as a Staff Cloud Engineer and be part of our Cloud Platform Team. In this hands-on technical role, you will develop and deploy cloud services, work closely with various teams, and contribute to the integration, validation, and optimization of our high-performance AI solutions. You will also have the opportunity to work with cutting-edge AI systems and make a significant impact in the field of AI technology.
- Contribuer au développement et au déploiement de services cloud, en travaillant en étroite collaboration avec les équipes de développement logiciel et d'exploitation des centres de données.
- Participer à l'intégration, à la validation, à l'optimisation et au développement de solutions d'IA haute performance, y compris les systèmes d'IA internes et les serveurs de haute performance.
- Développer et exploiter des services destinés aux utilisateurs finaux sur nos clouds, en transformant les exigences des utilisateurs finaux et des produits en services déployés.
- Strong proven Linux scripting ability (bash and python required)
- Strong proven Linux system administration (Ubuntu, RHEL and variants)
- Solid infrastructure or IT experience with a proven track record of delivering technical output as an individual contributor
- A solid understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems), virtual networks, block storage, resource management and monitoring
- Experience specifying, scoping, estimating and detailing work plans in an AGILE and SCRUM framework, including priorities, risks, issues, impacts and constraints
- Experience with IAC automation tools (e.g. Terraform/OpenTofu, Ansible, Packer)
- Hands-on experience deploying services into public or private clouds using Infrastructure-as-Code (IAC)
- Experience with container deployment and management tools (e.g. docker, podman, apptainer)
- Experience managing or operating on-premises or private-cloud environments
- Good communication and presentation skills, and experience dealing with end-users of IT or cloud services
- Experience with a version control system (preferably Git) and using it to manage system configuration or automation
- Bachelor's degree or equivalent practical experience in a relevant subject
- Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki, Mimir, OpenTelemetry, Fluentd ,Kafka
- An ability to work independently on critical infrastructure without oversight, and with a focus on end-user availability
- Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar
- Programming experience with Python3 utilising classes and inheritance
- Programming experience with Go
- Experience with managed switch configuration (e.g. EOS, SONiC, DNOS)
- Experience with managing production Kubernetes clusters and workloads
- Experience with workload queue management systems (SLURM, LSF, Kueue)
- Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions
- Strong skillset and experience in end-to-end deployment automation and CI of containerised services. Complete automation of pipelines for build, test, deploy, manage, alert, destroy, rebuild
- Experience with OpenStack deployments or the technologies they rely on (e.g. Ceph, Open vSwitch, KVM, QEMU )
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring