← Back to job listings
GR
Principal Cloud Engineer
Graphcore · Bristol, United Kingdom
About The Role
Join Graphcore, a leading AI technology company, as a Principal Cloud Engineer. In this hands-on technical role, you will work closely with various teams to develop and deploy cloud services on cutting-edge AI systems. You will be responsible for significant technical initiatives, mentoring junior engineers, and ensuring the peak performance of our AI systems. Enjoy a flexible work-life balance, private medical insurance, a pension plan, and opportunities for career progression and personal development.
- Contribuer au développement et au déploiement de services cloud, en travaillant en étroite collaboration avec les équipes de développement logiciel et d'exploitation des centres de données.
- Assurer l'intégration, la validation, l'optimisation et le développement de solutions d'IA haute performance, y compris les systèmes d'IA internes et les serveurs de haute performance.
- Mentorer une petite équipe d'ingénieurs moins expérimentés et être responsable d'initiatives techniques significatives.
- Expert-level, proven Linux system administration (Ubuntu, RHEL and variants)
- Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki, Mimir, OpenTelemetry, Fluentd ,Kafka
- Experience specifying, scoping, estimating and detailing work plans in an AGILE and SCRUM framework, including priorities, risks, issues, impacts and constraints
- Experience with Continuous Integration or testing pipelines using GitLab, GitHub or similar
- Experience with a version control system (preferably Git) and using it to manage system configuration or automation
- Experience with container deployment and management tools (e.g. docker, podman, apptainer)
- Bachelor's degree or equivalent practical experience in a relevant subject
- Hands-on experience deploying services into public or private clouds using Infrastructure-as-Code (IAC)
- A solid understanding of the technologies underpinning cloud services (APIs, virtualisation of CPUs, IO, systems), virtual networks, block storage, resource management and monitoring
- Expert with IAC automation tools (e.g. Terraform/OpenTofu, Ansible, Packer)
- Expert-level, proven Linux scripting ability (bash and python required)
- Excellent communication and presentation skills, and experience dealing with end-users of IT or cloud services
- Experience managing or operating on-premises or private-cloud environments
- Solid infrastructure or IT experience with a proven track record of delivering technical output as an individual contributor
- An ability to work independently and lead others on critical infrastructure without oversight, and with a focus on end-user availability
- Experience with OpenStack deployments or the technologies they rely on (e.g. Ceph, Open vSwitch, KVM, QEMU )
- Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions
- Strong skillset and experience in end-to-end deployment automation and CI of containerised services. Complete automation of pipelines for build, test, deploy, manage, alert, destroy, rebuild
- Experience with managing production Kubernetes clusters and workloads
- Experience with managed switch configuration (e.g. EOS, SONiC, DNOS)
- Experience with workload queue management systems (SLURM, LSF, Kueue)
- Programming experience with Python3 utilising classes and inheritance
- Programming experience with Go
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring