← Back to job listings
SU
Senior Platform Security Engineer (Zero Trust and Platform Security)
Submer · United Kingdom
About The Role
Join a fast-growing scale-up as a Senior Platform Security Engineer. In this role, you will design, build, and operate a GPU-native cloud platform for AI and high-performance workloads. You will bridge the current production platform with the next-generation orchestration architecture, maintain and evolve existing deployments, and work closely with networking, storage, and platform engineers. This is a full-time position with a focus on deep hands-on engineering and ownership of critical orchestration components.
- Concevoir, construire et exploiter la couche d'orchestration de calcul pour une plateforme cloud native GPU.
- Maintenir et faire évoluer les déploiements basés sur CloudStack tout en contribuant activement à la conception de la plateforme de calcul de nouvelle génération.
- Travailler en étroite collaboration avec les ingénieurs en réseau, de stockage et de plateforme pour intégrer les primitives de la plateforme.
- Experience running GPU-heavy infrastructure for AI training, inference, or HPC workloads
- Proven experience working with large-scale distributed compute environments at a neo-cloud, hyperscaler, or HPC provider
- Strong experience with CloudStack internals, including extending and maintaining platform functionality
- Experience operating cloud orchestration platforms in production environments
- Experience designing and maintaining control-plane services for infrastructure platforms
- Experience maintaining or extending large Java codebases, ideally within infrastructure platforms
- Strong programming skills in Go and Python, with experience building cloud-native platform components
- Familiarity with workflow orchestration systems such as Argo Workflows
- Deep practical knowledge of Kubernetes internals and Slurm scheduling systems
- Experience building or operating compute orchestration layers for large-scale clusters
- Experience working with GPU infrastructure, including passthrough, NVIDIA MIG, scheduling, and lifecycle management of GPUs in distributed clusters
- Understanding of GPU virtualization and passthrough mechanisms such as QEMU PCI passthrough and NVIDIA MIG
- Familiar with virtual networking and distributed networking technologies such as OVS, OVN, VPC networking, RDMA, RoCE, ECMP, EVPN/VXLAN, and leaf-spine fabrics
- Comfortable solving difficult implementation and operational problems across CloudStack, Kubernetes, Slurm, and workflow orchestration; improving orchestration quality through code, automation, and practical design decisions; collaborating effectively across compute, networking, storage, and platform teams; and influencing engineering practices through expertise and delivery
- Comfortable mentoring peers and improving implementation quality, documentation, operational workflows, and platform reliability within the compute orchestration domain
- Able to independently own major compute-orchestration initiatives from design through rollout and operational stabilization
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring