← Back to job listings
SU
Staff Storage Platform Engineer (AI Storage, Radian Arc)
Submer · United Kingdom
About The Role
Join our team as a Staff Storage Platform Engineer, where you will design, build, and operate the AI storage layer that powers our large-scale GPU infrastructure. You will play a key role in architecting and evolving the storage platform across edge and core deployments, supporting the full lifecycle of AI workloads. This position combines architectural ownership, technical direction, and cross-functional influence with hands-on execution across storage design, deployment, performance engineering, troubleshooting, platform integration, and operational improvement.
- Design, build, and operate the AI storage layer powering large-scale GPU infrastructure, ensuring high throughput and predictable latency.
- Architect and evolve the storage platform across edge and core deployments, supporting the full lifecycle of AI workloads.
- Define storage strategies for distributed inference, fine-tuning, and training workloads, and optimize storage throughput and latency for GPU-heavy clusters.
- Familiarity with Kubernetes storage integrations such as CSI
- Experience operating large-scale storage clusters
- Experience owning both architecture and direct implementation in lean or fast-scaling environments is strongly preferred
- Strong understanding of AI workload data access patterns
- Deep knowledge of the Linux storage and I/O stack
- Experience optimizing storage for GPU-accelerated workloads
- Proven experience designing storage architectures for large-scale AI inference or training platforms, including dataset distribution, checkpointing, and KV-cache storage patterns
- Strong hands-on experience designing and operating distributed storage systems for high-performance compute environments
- The candidate should have deep expertise in designing and operating storage platforms optimized for GPU-heavy environments and distributed AI workloads
- This includes a strong understanding of how training, fine-tuning, and inference systems interact with storage, and how storage architecture affects throughput, latency, concurrency, checkpoint recovery, dataset distribution, and serving performance
- Familiarity with storage patterns for KV-cache persistence and retrieval
- Experience optimizing data locality and reducing unnecessary network movement between storage and compute
- Practical experience tuning storage architectures for checkpointing, distributed file access, object access, and high-concurrency inference
- Understanding of how storage performance affects large-scale AI frameworks, model-serving systems, and inference orchestration layers
- Experience designing storage platforms that support large dataset ingestion and model artifact distribution at scale
- Strong understanding of storage access patterns for distributed inference and training
- Distributed workload behavior
- Storage hardware,
- Linux kernel and I/O paths,
- Strong ability to act as the senior escalation point for ambiguous, high-impact, and multi-domain technical issues
- Filesystems,
- Kubernetes integrations,
- Experience designing storage observability systems
- Object and block storage layers,
- Strong knowledge of storage hardware, NVMe devices, storage fabrics, and high-performance data paths
- Ability to debug complex cross-layer issues spanning:
- Networking,
- Experience applying software engineering practices to storage automation and operational tooling
- Strong automation skills using Python and/or Bash
- Experience building reusable tooling, standards, validation patterns, or lifecycle automation that increase leverage across teams
- Able to balance short-term execution needs with long-term platform design, operational sustainability, and cost efficiency
- Proven ability to lead complex technical initiatives across teams
- Strong mentoring capability and ability to raise the technical level of adjacent engineering teams
- Demonstrated ability to set architectural direction and drive adoption of engineering standards across an organization
- Strong systems-level thinking balancing performance, reliability, scalability, operability, and cost efficiency
- Proven ability to lead through technical influence across multiple teams and domains, without relying on formal people management authority
- Comfortable collaborating across engineering, operations, deployment teams, vendors, and platform stakeholders
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring