← Back to job listings
LA
Staff Software Engineer (Infrastructure Storage)
Lambda · United States
About The Role
Join Lambda, a leading company in the AI infrastructure space. As a Staff Software Engineer (Infrastructure Storage), you will play a crucial role in building the foundational infrastructure that powers advanced AI research and products. You will have the opportunity to work at the intersection of large-scale distributed systems and artificial intelligence, making a significant impact on the future of AI.
- Set technical direction for storage software architecture across the Infrastructure Engineering organization, influencing decisions that span petabyte-scale deployments.
- Design, develop, and maintain high-performance storage systems software with a focus on performance, scalability, reliability, and operational simplicity.
- Work closely with storage software and networking teams to execute cross-functional infrastructure initiatives and new data center deployments, including integration of storage protocols across a variety of on-prem solutions.
- 10+ years of experience in storage systems engineering, with at least 5 years in a technical lead or Staff+ IC role
- Background working in high-performance computing, AI/ML infrastructure, or large-scale cloud storage environments
- Experience leading technical projects end-to-end, from architecture through delivery with cross-functional stakeholders
- Proven track record designing and operating storage infrastructure at scale (multi-petabyte environments preferred) in production data center or cloud settings
- Demonstrated ability to write high-performance, concurrent, production-grade systems code and conduct thorough code reviews
- Experience with kernel-level storage drivers, user-space I/O frameworks, or storage daemon development is a strong plus
- Strong proficiency in one or more low-level systems programming languages: C, C++, Rust, or Go
- Familiarity with DPDK and SPDK and their role in building high-performance, kernel-bypass storage and networking data paths
- Deep hands-on experience with two or more storage protocols: object (S3 or similar), block (iSCSI, Fibre Channel, NVMe-oF), or file (NFS, SMB, Lustre, DAOS)
- Familiarity with storage API performance characteristics such as latency, throughput, IOPS and the ability to diagnose and resolve bottlenecks at the protocol level
- Experience implementing or maintaining storage protocol servers or clients in production, not just consuming them
- Working knowledge of DPUs (e.g., NVIDIA BlueField) and their role in offloading storage and networking data paths
- Experience with GPU-direct storage or similar zero-copy data paths is a plus
- Familiarity with NVMe, NVMe-oF, and RDMA (RoCE or InfiniBand) and their impact on storage system architecture
- Understanding of I/O scheduling, caching layers, write amplification, and related performance tradeoffs
- Familiarity with storage observability tooling — metrics pipelines (Prometheus, Grafana), log aggregation, and tracing in distributed storage environments
- Experience building and operating storage systems with strong reliability expectations: designing for failure, building runbooks, and driving incident response
- Comfort working in a physical data center environment — understanding rack-scale infrastructure, storage array hardware, cabling, and failure domains
- Familiarity with tools such as fio, blktrace, perf, eBPF/bpftrace, or equivalent for storage performance analysis
- Experience profiling and tuning storage systems for throughput, latency, and IOPS under real production workloads
- Experience with NVIDIA BlueField DPUs or SuperNICs for accelerated storage data paths, including GPUDirect Storage implementation
- Deep production experience with enterprise or HPC storage platforms: Vast Data, Weka, NetApp, or IBM Spectrum Scale
- Experience deploying and operating Ceph at scale (100PB+) in an HPC or AI infrastructure environment
- Familiarity with emerging storage technologies such as CXL memory pooling, computational storage, or ZNS (Zoned Namespace) SSDs
- Experience contributing to or maintaining open-source storage projects (e.g., Ceph, DAOS, Lustre, MinIO)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring