Skip to content
← Back to job listings

Staff Software Engineer (Infrastructure Storage)

Lambda · United States

External listingfull-timeabout 1 month ago

About The Role

Join Lambda, a leading company in the AI infrastructure space. As a Staff Software Engineer (Infrastructure Storage), you will play a crucial role in building the foundational infrastructure that powers advanced AI research and products. You will have the opportunity to work at the intersection of large-scale distributed systems and artificial intelligence, making a significant impact on the future of AI.

  • Set technical direction for storage software architecture across the Infrastructure Engineering organization, influencing decisions that span petabyte-scale deployments.
  • Design, develop, and maintain high-performance storage systems software with a focus on performance, scalability, reliability, and operational simplicity.
  • Work closely with storage software and networking teams to execute cross-functional infrastructure initiatives and new data center deployments, including integration of storage protocols across a variety of on-prem solutions.
  • 10+ years of experience in storage systems engineering, with at least 5 years in a technical lead or Staff+ IC role
  • Background working in high-performance computing, AI/ML infrastructure, or large-scale cloud storage environments
  • Experience leading technical projects end-to-end, from architecture through delivery with cross-functional stakeholders
  • Proven track record designing and operating storage infrastructure at scale (multi-petabyte environments preferred) in production data center or cloud settings
  • Demonstrated ability to write high-performance, concurrent, production-grade systems code and conduct thorough code reviews
  • Experience with kernel-level storage drivers, user-space I/O frameworks, or storage daemon development is a strong plus
  • Strong proficiency in one or more low-level systems programming languages: C, C++, Rust, or Go
  • Familiarity with DPDK and SPDK and their role in building high-performance, kernel-bypass storage and networking data paths
  • Deep hands-on experience with two or more storage protocols: object (S3 or similar), block (iSCSI, Fibre Channel, NVMe-oF), or file (NFS, SMB, Lustre, DAOS)
  • Familiarity with storage API performance characteristics such as latency, throughput, IOPS and the ability to diagnose and resolve bottlenecks at the protocol level
  • Experience implementing or maintaining storage protocol servers or clients in production, not just consuming them
  • Working knowledge of DPUs (e.g., NVIDIA BlueField) and their role in offloading storage and networking data paths
  • Experience with GPU-direct storage or similar zero-copy data paths is a plus
  • Familiarity with NVMe, NVMe-oF, and RDMA (RoCE or InfiniBand) and their impact on storage system architecture
  • Understanding of I/O scheduling, caching layers, write amplification, and related performance tradeoffs
  • Familiarity with storage observability tooling — metrics pipelines (Prometheus, Grafana), log aggregation, and tracing in distributed storage environments
  • Experience building and operating storage systems with strong reliability expectations: designing for failure, building runbooks, and driving incident response
  • Comfort working in a physical data center environment — understanding rack-scale infrastructure, storage array hardware, cabling, and failure domains
  • Familiarity with tools such as fio, blktrace, perf, eBPF/bpftrace, or equivalent for storage performance analysis
  • Experience profiling and tuning storage systems for throughput, latency, and IOPS under real production workloads
  • Experience with NVIDIA BlueField DPUs or SuperNICs for accelerated storage data paths, including GPUDirect Storage implementation
  • Deep production experience with enterprise or HPC storage platforms: Vast Data, Weka, NetApp, or IBM Spectrum Scale
  • Experience deploying and operating Ceph at scale (100PB+) in an HPC or AI infrastructure environment
  • Familiarity with emerging storage technologies such as CXL memory pooling, computational storage, or ZNS (Zoned Namespace) SSDs
  • Experience contributing to or maintaining open-source storage projects (e.g., Ceph, DAOS, Lustre, MinIO)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing