Skip to content
← Back to job listings

AI Infrastructure Engineer (Sandbox Platform)

Scale AI · New York, United States

External listingfull-time11 days ago

About The Role

Join our AI Infrastructure team as a Software Engineer, where you'll build and evolve our agent sandboxing platform. You'll focus on deep systems expertise, developer experience, and collaboration with internal teams. Your responsibilities will include designing and building the sandboxing platform, optimizing performance, and responding to incidents. You'll also have access to comprehensive health coverage, personal and career growth opportunities, and parental support.

  • Design and build the sandboxing platform, client library, and API surface for secure code execution across containerized and virtualized environments.
  • Ensure strong isolation, security, and reproducibility of execution across user sessions and workloads, while optimizing for cold-start latency, memory footprint, and resource utilization at scale.
  • Partner closely with internal teams using the platform to understand their needs, debug issues, and build tooling that serves their use cases.
  • This is a role for someone who cares as much about the experience of the engineers and researchers using this system as they do about the kernel internals underneath it
  • Comfort working across infrastructure layers, from kernel modules to orchestration frameworks (e.g., Kubernetes)
  • Deep understanding of Linux internals: process isolation, memory management, cgroups, namespaces, etc
  • A track record of obsessing over developer experience — API design, error propagation, documentation, and the small details that make a library feel well-crafted
  • Strong debugging skills and the ability to navigate performance/security tradeoffs in production systems
  • 4+ years of experience building high-performance systems software, with meaningful time spent maintaining libraries, SDKs, or developer-facing APIs
  • Comfort with ambiguity, and the ability to context-switch between reactive incident work and proactive product development
  • Experience with containerization and virtualization technologies (e.g., Docker, Firecracker, gVisor, QEMU, Kata Containers)
  • Proficiency in a systems programming language such as Go, Rust, or C/C++
  • Experience as a founder or early engineer at an infrastructure-focused startup, owning a product end-to-end
  • Familiarity with LLM agents and agent frameworks (e.g., OpenHands, Agent2Agent, MCP)
  • Exposure to snapshotting and restore techniques (e.g., CRIU, VM snapshots, overlayfs)
  • Experience running secure workloads in multi-tenant or untrusted environments (e.g., FaaS, CI sandboxes, remote notebooks)
  • History of on-call/incident response for production systems
  • Open-source contributions to systems or developer-tools projects

This is an external listing. JobSpring does not represent or verify the employer. Report this listing