← Back to job listings
XA
Software Engineer (Platform Infrastructure, Rust, C++)
xAI · Palo Alto, United States
About The Role
Join our team as a Software Engineer focused on Platform Infrastructure. You will design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters. Your role will involve profiling, debugging, and optimizing performance across diverse systems, collaborating on hardware, software, and algorithm co-design, maintaining and innovating on our codebase, and developing tools to enhance team productivity. You will benefit from comprehensive health insurance, flexible vacation, visa sponsorship, and a 401(k) plan.
- Design, build, and implement a large-scale distributed system that powers one of the world's largest supercomputing clusters.
- Collaborate on hardware, software, and algorithm co-design to push the boundaries of AI training, and maintain and innovate on the codebase to ensure scalability and reliability.
- Develop tools to enhance team productivity and streamline workflows, and contribute directly to the company’s mission.
- All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence
- Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates
- Systems programming experience in C, C++, or Rust
- Computer systems fundamentals with a grasp of how computers execute code from transistors to high-level applications
- Hands-on expertise with Kubernetes (K8s), including cluster architecture, pod lifecycle, networking (CNI), storage (CSI), service mesh, and production-grade operations
- Collaborate in a fast-paced, open environment to design and foundational systems
- Strong debugging skills across the full stack — from kernel and OS up through container orchestration layers
- Deep knowledge of operating systems internals (process scheduling, memory management, file systems, and synchronization primitives)
- Solid understanding of computer networks and the TCP/IP stack
- Proficiency in performance analysis, profiling, and low-level optimization techniques
- Experience working with Linux kernel concepts or systems-level debugging tools (e.g., perf, gdb, strace, Wireshark)
- Proficiency deploying and managing workloads using Kubernetes manifests, Helm, Operators, and GitOps workflows
- Solid understanding of containerization technologies (Docker, containerd, crio) and their interaction with the Linux kernel
- Experience with observability and monitoring in distributed systems (Prometheus, Grafana, VictoriaMetrics, OpenTelemetry, or similar)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring