← Back to job listings
CO
Senior Systems Engineer (OS Automation)
CoreWeave · Sunnyvale, United States
About The Role
Join our team as a Senior Systems Engineer (OS Automation) where you will design, build, and operate the services, APIs, and libraries that support our OS image, payload, and boot-configuration systems. You will work on a variety of projects, including a constraint-solver-based service, an end-to-end test framework, and a library suite for configuring node storage. This role requires strong proficiency in Go and/or Python, experience with Kubernetes, and a solid understanding of Linux systems.
- Concevoir, construire et exploiter les services, les API et les bibliothèques qui soutiennent nos systèmes d'image OS, de charge utile et de configuration de démarrage.
- Posséder et faire évoluer un service de configuration de démarrage qui modélise des relations de compatibilité et de dépendance complexes entre les images OS, les noyaux, les pilotes, les charges utiles et les types d'instance.
- Étendre notre cadre de test de bout en bout, natif de Kubernetes, qui valide les images et la configuration du système d'exploitation sur du matériel réel.
- Strong proficiency in Go and/or Python, with the ability to work fluently across both
- A collaborative, software-development-lifecycle mindset (sprints, planning, code review, design docs) and the judgment to refactor toward simplicity
- Comfort operating in a Kubernetes-based environment and reasoning about how software is built, packaged, deployed, and released
- Experience designing and maintaining APIs and service contracts (REST, gRPC, or similar), with an eye for clean, well-specified, versioned interfaces
- 3+ years of professional software engineering experience building and operating backend services, platforms, or developer/infrastructure tooling
- A demonstrated instinct for data modeling — representing relationships, constraints, and dependencies in code (graphs, constraint solving, relational models, or similar)
- A working understanding of how Linux systems boot and are configured (the OS image / cloud-init / provisioning lifecycle), even if you haven't owned it end to end
- Solid testing discipline: you write services that are testable, and you build the automation that proves they work
- Rust experience and/or workflow orchestration tools like Argo Workflows
- Experience modeling complex problems in novel ways
- Familiarity with bare-metal or node provisioning — PXE-style network boot, cloud-init, OS image building, firmware/driver enablement
- Fluency with NVIDIA GPU platforms
- To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency
- Interest in applying LLMs, RAG, and predictive modeling to large-scale infrastructure automation
- GRPC / Connect-RPC and Protobuf experience, including evolving service contracts safely over time
- Linux packaging and repository management, configuration management (e.g., Ansible), and shell-based build pipelines
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring