← Back to job listings
HA
Site Reliability Engineer (Robotics)
Hadrian · San Francisco, United States
About The Role
Join our team as a Site Reliability Engineer (Robotics) and take ownership of the reliability of our robotics systems. You will build interfaces to our observability system, write code frameworks and tools to support our controls and robotics systems, and partner with various engineering teams to ensure reliability is baked in early. This role offers a comprehensive benefits package, including 100% coverage of platinum medical, dental, vision, and life insurance plans, a 401k, a relocation stipend, and a flexible vacation policy.
- Assurer la fiabilité des systèmes robotiques, des PLCs à Kubernetes, en développant des interfaces pour l'observabilité.
- Écrire des frameworks et des outils pour soutenir les systèmes de contrôle et de robotique, y compris les outils de diagnostic.
- Collaborer avec les équipes d'ingénierie pour intégrer la fiabilité dès le début, en développant des SLOs et des SLIs.
- Strong Communication. You can run a war room, write a post-mortem, and explain a reliability tradeoff to a stakeholder
- Systems Thinker. Focused on understanding the relationship among various systems to design sustainable solutions, not one-time fixes
- Problem Solver. Solving complex puzzles excites and motivates you to find an efficient solution
- T-Shaped Skill Set. Comfortable with bare metal Kubernetes, networking, GitOps workflows, and Infrastructure as Code (IaC). Also skilled in programming in TypeScript, Python, Golang, or C++
- Ownership. Someone who has owned the reliability of a production system where downtime had physical or operational consequences (manufacturing line, autonomous vehicle, lab automation, network operations)
- You've built automated or self-healing remediation at scale. We want systems that remove humans from the loop
- Background in edge/on-prem infrastructure. You’ve run Kubernetes at the edge (k3s, k0s, k0smotron), managing on-prem clusters, time-series at the edge, or air-gapped deployments. A deep understanding of Linux operating system fundamentals such as cgroups, sockets, and system tuning, is a big plus
- Deep understanding of shipping and storing telemetry data at scale. Experience with Kafka/MQTT/RabbitMQ is a plus
- Direct robotics experience. ROS/ROS2, OPC UA, EtherCAT, motion controllers, or fleet management for autonomous systems
- An individual who is self-directed and can deliver with high velocity
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring