← Back to job listings
AR
Staff DevOps Engineer
Archer · San Jose, United States
About The Role
Join our team as a Staff DevOps Engineer, where you will architect, develop, and scale core services, real-time data pipelines, and distributed infrastructure for our cutting-edge AI products. You will work cross-functionally with AI Research, Engineering, and Data Science teams, and own the full development lifecycle. You should have a strong background in software engineering, cloud platforms, and infrastructure management.
- Architect, develop, and scale core services, real-time data pipelines, and distributed infrastructure for AI products.
- Leverage infrastructure-as-code (IaC) and cloud-native tooling to build, manage, and optimize production environments.
- Work cross-functionally with AI Research, Engineering, and Data Science teams to containerize, deploy, orchestrate, and monitor scalable ML models.
- Hands-on experience with modern GitOps practices (e.g., Argo CD) and managing CI/CD automation pipelines
- Solid experience with containerization and orchestration (Docker, Kubernetes), including configuring network policies, storage, and auto-scaling
- Strong proficiency in Java, Python and/or Rust, with a solid grasp of writing high-performance, concurrent backend code
- Production experience with stream-processing frameworks, specifically Apache Flink
- BS/MS/PhD degree in Computer Science, Software Engineering, or a related technical field
- Able to scale yourself with AI coding assistants while still fully understanding and taking absolute ownership of all code and infrastructure configurations you commit
- 5+ years of professional software engineering experience with a heavy focus on backend systems, distributed architectures, and infrastructure management
- Deep experience with cloud platforms (AWS, GCP, or Azure) and infrastructure-as-code tools such as OpenTofu/Terraform and Ansible
- Solid experience with relational databases (e.g., PostgreSQL), in-memory datastores (e.g., Valkey/Redis), and high-throughput data storage strategies
- Experience setting up observability stacks (eg Grafana, Prometheus, ELK, etc.)
- Excellent communication skills, with a proven ability to design resilient architectures and explain complex infrastructure paradigms to the broader team
- Deep experience with Apache Pulsar (or Kafka) handling high-throughput, event-driven streaming architectures
- Rust experience with async runtimes (Tokio) and message-driven architectures
- Experience architecting real-time media pipelines using WebRTC (LiveKit, RealtimeKit) with a focus on server-side jitter buffering and network impairment handling
- Deep understanding of audio codecs, specifically optimizing server-side Opus encoding/decoding for CPU efficiency and low-latency distribution
- Familiarity with MLOps frameworks, GPU orchestration in Kubernetes, and serving LLMs at scale
- Background in high-availability, real-time systems, telecommunications, or safety-critical networks
- Prior experience or a deep interest in aerospace, aviation, or autonomous tracking systems
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring