← Back to job listings
AX
AI-First SRE/DevOps Engineer
Axiad · San Jose, United States
About The Role
Join Axiad, a startup focused on identity visibility and intelligence. As an AI-First SRE/DevOps Engineer, you will own the reliability, observability, and delivery of our cloud-native Kubernetes platform. You will build CI/CD pipelines, infrastructure-as-code, and advocate for AI-First operations. The role requires deep operational expertise in Kubernetes, CI/CD, and infrastructure-as-code, along with practical experience running AI/LLM systems in production.
- Assumer la responsabilité de la fiabilité, de l'observabilité et de la livraison d'une plateforme Kubernetes cloud-native multi-tenant, de la conception à la production.
- Construire (et pas seulement exploiter) des pipelines CI/CD, une infrastructure en tant que code et une livraison progressive pilotée par GitOps.
- Automatiser la réponse aux incidents, les runbooks et la remédiation, et intégrer des agents d'IA pour trier, diagnostiquer et proposer des solutions.
- 5–8 years of professional experience in SRE, DevOps, or platform engineering roles
- Ownership: you take problems from ambiguity to resolution without waiting for a ticket, a spec, or permission. When something you own breaks, you're the first to know and the first to act
- Strong Kubernetes operational experience — running it in production, not just deploying to it
- Axiad is seeking a skilled AI-First SRE/DevOps Engineer with 5–8 years of hands-on infrastructure and platform engineering experience to help build and run Mesh, our Identity Visibility and Intelligence Platform (IVIP) — a cloud-native microservices platform on Kubernetes spanning human identity, non-human identity (NHI), post-quantum cryptography, and agentic AI identity risk
- Fluency with infrastructure-as-code, GitOps, and modern CI/CD; comfortable scripting and building tooling (Go or Python preferred)
- Builder mentality: you'd rather create a tool, platform, or automation than run a manual process twice. You ship things and stand behind them
- Solid observability expertise and SLO-driven operations experience
- The ideal candidate has a builder mentality and a strong AI-First mindset: automation and AI are the default, not the afterthought, and infrastructure is something you create, not just maintain
- Experience with containerization (Docker) and service mesh concepts
- Cloud-native depth on at least one major cloud provider
- Demonstrable adoption of an AI-First mindset and tools (Claude Code, Cursor, or Windsurf). Daily use of at least one AI development tool is a must
- A bias for shipping — startup pace energizes you rather than stresses you
- Strong problem-solving skills and a collaborative mindset; excellent communication within Agile teams
- Experience building or operating LLM infrastructure: inference gateways, eval/observability tooling, agentic orchestration
- Data-pipeline and streaming/CDC experience
- Security or identity background; familiarity with post-quantum cryptography or supply-chain security
- Prior experience at an early-stage startup
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring