← Back to job listings
NA
Senior Site Reliability Engineer
Navan · London, United Kingdom
About The Role
Join Navan, a fast-growing company revolutionizing travel and expense services for enterprises. As a Senior Site Reliability Engineer, you will design and develop tooling, automation, and infrastructure services that power Navan's services. You will work closely with development, release, productivity, and security teams to identify customer needs and build innovative solutions. You will also collaborate with backend and frontend engineering teams to ensure product solutions are scalable, efficient, and reliable.
- Design and develop tooling, automation, and infrastructure services that power Navan services.
- Collaborate with development teams, release and productivity teams, and security teams to identify customer needs and build innovative solutions.
- Drive the adoption of AI-assisted developer tools and platforms to increase engineering productivity, enforce code quality standards, and enable real-time architectural validation.
- Thrive in a fast-paced environment
- Hands-on operational experience with Java based applications and services including JVM profiling and performance tuning (python, Node.js and Go are a plus)
- Passionate about solving problems and learning new tools and technologies
- 2+ years of experience in working on a production, 24x7 product environment
- 5+ years of progressive experience as a Senior SRE or DevOps Lead (or equivalent role)
- Excellent communication skills working with stakeholders and domain experts across the company to design solutions to user problems
- Operate with a strong sense of ownership demonstrated through shipping production-quality code and infrastructure equipped with testing, monitoring and documentation
- Demonstrated experience mentoring and leading junior and mid-level engineers, and acting as a technical owner for cross-functional infrastructure projects
- Experience leveraging AI/LLM platforms (e.g., Gemini, Braintrust) and managing their secrets and infrastructure using Infrastructure as Code (Terraform) and AWS SSM
- Hands-on experience building and operating distributed systems in a public cloud environment (preferably AWS), using CI/CD to deploy, manage and operate production systems, focusing on tooling and automation using tools such as maven and Jenkins
- Built, using, and automating monitoring systems such as NewRelic, DataDog, SignalFX, Kibana,
- Demonstrated ability to integrate AI-specific telemetry and advanced observability practices to enable predictive insights and systemic root-cause analysis
- A passion for automating away everything, using scripting languages such as python, bash groovy (we prefer lazy engineers)
- Hands-on experience with writing Infrastructure as Code in Terraform or Cloudformation or similar tools
- Hands-on experience with microservice architecture and related reliability and resiliency patterns such as throttling, queueing, and retries
- Hands-on experience deploying, operating, and monitoring production-grade AI/ML microservices (e.g., RAG pipelines, agentic systems) on cloud platforms like AWS Fargate/ECS
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring