Senior Site Reliability Engineer
weekday-1 · Chennai, Tamil Nadu, India
About The Role
# Senior Site Reliability Engineer
> Weekday AI · Chennai, India · Full-time · Posted 2026-07-10
**Salary:** INR 3,000,000–4,500,000
**Workplace:** on_site
**Department:** Weekday's Client via platform
## Description
**This role is for one of the Weekday's clients**
**Salary range: Rs 3000000 - Rs 4500000 (ie INR 30-45 LPA)**
Min Experience: 7+ years
Location: Chennai
JobType: full-time
The Senior SRE is responsible for deployment, updates, and operational support for environments hosting our leading client’s cloud-based solutions. This role ensures operational excellence, a seamless client experience, and continuous improvement across infrastructure and delivery processes. The ideal candidate combines strong technical capabilities with the ability to lead delivery through influence and hands-on engineering expertise.
## Requirements
### **Key Responsibilities**
- Manage deployments, upgrades, maintenance, and operational support for cloud environments.
- Ensure high availability, scalability, performance, and reliability of production systems.
- Define, monitor, and improve SLAs, SLOs, and SLIs.
- Drive automation initiatives and Infrastructure as Code (IaC) adoption.
- Perform Root Cause Analysis (RCA) and implement preventive actions.
- Optimize cloud infrastructure, operational efficiency, and costs.
- Enhance monitoring, observability, security, and deployment processes.
- Collaborate with Engineering, Project Management, Customer Success, and cross-functional teams to deliver reliable services.
### Required Skills
- Strong hands-on experience with **AWS** cloud platforms.
- Expertise in **Kubernetes** for container orchestration and cluster management.
- Experience with **Terraform** for Infrastructure as Code (IaC).
- Proficiency in **Ansible** for configuration management and automation.
- Hands-on experience with **Helm** for Kubernetes application deployments.
- Experience managing **MariaDB** and **MongoDB** databases in production environments.
- Strong understanding of **CI/CD pipelines**, deployment automation, and DevOps practices.
- Experience with **monitoring, observability, logging, and alerting** tools (e.g., Prometheus, Grafana, ELK, CloudWatch, Azure Monitor).
- Good knowledge of **Linux administration, networking, DNS, load balancing, and cloud security** best practices.
- Experience troubleshooting production environments, conducting **Root Cause Analysis (RCA)**, and improving platform reliability.
- Understanding of **SRE principles**, including **SLAs, SLOs, and SLIs**.
- Scripting experience using **Bash, Python, or Shell** for automation.
- Excellent problem-solving, communication, and stakeholder management skills.
### Preferred Experience
- Experience working in Site Reliability Engineering, DevOps, or Cloud Operations roles.
- Experience managing large-scale, production cloud environments.
- Ability to thrive in a fast-paced, customer-focused environment.
- Strong analytical mindset with a proactive approach to continuous improvement.
### Must-have skills
AWS, Kubernetes, Site Reliability Engineering
### Good-to-have skills
Helm Charts, IaC, monitoring
## Apply
[Apply at Weekday AI](https://apply.workable.com/weekday-1/j/B1345C947C/apply)
---
Powered by [Workable](https://www.workable.com)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring