Skip to content
← Back to job listings

Systems Development Engineer (SRE/DevOps)

Model N · Hyderabad India, India

Software DevelopmentExternal listingfull-timeabout 15 hours ago

About The Role

Job Responsibilities

  • Design, build, and maintain automated CI/CD pipelines using tools such as Harness, GitHub Actions, and ArgoCD.
  • Develop and maintain infrastructure and configuration as code using CloudFormation, Terraform, Ansible , and related automation tools.
  • Administer and optimize AWS environments , including core services, networking, security, and architecture for availability, performance, and cost.
  • Manage and support Kubernetes clusters and containerized workloads, including configuration, scaling, and upgrades.
  • Design, implement, and evolve end‑to‑end monitoring and observability frameworks using tools such as Open Telemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic , or similar platforms.
  • Create and maintain dashboards, logs, traces, SLIs/SLOs, and automated alerting systems to ensure reliability and rapid detection of anomalies.
  • Embed observability, CI/CD best practices, and operational readiness into all stages of the software development lifecycle in partnership with engineering teams.
  • Lead or participate in incident response, troubleshooting, and root cause analysis for production incidents, using observability data to drive fast resolution.
  • Automate operational tasks, runbooks, and incident remediation workflows to reduce toil and improve service reliability.
  • Contribute to risk mitigation, backup, and disaster recovery strategies, including periodic testing and continuous improvement.
  • Participate in shared after hours support and project work as needed.

Job Qualification

  • 2-4 years of experience designing, implementing, and maintaining CI/CD pipelines (e.g., Harness, GitHub Actions, ArgoCD or similar tools).
  • Hands‑on experience with automation tools and Infrastructure as Code / Configuration as Code ( CloudFormation, Terraform, Ansible).
  • Strong understanding of Infrastructure as Code and Configuration as Code principles and patterns.
  • Solid grasp of the software development lifecycle and modern SRE/DevOps practices.
  • AWS administration and architecture experience , including networking, security, IAM, and core services.
  • Experience operating Kubernetes clusters (EKS or other distributions) and containerized workloads.
  • Deep experience with monitoring and observability tools such as OpenTelemetry, Groundcover, CloudWatch, Datadog, Prometheus, New Relic, or equivalent, i ncluding metrics, logs, and traces.
  • Ability to define and track SLIs/SLOs and use them to guide reliability improvements.
  • Proficiency in Linux administrati on, including system configuration, troubleshooting, and performance tuning.
  • Programming/scripting skills in at least one language such as Python, Go, or Rust for automation , tooling, and observability integrations.
  • Solid understanding of networking, load balancing, and performance tuning.
  • Experience troubleshooting complex distributed systems, supporting incident response, and driving root cause analysis.
  • Familiarity with risk mitigation, backup, and disaster recovery concepts.

Preferred

  • Experience building unified observability platforms or standardized dashboard s for multiple services/teams.
  • Experience with GitOps workflows and tools for declarative infrastructure and application delivery.
  • Background in incident command and post‑mortem frameworks.
  • Experience integrating observability and reliability practices into microservices and/or serverless architectures.
  • Experience integrating testing, security and compliance checks into CI/CD pipelines.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing