Skip to content
← Back to job listings

Senior Site Reliability Engineer (SRE / Backend)

nilo.health · Berlin, Germany

External listingfull-time7 days ago

About The Role

Join our team as a Senior Site Reliability Engineer (SRE / Backend) and take ownership of the infrastructure behind our platform. This hands-on role involves managing our AWS infrastructure, building and maintaining CI/CD pipelines, designing event-driven services for resilience, and contributing to backend development. You will also lead incident response, harden our security posture, and monitor and optimize AWS spend. Enjoy benefits such as equity options, 30 days of vacation, and the flexibility to work remotely or in-office.

  • Take ownership of the infrastructure behind the platform, managing AWS services and ensuring reliability and security.
  • Build and maintain CI/CD pipelines, design event-driven services for resilience, and lead incident response efforts.
  • Contribute to backend development, participate in architectural discussions, and mentor engineers on operational excellence.
  • Deep AWS experience across serverless and containers — Lambda, ECS Fargate, and debugging both under pressure
  • Effective communicator who can explain a technical tradeoff without jargon
  • Pragmatic about complexity — you reach for the simplest thing that meets the reliability bar
  • 5+ years in SRE, DevOps, platform, or backend engineering, with real production ownership
  • Production experience with Datadog or a comparable observability platform (Grafana, New Relic, Honeycomb)
  • Strong security fundamentals: IAM, least privilege, secrets, network isolation, common web vulnerabilities
  • Solid PostgreSQL: query tuning, indexing, connection management, and zero-downtime migrations
  • Genuine on-call and incident response experience — you've led an incident and written the postmortem
  • Strong Terraform skills, including module design and managing state across multiple environments
  • Hands-on experience with event-driven architecture (SQS, SNS, EventBridge) and a healthy respect for its failure modes
  • Comfortable writing production backend code in [Python / Node.js / Go]
  • AWS DMS or other data migration and replication tooling
  • Compliance experience (GDPR, SOC 2, ISO 27001)
  • Experience in a B2B SaaS environment or healthcare-related product
  • We would like to encourage you to apply even if the technical requirements cannot be met 100%

This is an external listing. JobSpring does not represent or verify the employer. Report this listing