← Back to job listings
SO
Senior Site Reliability Engineer
Solaris · Berlin, Germany
About The Role
Join our team as a Senior Site Reliability Engineer. In this role, you will be responsible for designing monitoring and alerting standards, creating frameworks for SLIs, SLOs, and SLAs, and championing incident response and reliability best practices. You will also design processes to improve resilience and reduce technical complexity, implement software and tooling to enhance resilience and automate operations, and create monitoring dashboards. This position offers a range of benefits, including a learning and development budget, remote working allowance, mental healthcare service, and additional vacation days.
- Designing and implementing monitoring and alerting standards, frameworks for SLIs, SLOs, and SLAs.
- Championing standards for incident response and reliability best practices, and designing processes to improve resilience.
- Implementing software and tooling to improve resilience and automate operations, and creating monitoring dashboards.
- Experience in the participation of a 24/7 on-call rotation
- Business proficient English (written and verbal). German language is a plus
- Curious self-starter who combines openness and creativity with a structured, hands-on mentality
- A degree in Computer Science, Software Engineering, Information Technology or equivalent professional experience
- Designing infrastructure with a focus to limit the blast radius of any single failure
- Depending on your level of experience, your responsibilities and scope of role will range. We don’t care much about fancy titles, but rather about real personal and professional development, as laid out in our learning framework. Let’s figure together out how you can contribute to our team
- You have spent 6+ years in DevOps, SRE, or Software Engineering roles in a high-growth environment
- Brings an agile mindset along with the drive to get things done and the self-reflection to keep improving
- Designing systems that automatically failover across different data centers
- Safely rolling out significant platform changes across the entire organization
- Break down high-level architectural goals into small, deliverable tasks
- Understand the unique challenges of the financial sector
- Programming experience in Python, Ruby, Java, or Go
- Expertise in tools like Chaos Mesh or Gremlin to proactively test system weaknesses
- Expertise in Service Discovery / Service Mesh to manage microservice communications
- Evaluate third-party tools and act as technical point of contact
- Experience in the analysis of capacity requirements and address exhaustion before it impacts SLOs
- While job ads usually paint an ideal picture of a candidate, studies show that most applicants meet an average of 60% of the criteria
- Unfortunately, many promising candidates tend to apply only if they meet all the criteria. So if you think you have what it takes, but don't necessarily meet every single item in the job description, please contact us anyway. We'd love to talk with you and find out if you might be a good fit for us
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring