← Back to job listings
AI
Senior Site Reliability Engineer (Forward Deployed, AWS & Databricks)
Abacus Insights · New York, United States
About The Role
Join Abacus Insights as a Senior Site Reliability Engineer. In this hands-on role, you will be responsible for production operations, incident response, and post-launch system reliability across our platform. You will work directly with customers on escalations and deployments, while ensuring learnings are translated back into durable product and platform improvements. You will troubleshoot and resolve complex issues, improve operational readiness, and contribute to reusable tools and best practices.
- Assumer la responsabilité des opérations de production, de la réponse aux incidents et de la fiabilité des systèmes après le lancement sur la plateforme d'Abacus Insights.
- Diriger la triage des incidents en temps réel, les efforts d'atténuation et de récupération, et conduire l'analyse des causes profondes avec un accent sur les corrections systémiques à long terme.
- Collaborer étroitement avec les équipes Produit, Ingénierie, Données et Client pour résoudre les problèmes techniques en production et améliorer la fiabilité et l'opérabilité des systèmes.
- Proficiency in Python and experience building production services or tooling
- Excellent communication skills, especially during incidents and customer escalations
- 10+ years of experience in software engineering, SRE, sustaining engineering, or production operations
- Distributed systems
- CI/CD Pipelines that leverage Infrastructure as Code
- Ability to work backward from customer impact to root cause across systems and codebases, delivering fixes in environments with minimal documentation
- Strong experience troubleshooting Databricks and large-scale data platforms
- Incident management and RCA practices
- Monitoring, alerting, and observability
- Strong instinct for operational risk, with the ability to proactively identify failure modes and harden systems before they impact customers
- Proven ability to own problems end-to-end, from detection to permanent resolution
- Deep hands-on experience operating production systems in AWS
- Strong understanding of:
- Experience in healthcare, health insurance, or regulated data environments
- Familiarity with:
- Kubernetes (EKS), EMR, Lambda
- Spark internals
- Snowflake or similar data warehouses
- Experience with FHIR, MDM systems, or entity resolution
- Experience contributing to or operating within SRE/on-call programs
- Prior experience in SWAT, escalation engineering, or tiger-team roles
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring