← Back to job listings
FE
Staff Reliability Engineer (Full Stack)
Feeld · London, United Kingdom
About The Role
Join Feeld, an inclusive and human-centered product focused on building intimate connections. As a Staff Reliability Engineer (Full Stack), you will improve the reliability and operability of our production systems, lead technical decisions, and strengthen engineering practices. This is a hands-on role with significant cross-team influence, and you will collaborate with various squads to enhance production ownership and reliability. Enjoy flexible working hours, unlimited paid time off, and a supportive remote-first work environment.
- Lead technical problem-solving during incidents, coordinating response, diagnosing root causes, communicating status, and driving to resolution.
- Build and evolve monitoring/observability (dashboards, alerts, tracing, logging) that enables fast detection and diagnosis.
- Drive post-incident reviews (blameless) and ensure learnings become durable fixes (tech changes, runbooks, automation, process updates).
- Demonstrated Staff-level IC leadership: influence through design reviews, technical direction, documentation, and cross-team alignment
- Experience collaborating with mobile teams and understanding mobile↔backend integration concerns (e.g., API compatibility, releases, feature flags)
- Proven incident response leadership: on-call participation, triage, mitigation, and root-cause analysis (RCA) with follow-through
- Significant experience building and operating production backend systems at scale, including debugging distributed systems and performance issues
- Solid observability skills: practical experience with logging/metrics/tracing and turning signals into actionable alerts and dashboards
- Strong TypeScript/Node.js (or equivalent) backend experience; comfort working across services and APIs
- React Native experience and/or strong understanding of mobile architecture patterns and release constraints
- AWS (or similar cloud) experience and familiarity with infrastructure-as-code, CI/CD, and production tooling
- Experience designing reliability programs (SLOs, error budgets, incident process) and running operational excellence improvements
- Experience with PostgreSQL / Redis and performance tuning in high-traffic systems
- Experience in a high-growth environment where prioritization and pragmatic trade-offs are essential
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring