Skip to content
← Back to job listings

Staff Back End Engineer (Evals, Hazel AI)

Altruist · San Francisco, United States

External listingfull-time2 months ago

About The Role

Join Altruist, a fast-growing company dedicated to making financial advice better, more affordable, and accessible to all. As a Staff Back End Engineer, you will architect our evaluation platform, design and build Hazel's evals platform, and develop LLM verification agents. You will work closely with backend engineers, product managers, and subject matter experts to translate fiduciary-grade requirements into automated quality signals. This is an opportunity to make a significant impact in the financial services industry while growing your career in a supportive and inclusive environment.

  • Architect the evaluation platform from first principles, focusing on observability, scoring, golden datasets, verification agents, and CI/CD integration.
  • Design and build Hazel's evaluation platform end-to-end, including online scoring, offline benchmarks, regression suites, and human-in-the-loop review workflows.
  • Build production observability and monitoring for AI quality, including hallucination rates, factual accuracy, refusal behavior, latency, cost, and domain-specific quality signals.
  • We're looking for exceptional talent to help us achieve our mission of making financial advice better, more affordable, and accessible to all. If you're passionate about challenging the status quo and want to do the most important work of your life, we'd love to meet you!
  • A bias toward shipping. You believe great evals enable speed, not just safety, and you build tools that engineers actually want to use
  • Comfort working across the stack – data engineering (SQL, dbt, warehouses), backend integration (APIs, async pipelines, queues), and observability tooling
  • 8+ years of engineering experience, with at least 2 years focused on evaluation infrastructure, model quality, fine-tuning, or ML platform work for production systems
  • Strong communication skills. You can translate fuzzy domain requirements from advisors and SMEs into precise, measurable, automatable eval criteria – and explain quality tradeoffs clearly to engineers, product managers, and leadership
  • Deep familiarity with evaluation and scoring methodologies for modern AI systems – RAG evaluation, document processing, fine-tuned model assessment, agentic and tool-use system evaluation, LLM-as-judge frameworks, and human evaluation protocols
  • Experience designing and curating golden datasets – sampling strategies, inter-rater agreement, dataset versioning, and managing the long tail of edge cases
  • Familiarity with frameworks like Braintrust, Langfuse or similar — including a clear point of view on when to use which
  • Prior experience at an applied AI company building evals, model quality, or applied research infrastructure
  • Experience evaluating multi-step agentic workflows, tool-use systems, or RAG pipelines in production
  • Domain knowledge of wealth management, tax planning, or financial planning — or genuine excitement to learn it deeply alongside our SME bench
  • Experience building human-in-the-loop labeling workflows, annotation tooling, or red-teaming programs
  • Background in regulated industries (financial services, healthcare, legal) where accuracy, auditability, and the cost of a wrong answer are unusually high

This is an external listing. JobSpring does not represent or verify the employer. Report this listing