Skip to content
← Back to job listings

User Researcher (AI Evaluations)

Notion · United States

External listingfull-timeabout 1 month ago

About The Role

Join Notion as a User Researcher focused on AI evaluations. In this role, you will define and scale the evaluation of Notion's AI-powered experiences, establish clear evaluation criteria, run recurring evaluations, anchor evaluation in real workflows, identify failure modes and recovery behavior, and operationalize evaluation with partners. You should have 5+ years of UX research experience, AI fluency, and a Master's or PhD in a related field.

  • Definir y escalar la evaluación de las experiencias impulsadas por IA de Notion, centrándose en lo que significa "bueno" tanto para la calidad de la salida del modelo como para la experiencia del producto de extremo a extremo.
  • Ejecutar estudios que descubran los modelos mentales de los usuarios, las expectativas y los comportamientos de falla/recuperación, y traducir esos conocimientos en rubricas reutilizables, flujos de trabajo y enfoques de medición.
  • Establecer criterios de evaluación claros y reutilizables que reflejen las expectativas reales de los usuarios, y traducir la información cualitativa en orientación de puntuación que se pueda aplicar de manera consistente en todos los equipos y a lo largo del tiempo.
  • Strong UX research craft (quant + qual): You can choose the right methods for the question— interviews, benchmarking, surveys, experiments—and synthesize into actionable guidance. You also can prioritize ruthlessly, work through ambiguity, and balance scrappy iteration with deep dives when needed
  • Clear communication and impact orientation: You can align diverse partners around shared definitions of quality and create artifacts that enable teams to act consistently. You tailor storytelling to different audiences, connect research to business outcomes, and drive follow-through so insights translate into product change
  • Experience: 5+ years doing UX research in industry
  • Ability to operationalize insight into measurement: You’re comfortable turning “soft” user expectations (trust, tone, usefulness, clarity) into concrete rubrics, scoring guidelines, and observable metrics
  • AI fluency and systems thinking: You’re curious and hands-on with AI products, and can reason about how model behavior, uncertainty, and system constraints shape user experience. You also have experience evaluating AI-enabled products (LLMs, agents, generative UI/workflow automation) and working with Data Science/ML partners on measurement strategy and evaluation tooling
  • Pragmatism in fast-moving environments: You can prioritize ruthlessly, work through ambiguity, and balance scrappy iteration with deep dives when needed
  • Master’s or PhD in HCI, Psychology, Behavioral Science, Anthropology, Sociology, or a related field
  • Experience using AI research tooling for rapid synthesis and communication (e.g., Dovetail, Listen Labs, Maze, Outset, etc.), as well as AI observability tooling like Braintrust
  • Familiarity with LLM-as-judge methods, prompt design for evaluators, or “golden dataset” creation
  • We hire talented people from a wide range of backgrounds. If you’re excited about this role but don’t meet every bullet, we still encourage you to apply
  • You’re familiar with the work of computing heroes like Douglas Engelbart, Alan Kay, Bret Victor, etc. — and understand why we're big fans
  • Experience using data querying languages (e.g., SQL), scripting languages (e.g., Python), or statistical/mathematical software (e.g., R, SAS, Matlab, etc.)

This is an external listing. JobSpring does not represent or verify the employer. Report this listing