Skip to content
← Back to job listings

Multilingual AI Quality Specialist

Spotify · Remote, Sweden

RemoteExternal listingpermanent13 days ago

About The Role

What You'll Do

  • Define quality frameworks, evaluation rubrics, thresholds, and methodologies for multilingual AI experiences.
  • Design and execute structured evaluations for AI-generated, AI-translated, AI-curated, and recommendation-driven experiences.
  • Lead multilingual dataset curation, annotation, enrichment, and ground-truth creation to support AI model development and evaluation.
  • Analyze evaluation results, identify quality gaps, and provide actionable recommendations to improve multilingual AI quality.
  • Support LLM-as-a-judge workflows, evaluator calibration, and human-AI agreement studies.
  • Partner closely with Product, Engineering, Data Science, Research, Localization, vendors, and market experts to improve AI quality signals and inform launch decisions.
  • Document best practices and help define quality standards across languages, markets, and AI use cases.
  • Contribute to building scalable evaluation capabilities that support the next generation of AI-powered experiences across Spotify.

Who You Are

  • You have experience in multilingual quality evaluation, localization, data curation, annotation, AI evaluation, or related fields, including text-to-text and text-to-speech experiences.
  • You understand language quality, cultural relevance, content quality, and user experience across multiple languages and markets.
  • You have experience designing or conducting structured evaluations using quality rubrics, audits, annotation projects, or review methodologies.
  • You are comfortable using qualitative and quantitative data to identify trends, measure quality, and make recommendations.
  • You are familiar with large language models (LLMs), generative AI evaluation, human-in-the-loop workflows, or LLM-as-a-judge methodologies.
  • You enjoy working through ambiguity and turning complex quality challenges into practical evaluation strategies.
  • You communicate effectively and thrive in highly cross-functional environments, collaborating with technical and non-technical partners alike.
  • Experience with recommendation systems, personalization, search, ranking, machine translation, generative AI, dataset creation, annotation operations, evaluator calibration, prompt testing, model evaluation, SQL, Python, dashboards, or annotation platforms is a plus.

Where You'll Be

  • This role is based in London or Stockholm.
  • We offer you the flexibility to work where you work best! There will be some in person meetings, but still allows for flexibility to work from home.

This is an external listing. JobSpring does not represent or verify the employer. Report this listing