RemoteHunt

Multilingual AI Quality Specialist

Spotify · Stockholm · posted Jul 14, 2026

mid

What this role actually asks for

Extracted by RemoteHunt

Must have

  • Multilingual AI data quality evaluation
  • Design and execute structured evaluations
  • Lead multilingual dataset curation
  • Analyze evaluation results and provide recommendations
  • Familiarity with LLMs and generative AI evaluation

Nice to have

  • Experience with recommendation systems
  • Experience with SQL or Python
  • Experience with annotation platforms

The full posting

About the Role

The Platform team creates the technology that enables Spotify to learn quickly and scale easily, enabling rapid growth in our users and our business around the globe. Spanning many disciplines, we work to make the business work; creating the infrastructure, tooling, frameworks, and capabilities needed to welcome a billion customers.

The Multilingual AI Data Quality space is part of Spotify's Global Language Quality Program within Localization. Bringing together multilingual language quality evaluation and AI data quality, the team helps ensure our AI-powered experiences are accurate, culturally relevant, and trustworthy across languages and markets. Working closely with Product, Engineering, Data Science, Research, Personalization, Localization, vendors, and market experts, the team develops evaluation methodologies, high-quality datasets, and quality signals that help improve AI experiences and inform product decisions at scale.

What You'll Do:

  • Define quality frameworks, evaluation rubrics, thresholds, and methodologies for multilingual AI experiences.
  • Design and execute structured evaluations for AI-generated, AI-translated, AI-curated, and recommendation-driven experiences.
  • Lead multilingual dataset curation, annotation, enrichment, and ground-truth creation to support AI model development and evaluation.
  • Analyze evaluation results, identify quality gaps, and provide actionable recommendations to improve multilingual AI quality.
  • Support LLM-as-a-judge workflows, evaluator calibration, and human-AI agreement studies.
  • Partner closely with Product, Engineering, Data Science, Research, Localization, vendors, and market experts to improve AI quality signals and inform launch decisions.
  • Document best practices and help define quality standards across languages, markets, and AI use cases.
  • Contribute to building scalable evaluation capabilities that support the next generation of AI-powered experiences across Spotify.

Who You Are:

  • You have experience in multilingual quality evaluation, localization, data curation, annotation, AI evaluation, or related fields, including text-to-text and text-to-speech experiences.
  • You understand language quality, cultural relevance, content quality, and user experience across multiple languages and markets.
  • You have experience designing or conducting structured evaluations using quality rubrics, audits, annotation projects, or review methodologies.
  • You are comfortable using qualitative and quantitative data to identify trends, measure quality, and make recommendations.
  • You are familiar with large language models (LLMs), generative AI evaluation, human-in-the-loop workflows, or LLM-as-a-judge methodologies.
  • You enjoy working through ambiguity and turning complex quality challenges into practical evaluation strategies.
  • You communicate effectively and thrive in highly cross-functional environments, collaborating with technical and non-technical partners alike.

Experience with recommendation systems, personalization, search, ranking, machine translation, generative AI, dataset creation, annotation operations, evaluator calibration, prompt testing, model evaluation, SQL, Python, dashboards, or annotation platforms is a plus.

Where You'll Be:

This role is based in London or Stockholm. We offer you the flexibility to work where you work best! There will be some in person meetings, but still allows for flexibility to work from home.

Is this one actually worth your time?

RemoteHunt scores every remote job 0–100 against your own resume, so you apply to the handful that fit instead of the hundred that don't. Free plan, no card required.