Principal AI Researcher

At Phrase, we help open the door to global business by providing the world’s leading Language Intelligence Platform. 

The Phrase Platform combines AI, agentic orchestration, and a headless, API-first architecture in one composable system. Beyond translation, it orchestrates and adapts content to culture, audience, channel, brand voice, and intended outcome in any language, for every audience. It applies the context that makes content perform in every market: quality standards, glossaries, prior translations, and cultural nuance. Every team, in every region, can ship content that is on-brand, on-point, and ready for any audience.

Phrase gives enterprises the intelligence to automate workflows, the freedom to connect their own tools and engines, and the control to govern global content at scale.

The AI Research team sits at the core of Phrase’s product differentiation. We build, train, and evaluate proprietary translation models; we design the agentic workflows that orchestrate translation quality end-to-end; and we run the research cadence that feeds product development with grounded, evidence-based decisions.

The team’s current research programmes span: fine-tuning LLMs for domain-specific machine translation, multi-agent evaluation pipelines, automated quality profiling from style guides, and active learning through feedback loops on real customer data. These are not exploratory proofs of concept; they are live systems serving production traffic.

The next phase requires someone who can take ownership of the most technically demanding of these streams, extend them, and help chart the research direction for the team as a whole.

What you’ll be responsible for:

Core Research

  • Lead design, training, and evaluation of LLMs for translation and language quality tasks, including work with fine-tuning techniques such as LoRA and DPO, and instruction-tuned models at various scales.
  • Design and implement robust evaluation frameworks for translation quality, moving beyond automatic metrics (BLEU, TER, ChrF, MQM, COMET), LLM-as-judge approaches, and hybrid evaluation pipelines.
  • Architect and evaluate complex agentic workflows for NLP problems: multi-step reasoning, tool use, structured output generation, and orchestration across multiple model providers.
  • Take technical ownership of the team’s flagship model development programme, from data curation and training pipeline design through to production integration and ongoing evaluation.
  • Design and run experiments using ML pipelines, maintaining reproducibility and clear documentation of results and decisions.

Applied NLP and Product Collaboration

  • Work closely with Product, Engineering, and Solutions to translate research findings into concrete, shippable capabilities.
  • Provide expert input on model provider strategy: evaluation of frontier models (OpenAI, Gemini, Claude, open-source), benchmarking against internal baselines, and recommendations on production use.
  • Contribute to the evolution of quality evaluation systems, including integration of style guides, and a focus on customer-led, outcome-oriented metrics and evaluation methodologies.
  • Support decisions on active learning, feedback loop design, and data pipeline strategy for continuous model improvement.

Team and Leadership

  • Act as a technical lead for the research team: setting direction on open problems, reviewing PRs and research write-ups, running or contributing to internal enablement sessions such as reading groups, deep dives etc.
  • Mentor junior and mid-level researchers; model good research hygiene and collaborative working practices.
  • Contribute to hiring decisions for the team, including take-home assignment design and technical interviews.
  • Represent the AI Research team in cross-functional forums and with external stakeholders where research credibility matters.

What We’re Looking For

Must have

  • Deep, hands-on experience building and training LLMs or large transformer-based models: pre-training, fine-tuning, RLHF/DPO, or equivalent.
  • Strong background in NLP, with demonstrable expertise in at least one of: machine translation, MT evaluation, information extraction, text generation, search optimization, or reinforcement learning.
  • Proven ability to design, implement, and critically evaluate agentic systems: tool-calling agents, multi-agent pipelines, or LLM-orchestrated workflows at scale.
  • Experience designing and running rigorous evaluation frameworks, including both automatic metrics and human-in-the-loop eval design.
  • Strong Python skills and experience with ML/NLP libraries (Hugging Face, PyTorch, vLLM, or similar); familiarity with workflow orchestration tools such as Flyte, Prefect, or Airflow.
  • PhD in Computer Science, Computational Linguistics, or a related field, OR equivalent research experience demonstrated through publications, open-source contributions, or significant applied research output.

Strongly Preferred

  • Industry experience working on production NLP systems, not just academic research: shipping models, handling data pipelines, monitoring quality in live traffic.
  • Familiarity with the localization and machine translation domain: translation quality metrics (MQM, COMET), post-editing workflows, segment-level quality signals, or TMS integrations.
  • Experience with multi-LLM ensemble architectures or model routing strategies.
  • Track record of publishing or presenting at top NLP venues (ACL, EMNLP, NAACL, AAAI, NeurIPS, ICLR).
  • Experience with experiment tracking tools (Weights & Biases), model serving frameworks (vLLM, TGI), and cloud-based ML infrastructure (AWS SageMaker, GCP Vertex AI).
  • Experience with PydanticAI, LangGraph, or similar frameworks for building structured agentic pipelines.
  • Prior experience in a technical lead or staff research role; evidence of mentoring or growing other researchers.

Character and Working Style

  • Product-minded: you think about research in terms of what it enables and who it serves, not just whether the numbers went up.
  • A natural driver: you identify open problems, propose directions, and move work forward without waiting to be pushed.
  • Collaborative and open: you share results early, invite challenge, give credit generously, and build on others’ work.
  • Precise but not precious: you care about rigour and reproducibility without losing sight of practical impact.
  • Growth-oriented: you are motivated by the prospect of building and leading a team, not just doing individual research.
  • Generative: you bring new ideas to the table, not just rigorous execution of known approaches; intellectual curiosity and a willingness to pursue directions that might not work are as important to us as technical depth.

 

What you’ll get:

  • Work experience in a successful and growing global SaaS company
  • Be part of an international team in Europe, APAC, and the Americas
  • Expert colleagues in their field who are determined to build the best localization platform on the market, creating a world where language never limits opportunity
  • An agile work environment, where it is encouraged to take smart risks
  • Take part in a culture full of trust, support and loyalty, where respectful and open feedback is valued, and diversity is fully embraced
  • A positive, open-minded, and innovative atmosphere
  • Support in your professional development and personal career goals

What’s on top:

  • 4 Company holidays additional to your regular holidays (1 day per quarter where the entire company is off to celebrate our achievements).
  • In addition, the company also provides employees with a Christmas break to allow you to spend time with family and friends without use of your vacation allocation. 
  • Your birthday is off because it is important to celebrate you as well.
  • 2 Giveback days where you can support the local community, volunteer, and/or participate in charity events and activities.
  • Professional and extensive onboarding.
  • Learning & development through Phrase learning 
  • Additional local benefits depending on the entity you’re hired at, just ask your Talent Acquisition Partner
 

At Phrase we believe in the critical importance of diversity in all its forms and intersectionalities, and are committed to ensuring that the people we interview reflect this diversity. We therefore strongly encourage people of any identity to apply for our exciting career opportunities. We have taken an active decision to make our work environment inclusive every day, starting at the very beginning of your Phrase experience.

We value and welcome different perspectives, experiences and backgrounds as we believe that these differences make our team even stronger on our mission of opening the door to global business by giving everybody access to the content they need in the language they speak.

 



我们的申请流程

联系我们

在招聘岗位页面提交申请。

申请小贴士:

  • 仔细阅读职位描述,选择投递与您的经历最为切合的职位。
  • 说说您为什么想要加入我们的团队,您是否认同我们的使命。
  • 确认您提交了我们需要了解的所有信息,包括但不限于您的资历、优势和理想。信息可以以简历、求职信、作品集或网站的形式展示。
  • 给我们留下您最新的联系方式。
  • 申请请用英文投递。我们虽然是一家跨国公司,但我们的工作语言是英文。用英文投递可以确保招聘环节中的每一个人都能充分了解您的信息。

与我们沟通

如果您的经历符合我们的需求,我们会与您联系。

首先,我们的招聘团队会与您联系,预约时间与您进行视屏通话,进一步了解您和您的经历。同时,我们也会进一步为您介绍我们的业务、团队和您应聘的岗位。

您应聘的团队和职位的级别会决定下一步具体的流程,但大体来说是这样的:

  • 认识未来的直属上司。
  • 进行能力测试,如要求您完成一项编程或设计任务、进行场景模拟或进行业务分析。
  • 认识未来主要的合作团队。招聘团队伙伴会全程引导您走完流程、沟通最新情况、告知您会面的对象和您当前的情况。

需要注意的事项:

  • 我们整个面试流程都是线上完成的,使用的工具是 Google Meet 和 Zoom。
  • 我们不仅要测试您的技能,还要考虑您的价值观和目标与我们是否相匹配。

拿下工作

双方互相了解之后,就该做决定了。

  • 我们希望尽快做出决定,但若有多位候选人同时应聘同一岗位,我们考虑的时间可能会延长。无论如何,我们都会随时与您沟通最新情况。我们也希望您能够及时与我们沟通您的最新情况。
  • 若您有任何疑问或情况有变,请随时与我们的招聘团队沟通。
  • 若我们决定拒绝您的申请,我们会给您反馈。如果您也对我们有任何反馈,那就太好了——我们乐于学习并不断优化流程。