Principal AI Researcher

At Phrase, we help open the door to global business by providing the world’s leading Language Intelligence Platform. 

The Phrase Platform combines AI, agentic orchestration, and a headless, API-first architecture in one composable system. Beyond translation, it orchestrates and adapts content to culture, audience, channel, brand voice, and intended outcome in any language, for every audience. It applies the context that makes content perform in every market: quality standards, glossaries, prior translations, and cultural nuance. Every team, in every region, can ship content that is on-brand, on-point, and ready for any audience.

Phrase gives enterprises the intelligence to automate workflows, the freedom to connect their own tools and engines, and the control to govern global content at scale.

The AI Research team sits at the core of Phrase’s product differentiation. We build, train, and evaluate proprietary translation models; we design the agentic workflows that orchestrate translation quality end-to-end; and we run the research cadence that feeds product development with grounded, evidence-based decisions.

The team’s current research programmes span: fine-tuning LLMs for domain-specific machine translation, multi-agent evaluation pipelines, automated quality profiling from style guides, and active learning through feedback loops on real customer data. These are not exploratory proofs of concept; they are live systems serving production traffic.

The next phase requires someone who can take ownership of the most technically demanding of these streams, extend them, and help chart the research direction for the team as a whole.

What you’ll be responsible for:

Core Research

  • Lead design, training, and evaluation of LLMs for translation and language quality tasks, including work with fine-tuning techniques such as LoRA and DPO, and instruction-tuned models at various scales.
  • Design and implement robust evaluation frameworks for translation quality, moving beyond automatic metrics (BLEU, TER, ChrF, MQM, COMET), LLM-as-judge approaches, and hybrid evaluation pipelines.
  • Architect and evaluate complex agentic workflows for NLP problems: multi-step reasoning, tool use, structured output generation, and orchestration across multiple model providers.
  • Take technical ownership of the team’s flagship model development programme, from data curation and training pipeline design through to production integration and ongoing evaluation.
  • Design and run experiments using ML pipelines, maintaining reproducibility and clear documentation of results and decisions.

Applied NLP and Product Collaboration

  • Work closely with Product, Engineering, and Solutions to translate research findings into concrete, shippable capabilities.
  • Provide expert input on model provider strategy: evaluation of frontier models (OpenAI, Gemini, Claude, open-source), benchmarking against internal baselines, and recommendations on production use.
  • Contribute to the evolution of quality evaluation systems, including integration of style guides, and a focus on customer-led, outcome-oriented metrics and evaluation methodologies.
  • Support decisions on active learning, feedback loop design, and data pipeline strategy for continuous model improvement.

Team and Leadership

  • Act as a technical lead for the research team: setting direction on open problems, reviewing PRs and research write-ups, running or contributing to internal enablement sessions such as reading groups, deep dives etc.
  • Mentor junior and mid-level researchers; model good research hygiene and collaborative working practices.
  • Contribute to hiring decisions for the team, including take-home assignment design and technical interviews.
  • Represent the AI Research team in cross-functional forums and with external stakeholders where research credibility matters.

What We’re Looking For

Must have

  • Deep, hands-on experience building and training LLMs or large transformer-based models: pre-training, fine-tuning, RLHF/DPO, or equivalent.
  • Strong background in NLP, with demonstrable expertise in at least one of: machine translation, MT evaluation, information extraction, text generation, search optimization, or reinforcement learning.
  • Proven ability to design, implement, and critically evaluate agentic systems: tool-calling agents, multi-agent pipelines, or LLM-orchestrated workflows at scale.
  • Experience designing and running rigorous evaluation frameworks, including both automatic metrics and human-in-the-loop eval design.
  • Strong Python skills and experience with ML/NLP libraries (Hugging Face, PyTorch, vLLM, or similar); familiarity with workflow orchestration tools such as Flyte, Prefect, or Airflow.
  • PhD in Computer Science, Computational Linguistics, or a related field, OR equivalent research experience demonstrated through publications, open-source contributions, or significant applied research output.

Strongly Preferred

  • Industry experience working on production NLP systems, not just academic research: shipping models, handling data pipelines, monitoring quality in live traffic.
  • Familiarity with the localization and machine translation domain: translation quality metrics (MQM, COMET), post-editing workflows, segment-level quality signals, or TMS integrations.
  • Experience with multi-LLM ensemble architectures or model routing strategies.
  • Track record of publishing or presenting at top NLP venues (ACL, EMNLP, NAACL, AAAI, NeurIPS, ICLR).
  • Experience with experiment tracking tools (Weights & Biases), model serving frameworks (vLLM, TGI), and cloud-based ML infrastructure (AWS SageMaker, GCP Vertex AI).
  • Experience with PydanticAI, LangGraph, or similar frameworks for building structured agentic pipelines.
  • Prior experience in a technical lead or staff research role; evidence of mentoring or growing other researchers.

Character and Working Style

  • Product-minded: you think about research in terms of what it enables and who it serves, not just whether the numbers went up.
  • A natural driver: you identify open problems, propose directions, and move work forward without waiting to be pushed.
  • Collaborative and open: you share results early, invite challenge, give credit generously, and build on others’ work.
  • Precise but not precious: you care about rigour and reproducibility without losing sight of practical impact.
  • Growth-oriented: you are motivated by the prospect of building and leading a team, not just doing individual research.
  • Generative: you bring new ideas to the table, not just rigorous execution of known approaches; intellectual curiosity and a willingness to pursue directions that might not work are as important to us as technical depth.

 

What you’ll get:

  • Work experience in a successful and growing global SaaS company
  • Be part of an international team in Europe, APAC, and the Americas
  • Expert colleagues in their field who are determined to build the best localization platform on the market, creating a world where language never limits opportunity
  • An agile work environment, where it is encouraged to take smart risks
  • Take part in a culture full of trust, support and loyalty, where respectful and open feedback is valued, and diversity is fully embraced
  • A positive, open-minded, and innovative atmosphere
  • Support in your professional development and personal career goals

What’s on top:

  • 4 Company holidays additional to your regular holidays (1 day per quarter where the entire company is off to celebrate our achievements).
  • In addition, the company also provides employees with a Christmas break to allow you to spend time with family and friends without use of your vacation allocation. 
  • Your birthday is off because it is important to celebrate you as well.
  • 2 Giveback days where you can support the local community, volunteer, and/or participate in charity events and activities.
  • Professional and extensive onboarding.
  • Learning & development through Phrase learning 
  • Additional local benefits depending on the entity you’re hired at, just ask your Talent Acquisition Partner
 

At Phrase we believe in the critical importance of diversity in all its forms and intersectionalities, and are committed to ensuring that the people we interview reflect this diversity. We therefore strongly encourage people of any identity to apply for our exciting career opportunities. We have taken an active decision to make our work environment inclusive every day, starting at the very beginning of your Phrase experience.

We value and welcome different perspectives, experiences and backgrounds as we believe that these differences make our team even stronger on our mission of opening the door to global business by giving everybody access to the content they need in the language they speak.

 



応募プロセス

お問い合わせ

求人広告からご応募ください。

応募書類を際立たせる方法:

  • あなたと連絡をとる際に使用するべき適切な代名詞(he/him、she/her、they/their等)をぜひお知らせください。
  • 職務内容の説明を熟読し、ご自身の経験に一番合ったポジション(複数応募可)に応募してください。
  • なぜ私たちと一緒に働きたいのか、私たちのミッションに共感するかどうかを教えてください。
  • 私たちが知る必要のある情報がすべてそろっていることを確認してください(資格、強み、目標など)。履歴書、カバーレター、ポートフォリオ、ウェブサイト、仕事のサンプルといったものを提出できます。
  • 最新の連絡先をお知らせください。
  • 英語でご応募ください。弊社はグローバル企業ですが、会社の公用語は英語です。このようにして、関係者全員があなたのスキルや経験を十分に理解することができます。

面談

弊社が必要とする経験を持ち得ている場合は、面談のご連絡を致します。

まず、人材採用チームのメンバーが連絡し、あなたとあなたの経験について詳しく教えていただくためにビデオ通話の予定を組みます。同時に、私たちの事業、チーム、募集ポジションについてもっと知っていただきます。

それ以降の流れは、ご応募いただくチームやポジションのレベルによって若干異なりますが、概ね以下に示す通りです。

  • 直属の上司となる人に会っていただきます。
  • コーディングやデザインタスク、ロールプレイ、ビジネスケースといった何らかの形のスキル評価が行われます。
  • 主なステークホルダーやチームメンバーと会っていただきます。人材採用パートナーは、応募から採用までの一連のプロセスをサポートし、最新情報の提供、次に面談する相手や先行の進捗状況などをお伝えします。

考慮すべき事項は次の通りです。

  • 面接はすべてバーチャル形式で行われます。Google MeetとZoomを使用します。
  • スキルセットだけでなく、あなたの価値観や目標が私たちの価値観や目標と一致するかどうかについても評価します。

ジョブの確保

お互いの理解が深まったら、採用担当者全員で検討し、決断を下します。

  • 採否の判断は数日以内に行いたいと思います。同じポジションに複数の応募者がいる場合、採用結果が遅れることがありますが、状況は随時お伝えします。皆さまも、何かあれば随時お伝えいただけると幸いです。
  • ご質問や変更事項がありましたら、人材採用パートナーまでお知らせください。
  • 採用に至らない場合も、フィードバックをさせていただきます。同様に、私たちに対するフィードバックがありましたら、ぜひお聞かせください。プロセスの理解、改善に役立てたいと思います。