Director, Model Research & Development
AI in this role
Direct model research and development for Thomson Reuters Labs, owning post-training pipelines, synthetic data, and LLM optimization.
Job Description
Thomson Reuters Labs is the global research and development division of Thomson Reuters. We bring together scientists, engineers, and domain experts to explore emerging technologies and develop compelling new AI products.
From research and prototyping to customer success, our work bridges frontier science with practical application. Our innovations help professionals make better decisions, faster β for example, with intelligent legal assistants, automated compliance workflows, or agentic tax engines.
We are always looking for creative, highly adaptable scientists and engineers who enjoy solving hard problems that matter.
Thomson Reuters Labs builds and ships custom models for professional workflows requiring fiduciary-grade AI. The Thomson LLM model family was developed by training leading open-weight models on decades of authoritative Thomson Reuters content, with thousands of hours of subject-matter-expert input. Owning the model layer is how we earn trust in the most regulated, highest-stakes professional workflows: it lets us tune for the accuracy, citation-grounding, and jurisdictional awareness our customers require, set our own training and evaluation priorities, and build the unit economics that let AI scale profitably across products used by more than a million professionals.
Reporting to the Vice President of Model Research & Development, the Director of Model Research & Development is responsible for hands-on technical direction of the research program, including the evolution of the post-training and data strategy that turns capable base models into accurate, well-calibrated, and trustworthy models for regulated use in professional workflows. Youβll work closely with the team designing and running experiments, building training and evaluation pipelines, and publishing findings at top tier conferences.
See the detailed technical report for more information on the latest model release: Thomson 1.0 Technical Report on Hugging Face.
Key Responsibilities
- Own the hands-on execution of post-training for Thomson's LLMs: supervised fine-tuning, preference optimization (e.g., DPO), and reinforcement learning, including RL in agentic, multi-step settings where models learn to use tools.
- Stand up and run the online, agentic reinforcement-learning pipeline, training with subject-matter experts in the loop.
- Own data selection, mixture optimization, synthetic-data generation, and evaluation design, and be able to point to a specific change and its measured effect on model behavior.
- Recognize when a training run is going wrong before it finishes, and know what to do about it.
- Work day-to-day with infrastructure and evaluation teams to keep training and eval pipelines reliable.
- Bring findings and recommendations forward to inform roadmap and prioritization decisions.
Required Qualifications
- Post-training and reinforcement learning depth. Hands-on command of supervised fine-tuning, preference optimization, and reinforcement learning, including RL in agentic, multi-step settings where models learn to use tools and complete tasks.
- Data-centric model development. You've personally shaped data selection, mixture design, synthetic-data generation, and evaluation design, and can point to a specific change and its measured effect.
- Engineering depth. Hands-on command of distributed training, data pipelines, and evaluation infrastructure, able to build and debug these systems yourself, not just direct others who do.
- Track record. Shipped models, widely-used open-source contributions, or peer-reviewed publications at top research venues (e.g., NeurIPS, ICML, ACL, EMNLP) that speak for themselves.
- Technical leadership. Experience leading a focused technical team (training and/or evaluation), even if smaller in scope than an org-wide leadership mandate.
Preferred Qualifications
- Experience applying LLMs in law, tax, or another comparably regulated professional domain, or clear evidence you can acquire that grounding quickly.
- Continued pre-training and domain adaptation of large base models.
- Experience with model safety, adversarial robustness, and red-teaming.
#LI-PPF
Β
Β
Whatβs in it For You?
- Hybrid Work Model: Weβve adopted a flexible hybrid working environment for our office-based roles while delivering a seamless experience that is digitally and physically connected.
- Flexibility & Work-Life Balance: Flex My Way is a set of supportive workplace policies designed to help manage personal and professional responsibilities, whether caring for family, giving back to the community, or finding time to refresh and reset. This builds upon our flexible work arrangements, including work from anywhere for up to 8 weeks per year, empowering employees to achieve a better work-life balance.
- Career Development and Growth: By fostering a culture of continuous learning and skill development, we prepare our talent to tackle tomorrowβs challenges and deliver real-world solutions. Our Grow My Way programming and skills-first approach ensures you have the tools and knowledge to grow, lead, and thrive in an AI-enabled future.
- Industry Competitive Benefits: We offer comprehensive benefit plans to include flexible vacation, two company-wide Mental Health Days off, access to the Headspace app, retirement savings, tuition reimbursement, employee incentive programs, and resources for mental, physical, and financial wellbeing.
- Culture: Globally recognized, award-winning reputation for inclusion and belonging, flexibility, work-life balance, and more. We live by our values: Obsess over our Customers, Compete to Win, Challenge (Y)our Thinking, Act Fast / Learn Fast, and Stronger Together.
- Social Impact: Make an impact in your community with our Social Impact Institute. We offer employees two paid volunteer days off annually and opportunities to get involved with pro-bono consulting projects and Environmental, Social, and Governance (ESG) initiatives.
- Making a Real-World Impact:β―We are one of the few companies globally that helps its customers pursue justice, truth, and transparency. Together, with the professionals and institutions we serve, we help uphold the rule of law, turn the wheels of commerce, catch bad actors, report the facts, and provide trusted, unbiased information to people all over the world.
DISCLAIMER
The above information in this description has been designed to indicate the general nature and level of work performed by employees within this classification. It is not designed to contain or be interpreted as a comprehensive inventory of all duties, responsibilities, and qualifications required of employees assigned to this job.
Β
Β
About Us
Thomson Reuters informs the way forward by bringing together the trusted content and technology that people and organizations need to make the right decisions. We serve professionals across legal, tax, accounting, compliance, government, and media. Our products combine highly specialized software and insights to empower professionals with the data, intelligence, and solutions needed to make informed decisions, and to help institutions in their pursuit of justice, truth, and transparency. Reuters, part of Thomson Reuters, is a world leading provider of trusted journalism and news.
We are powered by the talents of 26,000 employees across more than 70 countries, where everyone has a chance to contribute and grow professionally in flexible work environments. At a time when objectivity, accuracy, fairness, and transparency are under attack, we consider it our duty to pursue them. Sound exciting? Join us and help shape the industries that move society forward.
As a global business, we rely on the unique backgrounds, perspectives, and experiences of all employees to deliver on our business goals. To ensure we can do that, we seek talented, qualified employees in all our operations around the world regardless of race, color, sex/gender, including pregnancy, gender identity and expression, national origin, religion, sexual orientation, disability, age, marital status, citizen status, veteran status, or any other protected classification under applicable law. Thomson Reuters is proud to be an Equal Employment Opportunity Employer providing a drug-free workplace.
We also make reasonable accommodations for qualified individuals with disabilities and for sincerely held religious beliefs in accordance with applicable law. More information on requesting an accommodation here.
Learn more on how to protect yourself from fraudulent job postings here.
More information about Thomson Reuters can be found on thomsonreuters.com
How we rate this
Director, Model Research & Development at Thomson Reuters rates 100 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ββββ Builds AI80 to 100
- ββββ Works on AI60 to 79
- ββββ Uses AI40 to 59
- ββββ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Walk me through fine-tuning a model: what data did you use, and how did you check the result?
- Tell me about a project where machine learning was part of your work. What did you do?
- Tell me about a project where large language models was part of your work. What did you do?
- Tell me about a project where reinforcement learning was part of your work. What did you do?
- Tell me about a project where supervised fine tuning was part of your work. What did you do?
Adapt your resume
- List these exact terms on your resume: Fine Tuning, Machine Learning, Large Language Models, Reinforcement Learning, and Supervised Fine Tuning. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them β a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped β this role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new AI jobs (Builds AI ββββ) by email
One email a week with the new AI jobs (Builds AI ββββ), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Research roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step