Level

Mistral AI

Research Engineer - Eval Platform

Mistral AI is hiring a Research Engineer - Eval Platform in Paris, France. Level rates it ; you can apply on Level.

AI in this role

vllm

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector, co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

Evaluation is how we decide which models, checkpoints and recipes ship. As a Research Engineer on the Eval Platform team, you will build the infrastructure every science team relies on to measure model quality, and make it reliable, reproducible and fast.

You don't need to have designed benchmarks before. You do need to care about what a score means, and about when a difference between two runs is real.

What you will do

  • Build systems that keep eval results reproducible and comparable over time, as models, benchmarks and code evolve.

  • Run evaluations at scale across our GPU clusters, from model serving to scoring.

  • Make eval results easy to access, explore and trust, through APIs and dashboards that researchers use every day.

  • Catch broken or noisy evals before they mislead research decisions.

  • Support evaluation of agentic, multi-turn and tool-using models.

  • Work closely with researchers to turn new evaluation needs into robust, shared tooling.

What we're looking for

  • Master's or PhD in Computer Science, or equivalent experience.

  • 4+ years building production-grade software, ideally large-scale ML codebases or distributed systems.

  • Excellent Python and strong software-design instincts: testing, code review, CI/CD.

  • Experience running workloads on GPU clusters (Slurm, Kubernetes, Ray or similar).

  • Familiarity with LLM inference and evaluation.

  • A product mindset: researchers are your users.

  • Self-starter, low-ego, collaborative.

Nice to have

  • Experience building or maintaining evaluation harnesses or benchmarks.

  • Hands-on experience with inference engines such as vLLM or SGLang.

  • Experience with agentic or RL environments.

  • Statistics for experimentation: variance estimation, significance testing.

  • Open-source contributions to ML tooling.

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Privacy Policy

Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.

How we rate this

Research Engineer - Eval Platform at Mistral AI rates 99 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.

Classification

Builds AI. The job is building AI systems.

  1. ●●●● Builds AI80 to 100
  2. ●●●○ Works on AI60 to 79
  3. ●●○○ Uses AI40 to 59
  4. ●○○○ Little AI0 to 39

Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

vLLM

Questions you could be asked

  1. What's a project where you used vLLM hands-on?
  2. How would you decide a model or AI system is ready to ship?
  3. Tell me about a time a model underperformed in production. How did you find out, and what did you change?

Adapt your resume

  • List these exact terms on your resume: vLLM. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want an expert to read your CV for this job?

Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.

Get new research jobs (Builds AI ●●●●) by email

One email a week with the new research jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Research roles that build AI, at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Mistral AI

Related searches

Same AI level

Jobs by city