Level

Mercor

Research Intern

AI in this role

Research Scientist Intern focusing on post-training, reinforcement learning with verifiable rewards (RLVR), and large language model evaluation.

pythonpytorch
fine-tuningai-evaluationai-researchreinforcement-learningmachine-learningnlpllm-evaluation
About Mercor

Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.

 

Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.

About the Role

As a Research Scientist Intern at Mercor, you’ll work on research at the frontier of post-training, reinforcement learning with verifiable rewards (RLVR), data generation, and model evaluation.

You’ll investigate how datasets, rewards, and training methods affect the capabilities and behavior of large language models. This may include designing controlled experiments, developing new evaluation methodologies, conducting systematic failure analysis, and testing approaches to improve tool use, agentic behavior, and real-world reasoning.

You’ll work closely with research scientists, research engineers, and domain experts to turn open-ended questions into rigorous experiments. Your work will contribute to Mercor’s research agenda and may support external publications, benchmark releases, and the development of frontier AI systems.

What You’ll Do

  • Develop and investigate research questions related to post-training, RLVR, data quality, and model evaluation.

  • Design and run controlled experiments to understand how datasets, rewards, and training strategies affect model performance.

  • Study reward-shaping and post-training methods, including approaches such as GRPO and DAPO.

  • Develop methods for measuring data quality, usability, and performance uplift on key benchmarks.

  • Design and evaluate datasets, rubrics, evaluators, and scoring frameworks for complex model capabilities.

  • Conduct systematic error analysis to identify model failure modes and opportunities for improvement.

  • Analyze experimental results and communicate findings through clear reports, research artifacts, and presentations.

  • Build the research tooling and data pipelines needed to conduct experiments at scale.

  • Collaborate with research scientists, research engineers, applied AI teams, and domain experts producing training and evaluation data.

  • Contribute to research publications, benchmark releases, and other public research outputs where appropriate.

What We’re Looking For

  • Currently pursuing a master’s or PhD in computer science, machine learning, statistics, mathematics, or another relevant field.

  • Demonstrated ability to formulate research questions, design experiments, and draw sound conclusions from empirical results.

  • Demonstrated experience in at least one of the following:

    • Training, fine-tuning, or evaluating language models.

    • Agentic AI system, RL environments

    • Developing benchmarks, evaluation methodologies, or data-quality measures.

  • At least one publication or open source project.

  • Strong programming skills, particularly in Python, and the ability to write reliable research code.

  • Familiarity with machine learning fundamentals, experimental design, and statistical analysis.

  • Intellectual curiosity.

  • Comfort operating in a fast-paced research environment with rapid iteration and a high degree of ownership.

Nice to Have

  • Previous research experience in language models, reinforcement learning, model evaluation, or post-training.

  • Experience training, fine-tuning, or evaluating language models.

  • Familiarity with RLVR techniques, reward modeling, or agentic AI systems.

  • Experience developing benchmarks, evaluation methodologies, or data-quality measures.

  • Research publications or submissions at competitive CS conferences such as ACL, NeurIPS, ICML, ICLR, or EMNLP.

  • Research papers, technical reports, open-source projects, or other work samples demonstrating relevant skills.

Why Mercor

Impact: Your work powers how AI labs train and deploy their models

Learning: Get early exposure to frontier AI research and engineering

Growth: Work with a high-velocity team where interns ship to production

Benefits

  • Mentorship from experienced researchers.

  • Work on real, high-impact projects.

  • $1.5K monthly stipend for meals

  • $200 monthly laundry reimbursement

  • $200 monthly personal wellness reimbursement

  • Free Equinox membership

  • Team events and offsites.

  • Potential full-time return offer.

How we rate this

Research Intern at Mercor rates 95 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.

Classification

Builds AI. The job is building AI systems.

  1. ●●●● Builds AI80 to 100
  2. ●●●○ Works on AI60 to 79
  3. ●●○○ Uses AI40 to 59
  4. ●○○○ Little AI0 to 39

Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Fine TuningAI EvaluationAI ResearchReinforcement LearningMachine LearningNLPLLM EvaluationPython

Questions you could be asked

  1. Walk me through fine-tuning a model: what data did you use, and how did you check the result?
  2. How do you decide that one model's output is better than another's for a given task?
  3. Tell me about a research question you investigated. What did you find?
  4. Tell me about a project where reinforcement learning was part of your work. What did you do?
  5. Tell me about a project where machine learning was part of your work. What did you do?

Adapt your resume

  • List these exact terms on your resume: Fine Tuning, AI Evaluation, AI Research, Reinforcement Learning, and Machine Learning. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want an expert to read your CV for this job?

Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.

Get a free CV review

Get new AI jobs (Builds AI ●●●●) by email

One email a week with the new AI jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Research roles that build AI, at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Mercor

Related searches

Same AI level