Level

MirendilPosted 3mo ago

Member of Technical Staff, Model Evaluation

Member of Technical Staff, Model Evaluation at Mirendil scores 97 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.

San FranciscoleadFullTime$300k

AI in this role

ai-evaluationai-research

Mirendil

Mirendil is a tech-first company focused on solving core bottlenecks that unlock step-change acceleration across science and technology. Our first goal is to democratize frontier AI R&D across scientific disciplines. We are building a frontier AI research company and training our own models end-to-end.

The Role

We are looking for a research engineer to build the evaluation infrastructure that tells us whether our models are getting better in ways we care about. You'll own the frameworks, pipelines, and tooling that measure model behavior across capabilities. Some example areas you might work on (not limited to):

  • Design and build evaluation frameworks that measure model capabilities along realistic axes, beyond standard benchmarks.

  • Build automated eval pipelines and regression-detection systems that run continuously and surface signal quickly.

  • Develop agent-assisted workflows for humans to efficiently inspect model behavior.

  • Instrument training runs with observability tooling so researchers can understand what's changing in model behavior, and why.

  • Partner with post-training and RL teams to close the loop between eval signal and training decisions.

If you're excited about the hard problem of knowing whether a frontier AI system is actually improving, we'd love to hear from you.

We offer a base salary of $300,000–$400,000 USD and a meaningful equity grant, depending on experience and background, along with competitive benefits.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

AI EvaluationAI Research

Questions you could be asked

  1. How do you decide that one model's output is better than another's for a given task?
  2. Tell me about a research question you investigated. What did you find?
  3. How would you decide a model or AI system is ready to ship?
  4. Tell me about a time a model underperformed in production. How did you find out, and what did you change?

Adapt your resume

  • List these exact terms on your resume: AI Evaluation and AI Research. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Similar roles

Other roles rated AI Level 4 at other companies.

More jobs at Mirendil

Related searches

Same AI level