Level

Mechanize

Software Engineer

AI in this role

About Mechanize

Mechanize builds reinforcement learning environments that frontier AI labs use to train and evaluate their coding models. Learn more at mechanize.work.

Why the work matters

AI models have gotten good at narrow coding tasks but still fail at the complex, judgment-heavy parts of software engineering. We build the environments that expose those failures and help models improve.

What you'll do

You'll design, build, and quality-assure RL tasks. Each task is a self-contained software engineering challenge with a prompt, an environment, and an automated grader. You own the full lifecycle: ideation, grading infrastructure, running frontier models against the task, failure analysis, and iteration. At this level, we expect you to consistently produce tasks that target meaningful capability gaps in frontier models, and to develop a strong sense for what makes a task informative versus merely difficult.

You will use coding agents heavily, and a large part of the job is directing them well, evaluating their output, and knowing when they are failing in subtle ways. You may also contribute to shared infrastructure: improving our build pipeline, automating parts of QA, or building tooling for other engineers.

What makes someone good at this

Strong technical fundamentals combined with a well-calibrated intuition for AI model behavior. You need to anticipate where a model will take shortcuts, distinguish genuine capability gaps from grader issues, and understand how a model will interpret a prompt. At this level, we expect extensive familiarity with what frontier coding agents can and can't do.

Good fit if you:

  • Can code in Python

  • Are confident working independently at a consistent pace

  • Have developed an intuition for what coding agents can and can't do

  • No prior ML or AI experience required

Probably not a good fit if you:

  • Want a product engineering role building features for end users

  • Prefer a highly collaborative team environment with shared ownership

  • Want extensive structured mentorship

This is independent, high-ownership work. You own your tasks from start to finish, with regular check-ins and feedback.

Compensation

Compensation includes a $350,000 base salary, equity, and performance bonuses. Top performers can earn more in bonuses than in base salary.

Strong performers are recognized and promoted quickly. Benefits include health, dental, vision, and life insurance.

About Mechanize. ~20 person team in San Francisco. Backed by Patrick Collison, Nat Friedman, Daniel Gross, Jeff Dean, Dwarkesh Patel, and Sholto Douglas. Featured in the New York Times, the Dwarkesh Podcast and Hard Fork.

Learn more about the interview process: https://www.mechanize.work/how-our-interview-process-works

Learn more about the work: https://www.mechanize.work/what-working-here-is-like

How we score this

Software Engineer at Mechanize scores 86 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.

Classification

AI Level 4. Building AI systems is the job itself: without AI, the role would not exist.

  1. AI Level 480 to 100
  2. AI Level 360 to 79
  3. AI Level 240 to 59
  4. AI Level 10 to 39

Bands come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Get new AI jobs at AI Level 4+ by email

One email a week with the new AI jobs at AI Level 4+, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Software Engineering roles rated AI Level 4 at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Mechanize

Related searches

Same AI level