Senior Software Engineer - Paris
AI in this role
Senior software engineer building and scaling the evaluation framework for computer-use AI agents at H.
About H:
When we released Holo4 on 28 September, we published every trajectory behind its public benchmark scores at trajectories.hcompany.ai. For OSWorld that's 369 desktop tasks, each run 3 times, with the steps, tokens and time of every attempt. This role builds and runs the evaluation framework that produces runs like these.
H builds computer-use agents and the models behind them. Developers use them through a managed API, and our forward deployed engineers take them into enterprise workflows.
What this team owns
The evaluation framework: orchestration, runtimes and observability. Researchers and forward deployed engineers bring the benchmarks, across web apps, desktop applications and the command line. Your job is to make the framework that runs them reliable, fast and cheap, and to make adding a new one quick. Research uses the results to choose checkpoints and decide whether a model ships. Product and the forward deployed engineers use them to measure agents on customer workflows. It carries roughly 50 benchmarks now. That number should be between 100 and 200 soon, and the framework has to keep up.
What you'd be doing
Integration support for researchers and forward deployed engineers bringing in a benchmark, with a shorter path each time.
Setting the standard for how a benchmark enters the framework, and building the checks that enforce it.
Scheduling and observability, so cluster capacity isn't left idle while evaluation jobs queue.
Reproducible results across trials, so a release decision rests on numbers that hold.
Whatever stack a benchmark calls for. One week that's cluster tuning; the next it's a browser extension or desktop environments.
Time with customers, from single developers to large companies, to find out what they want measured, then automating it so the results flow back into our harnesses and models.
The first few months
By 3 months you'll have helped researchers or forward deployed engineers integrate 5 benchmarks, and started fixing what slows the framework down. By 6 months one part of it is yours, for example scaling the runs, observability, or a group of related benchmarks, and a release will have gone out on your numbers. By 12 months you'll know the design and trade-offs of the whole evaluation system, and be the person the rest of H asks about evaluations.
Who you'd work with
Ceiran Chapman, our VP Engineering, is hiring for this role. You'd join the evaluation team. The people relying on your work day to day are H's researchers and forward deployed engineers.
What we think it takes
Likely a good fit if you
Have spent 5+ years in backend development, with production Python at the core, and use coding agents to go faster without letting quality drop.
Have built test, QA or evaluation tooling that other teams depended on, and care whether a number is right.
Have operated distributed systems on Kubernetes in a public cloud. AWS experience helps most.
Have built and shipped systems end to end, including APIs (REST or GraphQL) and integrations with outside services.
Know relational and non-relational databases, and message queues such as SQS, RabbitMQ or Kafka.
Instrument what you build, with metrics, tracing and monitoring from the start.
Stronger still if you have
Measured LLM quality before, or built agents yourself.
Packaged and run workloads in Docker and on virtual machines.
Used Temporal, Dask, FastAPI, PostgreSQL, Grafana or Datadog.
Automated web or desktop software with Playwright, Selenium or a browser extension you wrote.
Set standards other engineers follow, through code review, design review or mentoring.
You do not need a background in machine learning. We'll work that out with you. If you match most of this but not all of it, apply anyway.
How we hire
A 30 minute call with our Talent team, a 60 minute technical challenge, a 60 minute system design interview, and a 30 minute final conversation with Ceiran. About 3.5 hours in total.
Practicalities
Paris posting: Hybrid in Paris. That means 3 office days a week and a London trip about once every 4 to 6 weeks. There is a London posting for the same role. We offer a competitive package.
How we rate this
Senior Software Engineer - Paris at H Company rates 90 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Tell me about a project where software engineering was part of your work. What did you do?
- Tell me about a project where evaluation was part of your work. What did you do?
- Tell me about a project where observability was part of your work. What did you do?
- Tell me about a project where ci cd was part of your work. What did you do?
- Walk me through how you've used Python in your day-to-day work.
Adapt your resume
- List these exact terms on your resume: Software Engineering, Evaluation, Observability, Ci Cd, and Python. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new remote software engineer jobs (Builds AI ●●●●) by email
One email a week with the new remote software engineer jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step