ML Engineer, Infrastructure
AI in this role
Who we are
Foundation models transformed text and images. Structured data - the largest and most consequential data format in the world - stayed untouched, until now. What LLMs did for language, we're doing for tables.
We pioneered tabular foundation models: TabPFN v2 was a Nature cover story, has passed 3.5M+ downloads and 7,500+ GitHub stars, and runs in production from detecting lung disease with Oxford Cancer Analytics to preventing train failures with Hitachi. The hardest problems - millions of rows, real-time inference, entirely new modalities - are still open, and no one else is working on them at this level.
We're a small, highly selective team of 40+ with backgrounds from Google, DeepMind, Meta, Apple, Amazon, Jane Street, and CERN, led by Frank Hutter, Noah Hollmann, and Sauraj Gambhir, and advised by Bernhard Schölkopf and Turing Award winner Yann LeCun.
In July 2026, less than 18 months after our €9M pre-seed, we joined SAP as an independent frontier AI lab - same team, mission, and open-weights models, now backed by more than €1 billion over four years.
About the Role
We spend tens of millions per year on GPU compute to train tabular foundation models. That's not a target, it's what we're running today, and it's growing. The person who owns this infrastructure makes decisions worth millions of dollars: cluster architecture, scheduling efficiency, provider strategy, hardware selection. A wrong call costs six figures.
Today we run Slurm on GCP across multiple clusters. We're scaling to multi-cluster, multi-provider infrastructure and evaluating new hardware generations as they come online. You own the full stack, from cluster operations and cost optimization to distributed training performance and the tooling layer that keeps researchers moving fast. You work directly with the research team and understand what they're doing well enough to make infrastructure decisions that actually help them. And this isn't a pure support role. We operate an open environment. If you've got the next SOTA tabular architecture up your sleeve, go ahead and train it.
What you'll work on:
Own and evolve multi-cluster GPU infrastructure. Slurm on GCP today, multi-provider and new hardware tomorrow. Architecture, scheduling, reliability, cost optimization
Drive GPU utilization and training throughput: profiling, memory optimization, communication bottlenecks, systems-level debugging of distributed training across large runs
Architect the next generation of our infrastructure: multi-cluster orchestration, new GPU generations, provider diversification, capacity planning against growing compute demands
Build the developer productivity layer: CI pipelines, experiment tracking, model registry, data processing, and internal tooling that keeps research iteration speed high
Own the compute budget. You understand cost per FLOP across providers and hardware, and you hate wasted compute
Tech stack: Slurm, GCP, Docker, wandb, GitHub Actions, uv, PyTorch, Triton
You may be a good fit if you have:
3+ years building and operating production GPU infrastructure or distributed training systems at scale. At a major AI lab, a well-funded ML startup, or an HPC environment
Deep hands-on experience with Slurm and cluster management. You've debugged scheduling failures, optimized utilization across multi-tenant GPU workloads, and operated infrastructure where downtime has real cost
Expert-level systems thinking: memory bandwidth, GPU profiling. You reason about hardware, not configs
Strong Python and genuine fluency with PyTorch internals. Enough to profile a training run and tell whether the bottleneck is data loading, communication, or compute
Track record of making infrastructure decisions that measurably improved training throughput or cost efficiency
Strong AI tooling skills. You use Claude Code, Cursor, or similar fluently to move fast without sacrificing quality
Bonus:
Experience operating at tens-of-millions-scale GPU spend
Multi-cloud or hybrid HPC/cloud infrastructure experience
Triton, CUDA, or custom kernel experience
Experience scaling from single cluster to multi-cluster orchestration
Background building experiment tracking, model registry, or ML pipeline tooling
Life at Prior Labs
You'll work alongside researchers and builders who hold themselves to a very high bar - in the quality of their work and in how they work with each other. We move fast and still take the time to do things right.
Our teams are based in Berlin, Freiburg, and New York - when you're working on something as hard as TabPFN, being in the same room matters. But great people come from everywhere, and in exceptional cases we're open to remote, which usually means frequent travel to one of our offices. Wherever you're based, the whole company comes together regularly for offsites to build and celebrate together.
Our Commitments
The best products and teams are built by people with a wide range of perspectives and backgrounds. We welcome applications from all identities and walks of life - especially if you've ever felt discouraged by "not checking every box" - and provide equal opportunities regardless of gender, sexual orientation, origin, disability, or any other trait that makes you who you are.
We care about how your data is handled - see our Recruiting Data Privacy page
How we score this
ML Engineer, Infrastructure at Prior Labs scores 93 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI Level 4. Building AI systems is the job itself: without AI, the role would not exist.
- AI Level 480 to 100
- AI Level 360 to 79
- AI Level 240 to 59
- AI Level 10 to 39
Bands come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- What's a project where you used Claude hands-on?
- Walk me through how you've used PyTorch in your day-to-day work.
- What are the limits of Weights And Biases that you've run into, and how did you work around them?
- What's a project where you used Cursor hands-on?
- Walk me through how you've used Claude Code in your day-to-day work.
Adapt your resume
- List these exact terms on your resume: Claude, PyTorch, Weights And Biases, Cursor, and Claude Code. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new machine learning engineer jobs at AI Level 4+ by email
One email a week with the new machine learning engineer jobs at AI Level 4+, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Data roles rated AI Level 4 at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step