Level

Sword HealthPosted 2w ago

L4

Senior ML Engineer

Senior ML Engineer at Sword Health scores 92 out of 100 on AI centrality, which makes it a Level 4 role on this board.

Remote (Remote - Europe)senior€60k-€85k

AI in this role

fine-tuning

At Sword, we’re building AI to heal billions and unlock humanity’s full potential. In doing so, we’re pioneering AI Care, a fundamentally new approach to healthcare built for medical reasoning, safety, and real-time treatment, not generic technology applied after the fact. As both a clinical-centric frontier AI lab and an applied AI platform, Sword is reimagining how care is delivered at scale, removing traditional barriers like appointments, waiting rooms, and stigma so more people can access the care they need—and ultimately get back to lives lived in full.

Since 2020, Sword has expanded across physical therapy, women’s health, cardiometabolic, and mental health, and is now moving beyond the session to a fully AI-native, 24/7 care program that brings physical activity, therapeutic exercise, psychotherapy, nutrition, and behavior change into one connected experience. More than 700,000 members across three continents have completed over 10 million AI sessions, helping 1,000+ enterprise clients avoid more than $1 billion in unnecessary healthcare costs. Backed by 42 clinical studies, 44+ patents, and more than $500 million raised from leading investors including Khosla Ventures, General Catalyst, and Founders Fund, Sword is defining a new standard for healthcare.

 

AI Proficiency at Sword

AI fluency is a core expectation at Sword. Every candidate is assessed against our three-level framework — be ready to share real examples of how AI is already part of how you work.

  • Explorer (Level 1) — Uses AI daily to boost personal productivity

  • Builder (Level 2) — Creates workflows and tools that elevate the whole team

  • Integrator (Level 3) — Embeds AI into products and processes at scale

Every hire must demonstrate at least Level 1. The expected level will vary depending on the seniority of the role.

 

What you’ll be doing

  • Own ML projects end-to-end: take problems from exploration through production deployment and keep iterating once real users are on them;
  • Build agentic LLM systems: design multi-step workflows with tool use, retrieval, and orchestration, and make them reliable enough for clinical settings;
  • Treat evaluation as core engineering work: build eval sets, offline and online harnesses, LLM-as-judge pipelines with human review, and regression tests that catch quality drops before they reach users;
  • Improve model quality with whatever fits the problem: prompting, retrieval, distillation, or fine-tuning, chosen on evidence rather than habit;
  • Work across the full AI stack: data prep, model adaptation, serving, monitoring, and the feedback loops that keep systems improving in production;
  • Partner with Product, Clinical, and Engineering: translate clinical requirements into technical decisions and surface tradeoffs early;
  • Help the team get better: review code, share what you learn, and mentor engineers earlier in their careers.

 

What you need to have

  • Experience shipping ML systems to production that people actually depend on;
  • Hands-on LLM work in production: prompting, retrieval, tool calling, and agent-style workflows;
  • A rigorous approach to evaluation: you've built eval datasets and frameworks, and you can tell a real improvement from noise;
  • Strong ML fundamentals: you know which approach fits which problem and can reason clearly about tradeoffs;
  • Comfort with ambiguity: you've taken loosely defined problems and turned them into something running in production;
  • Solid engineering skills: production-quality code, familiarity with distributed systems, and the patience to debug messy ML pipelines;
  • Clear communication with both technical and clinical stakeholders.

Bonus points

  • Experience with fine-tuning or preference optimization (RLHF, DPO, or similar);
  • Healthcare AI, or other high-stakes domains where errors carry real cost;
  • Built agent frameworks or evaluation tooling from scratch;
  • Open source contributions, technical writing, or other knowledge sharing.

 

*The range below reflects base, variable, and equity for this role.

This is a starting point, not a ceiling — once someone joins and proves they're outlier talent, we adjust quickly to make sure their compensation matches their impact.

Job titles may span more than one career level, so actual pay depends on skills, qualifications, experience, location, and market demand, among other factors. The range reflects base salary plus any variable, bonus, or sales incentives, and our estimate of the value of private company stock options where applicable. It's subject to change, and future stock value isn't guaranteed. Sword also offers additional benefits beyond total compensation.

Total Compensation Range€60.000—€85.000 EUR    Country-specific benefits may vary and will be detailed separately. Sword Health complies with applicable Federal and State civil rights laws and does not discriminate on the basis of Age, Ancestry, Color, Citizenship, Gender, Gender expression, Gender identity, Gender information, Marital status, Medical condition, National origin, Physical or mental disability, Pregnancy, Race, Religion, Caste, Sexual orientation, and Veteran status.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Fine Tuning

Questions you could be asked

  1. Walk me through fine-tuning a model: what data did you use, and how did you check the result?
  2. How would you decide a model or AI system is ready to ship?
  3. Tell me about a time a model underperformed in production. How did you find out, and what did you change?

Adapt your resume

  • List these exact terms on your resume: Fine Tuning. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Similar roles

Data roles rated Level 4 at other companies.

More jobs at Sword Health

More machine learning engineer jobs

Related searches

Same AI level