Level

Luma AI

Software Engineer, Inference

Luma AI is hiring a Software Engineer, Inference for a remote role open to applicants in United Kingdom. Level rates it ; you can apply on Level.

AI in this role

hugging-facevllmpytorch

You'll own how Luma's models get served — integrating new architectures into the inference engine, scaling deployments across thousands of machines, and keeping expensive GPU fleets busy while meeting internal SLOs.

This is large-scale inference systems work: scheduling, fleet management, deployment pipelines, and reliability across clusters and hardware providers. It fits a strong systems engineer comfortable with model serving and Kubernetes at scale. If you want pure modeling rather than the systems that run models, this is firmly the systems side.

What You'll Own

  • Ship new model architectures by integrating them into the inference engine.

  • Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.

  • Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.

  • Automate, test, and maintain inference services for maximum uptime and reliability.

  • Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.

  • Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.

First 90 Days

One way the first 90 could unfold.

  • Days 1–30 — Immerse & Diagnose: Learn the inference stack, the fleets, and where reliability or utilization break.

  • Days 30–60 — Ship & Validate: Integrate a model or ship tooling/scheduling that improves uptime or GPU utilization.

  • Days 60–90 — Scale & Systemize: Harden deployment pipelines and scheduling across clusters and providers.

What You Bring

  • Strong Python and system-architecture skills.

  • Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar.

  • Experience with queues, scheduling, traffic control, and fleet management at scale.

  • Experience with Linux, Docker, and Kubernetes, and with orchestration, deployment, and scheduling.

  • Familiarity with Redis and S3-compatible storage.

Nice to Have

  • Modern networking stacks including RDMA (RoCE, InfiniBand, NVLink).

  • High-performance large-scale ML systems (100+ GPUs).

  • CUDA, and FFmpeg or multimedia processing.

About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

How we rate this

Software Engineer, Inference at Luma AI rates 97 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.

Classification

Builds AI. The job is building AI systems.

  1. ●●●● Builds AI80 to 100
  2. ●●●○ Works on AI60 to 79
  3. ●●○○ Uses AI40 to 59
  4. ●○○○ Little AI0 to 39

Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Hugging FacevLLMPyTorch

Questions you could be asked

  1. What's a project where you used Hugging Face hands-on?
  2. Walk me through how you've used vLLM in your day-to-day work.
  3. What are the limits of PyTorch that you've run into, and how did you work around them?
  4. How would you decide a model or AI system is ready to ship?
  5. Tell me about a time a model underperformed in production. How did you find out, and what did you change?

Adapt your resume

  • List these exact terms on your resume: Hugging Face, vLLM, and PyTorch. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want an expert to read your CV for this job?

Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.

Get new remote software engineer jobs (Builds AI ●●●●) by email

One email a week with the new remote software engineer jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Software Engineering roles that build AI, at other companies.

Version 1

London, Birmingham, Manchester, Newcastle upon Tyne, Edinburgh, Belfast, England, United Kingdom2h ago

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Luma AI

More software engineer jobs

Related searches

Same AI level

Jobs by city