Level

Isomorphic Labs

Software Engineer (Compute Efficiency), London

AI in this role

Design observability systems and automated recovery mechanisms to ensure planet-scale accelerator fleet operates at peak efficiency for AI workloads.

gpu
compute-efficiencyinfrastructureobservabilitymachine-learning-platforms

Isomorphic Labs is applying frontier AI to help unlock deeper scientific insights, faster breakthroughs, and life-changing medicines with an ambition to solve all disease.

The future is coming. A future enabled and enriched by the incredible power of machine learning. A future in which diseases are curtailed or cured starting with better and faster drug discovery. 

Come and be part of an interdisciplinary team driving groundbreaking innovation and play a meaningful role in contributing towards us achieving our ambitious goals, while being a part of an inspiring and collaborative culture.

The world we want tomorrow is the one we’re building today. It starts with the culture at this company. It starts with you. 

 

About Iso

Isomorphic Labs (IsoLabs) was launched in 2021 to advance human health by building on and beyond the Nobel-winning AlphaFold system. Since then, our interdisciplinary team of drug discovery experts and machine learning specialists has built powerful new predictive and generative AI models that accelerate scientific discovery at digital speed.

Our name comes from the belief that there is an underlying symmetry between biology and information science. By harnessing AI’s powerful capabilities, we can use it to model complex biological phenomena to help design novel molecules, anticipate how drugs will perform and develop innovative medicines to treat and cure some of the world’s most devastating diseases.

We have built a world-leading drug design engine comprising AI models that are capable of working across multiple therapeutic areas and drug modalities. We are continually innovating on model architecture and developing cutting-edge capabilities to advance rational drug design.

Every day, and with each new breakthrough, we’re getting closer to the promise of digital biology, and achieving our ambitious mission to one day solve all disease with the help of AI.

 

Your impact 

We are building the largest foundation models in biotech and applying them immediately to cure disease. You will play a key role and work at a grand scale to deliver the foundations that make this happen. Joining the Compute Infrastructure team, you will ensure our planet-scale accelerator fleet operates at peak health and efficiency. By partnering with in-house machine learning platforms, performance & scaling team, and AI researchers, you will design the observability systems and automated recovery mechanisms that maximize scientific throughput on every paid GPU-hour.

What you will do 

  • Design, deploy, and scale robust observability systems and telemetry pipelines to monitor fleetwide compute efficiency, hardware health, and workload goodput across distributed clusters
  • Drive hardware efficiency and node reliability across our accelerator fleet, and integrating new hardware to leverage advancements
  • Identify compute waste and efficiency bottlenecks across the fleet, partnering with ML and platform teams to actively optimize accelerator utilization and improve workload goodput
  • Collaborate with teams in the AI org, and work closely with the ML Infrastructure team to identify, instrument, and improve canonical efficiency metrics across both training and inference runs
  • Contribute to the efforts for consistently improving the reliability of our ML runs.
  • Operate, maintain, and harden research, development, and production cloud infrastructure and cluster deployments
  • Partner and collaborate with a diverse set of teams incl. science, research, product, business development and operations
  • Contribute to core technical decisions (e.g. choice of tooling, infrastructure, and architectural design)

 

Skills and qualifications 

Essential:

  • Possess real world experience operating, monitoring, and debugging infrastructure for large-scale AI/ML workloads
  • Have experience working in cloud compute infrastructure design, preferably GCP
  • Possess strong programmings skills
  • Have significant experience working and deploying in Kubernetes at scale
  • Familiarity with the Nvidia GPU generations
  • Proven track record of building production observability and telemetry stacks

Nice to have:

  • Have a background in either ML SWE or infrastructure SRE work to build on
  • Conceptual understanding of ML workload efficiency paradigms
  • Have experience leading and delivering projects to multidisciplinary stakeholders
  • Familiarity with Google TPU generations
  • Familiarity with: workload scheduling; machine learning efficiency research; familiarity with ML-driven R&D cycles; familiarity with hardware benchmarking


Culture and values

We are guided by our shared values. It's not about finding people who think and act in the same way. These values help to guide our work and will continue to strengthen it. 

Thoughtful
Thoughtful at Iso is about curiosity, creativity and care. It is about good people doing good, rigorous and future-making science every single day.

Brave
Brave at Iso is about fearlessness, but it’s also about initiative and integrity. The scale of the challenge demands nothing less.

Determined
Determined at Iso is the way we pursue our goal. It’s a confidence in our hypothesis, as well as the urgency and agility needed to deliver on it. Because disease won’t wait, so neither should we.

Together
Together at Iso is about connection, collaboration across fields and catalytic relationships. It’s knowing that transformation is a group project, and remembering that what we’re doing will have a real impact on real people everywhere.


Creating an extraordinary company

We believe that to be successful we need a team with a range of skills and talents. We're building an environment where collaboration is fundamental, learning is shared and every employee feels supported and able to thrive. We value unique experiences, knowledge, backgrounds, and perspectives, and harness these qualities to create extraordinary impact.

We are committed to equal employment opportunities regardless of sex, race, religion or belief, ethnic or national origin, disability, age, citizenship, marital, domestic or civil partnership status, sexual orientation, gender identity, pregnancy or related condition (including breastfeeding) or any other basis protected by applicable law. If you have a disability or additional need that requires accommodation, please do not hesitate to let us know.


Hybrid working

It’s hugely important for us to share knowledge and build strong relationships with each other, and we find it easier to do this if we spend time together in person. This is why we follow a hybrid model, and would require you to be able to come into the office 3 days a week (currently Tuesday, Wednesday, and one other day depending on which team you’re in). If you have additional needs that would prevent you from following this hybrid approach, we’d be happy to talk through these if you’re selected for an initial screening call.

Please note that when you submit an application, your data will be processed in line with our privacy policy.


>> Click to view other open roles at Isomorphic Labs

How we rate this

Software Engineer (Compute Efficiency), London at Isomorphic Labs rates 85 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.

Classification

Builds AI. The job is building AI systems.

  1. ●●●● Builds AI80 to 100
  2. ●●●○ Works on AI60 to 79
  3. ●●○○ Uses AI40 to 59
  4. ●○○○ Little AI0 to 39

Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Compute EfficiencyInfrastructureObservabilityMachine Learning PlatformsGPU

Questions you could be asked

  1. Tell me about a project where compute efficiency was part of your work. What did you do?
  2. Tell me about a project where infrastructure was part of your work. What did you do?
  3. Tell me about a project where observability was part of your work. What did you do?
  4. Tell me about a project where machine learning platforms was part of your work. What did you do?
  5. Walk me through how you've used GPU in your day-to-day work.

Adapt your resume

  • List these exact terms on your resume: Compute Efficiency, Infrastructure, Observability, Machine Learning Platforms, and GPU. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want an expert to read your CV for this job?

Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.

Get a free CV review

Get new software engineer jobs (Builds AI ●●●●) by email

One email a week with the new software engineer jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Software Engineering roles that build AI, at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Isomorphic Labs

More software engineer jobs

Related searches

Same AI level