# Senior Infrastructure Engineer, Research at PhysicsX

PhysicsX is hiring a Senior Infrastructure Engineer, Research in Singapore, Singapore. Level rates it Builds AI ●●●●; you can [apply on Level](https://jobsbylevel.com/go/9b516650-a8f8-45af-85ee-b28024346b5b).

AI Level 4, AI centrality 92 out of 100. Singapore.

## Details

- Company: [PhysicsX](https://jobsbylevel.com/companies/physicsx)
- AI level: AI Level 4 (score 92 out of 100)
- Location: Singapore
- Posted: August 6, 2026
- Apply: https://jobsbylevel.com/go/9b516650-a8f8-45af-85ee-b28024346b5b

## Description

About us Re-architecting Engineering for the Age of Intelligence PhysicsX is the physics AI company for industrials. The company’s mission is to accelerate hardware innovation by overhauling what industrial engineering and manufacturing look like today. PhysicsX is building a new simulation software stack to deliver deep physics AI enablement across the entire engineering lifecycle. The company partners with leading organisations in aerospace & defence, automotive, semiconductors, materials, and energy & renewables, supporting them on some of their most critical and complex challenges. PhysicsX is headquartered in the United Kingdom, with offices in London, New York, and Singapore and an expanding presence in the Bay Area. Who we’re looking for Our Research team runs large-scale training jobs and data generation workflows for CAE and LPM training on an HPC cluster comprising of 304 B200s, orchestrated via SLURM on Kubernetes. We’re looking for a Senior Infrastructure Engineer to help us build, scale, and operate this platform - and to work directly with researchers to keep their jobs running fast, safely, and reliably. This role is split roughly evenly between platform engineering (building, hardening, and optimising the infrastructure everything runs on), and hands-on support (debugging job failures, optimising throughput, and helping Research get the most out of the cluster). What you will do Design, build, and maintain our SLURM-on-Kubernetes infrastructure, supporting large-scale distributed training and data generation workloads Own cluster reliability, efficiency, and utilization across hundreds of GPUs Partner directly with the Research team to debug job failures, diagnose performance bottlenecks, and improve iteration speed Build and maintain containerized job job environments (Docker, Enroot/Pyxis) for reproducible research workflows Manage high-performance storage systems to support data-intensive CAE workloads Develop tooling, automation, and observability to reduce operational overhead and catch issues before they impact the Research team Contribute to capacity planning and infrastructure roadmap as the cluster and team scale Build and implement a technical and scaling strategy for Research’s use of high-performance computing to achieve efficient large-scale model training and data generation workloads What you bring to the table 4-6 years of experience in infrastructure, platform, or HPC engineering roles Hands-on experience with SLURM and Kubernetes, ideally running them together in production environments Experience with containerization and orchestration for compute-heavy workloads (Docker, Enroot/Pyxis, etc) Working knowledge of high-performance networking (Infiniband, RDMA, NCCL) in the context of distributed training Experience managing HPC/ML storage systems at scale Strong debugging skills across the stack - from scheduling and networking to individual job failures Comfortable working directly with researchers and translating their needs into infrastructure improvements Solid scripting/automation skills Nice to have Experience operating GPU clusters at scale (hundreds of GPUs or more) Background in CAE or simulation-driven data generation Experience with GPU cluster health monitoring and fault tolerance for long-running training jobs Familiarity with infrastructure-as-code tools (Terraform, Helm, etc) What we offer Build what actually matters Help shape an AI-native engineering company at a formative stage, tackling problems that genuinely matter for industry and society. This is work with real-world impact - and something you can be proud to stand behind. Learn alongside exceptional people Work with a high-caliber, collaborative team of engineers, scientists, and operators who care deeply about doing great work, and about helping each other get better. We come from diverse backgrounds, but we share a commitment to operating at the highest level and addressing some of the most complex challenges out

The description is cut here. Read the full offer: https://jobsbylevel.com/jobs/senior-infrastructure-engineer-research-at-physicsx-5dc69e

Source: https://jobsbylevel.com/jobs/senior-infrastructure-engineer-research-at-physicsx-5dc69e

## Cite this page

Level. https://jobsbylevel.com/jobs/senior-infrastructure-engineer-research-at-physicsx-5dc69e.

Get job alerts: https://jobsbylevel.com/newsletter
