# Inference Systems Engineer at Modular

Modular is hiring an Inference Systems Engineer. It pays $148k-$270k a year and Level rates it Builds AI ●●●●; you can [apply on Level](https://jobsbylevel.com/go/6f39be41-f9b4-4116-90a8-931653e88619).

AI Level 4, AI centrality 90 out of 100. Remote (United States - Remote).

## Details

- Company: [Modular](https://jobsbylevel.com/companies/modular)
- AI level: AI Level 4 (score 90 out of 100)
- Location: Remote (United States - Remote)
- Salary: $148k-$270k
- Posted: August 27, 2026
- Apply: https://jobsbylevel.com/go/6f39be41-f9b4-4116-90a8-931653e88619

## Description

About Modular
At
Modular, a Qualcomm company
, we’re on a mission to revolutionize AI infrastructure by systematically rebuilding the AI software stack from the ground up. Our team, made up of industry leaders and experts, is building cutting-edge, modular infrastructure that simplifies AI development and deployment. By rethinking the complexities of AI systems, we’re empowering everyone to unlock AI’s full potential and tackle some of the world’s most pressing challenges.
If you’re passionate about shaping the future of AI and creating tools that make a real difference in people’s lives, we want you on our team. You can read about
our culture
and
careers
to understand how we work and what we value.
About the role:
ML developers today face significant friction in taking trained models into deployment. They work in a highly fragmented space, with incomplete and patchwork solutions that require significant performance tuning and non-generalizable, model-specific enhancements. At Modular, we are building the Modular platform: a next generation AI platform that will radically improve the way developers build and deploy AI models.
State-of-the-art inference is a fast moving target - every few months brings new model architectures, new attention variants, new parallelism strategies, and new optimization techniques (e.g. new variants of disaggregation, speculative decoding, hierarchical KV caching, constrained decoding, MoE routing, and more). Building each feature right in isolation is already a challenge, but designing for clean composability and long tail reliability brings it to another level.
LOCATION:
Candidates based in the
United States
are welcome to apply. To support growth and collaboration, those in earlier career stages work in a hybrid capacity at our Los Altos, CA. More senior staff can work out of our office in Los Altos, CA or remotely from home. Onboarding for new hires is conducted in-person in our Los Altos, CA office.
What you will do:

This team's job is to identify the opportunities for unifying structures underneath today’s inference features, and turn them into the robust and flexible distributed-systems abstractions and building blocks to make tomorrow's features ship faster
and
more robustly. Responsibilities would include:
Leading cross-team architecture-level decisions on concurrency, data movement, and system boundaries in a fast-moving in-house software stack spanning from the metal to the cloud.
Developing flexible and extensible interfaces for open-source model developers and community contributors that are not just enjoyable to use, but implicitly promote best practices.
Designing highly modular systems with clear contracts to create components with unambiguous ownership, testability, and scaffolding for robust agentic development.
What you bring to the table:

5+ years of systems programming experience with a focus on performance, concurrency, and distributed architectures
Strong instinct for API and abstraction design, and the ability to present a strong case for design proposals
Comfortable working in Python, Rust, Mojo, or similar (in the agentic-coding era, we believe deep understanding of software systems supersedes knowing language syntax)
Intuition and ability to reason about pathological performance bottlenecks in complex systems, and translate that into concrete diagnostics and solutions
Demonstrated ownership of software solutions from implementation through production support. You bring war stories from learning the hard way why defensive engineering is worth the extra investment.
Genuine interest in the research literature and solid instincts for what's worth building
Helpful Experience
Experience under the hood deep inside a high-performance ML inference system (vLLM, SGLang, TRT-LLM, etc.)
Broad systems programming experience: lock-free and wait-free data structures, POSIX APIs and architecture, memory-layout aware optimization, NUMA-aware design, zero-copy transport, etc.
Experience

The description is cut here. Read the full offer: https://jobsbylevel.com/jobs/inference-systems-engineer-at-modular-897a14

Source: https://jobsbylevel.com/jobs/inference-systems-engineer-at-modular-897a14

## Cite this page

Level. https://jobsbylevel.com/jobs/inference-systems-engineer-at-modular-897a14.

Get job alerts: https://jobsbylevel.com/newsletter
