Level

Modular

Software Engineer, Hardware Enablement

AI in this role

pytorch

About Modular


At Modular, a Qualcomm company, we’re on a mission to revolutionize AI infrastructure by systematically rebuilding the AI software stack from the ground up. Our team, made up of industry leaders and experts, is building cutting-edge, modular infrastructure that simplifies AI development and deployment. By rethinking the complexities of AI systems, we’re empowering everyone to unlock AI’s full potential and tackle some of the world’s most pressing challenges.
If you’re passionate about shaping the future of AI and creating tools that make a real difference in people’s lives, we want you on our team. You can read about our culture and careers to understand how we work and what we value.

About the Role:


We're looking for a motivated Hardware Enablement Engineer to help bring that vision to new accelerator platforms. You'll work across the Modular software stack — from Mojo kernels and compiler infrastructure to the graph compiler, runtime, and MAX model serving — to bring up and optimize support for new hardware architectures.
This is a highly cross-functional systems role. You'll learn how new accelerators execute computation, move data, synchronize work, and interact with vendor toolchains, then help translate those capabilities into a performant and reliable experience within Modular's stack. You'll collaborate closely with engineers across Modular as well as external hardware partners, developing deep expertise in novel architectures while contributing directly to our hardware portability story.
LOCATION: Candidates based in the US, Canada, and UK are welcome to apply. You can work in our office in Los Altos, CA, Edinburgh, or remotely from home. Onboarding for new hires is conducted in-person at the appropriate office.

What you will do:


  • Bring up and validate support for new hardware architectures across the Modular software stack, working across kernels, compiler infrastructure, runtime, graph execution, and model serving.
  • Write, port, and optimize Mojo kernels for novel accelerator architectures, establishing correctness first and iterating toward strong performance.
  • Investigate performance and correctness issues across layers of the stack, including kernel execution, memory movement, compiler lowering, graph decisions, synchronization, and vendor runtime behavior.
  • Develop a detailed understanding of new hardware platforms, including their ISA, execution model, memory hierarchy, synchronization mechanisms, compiler constraints, and vendor toolchains.
  • Help map important AI operators and workloads onto new architectures, understanding how operations such as matrix multiplication, convolution, reductions, and other model primitives execute efficiently on the target hardware.
  • Contribute to portability infrastructure, compiler integration, tooling, testing, and debugging workflows that make it easier to support additional hardware platforms over time.
  • Collaborate directly with hardware vendor engineers to understand platform capabilities, investigate issues, build integration tests, and improve end-to-end hardware/software integration.
  • Benchmark and profile workloads as hardware support matures, identifying bottlenecks and helping move new platforms from initial correctness toward production-quality performance.
  • Share knowledge about emerging architectures with the broader team through technical documentation, demos, design discussions, and engineering write-ups.
  • Participate in company events such as onsites and hackathons while contributing to Modular's collaborative and open engineering culture.


What you Bring to the table:


  • 5+ years of experience in high-performance computing, compiler engineering, accelerator software, systems engineering, or a closely related domain in industry or research.
  • Strong C++ skills and experience contributing to complex, multi-component software systems.
  • Hands-on experience with at least one heterogeneous programming model such as CUDA, SYCL, OpenCL, or a comparable accelerator programming environment.
  • Understanding of how AI operators are implemented at a low level, demonstrated through experience writing or modifying GPU kernels, custom operators, accelerator code, or working with ML frameworks such as PyTorch at the C++/systems layer.
  • Working knowledge of hardware concepts such as memory hierarchy, parallel execution, synchronization, data movement, and the relationship between software performance and the underlying architecture.
  • Experience debugging problems that cross software abstraction boundaries and an ability to systematically reason about correctness and performance.
  • Ability and enthusiasm to learn unfamiliar hardware platforms quickly, including becoming comfortable reading architecture manuals, ISA documentation, and vendor technical materials.
  • Strong communication and collaboration skills, including the ability to work effectively across compiler, runtime, kernel, and hardware teams.
  • A pragmatic engineering mindset that prioritizes correctness during initial bringup while understanding how to iterate methodically toward performance.


Helpful, but not required:


  • Experience with non-GPU accelerator architectures, such as DSPs, NPUs, AI ASICs, or other specialized compute hardware.
  • Familiarity with MLIR or LLVM compiler infrastructure.
  • Experience with GPU DSLs or libraries such as Triton, CUTLASS, or CuTe.
  • Previous experience bringing up software support for a new accelerator or hardware platform.
  • Experience working directly with hardware vendors or silicon engineering teams.
  • Experience profiling and optimizing kernels or accelerator workloads using hardware-specific performance tools.
  • Exposure to graph compilers and understanding how model operations are lowered, scheduled, fused, or mapped onto target hardware.
  • Experience with model serving or inference optimization workflows.
  • Familiarity with runtime concepts such as device management, memory allocation, queues, synchronization, or asynchronous execution.
  • Experience working on portability layers that target multiple hardware architectures.


What Modular brings to the table:


  • Amazing Team. We are a progressive and agile team with some of the industry’s best engineering and product leaders.
  • World-class Benefits. In order to attract the best, we need to offer the best. Your benefits package may include comprehensive healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities. Please note that specific benefit packages may vary based on your location, you can read more about benefits offered by Qualcomm here.
  • Competitive Compensation. We offer very strong compensation packages, including RSU grants. We want people to be focused on their best work and believe in tailoring compensation plans to meet the needs of our workforce. 
  • Team Building Events. We organize regular team onsites and local meetups in Los Altos, CA as well as different cities. Traveling 2-4 times a year is expected for all roles. 

Working at Modular will enable you to grow quickly as you work alongside incredibly motivated and talented people who have high standards, possess a growth mindset, and share a purpose to fundamentally change how software is built for AI and accelerated computing.
The estimated base salary range for this role to be performed in the US, regardless of the state, is The estimated base salary range for this role to be performed in the US, regardless of the state, is $180,000.00 - $270,000.00 USD. 
The estimated base salary range for this role to be performed in Canada, regardless of the province, is $172,400.00–$258,600.00 CAD.
The estimated base salary range for this role to be performed in the United Kingdom is £99,600.00 – £150,000.00 GBR.
The salary for the successful applicant will depend on a variety of permissible, non-discriminatory job-related factors, which include but are not limited to education, training, work experience, business needs, or market demands. This range may be modified in the future. Total compensation may also include annual target bonus, equity, and benefits, as applicable.
For candidates who fall outside of the listed requirements, we nevertheless encourage you to apply, as we may have upcoming openings at lower or higher levels than the one advertised.

How we rate this

Software Engineer, Hardware Enablement at Modular rates 87 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.

Classification

Builds AI. The job is building AI systems.

  1. ●●●● Builds AI80 to 100
  2. ●●●○ Works on AI60 to 79
  3. ●●○○ Uses AI40 to 59
  4. ●○○○ Little AI0 to 39

Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

PyTorch

Questions you could be asked

  1. What's a project where you used PyTorch hands-on?
  2. How would you decide a model or AI system is ready to ship?
  3. Tell me about a time a model underperformed in production. How did you find out, and what did you change?

Adapt your resume

  • List these exact terms on your resume: PyTorch. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want an expert to read your CV for this job?

Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.

Get a free CV review

Get new remote software engineer jobs (Builds AI ●●●●) by email

One email a week with the new remote software engineer jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Software Engineering roles that build AI, at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Modular

More software engineer jobs

Related searches

Same AI level