Software Engineer (C++ Systems)
AI in this role
Company
The world is building massive amounts of GPU capacity. Meanwhile, deployed GPUs are only 20% utilized.
This is because GPUs are not virtualized, while every other type of hardware is. For example CPUs and storage are allocated through virtual abstractions which efficiently manage the physical hardware, while GPUs are statically allocated on a one-to-one basis.
Thunder Compute is building this virtualization layer for GPUs. We have raised over $17.5M from Matrix Partners, Y Combinator, and leading angels from Coreweave, Microsoft, Cognition, and Anthropic.
Leading solutions for underutilization sit at the workload layer and are therefore only able to optimize specific use cases. We believe the ideal cluster optimization solution must be invisible to developers and compatible with all workloads; hence, it must sit at the systems layer.
We are a team of systems researchers productionizing cutting-edge GPU virtualization research to build this general-purpose optimization layer.
Concretely, our virtualization library abstracts GPUs across TCP networking. We use a userspace shim library, loaded through LD_PRELOAD, to intercept CUDA calls and send them over gRPC to a host server connected to a physical GPU elsewhere in the data center.
This enables something like “Ceph for GPUs”: GPUs become network resources that can be abstracted, pooled, and dynamically allocated across a cluster to improve utilization without requiring developers to modify their workloads.
Role
Your work will focus on building the core C++ systems behind our virtualization layer. This includes low-latency performance optimization, distributed systems debugging, production reliability, and research into new techniques for improving GPU utilization.
You will take ownership of complex systems from early experimentation through production deployment. Example projects may include:
Profiling and reducing latency across remote CUDA operations
Building high-performance networking and data-transfer paths
Debugging failures across customer processes, our userspace runtime, the network, and remote GPU servers
Improving support for process forking, signals, multithreading, dynamic linking, and unusual application behavior
Designing systems for GPU allocation, scheduling, failure recovery, and observability
Researching and productionizing new GPU virtualization and oversubscription techniques
Expanding compatibility across CUDA applications, frameworks, and GPU architectures
You will spend your days bouncing between the weeds of complex, performance-critical systems that are live in production. One week, you may be tracing a synchronization bug across a distributed CUDA workload; the next, you may be redesigning a hot data path to remove microseconds of overhead.
This work is not easy. It blends the hardest parts of systems research and production engineering.
We look for exceptional low-level engineering talent, strong work ethic, and extreme attention to detail. We must move quickly while shipping high-quality, reliable systems code.
Core Technical Skills
Exceptional modern C++ ability, including memory management, concurrency, performance optimization, and systems-level abstraction design
Deep understanding of operating systems, low-level networking, compilers, distributed systems, or computer architecture
Experience building and operating performance-critical C++ systems in production
Strong Linux systems programming and debugging ability
Ability to reason through unfamiliar systems across multiple layers of the stack
Must Haves
Strong work ethic and the ability to independently push a project from an experimental prototype through 100% completion under tight deadlines
Attention to detail and the ability to deliver production-ready, thoroughly tested code without significant oversight
Strong ownership over correctness, reliability, performance, and operational outcomes
Ability to debug ambiguous problems without a clear reproduction, existing playbook, or obvious owner
Willingness to work directly with customers and investigate difficult production failures
Preferred
Experience with CUDA, GPU systems, compilers, runtime interception, dynamic linking, high-performance networking, or distributed computing
Experience at a trading firm such as Citadel Securities or Jane Street; a hardware or AI infrastructure company such as NVIDIA or SambaNova; a systems research group; or a similarly demanding engineering environment
Strong computer science fundamentals demonstrated through academic work, systems research, competitive programming, open-source contributions, or exceptional professional experience
Experience taking new systems research from a paper or prototype into a reliable production system
Why Join
You will join early enough to meaningfully shape the architecture, engineering standards, and technical direction of the company.
You will work directly with the founders on a category-defining systems problem, with a short path between writing code and seeing it run in production. The systems you build will form the foundation of a new infrastructure layer for GPU computing.
Logistics
You will report to co-founder and CTO Brian Model, formerly a Quantitative Developer at Citadel Securities
This role is full-time and in person, five days per week, at our office in downtown San Francisco
Relocation support and visa sponsorship are available
Benefits
Competitive salary and meaningful equity
Daily lunch, snacks, and coffee
Team dinners and events
401(k)
Health, dental, and vision insurance
How we rate this
Software Engineer (C++ Systems) at Thunder Compute rates 90 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- What's a project where you used Anthropic hands-on?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: Anthropic. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new software engineer jobs (Builds AI ●●●●) by email
One email a week with the new software engineer jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step