Technical Program Manager, Supercomputing
AI in this role
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
About the Role
Compute allocation is a technical judgment problem. Research teams have different goals, workloads have different requirements, and available resources rarely match every need at once. Progress depends on understanding those constraints and finding ways to accomplish more within them.
In this role, you’ll own compute allocation and help us make better use of our supercomputing resources. You’ll work closely with researchers and infrastructure engineers to understand demand, evaluate competing needs, and translate research priorities into practical allocation decisions. You’ll investigate where constraints limit progress and help teams find workable alternatives.
Teaching is central to this work. You’ll help researchers plan their compute needs, understand allocation decisions, and reason through tradeoffs themselves. By working through real workloads, explaining your decisions, and developing useful tools and guidance, you’ll make specialized knowledge available to more of the team.
This role is a good fit if you enjoy working deeply with technical constraints, making decisions under uncertainty, and helping others develop their judgment.
What You’ll Do
Own compute allocation. Understand current and upcoming demand, develop allocation plans with technical leads, and carry decisions through to execution.
Optimize within constraints. Work through tradeoffs across capacity, hardware suitability, workload size, timing, and dependencies. Help teams adapt their plans as requirements and resources change.
Improve productive use of compute. Partner with researchers and engineers to investigate the gap between allocated resources and useful work. Identify opportunities to improve scheduling, workload placement, and utilization.
Teach teams to plan and use compute well. Help researchers estimate their needs, understand available options, and recognize the consequences of different choices. Build practical understanding through real workloads and decisions.
Make your reasoning transferable. Explain assumptions and tradeoffs clearly. Develop examples, guidance, and tools that enable others to handle recurring decisions and recognize when they need help.
Connect research and infrastructure. Keep research plans grounded in usable capacity and make future needs visible to supercomputing and infrastructure teams.
Learn from outcomes. Compare expected and actual usage, revisit assumptions, and improve allocation decisions as workloads, systems, and priorities change.
Skills and Qualifications
Minimum qualifications:
Experience making consequential resource allocation or capacity decisions in a complex computing environment, with clear ownership of the tradeoffs and outcomes.
Strong technical understanding of distributed systems, machine learning infrastructure, or high-performance computing. You can reason about how workload requirements interact with system constraints.
Demonstrated ability to improve useful output from limited resources. You can explain what constrained the work, which alternatives you considered, and why your approach helped.
Ability to teach complex technical ideas through clear explanations and practical examples. You have helped others become more capable and independent.
Comfort investigating usage and performance data directly, questioning assumptions, and changing your view when the evidence warrants it.
Sound judgment when teams have competing priorities. You earn trust, explain difficult decisions, and work constructively through disagreement.
A hands-on, collaborative approach with strong attention to detail and follow-through.
Preferred qualifications:
Experience owning compute allocation at an AI research lab or another organization with substantial shared computing resources.
Familiarity with GPU clusters, distributed training, scheduling, and workload placement.
Understanding of how networking, storage, reliability, and hardware differences affect usable capacity.
Experience building tools or analyses for capacity forecasting, allocation, or utilization.
Experience mentoring researchers, engineers, or technical program managers in compute planning and resource management.
We encourage you to apply even if you don’t meet every preferred qualification.
Logistics
Location: San Francisco, California.
Compensation: [Add approved salary range.]
Visa sponsorship: We support visa sponsorship and work with candidates through the process.
Benefits: Health, dental, and vision coverage; unlimited PTO; paid parental leave; and relocation support as needed.
How we rate this
Technical Program Manager, Supercomputing at Thinking Machines Lab rates 42 out of 100 for how much of the daily work is AI. That makes it Uses AI (AI Level 2 of 4). The level is about AI in the job, not seniority.
Uses AI. An ordinary role that requires AI tools.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Tell me about a research question you investigated. What did you find?
- This role expects you to use AI tools as part of the job. Which ones have you used, and for what?
- Tell me about a time an AI tool got something wrong. How did you catch it?
Adapt your resume
- List these exact terms on your resume: AI Research. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Put the AI tool in a bullet point about what you did, not just in a skills list — this role treats it as a required part of the job.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new AI jobs (Uses AI ●●○○ or higher) by email
One email a week with the new AI jobs (Uses AI ●●○○ or higher), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Product roles that use AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step