Level

AmazonPosted 1d ago

Sr Technical Program Manager - Hardware, AWS Generative AI & ML Servers

Sr Technical Program Manager - Hardware, AWS Generative AI & ML Servers at Amazon scores 65 out of 100 on AI centrality, which makes it AI Level 3 of 4 (Works on AI) on this board. The level measures how much of the work is AI, not seniority.

US, WA, Seattleseniorfull-time$171k-$231k

AI in this role

Drive end-to-end delivery of GPU-accelerated servers and infrastructure powering global AI and ML workloads at cloud scale.

awsgpus
program-managementhardwarefirmwaresupply-chain
AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML workloads at cloud scale. Our team designs, builds, and operates this fleet — solving systemic hardware issues and building systems that detect and prevent recurrence so customers experience the highest quality of service.

We are seeking a Senior Technical Program Manager to drive end-to-end delivery of GPU-accelerated servers across our global fleet. You will coordinate cross-functional engineering teams spanning hardware, firmware, and software, manage ODM partnerships across multiple continents, and establish closed-loop quality systems that drive continuous improvements. This role requires technical depth to translate engineering constraints into program risk, combined with program management excellence to deliver complex hardware at global scale.

What You Will Do
You will own programs where the critical path runs through silicon, firmware, and software teams simultaneously. You will translate ambiguity into structure: turning a fleet telemetry signal into a corrective action plan with quantified failure rates, a customer requirement into a new platform milestone with EVT/DVT/PVT gates, or a manufacturing escape into a design change with updated validation criteria. You will drive decisions on program trade-offs — adjusting scope when qualification gates slip, balancing deployment speed against fleet risk, and determining when to accept interim mitigations instead of holding for root-cause fixes. When a large scale of GPU servers depend on your program landing on time, you are the one ensuring hardware readiness, qualification completeness, and operational handoff happen without gaps.

Why You Will Love It
The world's most advanced frontier models are trained on the platforms you help build. Your programs launch the GPU servers that power the largest AI/ML workloads on the planet. You will see your decisions reflected in fleet reliability metrics within weeks of deployment. The team is small and high-trust — you own programs end to end from concept through production, with direct access to leadership and engineering alike.

The Ideal Candidate
You have deep technical intuition across hardware and software — enough to challenge engineering decisions, not just track them. You thrive in ambiguity, bringing structure to programs where requirements, timelines, and dependencies are still forming. You align priorities across teams in different organizations, and you escalate with data, not noise. You actively mentor and develop others — TPMs and engineers alike. You contribute to hiring, promotion assessments, and raising the bar for program management practices in your organization.


Key job responsibilities
Strategy & Mechanisms

* Define program strategy, objectives, and success criteria; influence resource allocation and priority decisions across engineering workstreams to align with organizational goals
* Build and own mechanisms for program visibility — defining metrics, dashboards, and review cadences that enable data-driven decisions and early risk detection
* Streamline delivery processes across teams; identify and eliminate dependencies, redundant gates, or coordination overhead that slow velocity

Requirements & Planning

* Facilitate requirements gathering with internal customers; develop Technical Requirements Documents (TRDs) covering server specs, rack configurations, PCIe topology, power/cooling topology, and SKU definitions
* Build program timelines aligned to different phases (Program Initiation, Design, Qualification, Pilot, Post-launch) with critical path analysis, risk identification with new & unique changes to the hardware, and milestone tracking for the program.
* Drive Program Initiation reviews: scope definition, preliminary annualized failure rate predictions, resource planning, supply chain long-lead identification, and RFP (Request for Proposal) issuance to ODMs

Execution & Coordination

* Drive cross-functional alignment across hardware, firmware (BIOS, BMC, CPLD), software, and operations teams through design reviews, manufacturing readiness and production readiness for fleet deployment.
* Manage ODM partnerships: track EVT/DVT/PVT builds, manufacturing readiness gates, Bill of Materials (BOM) management in PLM systems, and quality checkpoints
* Identify blockers early, escalate dependencies before they impact critical path, and facilitate technical trade-off decisions across engineering workstreams.
* Communicate program status to leadership with clear reporting on milestone progress, risk posture, and mitigation plans.

Risk & Quality

* Challenge technical workstreams to surface risks early; quantify impact to schedule and reliability (annualized failure rate targets, availability SLAs) with proposed mitigations
* Drive root cause analysis of fleet-wide hardware failures and ensure corrective actions flow back into qualification criteria and design requirements
* Define acceptance criteria, coordinate qualification testing at server and rack levels, and manage go/no-go decisions for mission-critical AI/ML workloads

Transition

* Conduct knowledge transfer and document lessons learned; ensure operational readiness including automation, monitoring, and runbook completeness for production handoff

May require occasional (<10%) regional and international travel to Design and Manufacturing Partner sites.

A day in the life
You start the day syncing with ODM partners across time zones on build status and open engineering actions. Mid-morning, you run an engineering review connecting firmware, software, and hardware teams to unblock a qualification gate. In the afternoon, you triage a fleet reliability issue with operations data, drive alignment on corrective actions, and update executive stakeholders on program risk posture. You end the day reviewing NPI milestone readiness and ensuring the next design review has clear entry criteria.

About the team
The Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end — from server conception through fleet-scale operations.

Basic qualifications

- Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering or a related discipline or equivalent
- Experience managing programs across cross-functional teams, building processes and coordinating release schedules
- 6+ years of technical product or program management experience, working directly with multiple engineering teams.
- 5+ years of experience driving hardware development programs (servers, racks, networking, or storage) through full product lifecycle including design, validation, manufacturing, and fleet deployment

Preferred qualifications

- Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or related fields
- Experience leading the design, automation, deployment, and support of large-scale infrastructure
- Experience facilitating discussions with senior leadership regarding technical / architectural trade-offs, best practices, and risk mitigation
- 5+ years of experience coordinating complex server programs with ODM/JDM partners across multiple geographies, managing design reviews, manufacturing readiness gates, and quality checkpoints
- 5+ years of experience working with GPU/accelerator server platforms, NPI (New Product Introduction), or hardware qualification programs across various phases (EVT, DVT, PVT)
- Experience driving root cause analysis and corrective action processes for hardware reliability issues at fleet scale
- Knowledge of server hardware development lifecycle: electrical/mechanical design, firmware, thermal/power validation, and manufacturing testing.

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.



USA, CA, Cupertino - 171,000.00 - 231,400.00 USD annually
USA, TX, Austin - 148,700.00 - 201,200.00 USD annually
USA, WA, Seattle - 148,700.00 - 201,200.00 USD annually

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Program ManagementHardwareFirmwareSupply ChainAwsGpus

Questions you could be asked

  1. Tell me about a project where program management was part of your work. What did you do?
  2. Tell me about a project where hardware was part of your work. What did you do?
  3. Tell me about a project where firmware was part of your work. What did you do?
  4. Tell me about a project where supply chain was part of your work. What did you do?
  5. Walk me through how you've used Aws in your day-to-day work.

Adapt your resume

  • List these exact terms on your resume: Program Management, Hardware, Firmware, Supply Chain, and Aws. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Show where AI is part of your daily process, not a one-off project — this role expects it to be a running habit.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Get new AI jobs at AI Level 3+ by email

One email a week with the new AI jobs at AI Level 3+, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Product roles rated AI Level 3 at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Amazon

Related searches

Same AI level