Sword HealthRemote · Remote - Portugal€50k-€72k
AmazonPosted 3mo ago
Post-Silicon Systems Validation Engineer, Annapurna Labs at Amazon scores 85 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI in this role
Join our Silicon Validation team to validate next-generation machine learning accelerators that power AWS's cloud computing infrastructure. You'll work in a fast-paced, startup-like environment alongside some of the brightest minds in the industry on cutting-edge, internet-scale technology that directly impacts how customers use Machine Learning acceleration. We are changing the landscape of cloud infrastructure by accelerating the development of custom silicon by moving beyond traditional partnerships to dominate in AI training and inference
Your work will span validation of the complete vertical stack—silicon, PCB, high-speed components (HBM, PCIe, chip-to-chip), inter-system connections, and system-to-system interfaces. You'll dive deep into new technology hardware components and scaling technologies that power our Machine Learning boards and servers at scale, ensuring every component of our hardware and software comes together into products our customers rely on.
Key job responsibilities
As a Validation Engineer on our Machine Learning Acceleration team, you'll own critical validation aspects across the entire product development lifecycle—from early design validation through emulation, silicon bring-up, post-silicon validation, and ongoing support of production systems deployed in AWS data centers. You'll collaborate deeply with architecture, RTL design, design verification, firmware, and software teams to ensure our next-generation AI/ML accelerators meet the highest standards of quality and performance. This role requires bridging multiple domains—from low-level hardware interfaces to high-level ML workloads—to deliver exceptional results.
We are looking for candidates with:
- Strong programming skills (Python, Lua, C/C++, Rust, Go, etc)
- A solid understanding of computer architecture
- Experience with AWS services, cloud infrastructure, firmware development (BIOS, BMC, drivers)
- Validation experience in any of these areas: PCIe, HBM, GPUs, neural networks, ML HW architecture, and/or CI/CD
- Familiarity with the validation lifecycle from RTL simulation (SystemVerilog/UVM, VCS, Questa, Xcelium) and emulation (Palladium, Zebu, Veloce) through silicon failure analysis and debug
A day in the life
- Developing comprehensive validation strategies and detailed test plans covering functional, performance, power, and stress testing from silicon bring-up to product release
- Executing complex test plans from RTL simulation and emulation environments through physical silicon validation
- Conducting hands-on silicon bring-up and debug in the lab using oscilloscopes, logic analyzers, and protocol analyzers
- Validating ML accelerator performance, accuracy, and reliability using real-world neural network workloads
- Building test infrastructure, CI/CD, and automated regression frameworks to enable efficient validation at scale
- Collaborating across architecture, design, firmware, and software teams to triage failures and drive root cause analysis to closure
- Reviewing test results, identifying patterns, and providing feedback to improve design quality and validation coverage
- Supporting production systems in AWS data centers and addressing field issues as they arise
Basic qualifications
- 3+ years of non-internship professional software development experience
- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Experience with Machine Learning and Large Language Model fundamentals, including architecture, training/inference lifecycles, and optimization of model execution, or experience working with PyTorch or JAX software
- Bachelor's degree in computer science, engineering, mathematics or equivalent, or experience in Java, C++, Python, or a related language
- 3+ years of experience with hardware performance counters and profiling tools for analyzing and optimizing system and application performance
- Strong understanding of computer architecture fundamentals including memory hierarchies (caches, DRAM, HBM), compute pipelines, and interconnect topologies
- Experience applying statistical methods, regression analysis, and data visualization techniques to interpret performance data and drive optimization decisions
Preferred qualifications
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Bachelor's degree in computer science or equivalent
- Experience with Machine Learning Hardware/Software Architecture
- Experience with CI/CD
- Experience with EDA Simulations or Emulation
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, TX, Austin - 143,700.00 - 194,400.00 USD annually
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- What's a project where you used PyTorch hands-on?
- Walk me through how you've used Jax in your day-to-day work.
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: PyTorch and Jax. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Similar roles
Software Engineering roles rated AI Level 4 at other companies.
GitLabRemote · Bangalore, India
NotionRemote · San Francisco, California$255k-$300k
Together AISan Francisco$260k-$300k
LambdaRemote · Bellevue Office$399k-$531k
More jobs at Amazon
AmazonIT, RM, Rome€50k
Meister / Techniker als Teamleiter Instandhaltung - Gattendorf bei Bayreuth / Coburg / Zwickau / Hof
AmazonDE, Gattendorf
AmazonGB, LEC, Coalville




