LambdaRemote · Bellevue Office$399k-$531k
AmazonPosted 5d ago
SDE II, ML Infra Services, Annapurna Labs
SDE II, ML Infra Services, Annapurna Labs at Amazon scores 90 out of 100 on AI centrality, which makes it a Level 4 role on this board.
AI in this role
Design and implement machine learning infrastructure platforms and tooling for capacity management and workload scheduling on ML accelerators.
AWS Neuron is the complete software stack for the AWS Inferentia and Trainium cloud-scale machine learning accelerators and the Trn1 and Inf1 servers that use them. This position is for a Software Engineer that will lead the development of machine learning tools to run, optimize, and analyze machine learning workloads. This candidate must have had experience leading machine learning tool projects, preferably starting from architecture through several generations of delivery to customers. Deep knowledge of profiling and optimization, resource management, scheduling, code generation are needed. The ideal candidate will have worked on new instruction set architectures, which may include CPU, NPU, GPU and other forms of compute.
Key job responsibilities
This engineer will lead the design and implementation of ML infrastructure platform, building systems for capacity management, workload scheduling, and fleet orchestration across ML accelerators. They will work with ML scientists, training infrastructure engineers, hardware teams, and internal customers to ensure the ML Infra service delivers seamless ML Accelerator access with low wait times, high utilization, and zero-config deployment from various environments.
A day in the life
As you design and code solutions to help our team drive efficiencies in software architecture, you’ll create metrics, implement automation and other improvements, and resolve the root cause of software defects. You’ll also:
Build high-impact solutions to deliver to our large customer base.
Participate in design discussions, code review, and communicate with internal and external stakeholders.
Work cross-functionally to help drive business decisions with your technical input.
Work in a startup-like development environment, where you’re always working on the most important stuff.
About the team
* High-impact, high-visibility: You'll directly accelerate every Neuron team's ability to ship — your work multiplies the output of 100+ engineers
* Greenfield opportunities: We're actively building new capabilities with significant design ownership for SDEs
* Small, senior team: where every person owns major components and drives architectural decisions
* AI infrastructure: Work at the intersection of Kubernetes, custom silicon, and large-scale ML workloads
Diverse Experiences
We value diverse experiences and non-traditional career paths. If your career is just starting or includes alternative experiences, we encourage you to apply.
Inclusive Team Culture
Our employee-led affinity groups foster inclusion. Events like CORE and AmazeCon inspire us to embrace our uniqueness.
Work/Life Balance
We strive for flexibility as part of our working culture, supporting you both at work and at home.
Mentorship & Career Growth
We offer knowledge-sharing, mentorship, and one-on-one code reviews to help you grow as a professional.
Basic qualifications
- 3+ years of non-internship professional software development experience
- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- Experience programming with at least one software programming language
Preferred qualifications
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Bachelor's degree in computer science or equivalent
- 2+ years of building large-scale machine-learning infrastructure for online recommendation, ads ranking, personalization or search experience
- Strong proficiency in Go/Java, Python and working knowledge Javascript/TypeScript
- Experience building and operating large-scale distributed systems on Kubernetes
- Experience designing, deploying, and maintaining production services at scale, including on-call ownership
- Experience with machine learning infrastructure — orchestration, scheduling, or resource management at scale
- Proficiency in application and kernel-level performance profiling and optimization
- Experience with integrated software/hardware performance analysis in heterogeneous compute environments
- Proficiency in observability and telemetry — instrumentation, metrics collection, alarming, dashboarding, and monitoring
- Experience debugging complex issues in large-scale distributed systems and driving best practices
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, WA, Seattle - 143,700.00 - 194,400.00 USD annually
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Tell me about a project where machine learning was part of your work. What did you do?
- Tell me about a project where infrastructure was part of your work. What did you do?
- Tell me about a project where distributed systems was part of your work. What did you do?
- Tell me about a project where performance optimization was part of your work. What did you do?
- Walk me through how you've used Aws Neuron in your day-to-day work.
Adapt your resume
- List these exact terms on your resume: Machine Learning, Infrastructure, Distributed Systems, Performance Optimization, and Aws Neuron. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Similar roles
Other roles rated Level 4 at other companies.


