AmazonPosted 6d ago
Applied Scientist II, Alexa Daily Essentials Science and Analytics at Amazon scores 90 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI in this role
Design and implement agentic AI systems and evaluation frameworks for Alexa daily essentials.
As an Applied Scientist II, you will design agentic AI systems that autonomously detect, diagnose, and act on signals across Alexa's product surface. You'll build the evaluation science that validates these systems, prototype novel AI-powered features that create new value for customers, and apply reinforcement learning to improve agent behavior through outcome signals. You'll own science end-to-end — from rapid prototyping through production deployment — working directly with engineers and product leaders to ship systems that impact how millions of customers interact with Alexa daily. This is a role for a scientist who brings rigor to evaluation, creativity to invention, and pragmatism to delivery.
Key job responsibilities
- Design and implement agentic AI systems — multi-step reasoning, tool use, planning, and orchestration — that autonomously surface intelligence and take action over complex, evolving data
- Apply reinforcement learning techniques to improve agent routing, decision-making, and task resolution quality from outcome signals
- Build evaluation and benchmarking frameworks — automated test suites, LLM-as-judge pipelines, competitive benchmarks, and regression detection across model versions and prompt strategies
- Prototype and develop novel AI-powered features — information extraction, proactive recommendations, multi-agent collaboration — from concept through production launch
- Build RAG pipelines and knowledge retrieval systems that ground agent reasoning in trusted data
- Design and execute experiments end-to-end — hypothesis generation, causal analysis, and results interpretation that directly shape product roadmaps
- Communicate findings to technical and non-technical audiences — influencing architecture decisions, model selection, and product strategy through rigorous evaluation
A day in the life
You might start by analyzing reward signal distributions from your latest RL experiment on agent routing — identifying where the policy improved and where exploration is still needed. Mid-morning, you prototype a new cross-feature experience, testing whether extraction models can reliably surface actionable items from customer content. After lunch, you run your eval suite against a new model version — comparing benchmark scores across reasoning depth, tool-use reliability, and output consistency to inform an architecture decision. Later, you pair with an engineer to add a new capability to an agent's tool registry, then validate it doesn't degrade downstream quality. You end the day designing an A/B test for a proactive feature your team is launching.
About the team
We're a small, high-ownership team building agentic intelligence systems for Alexa — one of the world's most widely used AI products, reaching hundreds of millions of customers across devices, languages, and contexts. We apply foundation models in novel ways: systems that reason over temporal data, learn from outcomes through reinforcement, autonomously detect regressions, and surface insights that previously required weeks of manual analysis. Our scientists prototype new features, benchmark rigorously, and ship from research through production. If you want to build at the frontier of applied AI — agentic systems, reinforcement learning, evaluation science — with real users and real impact at Alexa scale, this is the team.
Basic qualifications
- 3+ years of building machine learning models for business application experience
- PhD, or Master's degree and 4+ years of CS, CE, ML or related field experience
- Experience in patents or publications at top-tier peer-reviewed conferences or journals
- Experience programming in Java, C++, Python or related language
- Experience with one of the following areas: machine learning technologies, Reinforcement Learning, Deep Learning, Computer Vision, Natural Language Processing (NLP) or related applications
Preferred qualifications
- Experience in professional software development
- Experience creating novel algorithms and advancing the state of the art
- Experience designing or working with agentic AI systems (tool-use, multi-step reasoning, RAG, multi-agent orchestration)
- Experience building automated evaluation frameworks or LLM-as-judge systems
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
The base salary range for this position is listed below. As a total compensation company, Amazon's package may include other elements such as sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon offers comprehensive benefits including health insurance (medical, dental, vision, prescription, basic life & AD&D insurance), Registered Retirement Savings Plan (RRSP), Deferred Profit Sharing Plan (DPSP), paid time off, and other resources to improve health and well-being. We thank all applicants for their interest, however only those interviewed will be advised as to hiring status.
CAN, BC, Vancouver - 149,300.00 - 249,300.00 CAD annually
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How would you design a retrieval step so the model answers from real data instead of guessing?
- Walk me through a computer vision problem you solved, from raw data to a deployed model.
- What NLP problem have you worked on, and how did you measure whether it actually worked?
- Tell me about a project where machine learning was part of your work. What did you do?
- Tell me about a project where applied science was part of your work. What did you do?
Adapt your resume
- List these exact terms on your resume: Rag, Computer Vision, Nlp, Machine Learning, and Applied Science. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Similar roles
Research roles rated AI Level 4 at other companies.
NVIDIAUS, CA, Santa Clara$38-$94/hr
MercorRemote · San Francisco$5000k
WaymoRemote · Mountain View, CA, USA; San Francisco, CA, USA; New York, NY, USA$213k-$263k
Anduril IndustriesBroomfield, Colorado, United States; Fort Collins, Colorado, United States$190k-$252k
More jobs at Amazon
AmazonIT, RM, Rome€50k
Meister / Techniker als Teamleiter Instandhaltung - Gattendorf bei Bayreuth / Coburg / Zwickau / Hof
AmazonDE, Gattendorf
AmazonGB, LEC, Coalville






