Level

AmazonPosted 6d ago

Principal Applied Scientist, AAIS

Principal Applied Scientist, AAIS at Amazon scores 100 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.

US, WA, Seattleleadfull-time$199k-$269k

AI in this role

Own the scientific strategy and direction for advanced AI assistants, focusing on organizational knowledge representation, retrieval, and agentic behavior.

python
machine-learningknowledge-graphsinformation-retrievaltemporal-reasoningagentic-systems
AI assistants are getting genuinely good at remembering individuals: your preferences, your projects, the thread you left open last week. But that memory stops at the edge of one person's usage. It doesn't reach the level at which real work happens, where the knowledge that matters is spread across many people, where one person's decision changes what everyone else should do next, and where nobody has the full picture. We're building AI that operates at that level: a durable, accurate understanding of how a team works, used to make that team measurably faster.

We are looking for a Principal Applied Scientist to own the scientific direction of that work. This is a broad, ambiguous, high-leverage charter. The problems span knowledge representation, temporal reasoning, retrieval, agentic behavior, and the measurement science needed to know whether any of it is working. You will not be handed a well-posed problem. You will decide which problems are worth posing.

This is a science leadership role, not a solo research role. You will set direction and raise the scientific bar across a team of applied scientists and MLEs, while staying deep enough in the work to prototype an idea yourself and prove it on real data.

Key job responsibilities
Own the scientific strategy for how organizational knowledge is represented, kept current, and retrieved: extraction, entity resolution, deduplication, graph structure, and retrieval that unifies graph, semantic, keyword, and temporal search.

Advance temporal reasoning. Knowledge changes: facts are revised, decisions are reversed, priorities move. Representing what superseded what and when, and preserving the provenance to distinguish confirmed information from inferred information, is among the hardest open problems in this space.

Define the science of proactive behavior. When is it right for an AI system to interrupt a human? These are precision-critical problems where a false positive costs far more than a miss, and where the right threshold varies by team and by individual.

Lead our measurement science. Build evaluation for completeness and correctness across a multi-component agentic system, converging on a small number of trustworthy primary metrics rather than a sprawl of component scores. Judge honestly when an offline gain is real and when it is an artifact of a sparse dataset.

Build the data that doesn't exist. The most valuable phenomena in this domain are also the rarest, which makes naturally occurring examples too scarce to learn from. Design synthetic and simulated data pipelines that generate controlled, realistic scenarios so these capabilities can be developed and tested at all.

Own the learning loop. Turn human interaction into usable training signal, and set the direction for how the system improves from explicit feedback in the near term and from passive observation over the longer term.

Make the efficiency calls. Decide where frontier models are required and where a smaller domain-tuned model is sufficient, and build the cost and capacity measurement that makes it a data-driven decision rather than an opinion.

Raise the bar across the team. Mentor scientists, review designs, publish where the work merits it, and represent the science externally to customers and to the research community.

A day in the life
You might spend the morning in a design review arguing that a proposed approach won't survive contact with real data, the afternoon writing a prototype yourself to demonstrate the alternative, and the end of the day convincing an engineer that the capability is worth a sprint. Our sequencing is deliberate: try the idea on intuition, validate it on real data by inspection, then measure it, then operationalize it. Scientists here are expected to identify a problem, justify it, recruit others to it, and drive it into production, across whatever parts of the system that requires. Ownership follows the problem, not the org chart.

About the team
We are a combined science, product, and engineering team building one product together. Scientists own capabilities end to end rather than individual components, because these problems don't decompose cleanly: a single improvement typically touches extraction, storage, and retrieval at once. We invest in the tooling that makes that practical: local full-stack environments and sandboxed realistic data, so a scientist can go from idea to result in seconds rather than waiting on a deployment or on engineering support.

The work is grounded in real usage rather than benchmarks alone, which is a rare combination for science this early: real users, real data, real feedback, and a genuinely unsolved research agenda.

Basic qualifications

- PhD in Computer Science, Machine Learning, Statistics, or a related quantitative field; or a Master's degree with 8+ years of applied science experience
- 10+ years of experience building and shipping machine learning or AI systems that reached production users
- Deep expertise in large language models and at least two of: information retrieval, knowledge representation and graphs, reinforcement learning, agentic system design, or evaluation methodology for generative systems
- Demonstrated experience setting technical and scientific direction for a team of scientists, including mentoring senior scientists
- Hands-on proficiency in Python and the ability to prototype independently in a production codebase
- Track record of publications, patents, or equivalent evidence of original scientific contribution

Preferred qualifications

- Experience with agentic and multi-turn systems, including RL-based post-training, environment simulation, or agent harness evaluation
- Experience designing evaluation frameworks for open-ended or subjective tasks where ground truth is expensive or unavailable, including synthetic data generation
- Experience with memory, personalization, or long-horizon context systems for LLM applications
- Experience with temporal knowledge representation, entity resolution, or knowledge graph construction at scale
- Experience taking a product from prototype to launch under ambiguity, including making the judgment call on when quality is sufficient to ship
- Experience with model distillation or domain-specific tuning to reduce inference cost
- Scientific breadth across multiple ML domains, and comfort operating outside your original specialization

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.



USA, WA, Seattle - 198,900.00 - 269,000.00 USD annually

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Machine LearningKnowledge GraphsInformation RetrievalTemporal ReasoningAgentic SystemsPython

Questions you could be asked

  1. Tell me about a project where machine learning was part of your work. What did you do?
  2. Tell me about a project where knowledge graphs was part of your work. What did you do?
  3. Tell me about a project where information retrieval was part of your work. What did you do?
  4. Tell me about a project where temporal reasoning was part of your work. What did you do?
  5. Tell me about a project where agentic systems was part of your work. What did you do?

Adapt your resume

  • List these exact terms on your resume: Machine Learning, Knowledge Graphs, Information Retrieval, Temporal Reasoning, and Agentic Systems. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Similar roles

Research roles rated AI Level 4 at other companies.

More jobs at Amazon

More applied scientist jobs

Related searches

Same AI level