AmazonPosted 2mo ago
AI Language Engineer II, Alexa for Shopping Lang-Tech at Amazon scores 69 out of 100 on AI centrality, which makes it AI Level 3 of 4 (Works on AI) on this board. The level measures how much of the work is AI, not seniority.
AI in this role
We are looking for candidates who are passionate about the intersection of language and technology and who are keen to use their technical abilities to develop automated, scalable solutions to challenges in the Large Language Model (LLM) space. Applying a combination of expertise in LLMs, coding, and natural language, they will overcome complex problems in model evaluation, automation, and context engineering for agentic systems.
In this role within the AI Shopping org, they will contribute to our evaluation-driven product development strategy. They will work in close collaboration with Product Managers, Applied Scientists, Software Engineers, UX Researchers, and Editors on initiatives that drive quality, speed and consistency. They will be responsible for authoring, optimizing, and managing system prompts for customer-facing AI driven shopping experiences on both the mobile and web apps. They will define requirements for internal tooling by developing prototypes. They will employ their data processing and analysis skills to evaluate and report on model performance, producing insights that inform product decisions. By creating and synthesizing quality metrics, they will also support Conversational Shopping teams in delivering both internal stakeholder requirements and achieve the desired Amazon customer outcomes.
This role requires strong analytical and technical skills as well as experience in language technology to help us measure, analyze and solve complex problems. The candidate should have experience in creating technical solutions for automating and processing data workflows at scale and have the ability to do so while upholding the highest linguistic quality standards. They should also have exceptional writing and communication skills with the ability to interface between both technical and non-technical teams.
Key job responsibilities
- Develop LLM-as-a-judge systems to support Human-in-the-loop evaluations
- Automate operations and perform data analysis using scripting languages (e.g. Python)
- Author, optimize, and manage system prompts for customer-facing LLM systems
- Integrate API calls into Retrieval Augmented Generation (RAG) systems
- Evaluate model performance and annotation quality to produce reports for stakeholders
- Produce, process, and manipulate different types of language data
- Contribute to defining platform requirements for internal tooling by developing prototypes
- Raise the quality bar on editorial workflows and SOPs through standardization, documentation, and periodic audits and investigations
- Support processes and mechanisms to onboard and upskill Editors and AI Tutors on an ongoing basis
- Design, implement, and refine control mechanisms, metrics, and methodologies to ensure editorial and annotation quality
- Collaborate with editors, applied scientists, engineers, and product managers to deliver an optimal customer experience by defining metrics, guidelines, and workflows
- Deliver across parallel workstreams, balancing timelines, impact, and stakeholder requirements
About the team
The CMX-Lang-Tech team is a technical sub-team of the AI Shopping Content team. We are responsible for AI response quality both in terms of evaluating and prompting our AI models.
Basic qualifications
- Experience with Unix tools
- Experience that includes strong analytical skills, attention to detail, and effective communication abilities
- Experience in a fast paced, dynamic organization
- Experience prioritizing and handling multiple assignments at any given time while maintaining commitment to deadlines
- Master's Degree in Applied Linguistics, Computational Linguistics, Natural Language Processing (NLP), or related field.
- Experience with Large Language Models, NLP, or Machine Learning.
- Experience with Python libraries for data analysis such as pandas and scikit-learn.
- Familiarity with AI coding assistants.
Preferred qualifications
- PhD in Applied Linguistics, Computational Linguistics, Natural Language Processing (NLP), or related technical field.
- Experience with SQL and Git.
- Experience building RAG or agentic systems.
- Experience conducting quantitative analysis.
- Experience building data pipelines.
- Experience with AWS services (Bedrock, S3, EC2, etc.).
- Knowledge of user experience concepts and methods.
- Familiarity with online retail (e-commerce).
Amazon is an equal opportunities employer. We believe passionately that employing a diverse workforce is central to our success. We make recruiting decisions based on your experience and skills. We value your passion to discover, invent, simplify and build. Protecting your privacy and the security of your data is a longstanding top priority for Amazon. Please consult our Privacy Notice (https://www.amazon.jobs/en/privacy_page) to know more about how we collect, use and transfer the personal data of our candidates.
Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How would you design a retrieval step so the model answers from real data instead of guessing?
- How do you decide that one model's output is better than another's for a given task?
- What NLP problem have you worked on, and how did you measure whether it actually worked?
- What's a project where you used Bedrock hands-on?
- Walk me through how you've used scikit-learn in your day-to-day work.
Adapt your resume
- List these exact terms on your resume: Rag, AI Evaluation, Nlp, Bedrock, and scikit-learn. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Show where AI is part of your daily process, not a one-off project — this role expects it to be a running habit.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Similar roles
Software Engineering roles rated AI Level 3 at other companies.
LambdaRemote · San Francisco Office (Fremont St)$314k-$465k3h ago
OpenAIRemote · San Francisco$180k-$260k5h ago
Gong ioAustin | Chicago | New York City | Salt Lake City | San Francisco$142k-$205k8h ago
Solve IntelligenceNew York$100k-$220k9h ago
n8nRemote · Europe9h ago
SmartsheetRemote · -REMOTE, USA-$161k-$194k9h ago
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step






