AI Benchmarking Lead, Performance Benchmarking Evaluation
AI in this role
Lead AI model evaluations and benchmarking for Amazon's Seller Assistant, ensuring data quality and reliability across international markets.
About Seller Assistant
Seller Assistant is a conversational AI copilot that understands the full context of a seller's business. It intelligently orchestrates back end tools to deliver actionable, drilled-down responses and can independently complete complex tasks on behalf of sellers with their permission.
Our Scale and Impact:
- Expanded to 2.44MM sellers (45x growth vs. Dec 2024)
- Currently serving 61% of active sellers worldwide across 9 international stores (CN2XX, IN, UK, DE, JP, BR, MX, AE, SA)
- Supporting four languages: English, Chinese, German, and Japanese
- 2026 Goal: Scale to 90%+ active sellers WW with 5 new store launches (France, Italy, Spain, Canada, Australia)
As a AI Benchmarking Lead, you will benchmark Seller Assistant AI models for relevancy, correctness, and completeness. Your primary responsibilities include: 1) Evaluate audits performed by the core auditing team to increase confidence in evaluation metrics, 2) Improve audit reliability and consistency through systematic measurement of auditor accuracy,3) Conduct targeted calibration to ensure quality standards across the auditing function, 4) Enforce quality standards by quality-checking audits and providing actionable feedback to team members, 5) Drive continuous improvement in audit processes and methodologies.
- You conduct quality checks on audits performed by the core auditing team.
- You identify rubric gaps and evaluation ambiguities that lead to inconsistent audit outcomes.
- You surface high-confidence product issues earlier by validating and categorizing model failures.
- You serve as point of contact for annotation tasks across ML data process areas, ensuring quality execution and delivery
- You understand dependencies across ML data workflows and articulate customer impact effectively
- You modify existing annotation methods and update SOPs.
- You document SOP changes, secure approval, share knowledge with the team, and audit adoption and execution
- You test new SOPs and tools, providing feedback on quality and improvement recommendations to support onboarding
Key job responsibilities
- You structure data collection, analyse results and share inputs for SOP changes.
- You collate, track, and report progress on key metrics agreed to with respective stakeholders (e.g., Program managers, Applied Scientist) specific to your functional area.
- You identify operational issues related to process and tooling and recommend suggestions to improve key project metrics such as productivity and quality.
Basic qualifications
- Bachelor's degree or equivalent in a related field
- Experience in natural language data labeling, data annotation, linguistic annotation or other forms of data markup
- Technical Skills: Proficiency in MS Excel; basic understanding of SQL and Python
- Experience with Microsoft Office products and applications
- Communication Skills: Strong verbal and written communication skills in English
- Knowledge about SOA and process that deal with sellers.
Preferred qualifications
- 1 to 3 years of equivalent experience
- Performed annotation related tasks across ML data process areas.
- Strong knowledge of process documentation, analysis knowledge
- Technical proficiency in SQL querying and Python programming for data analysis
- Strong analytical and problem-solving skills
- Ability to work independently and as part of a team
Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.
How we rate this
AI Benchmarking Lead, Performance Benchmarking Evaluation at Amazon rates 65 out of 100 for how much of the daily work is AI. That makes it Works on AI (AI Level 3 of 4). The level is about AI in the job, not seniority.
Works on AI. The daily work is on AI products, without building the model.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you keep labeling instructions consistent across a large annotation team?
- How do you decide that one model's output is better than another's for a given task?
- Tell me about a project where data annotation was part of your work. What did you do?
- Tell me about a project where benchmarking was part of your work. What did you do?
- Tell me about a project where quality assurance was part of your work. What did you do?
Adapt your resume
- List these exact terms on your resume: AI Data Labeling, AI Evaluation, Data Annotation, Benchmarking, and Quality Assurance. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Show where AI is part of your daily process, not a one-off project — this role expects it to be a running habit.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new AI jobs (Works on AI ●●●○ or higher) by email
One email a week with the new AI jobs (Works on AI ●●●○ or higher), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Other roles that work on AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step