HelloFreshNew York, NY, United States$190k-$233k9h ago
Scale AIPosted 1mo ago
Senior AI Product Manager at Scale AI scores 75 out of 100 on AI centrality, which makes it AI Level 3 of 4 (Works on AI) on this board. The level measures how much of the work is AI, not seniority.
AI in this role
Lead AI product portfolios for coding benchmarks and leaderboards, shaping evaluation standards and driving adoption at Scale AI.
Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including frontier model training, enterprise adoption, defense applications, and autonomous vehicles. Our mission is to develop reliable AI systems for the world's most important decisions.
Scale's evaluation and data products shape how the industry trains and measures frontier models. Our benchmarks and leaderboards influence model development and purchasing decisions across the AI ecosystem, and our data and environments help leading labs build what comes next.
These roles are deeply cross-functional. You will work across AI Product Management, ML Research, Engineering, Operations, and Go-To-Market teams, and directly with frontier AI labs and enterprise customers, representing Scale as a thought leader in how models are trained and measured. Whichever portfolio you own, you will own it end-to-end: strategy, roadmap, infrastructure, governance, customer adoption, and business impact.
PMs here operate with significant autonomy, ship frequently, and are expected to be deeply analytical and hands-on. The ideal candidate combines strong product judgment, technical fluency, operational rigor, and customer-facing experience, with a passion for turning emerging model capabilities into scalable products.
Open Roles
- Senior AI PM, Coding: Own Scale's Coding portfolio, built on SWE-Bench Pro (which cut frontier scores from 70%+ to roughly 23%), SWE Atlas, and our FrontierBench contributions. You'll decide what comes next as agents saturate current tasks, and turn that benchmark credibility into revenue from SFT and preference data, RL environments, and agentic coding evals. You'll also own the coding infrastructure roadmap and our expert network of professional software engineers. This role requires hands-on software engineering depth.
- Senior AI PM, Leaderboards: Own and scale the SEAL Leaderboard portfolio across all of Scale's domains, turning cutting-edge evals into trusted industry benchmarks. You'll run the Leaderboard Steering Committee, decide which leaderboards get built, and set governance standards for evaluation integrity and release cadence. This is a horizontal role, so breadth across evaluation categories and a strong public-facing presence matter more than depth in any one domain.
You Will
- Own the roadmap and strategy for your portfolio, setting priorities across product development, launches, infrastructure investments, and expansion.
- Drive alignment across AI-PM, ML, Engineering, Operations, and GTM stakeholders.
- Evaluate, prioritize, and operationalize new product proposals, making sure they align with customer demand, model capability frontiers, and company strategy.
- Define and manage the end-to-end product lifecycle, from ideation and design to launch, scaling, maintenance, and sunset decisions.
- Partner with ML researchers and domain experts to develop trustworthy evaluation methodologies, task specifications, scoring and grader design, and quality bars.
- Drive the roadmap for infrastructure, automation, and operational tooling to improve scalability, reduce manual effort, and accelerate delivery.
- Establish governance processes for quality, evaluation integrity, contamination prevention, reproducibility, auditability, and release management.
- Work directly with frontier AI labs and enterprise customers to understand where their models fail, gather feedback, and shape future investments.
- Track adoption, usage, model impact, and business outcomes, using data to guide roadmap decisions and resource allocation.
- Identify new capability areas and product opportunities that strengthen Scale's leadership and create new revenue.
- Collaborate closely with GTM on customer engagements, thought leadership, product launches, and strategic partnerships.
Ideally, You'd Have
- 5+ years of experience in product management, technical program management, consulting, or customer-facing technical roles.
- Strong technical fluency with machine learning systems, AI model evaluation, benchmarking, or data products.
- Experience building and scaling products that require coordination across engineering, operations, and business teams.
- Excellent stakeholder management and executive communication skills, with a proven ability to drive alignment across cross-functional organizations.
- Strong analytical skills and the ability to translate ambiguous market and customer signals into clear product strategy.
- Experience working with AI researchers, ML teams, evaluation frameworks, developer tools, or human-in-the-loop data pipelines is strongly preferred.
- Entrepreneurial mindset with a track record of creating new products, programs, or business initiatives from the ground up.
- Bias for action and comfort operating in fast-moving, ambiguous environments.
- Passion for advancing trustworthy AI and helping define how the industry builds and measures frontier model capabilities.
- For the Coding role: strong software engineering depth, plus familiarity with how coding models are trained and evaluated (post-training methods, agentic scaffolds and harnesses, benchmarks like SWE-Bench and Terminal-Bench).
Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend.
Please reference the job posting's subtitle for where this position will be located. For pay transparency purposes, the base salary range for this full-time position in the locations of San Francisco, New York, Seattle is:$205,600—$257,000 USDPLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants.
About Us:
At Scale, our mission is to develop reliable AI systems for the world's most important decisions. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact. We work closely with industry leaders like Meta, Ernst & Young, Mayo Clinic, Time Inc., the Government of Qatar, and U.S. government agencies including the Army and Air Force. We are expanding our team to accelerate the development of AI applications.
We believe that everyone should be able to bring their whole selves to work, which is why we are proud to be an inclusive and equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability status, gender identity or Veteran status.
We are committed to working with and providing reasonable accommodations to applicants with physical and mental disabilities. If you need assistance and/or a reasonable accommodation in the application or recruiting process due to a disability, please contact us at [email protected]. Please see the United States Department of Labor's Know Your Rights poster for additional information.
We comply with the United States Department of Labor's Pay Transparency provision.
PLEASE NOTE: We collect, retain and use personal data for our professional business purposes, including notifying you of job opportunities that may be of interest and sharing with our affiliates. We limit the personal data we collect to that which we believe is appropriate and necessary to manage applicants’ needs, provide our services, and comply with applicable laws. Any information we collect in connection with your application will be treated in accordance with our internal policies and programs designed to protect personal data. Please see our privacy policy for additional information.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you decide that one model's output is better than another's for a given task?
- How do you decide which parts of a product should use AI versus a fixed set of rules?
- Tell me about a project where ai product management was part of your work. What did you do?
- Tell me about a project where llm evaluation was part of your work. What did you do?
- Tell me about a project where benchmarking was part of your work. What did you do?
Adapt your resume
- List these exact terms on your resume: AI Evaluation, AI Product, AI Product Management, Llm Evaluation, and Benchmarking. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Show where AI is part of your daily process, not a one-off project — this role expects it to be a running habit.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Similar roles
Product roles rated AI Level 3 at other companies.
TruvetaRemote · Hyderabad, India10h ago
AmazonUS, WA, Seattle$152k-$206k18h ago
DatabricksRemote · Remote - California; Remote - Virginia; Remote - Washington D.C.$182k-$250k21h ago
PinterestRemote · San Francisco, CA, US; Remote, US1d
DatadogNew York, New York, USA$192k-$240k1d



