Associate Director, Evaluation and Quality Standards
AI in this role
At Johnson & Johnson, we believe health is everything. Our strength in healthcare innovation empowers us to build a world where complex diseases are prevented, treated, and cured, where treatments are smarter and less invasive, and solutions are personal. Through our expertise in Innovative Medicine and MedTech, we are uniquely positioned to innovate across the full spectrum of healthcare solutions today to deliver the breakthroughs of tomorrow, and profoundly impact health for humanity. Learn more at jnj.com.
As guided by Our Credo, Johnson & Johnson is responsible to our employees who work with us throughout the world. We provide an inclusive work environment where each person is considered as an individual. At Johnson & Johnson, we respect the diversity and dignity of our employees and recognize their merit.
Job Function:
Data Analytics & Computational SciencesJob Sub Function:
Data ScienceJob Category:
People LeaderAll Job Posting Locations:
Cornellà de Llobregat, Barcelona, Spain, Madrid, SpainJob Description:
Evaluation & Quality Standards to join our Data Science and Digital Health team (DSDH). This is a newly created leadership role within the Generative AI organization, reporting directly to the Head of Generative AI.
The GenAI team is deploying AI across discovery, development, regulatory, and operations to create AI assistants and agentic systems across the therapeutic areas and functions. As we scale from proof-of-concept to enterprise deployment, we need a dedicated evaluation function that defines what quality means for generative AI in a regulated pharmaceutical R&D environment, communicates the metrics effectively, and enforces that standard across every platform, vendor, and capability we build or adopt. This role owns evaluation and governance for all generative AI work across R&D.
Key ResponsibilitiesEvaluation Framework Design
• Design and maintain evaluation frameworks for GenAI solutions spanning LLM quality, RAG performance, agent reliability, safety, and scientific accuracy.
• Develop therapeutic-area-specific evaluation criteria with business teams, reflecting domain-specific quality requirements per therapeutic area and application type.
• Define benchmarks for scientific validity, regulatory compliance, data quality, and operational reliability.
• Establish gold-standard validation datasets and automated evaluation pipelines.
Vendor and Solution Assessment
• Lead structured vendor assessments producing written Evaluation Reports
• Challenge vendor claims with independent benchmarking; distinguish genuine capability from demonstration performance.
• Assess integration complexity, total cost of ownership, regulatory posture, and vendor lock-in risk.
Governance and Standards
• Serve as technical lead for the Evaluation & Standards Board, a cross-functional governance body for GenAI quality.
• Present evaluation findings and recommendations to the GenAI Portfolio Steering Committee.
• Define what “production-ready” means for GenAI in a regulated R&D environment and enforce that standard.
• Set evaluation gates for product development stages: PoC, Limited Release, Scaled Product, and Product Ops.
Cross-Functional Collaboration
• Partner with QMS/MLOps on the handoff from pre-adoption evaluation to production quality monitoring.
• Work with therapeutic area teams on domain-specific evaluation standards.
• Collaborate with Data Strategy & Products on data quality as an input to the evaluation rubric.
• Engage JJIT architecture teams on platform-layer evaluation and security assessments.
Opinion Leadership and Capability Building
• Contribute to the weekly GenAI Outlook newsletter, specifically the evaluation implications of frontier AI developments.
• Build internal evaluation capability through training, documentation, and tooling for teams conducting assessments.
• Publish evaluation standards and rubrics as reference material for the broader organization.
Team Leadership
• Build and lead the GenAI Evaluation & Standards team
• Attract, develop, and retain top talent in AI evaluation, quality science, and governance.
• Establish a culture of scientific rigor and independent judgment within the team.
What This Role Is Not
• Not MLOps or production monitoring: the QMS function handles post-deployment operations.
• Not platform engineering.
• Not regulatory affairs or medical writing.
• Not vendor management or procurement: evaluation is independent of purchasing.
Key Qualifications
• Advanced degree (PhD strongly preferred) in computational biology, bioinformatics, data science, computer science, AI/ML, biomedical engineering, applied mathematics, or related discipline.
• Minimum 8 years of post-academic industry experience, with significant time in pharmaceutical or biotechnology R&D.
• Hands-on expertise with generative AI systems: large language models, retrieval-augmented generation, agentic frameworks, and prompt engineering.
• Demonstrated track record designing evaluation frameworks, scientific benchmarks, or quality standards for AI/ML systems.
• Strong people leadership experience, including building and managing technical teams in a matrixed organization.
• Excellent communication skills: ability to present complex technical findings clearly to non-technical senior stakeholders and defend unpopular conclusions.
• Strategic thinking capability: ability to connect technical evaluation to business impact and organizational decision-making.
Required Skills:
Preferred Skills:
Advanced Analytics, Budget Management, Compliance Management, Critical Thinking, Data Analysis, Data Privacy Standards, Data Quality, Data Reporting, Data Savvy, Data Science, Data Visualization, Developing Others, Digital Fluency, Inclusive Leadership, Leadership, Program Management, Strategic Thinking, Succession Planning
The anticipated base pay range for this position is:
€75,000.00 - €129,260.00
Benefits:
In addition to base pay, we offer the following benefits*: an annual bonus with set target (% of pay) depending on pay grade / location, where the actual amount is based on the employees’ and companies’ performance of the previous calendar year, or sales commissions. Moreover, we offer vacation days, parental leave for a minimum of 12 weeks, bereavement leave, caregiver leave, volunteer leave, well-being reimbursement, programs for financial, physical and mental health. We also offer service anniversary and recognition awards, and subject to the terms of their respective plans, employees - and in some location’s eligible dependents - can participate in several insurance plans. For more information, visit Employee benefits | Supporting well-being & career growth | Johnson & Johnson Careers.
*This is for informative purposes only. Amounts and actual benefits may vary by location and are subject to change.
How we rate this
Associate Director, Evaluation and Quality Standards at Johnson & Johnson rates 83 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you structure and test a prompt to get consistent output from a language model?
- How would you design a retrieval step so the model answers from real data instead of guessing?
- How do you monitor a model once it's live, and how do you know it needs retraining?
- How do you decide that one model's output is better than another's for a given task?
- How would you decide a model or AI system is ready to ship?
Adapt your resume
- List these exact terms on your resume: Prompt Engineering, RAG, ML Ops, and AI Evaluation. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new AI jobs (Builds AI ●●●●) by email
One email a week with the new AI jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Other roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step