Associate Director, Clinical AI Evaluation and Responsible Deployment
AstraZeneca is hiring an Associate Director, Clinical AI Evaluation and Responsible Deployment in Barcelona, Spain. Level rates it ; you can apply on Level.
AI in this role
Associate Director, Clinical AI Evaluation and Responsible Deployment
About AISI
AI Science & Innovation (AISI) sits at the centre of AstraZeneca’s R&D AI transformation within Enterprise AI Unit (EAI). Our remit is to build, buy, and deliver the AI models and agents that change pipeline outcomes across discovery, translational science, biomarkers, and clinical development, ultimately improve patients’ lives.
We are building an end-to-end Enterprise AI engine that unites data foundations, AI models, platforms, and business-facing applications to accelerate results across the value chain. Success comes from reusing what already works, sharing ideas across teams, and scaling impact rather than building in isolation.
Role overview
AstraZeneca is building an AI capability for Clinical Development that will improve how trials are designed, conducted, monitored, and analysed. We are hiring an Associate Director, Clinical AI Evaluation and Responsible Deployment to create the evidence systems that determine whether clinical AI is useful, reliable, safe, and ready to scale.
The role sits in the Applied Clinical AI team and engages directly with clinical stakeholders, Engineering, IT, and data teams to shape priorities, requirements, and delivery in collaboraiton with AstraZeneca’s existing Engineering, Product and Clinical Solutions teams.
Clinical AI rarely has simple ground truth. Expert judgments can differ, source data can be incomplete, and the consequence of an error depends on where an output appears in the workflow. You will turn these realities into rigorous evaluation environments, release criteria, monitoring strategies, and validation-ready evidence. You will work directly with clinical stakeholders to define what “good” means and directly with scientists and engineers to implement it in code.
This is a hands-on technical leadership role. You will build evaluation harnesses, analyse model and workflow behaviour, design experiments, and help teams diagnose failures—not merely review documents after development is complete. You will also connect scientific evaluation with Quality, Regulatory, and operational expectations so that evidence is useful both to builders and decision-makers.
What you’ll do
- Define and implement the evaluation strategy for clinical AI models and agents across development, release, monitoring, and change control.
- Build evaluation harnesses, curated test sets, simulation environments, automated regression suites, and analysis pipelines in Python and related technologies.
- Translate expert judgment from CRAs, medical monitors, clinical scientists, statisticians, and operations leaders into task definitions, scoring rubrics, error taxonomies, and clinically meaningful acceptance thresholds.
- Evaluate complete workflows—not only model outputs—including retrieval quality, tool selection, orchestration, source fidelity, abstention, human hand-offs, latency, and downstream operational impact.
- Design approaches for noisy, sparse, delayed, or expert-dependent ground truth, including adjudication, inter-rater agreement, challenge sets, prospective studies, and post-deployment surveillance.
- Lead failure analysis and red-teaming for hallucination, unsupported claims, automation bias, data leakage, prompt injection, subgroup performance, and unsafe workflow behaviour.
- Establish traceability from intended use and user need through requirements, test evidence, and release decisions, including how model, prompt, tool, data, and workflow changes are assessed, monitored, and revalidated throughout the product lifecycle.
- Work with Quality, Regulatory, Clinical Operations, Privacy, Security, and R&D IT to align evaluation evidence with GCP, GxP, data-integrity, validation, and inspection-readiness expectations.
- Advise teams on when evidence supports progression from prototype to controlled pilot, broader deployment, or regulated use—and when it does not.
- Create reusable evaluation components and standards that can be adopted across agentic workflows for improving clinical operations and the wider Clinical Development AI portfolio.
- Communicate findings and residual risks clearly to technical, clinical, quality, and executive audiences; ensure uncertainty is visible rather than hidden behind aggregate metrics.
- Mentor scientists and engineers in rigorous experimentation, reproducibility, and responsible clinical AI development.
Essential for the role
- PhD in Machine Learning, Computer Science, Statistics, Biostatistics, Biomedical Informatics, Computational Biology, or a related quantitative discipline; or an MD, master’s degree, or equivalent experience with a strong computational record.
- 4 to 7 years of post-PhD (or equivalent) experience evaluating, validating, or deploying AI/ML systems in healthcare, life sciences, clinical research, or another safety- or evidence-critical environment.
- Current, hands-on coding ability in Python and SQL, including experience building automated evaluation pipelines, analysing large datasets, writing tests, and working in version-controlled environments.
- Strong grounding in experimental design, statistical inference, uncertainty, error analysis, and measurement reliability.
- Experience evaluating LLMs or agentic systems, including retrieval, tool use, structured outputs, multi-step workflows, robustness, and human-in-the-loop performance.
- Demonstrated ability to construct useful evaluation approaches when labels are noisy, expert opinions differ, or outcomes are delayed.
- Experience producing audit-ready evidence under GCP and GxP, including model-risk management, audit trails, and inspection readiness for regulated AI systems.
- Ability to work directly with clinical stakeholders to define intended use, harmful failure modes, decision thresholds, and acceptable human oversight.
- Deep understanding of trustworthy AI principles and lifecycle governance, and the ability to turn them into concrete release and monitoring criteria.
- Excellent communication skills and the judgment to explain technical evidence and residual risk to clinical, quality, regulatory, and executive audiences.
- Experience leading cross-functional technical work through influence, with the judgment to make evidence-based go or no-go recommendations under ambiguity.
Desirable for the role
- Direct experience with clinical trial conduct, medical monitoring, pharmacovigilance, or clinical data management in a pharmaceutical, biotech, CRO, or health-system environment.
- Experience conducting prospective, silent-mode, shadow-mode, or human-factors evaluations in live clinical or healthcare workflows.
- Familiarity with causal inference, calibration, subgroup analysis, weak supervision, active learning, or methods for learning with imperfect labels.
- Experience with LLM evaluation platforms, observability, MLOps/LLMOps, reproducible experimentation, and production monitoring.
- Peer-reviewed publications, standards work, open-source contributions, or regulator/industry-consortium engagement related to AI evaluation or medical AI.
What success looks like
- Every clinical AI release has clear intended use, measurable acceptance criteria, and traceable evidence.
- Clinical experts recognize the evaluation as representative of real work and real failure modes.
- Scientists and engineers can detect regressions quickly and diagnose why a system failed.
- Quality and regulatory partners are engaged early, with evidence generated by design rather than assembled after the fact.
- Evaluation assets are reused across products, studies, therapeutic areas, and deployment environments.
Office working requirements
When we put unexpected teams in the same room, we unleash bold thinking with the power to inspire life-changing medicines. In-person working gives us the platform we need to connect, work at pace, and challenge perceptions. That’s why we work, on average, a minimum of three days per week from the office. We balance this expectation with individual flexibility. Join us in our unique and ambitious world.
#EAI
Date Posted
08-oct-2026Closing Date
19-oct-2026AstraZeneca embraces diversity and equality of opportunity. We are committed to building an inclusive and diverse team representing all backgrounds, with as wide a range of perspectives as possible, and harnessing industry-leading skills. We believe that the more inclusive we are, the better our work will be. We welcome and consider applications to join our team from all qualified candidates, regardless of their characteristics. We comply with all applicable laws and regulations on non-discrimination in employment (and recruitment), as well as work authorization and employment eligibility verification requirements.
How we rate this
Associate Director, Clinical AI Evaluation and Responsible Deployment at AstraZeneca rates 95 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you decide when an AI agent can act on its own versus asking for approval first?
- How do you monitor a model once it's live, and how do you know it needs retraining?
- How do you decide that one model's output is better than another's for a given task?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: AI agents, ML Ops, and AI Evaluation. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new healthcare jobs (Builds AI ●●●●) by email
One email a week with the new healthcare jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Healthcare roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step