In January 2026, the U.S. Food and Drug Administration and the European Medicines Agency jointly published the Guiding Principles of Good AI Practice in Drug Development, 10 high-level principles for the responsible use of artificial intelligence across the entire medicines lifecycle. These principles do not stand alone. They sit above an operational framework that the FDA had already advanced a year earlier, and above the analytical foundation the EMA had laid before that. Read together, the three instruments send one consistent message: artificial intelligence can accelerate the development of safe and effective medicines, but only when it operates inside architectures that keep human experts engaged and immersed at every stage. Models that relegate human involvement to a final review-and-sign-off will not satisfy what these regulators have articulated.
1. What the Frameworks Establish
The FDA - In January 2025, the FDA issued its first draft guidance on artificial intelligence for medicines: Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products. Developed across multiple centers and offices, the draft guidance introduces a seven-step, risk-based credibility assessment framework for any AI model whose outputs support a regulatory decision concerning a drug or biological product’s safety, effectiveness, or quality. A sponsor first articulates the specific question of interest, then precisely defines the context of use - the model’s role, scope, and how its outputs combine with other evidence. The sponsor next assesses model risk as a function of the model’s influence and the consequence of an incorrect decision. A credibility assessment plan follows, addressing model architecture, the fitness-for-use of training data, training methodology, and evaluation methods. The sponsor then executes that plan, documents the results in a credibility assessment report, and determines whether the model’s credibility is adequate for the stated context of use. A dedicated lifecycle-maintenance expectation recognizes that data drift and environmental change can degrade performance over time, calling for ongoing monitoring, periodic re-evaluation, and, where necessary, re-execution of credibility activities. Early engagement with the FDA is encouraged throughout.
The EMA - The EMA’s contribution rests on its Reflection paper on the use of Artificial Intelligence in the medicinal product lifecycle, adopted by the Committee for Medicinal Products for Human Use and the Committee for Veterinary Medicinal Products in September 2024. That paper set out the agency’s current thinking on AI from drug discovery and nonclinical development through clinical trials, manufacturing, and post-authorization surveillance, and consistently prioritized two commitments: a risk-based approach and a human-centered approach across all phases. It established the conceptual groundwork on which the joint principles now build.
The convergence - The Guiding Principles of Good AI Practice in Drug Development unify these positions. The 10 principles are non-binding and are intended to underpin forthcoming jurisdiction-specific guidance on both sides of the Atlantic. Taken together, they form 10 interlocking expectations whose common thread is that AI remains subordinate to human judgment and continuous oversight. They call for AI that is human-centric by design; a risk-based approach with validation, mitigation, and oversight proportionate to context of use and determined model risk; adherence to relevant legal, technical, scientific, and regulatory standards, including Good Practices; a clearly defined context of use; multidisciplinary expertise integrated throughout the technology’s lifecycle; rigorous data governance with documented, traceable, and verifiable provenance; sound model and system design that promotes transparency, reliability, and robustness; risk-based performance assessment of the complete system, expressly including human-AI interactions; lifecycle management with scheduled monitoring and periodic re-evaluation to address data drift; and clear, plain-language communication of context of use, performance, limitations, and updates. Across three instruments - the FDA framework, the EMA reflection paper, and the joint principles - the conclusion is identical: AI is a tool, and ultimate authority remains with human experts.
2. Human Engagement Means Immersion, Not a Final Review
A frequent misreading is that these frameworks can be satisfied by building or acquiring an AI system and then having qualified personnel review its outputs at the end. That interpretation falls short of what the regulators have written.
The expectation is genuine, continuous human immersion. Multidisciplinary teams are to participate from the outset - defining the question and the context of use, shaping model design and evaluation criteria, assessing human-AI team performance, interpreting outputs within their full clinical or manufacturing context, monitoring for drift or degradation, and retaining ultimate decision authority. The principle on performance assessment is explicit that evaluation addresses the complete system, including human-AI interactions, not an isolated model. The principle on lifecycle management calls for scheduled, human-led or human-overseen re-evaluation and adaptation. Human-centric design is not an afterthought; it is the first principle, and it runs through every requirement that follows.
Superficial oversight - endorsing AI-generated content without deep engagement - cannot deliver the proportionate risk mitigation and sustained accountability these frameworks describe. An organization that treats human involvement as a terminal checkpoint rather than an embedded, ongoing process will struggle to demonstrate alignment and, more consequentially, will expose itself to the very errors, biases, and performance shifts the regulators have identified as material risks.
3. What Life Sciences Organizations Should Do Now
Begin by recognizing what “AI” actually means in a regulated setting. Public discussion has collapsed the term into large language models, and those systems are doing genuinely impressive work. But generative models are one branch of a far larger field, and in pharmaceutical and medical-device environments, many tasks are served by other methods - computer vision for manufacturing inspection and medical imaging, classical and statistical machine learning for structured data, Bayesian and pharmacometric models for trial design, and purpose-built classifiers and anomaly-detection systems for pharmacovigilance and process monitoring. A Formula 1 car is a marvel of engineering, but no one plows a field with it, and a tractor will never win a Grand Prix - each is exceptional at the job it was built for. The same discipline applies here: when precision is required and patient safety is at stake, the question is not what a model can do in a demonstration, but what is best suited to the task, most controllable, and most readily proven. Some tasks call for a specialist model trained on narrow, well-characterized data; others may be well served by a language model. The regulators built for this distinction - the FDA framework and the joint principles are deliberately model-agnostic, defined by context of use and risk rather than by architecture. The instruction that follows is to match the model to the task and its risk, and to be able to prove the choice, rather than reaching for any one technology because it is familiar or impressive.
Treat AI as regulated infrastructure, not a productivity experiment. Bring each use inside the quality system; classify it by the influence of the model and the consequence of an error; and scale validation and oversight to that risk. Plan from the outset for the full lifecycle, because models drift and conditions change - schedule the monitoring, re-evaluation, and re-validation the principles call for rather than treating validation as a single event.
Build human immersion and multidisciplinary expertise into the process from the start. Engage regulatory, clinical, quality, data-science, and information-security functions from the framing of the question through deployment and monitoring, and assess performance on the human-AI team rather than on the model alone. For most organizations, this means new cross-functional governance, new competencies, and new standard operating procedures - not a single tool purchase.
Finally, govern data, design, and disclosure as rigorously as you govern the product itself. Document the provenance and fitness-for-use of training and input data; require that models be transparent, robust, and bounded to a defined context of use; and be able to explain, in plain language, what each system does, where its limits lie, and how it changes over time. In this environment, the selection of an AI system - and of an AI partner - is a regulatory decision in its own right. The test is no longer only whether a model performs, but whether you can prove how it was built, bounded, governed, and overseen.
The Responsible Path Forward
The FDA and the EMA have drawn a clear line. Artificial intelligence holds real promise for accelerating safe and effective medicines and for strengthening regulatory and manufacturing processes - but only when it is deployed inside architectures that keep human experts genuinely engaged across the lifecycle, supported by robust technical controls for traceability, monitoring, and risk mitigation.
Systems built on the premise of replacing human judgment, or that confine human involvement to a cursory end-of-pipeline sign-off, will struggle to demonstrate the human-AI interaction assessment, multidisciplinary integration, and continuous lifecycle oversight the principles now describe. By contrast, organizations that treat AI as a co-pilot within a human-immersed, risk-proportionate, and fully auditable process align with both the letter and the intent of Good AI Practice. That transparency is the answer to the black box: not an opaque output to be signed off, but a fully attributed process placed in front of the expert who decides.
This is the work we do at Ailethea. We understand AI in its full breadth - which methods fit which regulated tasks when precision and patient safety are at stake, how to keep human experts immersed, and how to make an entire process traceable and defensible to a regulator. We did not build our platform, TRIAD-AI, around a single technology; we built it around these principles, matching the right method to each task and keeping a human accountable for the result. Organizations working through what Good AI Practice requires of their programs do not have to do it alone; if you are weighing what to build, what to buy, and how to prove it, we can help.
This is not merely the path of least regulatory friction. It is the responsible path, the one that lets sponsors, regulators, and patients capture the benefits of AI without amplifying its risks. The future belongs to those who keep human experts at the center, supported, augmented, and fully accountable, rather than those who treat them as final reviewers of an opaque process.
Want to continue reading?
Subscribe to our mailing list for FREE access to the rest of the article.
