Neural Networks Are Now in Drug Registration Dossiers. How to Evaluate AI-Generated Evidence
Imagine this scenario: an expert at the National Centre for Expert Evaluation of Medicinal Products (NCESMP) opens a registration dossier for a new oncological drug and, in the pharmacokinetics section, encounters the phrase: «The dose selection is confirmed by a machine learning model trained on data from 12,000 patients.» What should be done with this? How can one verify whether the model is reliable? And the primary question: is this sufficient as an evidentiary basis for registration?
Five years ago, such wording in a dossier would have been an exception. Today it is no longer a rarity. Pharmaceutical companies are implementing neural network technologies across all stages of drug development, and regulatory authorities worldwide face a difficult choice: either develop clear rules for evaluating such data or risk having critical registration decisions made without an adequate understanding of their evidentiary value.
This article is primarily addressed to experts at regulatory bodies and regulatory affairs specialists in pharmaceutical companies. We will examine what has changed in international practice, what approach is currently taking shape, and specifically what needs to be done when algorithmic outputs appear in a dossier.
How It Used to Work
Classical expert evaluation of a registration dossier was built on the verification of reproducible, deterministic analytics. The biostatistical methods used by applicants were standardized. An expert could open the statistical section and verify the calculations using established formulas. If the data corresponded to the protocol and Type I and Type II errors were controlled, the evidentiary base could be assessed with sufficient certainty.
In the Chemistry, Manufacturing, and Controls (CMC) section, the verification of analytical control methods also relied on well-understood procedures: linearity, accuracy, reproducibility, and specificity. All of these are described in pharmacopoeial guidelines — Russian, European, and American.
This system was created for a specific type of evidence: deterministic algorithms, fixed methods, and static models. Neural networks do not fit into this framework. They are not deterministic in the traditional sense; they learn from data, and their performance depends on the quality of that data, the model architecture, and the training conditions. Moreover, the final result is sometimes impossible to explain through a logically traceable chain.
What Has Changed
Scale of Implementation
According to estimates from analytical firms, by the end of 2025, artificial intelligence (AI) technologies were applied in the development or repurposing of more than 3,000 medicinal products worldwide. Neural network models are used to predict toxicity, search for candidate molecules, optimize the design of clinical trials, and monitor manufacturing processes in real time.
Today, these technologies appear in every section of the dossier. The non-clinical block is supplemented by QSAR (Quantitative Structure-Activity Relationship) models that predict a compound’s toxicity from its structure. Clinical sections now feature algorithms that selected patients for trials or constructed synthetic control groups from historical data. In CMC, neural networks increasingly power continuous manufacturing monitoring systems: the algorithm analyzes sensor readings in real time and flags deviations.
Regulatory Response
In January 2026, the EMA (European Medicines Agency) and the FDA (U.S. Food and Drug Administration) published a joint document, «Principles for Good AI Practice.» This is the first genuinely harmonized international benchmark for evaluating AI technologies in the medicinal product lifecycle.
The document sets out ten principles. The underlying logic runs as follows: a human controls the system’s outputs and bears responsibility for decisions; the developer must fully document the data and training stages; and the scope of application for each specific model must be clearly defined. The position of the regulators remains consistent: AI serves as a support tool, while the final decision always rests with a qualified specialist.
In 2025, the FDA released a draft guidance, «Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making.» The central methodological instrument within it is a seven-step credibility assessment process (Credibility Assessment Framework), adapted from the ASME V&V40 engineering standard.
Table 1. Ten Principles of Good AI Practice (EMA/FDA, 2026)
| No. | Principle | Meaning for Expert Evaluation |
|---|---|---|
| 1 | Model Quality | Adherence to ethical and scientific standards during development |
| 2 | Human-Centricity | Human oversight of system outputs |
| 3 | Compliance | Adherence to data protection and cybersecurity requirements |
| 4 | Traceability | Full documentation of data sources and training stages |
| 5 | Role Definition | Clear boundaries for the application of a specific model |
| 6 | Result Accessibility | Model outputs must be interpretable by the expert |
| 7 | Representativeness | Training data must reflect the target patient population |
| 8 | Risk Management | Assessment of safety consequences in case of model error |
| 9 | Lifecycle Monitoring | Continuous tracking of model performance post-registration |
| 10 | International Collaboration | Exchange of experience between regulators |
Seven-Step Credibility Assessment
The FDA framework proposes a structured approach to one central question: how much can this model be trusted in a specific context? The key concept is the «Context of Use» (COU). The same neural network may be entirely acceptable for primary molecular screening yet insufficiently validated to justify a dose in a registration dossier.
The procedure begins with defining the scientific question and the context of use, followed by a regulatory risk assessment, development and execution of a verification plan, and documentation of deviations. The final step is the expert’s conclusion on the model’s adequacy for the stated purposes. This step cannot be delegated to an algorithm.
Situation in Russia and the EAEU
In the Russian Federation, this work is set out in the Strategic Direction for the Digital Transformation of Healthcare until 2030 (Order of the Government of the Russian Federation No. 959-r dated April 17, 2024). Among the tasks of this Strategic Direction is the implementation of neurotechnologies and AI: automating routine processes and supporting physicians and administrators in decision-making. Plans are also underway to simplify registration procedures for biosimilars, where demonstrating the comparability of complex biological molecules is practically impossible without advanced analytical tools.
The EAEU regulatory framework for medicinal product registration (Decision of the Council of the EEC No. 78 dated November 03, 2016, «On the Rules of Registration and Examination of Medicinal Products for Human Use,» hereinafter — Decision No. 78) contains no specific provisions on AI. This gap is currently being filled through the adaptation of international recommendations at the level of methodological clarifications and expert practice.
Three Major Challenges
The «Black Box» Problem
The main methodological difficulty when auditing AI data is that a neural network often does not allow one to trace the logic from input to output. Traditional evaluation of biostatistical data assumes reproducibility: an expert can recalculate. With a deep neural network, this is simply not possible.
Regulatory bodies propose several approaches. One is to require the use of Explainable AI (XAI) methods, which show exactly which data features influenced the result. Another path is to insist on model validation using external, independent data. If results are reproduced on material the algorithm did not see during training, this is a strong argument for reliability. Finally, one can examine the model’s performance in clinical practice directly: how closely do the model’s predictions match real-world observations?
Data Bias
A neural network learns from data. If the training sample is not representative — for example, if it predominantly includes male patients of European descent in a particular age range — the model will perform worse for all other groups. In clinical practice, this may lead to an underestimation of risk for specific patient populations. Principle 7 of the joint EMA/FDA document explicitly requires verification that training data are representative of the target population for a given medicinal product.
Change Management in Manufacturing Systems
A separate challenge arises in the CMC section when AI is embedded in continuous manufacturing monitoring systems. Neural networks are capable of continuing to learn during operation. If an algorithm approved at the time of registration is subsequently modified through additional training, a question emerges: must the regulator be notified? In international practice, the applicant pre-agrees with the regulator on a Pre-Defined Change Control Protocol (PCCP) to cover such scenarios. In Russian regulatory practice, this instrument is still being developed.
Recommended Actions
For Specialists at Expert Organizations
Define the Context of Use. When evaluating a dossier, first establish at which stage of development the AI model was applied and how its outputs influenced regulatory decisions. If the neural network was used only for primary screening and the final selection was confirmed by classical methods, validation requirements may be less stringent. If the model justifies a dosage, the bar is substantially higher.
Request training documentation. The applicant must describe the data sources, model architecture, quality metrics on both training and test samples, and validation procedures. The absence of this information is itself a signal of insufficient preparation.
Verify data representativeness. Does the demographic composition of the training sample match the target patient population for this drug? This is a foundational question explicitly required by the EMA/FDA principles.
Assess external validation. Results reproduced on an independent dataset carry substantially greater evidentiary weight than performance metrics obtained on the training sample alone.
Request a change management plan if AI is used in manufacturing. It is necessary to understand under what conditions the model may change and what the mechanism is for notifying the regulator.
For Regulatory Affairs Specialists at Pharmaceutical Companies
Document everything from the outset. Record data sources, model versions, training parameters, and validation results at every stage of development. Recovering this information retrospectively, just before filing a dossier, will be considerably harder.
Apply GxP principles to AI systems. GxP is the collective term for the family of Good Practice standards — GMP, GDP, GCP, and others. GMP requirements and data integrity principles extend to algorithmic systems. Records of AI model validation must be accessible to inspectors.
Prepare for interpretability questions. An expert who receives a neural network’s output may ask why the model produced a given result. If there is no answer — whether through Explainable AI or any other means — that is a vulnerability in the dossier.
Study the international guidelines. The 2026 EMA/FDA document and the 2025 FDA draft guidance on AI are required reading for any company working with these technologies.
Consider scientific consultations with the regulator before dossier submission. NCESMP, like the EMA, offers consultations on methodological questions prior to filing. This allows the approach to evaluating AI data to be aligned in advance, avoiding costly rework later.
Neural network technologies are already entering registration dossiers today. Some companies are submitting applications with AI-supported justifications right now, and regulatory authorities are evaluating them regardless of whether dedicated guidance exists.
The methodological foundation has been laid: the EMA/FDA principles and the FDA’s seven-step framework provide a workable basis. The task for Russian regulatory science is to adapt these approaches to EAEU conditions and close the gap in Decision No. 78. For the expert community, this is a time to build competence with new tools; regulatory affairs specialists, in turn, should be constructing their evidentiary base to meet requirements that will become mandatory tomorrow.
Regulatory Framework:
1. Decision of the Council of the EEC No. 78 dated November 03, 2016 «On the Rules of Registration and Examination of Medicinal Products for Human Use»
2. Strategic Direction for the Digital Transformation of Healthcare (Order of the Government of the Russian Federation No. 959-r dated April 17, 2024)
3. EMA/FDA Joint Principles for Good AI Practice (2026)
4. FDA Draft Guidance: Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products (2025)
5. EMA Reflection Paper on the use of Artificial Intelligence in the lifecycle of medicines (2023)
6. ASME V&V 40: Assessing Credibility of Computational Modeling through Verification and Validation