Back to Blog
Industry11 min read2024-07-15

Responsible AI in Drug Discovery: Beyond the Hype

AI in drug discovery carries unique ethical responsibilities. We discuss bias in training data, interpretability for regulators, and our framework for responsible deployment.

SL
Sarah Lindström
Chief Operating Officer

The application of AI to drug discovery has generated enormous excitement, and rightly so. Computational methods are accelerating target identification, molecular design, and clinical trial optimization in ways that were unimaginable a decade ago. But the excitement should be tempered by a sober assessment of the unique responsibilities that come with applying AI to human health.

At Gimerny AI, we believe that responsible AI is not a compliance checkbox but a competitive advantage. Pharmaceutical partners trust platforms that are transparent, well-validated, and honest about limitations. Regulators are more receptive to AI-assisted submissions that come with robust documentation and interpretability. This post outlines our framework for responsible AI in drug discovery.

Bias in Chemical and Biological Data

AI models are only as good as their training data, and chemical and biological datasets carry significant biases. Historical drug discovery has focused disproportionately on certain target classes (kinases represent 25% of all approved drug targets), certain disease areas (oncology and metabolic disease dominate clinical trial databases), and certain patient populations (clinical trials have historically underrepresented women, elderly patients, and non-European genetic backgrounds).

These biases propagate through AI models in subtle ways. A target identification model trained primarily on oncology data will be better at predicting oncology targets than rare disease targets, even if the underlying biology is equally well-characterized. A clinical trial digital twin trained on predominantly Caucasian patient data will make less accurate predictions for patients of other ancestries.

We address data bias through several mechanisms. First, bias auditing: before training any model, we conduct a systematic audit of the training data's coverage across target classes, disease areas, chemical scaffolds, and patient demographics, documenting gaps and their potential impact. Second, stratified evaluation: we evaluate model performance separately across subgroups (disease areas, scaffold types, demographic groups) rather than relying on aggregate metrics that can mask disparities. Third, active learning: we use uncertainty-guided data acquisition to prioritize collecting data in underrepresented regions of chemical and biological space.

Interpretability for Regulatory Submissions

Regulatory agencies (FDA, EMA, PMDA) are increasingly open to AI-assisted drug development, but they require understanding of how AI models arrive at their conclusions. A black-box prediction that a molecule is safe is not acceptable; regulators need to understand the basis for that prediction and assess its reliability.

We build interpretability into our models at three levels. Feature-level explanations identify which molecular features (substructures, physicochemical properties, binding interactions) drive each prediction, using attention weights, integrated gradients, and SHAP values. Mechanistic mapping links model predictions to known biological mechanisms wherever possible, connecting a hepatotoxicity prediction to specific structural alerts associated with reactive metabolite formation. Uncertainty quantification accompanies every prediction with calibrated confidence intervals, allowing regulators to assess reliability and identify predictions that require experimental confirmation.

Validation Standards

The drug discovery AI field suffers from a reproducibility problem. Published methods are often evaluated on different benchmarks, with different data splits, making fair comparison impossible. Worse, many published results use random data splits that leak information between training and test sets, inflating reported performance.

We adhere to strict validation protocols. Temporal splits ensure that test set compounds were published after training set compounds, simulating real-world prospective prediction. Scaffold splits ensure that no Murcko scaffold appears in both training and test sets, testing generalization to novel chemotypes. Multi-dataset evaluation requires that claims of improvement must hold across at least three independent datasets, not just one favorable benchmark. Prospective validation provides the ultimate test, where predictions are made before experimental results are available, and accuracy is assessed after the fact. We maintain a public tracker of our prospective prediction accuracy across all products.

The Human-in-the-Loop Imperative

AI in drug discovery must augment, not replace, human expertise. We design all our products with explicit human decision points. In GimernyDiscover, target recommendations are presented with evidence summaries that a biologist evaluates. In GimernyMolecule, generated molecules are reviewed by medicinal chemists who apply practical knowledge that no model captures. In GimernyTrial, trial design recommendations are discussed with clinical operations teams who understand site-level logistics.

This is not a limitation of AI but a design principle. The consequences of errors in drug discovery are measured in patient harm, and no AI system should make consequential decisions without human oversight.

Our Responsible AI Commitments

We have formalized our approach into five commitments. Transparency in every prediction, accompanied by confidence scores, applicability domain assessments, and explanations. Validation before deployment, requiring every model to pass prospective validation before being used in client projects. Bias monitoring through continuous auditing of model performance across subgroups, with automated alerts when disparities exceed thresholds. Human oversight, ensuring no AI output advances in the drug discovery pipeline without expert review. Continuous improvement through systematic collection of feedback from experimental validation, used to retrain and improve models.

Responsible AI is not about slowing down innovation. It is about building trust that accelerates adoption and, ultimately, gets better medicines to patients faster.

Share this article: