The Gimerny Blog
Research breakthroughs, engineering deep dives, and industry perspectives from our team.
Why Equivariant Diffusion Models Are Reshaping Drug Design
Traditional virtual screening searches existing chemical libraries. Generative models create entirely new molecules, but only if they respect the 3D symmetries of molecular interactions.
Building Biomedical Knowledge Graphs for Target Identification at Scale
How we constructed a 2.1 billion edge knowledge graph that connects proteins, genes, diseases, and drugs, and why heterogeneous GNNs are essential for reasoning over it.
Federated Learning in Pharma: Training on Patient Data Without Seeing It
How Gimerny enables multi-institutional model training while keeping sensitive patient data behind institutional firewalls, maintaining HIPAA and GDPR compliance.
A Practical Guide to Transformer-Based Retrosynthesis
How sequence-to-sequence transformers trained on 15M reactions predict synthesis routes, and what we learned about making them useful for real chemists.
Digital Twins in Clinical Trials: Simulating Before Recruiting
How patient-level digital twins trained on 2M+ historical records enable virtual trial simulations that reduce enrollment requirements by 35% and accelerate timelines.
Scaling Molecular Simulations: Our GPU Infrastructure Journey
From 8 GPUs to 2,000: how we built a distributed computing platform for molecular dynamics and generative chemistry on NVIDIA A100 and H100 clusters.
Predicting ADMET Properties with Deep Learning: What Works and What Does Not
ADMET prediction is the unsung hero of drug design. We evaluate 21 endpoints, compare architectures, and share which properties ML can reliably predict today.
Responsible AI in Drug Discovery: Beyond the Hype
AI in drug discovery carries unique ethical responsibilities. We discuss bias in training data, interpretability for regulators, and our framework for responsible deployment.
Foundation Models for Single-Cell Biology: Architecture and Training at Scale
How we trained a 1.2B parameter foundation model on 4.2PB of biological data, and why pre-trained representations are transforming single-cell analysis.