New Calibration Framework Improves Confidence in AI Virtual Cell Models
AI virtual cell models may be more capable of predicting cellular responses to genetic perturbations than earlier benchmarks suggested—but only when researchers use evaluation metrics sensitive enough to detect biologically meaningful signals, according to a new study from Shift Bioscience and collaborators.
Published in Nature Biotechnology, the paper, “Deep Learning Perturbation Models Can Outperform Baselines on Calibrated Metrics,” introduces a positive-control baseline and a framework for testing whether commonly used metrics can distinguish meaningful model predictions from less informative ones.
Genetic perturbation models, a subset of AI virtual cells, predict how activating or inhibiting genes will alter a cell’s transcriptome. If reliable, these models could support scalable in silico screens for therapeutic targets. However, recent studies found that sophisticated deep learning systems often failed to outperform simple baselines, raising doubts about whether the models captured perturbation-specific biology.
The researchers argued that part of the problem lies in the benchmarks themselves. “The promise of in silico perturbation models, or ‘virtual cells,’ is that the same scaling laws will apply, and that providing them with ever-increasing quantities of biological data and compute will improve their prediction of real-world experimental results,” Brendan Swain, PhD, CSO and founder of Shift Bioscience, told GEN. “However, early benchmarks of virtual cell models revealed apparently poor performance, casting doubt on their potential.”
To assess this issue, the team developed an “interpolated duplicate” positive control that combines the average perturbation profile with an independent technical duplicate, weighting each gene by the strength of evidence that it was affected. They then introduced the dynamic range fraction, or DRF, which measures how effectively a metric separates this positive control from a negative control.
Across 14 genetic perturbation datasets and 18 evaluation metrics, commonly used measures such as mean squared error and control-referenced Pearson correlation were often poorly calibrated, particularly in datasets with weaker perturbations. The authors also added that, “In contrast, weighted and rank-based metrics—including WMSE, weighted RΔ2, and normalized inverse rank (NIR)—consistently exhibited higher calibration across datasets, reflecting their shared design principle of emphasizing perturbation-specific signal.”
Using the better-calibrated measures, the researchers benchmarked nine deep learning models on tasks involving unseen single-gene perturbations and unseen gene combinations. Earlier models, including scGPT and GEARS, often surpassed uninformative baselines once their predictions were evaluated with calibrated metrics. More recent systems, including PRESAGE and scLambda, showed stronger performance, although results varied by dataset and metric.
The analysis also highlighted the importance of dataset design. In Norman19, where the training data covered only 0.63% of possible gene combinations, a simple additive baseline remained difficult to beat. In Wessels23, which covered 10.3% of possible combinations, multiple models surpassed the additive baseline, indicating that deep learning models can outperform additivity when trained on broader portions of the combinatorial landscape.
“Our findings show that by using well-calibrated metrics and the right dataset, virtual cell models can generate biologically meaningful insights,” said Swain. “As a result, we can use them with greater confidence to identify promising new targets that are relevant to aging and disease. We are applying this framework directly in our target identification program, focusing on targets whose inhibition can support both rejuvenation and treatment of age-related disease, giving us a clearly defined route towards clinical development.”
Shift Bioscience, a Cambridge, U.K.–based biotechnology company, plans to apply the framework in large-scale in vitro and in silico screens for inhibition targets relevant to cellular rejuvenation and age-related disease, initially focusing on fibrosis. “Ultimately, we hope these tools will guide researchers as they develop better and better virtual cell models,” Swain told GEN.
The post New Calibration Framework Improves Confidence in AI Virtual Cell Models appeared first on GEN - Genetic Engineering and Biotechnology News.
Apa Reaksi Anda?
Suka
0
Kurang Suka
0
Setuju
0
Tidak Setuju
0
Bagus
0
Berguna
0
Hebat
0
