Ссылка
click to show
click to show
OmicsPred catalog surpasses 3.3 million predictive models
On September 1, Nature Genetics, September 1, 2026 published a description of the OmicsPred catalog, which currently holds 3.3 million models that predict RNA, protein, and metabolite levels from DNA data. Each model is accompanied by metadata that lets researchers judge its suitability for a new study.
Models are trained on individuals who have both genotype and molecular measurements; they weigh DNA variant contributions to forecast a specific molecule’s level. The resulting formula can be applied to genotypes from another cohort or to summary statistics from genetic studies. However, a simple list of coefficients is insufficient for transfer.
To address this, OmicsPred stores the DNA variants and their weights, the genome‑build version, tissue type, training and validation sample details, participant ancestry, and prediction accuracy. This lets users compare their own data with the conditions under which a model was trained and tested.
By May 2026 the catalog contained 3,339,469 models, up from 17,227 models in the 2023 version. Most predict gene expression across 49 human tissues, with additional models for proteins and metabolites.
The authors applied the catalog to summary results from the Million Veteran Program, matching predicted RNA and protein levels in blood and plasma against 1,233 health conditions in African, admixed American, and European ancestry groups. This yielded about ~46 million comparisons; after correcting for multiple testing, >190,000 associations remained statistically significant.
For each association the catalog records the exact formula used, the training data, and the prediction accuracy, enabling other researchers to reuse the model and test the hypotheses on new datasets.
🔗 Read original →
5 ·