Public knowledge

Put omics knowledge on the research path

Browse GeniOmics platform knowledge and curated topic books about data, workflows, results, and reusable methods.

These pages use published public knowledge documents and curated books. Read them here, then open GeniOmics when you are ready to run an analysis.

Platform knowledge

Start with a concrete question

Use curated documents to understand omics data, analysis methods, and result evidence.

Platform article100/100

Meta-analysis of tumor- and T cell-intrinsic mechanisms of sensitization to checkpoint inhibition.

Checkpoint inhibitors (CPIs) augment adaptive immunity. Systematic pan-tumor analyses may reveal the relative importance of tumor-cell-intrinsic and microenvironmental features underpinning CPI sensitization. Here, we collated whole-exome and transcriptomic data for >1,000 CPI-treated patients across seven tumor types, utilizing standardized bioinformatics workflows and clinical outcome criteria to validate multivariable predictors of CPI sensitization. Clonal tumor mutation burden (TMB) was the strongest predictor of CPI response, followed by total TMB and CXCL9 expression. Subclonal TMB, somatic copy alteration burden, and histocompatibility leukocyte antigen (HLA) evolutionary divergence failed to attain pan-cancer significance. Dinucleotide variants were identified as a source of immunogenic epitopes associated with radical amino acid substitutions and enhanced peptide hydrophobicity/immunogenicity. Copy-number analysis revealed two additional determinants of CPI outcome supported

Aug 3, 2026Read article
Platform article100/100

Meta-analysis of the Parkinson's disease gut microbiome suggests alterations linked to intestinal inflammation.

The gut microbiota is emerging as an important modulator of neurodegenerative diseases, and accumulating evidence has linked gut microbes to Parkinson's disease (PD) symptomatology and pathophysiology. PD is often preceded by gastrointestinal symptoms and alterations of the enteric nervous system accompany the disease. Several studies have analyzed the gut microbiome in PD, but a consensus on the features of the PD-specific microbiota is missing. Here, we conduct a meta-analysis re-analyzing the ten currently available 16S microbiome datasets to investigate whether common alterations in the gut microbiota of PD patients exist across cohorts. We found significant alterations in the PD-associated microbiome, which are robust to study-specific technical heterogeneities, although differences in microbiome structure between PD and controls are small. Enrichment of the genera Lactobacillus, Akkermansia, and Bifidobacterium and depletion of bacteria belonging to the Lachnospiraceae family and

Aug 3, 2026Read article
Platform article100/100

Interpretation of T cell states from single-cell transcriptomics data using reference atlases.

Single-cell RNA sequencing (scRNA-seq) has revealed an unprecedented degree of immune cell diversity. However, consistent definition of cell subtypes and cell states across studies and diseases remains a major challenge. Here we generate reference T cell atlases for cancer and viral infection by multi-study integration, and develop ProjecTILs, an algorithm for reference atlas projection. In contrast to other methods, ProjecTILs allows not only accurate embedding of new scRNA-seq data into a reference without altering its structure, but also characterizing previously unknown cell states that "deviate" from the reference. ProjecTILs accurately predicts the effects of cell perturbations and identifies gene programs that are altered in different conditions and tissues. A meta-analysis of tumor-infiltrating T cells from several cohorts reveals a strong conservation of T cell subtypes between human and mouse, providing a consistent basis to describe T cell heterogeneity across studies, disea

Aug 3, 2026Read article
Platform article100/100

Long-term dietary patterns are associated with pro-inflammatory and anti-inflammatory features of the gut microbiome.

Objective The microbiome directly affects the balance of pro-inflammatory and anti-inflammatory responses in the gut. As microbes thrive on dietary substrates, the question arises whether we can nourish an anti-inflammatory gut ecosystem. We aim to unravel interactions between diet, gut microbiota and their functional ability to induce intestinal inflammation. Design We investigated the relation between 173 dietary factors and the microbiome of 1425 individuals spanning four cohorts: Crohn's disease, ulcerative colitis, irritable bowel syndrome and the general population. Shotgun metagenomic sequencing was performed to profile gut microbial composition and function. Dietary intake was assessed through food frequency questionnaires. We performed unsupervised clustering to identify dietary patterns and microbial clusters. Associations between diet and microbial features were explored per cohort, followed by a meta-analysis and heterogeneity estimation. Results We identified 38 associatio

Aug 3, 2026Read article
Platform article100/100

Single-cell and bulk transcriptome sequencing identifies two epithelial tumor cell states and refines the consensus molecular classification of colorectal cancer.

The consensus molecular subtype (CMS) classification of colorectal cancer is based on bulk transcriptomics. The underlying epithelial cell diversity remains unclear. We analyzed 373,058 single-cell transcriptomes from 63 patients, focusing on 49,155 epithelial cells. We identified a pervasive genetic and transcriptomic dichotomy of malignant cells, based on distinct gene expression, DNA copy number and gene regulatory network. We recapitulated these subtypes in bulk transcriptomes from 3,614 patients. The two intrinsic subtypes, iCMS2 and iCMS3, refine CMS. iCMS3 comprises microsatellite unstable (MSI-H) cancers and one-third of microsatellite-stable (MSS) tumors. iCMS3 MSS cancers are transcriptomically more similar to MSI-H cancers than to other MSS cancers. CMS4 cancers had either iCMS2 or iCMS3 epithelium; the latter had the worst prognosis. We defined the intrinsic epithelial axis of colorectal cancer and propose a refined 'IMF' classification with five subtypes, combining intrins

Aug 3, 2026Read article
Platform article100/100

A comprehensive benchmarking with practical guidelines for cellular deconvolution of spatial transcriptomics.

Spatial transcriptomics technologies are used to profile transcriptomes while preserving spatial information, which enables high-resolution characterization of transcriptional patterns and reconstruction of tissue architecture. Due to the existence of low-resolution spots in recent spatial transcriptomics technologies, uncovering cellular heterogeneity is crucial for disentangling the spatial patterns of cell types, and many related methods have been proposed. Here, we benchmark 18 existing methods resolving a cellular deconvolution task with 50 real-world and simulated datasets by evaluating the accuracy, robustness, and usability of the methods. We compare these methods comprehensively using different metrics, resolutions, spatial transcriptomics technologies, spot numbers, and gene numbers. In terms of performance, CARD, Cell2location, and Tangram are the best methods for conducting the cellular deconvolution task. To refine our comparative results, we provide decision-tree-style gu

Aug 3, 2026Read article
Platform article100/100

Oxidative stress gene expression, DNA methylation, and gut microbiota interaction trigger Crohn's disease: a multi-omics Mendelian randomization study.

Background Oxidative stress (OS) is a key pathophysiological mechanism in Crohn's disease (CD). OS-related genes can be affected by environmental factors, intestinal inflammation, gut microbiota, and epigenetic changes. However, the role of OS as a potential CD etiological factor or triggering factor is unknown, as differentially expressed OS genes in CD can be either a cause or a subsequent change of intestinal inflammation. Herein, we used a multi-omics summary data-based Mendelian randomization (SMR) approach to identify putative causal effects and underlying mechanisms of OS genes in CD. Methods OS-related genes were extracted from the GeneCards database. Intestinal transcriptome datasets were collected from the Gene Expression Omnibus (GEO) database and meta-analyzed to identify differentially expressed genes (DEGs) related to OS in CD. Integration analyses of the largest CD genome-wide association study (GWAS) summaries with expression quantitative trait loci (eQTLs) and DNA meth

Aug 3, 2026Read article
Platform article100/100

Profiling the heterogeneity of colorectal cancer consensus molecular subtypes using spatial transcriptomics.

The consensus molecular subtypes (CMS) of colorectal cancer (CRC) is the most widely-used gene expression-based classification and has contributed to a better understanding of disease heterogeneity and prognosis. Nevertheless, CMS intratumoral heterogeneity restricts its clinical application, stressing the necessity of further characterizing the composition and architecture of CRC. Here, we used Spatial Transcriptomics (ST) in combination with single-cell RNA sequencing (scRNA-seq) to decipher the spatially resolved cellular and molecular composition of CRC. In addition to mapping the intratumoral heterogeneity of CMS and their microenvironment, we identified cell communication events in the tumor-stroma interface of CMS2 carcinomas. This includes tumor growth-inhibiting as well as -activating signals, such as the potential regulation of the ETV4 transcriptional activity by DCN or the PLAU-PLAUR ligand-receptor interaction. Our study illustrates the potential of ST to resolve CRC molec

Aug 3, 2026Read article
Platform article98/100

The PRIDE database at 20 years: 2025 update.

The PRoteomics IDEntifications (PRIDE) database (https://www.ebi.ac.uk/pride/) is the world's leading mass spectrometry (MS)-based proteomics data repository and one of the founding members of the ProteomeXchange consortium. This manuscript summarizes the developments in PRIDE resources and related tools for the last three years. The number of submitted datasets to PRIDE Archive (the archival component of PRIDE) has reached on average around 534 datasets per month. This has been possible thanks to continuous improvements in infrastructure such as a new file transfer protocol for very large datasets (Globus), a new data resubmission pipeline and an automatic dataset validation process. Additionally, we will highlight novel activities such as the availability of the PRIDE chatbot (based on the use of open-source Large Language Models), and our work to improve support for MS crosslinking datasets. Furthermore, we will describe how we have increased our efforts to reuse, reanalyze and diss

Aug 3, 2026Read article
Platform article98/100

Definitions and guidelines for research on antibiotic persistence.

Increasing concerns about the rising rates of antibiotic therapy failure and advances in single-cell analyses have inspired a surge of research into antibiotic persistence. Bacterial persister cells represent a subpopulation of cells that can survive intensive antibiotic treatment without being resistant. Several approaches have emerged to define and measure persistence, and it is now time to agree on the basic definition of persistence and its relation to the other mechanisms by which bacteria survive exposure to bactericidal antibiotic treatments, such as antibiotic resistance, heteroresistance or tolerance. In this Consensus Statement, we provide definitions of persistence phenomena, distinguish between triggered and spontaneous persistence and provide a guide to measuring persistence. Antibiotic persistence is not only an interesting example of non-genetic single-cell heterogeneity, it may also have a role in the failure of antibiotic treatments. Therefore, it is our hope that the

Aug 3, 2026Read article
Platform article98/100

Meta-analysis of gut microbiome studies identifies disease-specific and shared responses.

Hundreds of clinical studies have demonstrated associations between the human microbiome and disease, yet fundamental questions remain on how we can generalize this knowledge. Results from individual studies can be inconsistent, and comparing published data is further complicated by a lack of standard processing and analysis methods. Here we introduce the MicrobiomeHD database, which includes 28 published case-control gut microbiome studies spanning ten diseases. We perform a cross-disease meta-analysis of these studies using standardized methods. We find consistent patterns characterizing disease-associated microbiome changes. Some diseases are associated with over 50 genera, while most show only 10-15 genus-level changes. Some diseases are marked by the presence of potentially pathogenic microbes, whereas others are characterized by a depletion of health-associated bacteria. Furthermore, we show that about half of genera associated with individual studies are bacteria that respond to

Aug 3, 2026Read article
Platform article98/100

Guidelines and considerations for the use of system suitability and quality control samples in mass spectrometry assays applied in untargeted clinical metabolomic studies.

Background Quality assurance (QA) and quality control (QC) are two quality management processes that are integral to the success of metabolomics including their application for the acquisition of high quality data in any high-throughput analytical chemistry laboratory. QA defines all the planned and systematic activities implemented before samples are collected, to provide confidence that a subsequent analytical process will fulfil predetermined requirements for quality. QC can be defined as the operational techniques and activities used to measure and report these quality requirements after data acquisition. Aim of review This tutorial review will guide the reader through the use of system suitability and QC samples, why these samples should be applied and how the quality of data can be reported. Key scientific concepts of review System suitability samples are applied to assess the operation and lack of contamination of the analytical platform prior to sample analysis. Isotopically-la

Aug 3, 2026Read article
Platform article98/100

Intestinal dysbiosis in preterm infants preceding necrotizing enterocolitis: a systematic review and meta-analysis.

Background Necrotizing enterocolitis (NEC) is a catastrophic disease of preterm infants, and microbial dysbiosis has been implicated in its pathogenesis. Studies evaluating the microbiome in NEC and preterm infants lack power and have reported inconsistent results. Methods and results Our objectives were to perform a systematic review and meta-analyses of stool microbiome profiles in preterm infants to discern and describe microbial dysbiosis prior to the onset of NEC and to explore heterogeneity among studies. We searched MEDLINE, PubMed, CINAHL, and conference abstracts from the proceedings of Pediatric Academic Societies and reference lists of relevant identified articles in April 2016. Studies comparing the intestinal microbiome in preterm infants who developed NEC to those of controls, using culture-independent molecular techniques and reported α and β-diversity metrics, and microbial profiles were included. In addition, 16S ribosomal ribonucleic acid (rRNA) sequence data with cli

Aug 3, 2026Read article
Platform article98/100

Looking for a Signal in the Noise: Revisiting Obesity and the Microbiome.

Unlabelled Two recent studies have reanalyzed previously published data and found that when data sets were analyzed independently, there was limited support for the widely accepted hypothesis that changes in the microbiome are associated with obesity. This hypothesis was reconsidered by increasing the number of data sets and pooling the results across the individual data sets. The preferred reporting items for systematic reviews and meta-analyses guidelines were used to identify 10 studies for an updated and more synthetic analysis. Alpha diversity metrics and the relative risk of obesity based on those metrics were used to identify a limited number of significant associations with obesity; however, when the results of the studies were pooled by using a random-effect model, significant associations were observed among Shannon diversity, the number of observed operational taxonomic units, Shannon evenness, and obesity status. They were not observed for the ratio of Bacteroidetes and Fir

Aug 3, 2026Read article
Platform article98/100

Machine Learning Meta-analysis of Large Metagenomic Datasets: Tools and Biological Insights.

Shotgun metagenomic analysis of the human associated microbiome provides a rich set of microbial features for prediction and biomarker discovery in the context of human diseases and health conditions. However, the use of such high-resolution microbial features presents new challenges, and validated computational tools for learning tasks are lacking. Moreover, classification rules have scarcely been validated in independent studies, posing questions about the generality and generalization of disease-predictive models across cohorts. In this paper, we comprehensively assess approaches to metagenomics-based prediction tasks and for quantitative assessment of the strength of potential microbiome-phenotype associations. We develop a computational framework for prediction tasks using quantitative microbiome profiles, including species-level relative abundances and presence of strain-specific markers. A comprehensive meta-analysis, with particular emphasis on generalization across cohorts, wa

Aug 3, 2026Read article
Platform article98/100

Genetic Predisposition to an Impaired Metabolism of the Branched-Chain Amino Acids and Risk of Type 2 Diabetes: A Mendelian Randomisation Analysis.

Background Higher circulating levels of the branched-chain amino acids (BCAAs; i.e., isoleucine, leucine, and valine) are strongly associated with higher type 2 diabetes risk, but it is not known whether this association is causal. We undertook large-scale human genetic analyses to address this question. Methods and findings Genome-wide studies of BCAA levels in 16,596 individuals revealed five genomic regions associated at genome-wide levels of significance (p < 5 × 10-8). The strongest signal was 21 kb upstream of the PPM1K gene (beta in standard deviations [SDs] of leucine per allele = 0.08, p = 3.9 × 10-25), encoding an activator of the mitochondrial branched-chain alpha-ketoacid dehydrogenase (BCKD) responsible for the rate-limiting step in BCAA catabolism. In another analysis, in up to 47,877 cases of type 2 diabetes and 267,694 controls, a genetically predicted difference of 1 SD in amino acid level was associated with an odds ratio for type 2 diabetes of 1.44 (95% CI 1.26-1.65,

Aug 3, 2026Read article
Platform article98/100

Estimating the total number of phosphoproteins and phosphorylation sites in eukaryotic proteomes.

Background Phosphorylation is the most frequent post-translational modification made to proteins and may regulate protein activity as either a molecular digital switch or a rheostat. Despite the cornucopia of high-throughput (HTP) phosphoproteomic data in the last decade, it remains unclear how many proteins are phosphorylated and how many phosphorylation sites (p-sites) can exist in total within a eukaryotic proteome. We present the first reliable estimates of the total number of phosphoproteins and p-sites for four eukaryotes (human, mouse, Arabidopsis, and yeast). Results In all, 187 HTP phosphoproteomic datasets were filtered, compiled, and studied along with two low-throughput (LTP) compendia. Estimates of the number of phosphoproteins and p-sites were inferred by two methods: Capture-Recapture, and fitting the saturation curve of cumulative redundant vs. cumulative non-redundant phosphoproteins/p-sites. Estimates were also adjusted for different levels of noise within the individ

Aug 3, 2026Read article
Platform article98/100

Identifying gene targets for brain-related traits using transcriptomic and methylomic data from blood.

Understanding the difference in genetic regulation of gene expression between brain and blood is important for discovering genes for brain-related traits and disorders. Here, we estimate the correlation of genetic effects at the top-associated cis-expression or -DNA methylation (DNAm) quantitative trait loci (cis-eQTLs or cis-mQTLs) between brain and blood (r b ). Using publicly available data, we find that genetic effects at the top cis-eQTLs or mQTLs are highly correlated between independent brain and blood samples ([Formula: see text] for cis-eQTLs and [Formula: see text] for cis-mQTLs). Using meta-analyzed brain cis-eQTL/mQTL data (n = 526 to 1194), we identify 61 genes and 167 DNAm sites associated with four brain-related phenotypes, most of which are a subset of the discoveries (97 genes and 295 DNAm sites) using data from blood with larger sample sizes (n = 1980 to 14,115). Our results demonstrate the gain of power in gene discovery for brain-related phenotypes using blood cis-e

Aug 3, 2026Read article
Platform article98/100

DNA methylation signatures of chronic low-grade inflammation are associated with complex diseases.

Background Chronic low-grade inflammation reflects a subclinical immune response implicated in the pathogenesis of complex diseases. Identifying genetic loci where DNA methylation is associated with chronic low-grade inflammation may reveal novel pathways or therapeutic targets for inflammation. Results We performed a meta-analysis of epigenome-wide association studies (EWAS) of serum C-reactive protein (CRP), which is a sensitive marker of low-grade inflammation, in a large European population (n = 8863) and trans-ethnic replication in African Americans (n = 4111). We found differential methylation at 218 CpG sites to be associated with CRP (P < 1.15 × 10 -7 ) in the discovery panel of European ancestry and replicated (P < 2.29 × 10 -4 ) 58 CpG sites (45 unique loci) among African Americans. To further characterize the molecular and clinical relevance of the findings, we examined the association with gene expression, genetic sequence variants, and clinical outcomes. DNA methylation at

Aug 3, 2026Read article
Platform article98/100

Database Resources of the National Genomics Data Center, China National Center for Bioinformation in 2025.

The National Genomics Data Center (NGDC), which is a part of the China National Center for Bioinformation (CNCB), offers a comprehensive suite of database resources to support the global scientific community. Amidst the unprecedented accumulation of multi-omics data, CNCB-NGDC is committed to continually evolving and updating its core database resources through big data archiving, integrative analysis and value-added curation. Over the past year, CNCB-NGDC has expanded its collaborations with international databases and established new subcenters focusing on biodiversity, traditional Chinese medicine and tumor genetics. Substantial efforts have been made toward encompassing a broad spectrum of multi-omics data, developing innovative resources and enhancing existing resources. Notably, new resources have been developed for single-cell omics (scTWAS Atlas), genome and variation (VDGE), health and disease (CVD Atlas, CPMKG, Immunosenescence Inventory, HemAtlas, Cyclicpepedia, IDeAS), biod

Aug 3, 2026Read article
Platform article98/100

Neoadjuvant therapy with immune checkpoint blockade, antiangiogenesis, and chemotherapy for locally advanced gastric cancer.

Despite neoadjuvant/conversion chemotherapy, the prognosis of cT4a/bN+ gastric cancer is poor. Immune checkpoint inhibitors (ICIs) and antiangiogenic agents have shown activity in late-stage gastric cancer, but their efficacy in the neoadjuvant/conversion setting is unclear. In this single-armed, phase II, exploratory trial (NCT03878472), we evaluate the efficacy of a combination of ICI (camrelizumab), antiangiogenesis (apatinib), and chemotherapy (S-1 ± oxaliplatin) for neoadjuvant/conversion treatment of cT4a/bN+ gastric cancer. The primary endpoints are pathological responses and their potential biomarkers. Secondary endpoints include safety, objective response, progression-free survival, and overall survival. Complete and major pathological response rates are 15.8% and 26.3%. Pathological responses correlate significantly with microsatellite instability status, PD-L1 expression, and tumor mutational burden. In addition, multi-omics examination reveals several putative biomarkers fo

Aug 3, 2026Read article
Platform article98/100

Maternal plasma folate impacts differential DNA methylation in an epigenome-wide meta-analysis of newborns.

Folate is vital for fetal development. Periconceptional folic acid supplementation and food fortification are recommended to prevent neural tube defects. Mechanisms whereby periconceptional folate influences normal development and disease are poorly understood: epigenetics may be involved. We examine the association between maternal plasma folate during pregnancy and epigenome-wide DNA methylation using Illumina's HumanMethyl450 Beadchip in 1,988 newborns from two European cohorts. Here we report the combined covariate-adjusted results using meta-analysis and employ pathway and gene expression analyses. Four-hundred forty-three CpGs (320 genes) are significantly associated with maternal plasma folate levels during pregnancy (false discovery rate 5%); 48 are significant after Bonferroni correction. Most genes are not known for folate biology, including APC2, GRM8, SLC16A12, OPCML, PRPH, LHX1, KLK4 and PRSS21. Some relate to birth defects other than neural tube defects, neurological func

Aug 3, 2026Read article
Platform article98/100

Evaluation of in silico algorithms for use with ACMG/AMP clinical variant interpretation guidelines.

Background The American College of Medical Genetics and American College of Pathologists (ACMG/AMP) variant classification guidelines for clinical reporting are widely used in diagnostic laboratories for variant interpretation. The ACMG/AMP guidelines recommend complete concordance of predictions among all in silico algorithms used without specifying the number or types of algorithms. The subjective nature of this recommendation contributes to discordance of variant classification among clinical laboratories and prevents definitive classification of variants. Results Using 14,819 benign or pathogenic missense variants from the ClinVar database, we compared performance of 25 algorithms across datasets differing in distinct biological and technical variables. There was wide variability in concordance among different combinations of algorithms with particularly low concordance for benign variants. We also identify a previously unreported source of error in variant interpretation (false co

Aug 3, 2026Read article
Platform article98/100

Tumor microenvironment evaluation promotes precise checkpoint immunotherapy of advanced gastric cancer.

Background Durable efficacy of immune checkpoint blockade (ICB) occurred in a small number of patients with metastatic gastric cancer (mGC) and the determinant biomarker of response to ICB remains unclear. Methods We developed an open-source TMEscore R package, to quantify the tumor microenvironment (TME) to aid in addressing this dilemma. Two advanced gastric cancer cohorts (RNAseq, N=45 and NanoString, N=48) and other advanced cancer (N=534) treated with ICB were leveraged to investigate the predictive value of TMEscore. Simultaneously, multi-omics data from The Cancer Genome Atlas of Stomach Adenocarcinoma (TCGA-STAD) and Asian Cancer Research Group (ACRG) were interrogated for underlying mechanisms. Results The predictive capacity of TMEscore was corroborated in patient with mGC cohorts treated with pembrolizumab in a prospective phase 2 clinical trial (NCT02589496, N=45, area under the curve (AUC)=0.891). Notably, TMEscore, which has a larger AUC than programmed death-ligand 1 com

Aug 3, 2026Read article

Data, analyses and reports in one workspace

Start with your next analysis

Choose a built-in analysis, add data, and review the run settings. Bring an existing Nextflow project through the import path.

Put omics knowledge on the research path · GeniOmics