Public knowledge document

Benchmarking of cell type deconvolution pipelines for transcriptomics data.

Many computational methods have been developed to infer cell type proportions from bulk transcriptomics data. However, an evaluation of the impact of data transformation, pre-processing, marker selection, cell type composition and choice of methodology on the deconvolution results is still lacking. Using five single-cell RNA-sequencing (scRNA-seq) datasets, we generate pseudo-bulk mixtures to evaluate the combined impact of these factors. Both bulk deconvolution methodologies and those that use scRNA-seq data as reference perform best when applied to data in linear scale and the choice of normalization has a dramatic impact on some, but not all methods. Overall, methods that use scRNA-seq data have comparable performance to the best performing bulk methods whereas semi-supervised approaches show higher error values. Moreover, failure to include cell types in the reference that are present in a mixture leads to substantially worse results, regardless of the previous choices. Altogether,

Benchmarking of cell type deconvolution pipelines for transcriptomics data.

> 商业许可源文 · EUROPE_PMC · [CC-BY](https://creativecommons.org/licenses/by/)

书目信息

  • 引用:Avila Cobos F, Alquicira-Hernandez J, Powell JE, Mestdagh P, De Preter K. (2020). Benchmarking of cell type deconvolution pipelines for transcriptomics data. Nature communications. PMID 33159064 · PMC7648640 · DOI 10.1038/s41467-020-19015-1
  • 证据类型:BENCHMARK
  • 主题:rna-seq、single-cell
  • 被引次数(采集时):366
  • 原始记录:[Europe PMC](https://europepmc.org/article/MED/33159064)
  • 来源许可:[CC-BY](https://creativecommons.org/licenses/by/)
  • 作者摘要(按来源许可复用)

    Many computational methods have been developed to infer cell type proportions from bulk transcriptomics data. However, an evaluation of the impact of data transformation, pre-processing, marker selection, cell type composition and choice of methodology on the deconvolution results is still lacking. Using five single-cell RNA-sequencing (scRNA-seq) datasets, we generate pseudo-bulk mixtures to evaluate the combined impact of these factors. Both bulk deconvolution methodologies and those that use scRNA-seq data as reference perform best when applied to data in linear scale and the choice of normalization has a dramatic impact on some, but not all methods. Overall, methods that use scRNA-seq data have comparable performance to the best performing bulk methods whereas semi-supervised approaches show higher error values. Moreover, failure to include cell types in the reference that are present in a mixture leads to substantially worse results, regardless of the previous choices. Altogether, we evaluate the combined impact of factors affecting the deconvolution task across different datasets and propose general guidelines to maximize its performance.

    合规说明

    本页保存的是来源文献书目信息及其在 CC-BY 许可下公开的作者摘要。除去除来源 HTML 标签和规范化空白外,摘要未作内容改写。本页不代表 GeniOmics 的医学建议;原文版权、署名和许可仍归原权利人,请通过原始记录核对最新版本、更正或撤稿状态。

    Benchmarking of cell type deconvolution pipelines for transcriptomics data. · GeniOmics