Public knowledge document

Comparison and evaluation of statistical error models for scRNA-seq.

Background Heterogeneity in single-cell RNA-seq (scRNA-seq) data is driven by multiple sources, including biological variation in cellular state as well as technical variation introduced during experimental processing. Deconvolving these effects is a key challenge for preprocessing workflows. Recent work has demonstrated the importance and utility of count models for scRNA-seq analysis, but there is a lack of consensus on which statistical distributions and parameter settings are appropriate. Results Here, we analyze 59 scRNA-seq datasets that span a wide range of technologies, systems, and sequencing depths in order to evaluate the performance of different error models. We find that while a Poisson error model appears appropriate for sparse datasets, we observe clear evidence of overdispersion for genes with sufficient sequencing depth in all biological systems, necessitating the use of a negative binomial model. Moreover, we find that the degree of overdispersion varies widely across

Comparison and evaluation of statistical error models for scRNA-seq.

> 商业许可源文 · EUROPE_PMC · [CC-BY](https://creativecommons.org/licenses/by/)

书目信息

  • 引用:Choudhary S, Satija R. (2022). Comparison and evaluation of statistical error models for scRNA-seq. Genome biology. PMID 35042561 · PMC8764781 · DOI 10.1186/s13059-021-02584-9
  • 证据类型:PRIMARY_RESEARCH
  • 主题:rna-seq、single-cell
  • 被引次数(采集时):618
  • 原始记录:[Europe PMC](https://europepmc.org/article/MED/35042561)
  • 来源许可:[CC-BY](https://creativecommons.org/licenses/by/)
  • 作者摘要(按来源许可复用)

    Background Heterogeneity in single-cell RNA-seq (scRNA-seq) data is driven by multiple sources, including biological variation in cellular state as well as technical variation introduced during experimental processing. Deconvolving these effects is a key challenge for preprocessing workflows. Recent work has demonstrated the importance and utility of count models for scRNA-seq analysis, but there is a lack of consensus on which statistical distributions and parameter settings are appropriate. Results Here, we analyze 59 scRNA-seq datasets that span a wide range of technologies, systems, and sequencing depths in order to evaluate the performance of different error models. We find that while a Poisson error model appears appropriate for sparse datasets, we observe clear evidence of overdispersion for genes with sufficient sequencing depth in all biological systems, necessitating the use of a negative binomial model. Moreover, we find that the degree of overdispersion varies widely across datasets, systems, and gene abundances, and argues for a data-driven approach for parameter estimation. Conclusions Based on these analyses, we provide a set of recommendations for modeling variation in scRNA-seq data, particularly when using generalized linear models or likelihood-based approaches for preprocessing and downstream analysis.

    合规说明

    本页保存的是来源文献书目信息及其在 CC-BY 许可下公开的作者摘要。除去除来源 HTML 标签和规范化空白外,摘要未作内容改写。本页不代表 GeniOmics 的医学建议;原文版权、署名和许可仍归原权利人,请通过原始记录核对最新版本、更正或撤稿状态。

    Comparison and evaluation of statistical error models for scRNA-seq. · GeniOmics