专题电子书
单细胞组学循证手册
按质量评分、证据类型和发表时间组织;优先收录明确允许商业复用的来源文献。
290 个章节
- 01Interpretation of T cell states from single-cell transcriptomics data using reference atlases.Single-cell RNA sequencing (scRNA-seq) has revealed an unprecedented degree of immune cell diversity. However, consistent definition of cell subtypes and cell states across studies and diseases remains a major challenge. Here we generate reference T cell atlases for cancer and viral infection by multi-study integration, and develop ProjecTILs, an algorithm for reference atlas projection. In contrast to other methods, ProjecTILs allows not only accurate embedding of new scRNA-seq data into a reference without altering its structure, but also characterizing previously unknown cell states that "deviate" from the reference. ProjecTILs accurately predicts the effects of cell perturbations and identifies gene programs that are altered in different conditions and tissues. A meta-analysis of tumor-infiltrating T cells from several cohorts reveals a strong conservation of T cell subtypes between human and mouse, providing a consistent basis to describe T cell heterogeneity across studies, disea
- 02Single-cell and bulk transcriptome sequencing identifies two epithelial tumor cell states and refines the consensus molecular classification of colorectal cancer.The consensus molecular subtype (CMS) classification of colorectal cancer is based on bulk transcriptomics. The underlying epithelial cell diversity remains unclear. We analyzed 373,058 single-cell transcriptomes from 63 patients, focusing on 49,155 epithelial cells. We identified a pervasive genetic and transcriptomic dichotomy of malignant cells, based on distinct gene expression, DNA copy number and gene regulatory network. We recapitulated these subtypes in bulk transcriptomes from 3,614 patients. The two intrinsic subtypes, iCMS2 and iCMS3, refine CMS. iCMS3 comprises microsatellite unstable (MSI-H) cancers and one-third of microsatellite-stable (MSS) tumors. iCMS3 MSS cancers are transcriptomically more similar to MSI-H cancers than to other MSS cancers. CMS4 cancers had either iCMS2 or iCMS3 epithelium; the latter had the worst prognosis. We defined the intrinsic epithelial axis of colorectal cancer and propose a refined 'IMF' classification with five subtypes, combining intrins
- 03Profiling the heterogeneity of colorectal cancer consensus molecular subtypes using spatial transcriptomics.The consensus molecular subtypes (CMS) of colorectal cancer (CRC) is the most widely-used gene expression-based classification and has contributed to a better understanding of disease heterogeneity and prognosis. Nevertheless, CMS intratumoral heterogeneity restricts its clinical application, stressing the necessity of further characterizing the composition and architecture of CRC. Here, we used Spatial Transcriptomics (ST) in combination with single-cell RNA sequencing (scRNA-seq) to decipher the spatially resolved cellular and molecular composition of CRC. In addition to mapping the intratumoral heterogeneity of CMS and their microenvironment, we identified cell communication events in the tumor-stroma interface of CMS2 carcinomas. This includes tumor growth-inhibiting as well as -activating signals, such as the potential regulation of the ETV4 transcriptional activity by DCN or the PLAU-PLAUR ligand-receptor interaction. Our study illustrates the potential of ST to resolve CRC molec
- 04Definitions and guidelines for research on antibiotic persistence.Increasing concerns about the rising rates of antibiotic therapy failure and advances in single-cell analyses have inspired a surge of research into antibiotic persistence. Bacterial persister cells represent a subpopulation of cells that can survive intensive antibiotic treatment without being resistant. Several approaches have emerged to define and measure persistence, and it is now time to agree on the basic definition of persistence and its relation to the other mechanisms by which bacteria survive exposure to bactericidal antibiotic treatments, such as antibiotic resistance, heteroresistance or tolerance. In this Consensus Statement, we provide definitions of persistence phenomena, distinguish between triggered and spontaneous persistence and provide a guide to measuring persistence. Antibiotic persistence is not only an interesting example of non-genetic single-cell heterogeneity, it may also have a role in the failure of antibiotic treatments. Therefore, it is our hope that the
- 05Guidelines for bioinformatics of single-cell sequencing data analysis in Alzheimer's disease: review, recommendation, implementation and application.Alzheimer's disease (AD) is the most common form of dementia, characterized by progressive cognitive impairment and neurodegeneration. Extensive clinical and genomic studies have revealed biomarkers, risk factors, pathways, and targets of AD in the past decade. However, the exact molecular basis of AD development and progression remains elusive. The emerging single-cell sequencing technology can potentially provide cell-level insights into the disease. Here we systematically review the state-of-the-art bioinformatics approaches to analyze single-cell sequencing data and their applications to AD in 14 major directions, including 1) quality control and normalization, 2) dimension reduction and feature extraction, 3) cell clustering analysis, 4) cell type inference and annotation, 5) differential expression, 6) trajectory inference, 7) copy number variation analysis, 8) integration of single-cell multi-omics, 9) epigenomic analysis, 10) gene network inference, 11) prioritization of cell s
- 06Benchmarking atlas-level data integration in single-cell genomics.Single-cell atlases often include samples that span locations, laboratories and conditions, leading to complex, nested batch effects in data. Thus, joint analysis of atlas datasets requires reliable data integration. To guide integration method choice, we benchmarked 68 method and preprocessing combinations on 85 batches of gene expression, chromatin accessibility and simulation data from 23 publications, altogether representing >1.2 million cells distributed in 13 atlas-level integration tasks. We evaluated methods according to scalability, usability and their ability to remove batch effects while retaining biological variation using 14 evaluation metrics. We show that highly variable gene selection improves the performance of data integration methods, whereas scaling pushes methods to prioritize batch removal over conservation of biological variation. Overall, scANVI, Scanorama, scVI and scGen perform well, particularly on complex integration tasks, while single-cell ATAC-sequencing
- 07An atlas of cells in the human tonsil.Palatine tonsils are secondary lymphoid organs (SLOs) representing the first line of immunological defense against inhaled or ingested pathogens. We generated an atlas of the human tonsil composed of >556,000 cells profiled across five different data modalities, including single-cell transcriptome, epigenome, proteome, and immune repertoire sequencing, as well as spatial transcriptomics. This census identified 121 cell types and states, defined developmental trajectories, and enabled an understanding of the functional units of the tonsil. Exemplarily, we stratified myeloid slan-like subtypes, established a BCL6 enhancer as locally active in follicle-associated T and B cells, and identified SIX5 as putative transcriptional regulator of plasma cell maturation. Analyses of a validation cohort confirmed the presence, annotation, and markers of tonsillar cell types and provided evidence of age-related compositional shifts. We demonstrate the value of this resource by annotating cells from B
- 08A human embryonic limb cell atlas resolved in space and time.Human limbs emerge during the fourth post-conception week as mesenchymal buds, which develop into fully formed limbs over the subsequent months 1 . This process is orchestrated by numerous temporally and spatially restricted gene expression programmes, making congenital alterations in phenotype common 2 . Decades of work with model organisms have defined the fundamental mechanisms underlying vertebrate limb development, but an in-depth characterization of this process in humans has yet to be performed. Here we detail human embryonic limb development across space and time using single-cell and spatial transcriptomics. We demonstrate extensive diversification of cells from a few multipotent progenitors to myriad differentiated cell states, including several novel cell populations. We uncover two waves of human muscle development, each characterized by different cell states regulated by separate gene expression programmes, and identify musculin (MSC) as a key transcriptional repressor mai
- 09A high-resolution transcriptomic and spatial atlas of cell types in the whole mouse brain.The mammalian brain consists of millions to billions of cells that are organized into many cell types with specific spatial distribution patterns and structural and functional properties 1-3 . Here we report a comprehensive and high-resolution transcriptomic and spatial cell-type atlas for the whole adult mouse brain. The cell-type atlas was created by combining a single-cell RNA-sequencing (scRNA-seq) dataset of around 7 million cells profiled (approximately 4.0 million cells passing quality control), and a spatial transcriptomic dataset of approximately 4.3 million cells using multiplexed error-robust fluorescence in situ hybridization (MERFISH). The atlas is hierarchically organized into 4 nested levels of classification: 34 classes, 338 subclasses, 1,201 supertypes and 5,322 clusters. We present an online platform, Allen Brain Cell Atlas, to visualize the mouse whole-brain cell-type atlas along with the single-cell RNA-sequencing and MERFISH datasets. We systematically analysed the
- 10CellMarker 2.0: an updated database of manually curated cell markers in human/mouse and web tools based on scRNA-seq data.CellMarker 2.0 (http://bio-bigdata.hrbmu.edu.cn/CellMarker or http://117.50.127.228/CellMarker/) is an updated database that provides a manually curated collection of experimentally supported markers of various cell types in different tissues of human and mouse. In addition, web tools for analyzing single cell sequencing data are described. We have updated CellMarker 2.0 with more data and several new features, including (i) Appending 36 300 tissue-cell type-maker entries, 474 tissues, 1901 cell types and 4566 markers over the previous version. The current release recruits 26 915 cell markers, 2578 cell types and 656 tissues, resulting in a total of 83 361 tissue-cell type-maker entries. (ii) There is new marker information from 48 sequencing technology sources, including 10X Chromium, Smart-Seq2 and Drop-seq, etc. (iii) Adding 29 types of cell markers, including protein-coding gene lncRNA and processed pseudogene, etc. Additionally, six flexible web tools, including cell annotation, c
- 11An integrated cell atlas of the lung in health and disease.Single-cell technologies have transformed our understanding of human tissues. Yet, studies typically capture only a limited number of donors and disagree on cell type definitions. Integrating many single-cell datasets can address these limitations of individual studies and capture the variability present in the population. Here we present the integrated Human Lung Cell Atlas (HLCA), combining 49 datasets of the human respiratory system into a single atlas spanning over 2.4 million cells from 486 individuals. The HLCA presents a consensus cell type re-annotation with matching marker genes, including annotations of rare and previously undescribed cell types. Leveraging the number and diversity of individuals in the HLCA, we identify gene modules that are associated with demographic covariates such as age, sex and body mass index, as well as gene modules changing expression along the proximal-to-distal axis of the bronchial tree. Mapping new data to the HLCA enables rapid data annotation
- 12An atlas of healthy and injured cell states and niches in the human kidney.Understanding kidney disease relies on defining the complexity of cell types and states, their associated molecular profiles and interactions within tissue neighbourhoods 1 . Here we applied multiple single-cell and single-nucleus assays (>400,000 nuclei or cells) and spatial imaging technologies to a broad spectrum of healthy reference kidneys (45 donors) and diseased kidneys (48 patients). This has provided a high-resolution cellular atlas of 51 main cell types, which include rare and previously undescribed cell populations. The multi-omic approach provides detailed transcriptomic profiles, regulatory factors and spatial localizations spanning the entire kidney. We also define 28 cellular states across nephron segments and interstitium that were altered in kidney injury, encompassing cycling, adaptive (successful or maladaptive repair), transitioning and degenerative states. Molecular signatures permitted the localization of these states within injury neighbourhoods using spatial tra
- 13Comparison of methods and resources for cell-cell communication inference from single-cell RNA-Seq data.The growing availability of single-cell data, especially transcriptomics, has sparked an increased interest in the inference of cell-cell communication. Many computational tools were developed for this purpose. Each of them consists of a resource of intercellular interactions prior knowledge and a method to predict potential cell-cell communication events. Yet the impact of the choice of resource and method on the resulting predictions is largely unknown. To shed light on this, we systematically compare 16 cell-cell communication inference resources and 7 methods, plus the consensus between the methods' predictions. Among the resources, we find few unique interactions, a varying degree of overlap, and an uneven coverage of specific pathways and tissue-enriched proteins. We then examine all possible combinations of methods and resources and show that both strongly influence the predicted intercellular interactions. Finally, we assess the agreement of cell-cell communication methods with
- 14Mapping single-cell data to reference atlases by transfer learning.Large single-cell atlases are now routinely generated to serve as references for analysis of smaller-scale studies. Yet learning from reference data is complicated by batch effects between datasets, limited availability of computational resources and sharing restrictions on raw data. Here we introduce a deep learning strategy for mapping query datasets on top of a reference called single-cell architectural surgery (scArches). scArches uses transfer learning and parameter optimization to enable efficient, decentralized, iterative reference building and contextualization of new datasets with existing references without sharing raw data. Using examples from mouse brain, pancreas, immune and whole-organism atlases, we show that scArches preserves biological state information while removing batch effects, despite using four orders of magnitude fewer parameters than de novo integration. scArches generalizes to multimodal reference mapping, allowing imputation of missing modalities. Finally,
- 15A multimodal cell census and atlas of the mammalian primary motor cortex.Here we report the generation of a multimodal cell census and atlas of the mammalian primary motor cortex as the initial product of the BRAIN Initiative Cell Census Network (BICCN). This was achieved by coordinated large-scale analyses of single-cell transcriptomes, chromatin accessibility, DNA methylomes, spatially resolved single-cell transcriptomes, morphological and electrophysiological properties and cellular resolution input-output mapping, integrated through cross-modal computational analysis. Our results advance the collective knowledge and understanding of brain cell-type organization 1-5 . First, our study reveals a unified molecular genetic landscape of cortical cell types that integrates their transcriptome, open chromatin and DNA methylation maps. Second, cross-species analysis achieves a consensus taxonomy of transcriptomic types and their hierarchical organization that is conserved from mouse to marmoset and human. Third, in situ single-cell transcriptomics provides a sp
- 16Spatially resolved cell atlas of the mouse primary motor cortex by MERFISH.A mammalian brain is composed of numerous cell types organized in an intricate manner to form functional neural circuits. Single-cell RNA sequencing allows systematic identification of cell types based on their gene expression profiles and has revealed many distinct cell populations in the brain 1,2 . Single-cell epigenomic profiling 3,4 further provides information on gene-regulatory signatures of different cell types. Understanding how different cell types contribute to brain function, however, requires knowledge of their spatial organization and connectivity, which is not preserved in sequencing-based methods that involve cell dissociation. Here we used a single-cell transcriptome-imaging method, multiplexed error-robust fluorescence in situ hybridization (MERFISH) 5 , to generate a molecularly defined and spatially resolved cell atlas of the mouse primary motor cortex. We profiled approximately 300,000 cells in the mouse primary motor cortex and its adjacent areas, identified 95 ne
- 17Molecularly defined and spatially resolved cell atlas of the whole mouse brain.In mammalian brains, millions to billions of cells form complex interaction networks to enable a wide range of functions. The enormous diversity and intricate organization of cells have impeded our understanding of the molecular and cellular basis of brain function. Recent advances in spatially resolved single-cell transcriptomics have enabled systematic mapping of the spatial organization of molecularly defined cell types in complex tissues 1-3 , including several brain regions (for example, refs. 1-11 ). However, a comprehensive cell atlas of the whole brain is still missing. Here we imaged a panel of more than 1,100 genes in approximately 10 million cells across the entire adult mouse brains using multiplexed error-robust fluorescence in situ hybridization 12 and performed spatially resolved, single-cell expression profiling at the whole-transcriptome scale by integrating multiplexed error-robust fluorescence in situ hybridization and single-cell RNA sequencing data. Using this appr
- 18Single-Cell Atlas of Lineage States, Tumor Microenvironment, and Subtype-Specific Expression Programs in Gastric Cancer.Gastric cancer heterogeneity represents a barrier to disease management. We generated a comprehensive single-cell atlas of gastric cancer (>200,000 cells) comprising 48 samples from 31 patients across clinical stages and histologic subtypes. We identified 34 distinct cell-lineage states including novel rare cell populations. Many lineage states exhibited distinct cancer-associated expression profiles, individually contributing to a combined tumor-wide molecular collage. We observed increased plasma cell proportions in diffuse-type tumors associated with epithelial-resident KLF2 and stage-wise accrual of cancer-associated fibroblast subpopulations marked by high INHBA and FAP coexpression. Single-cell comparisons between patient-derived organoids (PDO) and primary tumors highlighted inter- and intralineage similarities and differences, demarcating molecular boundaries of PDOs as experimental models. We complemented these findings by spatial transcriptomics, orthogonal validation in inde
- 19High-resolution single-cell atlas reveals diversity and plasticity of tissue-resident neutrophils in non-small cell lung cancer.Non-small cell lung cancer (NSCLC) is characterized by molecular heterogeneity with diverse immune cell infiltration patterns, which has been linked to therapy sensitivity and resistance. However, full understanding of how immune cell phenotypes vary across different patient subgroups is lacking. Here, we dissect the NSCLC tumor microenvironment at high resolution by integrating 1,283,972 single cells from 556 samples and 318 patients across 29 datasets, including our dataset capturing cells with low mRNA content. We stratify patients into immune-deserted, B cell, T cell, and myeloid cell subtypes. Using bulk samples with genomic and clinical information, we identify cellular components associated with tumor histology and genotypes. We then focus on the analysis of tissue-resident neutrophils (TRNs) and uncover distinct subpopulations that acquire new functional properties in the tissue microenvironment, providing evidence for the plasticity of TRNs. Finally, we show that a TRN-derived
- 20Single-cell atlas of early human brain development highlights heterogeneity of human neuroepithelial cells and early radial glia.The human cortex comprises diverse cell types that emerge from an initially uniform neuroepithelium that gives rise to radial glia, the neural stem cells of the cortex. To characterize the earliest stages of human brain development, we performed single-cell RNA-sequencing across regions of the developing human brain, including the telencephalon, diencephalon, midbrain, hindbrain and cerebellum. We identify nine progenitor populations physically proximal to the telencephalon, suggesting more heterogeneity than previously described, including a highly prevalent mesenchymal-like population that disappears once neurogenesis begins. Comparison of human and mouse progenitor populations at corresponding stages identifies two progenitor clusters that are enriched in the early stages of human cortical development. We also find that organoid systems display low fidelity to neuroepithelial and early radial glia cell types, but improve as neurogenesis progresses. Overall, we provide a comprehensiv
- 21A single-cell atlas of the multicellular ecosystem of primary and metastatic hepatocellular carcinoma.Hepatocellular carcinoma (HCC) represents a paradigm of the relation between tumor microenvironment (TME) and tumor development. Here, we generate a single-cell atlas of the multicellular ecosystem of HCC from four tissue sites. We show the enrichment of central memory T cells (T CM ) in the early tertiary lymphoid structures (E-TLSs) in HCC and assess the relationships between chronic HBV/HCV infection and T cell infiltration and exhaustion. We find the MMP9 + macrophages to be terminally differentiated tumor-associated macrophages (TAMs) and PPARγ to be the pivotal transcription factor driving their differentiation. We also characterize the heterogeneous subpopulations of malignant hepatocytes and their multifaceted functions in shaping the immune microenvironment of HCC. Finally, we identify seven microenvironment-based subtypes that can predict prognosis of HCC patients. Collectively, this large-scale atlas deepens our understanding of the HCC microenvironment, which might facilita
- 22DNA methylation atlas of the mouse brain at single-cell resolution.Mammalian brain cells show remarkable diversity in gene expression, anatomy and function, yet the regulatory DNA landscape underlying this extensive heterogeneity is poorly understood. Here we carry out a comprehensive assessment of the epigenomes of mouse brain cell types by applying single-nucleus DNA methylation sequencing 1,2 to profile 103,982 nuclei (including 95,815 neurons and 8,167 non-neuronal cells) from 45 regions of the mouse cortex, hippocampus, striatum, pallidum and olfactory areas. We identified 161 cell clusters with distinct spatial locations and projection targets. We constructed taxonomies of these epigenetic types, annotated with signature genes, regulatory elements and transcription factors. These features indicate the potential regulatory landscape supporting the assignment of putative cell types and reveal repetitive usage of regulators in excitatory and inhibitory cells for determining subtypes. The DNA methylation landscape of excitatory neurons in the cortex
- 23A single cell atlas of human cornea that defines its development, limbal progenitor cells and their interactions with the immune cells.Purpose Single cell (sc) analyses of key embryonic, fetal and adult stages were performed to generate a comprehensive single cell atlas of all the corneal and adjacent conjunctival cell types from development to adulthood. Methods Four human adult and seventeen embryonic and fetal corneas from 10 to 21 post conception week (PCW) specimens were dissociated to single cells and subjected to scRNA- and/or ATAC-Seq using the 10x Genomics platform. These were embedded using Uniform Manifold Approximation and Projection (UMAP) and clustered using Seurat graph-based clustering. Cluster identification was performed based on marker gene expression, bioinformatic data mining and immunofluorescence (IF) analysis. RNA interference, IF, colony forming efficiency and clonal assays were performed on cultured limbal epithelial cells (LECs). Results scRNA-Seq analysis of 21,343 cells from four adult human corneas and adjacent conjunctivas revealed the presence of 21 cell clusters, representing the proge
- 24Majorbio Cloud 2024: Update single-cell and multiomics workflows.Majorbio Cloud (https://cloud.majorbio.com/) is a one-stop online analytic platform aiming at promoting the development of bioinformatics services, narrowing the gap between wet and dry experiments, and accelerating the discoveries for the life sciences community. In 2024, three single-omics workflows, two multiomics workflows, and extensions were newly released to facilitate omics data mining and interpretation.
- 25Single-cell sequencing to multi-omics: technologies and applications.Cells, as the fundamental units of life, contain multidimensional spatiotemporal information. Single-cell RNA sequencing (scRNA-seq) is revolutionizing biomedical science by analyzing cellular state and intercellular heterogeneity. Undoubtedly, single-cell transcriptomics has emerged as one of the most vibrant research fields today. With the optimization and innovation of single-cell sequencing technologies, the intricate multidimensional details concealed within cells are gradually unveiled. The combination of scRNA-seq and other multi-omics is at the forefront of the single-cell field. This involves simultaneously measuring various omics data within individual cells, expanding our understanding across a broader spectrum of dimensions. Single-cell multi-omics precisely captures the multidimensional aspects of single-cell transcriptomes, immune repertoire, spatial information, temporal information, epitopes, and other omics in diverse spatiotemporal contexts. In addition to depicting t
- 26ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis.The advent of single-cell chromatin accessibility profiling has accelerated the ability to map gene regulatory landscapes but has outpaced the development of scalable software to rapidly extract biological meaning from these data. Here we present a software suite for single-cell analysis of regulatory chromatin in R (ArchR; https://www.archrproject.com/ ) that enables fast and comprehensive analysis of single-cell chromatin accessibility data. ArchR provides an intuitive, user-focused interface for complex single-cell analyses, including doublet removal, single-cell clustering and cell type identification, unified peak set generation, cellular trajectory identification, DNA element-to-gene linkage, transcription factor footprinting, mRNA expression level prediction from chromatin accessibility and multi-omic integration with single-cell RNA sequencing (scRNA-seq). Enabling the analysis of over 1.2 million single cells within 8 h on a standard Unix laptop, ArchR is a comprehensive softw
- 27Single-cell RNA sequencing technologies and applications: A brief overview.Single-cell RNA sequencing (scRNA-seq) technology has become the state-of-the-art approach for unravelling the heterogeneity and complexity of RNA transcripts within individual cells, as well as revealing the composition of different cell types and functions within highly organized tissues/organs/organisms. Since its first discovery in 2009, studies based on scRNA-seq provide massive information across different fields making exciting new discoveries in better understanding the composition and interaction of cells within humans, model animals and plants. In this review, we provide a concise overview about the scRNA-seq technology, experimental and computational procedures for transforming the biological and molecular processes into computational and statistical data. We also provide an explanation of the key technological steps in implementing the technology. We highlight a few examples on how scRNA-seq can provide unique information for better understanding health and diseases. One im
- 28An introduction to spatial transcriptomics for biomedical research.Single-cell transcriptomics (scRNA-seq) has become essential for biomedical research over the past decade, particularly in developmental biology, cancer, immunology, and neuroscience. Most commercially available scRNA-seq protocols require cells to be recovered intact and viable from tissue. This has precluded many cell types from study and largely destroys the spatial context that could otherwise inform analyses of cell identity and function. An increasing number of commercially available platforms now facilitate spatially resolved, high-dimensional assessment of gene transcription, known as 'spatial transcriptomics'. Here, we introduce different classes of method, which either record the locations of hybridized mRNA molecules in tissue, image the positions of cells themselves prior to assessment, or employ spatial arrays of mRNA probes of pre-determined location. We review sizes of tissue area that can be assessed, their spatial resolution, and the number and types of genes that can
- 29Statistics or biology: the zero-inflation controversy about scRNA-seq data.Researchers view vast zeros in single-cell RNA-seq data differently: some regard zeros as biological signals representing no or low gene expression, while others regard zeros as missing data to be corrected. To help address the controversy, here we discuss the sources of biological and non-biological zeros; introduce five mechanisms of adding non-biological zeros in computational benchmarking; evaluate the impacts of non-biological zeros on data analysis; benchmark three input data types: observed counts, imputed counts, and binarized counts; discuss the open questions regarding non-biological zeros; and advocate the importance of transparent analysis.
- 30Applications of single-cell sequencing in cancer research: progress and perspectives.Single-cell sequencing, including genomics, transcriptomics, epigenomics, proteomics and metabolomics sequencing, is a powerful tool to decipher the cellular and molecular landscape at a single-cell resolution, unlike bulk sequencing, which provides averaged data. The use of single-cell sequencing in cancer research has revolutionized our understanding of the biological characteristics and dynamics within cancer lesions. In this review, we summarize emerging single-cell sequencing technologies and recent cancer research progress obtained by single-cell sequencing, including information related to the landscapes of malignant cells and immune cells, tumor heterogeneity, circulating tumor cells and the underlying mechanisms of tumor biological behaviors. Overall, the prospects of single-cell sequencing in facilitating diagnosis, targeted therapy and prognostic prediction among a spectrum of tumors are bright. In the near future, advances in single-cell sequencing will undoubtedly improve
- 31From bulk, single-cell to spatial RNA sequencing.RNA sequencing (RNAseq) can reveal gene fusions, splicing variants, mutations/indels in addition to differential gene expression, thus providing a more complete genetic picture than DNA sequencing. This most widely used technology in genomics tool box has evolved from classic bulk RNA sequencing (RNAseq), popular single cell RNA sequencing (scRNAseq) to newly emerged spatial RNA sequencing (spRNAseq). Bulk RNAseq studies average global gene expression, scRNAseq investigates single cell RNA biology up to 20,000 individual cells simultaneously, while spRNAseq has ability to dissect RNA activities spatially, representing next generation of RNA sequencing. This article highlights these technologies, characteristic features and suitable applications in precision oncology.
- 32Macrophages and microglia in glioblastoma: heterogeneity, plasticity, and therapy.Glioblastoma (GBM) is the most aggressive tumor in the central nervous system and contains a highly immunosuppressive tumor microenvironment (TME). Tumor-associated macrophages and microglia (TAMs) are a dominant population of immune cells in the GBM TME that contribute to most GBM hallmarks, including immunosuppression. The understanding of TAMs in GBM has been limited by the lack of powerful tools to characterize them. However, recent progress on single-cell technologies offers an opportunity to precisely characterize TAMs at the single-cell level and identify new TAM subpopulations with specific tumor-modulatory functions in GBM. In this Review, we discuss TAM heterogeneity and plasticity in the TME and summarize current TAM-targeted therapeutic potential in GBM. We anticipate that the use of single-cell technologies followed by functional studies will accelerate the development of novel and effective TAM-targeted therapeutics for GBM patients.
- 33Single-cell RNA sequencing in cancer research.Single-cell RNA sequencing (scRNA-seq), a technology that analyzes transcriptomes of complex tissues at single-cell levels, can identify differential gene expression and epigenetic factors caused by mutations in unicellular genomes, as well as new cell-specific markers and cell types. scRNA-seq plays an important role in various aspects of tumor research. It reveals the heterogeneity of tumor cells and monitors the progress of tumor development, thereby preventing further cellular deterioration. Furthermore, the transcriptome analysis of immune cells in tumor tissue can be used to classify immune cells, their immune escape mechanisms and drug resistance mechanisms, and to develop effective clinical targeted therapies combined with immunotherapy. Moreover, this method enables the study of intercellular communication and the interaction of tumor cells and non-malignant cells to reveal their role in carcinogenesis. scRNA-seq provides new technical means for further development of tumor re
- 34Applications of multi-omics analysis in human diseases.Multi-omics usually refers to the crossover application of multiple high-throughput screening technologies represented by genomics, transcriptomics, single-cell transcriptomics, proteomics and metabolomics, spatial transcriptomics, and so on, which play a great role in promoting the study of human diseases. Most of the current reviews focus on describing the development of multi-omics technologies, data integration, and application to a particular disease; however, few of them provide a comprehensive and systematic introduction of multi-omics. This review outlines the existing technical categories of multi-omics, cautions for experimental design, focuses on the integrated analysis methods of multi-omics, especially the approach of machine learning and deep learning in multi-omics data integration and the corresponding tools, and the application of multi-omics in medical researches (e.g., cancer, neurodegenerative diseases, aging, and drug target discovery) as well as the corresponding
- 35Clinical and translational values of spatial transcriptomics.The combination of spatial transcriptomics (ST) and single cell RNA sequencing (scRNA-seq) acts as a pivotal component to bridge the pathological phenomes of human tissues with molecular alterations, defining in situ intercellular molecular communications and knowledge on spatiotemporal molecular medicine. The present article overviews the development of ST and aims to evaluate clinical and translational values for understanding molecular pathogenesis and uncovering disease-specific biomarkers. We compare the advantages and disadvantages of sequencing- and imaging-based technologies and highlight opportunities and challenges of ST. We also describe the bioinformatics tools necessary on dissecting spatial patterns of gene expression and cellular interactions and the potential applications of ST in human diseases for clinical practice as one of important issues in clinical and translational medicine, including neurology, embryo development, oncology, and inflammation. Thus, clear clinica
- 36Advances in single-cell omics and multiomics for high-resolution molecular profiling.Single-cell omics technologies have revolutionized molecular profiling by providing high-resolution insights into cellular heterogeneity and complexity. Traditional bulk omics approaches average signals from heterogeneous cell populations, thereby obscuring important cellular nuances. Single-cell omics studies enable the analysis of individual cells and reveal diverse cell types, dynamic cellular states, and rare cell populations. These techniques offer unprecedented resolution and sensitivity, enabling researchers to unravel the molecular landscape of individual cells. Furthermore, the integration of multimodal omics data within a single cell provides a comprehensive and holistic view of cellular processes. By combining multiple omics dimensions, multimodal omics approaches can facilitate the elucidation of complex cellular interactions, regulatory networks, and molecular mechanisms. This integrative approach enhances our understanding of cellular systems, from development to disease.
- 37Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression.Single-cell RNA-seq (scRNA-seq) data exhibits significant cell-to-cell variation due to technical factors, including the number of molecules detected in each cell, which can confound biological heterogeneity with technical effects. To address this, we present a modeling framework for the normalization and variance stabilization of molecular count data from scRNA-seq experiments. We propose that the Pearson residuals from "regularized negative binomial regression," where cellular sequencing depth is utilized as a covariate in a generalized linear model, successfully remove the influence of technical characteristics from downstream analyses while preserving biological heterogeneity. Importantly, we show that an unconstrained negative binomial model may overfit scRNA-seq data, and overcome this by pooling information across genes with similar abundances to obtain stable parameter estimates. Our procedure omits the need for heuristic steps including pseudocount addition or log-transformati
- 38Cell Hashing with barcoded antibodies enables multiplexing and doublet detection for single cell genomics.Despite rapid developments in single cell sequencing, sample-specific batch effects, detection of cell multiplets, and experimental costs remain outstanding challenges. Here, we introduce Cell Hashing, where oligo-tagged antibodies against ubiquitously expressed surface proteins uniquely label cells from distinct samples, which can be subsequently pooled. By sequencing these tags alongside the cellular transcriptome, we can assign each cell to its original sample, robustly identify cross-sample multiplets, and "super-load" commercial droplet-based systems for significant cost reduction. We validate our approach using a complementary genetic approach and demonstrate how hashing can generalize the benefits of single cell multiplexing to diverse samples and experimental designs.
- 39A benchmark of batch-effect correction methods for single-cell RNA sequencing data.Background Large-scale single-cell transcriptomic datasets generated using different technologies contain batch-specific systematic variations that present a challenge to batch-effect removal and data integration. With continued growth expected in scRNA-seq data, achieving effective batch integration with available computational resources is crucial. Here, we perform an in-depth benchmark study on available batch correction methods to determine the most suitable method for batch-effect removal. Results We compare 14 methods in terms of computational runtime, the ability to handle large datasets, and batch-effect correction efficacy while preserving cell type purity. Five scenarios are designed for the study: identical cell types with different technologies, non-identical cell types, multiple batches, big data, and simulated data. Performance is evaluated using four benchmarking metrics including kBET, LISI, ASW, and ARI. We also investigate the use of batch-corrected data to study diff
- 40MOFA+: a statistical framework for comprehensive integration of multi-modal single-cell data.Technological advances have enabled the profiling of multiple molecular layers at single-cell resolution, assaying cells from multiple samples or conditions. Consequently, there is a growing need for computational strategies to analyze data from complex experimental designs that include multiple data modalities and multiple groups of samples. We present Multi-Omics Factor Analysis v2 (MOFA+), a statistical framework for the comprehensive and scalable integration of single-cell multi-modal data. MOFA+ reconstructs a low-dimensional representation of the data using computationally efficient variational inference and supports flexible sparsity constraints, allowing to jointly model variation across multiple sample groups and data modalities.
- 41A single-cell and single-nucleus RNA-Seq toolbox for fresh and frozen human tumors.Single-cell genomics is essential to chart tumor ecosystems. Although single-cell RNA-Seq (scRNA-Seq) profiles RNA from cells dissociated from fresh tumors, single-nucleus RNA-Seq (snRNA-Seq) is needed to profile frozen or hard-to-dissociate tumors. Each requires customization to different tissue and tumor types, posing a barrier to adoption. Here, we have developed a systematic toolbox for profiling fresh and frozen clinical tumor samples using scRNA-Seq and snRNA-Seq, respectively. We analyzed 216,490 cells and nuclei from 40 samples across 23 specimens spanning eight tumor types of varying tissue and sample characteristics. We evaluated protocols by cell and nucleus quality, recovery rate and cellular composition. scRNA-Seq and snRNA-Seq from matched samples recovered the same cell types, but at different proportions. Our work provides guidance for studies in a broad range of tumors, including criteria for testing and selecting methods from the toolbox for other tumors, thus paving
- 42Systematic assessment of tissue dissociation and storage biases in single-cell and single-nucleus RNA-seq workflows.Background Single-cell RNA sequencing has been widely adopted to estimate the cellular composition of heterogeneous tissues and obtain transcriptional profiles of individual cells. Multiple approaches for optimal sample dissociation and storage of single cells have been proposed as have single-nuclei profiling methods. What has been lacking is a systematic comparison of their relative biases and benefits. Results Here, we compare gene expression and cellular composition of single-cell suspensions prepared from adult mouse kidney using two tissue dissociation protocols. For each sample, we also compare fresh cells to cryopreserved and methanol-fixed cells. Lastly, we compare this single-cell data to that generated using three single-nucleus RNA sequencing workflows. Our data confirms prior reports that digestion on ice avoids the stress response observed with 37 °C dissociation. It also reveals cell types more abundant either in the cold or warm dissociations that may represent populati
- 43An accurate and robust imputation method scImpute for single-cell RNA-seq data.The emerging single-cell RNA sequencing (scRNA-seq) technologies enable the investigation of transcriptomic landscapes at the single-cell resolution. ScRNA-seq data analysis is complicated by excess zero counts, the so-called dropouts due to low amounts of mRNA sequenced within individual cells. We introduce scImpute, a statistical method to accurately and robustly impute the dropouts in scRNA-seq data. scImpute automatically identifies likely dropouts, and only perform imputation on these values without introducing new biases to the rest data. scImpute also detects outlier cells and excludes them from imputation. Evaluation based on both simulated and real human and mouse scRNA-seq data suggests that scImpute is an effective tool to recover transcriptome dynamics masked by dropouts. scImpute is shown to identify likely dropouts, enhance the clustering of cell subpopulations, improve the accuracy of differential expression analysis, and aid the study of gene expression dynamics.
- 44scNMT-seq enables joint profiling of chromatin accessibility DNA methylation and transcription in single cells.Parallel single-cell sequencing protocols represent powerful methods for investigating regulatory relationships, including epigenome-transcriptome interactions. Here, we report a single-cell method for parallel chromatin accessibility, DNA methylation and transcriptome profiling. scNMT-seq (single-cell nucleosome, methylation and transcription sequencing) uses a GpC methyltransferase to label open chromatin followed by bisulfite and RNA sequencing. We validate scNMT-seq by applying it to differentiating mouse embryonic stem cells, finding links between all three molecular layers and revealing dynamic coupling between epigenomic layers during differentiation.
- 45A comparison of automatic cell identification methods for single-cell RNA sequencing data.Background Single-cell transcriptomics is rapidly advancing our understanding of the cellular composition of complex tissues and organisms. A major limitation in most analysis pipelines is the reliance on manual annotations to determine cell identities, which are time-consuming and irreproducible. The exponential growth in the number of cells and samples has prompted the adaptation and development of supervised classification methods for automatic cell identification. Results Here, we benchmarked 22 classification methods that automatically assign cell identities including single-cell-specific and general-purpose classifiers. The performance of the methods is evaluated using 27 publicly available single-cell RNA sequencing datasets of different sizes, technologies, species, and levels of complexity. We use 2 experimental setups to evaluate the performance of each method for within dataset predictions (intra-dataset) and across datasets (inter-dataset) based on accuracy, percentage of u
- 46scPred: accurate supervised method for cell-type classification from single-cell RNA-seq data.Single-cell RNA sequencing has enabled the characterization of highly specific cell types in many tissues, as well as both primary and stem cell-derived cell lines. An important facet of these studies is the ability to identify the transcriptional signatures that define a cell type or state. In theory, this information can be used to classify an individual cell based on its transcriptional profile. Here, we present scPred, a new generalizable method that is able to provide highly accurate classification of single cells, using a combination of unbiased feature selection from a reduced-dimension space, and machine-learning probability-based prediction method. We apply scPred to scRNA-seq data from pancreatic tissue, mononuclear cells, colorectal tumor biopsies, and circulating dendritic cells and show that scPred is able to classify individual cells with high accuracy. The generalized method is available at https://github.com/powellgenomicslab/scPred/.
- 47Feature selection and dimension reduction for single-cell RNA-Seq based on a multinomial model.Single-cell RNA-Seq (scRNA-Seq) profiles gene expression of individual cells. Recent scRNA-Seq datasets have incorporated unique molecular identifiers (UMIs). Using negative controls, we show UMI counts follow multinomial sampling with no zero inflation. Current normalization procedures such as log of counts per million and feature selection by highly variable genes produce false variability in dimension reduction. We propose simple multinomial methods, including generalized principal component analysis (GLM-PCA) for non-normal distributions, and feature selection using deviance. These methods outperform the current practice in a downstream clustering assessment using ground truth datasets.
- 48Benchmarking of cell type deconvolution pipelines for transcriptomics data.Many computational methods have been developed to infer cell type proportions from bulk transcriptomics data. However, an evaluation of the impact of data transformation, pre-processing, marker selection, cell type composition and choice of methodology on the deconvolution results is still lacking. Using five single-cell RNA-sequencing (scRNA-seq) datasets, we generate pseudo-bulk mixtures to evaluate the combined impact of these factors. Both bulk deconvolution methodologies and those that use scRNA-seq data as reference perform best when applied to data in linear scale and the choice of normalization has a dramatic impact on some, but not all methods. Overall, methods that use scRNA-seq data have comparable performance to the best performing bulk methods whereas semi-supervised approaches show higher error values. Moreover, failure to include cell types in the reference that are present in a mixture leads to substantially worse results, regardless of the previous choices. Altogether,
- 49Vireo: Bayesian demultiplexing of pooled single-cell RNA-seq data without genotype reference.Multiplexed single-cell RNA-seq analysis of multiple samples using pooling is a promising experimental design, offering increased throughput while allowing to overcome batch variation. To reconstruct the sample identify of each cell, genetic variants that segregate between the samples in the pool have been proposed as natural barcode for cell demultiplexing. Existing demultiplexing strategies rely on availability of complete genotype data from the pooled samples, which limits the applicability of such methods, in particular when genetic variation is not the primary object of study. To address this, we here present Vireo, a computationally efficient Bayesian model to demultiplex single-cell data from pooled experimental designs. Uniquely, our model can be applied in settings when only partial or no genotype information is available. Using pools based on synthetic mixtures and results on real data, we demonstrate the robustness of Vireo and illustrate the utility of multiplexed experimen
- 50MetaCell: analysis of single-cell RNA-seq data using K-nn graph partitions.scRNA-seq profiles each represent a highly partial sample of mRNA molecules from a unique cell that can never be resampled, and robust analysis must separate the sampling effect from biological variance. We describe a methodology for partitioning scRNA-seq datasets into metacells: disjoint and homogenous groups of profiles that could have been resampled from the same cell. Unlike clustering analysis, our algorithm specializes at obtaining granular as opposed to maximal groups. We show how to use metacells as building blocks for complex quantitative transcriptional maps while avoiding data smoothing. Our algorithms are implemented in the MetaCell R/C++ software package.
- 51CZ CELLxGENE Discover: a single-cell data platform for scalable exploration, analysis and modeling of aggregated data.Hundreds of millions of single cells have been analyzed using high-throughput transcriptomic methods. The cumulative knowledge within these datasets provides an exciting opportunity for unlocking insights into health and disease at the level of single cells. Meta-analyses that span diverse datasets building on recent advances in large language models and other machine-learning approaches pose exciting new directions to model and extract insight from single-cell data. Despite the promise of these and emerging analytical tools for analyzing large amounts of data, the sheer number of datasets, data models and accessibility remains a challenge. Here, we present CZ CELLxGENE Discover (cellxgene.cziscience.com), a data platform that provides curated and interoperable single-cell data. Available via a free-to-use online data portal, CZ CELLxGENE hosts a growing corpus of community-contributed data of over 93 million unique cells. Curated, standardized and associated with consistent cell-level
- 52Dictionary of immune responses to cytokines at single-cell resolution.Cytokines mediate cell-cell communication in the immune system and represent important therapeutic targets 1-3 . A myriad of studies have highlighted their central role in immune function 4-13 , yet we lack a global view of the cellular responses of each immune cell type to each cytokine. To address this gap, we created the Immune Dictionary, a compendium of single-cell transcriptomic profiles of more than 17 immune cell types in response to each of 86 cytokines (>1,400 cytokine-cell type combinations) in mouse lymph nodes in vivo. A cytokine-centric view of the dictionary revealed that most cytokines induce highly cell-type-specific responses. For example, the inflammatory cytokine interleukin-1β induces distinct gene programmes in almost every cell type. A cell-type-centric view of the dictionary identified more than 66 cytokine-driven cellular polarization states across immune cell types, including previously uncharacterized states such as an interleukin-18-induced polyfunctional na
- 53Nanopore long-read RNAseq reveals widespread transcriptional variation among the surface receptors of individual B cells.Understanding gene regulation and function requires a genome-wide method capable of capturing both gene expression levels and isoform diversity at the single-cell level. Short-read RNAseq is limited in its ability to resolve complex isoforms because it fails to sequence full-length cDNA copies of RNA molecules. Here, we investigate whether RNAseq using the long-read single-molecule Oxford Nanopore MinION sequencer is able to identify and quantify complex isoforms without sacrificing accurate gene expression quantification. After benchmarking our approach, we analyse individual murine B1a cells using a custom multiplexing strategy. We identify thousands of unannotated transcription start and end sites, as well as hundreds of alternative splicing events in these B1a cells. We also identify hundreds of genes expressed across B1a cells that display multiple complex isoforms, including several B cell-specific surface receptors. Our results show that we can identify and quantify complex isof
- 54scRNA-seq assessment of the human lung, spleen, and esophagus tissue stability after cold preservation.Background The Human Cell Atlas is a large international collaborative effort to map all cell types of the human body. Single-cell RNA sequencing can generate high-quality data for the delivery of such an atlas. However, delays between fresh sample collection and processing may lead to poor data and difficulties in experimental design. Results This study assesses the effect of cold storage on fresh healthy spleen, esophagus, and lung from ≥ 5 donors over 72 h. We collect 240,000 high-quality single-cell transcriptomes with detailed cell type annotations and whole genome sequences of donors, enabling future eQTL studies. Our data provide a valuable resource for the study of these 3 organs and will allow cross-organ comparison of cell types. We see little effect of cold ischemic time on cell yield, total number of reads per cell, and other quality control metrics in any of the tissues within the first 24 h. However, we observe a decrease in the proportions of lung T cells at 72 h, higher
- 55Assessment of computational methods for the analysis of single-cell ATAC-seq data.Background Recent innovations in single-cell Assay for Transposase Accessible Chromatin using sequencing (scATAC-seq) enable profiling of the epigenetic landscape of thousands of individual cells. scATAC-seq data analysis presents unique methodological challenges. scATAC-seq experiments sample DNA, which, due to low copy numbers (diploid in humans), lead to inherent data sparsity (1-10% of peaks detected per cell) compared to transcriptomic (scRNA-seq) data (10-45% of expressed genes detected per cell). Such challenges in data generation emphasize the need for informative features to assess cell heterogeneity at the chromatin level. Results We present a benchmarking framework that is applied to 10 computational methods for scATAC-seq on 13 synthetic and real datasets from different assays, profiling cell types from diverse tissues and organisms. Methods for processing and featurizing scATAC-seq data were compared by their ability to discriminate cell types when combined with common uns
- 56Integrative analyses of single-cell transcriptome and regulome using MAESTRO.We present Model-based AnalysEs of Transcriptome and RegulOme (MAESTRO), a comprehensive open-source computational workflow ( http://github.com/liulab-dfci/MAESTRO ) for the integrative analyses of single-cell RNA-seq (scRNA-seq) and ATAC-seq (scATAC-seq) data from multiple platforms. MAESTRO provides functions for pre-processing, alignment, quality control, expression and chromatin accessibility quantification, clustering, differential analysis, and annotation. By modeling gene regulatory potential from chromatin accessibilities at the single-cell level, MAESTRO outperforms the existing methods for integrating the cell clusters between scRNA-seq and scATAC-seq. Furthermore, MAESTRO supports automatic cell-type annotation using predefined cell type marker genes and identifies driver regulators from differential scRNA-seq genes and scATAC-seq peaks.
- 57Advances in spatial transcriptomics and related data analysis strategies.Spatial transcriptomics technologies developed in recent years can provide various information including tissue heterogeneity, which is fundamental in biological and medical research, and have been making significant breakthroughs. Single-cell RNA sequencing (scRNA-seq) cannot provide spatial information, while spatial transcriptomics technologies allow gene expression information to be obtained from intact tissue sections in the original physiological context at a spatial resolution. Various biological insights can be generated into tissue architecture and further the elucidation of the interaction between cells and the microenvironment. Thus, we can gain a general understanding of histogenesis processes and disease pathogenesis, etc. Furthermore, in silico methods involving the widely distributed R and Python packages for data analysis play essential roles in deriving indispensable bioinformation and eliminating technological limitations. In this review, we summarize available techno
- 58Evaluation of cell-cell interaction methods by integrating single-cell RNA sequencing data with spatial information.Background Cell-cell interactions are important for information exchange between different cells, which are the fundamental basis of many biological processes. Recent advances in single-cell RNA sequencing (scRNA-seq) enable the characterization of cell-cell interactions using computational methods. However, it is hard to evaluate these methods since no ground truth is provided. Spatial transcriptomics (ST) data profiles the relative position of different cells. We propose that the spatial distance suggests the interaction tendency of different cell types, thus could be used for evaluating cell-cell interaction tools. Results We benchmark 16 cell-cell interaction methods by integrating scRNA-seq with ST data. We characterize cell-cell interactions into short-range and long-range interactions using spatial distance distributions between ligands and receptors. Based on this classification, we define the distance enrichment score and apply an evaluation workflow to 16 cell-cell interactio
- 59Spatial multi-omics analyses of the tumor immune microenvironment.In the past decade, single-cell technologies have revealed the heterogeneity of the tumor-immune microenvironment at the genomic, transcriptomic, and proteomic levels and have furthered our understanding of the mechanisms of tumor development. Single-cell technologies have also been used to identify potential biomarkers. However, spatial information about the tumor-immune microenvironment such as cell locations and cell-cell interactomes is lost in these approaches. Recently, spatial multi-omics technologies have been used to study transcriptomes, proteomes, and metabolomes of tumor-immune microenvironments in several types of cancer, and the data obtained from these methods has been combined with immunohistochemistry and multiparameter analysis to yield markers of cancer progression. Here, we review numerous cutting-edge spatial 'omics techniques, their application to study of the tumor-immune microenvironment, and remaining technical challenges.
- 60Psoriatic Arthritis: Pathogenesis and Targeted Therapies.Psoriatic arthritis (PsA), a heterogeneous chronic inflammatory immune-mediated disease characterized by musculoskeletal inflammation (arthritis, enthesitis, spondylitis, and dactylitis), generally occurs in patients with psoriasis. PsA is also associated with uveitis and inflammatory bowel disease (Crohn's disease and ulcerative colitis). To capture these manifestations as well as the associated comorbidities, and to recognize their underlining common pathogenesis, the name of psoriatic disease was coined. The pathogenesis of PsA is complex and multifaceted, with an interplay of genetic predisposition, triggering environmental factors, and activation of the innate and adaptive immune system, although autoinflammation has also been implicated. Research has identified several immune-inflammatory pathways defined by cytokines (IL-23/IL-17, TNF), leading to the development of efficacious therapeutic targets. However, heterogeneous responses to these drugs occur in different patients and i
- 61Computational Approaches and Challenges in Spatial Transcriptomics.The development of spatial transcriptomics (ST) technologies has transformed genetic research from a single-cell data level to a two-dimensional spatial coordinate system and facilitated the study of the composition and function of various cell subsets in different environments and organs. The large-scale data generated by these ST technologies, which contain spatial gene expression information, have elicited the need for spatially resolved approaches to meet the requirements of computational and biological data interpretation. These requirements include dealing with the explosive growth of data to determine the cell-level and gene-level expression, correcting the inner batch effect and loss of expression to improve the data quality, conducting efficient interpretation and in-depth knowledge mining both at the single-cell and tissue-wide levels, and conducting multi-omics integration analysis to provide an extensible framework toward the in-depth understanding of biological processes.
- 62Spatial omics: Navigating to the golden era of cancer research.The idea that tumour microenvironment (TME) is organised in a spatial manner will not surprise many cancer biologists; however, systematically capturing spatial architecture of TME is still not possible until recent decade. The past five years have witnessed a boom in the research of high-throughput spatial techniques and algorithms to delineate TME at an unprecedented level. Here, we review the technological progress of spatial omics and how advanced computation methods boost multi-modal spatial data analysis. Then, we discussed the potential clinical translations of spatial omics research in precision oncology, and proposed a transfer of spatial ecological principles to cancer biology in spatial data interpretation. So far, spatial omics is placing us in the golden age of spatial cancer research. Further development and application of spatial omics may lead to a comprehensive decoding of the TME ecosystem and bring the current spatiotemporal molecular medical research into an entirel
- 63The Human Cell Atlas.The recent advent of methods for high-throughput single-cell molecular profiling has catalyzed a growing sense in the scientific community that the time is ripe to complete the 150-year-old effort to identify all cell types in the human body. The Human Cell Atlas Project is an international collaborative effort that aims to define all human cell types in terms of distinctive molecular profiles (such as gene expression profiles) and to connect this information with classical cellular descriptions (such as location and morphology). An open comprehensive reference map of the molecular state of cells in healthy human tissues would propel the systematic study of physiological states, developmental trajectories, regulatory circuitry and interactions of cells, and also provide a framework for understanding cellular dysregulation in human disease. Here we describe the idea, its potential utility, early proofs-of-concept, and some design considerations for the Human Cell Atlas, including a comm
- 64Collagen-producing lung cell atlas identifies multiple subsets with distinct localization and relevance to fibrosis.Collagen-producing cells maintain the complex architecture of the lung and drive pathologic scarring in pulmonary fibrosis. Here we perform single-cell RNA-sequencing to identify all collagen-producing cells in normal and fibrotic lungs. We characterize multiple collagen-producing subpopulations with distinct anatomical localizations in different compartments of murine lungs. One subpopulation, characterized by expression of Cthrc1 (collagen triple helix repeat containing 1), emerges in fibrotic lungs and expresses the highest levels of collagens. Single-cell RNA-sequencing of human lungs, including those from idiopathic pulmonary fibrosis and scleroderma patients, demonstrate similar heterogeneity and CTHRC1-expressing fibroblasts present uniquely in fibrotic lungs. Immunostaining and in situ hybridization show that these cells are concentrated within fibroblastic foci. We purify collagen-producing subpopulations and find disease-relevant phenotypes of Cthrc1-expressing fibroblasts in
- 65An atlas of the aging lung mapped by single cell transcriptomics and deep tissue proteomics.Aging promotes lung function decline and susceptibility to chronic lung diseases, which are the third leading cause of death worldwide. Here, we use single cell transcriptomics and mass spectrometry-based proteomics to quantify changes in cellular activity states across 30 cell types and chart the lung proteome of young and old mice. We show that aging leads to increased transcriptional noise, indicating deregulated epigenetic control. We observe cell type-specific effects of aging, uncovering increased cholesterol biosynthesis in type-2 pneumocytes and lipofibroblasts and altered relative frequency of airway epithelial cells as hallmarks of lung aging. Proteomic profiling reveals extracellular matrix remodeling in old mice, including increased collagen IV and XVI and decreased Fraser syndrome complex proteins and collagen XIV. Computational integration of the aging proteome with the single cell transcriptomes predicts the cellular source of regulated proteins and creates an unbiased r
- 66Expression Atlas update: from tissues to single cells.Expression Atlas is EMBL-EBI's resource for gene and protein expression. It sources and compiles data on the abundance and localisation of RNA and proteins in various biological systems and contexts and provides open access to this data for the research community. With the increased availability of single cell RNA-Seq datasets in the public archives, we have now extended Expression Atlas with a new added-value service to display gene expression in single cells. Single Cell Expression Atlas was launched in 2018 and currently includes 123 single cell RNA-Seq studies from 12 species. The website can be searched by genes within or across species to reveal experiments, tissues and cell types where this gene is expressed or under which conditions it is a marker gene. Within each study, cells can be visualized using a pre-calculated t-SNE plot and can be coloured by different features or by cell clusters based on gene expression. Within each experiment, there are links to downloadable files,
- 67Revealing the Critical Regulators of Cell Identity in the Mouse Cell Atlas.Recent progress in single-cell technologies has enabled the identification of all major cell types in mouse. However, for most cell types, the regulatory mechanism underlying their identity remains poorly understood. By computational analysis of the recently published mouse cell atlas data, we have identified 202 regulons whose activities are highly variable across different cell types, and more importantly, predicted a small set of essential regulators for each major cell type in mouse. Systematic validation by automated literature and data mining provides strong additional support for our predictions. Thus, these predictions serve as a valuable resource that would be useful for the broad biological community. Finally, we have built a user-friendly, interactive web portal to enable users to navigate this mouse cell network atlas.
- 68Multiomic spatial landscape of innate immune cells at human central nervous system borders.The innate immune compartment of the human central nervous system (CNS) is highly diverse and includes several immune-cell populations such as macrophages that are frequent in the brain parenchyma (microglia) and less numerous at the brain interfaces as CNS-associated macrophages (CAMs). Due to their scantiness and particular location, little is known about the presence of temporally and spatially restricted CAM subclasses during development, health and perturbation. Here we combined single-cell RNA sequencing, time-of-flight mass cytometry and single-cell spatial transcriptomics with fate mapping and advanced immunohistochemistry to comprehensively characterize the immune system at human CNS interfaces with over 356,000 analyzed transcriptomes from 102 individuals. We also provide a comprehensive analysis of resident and engrafted myeloid cells in the brains of 15 individuals with peripheral blood stem cell transplantation, revealing compartment-specific engraftment rates across diffe
- 69Single-cell and spatial transcriptomics analysis of non-small cell lung cancer.Lung cancer is the second most frequently diagnosed cancer and the leading cause of cancer-related mortality worldwide. Tumour ecosystems feature diverse immune cell types. Myeloid cells, in particular, are prevalent and have a well-established role in promoting the disease. In our study, we profile approximately 900,000 cells from 25 treatment-naive patients with adenocarcinoma and squamous-cell carcinoma by single-cell and spatial transcriptomics. We note an inverse relationship between anti-inflammatory macrophages and NK cells/T cells, and with reduced NK cell cytotoxicity within the tumour. While we observe a similar cell type composition in both adenocarcinoma and squamous-cell carcinoma, we detect significant differences in the co-expression of various immune checkpoint inhibitors. Moreover, we reveal evidence of a transcriptional "reprogramming" of macrophages in tumours, shifting them towards cholesterol export and adopting a foetal-like transcriptional signature which promote
- 70Unveiling inflammatory and prehypertrophic cell populations as key contributors to knee cartilage degeneration in osteoarthritis using multi-omics data integration.Objectives Single-cell and spatial transcriptomics analysis of human knee articular cartilage tissue to present a comprehensive transcriptome landscape and osteoarthritis (OA)-critical cell populations. Methods Single-cell RNA sequencing and spatially resolved transcriptomic technology have been applied to characterise the cellular heterogeneity of human knee articular cartilage which were collected from 8 OA donors, and 3 non-OA control donors, and a total of 19 samples. The novel chondrocyte population and marker genes of interest were validated by immunohistochemistry staining, quantitative real-time PCR, etc. The OA-critical cell populations were validated through integrative analyses of publicly available bulk RNA sequencing data and large-scale genome-wide association studies. Results We identified 33 cell population-specific marker genes that define 11 chondrocyte populations, including 9 known populations and 2 new populations, that is, pre-inflammatory chondrocyte population (
- 71Spatial transcriptomics reveal neuron-astrocyte synergy in long-term memory.Memory encodes past experiences, thereby enabling future plans. The basolateral amygdala is a centre of salience networks that underlie emotional experiences and thus has a key role in long-term fear memory formation 1 . Here we used spatial and single-cell transcriptomics to illuminate the cellular and molecular architecture of the role of the basolateral amygdala in long-term memory. We identified transcriptional signatures in subpopulations of neurons and astrocytes that were memory-specific and persisted for weeks. These transcriptional signatures implicate neuropeptide and BDNF signalling, MAPK and CREB activation, ubiquitination pathways, and synaptic connectivity as key components of long-term memory. Notably, upon long-term memory formation, a neuronal subpopulation defined by increased Penk and decreased Tac expression constituted the most prominent component of the memory engram of the basolateral amygdala. These transcriptional changes were observed both with single-cell RNA
- 72Flow Cytometry: The Next Revolution.Unmasking the subtleties of the immune system requires both a comprehensive knowledge base and the ability to interrogate that system with intimate sensitivity. That task, to a considerable extent, has been handled by an iterative expansion in flow cytometry methods, both in technological capability and also in accompanying advances in informatics. As the field of fluorescence-based cytomics matured, it reached a technological barrier at around 30 parameter analyses, which stalled the field until spectral flow cytometry created a fundamental transformation that will likely lead to the potential of 100 simultaneous parameter analyses within a few years. The simultaneous advance in informatics has now become a watershed moment for the field as it competes with mature systematic approaches such as genomics and proteomics, allowing cytomics to take a seat at the multi-omics table. In addition, recent technological advances try to combine the speed of flow systems with other detection metho
- 73Spatial omics technologies at multimodal and single cell/subcellular level.Spatial omics technologies enable a deeper understanding of cellular organizations and interactions within a tissue of interest. These assays can identify specific compartments or regions in a tissue with differential transcript or protein abundance, delineate their interactions, and complement other methods in defining cellular phenotypes. A variety of spatial methodologies are being developed and commercialized; however, these techniques differ in spatial resolution, multiplexing capability, scale/throughput, and coverage. Here, we review the current and prospective landscape of single cell to subcellular resolution spatial omics technologies and analysis tools to provide a comprehensive picture for both research and clinical applications.
- 74stPlus: a reference-based method for the accurate enhancement of spatial transcriptomics.Motivation Single-cell RNA sequencing (scRNA-seq) techniques have revolutionized the investigation of transcriptomic landscape in individual cells. Recent advancements in spatial transcriptomic technologies further enable gene expression profiling and spatial organization mapping of cells simultaneously. Among the technologies, imaging-based methods can offer higher spatial resolutions, while they are limited by either the small number of genes imaged or the low gene detection sensitivity. Although several methods have been proposed for enhancing spatially resolved transcriptomics, inadequate accuracy of gene expression prediction and insufficient ability of cell-population identification still impede the applications of these methods. Results We propose stPlus, a reference-based method that leverages information in scRNA-seq data to enhance spatial transcriptomics. Based on an auto-encoder with a carefully tailored loss function, stPlus performs joint embedding and predicts spatial ge
- 75Spatial Transcriptomics: A Powerful Tool in Disease Understanding and Drug Discovery.Recent advancements in modern science have provided robust tools for drug discovery. The rapid development of transcriptome sequencing technologies has given rise to single-cell transcriptomics and single-nucleus transcriptomics, increasing the accuracy of sequencing and accelerating the drug discovery process. With the evolution of single-cell transcriptomics, spatial transcriptomics (ST) technology has emerged as a derivative approach. Spatial transcriptomics has emerged as a hot topic in the field of omics research in recent years; it not only provides information on gene expression levels but also offers spatial information on gene expression. This technology has shown tremendous potential in research on disease understanding and drug discovery. In this article, we introduce the analytical strategies of spatial transcriptomics and review its applications in novel target discovery and drug mechanism unravelling. Moreover, we discuss the current challenges and issues in this research
- 76Advancements in single-cell RNA sequencing and spatial transcriptomics: transforming biomedical research.In recent years, significant advancements in biochemistry, materials science, engineering, and computer-aided testing have driven the development of high-throughput tools for profiling genetic information. Single-cell RNA sequencing (scRNA-seq) technologies have established themselves as key tools for dissecting genetic sequences at the level of single cells. These technologies reveal cellular diversity and allow for the exploration of cell states and transformations with exceptional resolution. Unlike bulk sequencing, which provides population-averaged data, scRNA-seq can detect cell subtypes or gene expression variations that would otherwise be overlooked. However, a key limitation of scRNA-seq is its inability to preserve spatial information about the RNA transcriptome, as the process requires tissue dissociation and cell isolation. Spatial transcriptomics is a pivotal advancement in medical biotechnology, facilitating the identification of molecules such as RNA in their original sp
- 77Single Cell Atlas: a single-cell multi-omics human cell encyclopedia.Single-cell sequencing datasets are key in biology and medicine for unraveling insights into heterogeneous cell populations with unprecedented resolution. Here, we construct a single-cell multi-omics map of human tissues through in-depth characterizations of datasets from five single-cell omics, spatial transcriptomics, and two bulk omics across 125 healthy adult and fetal tissues. We construct its complement web-based platform, the Single Cell Atlas (SCA, www.singlecellatlas.org ), to enable vast interactive data exploration of deep multi-omics signatures across human fetal and adult tissues. The atlas resources and database queries aspire to serve as a one-stop, comprehensive, and time-effective resource for various omics studies.
- 78Integrated analysis of multimodal single-cell data.The simultaneous measurement of multiple modalities represents an exciting frontier for single-cell genomics and necessitates computational methods that can define cellular states based on multimodal data. Here, we introduce "weighted-nearest neighbor" analysis, an unsupervised framework to learn the relative utility of each data type in each cell, enabling an integrative analysis of multiple modalities. We apply our procedure to a CITE-seq dataset of 211,000 human peripheral blood mononuclear cells (PBMCs) with panels extending to 228 antibodies to construct a multimodal reference atlas of the circulating immune system. Multimodal analysis substantially improves our ability to resolve cell states, allowing us to identify and validate previously unreported lymphoid subpopulations. Moreover, we demonstrate how to leverage this reference to rapidly map new datasets and to interpret immune responses to vaccination and coronavirus disease 2019 (COVID-19). Our approach represents a broadly
- 79Inference and analysis of cell-cell communication using CellChat.Understanding global communications among cells requires accurate representation of cell-cell signaling links and effective systems-level analyses of those links. We construct a database of interactions among ligands, receptors and their cofactors that accurately represent known heteromeric molecular complexes. We then develop CellChat, a tool that is able to quantitatively infer and analyze intercellular communication networks from single-cell RNA-sequencing (scRNA-seq) data. CellChat predicts major signaling inputs and outputs for cells and how those cells and signals coordinate for functions using network analysis and pattern recognition approaches. Through manifold learning and quantitative contrasts, CellChat classifies signaling pathways and delineates conserved and context-specific pathways across different datasets. Applying CellChat to mouse and human skin datasets shows its ability to extract complex signaling patterns. Our versatile and easy-to-use toolkit CellChat and a web
- 80The history and advances in cancer immunotherapy: understanding the characteristics of tumor-infiltrating immune cells and their therapeutic implications.Immunotherapy has revolutionized cancer treatment and rejuvenated the field of tumor immunology. Several types of immunotherapy, including adoptive cell transfer (ACT) and immune checkpoint inhibitors (ICIs), have obtained durable clinical responses, but their efficacies vary, and only subsets of cancer patients can benefit from them. Immune infiltrates in the tumor microenvironment (TME) have been shown to play a key role in tumor development and will affect the clinical outcomes of cancer patients. Comprehensive profiling of tumor-infiltrating immune cells would shed light on the mechanisms of cancer-immune evasion, thus providing opportunities for the development of novel therapeutic strategies. However, the highly heterogeneous and dynamic nature of the TME impedes the precise dissection of intratumoral immune cells. With recent advances in single-cell technologies such as single-cell RNA sequencing (scRNA-seq) and mass cytometry, systematic interrogation of the TME is feasible and
- 81Current best practices in single-cell RNA-seq analysis: a tutorial.Single-cell RNA-seq has enabled gene expression to be studied at an unprecedented resolution. The promise of this technology is attracting a growing user base for single-cell analysis methods. As more analysis tools are becoming available, it is becoming increasingly difficult to navigate this landscape and produce an up-to-date workflow to analyse one's data. Here, we detail the steps of a typical single-cell RNA-seq analysis, including pre-processing (quality control, normalization, data correction, feature selection, and dimensionality reduction) and cell- and gene-level downstream analysis. We formulate current best-practice recommendations for these steps based on independent comparison studies. We have integrated these best-practice recommendations into a workflow, which we apply to a public dataset to further illustrate how these steps work in practice. Our documented case study can be found at https://www.github.com/theislab/single-cell-tutorial This review will serve as a work
- 82A step-by-step workflow for low-level analysis of single-cell RNA-seq data with Bioconductor.Single-cell RNA sequencing (scRNA-seq) is widely used to profile the transcriptome of individual cells. This provides biological resolution that cannot be matched by bulk RNA sequencing, at the cost of increased technical noise and data complexity. The differences between scRNA-seq and bulk RNA-seq data mean that the analysis of the former cannot be performed by recycling bioinformatics pipelines for the latter. Rather, dedicated single-cell methods are required at various steps to exploit the cellular resolution while accounting for technical noise. This article describes a computational workflow for low-level analyses of scRNA-seq data, based primarily on software packages from the open-source Bioconductor project. It covers basic steps including quality control, data exploration and normalization, as well as more complex procedures such as cell cycle phase assignment, identification of highly variable and correlated genes, clustering into subpopulations and marker gene detection. An
- 83Doublet identification in single-cell sequencing data using scDblFinder .Doublets are prevalent in single-cell sequencing data and can lead to artifactual findings. A number of strategies have therefore been proposed to detect them. Building on the strengths of existing approaches, we developed scDblFinder , a fast, flexible and accurate Bioconductor-based doublet detection method. Here we present the method, justify its design choices, demonstrate its performance on both single-cell RNA and accessibility (ATAC) sequencing data, and provide some observations on doublet formation, detection, and enrichment analysis. Even in complex datasets, scDblFinder can accurately identify most heterotypic doublets, and was already found by an independent benchmark to outcompete alternatives.
- 84IOBR: Multi-Omics Immuno-Oncology Biological Research to Decode Tumor Microenvironment and Signatures.Recent advances in next-generation sequencing (NGS) technologies have triggered the rapid accumulation of publicly available multi-omics datasets. The application of integrated omics to explore robust signatures for clinical translation is increasingly emphasized, and this is attributed to the clinical success of immune checkpoint blockades in diverse malignancies. However, effective tools for comprehensively interpreting multi-omics data are still warranted to provide increased granularity into the intrinsic mechanism of oncogenesis and immunotherapeutic sensitivity. Therefore, we developed a computational tool for effective Immuno-Oncology Biological Research (IOBR), providing a comprehensive investigation of the estimation of reported or user-built signatures, TME deconvolution, and signature construction based on multi-omics data. Notably, IOBR offers batch analyses of these signatures and their correlations with clinical phenotypes, long non-coding RNA (lncRNA) profiling, genomic
- 85UCell: Robust and scalable single-cell gene signature scoring.UCell is an R package for evaluating gene signatures in single-cell datasets. UCell signature scores, based on the Mann-Whitney U statistic, are robust to dataset size and heterogeneity, and their calculation demands less computing time and memory than other available methods, enabling the processing of large datasets in a few minutes even on machines with limited computing power. UCell can be applied to any single-cell data matrix, and includes functions to directly interact with Seurat objects. The UCell package and documentation are available on GitHub at https://github.com/carmonalab/UCell.
- 86Eleven grand challenges in single-cell data science.The recent boom in microfluidics and combinatorial indexing strategies, combined with low sequencing costs, has empowered single-cell sequencing technology. Thousands-or even millions-of cells analyzed in a single experiment amount to a data revolution in single-cell biology and pose unique data science problems. Here, we outline eleven challenges that will be central to bringing this emerging field of single-cell data science forward. For each challenge, we highlight motivating research questions, review prior work, and formulate open problems. This compendium is for established researchers, newcomers, and students alike, highlighting interesting and rewarding problems for the coming years.
- 87A new gene set identifies senescent cells and predicts senescence-associated pathways across tissues.Although cellular senescence drives multiple age-related co-morbidities through the senescence-associated secretory phenotype, in vivo senescent cell identification remains challenging. Here, we generate a gene set (SenMayo) and validate its enrichment in bone biopsies from two aged human cohorts. We further demonstrate reductions in SenMayo in bone following genetic clearance of senescent cells in mice and in adipose tissue from humans following pharmacological senescent cell clearance. We next use SenMayo to identify senescent hematopoietic or mesenchymal cells at the single cell level from human and murine bone marrow/bone scRNA-seq data. Thus, SenMayo identifies senescent cells across tissues and species with high fidelity. Using this senescence panel, we are able to characterize senescent cells at the single cell level and identify key intercellular signaling pathways. SenMayo also represents a potentially clinically applicable panel for monitoring senescent cell burden with aging
- 88Confronting false discoveries in single-cell differential expression.Differential expression analysis in single-cell transcriptomics enables the dissection of cell-type-specific responses to perturbations such as disease, trauma, or experimental manipulations. While many statistical methods are available to identify differentially expressed genes, the principles that distinguish these methods and their performance remain unclear. Here, we show that the relative performance of these methods is contingent on their ability to account for variation between biological replicates. Methods that ignore this inevitable variation are biased and prone to false discoveries. Indeed, the most widely used methods can discover hundreds of differentially expressed genes in the absence of biological differences. To exemplify these principles, we exposed true and false discoveries of differentially expressed genes in the injured mouse spinal cord.
- 89A practical guide to single-cell RNA-sequencing for biomedical research and clinical applications.RNA sequencing (RNA-seq) is a genomic approach for the detection and quantitative analysis of messenger RNA molecules in a biological sample and is useful for studying cellular responses. RNA-seq has fueled much discovery and innovation in medicine over recent years. For practical reasons, the technique is usually conducted on samples comprising thousands to millions of cells. However, this has hindered direct assessment of the fundamental unit of biology-the cell. Since the first single-cell RNA-sequencing (scRNA-seq) study was published in 2009, many more have been conducted, mostly by specialist laboratories with unique skills in wet-lab single-cell genomics, bioinformatics, and computation. However, with the increasing commercial availability of scRNA-seq platforms, and the rapid ongoing maturation of bioinformatics approaches, a point has been reached where any biomedical researcher or clinician can use scRNA-seq to make exciting discoveries. In this review, we present a practical
- 90Deep learning and alignment of spatially resolved single-cell transcriptomes with Tangram.Charting an organs' biological atlas requires us to spatially resolve the entire single-cell transcriptome, and to relate such cellular features to the anatomical scale. Single-cell and single-nucleus RNA-seq (sc/snRNA-seq) can profile cells comprehensively, but lose spatial information. Spatial transcriptomics allows for spatial measurements, but at lower resolution and with limited sensitivity. Targeted in situ technologies solve both issues, but are limited in gene throughput. To overcome these limitations we present Tangram, a method that aligns sc/snRNA-seq data to various forms of spatial data collected from the same region, including MERFISH, STARmap, smFISH, Spatial Transcriptomics (Visium) and histological images. Tangram can map any type of sc/snRNA-seq data, including multimodal data such as those from SHARE-seq, which we used to reveal spatial patterns of chromatin accessibility. We demonstrate Tangram on healthy mouse brain tissue, by reconstructing a genome-wide anatomica
- 91GSEApy: a comprehensive package for performing gene set enrichment analysis in Python.Motivation Gene set enrichment analysis (GSEA) is a commonly used algorithm for characterizing gene expression changes. However, the currently available tools used to perform GSEA have a limited ability to analyze large datasets, which is particularly problematic for the analysis of single-cell data. To overcome this limitation, we developed a GSEA package in Python (GSEApy), which could efficiently analyze large single-cell datasets. Results We present a package (GSEApy) that performs GSEA in either the command line or Python environment. GSEApy uses a Rust implementation to enable it to calculate the same enrichment statistic as GSEA for a collection of pathways. The Rust implementation of GSEApy is 3-fold faster than the Numpy version of GSEApy (v0.10.8) and uses >4-fold less memory. GSEApy also provides an interface between Python and Enrichr web services, as well as for BioMart. The Enrichr application programming interface enables GSEApy to perform over-representation analysis fo
- 92Single-Cell RNA-Seq Technologies and Related Computational Data Analysis.Single-cell RNA sequencing (scRNA-seq) technologies allow the dissection of gene expression at single-cell resolution, which greatly revolutionizes transcriptomic studies. A number of scRNA-seq protocols have been developed, and these methods possess their unique features with distinct advantages and disadvantages. Due to technical limitations and biological factors, scRNA-seq data are noisier and more complex than bulk RNA-seq data. The high variability of scRNA-seq data raises computational challenges in data analysis. Although an increasing number of bioinformatics methods are proposed for analyzing and interpreting scRNA-seq data, novel algorithms are required to ensure the accuracy and reproducibility of results. In this review, we provide an overview of currently available single-cell isolation protocols and scRNA-seq technologies, and discuss the methods for diverse scRNA-seq data analyses including quality control, read mapping, gene expression quantification, batch effect corr
- 93Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq.A central goal of genetics is to define the relationships between genotypes and phenotypes. High-content phenotypic screens such as Perturb-seq (CRISPR-based screens with single-cell RNA-sequencing readouts) enable massively parallel functional genomic mapping but, to date, have been used at limited scales. Here, we perform genome-scale Perturb-seq targeting all expressed genes with CRISPR interference (CRISPRi) across >2.5 million human cells. We use transcriptional phenotypes to predict the function of poorly characterized genes, uncovering new regulators of ribosome biogenesis (including CCDC86, ZNF236, and SPATA5L1), transcription (C7orf26), and mitochondrial respiration (TMEM242). In addition to assigning gene function, single-cell transcriptional phenotypes allow for in-depth dissection of complex cellular phenomena-from RNA processing to differentiation. We leverage this ability to systematically identify genetic drivers and consequences of aneuploidy and to discover an unantici
- 94CellRank for directed single-cell fate mapping.Computational trajectory inference enables the reconstruction of cell state dynamics from single-cell RNA sequencing experiments. However, trajectory inference requires that the direction of a biological process is known, largely limiting its application to differentiating systems in normal development. Here, we present CellRank ( https://cellrank.org ) for single-cell fate mapping in diverse scenarios, including regeneration, reprogramming and disease, for which direction is unknown. Our approach combines the robustness of trajectory inference with directional information from RNA velocity, taking into account the gradual and stochastic nature of cellular fate decisions, as well as uncertainty in velocity vectors. On pancreas development data, CellRank automatically detects initial, intermediate and terminal populations, predicts fate potentials and visualizes continuous gene expression trends along individual lineages. Applied to lineage-traced cellular reprogramming data, predicted
- 95Fully-automated and ultra-fast cell-type identification using specific marker combinations from single-cell transcriptomic data.Identification of cell populations often relies on manual annotation of cell clusters using established marker genes. However, the selection of marker genes is a time-consuming process that may lead to sub-optimal annotations as the markers must be informative of both the individual cell clusters and various cell types present in the sample. Here, we developed a computational platform, ScType, which enables a fully-automated and ultra-fast cell-type identification based solely on a given scRNA-seq data, along with a comprehensive cell marker database as background information. Using six scRNA-seq datasets from various human and mouse tissues, we show how ScType provides unbiased and accurate cell type annotations by guaranteeing the specificity of positive and negative marker genes across cell clusters and cell types. We also demonstrate how ScType distinguishes between healthy and malignant cell populations, based on single-cell calling of single-nucleotide variants, making it a versa
- 96Single-cell profiling of tumor heterogeneity and the microenvironment in advanced non-small cell lung cancer.Lung cancer is a highly heterogeneous disease. Cancer cells and cells within the tumor microenvironment together determine disease progression, as well as response to or escape from treatment. To map the cell type-specific transcriptome landscape of cancer cells and their tumor microenvironment in advanced non-small cell lung cancer (NSCLC), we analyze 42 tissue biopsy samples from stage III/IV NSCLC patients by single cell RNA sequencing and present the large scale, single cell resolution profiles of advanced NSCLCs. In addition to cell types described in previous single cell studies of early stage lung cancer, we are able to identify rare cell types in tumors such as follicular dendritic cells and T helper 17 cells. Tumors from different patients display large heterogeneity in cellular composition, chromosomal structure, developmental trajectory, intercellular signaling network and phenotype dominance. Our study also reveals a correlation of tumor heterogeneity with tumor associated
- 97SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks.Joint profiling of chromatin accessibility and gene expression in individual cells provides an opportunity to decipher enhancer-driven gene regulatory networks (GRNs). Here we present a method for the inference of enhancer-driven GRNs, called SCENIC+. SCENIC+ predicts genomic enhancers along with candidate upstream transcription factors (TFs) and links these enhancers to candidate target genes. To improve both recall and precision of TF identification, we curated and clustered a motif collection with more than 30,000 motifs. We benchmarked SCENIC+ on diverse datasets from different species, including human peripheral blood mononuclear cells, ENCODE cell lines, melanoma cell states and Drosophila retinal development. Next, we exploit SCENIC+ predictions to study conserved TFs, enhancers and GRNs between human and mouse cell types in the cerebral cortex. Finally, we use SCENIC+ to study the dynamics of gene regulation along differentiation trajectories and the effect of TF perturbations
- 98Effect of the intratumoral microbiota on spatial and cellular heterogeneity in cancer.The tumour-associated microbiota is an intrinsic component of the tumour microenvironment across human cancer types 1,2 . Intratumoral host-microbiota studies have so far largely relied on bulk tissue analysis 1-3 , which obscures the spatial distribution and localized effect of the microbiota within tumours. Here, by applying in situ spatial-profiling technologies 4 and single-cell RNA sequencing 5 to oral squamous cell carcinoma and colorectal cancer, we reveal spatial, cellular and molecular host-microbe interactions. We adapted 10x Visium spatial transcriptomics to determine the identity and in situ location of intratumoral microbial communities within patient tissues. Using GeoMx digital spatial profiling 6 , we show that bacterial communities populate microniches that are less vascularized, highly immuno‑suppressive and associated with malignant cells with lower levels of Ki-67 as compared to bacteria-negative tumour regions. We developed a single-cell RNA-sequencing method that
- 99The MTT Assay: Utility, Limitations, Pitfalls, and Interpretation in Bulk and Single-Cell Analysis.The MTT assay for cellular metabolic activity is almost ubiquitous to studies of cell toxicity; however, it is commonly applied and interpreted erroneously. We investigated the applicability and limitations of the MTT assay in representing treatment toxicity, cell viability, and metabolic activity. We evaluated the effect of potential confounding variables on the MTT assay measurements on a prostate cancer cell line (PC-3) including cell seeding number, MTT concentration, MTT incubation time, serum starvation, cell culture media composition, released intracellular contents (cell lysate and secretome), and extrusion of formazan to the extracellular space. We also assessed the confounding effect of polyethylene glycol (PEG)-coated gold nanoparticles (Au-NPs) as a tested treatment in PC-3 cells on the assay measurements. We additionally evaluated the applicability of microscopic image cytometry as a tool for measuring intracellular MTT reduction at the single-cell level. Our findings show
- 100High resolution mapping of the tumor microenvironment using integrated single-cell, spatial and in situ analysis.Single-cell and spatial technologies that profile gene expression across a whole tissue are revolutionizing the resolution of molecular states in clinical samples. Current commercially available technologies provide whole transcriptome single-cell, whole transcriptome spatial, or targeted in situ gene expression analysis. Here, we combine these technologies to explore tissue heterogeneity in large, FFPE human breast cancer sections. This integrative approach allowed us to explore molecular differences that exist between distinct tumor regions and to identify biomarkers involved in the progression towards invasive carcinoma. Further, we study cell neighborhoods and identify rare boundary cells that sit at the critical myoepithelial border confining the spread of malignant cells. Here, we demonstrate that each technology alone provides information about molecular signatures relevant to understanding cancer heterogeneity; however, it is the integration of these technologies that leads to
- 101Probabilistic harmonization and annotation of single-cell transcriptomics data with deep generative models.As the number of single-cell transcriptomics datasets grows, the natural next step is to integrate the accumulating data to achieve a common ontology of cell types and states. However, it is not straightforward to compare gene expression levels across datasets and to automatically assign cell type labels in a new dataset based on existing annotations. In this manuscript, we demonstrate that our previously developed method, scVI, provides an effective and fully probabilistic approach for joint representation and analysis of scRNA-seq data, while accounting for uncertainty caused by biological and measurement noise. We also introduce single-cell ANnotation using Variational Inference (scANVI), a semi-supervised variant of scVI designed to leverage existing cell state annotations. We demonstrate that scVI and scANVI compare favorably to state-of-the-art methods for data integration and cell state annotation in terms of accuracy, scalability, and adaptability to challenging settings. In co
- 102Single-cell multi-omics analysis of the immune response in COVID-19.Analysis of human blood immune cells provides insights into the coordinated response to viral infections such as severe acute respiratory syndrome coronavirus 2, which causes coronavirus disease 2019 (COVID-19). We performed single-cell transcriptome, surface proteome and T and B lymphocyte antigen receptor analyses of over 780,000 peripheral blood mononuclear cells from a cross-sectional cohort of 130 patients with varying severities of COVID-19. We identified expansion of nonclassical monocytes expressing complement transcripts (CD16 + C1QA/B/C + ) that sequester platelets and were predicted to replenish the alveolar macrophage pool in COVID-19. Early, uncommitted CD34 + hematopoietic stem/progenitor cells were primed toward megakaryopoiesis, accompanied by expanded megakaryocyte-committed progenitors and increased platelet activation. Clonally expanded CD8 + T cells and an increased ratio of CD8 + effector T cells to effector memory T cells characterized severe disease, while circul
- 103Single-cell and spatial analysis reveal interaction of FAP + fibroblasts and SPP1 + macrophages in colorectal cancer.Colorectal cancer (CRC) is among the most common malignancies with limited treatments other than surgery. The tumor microenvironment (TME) profiling enables the discovery of potential therapeutic targets. Here, we profile 54,103 cells from tumor and adjacent tissues to characterize cellular composition and elucidate the potential origin and regulation of tumor-enriched cell types in CRC. We demonstrate that the tumor-specific FAP + fibroblasts and SPP1 + macrophages were positively correlated in 14 independent CRC cohorts containing 2550 samples and validate their close localization by immuno-fluorescent staining and spatial transcriptomics. This interaction might be regulated by chemerin, TGF-β, and interleukin-1, which would stimulate the formation of immune-excluded desmoplasic structure and limit the T cell infiltration. Furthermore, we find patients with high FAP or SPP1 expression achieved less therapeutic benefit from an anti-PD-L1 therapy cohort. Our results provide a potential
- 104Comparison and evaluation of statistical error models for scRNA-seq.Background Heterogeneity in single-cell RNA-seq (scRNA-seq) data is driven by multiple sources, including biological variation in cellular state as well as technical variation introduced during experimental processing. Deconvolving these effects is a key challenge for preprocessing workflows. Recent work has demonstrated the importance and utility of count models for scRNA-seq analysis, but there is a lack of consensus on which statistical distributions and parameter settings are appropriate. Results Here, we analyze 59 scRNA-seq datasets that span a wide range of technologies, systems, and sequencing depths in order to evaluate the performance of different error models. We find that while a Poisson error model appears appropriate for sparse datasets, we observe clear evidence of overdispersion for genes with sufficient sequencing depth in all biological systems, necessitating the use of a negative binomial model. Moreover, we find that the degree of overdispersion varies widely across
- 105hdWGCNA identifies co-expression networks in high-dimensional transcriptomics data.Biological systems are immensely complex, organized into a multi-scale hierarchy of functional units based on tightly regulated interactions between distinct molecules, cells, organs, and organisms. While experimental methods enable transcriptome-wide measurements across millions of cells, popular bioinformatic tools do not support systems-level analysis. Here we present hdWGCNA, a comprehensive framework for analyzing co-expression networks in high-dimensional transcriptomics data such as single-cell and spatial RNA sequencing (RNA-seq). hdWGCNA provides functions for network inference, gene module identification, gene enrichment analysis, statistical tests, and data visualization. Beyond conventional single-cell RNA-seq, hdWGCNA is capable of performing isoform-level network analysis using long-read single-cell data. We showcase hdWGCNA using data from autism spectrum disorder and Alzheimer's disease brain samples, identifying disease-relevant co-expression network modules. hdWGCNA i
- 106scMC learns biological variation through the alignment of multiple single-cell genomics datasets.Distinguishing biological from technical variation is crucial when integrating and comparing single-cell genomics datasets across different experiments. Existing methods lack the capability in explicitly distinguishing these two variations, often leading to the removal of both variations. Here, we present an integration method scMC to remove the technical variation while preserving the intrinsic biological variation. scMC learns biological variation via variance analysis to subtract technical variation inferred in an unsupervised manner. Application of scMC to both simulated and real datasets from single-cell RNA-seq and ATAC-seq experiments demonstrates its capability of detecting context-shared and context-specific biological signals via accurate alignment.
- 107Redefining Tumor-Associated Macrophage Subpopulations and Functions in the Tumor Microenvironment.The immunosuppressive status of the tumor microenvironment (TME) remains poorly defined due to a lack of understanding regarding the function of tumor-associated macrophages (TAMs), which are abundant in the TME. TAMs are crucial drivers of tumor progression, metastasis, and resistance to therapy. Intra- and inter-tumoral spatial heterogeneities are potential keys to understanding the relationships between subpopulations of TAMs and their functions. Antitumor M1-like and pro-tumor M2-like TAMs coexist within tumors, and the opposing effects of these M1/M2 subpopulations on tumors directly impact current strategies to improve antitumor immune responses. Recent studies have found significant differences among monocytes or macrophages from distinct tumors, and other investigations have explored the existence of diverse TAM subsets at the molecular level. In this review, we discuss emerging evidence highlighting the redefinition of TAM subpopulations and functions in the TME and the possib
- 108Cell type and gene expression deconvolution with BayesPrism enables Bayesian integrative analysis across bulk and single-cell RNA sequencing in oncology.Inferring single-cell compositions and their contributions to global gene expression changes from bulk RNA sequencing (RNA-seq) datasets is a major challenge in oncology. Here we develop Bayesian cell proportion reconstruction inferred using statistical marginalization (BayesPrism), a Bayesian method to predict cellular composition and gene expression in individual cell types from bulk RNA-seq, using patient-derived, scRNA-seq as prior information. We conduct integrative analyses in primary glioblastoma, head and neck squamous cell carcinoma and skin cutaneous melanoma to correlate cell type composition with clinical outcomes across tumor types, and explore spatial heterogeneity in malignant and nonmalignant cell states. We refine current cancer subtypes using gene expression annotation after exclusion of confounding nonmalignant cells. Finally, we identify genes whose expression in malignant cells correlates with macrophage infiltration, T cells, fibroblasts and endothelial cells acro
- 109IonQuant Enables Accurate and Sensitive Label-Free Quantification With FDR-Controlled Match-Between-Runs.Missing values weaken the power of label-free quantitative proteomic experiments to uncover true quantitative differences between biological samples or experimental conditions. Match-between-runs (MBR) has become a common approach to mitigate the missing value problem, where peptides identified by tandem mass spectra in one run are transferred to another by inference based on m/z, charge state, retention time, and ion mobility when applicable. Though tolerances are used to ensure such transferred identifications are reasonably located and meet certain quality thresholds, little work has been done to evaluate the statistical confidence of MBR. Here, we present a mixture model-based approach to estimate the false discovery rate (FDR) of peptide and protein identification transfer, which we implement in the label-free quantification tool IonQuant. Using several benchmarking datasets generated on both Orbitrap and timsTOF mass spectrometers, we demonstrate superior performance of IonQuant
- 110A general and flexible method for signal extraction from single-cell RNA-seq data.Single-cell RNA-sequencing (scRNA-seq) is a powerful high-throughput technique that enables researchers to measure genome-wide transcription levels at the resolution of single cells. Because of the low amount of RNA present in a single cell, some genes may fail to be detected even though they are expressed; these genes are usually referred to as dropouts. Here, we present a general and flexible zero-inflated negative binomial model (ZINB-WaVE), which leads to low-dimensional representations of the data that account for zero inflation (dropouts), over-dispersion, and the count nature of the data. We demonstrate, with simulated and real data, that the model and its associated estimation procedure are able to give a more stable and accurate low-dimensional representation of the data than principal component analysis (PCA) and zero-inflated factor analysis (ZIFA), without the need for a preliminary normalization step.
- 111From GWAS to Function: Using Functional Genomics to Identify the Mechanisms Underlying Complex Diseases.Genome-wide association studies (GWAS) have successfully mapped thousands of loci associated with complex traits. These associations could reveal the molecular mechanisms altered in common complex diseases and result in the identification of novel drug targets. However, GWAS have also left a number of outstanding questions. In particular, the majority of disease-associated loci lie in non-coding regions of the genome and, even though they are thought to play a role in gene expression regulation, it is unclear which genes they regulate and in which cell types or physiological contexts this regulation occurs. This has hindered the translation of GWAS findings into clinical interventions. In this review we summarize how these challenges have been addressed over the last decade, with a particular focus on the integration of GWAS results with functional genomics datasets. Firstly, we investigate how the tissues and cell types involved in diseases can be identified using methods that test fo
- 112Transcriptional programs of neoantigen-specific TIL in anti-PD-1-treated lung cancers.PD-1 blockade unleashes CD8 T cells 1 , including those specific for mutation-associated neoantigens (MANA), but factors in the tumour microenvironment can inhibit these T cell responses. Single-cell transcriptomics have revealed global T cell dysfunction programs in tumour-infiltrating lymphocytes (TIL). However, the majority of TIL do not recognize tumour antigens 2 , and little is known about transcriptional programs of MANA-specific TIL. Here, we identify MANA-specific T cell clones using the MANA functional expansion of specific T cells assay 3 in neoadjuvant anti-PD-1-treated non-small cell lung cancers (NSCLC). We use their T cell receptors as a 'barcode' to track and analyse their transcriptional programs in the tumour microenvironment using coupled single-cell RNA sequencing and T cell receptor sequencing. We find both MANA- and virus-specific clones in TIL, regardless of response, and MANA-, influenza- and Epstein-Barr virus-specific TIL each have unique transcriptional progr
- 113Ultra-high sensitivity mass spectrometry quantifies single-cell proteome changes upon perturbation.Single-cell technologies are revolutionizing biology but are today mainly limited to imaging and deep sequencing. However, proteins are the main drivers of cellular function and in-depth characterization of individual cells by mass spectrometry (MS)-based proteomics would thus be highly valuable and complementary. Here, we develop a robust workflow combining miniaturized sample preparation, very low flow-rate chromatography, and a novel trapped ion mobility mass spectrometer, resulting in a more than 10-fold improved sensitivity. We precisely and robustly quantify proteomes and their changes in single, FACS-isolated cells. Arresting cells at defined stages of the cell cycle by drug treatment retrieves expected key regulators. Furthermore, it highlights potential novel ones and allows cell phase prediction. Comparing the variability in more than 430 single-cell proteomes to transcriptome data revealed a stable-core proteome despite perturbation, while the transcriptome appears stochasti
- 114Single-cell sequencing of human midbrain reveals glial activation and a Parkinson-specific neuronal state.Idiopathic Parkinson's disease is characterized by a progressive loss of dopaminergic neurons, but the exact disease aetiology remains largely unknown. To date, Parkinson's disease research has mainly focused on nigral dopaminergic neurons, although recent studies suggest disease-related changes also in non-neuronal cells and in midbrain regions beyond the substantia nigra. While there is some evidence for glial involvement in Parkinson's disease, the molecular mechanisms remain poorly understood. The aim of this study was to characterize the contribution of all cell types of the midbrain to Parkinson's disease pathology by single-nuclei RNA sequencing and to assess the cell type-specific risk for Parkinson's disease using the latest genome-wide association study. We profiled >41 000 single-nuclei transcriptomes of post-mortem midbrain from six idiopathic Parkinson's disease patients and five age-/sex-matched controls. To validate our findings in a spatial context, we utilized immunola
- 115METABOLIC: high-throughput profiling of microbial genomes for functional traits, metabolism, biogeochemistry, and community-scale functional networks.Background Advances in microbiome science are being driven in large part due to our ability to study and infer microbial ecology from genomes reconstructed from mixed microbial communities using metagenomics and single-cell genomics. Such omics-based techniques allow us to read genomic blueprints of microorganisms, decipher their functional capacities and activities, and reconstruct their roles in biogeochemical processes. Currently available tools for analyses of genomic data can annotate and depict metabolic functions to some extent; however, no standardized approaches are currently available for the comprehensive characterization of metabolic predictions, metabolite exchanges, microbial interactions, and microbial contributions to biogeochemical cycling. Results We present METABOLIC (METabolic And BiogeOchemistry anaLyses In miCrobes), a scalable software to advance microbial ecology and biogeochemistry studies using genomes at the resolution of individual organisms and/or microbial
- 116Deep Visual Proteomics defines single-cell identity and heterogeneity.Despite the availabilty of imaging-based and mass-spectrometry-based methods for spatial proteomics, a key challenge remains connecting images with single-cell-resolution protein abundance measurements. Here, we introduce Deep Visual Proteomics (DVP), which combines artificial-intelligence-driven image analysis of cellular phenotypes with automated single-cell or single-nucleus laser microdissection and ultra-high-sensitivity mass spectrometry. DVP links protein abundance to complex cellular or subcellular phenotypes while preserving spatial context. By individually excising nuclei from cell culture, we classified distinct cell states with proteomic profiles defined by known and uncharacterized proteins. In an archived primary melanoma tissue, DVP identified spatially resolved proteome changes as normal melanocytes transition to fully invasive melanoma, revealing pathways that change in a spatial manner as cancer progresses, such as mRNA splicing dysregulation in metastatic vertical gr
- 117Spatiotemporal analysis of human intestinal development at single-cell resolution.Development of the human intestine is not well understood. Here, we link single-cell RNA sequencing and spatial transcriptomics to characterize intestinal morphogenesis through time. We identify 101 cell states including epithelial and mesenchymal progenitor populations and programs linked to key morphogenetic milestones. We describe principles of crypt-villus axis formation; neural, vascular, mesenchymal morphogenesis, and immune population of the developing gut. We identify the differentiation hierarchies of developing fibroblast and myofibroblast subtypes and describe diverse functions for these including as vascular niche cells. We pinpoint the origins of Peyer's patches and gut-associated lymphoid tissue (GALT) and describe location-specific immune programs. We use our resource to present an unbiased analysis of morphogen gradients that direct sequential waves of cellular differentiation and define cells and locations linked to rare developmental intestinal disorders. We compile a
- 118Emerging insights of tumor heterogeneity and drug resistance mechanisms in lung cancer targeted therapy.The biggest hurdle to targeted cancer therapy is the inevitable emergence of drug resistance. Tumor cells employ different mechanisms to resist the targeting agent. Most commonly in EGFR-mutant non-small cell lung cancer, secondary resistance mutations on the target kinase domain emerge to diminish the binding affinity of first- and second-generation inhibitors. Other alternative resistance mechanisms include activating complementary bypass pathways and phenotypic transformation. Sequential monotherapies promise to temporarily address the problem of acquired drug resistance, but evidently are limited by the tumor cells' ability to adapt and evolve new resistance mechanisms to persist in the drug environment. Recent studies have nominated a model of drug resistance and tumor progression under targeted therapy as a result of a small subpopulation of cells being able to endure the drug (minimal residual disease cells) and eventually develop further mutations that allow them to regrow and
- 119Multi-omics single-cell data integration and regulatory inference with graph-linked embedding.Despite the emergence of experimental methods for simultaneous measurement of multiple omics modalities in single cells, most single-cell datasets include only one modality. A major obstacle in integrating omics data from multiple modalities is that different omics layers typically have distinct feature spaces. Here, we propose a computational framework called GLUE (graph-linked unified embedding), which bridges the gap by modeling regulatory interactions across omics layers explicitly. Systematic benchmarking demonstrated that GLUE is more accurate, robust and scalable than state-of-the-art tools for heterogeneous single-cell multi-omics data. We applied GLUE to various challenging tasks, including triple-omics integration, integrative regulatory inference and multi-omics human cell atlas construction over millions of cells, where GLUE was able to correct previous annotations. GLUE features a modular design that can be flexibly extended and enhanced for new analysis tasks. The full pa
- 120Single cell transcriptional and chromatin accessibility profiling redefine cellular heterogeneity in the adult human kidney.The integration of single cell transcriptome and chromatin accessibility datasets enables a deeper understanding of cell heterogeneity. We performed single nucleus ATAC (snATAC-seq) and RNA (snRNA-seq) sequencing to generate paired, cell-type-specific chromatin accessibility and transcriptional profiles of the adult human kidney. We demonstrate that snATAC-seq is comparable to snRNA-seq in the assignment of cell identity and can further refine our understanding of functional heterogeneity in the nephron. The majority of differentially accessible chromatin regions are localized to promoters and a significant proportion are closely associated with differentially expressed genes. Cell-type-specific enrichment of transcription factor binding motifs implicates the activation of NF-κB that promotes VCAM1 expression and drives transition between a subpopulation of proximal tubule epithelial cells. Our multi-omics approach improves the ability to detect unique cell states within the kidney and
- 121Single-cell proteomic and transcriptomic analysis of macrophage heterogeneity using SCoPE2.Background Macrophages are innate immune cells with diverse functional and molecular phenotypes. This diversity is largely unexplored at the level of single-cell proteomes because of the limitations of quantitative single-cell protein analysis. Results To overcome this limitation, we develop SCoPE2, which substantially increases quantitative accuracy and throughput while lowering cost and hands-on time by introducing automated and miniaturized sample preparation. These advances enable us to analyze the emergence of cellular heterogeneity as homogeneous monocytes differentiate into macrophage-like cells in the absence of polarizing cytokines. SCoPE2 quantifies over 3042 proteins in 1490 single monocytes and macrophages in 10 days of instrument time, and the quantified proteins allow us to discern single cells by cell type. Furthermore, the data uncover a continuous gradient of proteome states for the macrophages, suggesting that macrophage heterogeneity may emerge in the absence of pola
- 122Single-cell analysis of human glioma and immune cells identifies S100A4 as an immunotherapy target.A major rate-limiting step in developing more effective immunotherapies for GBM is our inadequate understanding of the cellular complexity and the molecular heterogeneity of immune infiltrates in gliomas. Here, we report an integrated analysis of 201,986 human glioma, immune, and other stromal cells at the single cell level. In doing so, we discover extensive spatial and molecular heterogeneity in immune infiltrates. We identify molecular signatures for nine distinct myeloid cell subtypes, of which five are independent prognostic indicators of glioma patient survival. Furthermore, we identify S100A4 as a regulator of immune suppressive T and myeloid cells in GBM and demonstrate that deleting S100a4 in non-cancer cells is sufficient to reprogram the immune landscape and significantly improve survival. This study provides insights into spatial, molecular, and functional heterogeneity of glioma and glioma-associated immune cells and demonstrates the utility of this dataset for discovering
- 123Single-cell genomic profiling of human dopamine neurons identifies a population that selectively degenerates in Parkinson's disease.The loss of dopamine (DA) neurons within the substantia nigra pars compacta (SNpc) is a defining pathological hallmark of Parkinson's disease (PD). Nevertheless, the molecular features associated with DA neuron vulnerability have not yet been fully identified. Here, we developed a protocol to enrich and transcriptionally profile DA neurons from patients with PD and matched controls, sampling a total of 387,483 nuclei, including 22,048 DA neuron profiles. We identified ten populations and spatially localized each within the SNpc using Slide-seq. A single subtype, marked by the expression of the gene AGTR1 and spatially confined to the ventral tier of SNpc, was highly susceptible to loss in PD and showed the strongest upregulation of targets of TP53 and NR2F2, nominating molecular processes associated with degeneration. This same vulnerable population was specifically enriched for the heritable risk associated with PD, highlighting the importance of cell-intrinsic processes in determinin
- 124Pan-cancer single-cell analysis reveals the heterogeneity and plasticity of cancer-associated fibroblasts in the tumor microenvironment.Cancer-associated fibroblasts (CAFs) are the predominant components of the tumor microenvironment (TME) and influence cancer hallmarks, but without systematic investigation on their ubiquitous characteristics across different cancer types. Here, we perform pan-cancer analysis on 226 samples across 10 solid cancer types to profile the TME at single-cell resolution, illustrating the commonalities/plasticity of heterogenous CAFs. Activation trajectory of the major CAF types is divided into three states, exhibiting distinct interactions with other cell components, and relating to prognosis of immunotherapy. Moreover, minor CAF components represent the alternative origin from other TME components (e.g., endothelia and macrophages). Particularly, the ubiquitous presentation of endothelial-to-mesenchymal transition CAF, which may interact with proximal SPP1 + tumor-associated macrophages, is implicated in endothelial-to-mesenchymal transition and survival stratifications. Our study comprehens
- 125Single-cell RNA-seq reveals fibroblast heterogeneity and increased mesenchymal fibroblasts in human fibrotic skin diseases.Fibrotic skin disease represents a major global healthcare burden, characterized by fibroblast hyperproliferation and excessive accumulation of extracellular matrix. Fibroblasts are found to be heterogeneous in multiple fibrotic diseases, but fibroblast heterogeneity in fibrotic skin diseases is not well characterized. In this study, we explore fibroblast heterogeneity in keloid, a paradigm of fibrotic skin diseases, by using single-cell RNA-seq. Our results indicate that keloid fibroblasts can be divided into 4 subpopulations: secretory-papillary, secretory-reticular, mesenchymal and pro-inflammatory. Interestingly, the percentage of mesenchymal fibroblast subpopulation is significantly increased in keloid compared to normal scar. Functional studies indicate that mesenchymal fibroblasts are crucial for collagen overexpression in keloid. Increased mesenchymal fibroblast subpopulation is also found in another fibrotic skin disease, scleroderma, suggesting this is a broad mechanism for s
- 126Mapping transcriptomic vector fields of single cells.Single-cell (sc)RNA-seq, together with RNA velocity and metabolic labeling, reveals cellular states and transitions at unprecedented resolution. Fully exploiting these data, however, requires kinetic models capable of unveiling governing regulatory functions. Here, we introduce an analytical framework dynamo (https://github.com/aristoteleo/dynamo-release), which infers absolute RNA velocity, reconstructs continuous vector fields that predict cell fates, employs differential geometry to extract underlying regulations, and ultimately predicts optimal reprogramming paths and perturbation outcomes. We highlight dynamo's power to overcome fundamental limitations of conventional splicing-based RNA velocity analyses to enable accurate velocity estimations on a metabolically labeled human hematopoiesis scRNA-seq dataset. Furthermore, differential geometry analyses reveal mechanisms driving early megakaryocyte appearance and elucidate asymmetrical regulation within the PU.1-GATA1 circuit. Lever
- 127Molecular and spatial signatures of mouse brain aging at single-cell resolution.The diversity and complex organization of cells in the brain have hindered systematic characterization of age-related changes in its cellular and molecular architecture, limiting our ability to understand the mechanisms underlying its functional decline during aging. Here, we generated a high-resolution cell atlas of brain aging within the frontal cortex and striatum using spatially resolved single-cell transcriptomics and quantified changes in gene expression and spatial organization of major cell types in these regions over the mouse lifespan. We observed substantially more pronounced changes in cell state, gene expression, and spatial organization of non-neuronal cells over neurons. Our data revealed molecular and spatial signatures of glial and immune cell activation during aging, particularly enriched in the subcortical white matter, and identified both similarities and notable differences in cell-activation patterns induced by aging and systemic inflammatory challenge. These resu
- 128Single-cell transcriptomics reveals cell-type-specific diversification in human heart failure.Heart failure represents a major cause of morbidity and mortality worldwide. Single-cell transcriptomics have revolutionized our understanding of cell composition and associated gene expression. Through integrated analysis of single-cell and single-nucleus RNA-sequencing data generated from 27 healthy donors and 18 individuals with dilated cardiomyopathy, here we define the cell composition of the healthy and failing human heart. We identify cell-specific transcriptional signatures associated with age and heart failure and reveal the emergence of disease-associated cell states. Notably, cardiomyocytes converge toward common disease-associated cell states, whereas fibroblasts and myeloid cells undergo dramatic diversification. Endothelial cells and pericytes display global transcriptional shifts without changes in cell complexity. Collectively, our findings provide a comprehensive analysis of the cellular and transcriptomic landscape of human heart failure, identify cell type-specific t
- 129scCODA is a Bayesian model for compositional single-cell data analysis.Compositional changes of cell types are main drivers of biological processes. Their detection through single-cell experiments is difficult due to the compositionality of the data and low sample sizes. We introduce scCODA ( https://github.com/theislab/scCODA ), a Bayesian model addressing these issues enabling the study of complex cell type effects in disease, and other stimuli. scCODA demonstrated excellent detection performance, while reliably controlling for false discoveries, and identified experimentally verified cell type changes that were missed in original analyses.
- 130Single-cell RNA sequencing reveals functional heterogeneity of glioma-associated brain macrophages.Microglia are resident myeloid cells in the central nervous system (CNS) that control homeostasis and protect CNS from damage and infections. Microglia and peripheral myeloid cells accumulate and adapt tumor supporting roles in human glioblastomas that show prevalence in men. Cell heterogeneity and functional phenotypes of myeloid subpopulations in gliomas remain elusive. Here we show single-cell RNA sequencing (scRNA-seq) of CD11b + myeloid cells in naïve and GL261 glioma-bearing mice that reveal distinct profiles of microglia, infiltrating monocytes/macrophages and CNS border-associated macrophages. We demonstrate an unforeseen molecular heterogeneity among myeloid cells in naïve and glioma-bearing brains, validate selected marker proteins and show distinct spatial distribution of identified subsets in experimental gliomas. We find higher expression of MHCII encoding genes in glioma-activated male microglia, which was corroborated in bulk and scRNA-seq data from human diffuse gliomas
- 131Single-cell multiomics: technologies and data analysis methods.Advances in single-cell isolation and barcoding technologies offer unprecedented opportunities to profile DNA, mRNA, and proteins at a single-cell resolution. Recently, bulk multiomics analyses, such as multidimensional genomic and proteogenomic analyses, have proven beneficial for obtaining a comprehensive understanding of cellular events. This benefit has facilitated the development of single-cell multiomics analysis, which enables cell type-specific gene regulation to be examined. The cardinal features of single-cell multiomics analysis include (1) technologies for single-cell isolation, barcoding, and sequencing to measure multiple types of molecules from individual cells and (2) the integrative analysis of molecules to characterize cell types and their functions regarding pathophysiological processes based on molecular signatures. Here, we summarize the technologies for single-cell multiomics analyses (mRNA-genome, mRNA-DNA methylation, mRNA-chromatin accessibility, and mRNA-prote
- 132Morphological diversity of single neurons in molecularly defined cell types.Dendritic and axonal morphology reflects the input and output of neurons and is a defining feature of neuronal types 1,2 , yet our knowledge of its diversity remains limited. Here, to systematically examine complete single-neuron morphologies on a brain-wide scale, we established a pipeline encompassing sparse labelling, whole-brain imaging, reconstruction, registration and analysis. We fully reconstructed 1,741 neurons from cortex, claustrum, thalamus, striatum and other brain regions in mice. We identified 11 major projection neuron types with distinct morphological features and corresponding transcriptomic identities. Extensive projectional diversity was found within each of these major types, on the basis of which some types were clustered into more refined subtypes. This diversity follows a set of generalizable principles that govern long-range axonal projections at different levels, including molecular correspondence, divergent or convergent projection, axon termination pattern,
- 133Deconstruction of rheumatoid arthritis synovium defines inflammatory subtypes.Rheumatoid arthritis is a prototypical autoimmune disease that causes joint inflammation and destruction 1 . There is currently no cure for rheumatoid arthritis, and the effectiveness of treatments varies across patients, suggesting an undefined pathogenic diversity 1,2 . Here, to deconstruct the cell states and pathways that characterize this pathogenic heterogeneity, we profiled the full spectrum of cells in inflamed synovium from patients with rheumatoid arthritis. We used multi-modal single-cell RNA-sequencing and surface protein data coupled with histology of synovial tissue from 79 donors to build single-cell atlas of rheumatoid arthritis synovial tissue that includes more than 314,000 cells. We stratified tissues into six groups, referred to as cell-type abundance phenotypes (CTAPs), each characterized by selectively enriched cell states. These CTAPs demonstrate the diversity of synovial inflammation in rheumatoid arthritis, ranging from samples enriched for T and B cells to tho
- 134Comprehensive analysis of single cell ATAC-seq data with SnapATAC.Identification of the cis-regulatory elements controlling cell-type specific gene expression patterns is essential for understanding the origin of cellular diversity. Conventional assays to map regulatory elements via open chromatin analysis of primary tissues is hindered by sample heterogeneity. Single cell analysis of accessible chromatin (scATAC-seq) can overcome this limitation. However, the high-level noise of each single cell profile and the large volume of data pose unique computational challenges. Here, we introduce SnapATAC, a software package for analyzing scATAC-seq datasets. SnapATAC dissects cellular heterogeneity in an unbiased manner and map the trajectories of cellular states. Using the Nyström method, SnapATAC can process data from up to a million cells. Furthermore, SnapATAC incorporates existing tools into a comprehensive package for analyzing single cell ATAC-seq dataset. As demonstration of its utility, SnapATAC is applied to 55,592 single-nucleus ATAC-seq profiles
- 135Single-cell roadmap of human gonadal development.Gonadal development is a complex process that involves sex determination followed by divergent maturation into either testes or ovaries 1 . Historically, limited tissue accessibility, a lack of reliable in vitro models and critical differences between humans and mice have hampered our knowledge of human gonadogenesis, despite its importance in gonadal conditions and infertility. Here, we generated a comprehensive map of first- and second-trimester human gonads using a combination of single-cell and spatial transcriptomics, chromatin accessibility assays and fluorescent microscopy. We extracted human-specific regulatory programmes that control the development of germline and somatic cell lineages by profiling equivalent developmental stages in mice. In both species, we define the somatic cell states present at the time of sex specification, including the bipotent early supporting population that, in males, upregulates the testis-determining factor SRY and sPAX8s, a gonadal lineage locat
- 136Dissection of artifactual and confounding glial signatures by single-cell sequencing of mouse and human brain.A key aspect of nearly all single-cell sequencing experiments is dissociation of intact tissues into single-cell suspensions. While many protocols have been optimized for optimal cell yield, they have often overlooked the effects that dissociation can have on ex vivo gene expression. Here, we demonstrate that use of enzymatic dissociation on brain tissue induces an aberrant ex vivo gene expression signature, most prominently in microglia, which is prevalent in published literature and can substantially confound downstream analyses. To address this issue, we present a rigorously validated protocol that preserves both in vivo transcriptional profiles and cell-type diversity and yield across tissue types and species. We also identify a similar signature in postmortem human brain single-nucleus RNA-sequencing datasets, and show that this signature is induced in freshly isolated human tissue by exposure to elevated temperatures ex vivo. Together, our results provide a methodological solutio
- 137Single cell transcriptomic landscape of diabetic foot ulcers.Diabetic foot ulceration (DFU) is a devastating complication of diabetes whose pathogenesis remains incompletely understood. Here, we profile 174,962 single cells from the foot, forearm, and peripheral blood mononuclear cells using single-cell RNA sequencing. Our analysis shows enrichment of a unique population of fibroblasts overexpressing MMP1, MMP3, MMP11, HIF1A, CHI3L1, and TNFAIP6 and increased M1 macrophage polarization in the DFU patients with healing wounds. Further, analysis of spatially separated samples from the same patient and spatial transcriptomics reveal preferential localization of these healing associated fibroblasts toward the wound bed as compared to the wound edge or unwounded skin. Spatial transcriptomics also validates our findings of higher abundance of M1 macrophages in healers and M2 macrophages in non-healers. Our analysis provides deep insights into the wound healing microenvironment, identifying cell types that could be critical in promoting DFU healing, an
- 138Spatially resolved multiomics of human cardiac niches.The function of a cell is defined by its intrinsic characteristics and its niche: the tissue microenvironment in which it dwells. Here we combine single-cell and spatial transcriptomics data to discover cellular niches within eight regions of the human heart. We map cells to microanatomical locations and integrate knowledge-based and unsupervised structural annotations. We also profile the cells of the human cardiac conduction system 1 . The results revealed their distinctive repertoire of ion channels, G-protein-coupled receptors (GPCRs) and regulatory networks, and implicated FOXP2 in the pacemaker phenotype. We show that the sinoatrial node is compartmentalized, with a core of pacemaker cells, fibroblasts and glial cells supporting glutamatergic signalling. Using a custom CellPhoneDB.org module, we identify trans-synaptic pacemaker cell interactions with glia. We introduce a druggable target prediction tool, drug2cell, which leverages single-cell profiles and drug-target interaction
- 139Design and computational analysis of single-cell RNA-sequencing experiments.Single-cell RNA-sequencing (scRNA-seq) has emerged as a revolutionary tool that allows us to address scientific questions that eluded examination just a few years ago. With the advantages of scRNA-seq come computational challenges that are just beginning to be addressed. In this article, we highlight the computational methods available for the design and analysis of scRNA-seq experiments, their advantages and disadvantages in various settings, the open questions for which novel methods are needed, and expected future developments in this exciting area.
- 140Astaxanthin-Producing Green Microalga Haematococcus pluvialis: From Single Cell to High Value Commercial Products.Many species of microalgae have been used as source of nutrient rich food, feed, and health promoting compounds. Among the commercially important microalgae, Haematococcus pluvialis is the richest source of natural astaxanthin which is considered as "super anti-oxidant." Natural astaxanthin produced by H. pluvialis has significantly greater antioxidant capacity than the synthetic one. Astaxanthin has important applications in the nutraceuticals, cosmetics, food, and aquaculture industries. It is now evident that, astaxanthin can significantly reduce free radicals and oxidative stress and help human body maintain a healthy state. With extraordinary potency and increase in demand, astaxanthin is one of the high-value microalgal products of the future.This comprehensive review summarizes the most important aspects of the biology, biochemical composition, biosynthesis, and astaxanthin accumulation in the cells of H. pluvialis and its wide range of applications for humans and animals. In th
- 141Mapping the immune environment in clear cell renal carcinoma by single-cell genomics.Clear cell renal cell carcinoma (ccRCC) is one of the most immunologically distinct tumor types due to high response rate to immunotherapies, despite low tumor mutational burden. To characterize the tumor immune microenvironment of ccRCC, we applied single-cell-RNA sequencing (SCRS) along with T-cell-receptor (TCR) sequencing to map the transcriptomic heterogeneity of 25,688 individual CD45 + lymphoid and myeloid cells in matched tumor and blood from three patients with ccRCC. We also included 11,367 immune cells from four other individuals derived from the kidney and peripheral blood to facilitate the identification and assessment of ccRCC-specific differences. There is an overall increase in CD8 + T-cell and macrophage populations in tumor-infiltrated immune cells compared to normal renal tissue. We further demonstrate the divergent cell transcriptional states for tumor-infiltrating CD8 + T cells and identify a MKI67 + proliferative subpopulation being a potential culprit for the pro
- 142Treg Heterogeneity, Function, and Homeostasis.T-regulatory cells (Tregs) represent a unique subpopulation of helper T-cells by maintaining immune equilibrium using various mechanisms. The role of T-cell receptors (TCR) in providing homeostasis and activation of conventional T-cells is well-known; however, for Tregs, this area is understudied. In the last two decades, evidence has accumulated to confirm the importance of the TCR in Treg homeostasis and antigen-specific immune response regulation. In this review, we describe the current view of Treg subset heterogeneity, homeostasis and function in the context of TCR involvement. Recent studies of the TCR repertoire of Tregs, combined with single-cell gene expression analysis, revealed the importance of TCR specificity in shaping Treg phenotype diversity, their functions and homeostatic maintenance in various tissues. We propose that Tregs, like conventional T-helper cells, act to a great extent in an antigen-specific manner, which is provided by a specific distribution of Tregs in
- 143zUMIs - A fast and flexible pipeline to process RNA sequencing data with UMIs.Background Single-cell RNA-sequencing (scRNA-seq) experiments typically analyze hundreds or thousands of cells after amplification of the cDNA. The high throughput is made possible by the early introduction of sample-specific bar codes (BCs), and the amplification bias is alleviated by unique molecular identifiers (UMIs). Thus, the ideal analysis pipeline for scRNA-seq data needs to efficiently tabulate reads according to both BC and UMI. Findings zUMIs is a pipeline that can handle both known and random BCs and also efficiently collapse UMIs, either just for exon mapping reads or for both exon and intron mapping reads. If BC annotation is missing, zUMIs can accurately detect intact cells from the distribution of sequencing reads. Another unique feature of zUMIs is the adaptive downsampling function that facilitates dealing with hugely varying library sizes but also allows the user to evaluate whether the library has been sequenced to saturation. To illustrate the utility of zUMIs, we
- 144Single-Cell Multiomics: Multiple Measurements from Single Cells.Single-cell sequencing provides information that is not confounded by genotypic or phenotypic heterogeneity of bulk samples. Sequencing of one molecular type (RNA, methylated DNA or open chromatin) in a single cell, furthermore, provides insights into the cell's phenotype and links to its genotype. Nevertheless, only by taking measurements of these phenotypes and genotypes from the same single cells can such inferences be made unambiguously. In this review, we survey the first experimental approaches that assay, in parallel, multiple molecular types from the same single cell, before considering the challenges and opportunities afforded by these and future technologies.
- 145Impaired local intrinsic immunity to SARS-CoV-2 infection in severe COVID-19.SARS-CoV-2 infection can cause severe respiratory COVID-19. However, many individuals present with isolated upper respiratory symptoms, suggesting potential to constrain viral pathology to the nasopharynx. Which cells SARS-CoV-2 primarily targets and how infection influences the respiratory epithelium remains incompletely understood. We performed scRNA-seq on nasopharyngeal swabs from 58 healthy and COVID-19 participants. During COVID-19, we observe expansion of secretory, loss of ciliated, and epithelial cell repopulation via deuterosomal cell expansion. In mild and moderate COVID-19, epithelial cells express anti-viral/interferon-responsive genes, while cells in severe COVID-19 have muted anti-viral responses despite equivalent viral loads. SARS-CoV-2 RNA + host-target cells are highly heterogenous, including developing ciliated, interferon-responsive ciliated, AZGP1 high goblet, and KRT13 + "hillock"-like cells, and we identify genes associated with susceptibility, resistance, or in
- 146Cancer-associated fibroblast classification in single-cell and spatial proteomics data.Cancer-associated fibroblasts (CAFs) are a diverse cell population within the tumour microenvironment, where they have critical effects on tumour evolution and patient prognosis. To define CAF phenotypes, we analyse a single-cell RNA sequencing (scRNA-seq) dataset of over 16,000 stromal cells from tumours of 14 breast cancer patients, based on which we define and functionally annotate nine CAF phenotypes and one class of pericytes. We validate this classification system in four additional cancer types and use highly multiplexed imaging mass cytometry on matched breast cancer samples to confirm our defined CAF phenotypes at the protein level and to analyse their spatial distribution within tumours. This general CAF classification scheme will allow comparison of CAF phenotypes across studies, facilitate analysis of their functional roles, and potentially guide development of new treatment strategies in the future.
- 147Organization of the human intestine at single-cell resolution.The intestine is a complex organ that promotes digestion, extracts nutrients, participates in immune surveillance, maintains critical symbiotic relationships with microbiota and affects overall health 1 . The intesting has a length of over nine metres, along which there are differences in structure and function 2 . The localization of individual cell types, cell type development trajectories and detailed cell transcriptional programs probably drive these differences in function. Here, to better understand these differences, we evaluated the organization of single cells using multiplexed imaging and single-nucleus RNA and open chromatin assays across eight different intestinal sites from nine donors. Through systematic analyses, we find cell compositions that differ substantially across regions of the intestine and demonstrate the complexity of epithelial subtypes, and find that the same cell types are organized into distinct neighbourhoods and communities, highlighting distinct immunol
- 148Integration of spatial and single-cell transcriptomic data elucidates mouse organogenesis.Molecular profiling of single cells has advanced our knowledge of the molecular basis of development. However, current approaches mostly rely on dissociating cells from tissues, thereby losing the crucial spatial context of regulatory processes. Here, we apply an image-based single-cell transcriptomics method, sequential fluorescence in situ hybridization (seqFISH), to detect mRNAs for 387 target genes in tissue sections of mouse embryos at the 8-12 somite stage. By integrating spatial context and multiplexed transcriptional measurements with two single-cell transcriptome atlases, we characterize cell types across the embryo and demonstrate that spatially resolved expression of genes not profiled by seqFISH can be imputed. We use this high-resolution spatial map to characterize fundamental steps in the patterning of the midbrain-hindbrain boundary (MHB) and the developing gut tube. We uncover axes of cell differentiation that are not apparent from single-cell RNA-sequencing (scRNA-seq)
- 149Quantitative single-cell proteomics as a tool to characterize cellular hierarchies.Large-scale single-cell analyses are of fundamental importance in order to capture biological heterogeneity within complex cell systems, but have largely been limited to RNA-based technologies. Here we present a comprehensive benchmarked experimental and computational workflow, which establishes global single-cell mass spectrometry-based proteomics as a tool for large-scale single-cell analyses. By exploiting a primary leukemia model system, we demonstrate both through pre-enrichment of cell populations and through a non-enriched unbiased approach that our workflow enables the exploration of cellular heterogeneity within this aberrant developmental hierarchy. Our approach is capable of consistently quantifying ~1000 proteins per cell across thousands of individual cells using limited instrument time. Furthermore, we develop a computational workflow (SCeptre) that effectively normalizes the data, integrates available FACS data and facilitates downstream analysis. The approach presented
- 150Single-cell spatial landscapes of the lung tumour immune microenvironment.Single-cell technologies have revealed the complexity of the tumour immune microenvironment with unparalleled resolution 1-9 . Most clinical strategies rely on histopathological stratification of tumour subtypes, yet the spatial context of single-cell phenotypes within these stratified subgroups is poorly understood. Here we apply imaging mass cytometry to characterize the tumour and immunological landscape of samples from 416 patients with lung adenocarcinoma across five histological patterns. We resolve more than 1.6 million cells, enabling spatial analysis of immune lineages and activation states with distinct clinical correlates, including survival. Using deep learning, we can predict with high accuracy those patients who will progress after surgery using a single 1-mm 2 tumour core, which could be informative for clinical management following surgical resection. Our dataset represents a valuable resource for the non-small cell lung cancer research community and exemplifies the uti
- 151SCODE: an efficient regulatory network inference algorithm from single-cell RNA-Seq during differentiation.Motivation The analysis of RNA-Seq data from individual differentiating cells enables us to reconstruct the differentiation process and the degree of differentiation (in pseudo-time) of each cell. Such analyses can reveal detailed expression dynamics and functional relationships for differentiation. To further elucidate differentiation processes, more insight into gene regulatory networks is required. The pseudo-time can be regarded as time information and, therefore, single-cell RNA-Seq data are time-course data with high time resolution. Although time-course data are useful for inferring networks, conventional inference algorithms for such data suffer from high time complexity when the number of samples and genes is large. Therefore, a novel algorithm is necessary to infer networks from single-cell RNA-Seq during differentiation. Results In this study, we developed the novel and efficient algorithm SCODE to infer regulatory networks, based on ordinary differential equations. We appli
- 152Single-cell spatial immune landscapes of primary and metastatic brain tumours.Single-cell technologies have enabled the characterization of the tumour microenvironment at unprecedented depth and have revealed vast cellular diversity among tumour cells and their niche. Anti-tumour immunity relies on cell-cell relationships within the tumour microenvironment 1,2 , yet many single-cell studies lack spatial context and rely on dissociated tissues 3 . Here we applied imaging mass cytometry to characterize the immunological landscape of 139 high-grade glioma and 46 brain metastasis tumours from patients. Single-cell analysis of more than 1.1 million cells across 389 high-dimensional histopathology images enabled the spatial resolution of immune lineages and activation states, revealing differences in immune landscapes between primary tumours and brain metastases from diverse solid cancers. These analyses revealed cellular neighbourhoods associated with survival in patients with glioblastoma, which we leveraged to identify a unique population of myeloperoxidase (MPO)-p
- 153Discriminating mild from critical COVID-19 by innate and adaptive immune single-cell profiling of bronchoalveolar lavages.How the innate and adaptive host immune system miscommunicate to worsen COVID-19 immunopathology has not been fully elucidated. Here, we perform single-cell deep-immune profiling of bronchoalveolar lavage (BAL) samples from 5 patients with mild and 26 with critical COVID-19 in comparison to BALs from non-COVID-19 pneumonia and normal lung. We use pseudotime inference to build T-cell and monocyte-to-macrophage trajectories and model gene expression changes along them. In mild COVID-19, CD8 + resident-memory (T RM ) and CD4 + T-helper-17 (T H17 ) cells undergo active (presumably antigen-driven) expansion towards the end of the trajectory, and are characterized by good effector functions, while in critical COVID-19 they remain more naïve. Vice versa, CD4 + T-cells with T-helper-1 characteristics (T H1 -like) and CD8 + T-cells expressing exhaustion markers (T EX -like) are enriched halfway their trajectories in mild COVID-19, where they also exhibit good effector functions, while in critic
- 154T-cell dysfunction in the glioblastoma microenvironment is mediated by myeloid cells releasing interleukin-10.Despite recent advances in cancer immunotherapy, certain tumor types, such as Glioblastomas, are highly resistant due to their tumor microenvironment disabling the anti-tumor immune response. Here we show, by applying an in-silico multidimensional model integrating spatially resolved and single-cell gene expression data of 45,615 immune cells from 12 tumor samples, that a subset of Interleukin-10-releasing HMOX1 + myeloid cells, spatially localizing to mesenchymal-like tumor regions, drive T-cell exhaustion and thus contribute to the immunosuppressive tumor microenvironment. These findings are validated using a human ex-vivo neocortical glioblastoma model inoculated with patient derived peripheral T-cells to simulate the immune compartment. This model recapitulates the dysfunctional transformation of tumor infiltrating T-cells. Inhibition of the JAK/STAT pathway rescues T-cell functionality both in our model and in-vivo, providing further evidence of IL-10 release being an important dr
- 155Single-cell epigenomics reveals mechanisms of human cortical development.During mammalian development, differences in chromatin state coincide with cellular differentiation and reflect changes in the gene regulatory landscape 1 . In the developing brain, cell fate specification and topographic identity are important for defining cell identity 2 and confer selective vulnerabilities to neurodevelopmental disorders 3 . Here, to identify cell-type-specific chromatin accessibility patterns in the developing human brain, we used a single-cell assay for transposase accessibility by sequencing (scATAC-seq) in primary tissue samples from the human forebrain. We applied unbiased analyses to identify genomic loci that undergo extensive cell-type- and brain-region-specific changes in accessibility during neurogenesis, and an integrative analysis to predict cell-type-specific candidate regulatory elements. We found that cerebral organoids recapitulate most putative cell-type-specific enhancer accessibility patterns but lack many cell-type-specific open chromatin regions
- 156Profiling genome-wide DNA methylation.DNA methylation is an epigenetic modification that plays an important role in regulating gene expression and therefore a broad range of biological processes and diseases. DNA methylation is tissue-specific, dynamic, sequence-context-dependent and trans-generationally heritable, and these complex patterns of methylation highlight the significance of profiling DNA methylation to answer biological questions. In this review, we surveyed major methylation assays, along with comparisons and biological examples, to provide an overview of DNA methylation profiling techniques. The advances in microarray and sequencing technologies make genome-wide profiling possible at a single-nucleotide or even a single-cell resolution. These profiling approaches vary in many aspects, such as DNA input, resolution, genomic region coverage, and bioinformatics analysis, and selecting a feasible method requires knowledge of these methods. We first introduce the biological background of DNA methylation and its pa
- 157Single-cell genomics and spatial transcriptomics: Discovery of novel cell states and cellular interactions in liver physiology and disease biology.Transcriptome analysis enables the study of gene expression in human tissues and is a valuable tool to characterise liver function and gene expression dynamics during liver disease, as well as to identify prognostic markers or signatures, and to facilitate discovery of new therapeutic targets. In contrast to whole tissue RNA sequencing analysis, single-cell RNA-sequencing (scRNA-seq) and spatial transcriptomics enables the study of transcriptional activity at the single cell or spatial level. ScRNA-seq has paved the way for the discovery of previously unknown cell types and subtypes in normal and diseased liver, facilitating the study of rare cells (such as liver progenitor cells) and the functional roles of non-parenchymal cells in chronic liver disease and cancer. By adding spatial information to scRNA-seq data, spatial transcriptomics has transformed our understanding of tissue functional organisation and cell-to-cell interactions in situ. These approaches have recently been applied
- 158Single-cell epigenomics: powerful new methods for understanding gene regulation and cell identity.Emerging single-cell epigenomic methods are being developed with the exciting potential to transform our knowledge of gene regulation. Here we review available techniques and future possibilities, arguing that the full potential of single-cell epigenetic studies will be realized through parallel profiling of genomic, transcriptional, and epigenetic information.
- 159Simultaneous trimodal single-cell measurement of transcripts, epitopes, and chromatin accessibility using TEA-seq.Single-cell measurements of cellular characteristics have been instrumental in understanding the heterogeneous pathways that drive differentiation, cellular responses to signals, and human disease. Recent advances have allowed paired capture of protein abundance and transcriptomic state, but a lack of epigenetic information in these assays has left a missing link to gene regulation. Using the heterogeneous mixture of cells in human peripheral blood as a test case, we developed a novel scATAC-seq workflow that increases signal-to-noise and allows paired measurement of cell surface markers and chromatin accessibility: integrated cellular indexing of chromatin landscape and epitopes, called ICICLE-seq. We extended this approach using a droplet-based multiomics platform to develop a trimodal assay that simultaneously measures transcriptomics (scRNA-seq), epitopes, and chromatin accessibility (scATAC-seq) from thousands of single cells, which we term TEA-seq. Together, these multimodal sing
- 160Single-cell sequencing techniques from individual to multiomics analyses.Here, we review single-cell sequencing techniques for individual and multiomics profiling in single cells. We mainly describe single-cell genomic, epigenomic, and transcriptomic methods, and examples of their applications. For the integration of multilayered data sets, such as the transcriptome data derived from single-cell RNA sequencing and chromatin accessibility data derived from single-cell ATAC-seq, there are several computational integration methods. We also describe single-cell experimental methods for the simultaneous measurement of two or more omics layers. We can achieve a detailed understanding of the basic molecular profiles and those associated with disease in each cell by utilizing a large number of single-cell sequencing techniques and the accumulated data sets.
- 161Reference-free cell type deconvolution of multi-cellular pixel-resolution spatially resolved transcriptomics data.Recent technological advancements have enabled spatially resolved transcriptomic profiling but at multi-cellular pixel resolution, thereby hindering the identification of cell-type-specific spatial patterns and gene expression variation. To address this challenge, we develop STdeconvolve as a reference-free approach to deconvolve underlying cell types comprising such multi-cellular pixel resolution spatial transcriptomics (ST) datasets. Using simulated as well as real ST datasets from diverse spatial transcriptomics technologies comprising a variety of spatial resolutions such as Spatial Transcriptomics, 10X Visium, DBiT-seq, and Slide-seq, we show that STdeconvolve can effectively recover cell-type transcriptional profiles and their proportional representation within pixels without reliance on external single-cell transcriptomics references. STdeconvolve provides comparable performance to existing reference-based methods when suitable single-cell references are available, as well as p
- 162A rapid and robust method for single cell chromatin accessibility profiling.The assay for transposase-accessible chromatin using sequencing (ATAC-seq) is widely used to identify regulatory regions throughout the genome. However, very few studies have been performed at the single cell level (scATAC-seq) due to technical challenges. Here we developed a simple and robust plate-based scATAC-seq method, combining upfront bulk Tn5 tagging with single-nuclei sorting. We demonstrate that our method works robustly across various systems, including fresh and cryopreserved cells from primary tissues. By profiling over 3000 splenocytes, we identify distinct immune cell types and reveal cell type-specific regulatory regions and related transcription factors.
- 163Accurate genomic variant detection in single cells with primary template-directed amplification.Improvements in whole genome amplification (WGA) would enable new types of basic and applied biomedical research, including studies of intratissue genetic diversity that require more accurate single-cell genotyping. Here, we present primary template-directed amplification (PTA), an isothermal WGA method that reproducibly captures >95% of the genomes of single cells in a more uniform and accurate manner than existing approaches, resulting in significantly improved variant calling sensitivity and precision. To illustrate the types of studies that are enabled by PTA, we developed direct measurement of environmental mutagenicity (DMEM), a tool for mapping genome-wide interactions of mutagens with single living human cells at base-pair resolution. In addition, we utilized PTA for genome-wide off-target indel and structural variant detection in cells that had undergone CRISPR-mediated genome editing, establishing the feasibility for performing single-cell evaluations of biopsies from edited
- 164Mapping gene regulatory networks from single-cell omics data.Single-cell techniques are advancing rapidly and are yielding unprecedented insight into cellular heterogeneity. Mapping the gene regulatory networks (GRNs) underlying cell states provides attractive opportunities to mechanistically understand this heterogeneity. In this review, we discuss recently emerging methods to map GRNs from single-cell transcriptomics data, tackling the challenge of increased noise levels and data sparsity compared with bulk data, alongside increasing data volumes. Next, we discuss how new techniques for single-cell epigenomics, such as single-cell ATAC-seq and single-cell DNA methylation profiling, can be used to decipher gene regulatory programmes. We finally look forward to the application of single-cell multi-omics and perturbation techniques that will likely play important roles for GRN inference in the future.
- 165SEACells infers transcriptional and epigenomic cellular states from single-cell genomics data.Metacells are cell groupings derived from single-cell sequencing data that represent highly granular, distinct cell states. Here we present single-cell aggregation of cell states (SEACells), an algorithm for identifying metacells that overcome the sparsity of single-cell data while retaining heterogeneity obscured by traditional cell clustering. SEACells outperforms existing algorithms in identifying comprehensive, compact and well-separated metacells in both RNA and assay for transposase-accessible chromatin (ATAC) modalities across datasets with discrete cell types and continuous trajectories. We demonstrate the use of SEACells to improve gene-peak associations, compute ATAC gene scores and infer the activities of critical regulators during differentiation. Metacell-level analysis scales to large datasets and is particularly well suited for patient cohorts, where per-patient aggregation provides more robust units for data integration. We use our metacells to reveal expression dynamic
- 166Single-cell multi-omics reveals dyssynchrony of the innate and adaptive immune system in progressive COVID-19.Dysregulated immune responses against the SARS-CoV-2 virus are instrumental in severe COVID-19. However, the immune signatures associated with immunopathology are poorly understood. Here we use multi-omics single-cell analysis to probe the dynamic immune responses in hospitalized patients with stable or progressive course of COVID-19, explore V(D)J repertoires, and assess the cellular effects of tocilizumab. Coordinated profiling of gene expression and cell lineage protein markers shows that S100A hi /HLA-DR lo classical monocytes and activated LAG-3 hi T cells are hallmarks of progressive disease and highlights the abnormal MHC-II/LAG-3 interaction on myeloid and T cells, respectively. We also find skewed T cell receptor repertories in expanded effector CD8 + clones, unmutated IGHG + B cell clones, and mutated B cell clones with stable somatic hypermutation frequency over time. In conclusion, our in-depth immune profiling reveals dyssynchrony of the innate and adaptive immune interact
- 167High-resolution alignment of single-cell and spatial transcriptomes with CytoSPACE.Recent studies have emphasized the importance of single-cell spatial biology, yet available assays for spatial transcriptomics have limited gene recovery or low spatial resolution. Here we introduce CytoSPACE, an optimization method for mapping individual cells from a single-cell RNA sequencing atlas to spatial expression profiles. Across diverse platforms and tissue types, we show that CytoSPACE outperforms previous methods with respect to noise tolerance and accuracy, enabling tissue cartography at single-cell resolution.
- 168Delineating the dynamic evolution from preneoplasia to invasive lung adenocarcinoma by integrating single-cell RNA sequencing and spatial transcriptomics.The cell ecology and spatial niche implicated in the dynamic and sequential process of lung adenocarcinoma (LUAD) from adenocarcinoma in situ (AIS) to minimally invasive adenocarcinoma (MIA) and subsequent invasive adenocarcinoma (IAC) have not yet been elucidated. Here, we performed an integrative analysis of single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) to characterize the cell atlas of the invasion trajectory of LUAD. We found that the UBE2C + cancer cell subpopulation constantly increased during the invasive process of LUAD with remarkable elevation in IAC, and its spatial distribution was in the peripheral cancer region of the IAC, representing a more malignant phenotype. Furthermore, analysis of the TME cell type subpopulation showed a constant decrease in mast cells, monocytes, and lymphatic endothelial cells, which were implicated in the whole process of invasive LUAD, accompanied by an increase in NK cells and MALT B cells from AIS to MIA and an incre
- 169Single-cell RNA sequencing and spatial transcriptomics reveal cancer-associated fibroblasts in glioblastoma with protumoral effects.Cancer-associated fibroblasts (CAFs) were presumed absent in glioblastoma given the lack of brain fibroblasts. Serial trypsinization of glioblastoma specimens yielded cells with CAF morphology and single-cell transcriptomic profiles based on their lack of copy number variations (CNVs) and elevated individual cell CAF probability scores derived from the expression of 9 CAF markers and absence of 5 markers from non-CAF stromal cells sharing features with CAFs. Cells without CNVs and with high CAF probability scores were identified in single-cell RNA-Seq of 12 patient glioblastomas. Pseudotime reconstruction revealed that immature CAFs evolved into subtypes, with mature CAFs expressing actin alpha 2, smooth muscle (ACTA2). Spatial transcriptomics from 16 patient glioblastomas confirmed CAF proximity to mesenchymal glioblastoma stem cells (GSCs), endothelial cells, and M2 macrophages. CAFs were chemotactically attracted to GSCs, and CAFs enriched GSCs. We created a resource of inferred cro
- 170Single-Cell RNA Sequencing with Spatial Transcriptomics of Cancer Tissues.Single-cell RNA sequencing (RNA-seq) techniques can perform analysis of transcriptome at the single-cell level and possess an unprecedented potential for exploring signatures involved in tumor development and progression. These techniques can perform sequence analysis of transcripts with a better resolution that could increase understanding of the cellular diversity found in the tumor microenvironment and how the cells interact with each other in complex heterogeneous cancerous tissues. Identifying the changes occurring in the genome and transcriptome in the spatial context is considered to increase knowledge of molecular factors fueling cancers. It may help develop better monitoring strategies and innovative approaches for cancer treatment. Recently, there has been a growing trend in the integration of RNA-seq techniques with contemporary omics technologies to study the tumor microenvironment. There has been a realization that this area of research has a huge scope of application in t
- 171Spatial Transcriptomic Technologies.Spatial transcriptomic technologies enable measurement of expression levels of genes systematically throughout tissue space, deepening our understanding of cellular organizations and interactions within tissues as well as illuminating biological insights in neuroscience, developmental biology and a range of diseases, including cancer. A variety of spatial technologies have been developed and/or commercialized, differing in spatial resolution, sensitivity, multiplexing capability, throughput and coverage. In this paper, we review key enabling spatial transcriptomic technologies and their applications as well as the perspective of the techniques and new emerging technologies that are developed to address current limitations of spatial methodologies. In addition, we describe how spatial transcriptomics data can be integrated with other omics modalities, complementing other methods in deciphering cellar interactions and phenotypes within tissues as well as providing novel insight into tiss
- 172The promising application of cell-cell interaction analysis in cancer from single-cell and spatial transcriptomics.Cell-cell interactions instruct cell fate and function. These interactions are hijacked to promote cancer development. Single-cell transcriptomics and spatial transcriptomics have become powerful new tools for researchers to profile the transcriptional landscape of cancer at unparalleled genetic depth. In this review, we discuss the rapidly growing array of computational tools to infer cell-cell interactions from non-spatial single-cell RNA-sequencing and the limited but growing number of methods for spatial transcriptomics data. Downstream analyses of these computational tools and applications to cancer studies are highlighted. We finish by suggesting several directions for further extensions that anticipate the increasing availability of multi-omics cancer data.
- 173Spatial Transcriptomics: Emerging Technologies in Tissue Gene Expression Profiling.In this Perspective, we discuss the current status and advances in spatial transcriptomics technologies, which allow high-resolution mapping of gene expression in intact cell and tissue samples. Spatial transcriptomics enables the creation of high-resolution maps of gene expression patterns within their native spatial context, adding an extra layer of information to the bulk sequencing data. Spatial transcriptomics has expanded significantly in recent years and is making a notable impact on a range of fields, including tissue architecture, developmental biology, cancer, and neurodegenerative and infectious diseases. The latest advancements in spatial transcriptomics have resulted in the development of highly multiplexed methods, transcriptomic-wide analysis, and single-cell resolution utilizing diverse technological approaches. In this Perspective, we provide a detailed analysis of the molecular foundations behind the main spatial transcriptomics technologies, including methods based o
- 174Advances and Challenges in Spatial Transcriptomics for Developmental Biology.Development from single cells to multicellular tissues and organs involves more than just the exact replication of cells, which is known as differentiation. The primary focus of research into the mechanism of differentiation has been differences in gene expression profiles between individual cells. However, it has predominantly been conducted at low throughput and bulk levels, challenging the efforts to understand molecular mechanisms of differentiation during the developmental process in animals and humans. During the last decades, rapid methodological advancements in genomics facilitated the ability to study developmental processes at a genome-wide level and finer resolution. Particularly, sequencing transcriptomes at single-cell resolution, enabled by single-cell RNA-sequencing (scRNA-seq), was a breath-taking innovation, allowing scientists to gain a better understanding of differentiation and cell lineage during the developmental process. However, single-cell isolation during scRN
- 175Exploring the untapped potential of single-cell and spatial omics in plant biology.Advances in single-cell and spatial omics technologies have revolutionised biology by revealing the diverse molecular states of individual cells and their spatial organization within tissues. The field of plant biology has widely adopted single-cell transcriptome and chromatin accessibility profiling and spatial transcriptomics, which extend traditional cell biology and genomics analyses and provide unique opportunities to reveal molecular and cellular dynamics of tissues. Using these technologies, comprehensive cell atlases have been generated in several model plant species, providing valuable platforms for discovery and tool development. Other emerging technologies related to single-cell and spatial omics, such as multiomics, lineage tracing, molecular recording, and high-content genetic and chemical perturbation phenotyping, offer immense potential for deepening our understanding of plant biology yet remain underutilised due to unique technical challenges and resource availability.
- 176Two-dimensional single-cell patterning with one cell per well driven by surface acoustic waves.In single-cell analysis, cellular activity and parameters are assayed on an individual, rather than population-average basis. Essential to observing the activity of these cells over time is the ability to trap, pattern and retain them, for which previous single-cell-patterning work has principally made use of mechanical methods. While successful as a long-term cell-patterning strategy, these devices remain essentially single use. Here we introduce a new method for the patterning of multiple spatially separated single particles and cells using high-frequency acoustic fields with one cell per acoustic well. We characterize and demonstrate patterning for both a range of particle sizes and the capture and patterning of cells, including human lymphocytes and red blood cells infected by the malarial parasite Plasmodium falciparum. This ability is made possible by a hitherto unexplored regime where the acoustic wavelength is on the same order as the cell dimensions.
- 177APOE modulates microglial immunometabolism in response to age, amyloid pathology, and inflammatory challenge.The E4 allele of Apolipoprotein E (APOE) is associated with both metabolic dysfunction and a heightened pro-inflammatory response: two findings that may be intrinsically linked through the concept of immunometabolism. Here, we combined bulk, single-cell, and spatial transcriptomics with cell-specific and spatially resolved metabolic analyses in mice expressing human APOE to systematically address the role of APOE across age, neuroinflammation, and AD pathology. RNA sequencing (RNA-seq) highlighted immunometabolic changes across the APOE4 glial transcriptome, specifically in subsets of metabolically distinct microglia enriched in the E4 brain during aging or following an inflammatory challenge. E4 microglia display increased Hif1α expression and a disrupted tricarboxylic acid (TCA) cycle and are inherently pro-glycolytic, while spatial transcriptomics and mass spectrometry imaging highlight an E4-specific response to amyloid that is characterized by widespread alterations in lipid metab
- 178Single nucleus multi-omics identifies human cortical cell regulatory genome diversity.Single-cell technologies measure unique cellular signatures but are typically limited to a single modality. Computational approaches allow the fusion of diverse single-cell data types, but their efficacy is difficult to validate in the absence of authentic multi-omic measurements. To comprehensively assess the molecular phenotypes of single cells, we devised single-nucleus methylcytosine, chromatin accessibility, and transcriptome sequencing (snmCAT-seq) and applied it to postmortem human frontal cortex tissue. We developed a cross-validation approach using multi-modal information to validate fine-grained cell types and assessed the effectiveness of computational data fusion methods. Correlation analysis in individual cells revealed distinct relations between methylation and gene expression. Our integrative approach enabled joint analyses of the methylome, transcriptome, chromatin accessibility, and conformation for 63 human cortical cell types. We reconstructed regulatory lineages for
- 179Integration of spatial and single-cell transcriptomics localizes epithelial cell-immune cross-talk in kidney injury.Single-cell sequencing studies have characterized the transcriptomic signature of cell types within the kidney. However, the spatial distribution of acute kidney injury (AKI) is regional and affects cells heterogeneously. We first optimized coordination of spatial transcriptomics and single-nuclear sequencing data sets, mapping 30 dominant cell types to a human nephrectomy. The predicted cell-type spots corresponded with the underlying histopathology. To study the implications of AKI on transcript expression, we then characterized the spatial transcriptomic signature of 2 murine AKI models: ischemia/reperfusion injury (IRI) and cecal ligation puncture (CLP). Localized regions of reduced overall expression were associated with injury pathways. Using single-cell sequencing, we deconvoluted the signature of each spatial transcriptomic spot, identifying patterns of colocalization between immune and epithelial cells. Neutrophils infiltrated the renal medulla in the ischemia model. Atf3 was
- 180Single-cell multi-omics identifies chronic inflammation as a driver of TP53-mutant leukemic evolution.Understanding the genetic and nongenetic determinants of tumor protein 53 (TP53)-mutation-driven clonal evolution and subsequent transformation is a crucial step toward the design of rational therapeutic strategies. Here we carry out allelic resolution single-cell multi-omic analysis of hematopoietic stem/progenitor cells (HSPCs) from patients with a myeloproliferative neoplasm who transform to TP53-mutant secondary acute myeloid leukemia (sAML). All patients showed dominant TP53 'multihit' HSPC clones at transformation, with a leukemia stem cell transcriptional signature strongly predictive of adverse outcomes in independent cohorts, across both TP53-mutant and wild-type (WT) AML. Through analysis of serial samples, antecedent TP53-heterozygous clones and in vivo perturbations, we demonstrate a hitherto unrecognized effect of chronic inflammation, which suppressed TP53 WT HSPCs while enhancing the fitness advantage of TP53-mutant cells and promoted genetic evolution. Our findings will
- 181Single-cell multiomics sequencing reveals the functional regulatory landscape of early embryos.Extensive epigenetic reprogramming occurs during preimplantation embryo development. However, it remains largely unclear how the drastic epigenetic reprogramming contributes to transcriptional regulatory network during this period. Here, we develop a single-cell multiomics sequencing technology (scNOMeRe-seq) that enables profiling of genome-wide chromatin accessibility, DNA methylation and RNA expression in the same individual cell. We apply this method to depict a single-cell multiomics map of mouse preimplantation development. We find that genome-wide DNA methylation remodeling facilitates the reconstruction of genetic lineages in early embryos. Further, we construct a zygotic genome activation (ZGA)-associated regulatory network and reveal coordination among multiple epigenetic layers, transcription factors and repeat elements that instruct proper ZGA. Cell fates associated cis-regulatory elements are activated stepwise in post-ZGA stages. Trophectoderm (TE)-specific transcription
- 182Single-cell multi-omics in the medicinal plant Catharanthus roseus.Advances in omics technologies now permit the generation of highly contiguous genome assemblies, detection of transcripts and metabolites at the level of single cells and high-resolution determination of gene regulatory features. Here, using a complementary, multi-omics approach, we interrogated the monoterpene indole alkaloid (MIA) biosynthetic pathway in Catharanthus roseus, a source of leading anticancer drugs. We identified clusters of genes involved in MIA biosynthesis on the eight C. roseus chromosomes and extensive gene duplication of MIA pathway genes. Clustering was not limited to the linear genome, and through chromatin interaction data, MIA pathway genes were present within the same topologically associated domain, permitting the identification of a secologanin transporter. Single-cell RNA-sequencing revealed sequential cell-type-specific partitioning of the leaf MIA biosynthetic pathway that, when coupled with a single-cell metabolomics approach, permitted the identificatio
- 183Single-cell biological network inference using a heterogeneous graph transformer.Single-cell multi-omics (scMulti-omics) allows the quantification of multiple modalities simultaneously to capture the intricacy of complex molecular mechanisms and cellular heterogeneity. Existing tools cannot effectively infer the active biological networks in diverse cell types and the response of these networks to external stimuli. Here we present DeepMAPS for biological network inference from scMulti-omics. It models scMulti-omics in a heterogeneous graph and learns relations among cells and genes within both local and global contexts in a robust manner using a multi-head graph transformer. Benchmarking results indicate DeepMAPS performs better than existing tools in cell clustering and biological network construction. It also showcases competitive capability in deriving cell-type-specific biological networks in lung tumor leukocyte CITE-seq data and matched diffuse small lymphocytic lymphoma scRNA-seq and scATAC-seq data. In addition, we deploy a DeepMAPS webserver equipped with
- 184BASS: multi-scale and multi-sample analysis enables accurate cell type clustering and spatial domain detection in spatial transcriptomic studies.Spatial transcriptomic studies are reaching single-cell spatial resolution, with data often collected from multiple tissue sections. Here, we present a computational method, BASS, that enables multi-scale and multi-sample analysis for single-cell resolution spatial transcriptomics. BASS performs cell type clustering at the single-cell scale and spatial domain detection at the tissue regional scale, with the two tasks carried out simultaneously within a Bayesian hierarchical modeling framework. We illustrate the benefits of BASS through comprehensive simulations and applications to three datasets. The substantial power gain brought by BASS allows us to reveal accurate transcriptomic and cellular landscape in both cortex and hypothalamus.
- 185Single-cell and spatial transcriptomics identify a macrophage population associated with skeletal muscle fibrosis.Macrophages are essential for skeletal muscle homeostasis, but how their dysregulation contributes to the development of fibrosis in muscle disease remains unclear. Here, we used single-cell transcriptomics to determine the molecular attributes of dystrophic and healthy muscle macrophages. We identified six clusters and unexpectedly found that none corresponded to traditional definitions of M1 or M2 macrophages. Rather, the predominant macrophage signature in dystrophic muscle was characterized by high expression of fibrotic factors, galectin-3 (gal-3) and osteopontin ( Spp1 ). Spatial transcriptomics, computational inferences of intercellular communication, and in vitro assays indicated that macrophage-derived Spp1 regulates stromal progenitor differentiation. Gal-3 + macrophages were chronically activated in dystrophic muscle, and adoptive transfer assays showed that the gal-3 + phenotype was the dominant molecular program induced within the dystrophic milieu. Gal-3 + macrophages wer
- 186Single-Nucleus RNA Sequencing and Spatial Transcriptomics Reveal the Immunological Microenvironment of Cervical Squamous Cell Carcinoma.The effective treatment of advanced cervical cancer remains challenging. Herein, single-nucleus RNA sequencing (snRNA-seq) and SpaTial enhanced resolution omics-sequencing (Stereo-seq) are used to investigate the immunological microenvironment of cervical squamous cell carcinoma (CSCC). The expression levels of most immune suppressive genes in the tumor and inflammation areas of CSCC are not significantly higher than those in the non-cancer samples, except for LGALS9 and IDO1. Stronger signals of CD56 + NK cells and immature dendritic cells are found in the hypermetabolic tumor areas, whereas more eosinophils, immature B cells, and Treg cells are found in the hypometabolic tumor areas. Moreover, a cluster of pro-tumorigenic cancer-associated myofibroblasts (myCAFs) are identified. The myCAFs may support the growth and metastasis of tumors by inhibiting lymphocyte infiltration and remodeling of the tumor extracellular matrix. Furthermore, these myCAFs are associated with poorer survival
- 187Single Cell Multi-Omics Technology: Methodology and Application.In the era of precision medicine, multi-omics approaches enable the integration of data from diverse omics platforms, providing multi-faceted insight into the interrelation of these omics layers on disease processes. Single cell sequencing technology can dissect the genotypic and phenotypic heterogeneity of bulk tissue and promises to deepen our understanding of the underlying mechanisms governing both health and disease. Through modification and combination of single cell assays available for transcriptome, genome, epigenome, and proteome profiling, single cell multi-omics approaches have been developed to simultaneously and comprehensively study not only the unique genotypic and phenotypic characteristics of single cells, but also the combined regulatory mechanisms evident only at single cell resolution. In this review, we summarize the state-of-the-art single cell multi-omics methods and discuss their applications, challenges, and future directions.
- 188Multi-Omics of Single Cells: Strategies and Applications.Most genome-wide assays provide averages across large numbers of cells, but recent technological advances promise to overcome this limitation. Pioneering single-cell assays are now available for genome, epigenome, transcriptome, proteome, and metabolome profiling. Here, we describe how these different dimensions can be combined into multi-omics assays that provide comprehensive profiles of the same cell.
- 189Combined Single-Cell and Spatial Transcriptomics Reveal the Metabolic Evolvement of Breast Cancer during Early Dissemination.Breast cancer is now the most frequently diagnosed malignancy, and metastasis remains the leading cause of death in breast cancer. However, little is known about the dynamic changes during the evolvement of dissemination. In this study, 65 968 cells from four patients with breast cancer and paired metastatic axillary lymph nodes are profiled using single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics. A disseminated cancer cell cluster with high levels of oxidative phosphorylation (OXPHOS), including the upregulation of cytochrome C oxidase subunit 6C and dehydrogenase/reductase 2, is identified. The transition between glycolysis and OXPHOS when dissemination initiates is noticed. Furthermore, this distinct cell cluster is distributed along the tumor's leading edge. The findings here are verified in three different cohorts of breast cancer patients and an external scRNA-seq dataset, which includes eight patients with breast cancer and paired metastatic axillary lymph nodes
- 190Simultaneous CRISPR screening and spatial transcriptomics reveal intracellular, intercellular, and functional transcriptional circuits.Pooled optical screens have enabled the study of cellular interactions, morphology, or dynamics at massive scale, but they have not yet leveraged the power of highly plexed single-cell resolved transcriptomic readouts to inform molecular pathways. Here, we present a combination of imaging spatial transcriptomics with parallel optical detection of in situ amplified guide RNAs (Perturb-FISH). Perturb-FISH recovers intracellular effects that are consistent with single-cell RNA-sequencing-based readouts of perturbation effects (Perturb-seq) in a screen of lipopolysaccharide response in cultured monocytes, and it uncovers intercellular and density-dependent regulation of the innate immune response. Similarly, in three-dimensional xenograft models, Perturb-FISH identifies tumor-immune interactions altered by genetic knockout. When paired with a functional readout in a separate screen of autism spectrum disorder risk genes in human-induced pluripotent stem cell (hIPSC) astrocytes, Perturb-FIS
- 191Integration Analysis of Single-Cell Multi-Omics Reveals Prostate Cancer Heterogeneity.Prostate cancer (PCa) is an extensive heterogeneous disease with a complex cellular ecosystem in the tumor microenvironment (TME). However, the manner in which heterogeneity is shaped by tumors and stromal cells, or vice versa, remains poorly understood. In this study, single-cell RNA sequencing, spatial transcriptomics, and bulk ATAC-sequence are integrated from a series of patients with PCa and healthy controls. A stemness subset of club cells marked with SOX9 high AR low expression is identified, which is markedly enriched after neoadjuvant androgen-deprivation therapy (ADT). Furthermore, a subset of CD8 + CXCR6 + T cells that function as effector T cells is markedly reduced in patients with malignant PCa. For spatial transcriptome analysis, machine learning and computational intelligence are comprehensively utilized to identify the cellular diversity of prostate cancer cells and cell-cell communication in situ. Macrophage and neutrophil state transitions along the trajectory of can
- 192BIDCell: Biologically-informed self-supervised learning for segmentation of subcellular spatial transcriptomics data.Recent advances in subcellular imaging transcriptomics platforms have enabled high-resolution spatial mapping of gene expression, while also introducing significant analytical challenges in accurately identifying cells and assigning transcripts. Existing methods grapple with cell segmentation, frequently leading to fragmented cells or oversized cells that capture contaminated expression. To this end, we present BIDCell, a self-supervised deep learning-based framework with biologically-informed loss functions that learn relationships between spatially resolved gene expression and cell morphology. BIDCell incorporates cell-type data, including single-cell transcriptomics data from public repositories, with cell morphology information. Using a comprehensive evaluation framework consisting of metrics in five complementary categories for cell segmentation performance, we demonstrate that BIDCell outperforms other state-of-the-art methods according to many metrics across a variety of tissue
- 193Single cell and spatial transcriptomics highlight the interaction of club-like cells with immunosuppressive myeloid cells in prostate cancer.Prostate cancer treatment resistance is a significant challenge facing the field. Genomic and transcriptomic profiling have partially elucidated the mechanisms through which cancer cells escape treatment, but their relation toward the tumor microenvironment (TME) remains elusive. Here we present a comprehensive transcriptomic landscape of the prostate TME at multiple points in the standard treatment timeline employing single-cell RNA-sequencing and spatial transcriptomics data from 120 patients. We identify club-like cells as a key epithelial cell subtype that acts as an interface between the prostate and the immune system. Tissue areas enriched with club-like cells have depleted androgen signaling and upregulated expression of luminal progenitor cell markers. Club-like cells display a senescence-associated secretory phenotype and their presence is linked to increased polymorphonuclear myeloid-derived suppressor cell (PMN-MDSC) activity. Our results indicate that club-like cells are as
- 194Deciphering cell-cell communication at single-cell resolution for spatial transcriptomics with subgraph-based graph attention network.The inference of cell-cell communication (CCC) is crucial for a better understanding of complex cellular dynamics and regulatory mechanisms in biological systems. However, accurately inferring spatial CCCs at single-cell resolution remains a significant challenge. To address this issue, we present a versatile method, called DeepTalk, to infer spatial CCC at single-cell resolution by integrating single-cell RNA sequencing (scRNA-seq) data and spatial transcriptomics (ST) data. DeepTalk utilizes graph attention network (GAT) to integrate scRNA-seq and ST data, which enables accurate cell-type identification for single-cell ST data and deconvolution for spot-based ST data. Then, DeepTalk can capture the connections among cells at multiple levels using subgraph-based GAT, and further achieve spatially resolved CCC inference at single-cell resolution. DeepTalk achieves excellent performance in discovering meaningful spatial CCCs on multiple cross-platform datasets, which demonstrates its su
- 195Spatial transcriptomics reveals human cortical layer and area specification.The human cerebral cortex is composed of six layers and dozens of areas that are molecularly and structurally distinct 1-4 . Although single-cell transcriptomic studies have advanced the molecular characterization of human cortical development, a substantial gap exists owing to the loss of spatial context during cell dissociation 5-8 . Here we used multiplexed error-robust fluorescence in situ hybridization (MERFISH) 9 , augmented with deep-learning-based nucleus segmentation, to examine the molecular, cellular and cytoarchitectural development of the human fetal cortex with spatially resolved single-cell resolution. Our extensive spatial atlas, encompassing more than 18 million single cells, spans eight cortical areas across seven developmental time points. We uncovered the early establishment of the six-layer structure, identifiable by the laminar distribution of excitatory neuron subtypes, 3 months before the emergence of cytoarchitectural layers. Notably, we discovered two distinct
- 196Analysis and Visualization of Spatial Transcriptomic Data.Human and animal tissues consist of heterogeneous cell types that organize and interact in highly structured manners. Bulk and single-cell sequencing technologies remove cells from their original microenvironments, resulting in a loss of spatial information. Spatial transcriptomics is a recent technological innovation that measures transcriptomic information while preserving spatial information. Spatial transcriptomic data can be generated in several ways. RNA molecules are measured by in situ sequencing, in situ hybridization, or spatial barcoding to recover original spatial coordinates. The inclusion of spatial information expands the range of possibilities for analysis and visualization, and spurred the development of numerous novel methods. In this review, we summarize the core concepts of spatial genomics technology and provide a comprehensive review of current analysis and visualization methods for spatial transcriptomics.
- 197Isoform Age - Splice Isoform Profiling Using Long-Read Technologies.Alternative splicing (AS) of RNA is a key mechanism that results in the expression of multiple transcript isoforms from single genes and leads to an increase in the complexity of both the transcriptome and proteome. Regulation of AS is critical for the correct functioning of many biological pathways, while disruption of AS can be directly pathogenic in diseases such as cancer or cause risk for complex disorders. Current short-read sequencing technologies achieve high read depth but are limited in their ability to resolve complex isoforms. In this review we examine how long-read sequencing (LRS) technologies can address this challenge by covering the entire RNA sequence in a single read and thereby distinguish isoform changes that could impact RNA regulation or protein function. Coupling LRS with technologies such as single cell sequencing, targeted sequencing and spatial transcriptomics is producing a rapidly expanding suite of technological approaches to profile alternative splicing a
- 198Spatial transcriptomics in neuroscience.The brain is one of the most complex living tissue types and is composed of an exceptional diversity of cell types displaying unique functional connectivity. Single-cell RNA sequencing (scRNA-seq) can be used to efficiently map the molecular identities of the various cell types in the brain by providing the transcriptomic profiles of individual cells isolated from the tissue. However, the lack of spatial context in scRNA-seq prevents a comprehensive understanding of how different configurations of cell types give rise to specific functions in individual brain regions and how each distinct cell is connected to form a functional unit. To understand how the various cell types contribute to specific brain functions, it is crucial to correlate the identities of individual cells obtained through scRNA-seq with their spatial information in intact tissue. Spatial transcriptomics (ST) can resolve the complex spatial organization of cell types in the brain and their connectivity. Various ST tool
- 199Single-cell omics traces the heterogeneity of prostate cancer cells and the tumor microenvironment.Prostate cancer is one of the more heterogeneous tumour types. In recent years, with the rapid development of single-cell sequencing and spatial transcriptome technologies, researchers have gained a more intuitive and comprehensive understanding of the heterogeneity of prostate cancer. Tumour-associated epithelial cells; cancer-associated fibroblasts; the complexity of the immune microenvironment, and the heterogeneity of the spatial distribution of tumour cells and other cancer-promoting molecules play a crucial role in the growth, invasion, and metastasis of prostate cancer. Single-cell multi-omics biotechnology, especially single-cell transcriptome sequencing, reveals the expression level of single cells with higher resolution and finely dissects the molecular characteristics of different tumour cells. We reviewed the recent literature on prostate cancer cells, focusing on single-cell RNA sequencing. And we analysed the heterogeneity and spatial distribution differences of different
- 200A review of spatial profiling technologies for characterizing the tumor microenvironment in immuno-oncology.Interpreting the mechanisms and principles that govern gene activity and how these genes work according to -their cellular distribution in organisms has profound implications for cancer research. The latest technological advancements, such as imaging-based approaches and next-generation single-cell sequencing technologies, have established a platform for spatial transcriptomics to systematically quantify the expression of all or most genes in the entire tumor microenvironment and explore an array of disease milieus, particularly in tumors. Spatial profiling technologies permit the study of transcriptional activity at the spatial or single-cell level. This multidimensional classification of the transcriptomic and proteomic signatures of tumors, especially the associated immune and stromal cells, facilitates evaluation of tumor heterogeneity, details of the evolutionary trajectory of each tumor, and multifaceted interactions between each tumor cell and its microenvironment. Therefore, sp
- 201A spatial sequencing atlas of age-induced changes in the lung during influenza infection.Influenza virus infection causes increased morbidity and mortality in the elderly. Aging impairs the immune response to influenza, both intrinsically and because of altered interactions with endothelial and pulmonary epithelial cells. To characterize these changes, we performed single-cell RNA sequencing (scRNA-seq), spatial transcriptomics, and bulk RNA sequencing (bulk RNA-seq) on lung tissue from young and aged female mice at days 0, 3, and 9 post-influenza infection. Our analyses identified dozens of key genes differentially expressed in kinetic, age-dependent, and cell type-specific manners. Aged immune cells exhibited altered inflammatory, memory, and chemotactic profiles. Aged endothelial cells demonstrated characteristics of reduced vascular wound healing and a prothrombotic state. Spatial transcriptomics identified novel profibrotic and antifibrotic markers expressed by epithelial and non-epithelial cells, highlighting the complex networks that promote fibrosis in aged lungs.
- 202Comprehensive visualization of cell-cell interactions in single-cell and spatial transcriptomics with NICHES.Motivation Recent years have seen the release of several toolsets that reveal cell-cell interactions from single-cell data. However, all existing approaches leverage mean celltype gene expression values, and do not preserve the single-cell fidelity of the original data. Here, we present NICHES (Niche Interactions and Communication Heterogeneity in Extracellular Signaling), a tool to explore extracellular signaling at the truly single-cell level. Results NICHES allows embedding of ligand-receptor signal proxies to visualize heterogeneous signaling archetypes within cell clusters, between cell clusters and across experimental conditions. When applied to spatial transcriptomic data, NICHES can be used to reflect local cellular microenvironment. NICHES can operate with any list of ligand-receptor signaling mechanisms, is compatible with existing single-cell packages, and allows rapid, flexible analysis of cell-cell signaling at single-cell resolution. Availability and implementation NICHES
- 203Clonal relations in the mouse brain revealed by single-cell and spatial transcriptomics.The mammalian brain contains many specialized cells that develop from a thin sheet of neuroepithelial progenitor cells. Single-cell transcriptomics revealed hundreds of molecularly diverse cell types in the nervous system, but the lineage relationships between mature cell types and progenitor cells are not well understood. Here we show in vivo barcoding of early progenitors to simultaneously profile cell phenotypes and clonal relations in the mouse brain using single-cell and spatial transcriptomics. By reconstructing thousands of clones, we discovered fate-restricted progenitor cells in the mouse hippocampal neuroepithelium and show that microglia are derived from few primitive myeloid precursors that massively expand to generate widely dispersed progeny. We combined spatial transcriptomics with clonal barcoding and disentangled migration patterns of clonally related cells in densely labeled tissue sections. Our approach enables high-throughput dense reconstruction of cell phenotypes
- 204Innovative breakthroughs facilitated by single-cell multi-omics: manipulating natural killer cell functionality correlates with a novel subcategory of melanoma cells.Background Melanoma is typically regarded as the most dangerous form of skin cancer. Although surgical removal of in situ lesions can be used to effectively treat metastatic disease, this condition is still difficult to cure. Melanoma cells are removed in great part due to the action of natural killer (NK) and T cells on the immune system. Still, not much is known about how the activity of NK cell-related pathways changes in melanoma tissue. Thus, we performed a single-cell multi-omics analysis on human melanoma cells in this study to explore the modulation of NK cell activity. Materials and methods Cells in which mitochondrial genes comprised > 20% of the total number of expressed genes were removed. Gene ontology (GO), gene set enrichment analysis (GSEA), gene set variation analysis (GSVA), and AUCcell analysis of differentially expressed genes (DEGs) in melanoma subtypes were performed. The CellChat package was used to predict cell-cell contact between NK cell and melanoma cell subt
- 205Joint cell segmentation and cell type annotation for spatial transcriptomics.RNA hybridization-based spatial transcriptomics provides unparalleled detection sensitivity. However, inaccuracies in segmentation of image volumes into cells cause misassignment of mRNAs which is a major source of errors. Here, we develop JSTA, a computational framework for joint cell segmentation and cell type annotation that utilizes prior knowledge of cell type-specific gene expression. Simulation results show that leveraging existing cell type taxonomy increases RNA assignment accuracy by more than 45%. Using JSTA, we were able to classify cells in the mouse hippocampus into 133 (sub)types revealing the spatial organization of CA1, CA3, and Sst neuron subtypes. Analysis of within cell subtype spatial differential gene expression of 80 candidate genes identified 63 with statistically significant spatial differential gene expression across 61 (sub)types. Overall, our work demonstrates that known cell type expression patterns can be leveraged to improve the accuracy of RNA hybridizat
- 206Inferring Interaction Networks From Multi-Omics Data.A major goal in systems biology is a comprehensive description of the entirety of all complex interactions between different types of biomolecules-also referred to as the interactome-and how these interactions give rise to higher, cellular and organism level functions or diseases. Numerous efforts have been undertaken to define such interactomes experimentally, for example yeast-two-hybrid based protein-protein interaction networks or ChIP-seq based protein-DNA interactions for individual proteins. To complement these direct measurements, genome-scale quantitative multi-omics data (transcriptomics, proteomics, metabolomics, etc.) enable researchers to predict novel functional interactions between molecular species. Moreover, these data allow to distinguish relevant functional from non-functional interactions in specific biological contexts. However, integration of multi-omics data is not straight forward due to their heterogeneity. Numerous methods for the inference of interaction netw
- 207STRIDE: accurately decomposing and integrating spatial transcriptomics using single-cell RNA sequencing.The recent advances in spatial transcriptomics have brought unprecedented opportunities to understand the cellular heterogeneity in the spatial context. However, the current limitations of spatial technologies hamper the exploration of cellular localizations and interactions at single-cell level. Here, we present spatial transcriptomics deconvolution by topic modeling (STRIDE), a computational method to decompose cell types from spatial mixtures by leveraging topic profiles trained from single-cell transcriptomics. STRIDE accurately estimated the cell-type proportions and showed balanced specificity and sensitivity compared to existing methods. We demonstrated STRIDE's utility by applying it to different spatial platforms and biological systems. Deconvolution by STRIDE not only mapped rare cell types to spatial locations but also improved the identification of spatially localized genes and domains. Moreover, topics discovered by STRIDE were associated with cell-type-specific functions
- 208Single-cell multiomics analysis reveals regulatory programs in clear cell renal cell carcinoma.The clear cell renal cell carcinoma (ccRCC) microenvironment consists of many different cell types and structural components that play critical roles in cancer progression and drug resistance, but the cellular architecture and underlying gene regulatory features of ccRCC have not been fully characterized. Here, we applied single-cell RNA sequencing (scRNA-seq) and single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq) to generate transcriptional and epigenomic landscapes of ccRCC. We identified tumor cell-specific regulatory programs mediated by four key transcription factors (TFs) (HOXC5, VENTX, ISL1, and OTP), and these TFs have prognostic significance in The Cancer Genome Atlas (TCGA) database. Targeting these TFs via short hairpin RNAs (shRNAs) or small molecule inhibitors decreased tumor cell proliferation. We next performed an integrative analysis of chromatin accessibility and gene expression for CD8 + T cells and macrophages to reveal the different regul
- 209Spatiotemporally deciphering the mysterious mechanism of persistent HPV-induced malignant transition and immune remodelling from HPV-infected normal cervix, precancer to cervical cancer: Integrating single-cell RNA-sequencing and spatial transcriptome.Background The mechanism underlying cervical carcinogenesis that is mediated by persistent human papillomavirus (HPV) infection remains elusive. Aims Here, for the first time, we deciphered both the temporal transition and spatial distribution of cellular subsets during disease progression from normal cervix tissues to precursor lesions to cervical cancer. Materials & methods We generated scRNA-seq profiles and spatial transcriptomics data from nine patient samples, including two HPV-negative normal, two HPV-positive normal, two HPV-positive HSIL and three HPV-positive cancer samples. Results We not only identified three 'HPV-related epithelial clusters' that are unique to normal, high-grade squamous intraepithelial lesions (HSIL) and cervical cancer tissues but also discovered node genes that potentially regulate disease progression. Moreover, we observed the gradual transition of multiple immune cells that exhibited positive immune responses, followed by dysregulation and exhaustion,
- 210Single-cell spatial transcriptomics reveals a dynamic control of metabolic zonation and liver regeneration by endothelial cell Wnt2 and Wnt9b.The conclusive identity of Wnts regulating liver zonation (LZ) and regeneration (LR) remains unclear despite an undisputed role of β-catenin. Using single-cell analysis, we identified a conserved Wnt2 and Wnt9b expression in endothelial cells (ECs) in zone 3. EC-elimination of Wnt2 and Wnt9b led to both loss of β-catenin targets in zone 3, and re-appearance of zone 1 genes in zone 3, unraveling dynamicity in the LZ process. Impaired LR observed in the knockouts phenocopied models of defective hepatic Wnt signaling. Administration of a tetravalent antibody to activate Wnt signaling rescued LZ and LR in the knockouts and induced zone 3 gene expression and LR in controls. Administration of the agonist also promoted LR in acetaminophen overdose acute liver failure (ALF) fulfilling an unmet clinical need. Overall, we report an unequivocal role of EC-Wnt2 and Wnt9b in LZ and LR and show the role of Wnt activators as regenerative therapy for ALF.
- 211Single cell multi-omics reveal intra-cell-line heterogeneity across human cancer cell lines.Human cancer cell lines have long served as tools for cancer research and drug discovery, but the presence and the source of intra-cell-line heterogeneity remain elusive. Here, we perform single-cell RNA-sequencing and ATAC-sequencing on 42 and 39 human cell lines, respectively, to illustrate both transcriptomic and epigenetic heterogeneity within individual cell lines. Our data reveal that transcriptomic heterogeneity is frequently observed in cancer cell lines of different tissue origins, often driven by multiple common transcriptional programs. Copy number variation, as well as epigenetic variation and extrachromosomal DNA distribution all contribute to the detected intra-cell-line heterogeneity. Using hypoxia treatment as an example, we demonstrate that transcriptomic heterogeneity could be reshaped by environmental stress. Overall, our study performs single-cell multi-omics of commonly used human cancer cell lines and offers mechanistic insights into the intra-cell-line heterogene
- 212Changing Technologies of RNA Sequencing and Their Applications in Clinical Oncology.RNA sequencing (RNAseq) is one of the most commonly used techniques in life sciences, and has been widely used in cancer research, drug development, and cancer diagnosis and prognosis. Driven by various biological and technical questions, the techniques of RNAseq have progressed rapidly from bulk RNAseq, laser-captured micro-dissected RNAseq, and single-cell RNAseq to digital spatial RNA profiling, spatial transcriptomics, and direct in situ sequencing. These different technologies have their unique strengths, weaknesses, and suitable applications in the field of clinical oncology. To guide cancer researchers to select the most appropriate RNAseq technique for their biological questions, we will discuss each of these technologies, technical features, and clinical applications in cancer. We will help cancer researchers to understand the key differences of these RNAseq technologies and their optimal applications.
- 213Concordance of MERFISH spatial transcriptomics with bulk and single-cell RNA sequencing.Spatial transcriptomics extends single-cell RNA sequencing (scRNA-seq) by providing spatial context for cell type identification and analysis. Imaging-based spatial technologies such as multiplexed error-robust fluorescence in situ hybridization (MERFISH) can achieve single-cell resolution, directly mapping single-cell identities to spatial positions. MERFISH produces a different data type than scRNA-seq, and a technical comparison between the two modalities is necessary to ascertain how to best integrate them. We performed MERFISH on the mouse liver and kidney and compared the resulting bulk and single-cell RNA statistics with those from the Tabula Muris Senis cell atlas and from two Visium datasets. MERFISH quantitatively reproduced the bulk RNA-seq and scRNA-seq results with improvements in overall dropout rates and sensitivity. Finally, we found that MERFISH independently resolved distinct cell types and spatial structure in both the liver and kidney. Computational integration with
- 214Spatial transcriptomic characterization of pathologic niches in IPF.Despite advancements in antifibrotic therapy, idiopathic pulmonary fibrosis (IPF) remains a medical condition with unmet needs. Single-cell RNA sequencing (scRNA-seq) has enhanced our understanding of IPF but lacks the cellular tissue context and gene expression localization that spatial transcriptomics provides. To bridge this gap, we profiled IPF and control patient lung tissue using spatial transcriptomics, integrating the data with an IPF scRNA-seq atlas. We identified three disease-associated niches with unique cellular compositions and localizations. These include a fibrotic niche, consisting of myofibroblasts and aberrant basaloid cells, located around airways and adjacent to an airway macrophage niche in the lumen, containing SPP1 + macrophages. In addition, we identified an immune niche, characterized by distinct lymphoid cell foci in fibrotic tissue, surrounded by remodeled endothelial vessels. This spatial characterization of IPF niches will facilitate the identification of
- 215Library size confounds biology in spatial transcriptomics data.Spatial molecular data has transformed the study of disease microenvironments, though, larger datasets pose an analytics challenge prompting the direct adoption of single-cell RNA-sequencing tools including normalization methods. Here, we demonstrate that library size is associated with tissue structure and that normalizing these effects out using commonly applied scRNA-seq normalization methods will negatively affect spatial domain identification. Spatial data should not be specifically corrected for library size prior to analysis, and algorithms designed for scRNA-seq data should be adopted with caution.
- 216Nicheformer: a foundation model for single-cell and spatial omics.Tissue makeup depends on the local cellular microenvironment. Spatial single-cell genomics enables scalable and unbiased interrogation of these interactions. Here we introduce Nicheformer, a transformer-based foundation model trained on both human and mouse dissociated single-cell and targeted spatial transcriptomics data. Pretrained on SpatialCorpus-110M, a curated collection of over 57 million dissociated and 53 million spatially resolved cells across 73 tissues on cellular reconstruction, Nicheformer learns cell representations that capture spatial context. It excels in linear-probing and fine-tuning scenarios for a newly designed set of downstream tasks, in particular spatial composition prediction and spatial label prediction. Critically, we show that models trained only on dissociated data fail to recover the complexity of spatial microenvironments, underscoring the need for multiscale integration. Nicheformer enables the prediction of the spatial context of dissociated cells, al
- 217Single-cell and Spatial Transcriptomics Reveals Ferroptosis as The Most Enriched Programmed Cell Death Process in Hemorrhage Stroke-induced Oligodendrocyte-mediated White Matter Injury.Intracerebral hemorrhage (ICH) is a severe stroke subtype with limited therapeutic options. Programmed cell death (PCD) is crucial for immunological balance, and includes necroptosis, pyroptosis, apoptosis, ferroptosis, and necrosis. However, the distinctions between these programmed cell death modalities after ICH remain to be further investigated. We used single-cell transcriptome (single-cell RNA sequencing) and spatial transcriptome (spatial RNA sequencing) techniques to investigate PCD-related gene expression trends in the rat brain following hemorrhagic stroke. Ferroptosis was the main PCD process after ICH, and primarily affected mature oligodendrocytes. Its onset occurred as early as 1 hour post-ICH, peaking at 24 hours post-ICH. Additionally, ferroptosis-related genes were distributed in the hippocampus and choroid plexus. We also elucidated a specific interaction between lipocalin-2 (LCN2)-positive microglia and oligodendrocytes that was mediated by the colony stimulating fac
- 218Single-cell tumor heterogeneity landscape of hepatocellular carcinoma: unraveling the pro-metastatic subtype and its interaction loop with fibroblasts.Background Tumor heterogeneity presents a formidable challenge in understanding the mechanisms driving tumor progression and metastasis. The heterogeneity of hepatocellular carcinoma (HCC) in cellular level is not clear. Methods Integration analysis of single-cell RNA sequencing data and spatial transcriptomics data was performed. Multiple methods were applied to investigate the subtype of HCC tumor cells. The functional characteristics, translation factors, clinical implications and microenvironment associations of different subtypes of tumor cells were analyzed. The interaction of subtype and fibroblasts were analyzed. Results We established a heterogeneity landscape of HCC malignant cells by integrated 52 single-cell RNA sequencing data and 5 spatial transcriptomics data. We identified three subtypes in tumor cells, including ARG1 + metabolism subtype (Metab-subtype), TOP2A + proliferation phenotype (Prol-phenotype), and S100A6 + pro-metastatic subtype (EMT-subtype). Enrichment anal
- 219Systematic dissection of tumor-normal single-cell ecosystems across a thousand tumors of 30 cancer types.The complexity of the tumor microenvironment poses significant challenges in cancer therapy. Here, to comprehensively investigate the tumor-normal ecosystems, we perform an integrative analysis of 4.9 million single-cell transcriptomes from 1070 tumor and 493 normal samples in combination with pan-cancer 137 spatial transcriptomics, 8887 TCGA, and 1261 checkpoint inhibitor-treated bulk tumors. We define a myriad of cell states constituting the tumor-normal ecosystems and also identify hallmark gene signatures across different cell types and organs. Our atlas characterizes distinctions between inflammatory fibroblasts marked by AKR1C1 or WNT5A in terms of cellular interactions and spatial co-localization patterns. Co-occurrence analysis reveals interferon-enriched community states including tertiary lymphoid structure (TLS) components, which exhibit differential rewiring between tumor, adjacent normal, and healthy normal tissues. The favorable response of interferon-enriched community s
- 220Single-cell and spatial transcriptomics reveal metastasis mechanism and microenvironment remodeling of lymph node in osteosarcoma.Background Osteosarcoma (OS) is the most common primary malignant bone tumor and is highly prone to metastasis. OS can metastasize to the lymph node (LN) through the lymphatics, and the metastasis of tumor cells reestablishes the immune landscape of the LN, which is conducive to the growth of tumor cells. However, the mechanism of LN metastasis of osteosarcoma and remodeling of the metastatic lymph node (MLN) microenvironment is not clear. Methods Single-cell RNA sequencing of 18 samples from paracancerous, primary tumor, and lymph nodes was performed. Then, new signaling axes closely related to metastasis were identified using bioinformatics, in vitro experiments, and immunohistochemistry. The mechanism of remodeling of the LN microenvironment in tumor cells was investigated by integrating single-cell and spatial transcriptomics. Results From 18 single-cell sequencing samples, we obtained 117,964 cells. The pseudotime analysis revealed that osteoblast(OB) cells may follow a differenti
- 221Domain generalization enables general cancer cell annotation in single-cell and spatial transcriptomics.Single-cell and spatial transcriptome sequencing, two recently optimized transcriptome sequencing methods, are increasingly used to study cancer and related diseases. Cell annotation, particularly for malignant cell annotation, is essential and crucial for in-depth analyses in these studies. However, current algorithms lack accuracy and generalization, making it difficult to consistently and rapidly infer malignant cells from pan-cancer data. To address this issue, we present Cancer-Finder, a domain generalization-based deep-learning algorithm that can rapidly identify malignant cells in single-cell data with an average accuracy of 95.16%. More importantly, by replacing the single-cell training data with spatial transcriptomic datasets, Cancer-Finder can accurately identify malignant spots on spatial slides. Applying Cancer-Finder to 5 clear cell renal cell carcinoma spatial transcriptomic samples, Cancer-Finder demonstrates a good ability to identify malignant spots and identifies a g
- 222Heterogeneous Skeletal Muscle Cell and Nucleus Populations Identified by Single-Cell and Single-Nucleus Resolution Transcriptome Assays.Single-cell RNA-seq (scRNA-seq) has revolutionized modern genomics, but the large size of myotubes and myofibers has restricted use of scRNA-seq in skeletal muscle. For the study of muscle, single-nucleus RNA-seq (snRNA-seq) has emerged not only as an alternative to scRNA-seq, but as a novel method providing valuable insights into multinucleated cells such as myofibers. Nuclei within myofibers specialize at junctions with other cell types such as motor neurons. Nuclear heterogeneity plays important roles in certain diseases such as muscular dystrophies. We survey current methods of high-throughput single cell and subcellular resolution transcriptomics, including single-cell and single-nucleus RNA-seq and spatial transcriptomics, applied to satellite cells, myoblasts, myotubes and myofibers. We summarize the major myonuclei subtypes identified in homeostatic and regenerating tissue including those specific to fiber type or at junctions with other cell types. Disease-specific nucleus pop
- 223Important Cells and Factors from Tumor Microenvironment Participated in Perineural Invasion.Perineural invasion (PNI) as the fourth way for solid tumors metastasis and invasion has attracted a lot of attention, recent research reported a new point that PNI starts to include axon growth and possible nerve "invasion" to tumors as the component. More and more tumor-nerve crosstalk has been explored to explain the internal mechanism for tumor microenvironment (TME) of some types of tumors tends to observe nerve infiltration. As is well known, the interaction of tumor cells, peripheral blood vessels, extracellular matrix, other non-malignant cells, and signal molecules in TME plays a key role in the occurrence, development, and metastasis of cancer, as to the occurrence and development of PNI. We aim to summarize the current theories on the molecular mediators and pathogenesis of PNI, add the latest scientific research progress, and explore the use of single-cell spatial transcriptomics in this invasion way. A better understanding of PNI may help to understand tumor metastasis and
- 224SCANPY: large-scale single-cell gene expression data analysis.SCANPY is a scalable toolkit for analyzing single-cell gene expression data. It includes methods for preprocessing, visualization, clustering, pseudotime and trajectory inference, differential expression testing, and simulation of gene regulatory networks. Its Python-based implementation efficiently deals with data sets of more than one million cells ( https://github.com/theislab/Scanpy ). Along with SCANPY, we present ANNDATA, a generic class for handling annotated data matrices ( https://github.com/theislab/anndata ).
- 225Massively parallel digital transcriptional profiling of single cells.Characterizing the transcriptome of individual cells is fundamental to understanding complex biological systems. We describe a droplet-based system that enables 3' mRNA counting of tens of thousands of single cells per sample. Cell encapsulation, of up to 8 samples at a time, takes place in ∼6 min, with ∼50% cell capture efficiency. To demonstrate the system's technical performance, we collected transcriptome data from ∼250k single cells across 29 samples. We validated the sensitivity of the system and its ability to detect rare populations using cell lines and synthetic RNAs. We profiled 68k peripheral blood mononuclear cells to demonstrate the system's ability to characterize large immune populations. Finally, we used sequence variation in the transcriptome data to determine host and donor chimerism at single-cell resolution from bone marrow mononuclear cells isolated from transplant patients.
- 226Molecular Architecture of the Mouse Nervous System.The mammalian nervous system executes complex behaviors controlled by specialized, precisely positioned, and interacting cell types. Here, we used RNA sequencing of half a million single cells to create a detailed census of cell types in the mouse nervous system. We mapped cell types spatially and derived a hierarchical, data-driven taxonomy. Neurons were the most diverse and were grouped by developmental anatomical units and by the expression of neurotransmitters and neuropeptides. Neuronal diversity was driven by genes encoding cell identity, synaptic connectivity, neurotransmission, and membrane conductance. We discovered seven distinct, regionally restricted astrocyte types that obeyed developmental boundaries and correlated with the spatial distribution of key glutamate and glycine neurotransmitters. In contrast, oligodendrocytes showed a loss of regional identity followed by a secondary diversification. The resource presented here lays a solid foundation for understanding the mol
- 227Slingshot: cell lineage and pseudotime inference for single-cell transcriptomics.Background Single-cell transcriptomics allows researchers to investigate complex communities of heterogeneous cells. It can be applied to stem cells and their descendants in order to chart the progression from multipotent progenitors to fully differentiated cells. While a variety of statistical and computational methods have been proposed for inferring cell lineages, the problem of accurately characterizing multiple branching lineages remains difficult to solve. Results We introduce Slingshot, a novel method for inferring cell lineages and pseudotimes from single-cell gene expression data. In previously published datasets, Slingshot correctly identifies the biological signal for one to three branching trajectories. Additionally, our simulation study shows that Slingshot infers more accurate pseudotimes than other leading methods. Conclusions Slingshot is a uniquely robust and flexible tool which combines the highly stable techniques necessary for noisy single-cell data with the ability
- 228UMI-tools: modeling sequencing errors in Unique Molecular Identifiers to improve quantification accuracy.Unique Molecular Identifiers (UMIs) are random oligonucleotide barcodes that are increasingly used in high-throughput sequencing experiments. Through a UMI, identical copies arising from distinct molecules can be distinguished from those arising through PCR amplification of the same molecule. However, bioinformatic methods to leverage the information from UMIs have yet to be formalized. In particular, sequencing errors in the UMI sequence are often ignored or else resolved in an ad hoc manner. We show that errors in the UMI sequence are common and introduce network-based methods to account for these errors when identifying PCR duplicates. Using these methods, we demonstrate improved quantification accuracy both under simulated conditions and real iCLIP and single-cell RNA-seq data sets. Reproducibility between iCLIP replicates and single-cell RNA-seq clustering are both improved using our proposed network-based method, demonstrating the value of properly accounting for errors in UMIs.
- 229SARS-CoV-2 Receptor ACE2 Is an Interferon-Stimulated Gene in Human Airway Epithelial Cells and Is Detected in Specific Cell Subsets across Tissues.There is pressing urgency to understand the pathogenesis of the severe acute respiratory syndrome coronavirus clade 2 (SARS-CoV-2), which causes the disease COVID-19. SARS-CoV-2 spike (S) protein binds angiotensin-converting enzyme 2 (ACE2), and in concert with host proteases, principally transmembrane serine protease 2 (TMPRSS2), promotes cellular entry. The cell subsets targeted by SARS-CoV-2 in host tissues and the factors that regulate ACE2 expression remain unknown. Here, we leverage human, non-human primate, and mouse single-cell RNA-sequencing (scRNA-seq) datasets across health and disease to uncover putative targets of SARS-CoV-2 among tissue-resident cell subsets. We identify ACE2 and TMPRSS2 co-expressing cells within lung type II pneumocytes, ileal absorptive enterocytes, and nasal goblet secretory cells. Strikingly, we discovered that ACE2 is a human interferon-stimulated gene (ISG) in vitro using airway epithelial cells and extend our findings to in vivo viral infections.
- 230CUT&Tag for efficient epigenomic profiling of small samples and single cells.Many chromatin features play critical roles in regulating gene expression. A complete understanding of gene regulation will require the mapping of specific chromatin features in small samples of cells at high resolution. Here we describe Cleavage Under Targets and Tagmentation (CUT&Tag), an enzyme-tethering strategy that provides efficient high-resolution sequencing libraries for profiling diverse chromatin components. In CUT&Tag, a chromatin protein is bound in situ by a specific antibody, which then tethers a protein A-Tn5 transposase fusion protein. Activation of the transposase efficiently generates fragment libraries with high resolution and exceptionally low background. All steps from live cells to sequencing-ready libraries can be performed in a single tube on the benchtop or a microwell in a high-throughput pipeline, and the entire procedure can be performed in one day. We demonstrate the utility of CUT&Tag by profiling histone modifications, RNA Polymerase II and transcription
- 231Single-cell RNA-seq data analysis on the receptor ACE2 expression reveals the potential risk of different human organs vulnerable to 2019-nCoV infection.It has been known that, the novel coronavirus, 2019-nCoV, which is considered similar to SARS-CoV, invades human cells via the receptor angiotensin converting enzyme II (ACE2). Moreover, lung cells that have ACE2 expression may be the main target cells during 2019-nCoV infection. However, some patients also exhibit non-respiratory symptoms, such as kidney failure, implying that 2019-nCoV could also invade other organs. To construct a risk map of different human organs, we analyzed the single-cell RNA sequencing (scRNA-seq) datasets derived from major human physiological systems, including the respiratory, cardiovascular, digestive, and urinary systems. Through scRNA-seq data analyses, we identified the organs at risk, such as lung, heart, esophagus, kidney, bladder, and ileum, and located specific cell types (i.e., type II alveolar cells (AT2), myocardial cells, proximal tubule cells of the kidney, ileum and esophagus epithelial cells, and bladder urothelial cells), which are vulnerabl
- 232Scater: pre-processing, quality control, normalization and visualization of single-cell RNA-seq data in R.Motivation Single-cell RNA sequencing (scRNA-seq) is increasingly used to study gene expression at the level of individual cells. However, preparing raw sequence data for further analysis is not a straightforward process. Biases, artifacts and other sources of unwanted variation are present in the data, requiring substantial time and effort to be spent on pre-processing, quality control (QC) and normalization. Results We have developed the R/Bioconductor package scater to facilitate rigorous pre-processing, quality control, normalization and visualization of scRNA-seq data. The package provides a convenient, flexible workflow to process raw sequencing reads into a high-quality expression dataset ready for downstream analysis. scater provides a rich suite of plotting tools for single-cell data and a flexible data structure that is compatible with existing tools and can be used as infrastructure for future software development. Availability and implementation The open-source code, along
- 233Cells of the adult human heart.Cardiovascular disease is the leading cause of death worldwide. Advanced insights into disease mechanisms and therapeutic strategies require a deeper understanding of the molecular processes involved in the healthy heart. Knowledge of the full repertoire of cardiac cells and their gene expression profiles is a fundamental first step in this endeavour. Here, using state-of-the-art analyses of large-scale single-cell and single-nucleus transcriptomes, we characterize six anatomical adult heart regions. Our results highlight the cellular heterogeneity of cardiomyocytes, pericytes and fibroblasts, and reveal distinct atrial and ventricular subsets of cells with diverse developmental origins and specialized properties. We define the complexity of the cardiac vasculature and its changes along the arterio-venous axis. In the immune compartment, we identify cardiac-resident macrophages with inflammatory and protective transcriptional signatures. Furthermore, analyses of cell-to-cell interactio
- 234SoupX removes ambient RNA contamination from droplet-based single-cell RNA sequencing data.Background Droplet-based single-cell RNA sequence analyses assume that all acquired RNAs are endogenous to cells. However, any cell-free RNAs contained within the input solution are also captured by these assays. This sequencing of cell-free RNA constitutes a background contamination that confounds the biological interpretation of single-cell transcriptomic data. Results We demonstrate that contamination from this "soup" of cell-free RNAs is ubiquitous, with experiment-specific variations in composition and magnitude. We present a method, SoupX, for quantifying the extent of the contamination and estimating "background-corrected" cell expression profiles that seamlessly integrate with existing downstream analysis tools. Applying this method to several datasets using multiple droplet sequencing technologies, we demonstrate that its application improves biological interpretation of otherwise misleading data, as well as improving quality control metrics. Conclusions We present SoupX, a to
- 235PAGA: graph abstraction reconciles clustering with trajectory inference through a topology preserving map of single cells.Single-cell RNA-seq quantifies biological heterogeneity across both discrete cell types and continuous cell transitions. Partition-based graph abstraction (PAGA) provides an interpretable graph-like map of the arising data manifold, based on estimating connectivity of manifold partitions ( https://github.com/theislab/paga ). PAGA maps preserve the global topology of data, allow analyzing data at different resolutions, and result in much higher computational efficiency of the typical exploratory data analysis workflow. We demonstrate the method by inferring structure-rich cell maps with consistent topology across four hematopoietic datasets, adult planaria and the zebrafish embryo and benchmark computational performance on one million neurons.
- 236PanglaoDB: a web server for exploration of mouse and human single-cell RNA sequencing data.Single-cell RNA sequencing is an increasingly used method to measure gene expression at the single cell level and build cell-type atlases of tissues. Hundreds of single-cell sequencing datasets have already been published. However, studies are frequently deposited as raw data, a format difficult to access for biological researchers due to the need for data processing using complex computational pipelines. We have implemented an online database, PanglaoDB, accessible through a user-friendly interface that can be used to explore published mouse and human single cell RNA sequencing studies. PanglaoDB contains pre-processed and pre-computed analyses from more than 1054 single-cell experiments covering most major single cell platforms and protocols, based on more than 4 million cells from a wide range of tissues and organs. The online interface allows users to query and explore cell types, genetic pathways and regulatory networks. In addition, we have established a community-curated cell-ty
- 237Single cell RNA sequencing of human liver reveals distinct intrahepatic macrophage populations.The liver is the largest solid organ in the body and is critical for metabolic and immune functions. However, little is known about the cells that make up the human liver and its immune microenvironment. Here we report a map of the cellular landscape of the human liver using single-cell RNA sequencing. We provide the transcriptional profiles of 8444 parenchymal and non-parenchymal cells obtained from the fractionation of fresh hepatic tissue from five human livers. Using gene expression patterns, flow cytometry, and immunohistochemical examinations, we identify 20 discrete cell populations of hepatocytes, endothelial cells, cholangiocytes, hepatic stellate cells, B cells, conventional and non-conventional T cells, NK-like cells, and distinct intrahepatic monocyte/macrophage populations. Together, our study presents a comprehensive view of the human liver at single-cell resolution that outlines the characteristics of resident cells in the liver, and in particular provides a map of the h
- 238Deep Profiling of Mouse Splenic Architecture with CODEX Multiplexed Imaging.A highly multiplexed cytometric imaging approach, termed co-detection by indexing (CODEX), is used here to create multiplexed datasets of normal and lupus (MRL/lpr) murine spleens. CODEX iteratively visualizes antibody binding events using DNA barcodes, fluorescent dNTP analogs, and an in situ polymerization-based indexing procedure. An algorithmic pipeline for single-cell antigen quantification in tightly packed tissues was developed and used to overlay well-known morphological features with de novo characterization of lymphoid tissue architecture at a single-cell and cellular neighborhood levels. We observed an unexpected, profound impact of the cellular neighborhood on the expression of protein receptors on immune cells. By comparing normal murine spleen to spleens from animals with systemic autoimmune disease (MRL/lpr), extensive and previously uncharacterized splenic cell-interaction dynamics in the healthy versus diseased state was observed. The fidelity of multiplexed spatial cy
- 239Single-cell RNA sequencing demonstrates the molecular and cellular reprogramming of metastatic lung adenocarcinoma.Advanced metastatic cancer poses utmost clinical challenges and may present molecular and cellular features distinct from an early-stage cancer. Herein, we present single-cell transcriptome profiling of metastatic lung adenocarcinoma, the most prevalent histological lung cancer type diagnosed at stage IV in over 40% of all cases. From 208,506 cells populating the normal tissues or early to metastatic stage cancer in 44 patients, we identify a cancer cell subtype deviating from the normal differentiation trajectory and dominating the metastatic stage. In all stages, the stromal and immune cell dynamics reveal ontological and functional changes that create a pro-tumoral and immunosuppressive microenvironment. Normal resident myeloid cell populations are gradually replaced with monocyte-derived macrophages and dendritic cells, along with T-cell exhaustion. This extensive single-cell analysis enhances our understanding of molecular and cellular dynamics in metastatic lung cancer and reveal
- 240CEL-Seq2: sensitive highly-multiplexed single-cell RNA-Seq.Single-cell transcriptomics requires a method that is sensitive, accurate, and reproducible. Here, we present CEL-Seq2, a modified version of our CEL-Seq method, with threefold higher sensitivity, lower costs, and less hands-on time. We implemented CEL-Seq2 on Fluidigm's C1 system, providing its first single-cell, on-chip barcoding method, and we detected gene expression changes accompanying the progression through the cell cycle in mouse fibroblast cells. We also compare with Smart-Seq to demonstrate CEL-Seq2's increased sensitivity relative to other available methods. Collectively, the improvements make CEL-Seq2 uniquely suited to single-cell RNA-Seq analysis in terms of economics, resolution, and ease of use.
- 241EmptyDrops: distinguishing cells from empty droplets in droplet-based single-cell RNA sequencing data.Droplet-based single-cell RNA sequencing protocols have dramatically increased the throughput of single-cell transcriptomics studies. A key computational challenge when processing these data is to distinguish libraries for real cells from empty droplets. Here, we describe a new statistical method for calling cells from droplet-based data, based on detecting significant deviations from the expression profile of the ambient solution. Using simulations, we demonstrate that EmptyDrops has greater power than existing approaches while controlling the false discovery rate among detected cells. Our method also retains distinct cell types that would have been discarded by existing methods in several real data sets.
- 242Pooling across cells to normalize single-cell RNA sequencing data with many zero counts.Normalization of single-cell RNA sequencing data is necessary to eliminate cell-specific biases prior to downstream analyses. However, this is not straightforward for noisy single-cell data where many counts are zero. We present a novel approach where expression values are summed across pools of cells, and the summed values are used for normalization. Pool-based size factors are then deconvolved to yield cell-based factors. Our deconvolution approach outperforms existing methods for accurate normalization of cell-specific biases in simulated data. Similar behavior is observed in real data, where deconvolution improves the relevance of results of downstream analyses.
- 243Clustering trees: a visualization for evaluating clusterings at multiple resolutions.Clustering techniques are widely used in the analysis of large datasets to group together samples with similar properties. For example, clustering is often used in the field of single-cell RNA-sequencing in order to identify different cell types present in a tissue sample. There are many algorithms for performing clustering, and the results can vary substantially. In particular, the number of groups present in a dataset is often unknown, and the number of clusters identified by an algorithm can change based on the parameters used. To explore and examine the impact of varying clustering resolution, we present clustering trees. This visualization shows the relationships between clusters at multiple resolutions, allowing researchers to see how samples move as the number of clusters increases. In addition, meta-information can be overlaid on the tree to inform the choice of resolution and guide in identification of clusters. We illustrate the features of clustering trees using a series of
- 244Multimodal Analysis of Composition and Spatial Architecture in Human Squamous Cell Carcinoma.To define the cellular composition and architecture of cutaneous squamous cell carcinoma (cSCC), we combined single-cell RNA sequencing with spatial transcriptomics and multiplexed ion beam imaging from a series of human cSCCs and matched normal skin. cSCC exhibited four tumor subpopulations, three recapitulating normal epidermal states, and a tumor-specific keratinocyte (TSK) population unique to cancer, which localized to a fibrovascular niche. Integration of single-cell and spatial data mapped ligand-receptor networks to specific cell types, revealing TSK cells as a hub for intercellular communication. Multiple features of potential immunosuppression were observed, including T regulatory cell (Treg) co-localization with CD8 T cells in compartmentalized tumor stroma. Finally, single-cell characterization of human tumor xenografts and in vivo CRISPR screens identified essential roles for specific tumor subpopulation-enriched gene networks in tumorigenesis. These data define cSCC tumor
- 245Bulk tissue cell type deconvolution with multi-subject single-cell expression reference.Knowledge of cell type composition in disease relevant tissues is an important step towards the identification of cellular targets of disease. We present MuSiC, a method that utilizes cell-type specific gene expression from single-cell RNA sequencing (RNA-seq) data to characterize cell type compositions from bulk RNA-seq data in complex tissues. By appropriate weighting of genes showing cross-subject and cross-cell consistency, MuSiC enables the transfer of cell type-specific gene expression information from one dataset to another. When applied to pancreatic islet and whole kidney expression data in human, mouse, and rats, MuSiC outperformed existing methods, especially for tissues with closely related cell types. MuSiC enables the characterization of cellular heterogeneity of complex tissues for understanding of disease mechanisms. As bulk tissue data are more easily accessible than single-cell RNA-seq, MuSiC allows the utilization of the vast amounts of disease relevant bulk tissue R
- 246Single-cell RNA-seq enables comprehensive tumour and immune cell profiling in primary breast cancer.Single-cell transcriptome profiling of tumour tissue isolates allows the characterization of heterogeneous tumour cells along with neighbouring stromal and immune cells. Here we adopt this powerful approach to breast cancer and analyse 515 cells from 11 patients. Inferred copy number variations from the single-cell RNA-seq data separate carcinoma cells from non-cancer cells. At a single-cell resolution, carcinoma cells display common signatures within the tumour as well as intratumoral heterogeneity regarding breast cancer subtype and crucial cancer-related pathways. Most of the non-cancer cells are immune cells, with three distinct clusters of T lymphocytes, B lymphocytes and macrophages. T lymphocytes and macrophages both display immunosuppressive characteristics: T cells with a regulatory or an exhausted phenotype and macrophages with an M2 phenotype. These results illustrate that the breast cancer transcriptome has a wide range of intratumoral heterogeneity, which is shaped by the
- 247Single-cell RNA-seq denoising using a deep count autoencoder.Single-cell RNA sequencing (scRNA-seq) has enabled researchers to study gene expression at a cellular resolution. However, noise due to amplification and dropout may obstruct analyses, so scalable denoising methods for increasingly large but sparse scRNA-seq data are needed. We propose a deep count autoencoder network (DCA) to denoise scRNA-seq datasets. DCA takes the count distribution, overdispersion and sparsity of the data into account using a negative binomial noise model with or without zero-inflation, and nonlinear gene-gene dependencies are captured. Our method scales linearly with the number of cells and can, therefore, be applied to datasets of millions of cells. We demonstrate that DCA denoising improves a diverse set of typical scRNA-seq data analyses using simulated and real datasets. DCA outperforms existing methods for data imputation in quality and speed, enhancing biological discovery.
- 248Single-Cell RNA-Seq Reveals Lineage and X Chromosome Dynamics in Human Preimplantation Embryos.Mouse studies have been instrumental in forming our current understanding of early cell-lineage decisions; however, similar insights into the early human development are severely limited. Here, we present a comprehensive transcriptional map of human embryo development, including the sequenced transcriptomes of 1,529 individual cells from 88 human preimplantation embryos. These data show that cells undergo an intermediate state of co-expression of lineage-specific genes, followed by a concurrent establishment of the trophectoderm, epiblast, and primitive endoderm lineages, which coincide with blastocyst formation. Female cells of all three lineages achieve dosage compensation of X chromosome RNA levels prior to implantation. However, in contrast to the mouse, XIST is transcribed from both alleles throughout the progression of this expression dampening, and X chromosome genes maintain biallelic expression while dosage compensation proceeds. We envision broad utility of this transcription
- 249BBKNN: fast batch alignment of single cell transcriptomes.Motivation Increasing numbers of large scale single cell RNA-Seq projects are leading to a data explosion, which can only be fully exploited through data integration. A number of methods have been developed to combine diverse datasets by removing technical batch effects, but most are computationally intensive. To overcome the challenge of enormous datasets, we have developed BBKNN, an extremely fast graph-based data integration algorithm. We illustrate the power of BBKNN on large scale mouse atlasing data, and favourably benchmark its run time against a number of competing methods. Availability and implementation BBKNN is available at https://github.com/Teichlab/bbknn, along with documentation and multiple example notebooks, and can be installed from pip. Supplementary information Supplementary data are available at Bioinformatics online.
- 250Molecular Diversity of Midbrain Development in Mouse, Human, and Stem Cells.Understanding human embryonic ventral midbrain is of major interest for Parkinson's disease. However, the cell types, their gene expression dynamics, and their relationship to commonly used rodent models remain to be defined. We performed single-cell RNA sequencing to examine ventral midbrain development in human and mouse. We found 25 molecularly defined human cell types, including five subtypes of radial glia-like cells and four progenitors. In the mouse, two mature fetal dopaminergic neuron subtypes diversified into five adult classes during postnatal development. Cell types and gene expression were generally conserved across species, but with clear differences in cell proliferation, developmental timing, and dopaminergic neuron development. Additionally, we developed a method to quantitatively assess the fidelity of dopaminergic neurons derived from human pluripotent stem cells, at a single-cell level. Thus, our study provides insight into the molecular programs controlling human m
- 251Splatter: simulation of single-cell RNA sequencing data.As single-cell RNA sequencing (scRNA-seq) technologies have rapidly developed, so have analysis methods. Many methods have been tested, developed, and validated using simulated datasets. Unfortunately, current simulations are often poorly documented, their similarity to real data is not demonstrated, or reproducible code is not available. Here, we present the Splatter Bioconductor package for simple, reproducible, and well-documented simulation of scRNA-seq data. Splatter provides an interface to multiple simulation methods including Splat, our own simulation, based on a gamma-Poisson distribution. Splat can simulate single populations of cells, populations with multiple cell types, or differentiation paths.
- 252Identification of region-specific astrocyte subtypes at single cell resolution.Astrocytes, a major cell type found throughout the central nervous system, have general roles in the modulation of synapse formation and synaptic transmission, blood-brain barrier formation, and regulation of blood flow, as well as metabolic support of other brain resident cells. Crucially, emerging evidence shows specific adaptations and astrocyte-encoded functions in regions, such as the spinal cord and cerebellum. To investigate the true extent of astrocyte molecular diversity across forebrain regions, we used single-cell RNA sequencing. Our analysis identifies five transcriptomically distinct astrocyte subtypes in adult mouse cortex and hippocampus. Validation of our data in situ reveals distinct spatial positioning of defined subtypes, reflecting the distribution of morphologically and physiologically distinct astrocyte populations. Our findings are evidence for specialized astrocyte subtypes between and within brain regions. The data are available through an online database (http
- 253Trajectory-based differential expression analysis for single-cell sequencing data.Trajectory inference has radically enhanced single-cell RNA-seq research by enabling the study of dynamic changes in gene expression. Downstream of trajectory inference, it is vital to discover genes that are (i) associated with the lineages in the trajectory, or (ii) differentially expressed between lineages, to illuminate the underlying biological processes. Current data analysis procedures, however, either fail to exploit the continuous resolution provided by trajectory inference, or fail to pinpoint the exact types of differential expression. We introduce tradeSeq, a powerful generalized additive model framework based on the negative binomial distribution that allows flexible inference of both within-lineage and between-lineage differential expression. By incorporating observation-level weights, the model additionally allows to account for zero inflation. We evaluate the method on simulated datasets and on real datasets from droplet-based and full-length protocols, and show that it
- 254Single-cell transcriptomics of human T cells reveals tissue and activation signatures in health and disease.Human T cells coordinate adaptive immunity in diverse anatomic compartments through production of cytokines and effector molecules, but it is unclear how tissue site influences T cell persistence and function. Here, we use single cell RNA-sequencing (scRNA-seq) to define the heterogeneity of human T cells isolated from lungs, lymph nodes, bone marrow and blood, and their functional responses following stimulation. Through analysis of >50,000 resting and activated T cells, we reveal tissue T cell signatures in mucosal and lymphoid sites, and lineage-specific activation states across all sites including distinct effector states for CD8 + T cells and an interferon-response state for CD4 + T cells. Comparing scRNA-seq profiles of tumor-associated T cells to our dataset reveals predominant activated CD8 + compared to CD4 + T cell states within multiple tumor types. Our results therefore establish a high dimensional reference map of human T cell activation in health for analyzing T cells in
- 255SCoPE-MS: mass spectrometry of single mammalian cells quantifies proteome heterogeneity during cell differentiation.Some exciting biological questions require quantifying thousands of proteins in single cells. To achieve this goal, we develop Single Cell ProtEomics by Mass Spectrometry (SCoPE-MS) and validate its ability to identify distinct human cancer cell types based on their proteomes. We use SCoPE-MS to quantify over a thousand proteins in differentiating mouse embryonic stem cells. The single-cell proteomes enable us to deconstruct cell populations and infer protein abundance relationships. Comparison between single-cell proteomes and transcriptomes indicates coordinated mRNA and protein covariation, yet many genes exhibit functionally concerted and distinct regulatory patterns at the mRNA and the protein level.
- 256Spatially and functionally distinct subclasses of breast cancer-associated fibroblasts revealed by single cell RNA sequencing.Cancer-associated fibroblasts (CAFs) are a major constituent of the tumor microenvironment, although their origin and roles in shaping disease initiation, progression and treatment response remain unclear due to significant heterogeneity. Here, following a negative selection strategy combined with single-cell RNA sequencing of 768 transcriptomes of mesenchymal cells from a genetically engineered mouse model of breast cancer, we define three distinct subpopulations of CAFs. Validation at the transcriptional and protein level in several experimental models of cancer and human tumors reveal spatial separation of the CAF subclasses attributable to different origins, including the peri-vascular niche, the mammary fat pad and the transformed epithelium. Gene profiles for each CAF subtype correlate to distinctive functional programs and hold independent prognostic capability in clinical cohorts by association to metastatic disease. In conclusion, the improved resolution of the widely defined
- 257The art of using t-SNE for single-cell transcriptomics.Single-cell transcriptomics yields ever growing data sets containing RNA expression levels for thousands of genes from up to millions of cells. Common data analysis pipelines include a dimensionality reduction step for visualising the data in two dimensions, most frequently performed using t-distributed stochastic neighbour embedding (t-SNE). It excels at revealing local structure in high-dimensional data, but naive applications often suffer from severe shortcomings, e.g. the global structure of the data is not represented accurately. Here we describe how to circumvent such pitfalls, and develop a protocol for creating more faithful t-SNE visualisations. It includes PCA initialisation, a high learning rate, and multi-scale similarity kernels; for very large data sets, we additionally use exaggeration and downsampling-based initialisation. We use published single-cell RNA-seq data sets to demonstrate that this protocol yields superior results compared to the naive application of t-SNE.
- 258Single-cell RNA-seq: advances and future challenges.Phenotypically identical cells can dramatically vary with respect to behavior during their lifespan and this variation is reflected in their molecular composition such as the transcriptomic landscape. Single-cell transcriptomics using next-generation transcript sequencing (RNA-seq) is now emerging as a powerful tool to profile cell-to-cell variability on a genomic scale. Its application has already greatly impacted our conceptual understanding of diverse biological processes with broad implications for both basic and clinical research. Different single-cell RNA-seq protocols have been introduced and are reviewed here-each one with its own strengths and current limitations. We further provide an overview of the biological questions single-cell RNA-seq has been used to address, the major findings obtained from such studies, and current challenges and expected future developments in this booming field.
- 259Single cell RNA sequencing of human microglia uncovers a subset associated with Alzheimer's disease.The extent of microglial heterogeneity in humans remains a central yet poorly explored question in light of the development of therapies targeting this cell type. Here, we investigate the population structure of live microglia purified from human cerebral cortex samples obtained at autopsy and during neurosurgical procedures. Using single cell RNA sequencing, we find that some subsets are enriched for disease-related genes and RNA signatures. We confirm the presence of four of these microglial subpopulations histologically and illustrate the utility of our data by characterizing further microglial cluster 7, enriched for genes depleted in the cortex of individuals with Alzheimer's disease (AD). Histologically, these cluster 7 microglia are reduced in frequency in AD tissue, and we validate this observation in an independent set of single nucleus data. Thus, our live human microglia identify a range of subtypes, and we prioritize one of these as being altered in AD.
- 260Immune cell profiling of COVID-19 patients in the recovery stage by single-cell sequencing.COVID-19, caused by SARS-CoV-2, has recently affected over 1,200,000 people and killed more than 60,000. The key immune cell subsets change and their states during the course of COVID-19 remain unclear. We sought to comprehensively characterize the transcriptional changes in peripheral blood mononuclear cells during the recovery stage of COVID-19 by single-cell RNA sequencing technique. It was found that T cells decreased remarkably, whereas monocytes increased in patients in the early recovery stage (ERS) of COVID-19. There was an increased ratio of classical CD14 ++ monocytes with high inflammatory gene expression as well as a greater abundance of CD14 ++ IL1β + monocytes in the ERS. CD4 + T cells and CD8 + T cells decreased significantly and expressed high levels of inflammatory genes in the ERS. Among the B cells, the plasma cells increased remarkably, whereas the naïve B cells decreased. Several novel B cell-receptor (BCR) changes were identified, such as IGHV3-23 and IGHV3-7, and
- 261Gene expression markers of Tumor Infiltrating Leukocytes.Background Assays of the abundance of immune cell populations in the tumor microenvironment promise to inform immune oncology research and the choice of immunotherapy for individual patients. We propose to measure the intratumoral abundance of various immune cell populations with gene expression. In contrast to IHC and flow cytometry, gene expression assays yield high information content from a clinically practical workflow. Previous studies of gene expression in purified immune cells have reported hundreds of genes showing enrichment in a single cell type, but the utility of these genes in tumor samples is unknown. We use co-expression patterns in large tumor gene expression datasets to evaluate previously reported candidate cell type marker genes lists, eliminate numerous false positives and identify a subset of high confidence marker genes. Methods Using a novel statistical tool, we use co-expression patterns in 9986 samples from The Cancer Genome Atlas (TCGA) to evaluate previously
- 262Structural Remodeling of the Human Colonic Mesenchyme in Inflammatory Bowel Disease.Intestinal mesenchymal cells play essential roles in epithelial homeostasis, matrix remodeling, immunity, and inflammation. But the extent of heterogeneity within the colonic mesenchyme in these processes remains unknown. Using unbiased single-cell profiling of over 16,500 colonic mesenchymal cells, we reveal four subsets of fibroblasts expressing divergent transcriptional regulators and functional pathways, in addition to pericytes and myofibroblasts. We identified a niche population located in proximity to epithelial crypts expressing SOX6, F3 (CD142), and WNT genes essential for colonic epithelial stem cell function. In colitis, we observed dysregulation of this niche and emergence of an activated mesenchymal population. This subset expressed TNF superfamily member 14 (TNFSF14), fibroblastic reticular cell-associated genes, IL-33, and Lysyl oxidases. Further, it induced factors that impaired epithelial proliferation and maturation and contributed to oxidative stress and disease seve
- 263Therapy-Induced Evolution of Human Lung Cancer Revealed by Single-Cell RNA Sequencing.Lung cancer, the leading cause of cancer mortality, exhibits heterogeneity that enables adaptability, limits therapeutic success, and remains incompletely understood. Single-cell RNA sequencing (scRNA-seq) of metastatic lung cancer was performed using 49 clinical biopsies obtained from 30 patients before and during targeted therapy. Over 20,000 cancer and tumor microenvironment (TME) single-cell profiles exposed a rich and dynamic tumor ecosystem. scRNA-seq of cancer cells illuminated targetable oncogenes beyond those detected clinically. Cancer cells surviving therapy as residual disease (RD) expressed an alveolar-regenerative cell signature suggesting a therapy-induced primitive cell-state transition, whereas those present at on-therapy progressive disease (PD) upregulated kynurenine, plasminogen, and gap-junction pathways. Active T-lymphocytes and decreased macrophages were present at RD and immunosuppressive cell states characterized PD. Biological features revealed by scRNA-seq we
- 264Highly multiplexed immunofluorescence imaging of human tissues and tumors using t-CyCIF and conventional optical microscopes.The architecture of normal and diseased tissues strongly influences the development and progression of disease as well as responsiveness and resistance to therapy. We describe a tissue-based cyclic immunofluorescence (t-CyCIF) method for highly multiplexed immuno-fluorescence imaging of formalin-fixed, paraffin-embedded (FFPE) specimens mounted on glass slides, the most widely used specimens for histopathological diagnosis of cancer and other diseases. t-CyCIF generates up to 60-plex images using an iterative process (a cycle) in which conventional low-plex fluorescence images are repeatedly collected from the same sample and then assembled into a high-dimensional representation. t-CyCIF requires no specialized instruments or reagents and is compatible with super-resolution imaging; we demonstrate its application to quantifying signal transduction cascades, tumor antigens and immune markers in diverse tissues and tumors. The simplicity and adaptability of t-CyCIF makes it an effective
- 265Single-cell analysis uncovers fibroblast heterogeneity and criteria for fibroblast and mural cell identification and discrimination.Many important cell types in adult vertebrates have a mesenchymal origin, including fibroblasts and vascular mural cells. Although their biological importance is undisputed, the level of mesenchymal cell heterogeneity within and between organs, while appreciated, has not been analyzed in detail. Here, we compare single-cell transcriptional profiles of fibroblasts and vascular mural cells across four murine muscular organs: heart, skeletal muscle, intestine and bladder. We reveal gene expression signatures that demarcate fibroblasts from mural cells and provide molecular signatures for cell subtype identification. We observe striking inter- and intra-organ heterogeneity amongst the fibroblasts, primarily reflecting differences in the expression of extracellular matrix components. Fibroblast subtypes localize to discrete anatomical positions offering novel predictions about physiological function(s) and regulatory signaling circuits. Our data shed new light on the diversity of poorly def
- 266Cell Types of the Human Retina and Its Organoids at Single-Cell Resolution.Human organoids recapitulating the cell-type diversity and function of their target organ are valuable for basic and translational research. We developed light-sensitive human retinal organoids with multiple nuclear and synaptic layers and functional synapses. We sequenced the RNA of 285,441 single cells from these organoids at seven developmental time points and from the periphery, fovea, pigment epithelium and choroid of light-responsive adult human retinas, and performed histochemistry. Cell types in organoids matured in vitro to a stable "developed" state at a rate similar to human retina development in vivo. Transcriptomes of organoid cell types converged toward the transcriptomes of adult peripheral retinal cell types. Expression of disease-associated genes was cell-type-specific in adult retina, and cell-type specificity was retained in organoids. We implicate unexpected cell types in diseases such as macular degeneration. This resource identifies cellular targets for studying d
- 267Decontamination of ambient RNA in single-cell RNA-seq with DecontX.Droplet-based microfluidic devices have become widely used to perform single-cell RNA sequencing (scRNA-seq). However, ambient RNA present in the cell suspension can be aberrantly counted along with a cell's native mRNA and result in cross-contamination of transcripts between different cell populations. DecontX is a novel Bayesian method to estimate and remove contamination in individual cells. DecontX accurately predicts contamination levels in a mouse-human mixture dataset and removes aberrant expression of marker genes in PBMC datasets. We also compare the contamination levels between four different scRNA-seq protocols. Overall, DecontX can be incorporated into scRNA-seq workflows to improve downstream analyses.
- 268Classification of low quality cells from single-cell RNA-seq data.Single-cell RNA sequencing (scRNA-seq) has broad applications across biomedical research. One of the key challenges is to ensure that only single, live cells are included in downstream analysis, as the inclusion of compromised cells inevitably affects data interpretation. Here, we present a generic approach for processing scRNA-seq data and detecting low quality cells, using a curated set of over 20 biological and technical features. Our approach improves classification accuracy by over 30 % compared to traditional methods when tested on over 5,000 cells, including CD4+ T cells, bone marrow dendritic cells, and mouse embryonic stem cells.
- 269Single-cell profiling of human gliomas reveals macrophage ontogeny as a basis for regional differences in macrophage activation in the tumor microenvironment.Background Tumor-associated macrophages (TAMs) are abundant in gliomas and immunosuppressive TAMs are a barrier to emerging immunotherapies. It is unknown to what extent macrophages derived from peripheral blood adopt the phenotype of brain-resident microglia in pre-treatment gliomas. The relative proportions of blood-derived macrophages and microglia have been poorly quantified in clinical samples due to a paucity of markers that distinguish these cell types in malignant tissue. Results We perform single-cell RNA-sequencing of human gliomas and identify phenotypic differences in TAMs of distinct lineages. We isolate TAMs from patient biopsies and compare them with macrophages from non-malignant human tissue, glioma atlases, and murine glioma models. We present a novel signature that distinguishes TAMs by ontogeny in human gliomas. Blood-derived TAMs upregulate immunosuppressive cytokines and show an altered metabolism compared to microglial TAMs. They are also enriched in perivascular
- 270Identifying gene expression programs of cell-type identity and cellular activity with single-cell RNA-Seq.Identifying gene expression programs underlying both cell-type identity and cellular activities (e.g. life-cycle processes, responses to environmental cues) is crucial for understanding the organization of cells and tissues. Although single-cell RNA-Seq (scRNA-Seq) can quantify transcripts in individual cells, each cell's expression profile may be a mixture of both types of programs, making them difficult to disentangle. Here, we benchmark and enhance the use of matrix factorization to solve this problem. We show with simulations that a method we call consensus non-negative matrix factorization (cNMF) accurately infers identity and activity programs, including their relative contributions in each cell. To illustrate the insights this approach enables, we apply it to published brain organoid and visual cortex scRNA-Seq datasets; cNMF refines cell types and identifies both expected (e.g. cell cycle and hypoxia) and novel activity programs, including programs that may underlie a neurosecr
- 271Single-cell RNA sequencing highlights the role of inflammatory cancer-associated fibroblasts in bladder urothelial carcinoma.Although substantial progress has been made in cancer biology and treatment, clinical outcomes of bladder carcinoma (BC) patients are still not satisfactory. The tumor microenvironment (TME) is a potential target. Here, by single-cell RNA sequencing on 8 BC tumor samples and 3 para tumor samples, we identify 19 different cell types in the BC microenvironment, indicating high intra-tumoral heterogeneity. We find that tumor cells down regulated MHC-II molecules, suggesting that the downregulated immunogenicity of cancer cells may contribute to the formation of an immunosuppressive microenvironment. We also find that monocytes undergo M2 polarization in the tumor region and differentiate. Furthermore, the LAMP3 + DC subgroup may be able to recruit regulatory T cells, potentially taking part in the formation of an immunosuppressive TME. Through correlation analysis using public datasets containing over 3000 BC samples, we identify a role for inflammatory cancer-associated fibroblasts (iCAF
- 272Single-nucleus and single-cell transcriptomes compared in matched cortical cell types.Transcriptomic profiling of complex tissues by single-nucleus RNA-sequencing (snRNA-seq) affords some advantages over single-cell RNA-sequencing (scRNA-seq). snRNA-seq provides less biased cellular coverage, does not appear to suffer cell isolation-based transcriptional artifacts, and can be applied to archived frozen specimens. We used well-matched snRNA-seq and scRNA-seq datasets from mouse visual cortex to compare cell type detection. Although more transcripts are detected in individual whole cells (~11,000 genes) than nuclei (~7,000 genes), we demonstrate that closely related neuronal cell types can be similarly discriminated with both methods if intronic sequences are included in snRNA-seq analysis. We estimate that the nuclear proportion of total cellular mRNA varies from 20% to over 50% for large and small pyramidal neurons, respectively. Together, these results illustrate the high information content of nuclear RNA for characterization of cellular diversity in brain tissues.
- 273Single-cell RNA landscape of intratumoral heterogeneity and immunosuppressive microenvironment in advanced osteosarcoma.Osteosarcoma is the most frequent primary bone tumor with poor prognosis. Through RNA-sequencing of 100,987 individual cells from 7 primary, 2 recurrent, and 2 lung metastatic osteosarcoma lesions, 11 major cell clusters are identified based on unbiased clustering of gene expression profiles and canonical markers. The transcriptomic properties, regulators and dynamics of osteosarcoma malignant cells together with their tumor microenvironment particularly stromal and immune cells are characterized. The transdifferentiation of malignant osteoblastic cells from malignant chondroblastic cells is revealed by analyses of inferred copy-number variation and trajectory. A proinflammatory FABP4 + macrophages infiltration is noticed in lung metastatic osteosarcoma lesions. Lower osteoclasts infiltration is observed in chondroblastic, recurrent and lung metastatic osteosarcoma lesions compared to primary osteoblastic osteosarcoma lesions. Importantly, TIGIT blockade enhances the cytotoxicity effec
- 274Highly multiplexed imaging of single cells using a high-throughput cyclic immunofluorescence method.Single-cell analysis reveals aspects of cellular physiology not evident from population-based studies, particularly in the case of highly multiplexed methods such as mass cytometry (CyTOF) able to correlate the levels of multiple signalling, differentiation and cell fate markers. Immunofluorescence (IF) microscopy adds information on cell morphology and the microenvironment that are not obtained using flow-based techniques, but the multiplicity of conventional IF is limited. This has motivated development of imaging methods that require specialized instrumentation, exotic reagents or proprietary protocols that are difficult to reproduce in most laboratories. Here we report a public-domain method for achieving high multiplicity single-cell IF using cyclic immunofluorescence (CycIF), a simple and versatile procedure in which four-colour staining alternates with chemical inactivation of fluorophores to progressively build a multichannel image. Because CycIF uses standard reagents and inst
- 275TSCAN: Pseudo-time reconstruction and evaluation in single-cell RNA-seq analysis.When analyzing single-cell RNA-seq data, constructing a pseudo-temporal path to order cells based on the gradual transition of their transcriptomes is a useful way to study gene expression dynamics in a heterogeneous cell population. Currently, a limited number of computational tools are available for this task, and quantitative methods for comparing different tools are lacking. Tools for Single Cell Analysis (TSCAN) is a software tool developed to better support in silico pseudo-Time reconstruction in Single-Cell RNA-seq ANalysis. TSCAN uses a cluster-based minimum spanning tree (MST) approach to order cells. Cells are first grouped into clusters and an MST is then constructed to connect cluster centers. Pseudo-time is obtained by projecting each cell onto the tree, and the ordered sequence of cells can be used to study dynamic changes of gene expression along the pseudo-time. Clustering cells before MST construction reduces the complexity of the tree space. This often leads to improv
- 276ArrayExpress update - from bulk to single-cell expression data.ArrayExpress (https://www.ebi.ac.uk/arrayexpress) is an archive of functional genomics data from a variety of technologies assaying functional modalities of a genome, such as gene expression or promoter occupancy. The number of experiments based on sequencing technologies, in particular RNA-seq experiments, has been increasing over the last few years and submissions of sequencing data have overtaken microarray experiments in the last 12 months. Additionally, there is a significant increase in experiments investigating single cells, rather than bulk samples, known as single-cell RNA-seq. To accommodate these trends, we have substantially changed our submission tool Annotare which, along with raw and processed data, collects all metadata necessary to interpret these experiments. Selected datasets are re-processed and loaded into our sister resource, the value-added Expression Atlas (and its component Single Cell Expression Atlas), which not only enables users to interpret the data easily
- 277Microanatomy of the Human Atherosclerotic Plaque by Single-Cell Transcriptomics.Rationale Atherosclerotic lesions are known for their cellular heterogeneity, yet the molecular complexity within the cells of human plaques has not been fully assessed. Objective Using single-cell transcriptomics and chromatin accessibility, we gained a better understanding of the pathophysiology underlying human atherosclerosis. Methods and results We performed single-cell RNA and single-cell ATAC sequencing on human carotid atherosclerotic plaques to define the cells at play and determine their transcriptomic and epigenomic characteristics. We identified 14 distinct cell populations including endothelial cells, smooth muscle cells, mast cells, B cells, myeloid cells, and T cells and identified multiple cellular activation states and suggested cellular interconversions. Within the endothelial cell population, we defined subsets with angiogenic capacity plus clear signs of endothelial to mesenchymal transition. CD4 + and CD8 + T cells showed activation-based subclasses, each with a gr
- 278Single-cell triple omics sequencing reveals genetic, epigenetic, and transcriptomic heterogeneity in hepatocellular carcinomas.Single-cell genome, DNA methylome, and transcriptome sequencing methods have been separately developed. However, to accurately analyze the mechanism by which transcriptome, genome and DNA methylome regulate each other, these omic methods need to be performed in the same single cell. Here we demonstrate a single-cell triple omics sequencing technique, scTrio-seq, that can be used to simultaneously analyze the genomic copy-number variations (CNVs), DNA methylome, and transcriptome of an individual mammalian cell. We show that large-scale CNVs cause proportional changes in RNA expression of genes within the gained or lost genomic regions, whereas these CNVs generally do not affect DNA methylation in these regions. Furthermore, we applied scTrio-seq to 25 single cancer cells derived from a human hepatocellular carcinoma tissue sample. We identified two subpopulations within these cells based on CNVs, DNA methylome, or transcriptome of individual cells. Our work offers a new avenue of disse
- 279Single-cell RNA-seq reveals that glioblastoma recapitulates a normal neurodevelopmental hierarchy.Cancer stem cells are critical for cancer initiation, development, and treatment resistance. Our understanding of these processes, and how they relate to glioblastoma heterogeneity, is limited. To overcome these limitations, we performed single-cell RNA sequencing on 53586 adult glioblastoma cells and 22637 normal human fetal brain cells, and compared the lineage hierarchy of the developing human brain to the transcriptome of cancer cells. We find a conserved neural tri-lineage cancer hierarchy centered around glial progenitor-like cells. We also find that this progenitor population contains the majority of the cancer's cycling cells, and, using RNA velocity, is often the originator of the other cell types. Finally, we show that this hierarchal map can be used to identify therapeutic targets specific to progenitor cancer stem cells. Our analyses show that normal brain development reconciles glioblastoma development, suggests a possible origin for glioblastoma hierarchy, and helps to id
- 280Single-cell expression profiling reveals dynamic flux of cardiac stromal, vascular and immune cells in health and injury.Besides cardiomyocytes (CM), the heart contains numerous interstitial cell types which play key roles in heart repair, regeneration and disease, including fibroblast, vascular and immune cells. However, a comprehensive understanding of this interactive cell community is lacking. We performed single-cell RNA-sequencing of the total non-CM fraction and enriched ( Pdgfra -GFP + ) fibroblast lineage cells from murine hearts at days 3 and 7 post-sham or myocardial infarction (MI) surgery. Clustering of >30,000 single cells identified >30 populations representing nine cell lineages, including a previously undescribed fibroblast lineage trajectory present in both sham and MI hearts leading to a uniquely activated cell state defined in part by a strong anti-WNT transcriptome signature. We also uncovered novel myofibroblast subtypes expressing either pro-fibrotic or anti-fibrotic signatures. Our data highlight non-linear dynamics in myeloid and fibroblast lineages after cardiac injury, and prov
- 281Single-cell transcriptomes identify human islet cell signatures and reveal cell-type-specific expression changes in type 2 diabetes.Blood glucose levels are tightly controlled by the coordinated action of at least four cell types constituting pancreatic islets. Changes in the proportion and/or function of these cells are associated with genetic and molecular pathophysiology of monogenic, type 1, and type 2 (T2D) diabetes. Cellular heterogeneity impedes precise understanding of the molecular components of each islet cell type that govern islet (dys)function, particularly the less abundant delta and gamma/pancreatic polypeptide (PP) cells. Here, we report single-cell transcriptomes for 638 cells from nondiabetic (ND) and T2D human islet samples. Analyses of ND single-cell transcriptomes identified distinct alpha, beta, delta, and PP/gamma cell-type signatures. Genes linked to rare and common forms of islet dysfunction and diabetes were expressed in the delta and PP/gamma cell types. Moreover, this study revealed that delta cells specifically express receptors that receive and coordinate systemic cues from the leptin,
- 282Single-cell analysis reveals fibroblast heterogeneity and myeloid-derived adipocyte progenitors in murine skin wounds.During wound healing in adult mouse skin, hair follicles and then adipocytes regenerate. Adipocytes regenerate from myofibroblasts, a specialized contractile wound fibroblast. Here we study wound fibroblast diversity using single-cell RNA-sequencing. On analysis, wound fibroblasts group into twelve clusters. Pseudotime and RNA velocity analyses reveal that some clusters likely represent consecutive differentiation states toward a contractile phenotype, while others appear to represent distinct fibroblast lineages. One subset of fibroblasts expresses hematopoietic markers, suggesting their myeloid origin. We validate this finding using single-cell western blot and single-cell RNA-sequencing on genetically labeled myofibroblasts. Using bone marrow transplantation and Cre recombinase-based lineage tracing experiments, we rule out cell fusion events and confirm that hematopoietic lineage cells give rise to a subset of myofibroblasts and rare regenerated adipocytes. In conclusion, our study
- 283Single-cell transcriptomes of the human skin reveal age-related loss of fibroblast priming.Fibroblasts are an essential cell population for human skin architecture and function. While fibroblast heterogeneity is well established, this phenomenon has not been analyzed systematically yet. We have used single-cell RNA sequencing to analyze the transcriptomes of more than 5,000 fibroblasts from a sun-protected area in healthy human donors. Our results define four main subpopulations that can be spatially localized and show differential secretory, mesenchymal and pro-inflammatory functional annotations. Importantly, we found that this fibroblast 'priming' becomes reduced with age. We also show that aging causes a substantial reduction in the predicted interactions between dermal fibroblasts and other skin cells, including undifferentiated keratinocytes at the dermal-epidermal junction. Our work thus provides evidence for a functional specialization of human dermal fibroblasts and identifies the partial loss of cellular identity as an important age-related change in the human derm
- 284Single-nucleus transcriptome analysis reveals dysregulation of angiogenic endothelial cells and neuroprotective glia in Alzheimer's disease.Alzheimer's disease (AD) is the most common form of dementia but has no effective treatment. A comprehensive investigation of cell type-specific responses and cellular heterogeneity in AD is required to provide precise molecular and cellular targets for therapeutic development. Accordingly, we perform single-nucleus transcriptome analysis of 169,496 nuclei from the prefrontal cortical samples of AD patients and normal control (NC) subjects. Differential analysis shows that the cell type-specific transcriptomic changes in AD are associated with the disruption of biological processes including angiogenesis, immune activation, synaptic signaling, and myelination. Subcluster analysis reveals that compared to NC brains, AD brains contain fewer neuroprotective astrocytes and oligodendrocytes. Importantly, our findings show that a subpopulation of angiogenic endothelial cells is induced in the brain in patients with AD. These angiogenic endothelial cells exhibit increased expression of angiog
- 285muscat detects subpopulation-specific state transitions from multi-sample multi-condition single-cell transcriptomics data.Single-cell RNA sequencing (scRNA-seq) has become an empowering technology to profile the transcriptomes of individual cells on a large scale. Early analyses of differential expression have aimed at identifying differences between subpopulations to identify subpopulation markers. More generally, such methods compare expression levels across sets of cells, thus leading to cross-condition analyses. Given the emergence of replicated multi-condition scRNA-seq datasets, an area of increasing focus is making sample-level inferences, termed here as differential state analysis; however, it is not clear which statistical framework best handles this situation. Here, we surveyed methods to perform cross-condition differential state analyses, including cell-level mixed models and methods based on aggregated pseudobulk data. To evaluate method performance, we developed a flexible simulation that mimics multi-sample scRNA-seq data. We analyzed scRNA-seq data from mouse cortex cells to uncover subpop
- 286Metabolic landscape of the tumor microenvironment at single cell resolution.The tumor milieu consists of numerous cell types each existing in a different environment. However, a characterization of metabolic heterogeneity at single-cell resolution is not established. Here, we develop a computational pipeline to study metabolic programs in single cells. In two representative human cancers, melanoma and head and neck, we apply this algorithm to define the intratumor metabolic landscape. We report an overall discordance between analyses of single cells and those of bulk tumors with higher metabolic activity in malignant cells than previously appreciated. Variation in mitochondrial programs is found to be the major contributor to metabolic heterogeneity. Surprisingly, the expression of both glycolytic and mitochondrial programs strongly correlates with hypoxia in all cell types. Immune and stromal cells could also be distinguished by their metabolic features. Taken together this analysis establishes a computational framework for characterizing metabolism using sin
- 287An Optimized Shotgun Strategy for the Rapid Generation of Comprehensive Human Proteomes.This study investigates the challenge of comprehensively cataloging the complete human proteome from a single-cell type using mass spectrometry (MS)-based shotgun proteomics. We modify a classical two-dimensional high-resolution reversed-phase peptide fractionation scheme and optimize a protocol that provides sufficient peak capacity to saturate the sequencing speed of modern MS instruments. This strategy enables the deepest proteome of a human single-cell type to date, with the HeLa proteome sequenced to a depth of ∼584,000 unique peptide sequences and ∼14,200 protein isoforms (∼12,200 protein-coding genes). This depth is comparable with next-generation RNA sequencing and enables the identification of post-translational modifications, including ∼7,000 N-acetylation sites and ∼10,000 phosphorylation sites, without the need for enrichment. We further demonstrate the general applicability and clinical potential of this proteomics strategy by comprehensively quantifying global proteome ex
- 288Gene Regulatory Network Inference from Single-Cell Data Using Multivariate Information Measures.While single-cell gene expression experiments present new challenges for data processing, the cell-to-cell variability observed also reveals statistical relationships that can be used by information theory. Here, we use multivariate information theory to explore the statistical dependencies between triplets of genes in single-cell gene expression datasets. We develop PIDC, a fast, efficient algorithm that uses partial information decomposition (PID) to identify regulatory relationships between genes. We thoroughly evaluate the performance of our algorithm and demonstrate that the higher-order information captured by PIDC allows it to outperform pairwise mutual information-based algorithms when recovering true relationships present in simulated data. We also infer gene regulatory networks from three experimental single-cell datasets and illustrate how network context, choices made during analysis, and sources of variability affect network inference. PIDC tutorials and open-source softwa
- 289scRNA-seq Profiling of Human Testes Reveals the Presence of the ACE2 Receptor, A Target for SARS-CoV-2 Infection in Spermatogonia, Leydig and Sertoli Cells.In December 2019, a novel coronavirus (SARS-CoV-2) was identified in COVID-19 patients in Wuhan, Hubei Province, China. SARS-CoV-2 shares both high sequence similarity and the use of the same cell entry receptor, angiotensin-converting enzyme 2 (ACE2), with severe acute respiratory syndrome coronavirus (SARS-CoV). Several studies have provided bioinformatic evidence of potential routes of SARS-CoV-2 infection in respiratory, cardiovascular, digestive and urinary systems. However, whether the reproductive system is a potential target of SARS-CoV-2 infection has not yet been determined. Here, we investigate the expression pattern of ACE2 in adult human testes at the level of single-cell transcriptomes. The results indicate that ACE2 is predominantly enriched in spermatogonia and Leydig and Sertoli cells. Gene Set Enrichment Analysis (GSEA) indicates that Gene Ontology (GO) categories associated with viral reproduction and transmission are highly enriched in ACE2-positive spermatogonia, w
- 290Unravelling subclonal heterogeneity and aggressive disease states in TNBC through single-cell RNA-seq.Triple-negative breast cancer (TNBC) is an aggressive subtype characterized by extensive intratumoral heterogeneity. To investigate the underlying biology, we conducted single-cell RNA-sequencing (scRNA-seq) of >1500 cells from six primary TNBC. Here, we show that intercellular heterogeneity of gene expression programs within each tumor is variable and largely correlates with clonality of inferred genomic copy number changes, suggesting that genotype drives the gene expression phenotype of individual subpopulations. Clustering of gene expression profiles identified distinct subgroups of malignant cells shared by multiple tumors, including a single subpopulation associated with multiple signatures of treatment resistance and metastasis, and characterized functionally by activation of glycosphingolipid metabolism and associated innate immunity pathways. A novel signature defining this subpopulation predicts long-term outcomes for TNBC patients in a large cohort. Collectively, this analys