专题电子书
表观组学循证手册
按质量评分、证据类型和发表时间组织;优先收录明确允许商业复用的来源文献。
72 个章节
- 01Identifying gene targets for brain-related traits using transcriptomic and methylomic data from blood.Understanding the difference in genetic regulation of gene expression between brain and blood is important for discovering genes for brain-related traits and disorders. Here, we estimate the correlation of genetic effects at the top-associated cis-expression or -DNA methylation (DNAm) quantitative trait loci (cis-eQTLs or cis-mQTLs) between brain and blood (r b ). Using publicly available data, we find that genetic effects at the top cis-eQTLs or mQTLs are highly correlated between independent brain and blood samples ([Formula: see text] for cis-eQTLs and [Formula: see text] for cis-mQTLs). Using meta-analyzed brain cis-eQTL/mQTL data (n = 526 to 1194), we identify 61 genes and 167 DNAm sites associated with four brain-related phenotypes, most of which are a subset of the discoveries (97 genes and 295 DNAm sites) using data from blood with larger sample sizes (n = 1980 to 14,115). Our results demonstrate the gain of power in gene discovery for brain-related phenotypes using blood cis-e
- 02DNA methylation signatures of chronic low-grade inflammation are associated with complex diseases.Background Chronic low-grade inflammation reflects a subclinical immune response implicated in the pathogenesis of complex diseases. Identifying genetic loci where DNA methylation is associated with chronic low-grade inflammation may reveal novel pathways or therapeutic targets for inflammation. Results We performed a meta-analysis of epigenome-wide association studies (EWAS) of serum C-reactive protein (CRP), which is a sensitive marker of low-grade inflammation, in a large European population (n = 8863) and trans-ethnic replication in African Americans (n = 4111). We found differential methylation at 218 CpG sites to be associated with CRP (P < 1.15 × 10 -7 ) in the discovery panel of European ancestry and replicated (P < 2.29 × 10 -4 ) 58 CpG sites (45 unique loci) among African Americans. To further characterize the molecular and clinical relevance of the findings, we examined the association with gene expression, genetic sequence variants, and clinical outcomes. DNA methylation at
- 03Maternal plasma folate impacts differential DNA methylation in an epigenome-wide meta-analysis of newborns.Folate is vital for fetal development. Periconceptional folic acid supplementation and food fortification are recommended to prevent neural tube defects. Mechanisms whereby periconceptional folate influences normal development and disease are poorly understood: epigenetics may be involved. We examine the association between maternal plasma folate during pregnancy and epigenome-wide DNA methylation using Illumina's HumanMethyl450 Beadchip in 1,988 newborns from two European cohorts. Here we report the combined covariate-adjusted results using meta-analysis and employ pathway and gene expression analyses. Four-hundred forty-three CpGs (320 genes) are significantly associated with maternal plasma folate levels during pregnancy (false discovery rate 5%); 48 are significant after Bonferroni correction. Most genes are not known for folate biology, including APC2, GRM8, SLC16A12, OPCML, PRPH, LHX1, KLK4 and PRSS21. Some relate to birth defects other than neural tube defects, neurological func
- 04Maternal BMI at the start of pregnancy and offspring epigenome-wide DNA methylation: findings from the pregnancy and childhood epigenetics (PACE) consortium.Pre-pregnancy maternal obesity is associated with adverse offspring outcomes at birth and later in life. Individual studies have shown that epigenetic modifications such as DNA methylation could contribute. Within the Pregnancy and Childhood Epigenetics (PACE) Consortium, we meta-analysed the association between pre-pregnancy maternal BMI and methylation at over 450,000 sites in newborn blood DNA, across 19 cohorts (9,340 mother-newborn pairs). We attempted to infer causality by comparing the effects of maternal versus paternal BMI and incorporating genetic variation. In four additional cohorts (1,817 mother-child pairs), we meta-analysed the association between maternal BMI at the start of pregnancy and blood methylation in adolescents. In newborns, maternal BMI was associated with small (<0.2% per BMI unit (1 kg/m2), P < 1.06 × 10-7) methylation variation at 9,044 sites throughout the genome. Adjustment for estimated cell proportions greatly attenuated the number of significant CpGs
- 05Epigenome-Wide Meta-Analysis of Methylation in Children Related to Prenatal NO2 Air Pollution Exposure.Background Prenatal exposure to air pollution is considered to be associated with adverse effects on child health. This may partly be mediated by mechanisms related to DNA methylation. Objectives We investigated associations between exposure to air pollution, using nitrogen dioxide (NO2) as marker, and epigenome-wide cord blood DNA methylation. Methods We meta-analyzed the associations between NO2 exposure at residential addresses during pregnancy and cord blood DNA methylation (Illumina 450K) in four European and North American studies (n = 1,508) with subsequent look-up analyses in children ages 4 (n = 733) and 8 (n = 786) years. Additionally, we applied a literature-based candidate approach for antioxidant and anti-inflammatory genes. To assess influence of exposure at the transcriptomics level, we related mRNA expression in blood cells to NO2 exposure in 4- (n = 111) and 16-year-olds (n = 239). Results We found epigenome-wide significant associations [false discovery rate (FDR) p <
- 06Guidelines for bioinformatics of single-cell sequencing data analysis in Alzheimer's disease: review, recommendation, implementation and application.Alzheimer's disease (AD) is the most common form of dementia, characterized by progressive cognitive impairment and neurodegeneration. Extensive clinical and genomic studies have revealed biomarkers, risk factors, pathways, and targets of AD in the past decade. However, the exact molecular basis of AD development and progression remains elusive. The emerging single-cell sequencing technology can potentially provide cell-level insights into the disease. Here we systematically review the state-of-the-art bioinformatics approaches to analyze single-cell sequencing data and their applications to AD in 14 major directions, including 1) quality control and normalization, 2) dimension reduction and feature extraction, 3) cell clustering analysis, 4) cell type inference and annotation, 5) differential expression, 6) trajectory inference, 7) copy number variation analysis, 8) integration of single-cell multi-omics, 9) epigenomic analysis, 10) gene network inference, 11) prioritization of cell s
- 07DNA methylation-calling tools for Oxford Nanopore sequencing: a survey and human epigenome-wide evaluation.Background Nanopore long-read sequencing technology greatly expands the capacity of long-range, single-molecule DNA-modification detection. A growing number of analytical tools have been developed to detect DNA methylation from nanopore sequencing reads. Here, we assess the performance of different methylation-calling tools to provide a systematic evaluation to guide researchers performing human epigenome-wide studies. Results We compare seven analytic tools for detecting DNA methylation from nanopore long-read sequencing data generated from human natural DNA at a whole-genome scale. We evaluate the per-read and per-site performance of CpG methylation prediction across different genomic contexts, CpG site coverage, and computational resources consumed by each tool. The seven tools exhibit different performances across the evaluation criteria. We show that the methylation prediction at regions with discordant DNA methylation patterns, intergenic regions, low CG density regions, and repe
- 08A multimodal cell census and atlas of the mammalian primary motor cortex.Here we report the generation of a multimodal cell census and atlas of the mammalian primary motor cortex as the initial product of the BRAIN Initiative Cell Census Network (BICCN). This was achieved by coordinated large-scale analyses of single-cell transcriptomes, chromatin accessibility, DNA methylomes, spatially resolved single-cell transcriptomes, morphological and electrophysiological properties and cellular resolution input-output mapping, integrated through cross-modal computational analysis. Our results advance the collective knowledge and understanding of brain cell-type organization 1-5 . First, our study reveals a unified molecular genetic landscape of cortical cell types that integrates their transcriptome, open chromatin and DNA methylation maps. Second, cross-species analysis achieves a consensus taxonomy of transcriptomic types and their hierarchical organization that is conserved from mouse to marmoset and human. Third, in situ single-cell transcriptomics provides a sp
- 09A transcriptomic and epigenomic cell atlas of the mouse primary motor cortex.Single-cell transcriptomics can provide quantitative molecular signatures for large, unbiased samples of the diverse cell types in the brain 1-3 . With the proliferation of multi-omics datasets, a major challenge is to validate and integrate results into a biological understanding of cell-type organization. Here we generated transcriptomes and epigenomes from more than 500,000 individual cells in the mouse primary motor cortex, a structure that has an evolutionarily conserved role in locomotion. We developed computational and statistical methods to integrate multimodal data and quantitatively validate cell-type reproducibility. The resulting reference atlas-containing over 56 neuronal cell types that are highly replicable across analysis methods, sequencing technologies and modalities-is a comprehensive molecular and genomic account of the diverse neuronal and non-neuronal cell types in the mouse primary motor cortex. The atlas includes a population of excitatory neurons that resemble
- 10ChIP-Atlas 2021 update: a data-mining suite for exploring epigenomic landscapes by fully integrating ChIP-seq, ATAC-seq and Bisulfite-seq data.ChIP-Atlas (https://chip-atlas.org) is a web service providing both GUI- and API-based data-mining tools to reveal the architecture of the transcription regulatory landscape. ChIP-Atlas is powered by comprehensively integrating all data sets from high-throughput ChIP-seq and DNase-seq, a method for profiling chromatin regions accessible to DNase. In this update, we further collected all the ATAC-seq and whole-genome bisulfite-seq data for six model organisms (human, mouse, rat, fruit fly, nematode, and budding yeast) with the latest genome assemblies. These together with ChIP-seq data can be visualized with the Peak Browser tool and a genome browser to explore the epigenomic landscape of a query genomic locus, such as its chromatin accessibility, DNA methylation status, and protein-genome interactions. This epigenomic landscape can also be characterized for multiple genes and genomic loci by querying with the Enrichment Analysis tool, which, for example, revealed that inflammatory bowe
- 11DNA methylation atlas of the mouse brain at single-cell resolution.Mammalian brain cells show remarkable diversity in gene expression, anatomy and function, yet the regulatory DNA landscape underlying this extensive heterogeneity is poorly understood. Here we carry out a comprehensive assessment of the epigenomes of mouse brain cell types by applying single-nucleus DNA methylation sequencing 1,2 to profile 103,982 nuclei (including 95,815 neurons and 8,167 non-neuronal cells) from 45 regions of the mouse cortex, hippocampus, striatum, pallidum and olfactory areas. We identified 161 cell clusters with distinct spatial locations and projection targets. We constructed taxonomies of these epigenetic types, annotated with signature genes, regulatory elements and transcription factors. These features indicate the potential regulatory landscape supporting the assignment of putative cell types and reveal repetitive usage of regulators in excitatory and inhibitory cells for determining subtypes. The DNA methylation landscape of excitatory neurons in the cortex
- 12The EWAS Catalog: a database of epigenome-wide association studies.Epigenome-wide association studies (EWAS) seek to quantify associations between traits/exposures and DNA methylation measured at thousands or millions of CpG sites across the genome. In recent years, the increase in availability of DNA methylation measures in population-based cohorts and case-control studies has resulted in a dramatic expansion of the number of EWAS being performed and published. To make this rich source of results more accessible, we have manually curated a database of CpG-trait associations (with p<1x10 -4 ) from published EWAS, each assaying over 100,000 CpGs in at least 100 individuals. From January 7, 2022, The EWAS Catalog contained 1,737,746 associations from 2,686 EWAS. This includes 1,345,398 associations from 342 peer-reviewed publications. In addition, it also contains summary statistics for 392,348 associations from 427 EWAS, performed on data from the Avon Longitudinal Study of Parents and Children (ALSPAC) and the Gene Expression Omnibus (GEO). The databa
- 13A single cell atlas of human cornea that defines its development, limbal progenitor cells and their interactions with the immune cells.Purpose Single cell (sc) analyses of key embryonic, fetal and adult stages were performed to generate a comprehensive single cell atlas of all the corneal and adjacent conjunctival cell types from development to adulthood. Methods Four human adult and seventeen embryonic and fetal corneas from 10 to 21 post conception week (PCW) specimens were dissociated to single cells and subjected to scRNA- and/or ATAC-Seq using the 10x Genomics platform. These were embedded using Uniform Manifold Approximation and Projection (UMAP) and clustered using Seurat graph-based clustering. Cluster identification was performed based on marker gene expression, bioinformatic data mining and immunofluorescence (IF) analysis. RNA interference, IF, colony forming efficiency and clonal assays were performed on cultured limbal epithelial cells (LECs). Results scRNA-Seq analysis of 21,343 cells from four adult human corneas and adjacent conjunctivas revealed the presence of 21 cell clusters, representing the proge
- 14Practical guidelines for the comprehensive analysis of ChIP-seq data.Mapping the chromosomal locations of transcription factors, nucleosomes, histone modifications, chromatin remodeling enzymes, chaperones, and polymerases is one of the key tasks of modern biology, as evidenced by the Encyclopedia of DNA Elements (ENCODE) Project. To this end, chromatin immunoprecipitation followed by high-throughput sequencing (ChIP-seq) is the standard methodology. Mapping such protein-DNA interactions in vivo using ChIP-seq presents multiple challenges not only in sample preparation and sequencing but also for computational analysis. Here, we present step-by-step guidelines for the computational analysis of ChIP-seq data. We address all the major steps in the analysis of ChIP-seq data: sequencing depth selection, quality checking, mapping, data normalization, assessment of reproducibility, peak calling, differential binding analysis, controlling the false discovery rate, peak annotation, visualization, and motif analysis. At each step in our guidelines we discuss som
- 15jMorp updates in 2020: large enhancement of multi-omics data resources on the general Japanese population.In the Tohoku Medical Megabank project, genome and omics analyses of participants in two cohort studies were performed. A part of the data is available at the Japanese Multi Omics Reference Panel (jMorp; https://jmorp.megabank.tohoku.ac.jp) as a web-based database, as reported in our previous manuscript published in Nucleic Acid Research in 2018. At that time, jMorp mainly consisted of metabolome data; however, now genome, methylome, and transcriptome data have been integrated in addition to the enhancement of the number of samples for the metabolome data. For genomic data, jMorp provides a Japanese reference sequence obtained using de novo assembly of sequences from three Japanese individuals and allele frequencies obtained using whole-genome sequencing of 8,380 Japanese individuals. In addition, the omics data include methylome and transcriptome data from ∼300 samples and distribution of concentrations of more than 755 metabolites obtained using high-throughput nuclear magnetic reson
- 16Revolutionizing Personalized Medicine: Synergy with Multi-Omics Data Generation, Main Hurdles, and Future Perspectives.The field of personalized medicine is undergoing a transformative shift through the integration of multi-omics data, which mainly encompasses genomics, transcriptomics, proteomics, and metabolomics. This synergy allows for a comprehensive understanding of individual health by analyzing genetic, molecular, and biochemical profiles. The generation and integration of multi-omics data enable more precise and tailored therapeutic strategies, improving the efficacy of treatments and reducing adverse effects. However, several challenges hinder the full realization of personalized medicine. Key hurdles include the complexity of data integration across different omics layers, the need for advanced computational tools, and the high cost of comprehensive data generation. Additionally, issues related to data privacy, standardization, and the need for robust validation in diverse populations remain significant obstacles. Looking ahead, the future of personalized medicine promises advancements in te
- 17scNMT-seq enables joint profiling of chromatin accessibility DNA methylation and transcription in single cells.Parallel single-cell sequencing protocols represent powerful methods for investigating regulatory relationships, including epigenome-transcriptome interactions. Here, we report a single-cell method for parallel chromatin accessibility, DNA methylation and transcriptome profiling. scNMT-seq (single-cell nucleosome, methylation and transcription sequencing) uses a GpC methyltransferase to label open chromatin followed by bisulfite and RNA sequencing. We validate scNMT-seq by applying it to differentiating mouse embryonic stem cells, finding links between all three molecular layers and revealing dynamic coupling between epigenomic layers during differentiation.
- 18Identification of transcription factor binding sites using ATAC-seq.Transposase-Accessible Chromatin followed by sequencing (ATAC-seq) is a simple protocol for detection of open chromatin. Computational footprinting, the search for regions with depletion of cleavage events due to transcription factor binding, is poorly understood for ATAC-seq. We propose the first footprinting method considering ATAC-seq protocol artifacts. HINT-ATAC uses a position dependency model to learn the cleavage preferences of the transposase. We observe strand-specific cleavage patterns around transcription factor binding sites, which are determined by local nucleosome architecture. By incorporating all these biases, HINT-ATAC is able to significantly outperform competing methods in the prediction of transcription factor binding sites with footprints.
- 19Assessment of computational methods for the analysis of single-cell ATAC-seq data.Background Recent innovations in single-cell Assay for Transposase Accessible Chromatin using sequencing (scATAC-seq) enable profiling of the epigenetic landscape of thousands of individual cells. scATAC-seq data analysis presents unique methodological challenges. scATAC-seq experiments sample DNA, which, due to low copy numbers (diploid in humans), lead to inherent data sparsity (1-10% of peaks detected per cell) compared to transcriptomic (scRNA-seq) data (10-45% of expressed genes detected per cell). Such challenges in data generation emphasize the need for informative features to assess cell heterogeneity at the chromatin level. Results We present a benchmarking framework that is applied to 10 computational methods for scATAC-seq on 13 synthetic and real datasets from different assays, profiling cell types from diverse tissues and organisms. Methods for processing and featurizing scATAC-seq data were compared by their ability to discriminate cell types when combined with common uns
- 20Integrative analyses of single-cell transcriptome and regulome using MAESTRO.We present Model-based AnalysEs of Transcriptome and RegulOme (MAESTRO), a comprehensive open-source computational workflow ( http://github.com/liulab-dfci/MAESTRO ) for the integrative analyses of single-cell RNA-seq (scRNA-seq) and ATAC-seq (scATAC-seq) data from multiple platforms. MAESTRO provides functions for pre-processing, alignment, quality control, expression and chromatin accessibility quantification, clustering, differential analysis, and annotation. By modeling gene regulatory potential from chromatin accessibilities at the single-cell level, MAESTRO outperforms the existing methods for integrating the cell clusters between scRNA-seq and scATAC-seq. Furthermore, MAESTRO supports automatic cell-type annotation using predefined cell type marker genes and identifies driver regulators from differential scRNA-seq genes and scATAC-seq peaks.
- 21ChIP-Atlas: a data-mining suite powered by full integration of public ChIP-seq data.We have fully integrated public chromatin chromatin immunoprecipitation sequencing (ChIP-seq) and DNase-seq data ( n > 70,000) derived from six representative model organisms (human, mouse, rat, fruit fly, nematode, and budding yeast), and have devised a data-mining platform-designated ChIP-Atlas (http://chip-atlas.org). ChIP-Atlas is able to show alignment and peak-call results for all public ChIP-seq and DNase-seq data archived in the NCBI Sequence Read Archive (SRA), which encompasses data derived from GEO, ArrayExpress, DDBJ, ENCODE, Roadmap Epigenomics, and the scientific literature. All peak-call data are integrated to visualize multiple histone modifications and binding sites of transcriptional regulators (TRs) at given genomic loci. The integrated data can be further analyzed to show TR-gene and TR-TR interactions, as well as to examine enrichment of protein binding for given multiple genomic coordinates or gene names. ChIP-Atlas is superior to other platforms in terms of data
- 22An atlas of dynamic chromatin landscapes in mouse fetal development.The Encyclopedia of DNA Elements (ENCODE) project has established a genomic resource for mammalian development, profiling a diverse panel of mouse tissues at 8 developmental stages from 10.5 days after conception until birth, including transcriptomes, methylomes and chromatin states. Here we systematically examined the state and accessibility of chromatin in the developing mouse fetus. In total we performed 1,128 chromatin immunoprecipitation with sequencing (ChIP-seq) assays for histone modifications and 132 assay for transposase-accessible chromatin using sequencing (ATAC-seq) assays for chromatin accessibility across 72 distinct tissue-stages. We used integrative analysis to develop a unified set of chromatin state annotations, infer the identities of dynamic enhancers and key transcriptional regulators, and characterize the relationship between chromatin state and accessibility during developmental gene regulation. We also leveraged these data to link enhancers to putative target g
- 23GTRD: a database on gene transcription regulation-2019 update.The current version of the Gene Transcription Regulation Database (GTRD; http://gtrd.biouml.org) contains information about: (i) transcription factor binding sites (TFBSs) and transcription coactivators identified by ChIP-seq experiments for Homo sapiens, Mus musculus, Rattus norvegicus, Danio rerio, Caenorhabditis elegans, Drosophila melanogaster, Saccharomyces cerevisiae, Schizosaccharomyces pombe and Arabidopsis thaliana; (ii) regions of open chromatin and TFBSs (DNase footprints) identified by DNase-seq; (iii) unmappable regions where TFBSs cannot be identified due to repeats; (iv) potential TFBSs for both human and mouse using position weight matrices from the HOCOMOCO database. Raw ChIP-seq and DNase-seq data were obtained from ENCODE and SRA, and uniformly processed. ChIP-seq peaks were called using four different methods: MACS, SISSRs, GEM and PICS. Moreover, peaks for the same factor and peak calling method, albeit using different experiment conditions (cell line, treatment, e
- 24GTRD: a database of transcription factor binding sites identified by ChIP-seq experiments.GTRD-Gene Transcription Regulation Database (http://gtrd.biouml.org)-is a database of transcription factor binding sites (TFBSs) identified by ChIP-seq experiments for human and mouse. Raw ChIP-seq data were obtained from ENCODE and SRA and uniformly processed: (i) reads were aligned using Bowtie2; (ii) ChIP-seq peaks were called using peak callers MACS, SISSRs, GEM and PICS; (iii) peaks for the same factor and peak callers, but different experiment conditions (cell line, treatment, etc.), were merged into clusters; (iv) such clusters for different peak callers were merged into metaclusters that were considered as non-redundant sets of TFBSs. In addition to information on location in genome, the sets contain structured information about cell lines and experimental conditions extracted from descriptions of corresponding ChIP-seq experiments. A web interface to access GTRD was developed using the BioUML platform. It provides: (i) browsing and displaying information; (ii) advanced search
- 25Atlas of quantitative single-base-resolution N 6 -methyl-adenine methylomes.Various methyltransferases and demethylases catalyse methylation and demethylation of N 6 -methyladenosine (m6A) and N 6 ,2'-O-dimethyladenosine (m6Am) but precise methylomes uniquely mediated by each methyltransferase/demethylase are still lacking. Here, we develop m6A-Crosslinking-Exonuclease-sequencing (m6ACE-seq) to map transcriptome-wide m6A and m6Am at quantitative single-base-resolution. This allows for the generation of a comprehensive atlas of distinct methylomes uniquely mediated by every individual known methyltransferase or demethylase. Our atlas reveals METTL16 to indirectly impact manifold methylation targets beyond its consensus target motif and highlights the importance of precision in mapping PCIF1-dependent m6Am. Rather than reverse RNA methylation, we find that both ALKBH5 and FTO instead maintain their regulated sites in an unmethylated steady-state. In FTO's absence, anomalous m6Am disrupts snRNA interaction with nuclear export machinery, potentially causing aberra
- 26Multi-Omics Pipeline and Omics-Integration Approach to Decipher Plant's Abiotic Stress Tolerance Responses.The present day's ongoing global warming and climate change adversely affect plants through imposing environmental (abiotic) stresses and disease pressure. The major abiotic factors such as drought, heat, cold, salinity, etc., hamper a plant's innate growth and development, resulting in reduced yield and quality, with the possibility of undesired traits. In the 21st century, the advent of high-throughput sequencing tools, state-of-the-art biotechnological techniques and bioinformatic analyzing pipelines led to the easy characterization of plant traits for abiotic stress response and tolerance mechanisms by applying the 'omics' toolbox. Panomics pipeline including genomics, transcriptomics, proteomics, metabolomics, epigenomics, proteogenomics, interactomics, ionomics, phenomics, etc., have become very handy nowadays. This is important to produce climate-smart future crops with a proper understanding of the molecular mechanisms of abiotic stress responses by the plant's genes, transcrip
- 27Single Cell Atlas: a single-cell multi-omics human cell encyclopedia.Single-cell sequencing datasets are key in biology and medicine for unraveling insights into heterogeneous cell populations with unprecedented resolution. Here, we construct a single-cell multi-omics map of human tissues through in-depth characterizations of datasets from five single-cell omics, spatial transcriptomics, and two bulk omics across 125 healthy adult and fetal tissues. We construct its complement web-based platform, the Single Cell Atlas (SCA, www.singlecellatlas.org ), to enable vast interactive data exploration of deep multi-omics signatures across human fetal and adult tissues. The atlas resources and database queries aspire to serve as a one-stop, comprehensive, and time-effective resource for various omics studies.
- 28Comparative cellular analysis of motor cortex in human, marmoset and mouse.The primary motor cortex (M1) is essential for voluntary fine-motor control and is functionally conserved across mammals 1 . Here, using high-throughput transcriptomic and epigenomic profiling of more than 450,000 single nuclei in humans, marmoset monkeys and mice, we demonstrate a broadly conserved cellular makeup of this region, with similarities that mirror evolutionary distance and are consistent between the transcriptome and epigenome. The core conserved molecular identities of neuronal and non-neuronal cell types allow us to generate a cross-species consensus classification of cell types, and to infer conserved properties of cell types across species. Despite the overall conservation, however, many species-dependent specializations are apparent, including differences in cell-type proportions, gene expression, DNA methylation and chromatin state. Few cell-type marker genes are conserved across species, revealing a short list of candidate genes and regulatory mechanisms that are re
- 29Single cell transcriptional and chromatin accessibility profiling redefine cellular heterogeneity in the adult human kidney.The integration of single cell transcriptome and chromatin accessibility datasets enables a deeper understanding of cell heterogeneity. We performed single nucleus ATAC (snATAC-seq) and RNA (snRNA-seq) sequencing to generate paired, cell-type-specific chromatin accessibility and transcriptional profiles of the adult human kidney. We demonstrate that snATAC-seq is comparable to snRNA-seq in the assignment of cell identity and can further refine our understanding of functional heterogeneity in the nephron. The majority of differentially accessible chromatin regions are localized to promoters and a significant proportion are closely associated with differentially expressed genes. Cell-type-specific enrichment of transcription factor binding motifs implicates the activation of NF-κB that promotes VCAM1 expression and drives transition between a subpopulation of proximal tubule epithelial cells. Our multi-omics approach improves the ability to detect unique cell states within the kidney and
- 30A comparison of reference-based algorithms for correcting cell-type heterogeneity in Epigenome-Wide Association Studies.Background Intra-sample cellular heterogeneity presents numerous challenges to the identification of biomarkers in large Epigenome-Wide Association Studies (EWAS). While a number of reference-based deconvolution algorithms have emerged, their potential remains underexplored and a comparative evaluation of these algorithms beyond tissues such as blood is still lacking. Results Here we present a novel framework for reference-based inference, which leverages cell-type specific DNAse Hypersensitive Site (DHS) information from the NIH Epigenomics Roadmap to construct an improved reference DNA methylation database. We show that this leads to a marginal but statistically significant improvement of cell-count estimates in whole blood as well as in mixtures involving epithelial cell-types. Using this framework we compare a widely used state-of-the-art reference-based algorithm (called constrained projection) to two non-constrained approaches including CIBERSORT and a method based on robust part
- 31Regulatory genomic circuitry of human disease loci by integrative epigenomics.Annotating the molecular basis of human disease remains an unsolved challenge, as 93% of disease loci are non-coding and gene-regulatory annotations are highly incomplete 1-3 . Here we present EpiMap, a compendium comprising 10,000 epigenomic maps across 800 samples, which we used to define chromatin states, high-resolution enhancers, enhancer modules, upstream regulators and downstream target genes. We used this resource to annotate 30,000 genetic loci that were associated with 540 traits 4 , predicting trait-relevant tissues, putative causal nucleotide variants in enriched tissue enhancers and candidate tissue-specific target genes for each. We partitioned multifactorial traits into tissue-specific contributing factors with distinct functional enrichments and disease comorbidity patterns, and revealed both single-factor monotropic and multifactor pleiotropic loci. Top-scoring loci frequently had multiple predicted driver variants, converging through multiple enhancers with a common t
- 32From reads to insight: a hitchhiker's guide to ATAC-seq data analysis.Assay of Transposase Accessible Chromatin sequencing (ATAC-seq) is widely used in studying chromatin biology, but a comprehensive review of the analysis tools has not been completed yet. Here, we discuss the major steps in ATAC-seq data analysis, including pre-analysis (quality check and alignment), core analysis (peak calling), and advanced analysis (peak differential analysis and annotation, motif enrichment, footprinting, and nucleosome position analysis). We also review the reconstruction of transcriptional regulatory networks with multiomics data and highlight the current challenges of each step. Finally, we describe the potential of single-cell ATAC-seq and highlight the necessity of developing ATAC-seq specific analysis tools to obtain biologically meaningful insights.
- 33Comprehensive analysis of single cell ATAC-seq data with SnapATAC.Identification of the cis-regulatory elements controlling cell-type specific gene expression patterns is essential for understanding the origin of cellular diversity. Conventional assays to map regulatory elements via open chromatin analysis of primary tissues is hindered by sample heterogeneity. Single cell analysis of accessible chromatin (scATAC-seq) can overcome this limitation. However, the high-level noise of each single cell profile and the large volume of data pose unique computational challenges. Here, we introduce SnapATAC, a software package for analyzing scATAC-seq datasets. SnapATAC dissects cellular heterogeneity in an unbiased manner and map the trajectories of cellular states. Using the Nyström method, SnapATAC can process data from up to a million cells. Furthermore, SnapATAC incorporates existing tools into a comprehensive package for analyzing single cell ATAC-seq dataset. As demonstration of its utility, SnapATAC is applied to 55,592 single-nucleus ATAC-seq profiles
- 34Spatial epigenome-transcriptome co-profiling of mammalian tissues.Emerging spatial technologies, including spatial transcriptomics and spatial epigenomics, are becoming powerful tools for profiling of cellular states in the tissue context 1-5 . However, current methods capture only one layer of omics information at a time, precluding the possibility of examining the mechanistic relationship across the central dogma of molecular biology. Here, we present two technologies for spatially resolved, genome-wide, joint profiling of the epigenome and transcriptome by cosequencing chromatin accessibility and gene expression, or histone modifications (H3K27me3, H3K27ac or H3K4me3) and gene expression on the same tissue section at near-single-cell resolution. These were applied to embryonic and juvenile mouse brain, as well as adult human brain, to map how epigenetic mechanisms control transcriptional phenotype and cell dynamics in tissue. Although highly concordant tissue features were identified by either spatial epigenome or spatial transcriptome we also obs
- 35Spatial profiling of chromatin accessibility in mouse and human tissues.Cellular function in tissue is dependent on the local environment, requiring new methods for spatial mapping of biomolecules and cells in the tissue context 1 . The emergence of spatial transcriptomics has enabled genome-scale gene expression mapping 2-5 , but the ability to capture spatial epigenetic information of tissue at the cellular level and genome scale is lacking. Here we describe a method for spatially resolved chromatin accessibility profiling of tissue sections using next-generation sequencing (spatial-ATAC-seq) by combining in situ Tn5 transposition chemistry 6 and microfluidic deterministic barcoding 5 . Profiling mouse embryos using spatial-ATAC-seq delineated tissue-region-specific epigenetic landscapes and identified gene regulators involved in the development of the central nervous system. Mapping the accessible genome in the mouse and human brain revealed the intricate arealization of brain regions. Applying spatial-ATAC-seq to tonsil tissue resolved the spatially di
- 36The Need for Multi-Omics Biomarker Signatures in Precision Medicine.Recent advances in omics technologies have led to unprecedented efforts characterizing the molecular changes that underlie the development and progression of a wide array of complex human diseases, including cancer. As a result, multi-omics analyses-which take advantage of these technologies in genomics, transcriptomics, epigenomics, proteomics, metabolomics, and other omics areas-have been proposed and heralded as the key to advancing precision medicine in the clinic. In the field of precision oncology, genomics approaches, and, more recently, other omics analyses have helped reveal several key mechanisms in cancer development, treatment resistance, and recurrence risk, and several of these findings have been implemented in clinical oncology to help guide treatment decisions. However, truly integrated multi-omics analyses have not been applied widely, preventing further advances in precision medicine. Additional efforts are needed to develop the analytical infrastructure necessary to
- 37Knee Osteoarthritis: A Review of Pathogenesis and State-Of-The-Art Non-Operative Therapeutic Considerations.Being the most common musculoskeletal progressive condition, osteoarthritis is an interesting target for research. It is estimated that the prevalence of knee osteoarthritis (OA) among adults 60 years of age or older is approximately 10% in men and 13% in women, making knee OA one of the leading causes of disability in elderly population. Today, we know that osteoarthritis is not a disease characterized by loss of cartilage due to mechanical loading only, but a condition that affects all of the tissues in the joint, causing detectable changes in tissue architecture, its metabolism and function. All of these changes are mediated by a complex and not yet fully researched interplay of proinflammatory and anti-inflammatory cytokines, chemokines, growth factors and adipokines, all of which can be measured in the serum, synovium and histological samples, potentially serving as biomarkers of disease stage and progression. Another key aspect of disease progression is the epigenome that regulat
- 38Single-Cell Multiomics: Multiple Measurements from Single Cells.Single-cell sequencing provides information that is not confounded by genotypic or phenotypic heterogeneity of bulk samples. Sequencing of one molecular type (RNA, methylated DNA or open chromatin) in a single cell, furthermore, provides insights into the cell's phenotype and links to its genotype. Nevertheless, only by taking measurements of these phenotypes and genotypes from the same single cells can such inferences be made unambiguously. In this review, we survey the first experimental approaches that assay, in parallel, multiple molecular types from the same single cell, before considering the challenges and opportunities afforded by these and future technologies.
- 39The Impact of Nutrition and Environmental Epigenetics on Human Health and Disease.Environmental epigenetics describes how environmental factors affect cellular epigenetics and, hence, human health. Epigenetic marks alter the spatial conformation of chromatin to regulate gene expression. Environmental factors with epigenetic effects include behaviors, nutrition, and chemicals and industrial pollutants. Epigenetic mechanisms are also implicated during development in utero and at the cellular level, so environmental exposures may harm the fetus by impairing the epigenome of the developing organism to modify disease risk later in life. By contrast, bioactive food components may trigger protective epigenetic modifications throughout life, with early life nutrition being particularly important. Beyond their genetics, the overall health status of an individual may be regarded as an integration of many environmental signals starting at gestation and acting through epigenetic modifications. This review explores how the environment affects the epigenome in health and disease,
- 40Single-cell epigenomics reveals mechanisms of human cortical development.During mammalian development, differences in chromatin state coincide with cellular differentiation and reflect changes in the gene regulatory landscape 1 . In the developing brain, cell fate specification and topographic identity are important for defining cell identity 2 and confer selective vulnerabilities to neurodevelopmental disorders 3 . Here, to identify cell-type-specific chromatin accessibility patterns in the developing human brain, we used a single-cell assay for transposase accessibility by sequencing (scATAC-seq) in primary tissue samples from the human forebrain. We applied unbiased analyses to identify genomic loci that undergo extensive cell-type- and brain-region-specific changes in accessibility during neurogenesis, and an integrative analysis to predict cell-type-specific candidate regulatory elements. We found that cerebral organoids recapitulate most putative cell-type-specific enhancer accessibility patterns but lack many cell-type-specific open chromatin regions
- 41Mutation bias reflects natural selection in Arabidopsis thaliana.Since the first half of the twentieth century, evolutionary theory has been dominated by the idea that mutations occur randomly with respect to their consequences 1 . Here we test this assumption with large surveys of de novo mutations in the plant Arabidopsis thaliana. In contrast to expectations, we find that mutations occur less often in functionally constrained regions of the genome-mutation frequency is reduced by half inside gene bodies and by two-thirds in essential genes. With independent genomic mutation datasets, including from the largest Arabidopsis mutation accumulation experiment conducted to date, we demonstrate that epigenomic and physical features explain over 90% of variance in the genome-wide pattern of mutation bias surrounding genes. Observed mutation frequencies around genes in turn accurately predict patterns of genetic polymorphisms in natural Arabidopsis accessions (r = 0.96). That mutation bias is the primary force behind patterns of sequence evolution around
- 42GTRD: an integrated view of transcription regulation.The Gene Transcription Regulation Database (GTRD; http://gtrd.biouml.org/) contains uniformly annotated and processed NGS data related to gene transcription regulation: ChIP-seq, ChIP-exo, DNase-seq, MNase-seq, ATAC-seq and RNA-seq. With the latest release, the database has reached a new level of data integration. All cell types (cell lines and tissues) presented in the GTRD were arranged into a dictionary and linked with different ontologies (BRENDA, Cell Ontology, Uberon, Cellosaurus and Experimental Factor Ontology) and with related experiments in specialized databases on transcription regulation (FANTOM5, ENCODE and GTEx). The updated version of the GTRD provides an integrated view of transcription regulation through a dedicated web interface with advanced browsing and search capabilities, an integrated genome browser, and table reports by cell types, transcription factors, and genes of interest.
- 43Small-Magnitude Effect Sizes in Epigenetic End Points are Important in Children's Environmental Health Studies: The Children's Environmental Health and Disease Prevention Research Center's Epigenetics Working Group.Background Characterization of the epigenome is a primary interest for children's environmental health researchers studying the environmental influences on human populations, particularly those studying the role of pregnancy and early-life exposures on later-in-life health outcomes. Objectives Our objective was to consider the state of the science in environmental epigenetics research and to focus on DNA methylation and the collective observations of many studies being conducted within the Children's Environmental Health and Disease Prevention Research Centers, as they relate to the Developmental Origins of Health and Disease (DOHaD) hypothesis. Methods We address the current laboratory and statistical tools available for epigenetic analyses, discuss methods for validation and interpretation of findings, particularly when magnitudes of effect are small, question the functional relevance of findings, and discuss the future for environmental epigenetics research. Discussion A common find
- 44Integration of Alzheimer's disease genetics and myeloid genomics identifies disease risk regulatory elements and genes.Genome-wide association studies (GWAS) have identified more than 40 loci associated with Alzheimer's disease (AD), but the causal variants, regulatory elements, genes and pathways remain largely unknown, impeding a mechanistic understanding of AD pathogenesis. Previously, we showed that AD risk alleles are enriched in myeloid-specific epigenomic annotations. Here, we show that they are specifically enriched in active enhancers of monocytes, macrophages and microglia. We integrated AD GWAS with myeloid epigenomic and transcriptomic datasets using analytical approaches to link myeloid enhancer activity to target gene expression regulation and AD risk modification. We identify AD risk enhancers and nominate candidate causal genes among their likely targets (including AP4E1, AP4M1, APBB3, BIN1, MS4A4A, MS4A6A, PILRA, RABEP1, SPI1, TP53INP1, and ZYX) in twenty loci. Fine-mapping of these enhancers nominates candidate functional variants that likely modify AD risk by regulating gene expressi
- 45Profiling genome-wide DNA methylation.DNA methylation is an epigenetic modification that plays an important role in regulating gene expression and therefore a broad range of biological processes and diseases. DNA methylation is tissue-specific, dynamic, sequence-context-dependent and trans-generationally heritable, and these complex patterns of methylation highlight the significance of profiling DNA methylation to answer biological questions. In this review, we surveyed major methylation assays, along with comparisons and biological examples, to provide an overview of DNA methylation profiling techniques. The advances in microarray and sequencing technologies make genome-wide profiling possible at a single-nucleotide or even a single-cell resolution. These profiling approaches vary in many aspects, such as DNA input, resolution, genomic region coverage, and bioinformatics analysis, and selecting a feasible method requires knowledge of these methods. We first introduce the biological background of DNA methylation and its pa
- 46Single-cell epigenomics: powerful new methods for understanding gene regulation and cell identity.Emerging single-cell epigenomic methods are being developed with the exciting potential to transform our knowledge of gene regulation. Here we review available techniques and future possibilities, arguing that the full potential of single-cell epigenetic studies will be realized through parallel profiling of genomic, transcriptional, and epigenetic information.
- 47Simultaneous trimodal single-cell measurement of transcripts, epitopes, and chromatin accessibility using TEA-seq.Single-cell measurements of cellular characteristics have been instrumental in understanding the heterogeneous pathways that drive differentiation, cellular responses to signals, and human disease. Recent advances have allowed paired capture of protein abundance and transcriptomic state, but a lack of epigenetic information in these assays has left a missing link to gene regulation. Using the heterogeneous mixture of cells in human peripheral blood as a test case, we developed a novel scATAC-seq workflow that increases signal-to-noise and allows paired measurement of cell surface markers and chromatin accessibility: integrated cellular indexing of chromatin landscape and epitopes, called ICICLE-seq. We extended this approach using a droplet-based multiomics platform to develop a trimodal assay that simultaneously measures transcriptomics (scRNA-seq), epitopes, and chromatin accessibility (scATAC-seq) from thousands of single cells, which we term TEA-seq. Together, these multimodal sing
- 48TCGA Workflow : Analyze cancer genomics and epigenomics data using Bioconductor packages.Biotechnological advances in sequencing have led to an explosion of publicly available data via large international consortia such as The Cancer Genome Atlas (TCGA), The Encyclopedia of DNA Elements (ENCODE), and The NIH Roadmap Epigenomics Mapping Consortium (Roadmap). These projects have provided unprecedented opportunities to interrogate the epigenome of cultured cancer cell lines as well as normal and tumor tissues with high genomic resolution. The Bioconductor project offers more than 1,000 open-source software and statistical packages to analyze high-throughput genomic data. However, most packages are designed for specific data types (e.g. expression, epigenetics, genomics) and there is no one comprehensive tool that provides a complete integrative analysis of the resources and data provided by all three public projects. A need to create an integration of these different analyses was recently proposed. In this workflow, we provide a series of biologically focused integrative anal
- 49Single-cell sequencing techniques from individual to multiomics analyses.Here, we review single-cell sequencing techniques for individual and multiomics profiling in single cells. We mainly describe single-cell genomic, epigenomic, and transcriptomic methods, and examples of their applications. For the integration of multilayered data sets, such as the transcriptome data derived from single-cell RNA sequencing and chromatin accessibility data derived from single-cell ATAC-seq, there are several computational integration methods. We also describe single-cell experimental methods for the simultaneous measurement of two or more omics layers. We can achieve a detailed understanding of the basic molecular profiles and those associated with disease in each cell by utilizing a large number of single-cell sequencing techniques and the accumulated data sets.
- 50Fast alignment and preprocessing of chromatin profiles with Chromap.As sequencing depth of chromatin studies continually grows deeper for sensitive profiling of regulatory elements or chromatin spatial structures, aligning and preprocessing of these sequencing data have become the bottleneck for analysis. Here we present Chromap, an ultrafast method for aligning and preprocessing high throughput chromatin profiles. Chromap is comparable to BWA-MEM and Bowtie2 in alignment accuracy and is over 10 times faster than traditional workflows on bulk ChIP-seq/Hi-C profiles and than 10x Genomics' CellRanger v2.0.0 pipeline on single-cell ATAC-seq profiles.
- 51A rapid and robust method for single cell chromatin accessibility profiling.The assay for transposase-accessible chromatin using sequencing (ATAC-seq) is widely used to identify regulatory regions throughout the genome. However, very few studies have been performed at the single cell level (scATAC-seq) due to technical challenges. Here we developed a simple and robust plate-based scATAC-seq method, combining upfront bulk Tn5 tagging with single-nuclei sorting. We demonstrate that our method works robustly across various systems, including fresh and cryopreserved cells from primary tissues. By profiling over 3000 splenocytes, we identify distinct immune cell types and reveal cell type-specific regulatory regions and related transcription factors.
- 52The role of ABCA7 in Alzheimer's disease: evidence from genomics, transcriptomics and methylomics.Genome-wide association studies (GWAS) originally identified ATP-binding cassette, sub-family A, member 7 (ABCA7), as a novel risk gene of Alzheimer's disease (AD). Since then, accumulating evidence from in vitro, in vivo, and human-based studies has corroborated and extended this association, promoting ABCA7 as one of the most important risk genes of both early-onset and late-onset AD, harboring both common and rare risk variants with relatively large effect on AD risk. Within this review, we provide a comprehensive assessment of the literature on ABCA7, with a focus on AD-related human -omics studies (e.g. genomics, transcriptomics, and methylomics). In European and African American populations, indirect ABCA7 GWAS associations are explained by expansion of an ABCA7 variable number tandem repeat (VNTR), and a common premature termination codon (PTC) variant, respectively. Rare ABCA7 PTC variants are strongly enriched in AD patients, and some of these have displayed inheritance patter
- 53Microglial PGC-1α protects against ischemic brain injury by suppressing neuroinflammation.Background Neuroinflammation and immune responses occurring minutes to hours after stroke are associated with brain injury after acute ischemic stroke (AIS). PPARγ coactivator-1α (PGC-1α), as a master coregulator of gene expression in mitochondrial biogenesis, was found to be transiently upregulated in microglia after AIS. However, the role of microglial PGC-1α in poststroke immune modulation remains unknown. Methods PGC-1α expression in microglia from human and mouse brain samples following ischemic stroke was first determined. Subsequently, we employed transgenic mice with microglia-specific overexpression of PGC-1α for middle cerebral artery occlusion (MCAO). The morphology and gene expression profile of microglia with PGC-1α overexpression were evaluated. Downstream inflammatory cytokine production and NLRP3 activation were also determined. ChIP-Seq analysis was performed to detect PGC-1α-binding sites in microglia. Autophagic and mitophagic activity was further monitored by immuno
- 54Mapping gene regulatory networks from single-cell omics data.Single-cell techniques are advancing rapidly and are yielding unprecedented insight into cellular heterogeneity. Mapping the gene regulatory networks (GRNs) underlying cell states provides attractive opportunities to mechanistically understand this heterogeneity. In this review, we discuss recently emerging methods to map GRNs from single-cell transcriptomics data, tackling the challenge of increased noise levels and data sparsity compared with bulk data, alongside increasing data volumes. Next, we discuss how new techniques for single-cell epigenomics, such as single-cell ATAC-seq and single-cell DNA methylation profiling, can be used to decipher gene regulatory programmes. We finally look forward to the application of single-cell multi-omics and perturbation techniques that will likely play important roles for GRN inference in the future.
- 55SEACells infers transcriptional and epigenomic cellular states from single-cell genomics data.Metacells are cell groupings derived from single-cell sequencing data that represent highly granular, distinct cell states. Here we present single-cell aggregation of cell states (SEACells), an algorithm for identifying metacells that overcome the sparsity of single-cell data while retaining heterogeneity obscured by traditional cell clustering. SEACells outperforms existing algorithms in identifying comprehensive, compact and well-separated metacells in both RNA and assay for transposase-accessible chromatin (ATAC) modalities across datasets with discrete cell types and continuous trajectories. We demonstrate the use of SEACells to improve gene-peak associations, compute ATAC gene scores and infer the activities of critical regulators during differentiation. Metacell-level analysis scales to large datasets and is particularly well suited for patient cohorts, where per-patient aggregation provides more robust units for data integration. We use our metacells to reveal expression dynamic
- 56Integrated spatial multiomics reveals fibroblast fate during tissue repair.In the skin, tissue injury results in fibrosis in the form of scars composed of dense extracellular matrix deposited by fibroblasts. The therapeutic goal of regenerative wound healing has remained elusive, in part because principles of fibroblast programming and adaptive response to injury remain incompletely understood. Here, we present a multimodal -omics platform for the comprehensive study of cell populations in complex tissue, which has allowed us to characterize the cells involved in wound healing across both time and space. We employ a stented wound model that recapitulates human tissue repair kinetics and multiple Rainbow transgenic lines to precisely track fibroblast fate during the physiologic response to skin injury. Through integrated analysis of single cell chromatin landscapes and gene expression states, coupled with spatial transcriptomic profiling, we are able to impute fibroblast epigenomes with temporospatial resolution. This has allowed us to reveal potential mechani
- 57Single Cell Multi-Omics Technology: Methodology and Application.In the era of precision medicine, multi-omics approaches enable the integration of data from diverse omics platforms, providing multi-faceted insight into the interrelation of these omics layers on disease processes. Single cell sequencing technology can dissect the genotypic and phenotypic heterogeneity of bulk tissue and promises to deepen our understanding of the underlying mechanisms governing both health and disease. Through modification and combination of single cell assays available for transcriptome, genome, epigenome, and proteome profiling, single cell multi-omics approaches have been developed to simultaneously and comprehensively study not only the unique genotypic and phenotypic characteristics of single cells, but also the combined regulatory mechanisms evident only at single cell resolution. In this review, we summarize the state-of-the-art single cell multi-omics methods and discuss their applications, challenges, and future directions.
- 58Multi-Omics of Single Cells: Strategies and Applications.Most genome-wide assays provide averages across large numbers of cells, but recent technological advances promise to overcome this limitation. Pioneering single-cell assays are now available for genome, epigenome, transcriptome, proteome, and metabolome profiling. Here, we describe how these different dimensions can be combined into multi-omics assays that provide comprehensive profiles of the same cell.
- 59A reference methylome database and analysis pipeline to facilitate integrative and comparative epigenomics.DNA methylation is implicated in a surprising diversity of regulatory, evolutionary processes and diseases in eukaryotes. The introduction of whole-genome bisulfite sequencing has enabled the study of DNA methylation at a single-base resolution, revealing many new aspects of DNA methylation and highlighting the usefulness of methylome data in understanding a variety of genomic phenomena. As the number of publicly available whole-genome bisulfite sequencing studies reaches into the hundreds, reliable and convenient tools for comparing and analyzing methylomes become increasingly important. We present MethPipe, a pipeline for both low and high-level methylome analysis, and MethBase, an accompanying database of annotated methylomes from the public domain. Together these resources enable researchers to extract interesting features from methylomes and compare them with those identified in public methylomes in our database.
- 60Single-cell multiomics analysis reveals regulatory programs in clear cell renal cell carcinoma.The clear cell renal cell carcinoma (ccRCC) microenvironment consists of many different cell types and structural components that play critical roles in cancer progression and drug resistance, but the cellular architecture and underlying gene regulatory features of ccRCC have not been fully characterized. Here, we applied single-cell RNA sequencing (scRNA-seq) and single-cell assay for transposase-accessible chromatin sequencing (scATAC-seq) to generate transcriptional and epigenomic landscapes of ccRCC. We identified tumor cell-specific regulatory programs mediated by four key transcription factors (TFs) (HOXC5, VENTX, ISL1, and OTP), and these TFs have prognostic significance in The Cancer Genome Atlas (TCGA) database. Targeting these TFs via short hairpin RNAs (shRNAs) or small molecule inhibitors decreased tumor cell proliferation. We next performed an integrative analysis of chromatin accessibility and gene expression for CD8 + T cells and macrophages to reveal the different regul
- 61CUT&Tag for efficient epigenomic profiling of small samples and single cells.Many chromatin features play critical roles in regulating gene expression. A complete understanding of gene regulation will require the mapping of specific chromatin features in small samples of cells at high resolution. Here we describe Cleavage Under Targets and Tagmentation (CUT&Tag), an enzyme-tethering strategy that provides efficient high-resolution sequencing libraries for profiling diverse chromatin components. In CUT&Tag, a chromatin protein is bound in situ by a specific antibody, which then tethers a protein A-Tn5 transposase fusion protein. Activation of the transposase efficiently generates fragment libraries with high resolution and exceptionally low background. All steps from live cells to sequencing-ready libraries can be performed in a single tube on the benchtop or a microwell in a high-throughput pipeline, and the entire procedure can be performed in one day. We demonstrate the utility of CUT&Tag by profiling histone modifications, RNA Polymerase II and transcription
- 62An efficient targeted nuclease strategy for high-resolution mapping of DNA binding sites.We describe Cleavage Under Targets and Release Using Nuclease (CUT&RUN), a chromatin profiling strategy in which antibody-targeted controlled cleavage by micrococcal nuclease releases specific protein-DNA complexes into the supernatant for paired-end DNA sequencing. Unlike Chromatin Immunoprecipitation (ChIP), which fragments and solubilizes total chromatin, CUT&RUN is performed in situ, allowing for both quantitative high-resolution chromatin mapping and probing of the local chromatin environment. When applied to yeast and human nuclei, CUT&RUN yielded precise transcription factor profiles while avoiding crosslinking and solubilization issues. CUT&RUN is simple to perform and is inherently robust, with extremely low backgrounds requiring only ~1/10th the sequencing depth as ChIP, making CUT&RUN especially cost-effective for transcription factor and chromatin profiling. When used in conjunction with native ChIP-seq and applied to human CTCF, CUT&RUN mapped directional long range contac
- 63ChIPpeakAnno: a Bioconductor package to annotate ChIP-seq and ChIP-chip data.Background Chromatin immunoprecipitation (ChIP) followed by high-throughput sequencing (ChIP-seq) or ChIP followed by genome tiling array analysis (ChIP-chip) have become standard technologies for genome-wide identification of DNA-binding protein target sites. A number of algorithms have been developed in parallel that allow identification of binding sites from ChIP-seq or ChIP-chip datasets and subsequent visualization in the University of California Santa Cruz (UCSC) Genome Browser as custom annotation tracks. However, summarizing these tracks can be a daunting task, particularly if there are a large number of binding sites or the binding sites are distributed widely across the genome. Results We have developed ChIPpeakAnno as a Bioconductor package within the statistical programming environment R to facilitate batch annotation of enriched peaks identified from ChIP-seq, ChIP-chip, cap analysis of gene expression (CAGE) or any experiments resulting in a large number of enriched genom
- 64Three-dimensional Epigenome Statistical Model: Genome-wide Chromatin Looping Prediction.This study aims to understand through statistical learning the basic biophysical mechanisms behind three-dimensional folding of epigenomes. The 3DEpiLoop algorithm predicts three-dimensional chromatin looping interactions within topologically associating domains (TADs) from one-dimensional epigenomics and transcription factor profiles using the statistical learning. The predictions obtained by 3DEpiLoop are highly consistent with the reported experimental interactions. The complex signatures of epigenomic and transcription factors within the physically interacting chromatin regions (anchors) are similar across all genomic scales: genomic domains, chromosomal territories, cell types, and different individuals. We report the most important epigenetic and transcription factor features used for interaction identification either shared, or unique for each of sixteen (16) cell lines. The analysis shows that CTCF interaction anchors are enriched by transcription factors yet deficient in histo
- 65ATAC-seq footprinting unravels kinetics of transcription factor binding during zygotic genome activation.While footprinting analysis of ATAC-seq data can theoretically enable investigation of transcription factor (TF) binding, the lack of a computational tool able to conduct different levels of footprinting analysis has so-far hindered the widespread application of this method. Here we present TOBIAS, a comprehensive, accurate, and fast footprinting framework enabling genome-wide investigation of TF binding dynamics for hundreds of TFs simultaneously. We validate TOBIAS using paired ATAC-seq and ChIP-seq data, and find that TOBIAS outperforms existing methods for bias correction and footprinting. As a proof-of-concept, we illustrate how TOBIAS can unveil complex TF dynamics during zygotic genome activation in both humans and mice, and propose how zygotic Dux activates cascades of TFs, binds to repeat elements and induces expression of novel genetic elements.
- 66Cistrome: an integrative platform for transcriptional regulation studies.The increasing volume of ChIP-chip and ChIP-seq data being generated creates a challenge for standard, integrative and reproducible bioinformatics data analysis platforms. We developed a web-based application called Cistrome, based on the Galaxy open source framework. In addition to the standard Galaxy functions, Cistrome has 29 ChIP-chip- and ChIP-seq-specific tools in three major categories, from preliminary peak calling and correlation analyses to downstream genome feature association, gene expression analyses, and motif discovery. Cistrome is available at http://cistrome.org/ap/.
- 67Peak calling by Sparse Enrichment Analysis for CUT&RUN chromatin profiling.Background CUT&RUN is an efficient epigenome profiling method that identifies sites of DNA binding protein enrichment genome-wide with high signal to noise and low sequencing requirements. Currently, the analysis of CUT&RUN data is complicated by its exceptionally low background, which renders programs designed for analysis of ChIP-seq data vulnerable to oversensitivity in identifying sites of protein binding. Results Here we introduce Sparse Enrichment Analysis for CUT&RUN (SEACR), an analysis strategy that uses the global distribution of background signal to calibrate a simple threshold for peak calling. SEACR discriminates between true and false-positive peaks with near-perfect specificity from "gold standard" CUT&RUN datasets and efficiently identifies enriched regions for several different protein targets. We also introduce a web server ( http://seacr.fredhutch.org ) for plug-and-play analysis with SEACR that facilitates maximum accessibility across users of all skill levels. Conc
- 68Microanatomy of the Human Atherosclerotic Plaque by Single-Cell Transcriptomics.Rationale Atherosclerotic lesions are known for their cellular heterogeneity, yet the molecular complexity within the cells of human plaques has not been fully assessed. Objective Using single-cell transcriptomics and chromatin accessibility, we gained a better understanding of the pathophysiology underlying human atherosclerosis. Methods and results We performed single-cell RNA and single-cell ATAC sequencing on human carotid atherosclerotic plaques to define the cells at play and determine their transcriptomic and epigenomic characteristics. We identified 14 distinct cell populations including endothelial cells, smooth muscle cells, mast cells, B cells, myeloid cells, and T cells and identified multiple cellular activation states and suggested cellular interconversions. Within the endothelial cell population, we defined subsets with angiogenic capacity plus clear signs of endothelial to mesenchymal transition. CD4 + and CD8 + T cells showed activation-based subclasses, each with a gr
- 69Single-cell triple omics sequencing reveals genetic, epigenetic, and transcriptomic heterogeneity in hepatocellular carcinomas.Single-cell genome, DNA methylome, and transcriptome sequencing methods have been separately developed. However, to accurately analyze the mechanism by which transcriptome, genome and DNA methylome regulate each other, these omic methods need to be performed in the same single cell. Here we demonstrate a single-cell triple omics sequencing technique, scTrio-seq, that can be used to simultaneously analyze the genomic copy-number variations (CNVs), DNA methylome, and transcriptome of an individual mammalian cell. We show that large-scale CNVs cause proportional changes in RNA expression of genes within the gained or lost genomic regions, whereas these CNVs generally do not affect DNA methylation in these regions. Furthermore, we applied scTrio-seq to 25 single cancer cells derived from a human hepatocellular carcinoma tissue sample. We identified two subpopulations within these cells based on CNVs, DNA methylome, or transcriptome of individual cells. Our work offers a new avenue of disse
- 70Improved CUT&RUN chromatin profiling tools.Previously, we described a novel alternative to chromatin immunoprecipitation, CUT&RUN, in which unfixed permeabilized cells are incubated with antibody, followed by binding of a protein A-Micrococcal Nuclease (pA/MNase) fusion protein (Skene and Henikoff, 2017). Here we introduce three enhancements to CUT&RUN: A hybrid protein A-Protein G-MNase construct that expands antibody compatibility and simplifies purification, a modified digestion protocol that inhibits premature release of the nuclease-bound complex, and a calibration strategy based on carry-over of E. coli DNA introduced with the fusion protein. These new features, coupled with the previously described low-cost, high efficiency, high reproducibility and high-throughput capability of CUT&RUN make it the method of choice for routine epigenomic profiling.
- 71Sexual-dimorphism in human immune system aging.Differences in immune function and responses contribute to health- and life-span disparities between sexes. However, the role of sex in immune system aging is not well understood. Here, we characterize peripheral blood mononuclear cells from 172 healthy adults 22-93 years of age using ATAC-seq, RNA-seq, and flow cytometry. These data reveal a shared epigenomic signature of aging including declining naïve T cell and increasing monocyte and cytotoxic cell functions. These changes are greater in magnitude in men and accompanied by a male-specific decline in B-cell specific loci. Age-related epigenomic changes first spike around late-thirties with similar timing and magnitude between sexes, whereas the second spike is earlier and stronger in men. Unexpectedly, genomic differences between sexes increase after age 65, with men having higher innate and pro-inflammatory activity and lower adaptive activity. Impact of age and sex on immune phenotypes can be visualized at https://immune-aging.ja
- 72Widespread natural variation of DNA methylation within angiosperms.Background DNA methylation is an important feature of plant epigenomes, involved in the formation of heterochromatin and affecting gene expression. Extensive variation of DNA methylation patterns within a species has been uncovered from studies of natural variation. However, the extent to which DNA methylation varies between flowering plant species is still unclear. To understand the variation in genomic patterning of DNA methylation across flowering plant species, we compared single base resolution DNA methylomes of 34 diverse angiosperm species. Results By analyzing whole-genome bisulfite sequencing data in a phylogenetic context, it becomes clear that there is extensive variation throughout angiosperms in gene body DNA methylation, euchromatic silencing of transposons and repeats, as well as silencing of heterochromatic transposons. The Brassicaceae have reduced CHG methylation levels and also reduced or loss of CG gene body methylation. The Poaceae are characterized by a lack or re