专题电子书

蛋白质组循证手册

按质量评分、证据类型和发表时间组织;优先收录明确允许商业复用的来源文献。

135 个章节

  1. 01The PRIDE database at 20 years: 2025 update.The PRoteomics IDEntifications (PRIDE) database (https://www.ebi.ac.uk/pride/) is the world's leading mass spectrometry (MS)-based proteomics data repository and one of the founding members of the ProteomeXchange consortium. This manuscript summarizes the developments in PRIDE resources and related tools for the last three years. The number of submitted datasets to PRIDE Archive (the archival component of PRIDE) has reached on average around 534 datasets per month. This has been possible thanks to continuous improvements in infrastructure such as a new file transfer protocol for very large datasets (Globus), a new data resubmission pipeline and an automatic dataset validation process. Additionally, we will highlight novel activities such as the availability of the PRIDE chatbot (based on the use of open-source Large Language Models), and our work to improve support for MS crosslinking datasets. Furthermore, we will describe how we have increased our efforts to reuse, reanalyze and diss
  2. 02Guidelines and considerations for the use of system suitability and quality control samples in mass spectrometry assays applied in untargeted clinical metabolomic studies.Background Quality assurance (QA) and quality control (QC) are two quality management processes that are integral to the success of metabolomics including their application for the acquisition of high quality data in any high-throughput analytical chemistry laboratory. QA defines all the planned and systematic activities implemented before samples are collected, to provide confidence that a subsequent analytical process will fulfil predetermined requirements for quality. QC can be defined as the operational techniques and activities used to measure and report these quality requirements after data acquisition. Aim of review This tutorial review will guide the reader through the use of system suitability and QC samples, why these samples should be applied and how the quality of data can be reported. Key scientific concepts of review System suitability samples are applied to assess the operation and lack of contamination of the analytical platform prior to sample analysis. Isotopically-la
  3. 03Estimating the total number of phosphoproteins and phosphorylation sites in eukaryotic proteomes.Background Phosphorylation is the most frequent post-translational modification made to proteins and may regulate protein activity as either a molecular digital switch or a rheostat. Despite the cornucopia of high-throughput (HTP) phosphoproteomic data in the last decade, it remains unclear how many proteins are phosphorylated and how many phosphorylation sites (p-sites) can exist in total within a eukaryotic proteome. We present the first reliable estimates of the total number of phosphoproteins and p-sites for four eukaryotes (human, mouse, Arabidopsis, and yeast). Results In all, 187 HTP phosphoproteomic datasets were filtered, compiled, and studied along with two low-throughput (LTP) compendia. Estimates of the number of phosphoproteins and p-sites were inferred by two methods: Capture-Recapture, and fitting the saturation curve of cumulative redundant vs. cumulative non-redundant phosphoproteins/p-sites. Estimates were also adjusted for different levels of noise within the individ
  4. 04The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences.The PRoteomics IDEntifications (PRIDE) database (https://www.ebi.ac.uk/pride/) is the world's largest data repository of mass spectrometry-based proteomics data. PRIDE is one of the founding members of the global ProteomeXchange (PX) consortium and an ELIXIR core data resource. In this manuscript, we summarize the developments in PRIDE resources and related tools since the previous update manuscript was published in Nucleic Acids Research in 2019. The number of submitted datasets to PRIDE Archive (the archival component of PRIDE) has reached on average around 500 datasets per month during 2021. In addition to continuous improvements in PRIDE Archive data pipelines and infrastructure, the PRIDE Spectra Archive has been developed to provide direct access to the submitted mass spectra using Universal Spectrum Identifiers. As a key point, the file format MAGE-TAB for proteomics has been developed to enable the improvement of sample metadata annotation. Additionally, the resource PRIDE Pept
  5. 05A blood atlas of COVID-19 defines hallmarks of disease severity and specificity.Treatment of severe COVID-19 is currently limited by clinical heterogeneity and incomplete description of specific immune biomarkers. We present here a comprehensive multi-omic blood atlas for patients with varying COVID-19 severity in an integrated comparison with influenza and sepsis patients versus healthy volunteers. We identify immune signatures and correlates of host response. Hallmarks of disease severity involved cells, their inflammatory mediators and networks, including progenitor cells and specific myeloid and lymphocyte subsets, features of the immune repertoire, acute phase response, metabolism, and coagulation. Persisting immune activation involving AP-1/p38MAPK was a specific feature of COVID-19. The plasma proteome enabled sub-phenotyping into patient clusters, predictive of severity and outcome. Systems-based integrative analyses including tensor and matrix decomposition of all modalities revealed feature groupings linked with severity and specificity compared to influ
  6. 06Single-cell sequencing to multi-omics: technologies and applications.Cells, as the fundamental units of life, contain multidimensional spatiotemporal information. Single-cell RNA sequencing (scRNA-seq) is revolutionizing biomedical science by analyzing cellular state and intercellular heterogeneity. Undoubtedly, single-cell transcriptomics has emerged as one of the most vibrant research fields today. With the optimization and innovation of single-cell sequencing technologies, the intricate multidimensional details concealed within cells are gradually unveiled. The combination of scRNA-seq and other multi-omics is at the forefront of the single-cell field. This involves simultaneously measuring various omics data within individual cells, expanding our understanding across a broader spectrum of dimensions. Single-cell multi-omics precisely captures the multidimensional aspects of single-cell transcriptomes, immune repertoire, spatial information, temporal information, epitopes, and other omics in diverse spatiotemporal contexts. In addition to depicting t
  7. 07The TOPCONS web server for consensus prediction of membrane protein topology and signal peptides.TOPCONS (http://topcons.net/) is a widely used web server for consensus prediction of membrane protein topology. We hereby present a major update to the server, with some substantial improvements, including the following: (i) TOPCONS can now efficiently separate signal peptides from transmembrane regions. (ii) The server can now differentiate more successfully between globular and membrane proteins. (iii) The server now is even slightly faster, although a much larger database is used to generate the multiple sequence alignments. For most proteins, the final prediction is produced in a matter of seconds. (iv) The user-friendly interface is retained, with the additional feature of submitting batch files and accessing the server programmatically using standard interfaces, making it thus ideal for proteome-wide analyses. Indicatively, the user can now scan the entire human proteome in a few days. (v) For proteins with homology to a known 3D structure, the homology-inferred topology is also
  8. 08Proteome Discoverer-A Community Enhanced Data Processing Suite for Protein Informatics.Proteomics researchers today face an interesting challenge: how to choose among the dozens of data processing and analysis pipelines available for converting tandem mass spectrometry files to protein identifications. Due to the dominance of Orbitrap technology in proteomics in recent history, many researchers have defaulted to the vendor software Proteome Discoverer. Over the fourteen years since the initial release of the software, it has evolved in parallel with the increasingly complex demands faced by proteomics researchers. Today, Proteome Discoverer exists in two distinct forms with both powerful commercial versions and fully functional free versions in use in many labs today. Throughout the 11 main versions released to date, a central theme of the software has always been the ability to easily view and verify the spectra from which identifications are made. This ability is, even today, a key differentiator from other data analysis solutions. In this review I will attempt to summ
  9. 09Molecular evaluation of five different isolation methods for extracellular vesicles reveals different clinical applicability and subcellular origin.Extracellular vesicles (EVs) are increasingly tested as therapeutic vehicles and biomarkers, but still EV subtypes are not fully characterised. To isolate EVs with few co-isolated entities, a combination of methods is needed. However, this is time-consuming and requires large sample volumes, often not feasible in most clinical studies or in studies where small sample volumes are available. Therefore, we compared EVs rendered by five commonly used methods based on different principles from conditioned cell medium and 250 μl or 3 ml plasma, that is, precipitation (ExoQuick ULTRA), membrane affinity (exoEasy Maxi Kit), size-exclusion chromatography (qEVoriginal), iodixanol gradient (OptiPrep), and phosphatidylserine affinity (MagCapture). EVs were characterised by electron microscopy, Nanoparticle Tracking Analysis, Bioanalyzer, flow cytometry, and LC-MS/MS. The different methods yielded samples of different morphology, particle size, and proteomic profile. For the conditioned medium, Izo
  10. 10Advances and Trends in Omics Technology Development.The human history has witnessed the rapid development of technologies such as high-throughput sequencing and mass spectrometry that led to the concept of "omics" and methodological advancement in systematically interrogating a cellular system. Yet, the ever-growing types of molecules and regulatory mechanisms being discovered have been persistently transforming our understandings on the cellular machinery. This renders cell omics seemingly, like the universe, expand with no limit and our goal toward the complete harness of the cellular system merely impossible. Therefore, it is imperative to review what has been done and is being done to predict what can be done toward the translation of omics information to disease control with minimal cell perturbation. With a focus on the "four big omics," i.e., genomics, transcriptomics, proteomics, metabolomics, we delineate hierarchies of these omics together with their epiomics and interactomics, and review technologies developed for interrogati
  11. 11Multi-Omics Profiling for Health.The world has witnessed a steady rise in both non-infectious and infectious chronic diseases, prompting a cross-disciplinary approach to understand and treating disease. Current medical care focuses on treating people after they become patients rather than preventing illness, leading to high costs in treating chronic and late-stage diseases. Additionally, a "one-size-fits all" approach to health care does not take into account individual differences in genetics, environment, or lifestyle factors, decreasing the number of people benefiting from interventions. Rapid advances in omics technologies and progress in computational capabilities have led to the development of multi-omics deep phenotyping, which profiles the interaction of multiple levels of biology over time and empowers precision health approaches. This review highlights current and emerging multi-omics modalities for precision health and discusses applications in the following areas: genetic variation, cardio-metabolic diseas
  12. 12Integrating Molecular Perspectives: Strategies for Comprehensive Multi-Omics Integrative Data Analysis and Machine Learning Applications in Transcriptomics, Proteomics, and Metabolomics.With the advent of high-throughput technologies, the field of omics has made significant strides in characterizing biological systems at various levels of complexity. Transcriptomics, proteomics, and metabolomics are the three most widely used omics technologies, each providing unique insights into different layers of a biological system. However, analyzing each omics data set separately may not provide a comprehensive understanding of the subject under study. Therefore, integrating multi-omics data has become increasingly important in bioinformatics research. In this article, we review strategies for integrating transcriptomics, proteomics, and metabolomics data, including co-expression analysis, metabolite-gene networks, constraint-based models, pathway enrichment analysis, and interactome analysis. We discuss combined omics integration approaches, correlation-based strategies, and machine learning techniques that utilize one or more types of omics data. By presenting these methods,
  13. 13Revolutionizing Personalized Medicine: Synergy with Multi-Omics Data Generation, Main Hurdles, and Future Perspectives.The field of personalized medicine is undergoing a transformative shift through the integration of multi-omics data, which mainly encompasses genomics, transcriptomics, proteomics, and metabolomics. This synergy allows for a comprehensive understanding of individual health by analyzing genetic, molecular, and biochemical profiles. The generation and integration of multi-omics data enable more precise and tailored therapeutic strategies, improving the efficacy of treatments and reducing adverse effects. However, several challenges hinder the full realization of personalized medicine. Key hurdles include the complexity of data integration across different omics layers, the need for advanced computational tools, and the high cost of comprehensive data generation. Additionally, issues related to data privacy, standardization, and the need for robust validation in diverse populations remain significant obstacles. Looking ahead, the future of personalized medicine promises advancements in te
  14. 14A high-speed search engine pLink 2 with systematic evaluation for proteome-scale identification of cross-linked peptides.We describe pLink 2, a search engine with higher speed and reliability for proteome-scale identification of cross-linked peptides. With a two-stage open search strategy facilitated by fragment indexing, pLink 2 is ~40 times faster than pLink 1 and 3~10 times faster than Kojak. Furthermore, using simulated datasets, synthetic datasets, 15 N metabolically labeled datasets, and entrapment databases, four analysis methods were designed to evaluate the credibility of ten state-of-the-art search engines. This systematic evaluation shows that pLink 2 outperforms these methods in precision and sensitivity, especially at proteome scales. Lastly, re-analysis of four published proteome-scale cross-linking datasets with pLink 2 required only a fraction of the time used by pLink 1, with up to 27% more cross-linked residue pairs identified. pLink 2 is therefore an efficient and reliable tool for cross-linking mass spectrometry analysis, and the systematic evaluation methods described here will be us
  15. 15A systematic evaluation of normalization methods in quantitative label-free proteomics.To date, mass spectrometry (MS) data remain inherently biased as a result of reasons ranging from sample handling to differences caused by the instrumentation. Normalization is the process that aims to account for the bias and make samples more comparable. The selection of a proper normalization method is a pivotal task for the reliability of the downstream analysis and results. Many normalization methods commonly used in proteomics have been adapted from the DNA microarray techniques. Previous studies comparing normalization methods in proteomics have focused mainly on intragroup variation. In this study, several popular and widely used normalization methods representing different strategies in normalization are evaluated using three spike-in and one experimental mouse label-free proteomic data sets. The normalization methods are evaluated in terms of their ability to reduce variation between technical replicates, their effect on differential expression analysis and their effect on th
  16. 16Ultra-fast label-free quantification and comprehensive proteome coverage with narrow-window data-independent acquisition.Mass spectrometry (MS)-based proteomics aims to characterize comprehensive proteomes in a fast and reproducible manner. Here we present the narrow-window data-independent acquisition (nDIA) strategy consisting of high-resolution MS1 scans with parallel tandem MS (MS/MS) scans of ~200 Hz using 2-Th isolation windows, dissolving the differences between data-dependent and -independent methods. This is achieved by pairing a quadrupole Orbitrap mass spectrometer with the asymmetric track lossless (Astral) analyzer which provides >200-Hz MS/MS scanning speed, high resolving power and sensitivity, and low-ppm mass accuracy. The nDIA strategy enables profiling of >100 full yeast proteomes per day, or 48 human proteomes per day at the depth of ~10,000 human protein groups in half-an-hour or ~7,000 proteins in 5 min, representing 3× higher coverage compared with current state-of-the-art MS. Multi-shot acquisition of offline fractionated samples provides comprehensive coverage of human proteomes
  17. 17Best practices and benchmarks for intact protein analysis for top-down mass spectrometry.One gene can give rise to many functionally distinct proteoforms, each of which has a characteristic molecular mass. Top-down mass spectrometry enables the analysis of intact proteins and proteoforms. Here members of the Consortium for Top-Down Proteomics provide a decision tree that guides researchers to robust protocols for mass analysis of intact proteins (antibodies, membrane proteins and others) from mixtures of varying complexity. We also present cross-platform analytical benchmarks using a protein standard sample, to allow users to gauge their proficiency.
  18. 18Proteomic aging clock predicts mortality and risk of common age-related diseases in diverse populations.Circulating plasma proteins play key roles in human health and can potentially be used to measure biological age, allowing risk prediction for age-related diseases, multimorbidity and mortality. Here we developed a proteomic age clock in the UK Biobank (n = 45,441) using a proteomic platform comprising 2,897 plasma proteins and explored its utility to predict major disease morbidity and mortality in diverse populations. We identified 204 proteins that accurately predict chronological age (Pearson r = 0.94) and found that proteomic aging was associated with the incidence of 18 major chronic diseases (including diseases of the heart, liver, kidney and lung, diabetes, neurodegeneration and cancer), as well as with multimorbidity and all-cause mortality risk. Proteomic aging was also associated with age-related measures of biological, physical and cognitive function, including telomere length, frailty index and reaction time. Proteins contributing most substantially to the proteomic age cl
  19. 19Applications of Multi-Omics Technologies for Crop Improvement.Multiple "omics" approaches have emerged as successful technologies for plant systems over the last few decades. Advances in next-generation sequencing (NGS) have paved a way for a new generation of different omics, such as genomics, transcriptomics, and proteomics. However, metabolomics, ionomics, and phenomics have also been well-documented in crop science. Multi-omics approaches with high throughput techniques have played an important role in elucidating growth, senescence, yield, and the responses to biotic and abiotic stress in numerous crops. These omics approaches have been implemented in some important crops including wheat ( Triticum aestivum L.), soybean ( Glycine max ), tomato ( Solanum lycopersicum ), barley ( Hordeum vulgare L.), maize ( Zea mays L.), millet ( Setaria italica L.), cotton ( Gossypium hirsutum L.), Medicago truncatula , and rice ( Oryza sativa L.). The integration of functional genomics with other omics highlights the relationships between crop genomes and p
  20. 20Spatial multi-omics analyses of the tumor immune microenvironment.In the past decade, single-cell technologies have revealed the heterogeneity of the tumor-immune microenvironment at the genomic, transcriptomic, and proteomic levels and have furthered our understanding of the mechanisms of tumor development. Single-cell technologies have also been used to identify potential biomarkers. However, spatial information about the tumor-immune microenvironment such as cell locations and cell-cell interactomes is lost in these approaches. Recently, spatial multi-omics technologies have been used to study transcriptomes, proteomes, and metabolomes of tumor-immune microenvironments in several types of cancer, and the data obtained from these methods has been combined with immunohistochemistry and multiparameter analysis to yield markers of cancer progression. Here, we review numerous cutting-edge spatial 'omics techniques, their application to study of the tumor-immune microenvironment, and remaining technical challenges.
  21. 21The PRIDE database and related tools and resources in 2019: improving support for quantification data.The PRoteomics IDEntifications (PRIDE) database (https://www.ebi.ac.uk/pride/) is the world's largest data repository of mass spectrometry-based proteomics data, and is one of the founding members of the global ProteomeXchange (PX) consortium. In this manuscript, we summarize the developments in PRIDE resources and related tools since the previous update manuscript was published in Nucleic Acids Research in 2016. In the last 3 years, public data sharing through PRIDE (as part of PX) has definitely become the norm in the field. In parallel, data re-use of public proteomics data has increased enormously, with multiple applications. We first describe the new architecture of PRIDE Archive, the archival component of PRIDE. PRIDE Archive and the related data submission framework have been further developed to support the increase in submitted data volumes and additional data types. A new scalable and fault tolerant storage backend, Application Programming Interface and web interface have bee
  22. 222016 update of the PRIDE database and its related tools.The PRoteomics IDEntifications (PRIDE) database is one of the world-leading data repositories of mass spectrometry (MS)-based proteomics data. Since the beginning of 2014, PRIDE Archive (http://www.ebi.ac.uk/pride/archive/) is the new PRIDE archival system, replacing the original PRIDE database. Here we summarize the developments in PRIDE resources and related tools since the previous update manuscript in the Database Issue in 2013. PRIDE Archive constitutes a complete redevelopment of the original PRIDE, comprising a new storage backend, data submission system and web interface, among other components. PRIDE Archive supports the most-widely used PSI (Proteomics Standards Initiative) data standard formats (mzML and mzIdentML) and implements the data requirements and guidelines of the ProteomeXchange Consortium. The wide adoption of ProteomeXchange within the community has triggered an unprecedented increase in the number of submitted data sets (around 150 data sets per month). We outli
  23. 23The MEROPS database of proteolytic enzymes, their substrates and inhibitors in 2017 and a comparison with peptidases in the PANTHER database.The MEROPS database (http://www.ebi.ac.uk/merops/) is an integrated source of information about peptidases, their substrates and inhibitors. The hierarchical classification is: protein-species, family, clan, with an identifier at each level. The MEROPS website moved to the EMBL-EBI in 2017, requiring refactoring of the code-base and services provided. The interface to sequence searching has changed and the MEROPS protein sequence libraries can be searched at the EMBL-EBI with HMMER, FastA and BLASTP. Cross-references have been established between MEROPS and the PANTHER database at both the family and protein-species level, which will help to improve curation and coverage between the resources. Because of the increasing size of the MEROPS sequence collection, in future only sequences of characterized proteins, and from completely sequenced genomes of organisms of evolutionary, medical or commercial significance will be added. As an example, peptidase homologues in four proteomes from th
  24. 24A proteomic atlas of senescence-associated secretomes for aging biomarker development.The senescence-associated secretory phenotype (SASP) has recently emerged as a driver of and promising therapeutic target for multiple age-related conditions, ranging from neurodegeneration to cancer. The complexity of the SASP, typically assessed by a few dozen secreted proteins, has been greatly underestimated, and a small set of factors cannot explain the diverse phenotypes it produces in vivo. Here, we present the "SASP Atlas," a comprehensive proteomic database of soluble proteins and exosomal cargo SASP factors originating from multiple senescence inducers and cell types. Each profile consists of hundreds of largely distinct proteins but also includes a subset of proteins elevated in all SASPs. Our analyses identify several candidate biomarkers of cellular senescence that overlap with aging markers in human plasma, including Growth/differentiation factor 15 (GDF15), stanniocalcin 1 (STC1), and serine protease inhibitors (SERPINs), which significantly correlated with age in plasma
  25. 25A deep proteome and transcriptome abundance atlas of 29 healthy human tissues.Genome-, transcriptome- and proteome-wide measurements provide insights into how biological systems are regulated. However, fundamental aspects relating to which human proteins exist, where they are expressed and in which quantities are not fully understood. Therefore, we generated a quantitative proteome and transcriptome abundance atlas of 29 paired healthy human tissues from the Human Protein Atlas project representing human genes by 18,072 transcripts and 13,640 proteins including 37 without prior protein-level evidence. The analysis revealed that hundreds of proteins, particularly in testis, could not be detected even for highly expressed mRNAs, that few proteins show tissue-specific expression, that strong differences between mRNA and protein quantities within and across tissues exist and that protein expression is often more stable across tissues than that of transcripts. Only 238 of 9,848 amino acid variants found by exome sequencing could be confidently detected at the protein
  26. 26An atlas of the aging lung mapped by single cell transcriptomics and deep tissue proteomics.Aging promotes lung function decline and susceptibility to chronic lung diseases, which are the third leading cause of death worldwide. Here, we use single cell transcriptomics and mass spectrometry-based proteomics to quantify changes in cellular activity states across 30 cell types and chart the lung proteome of young and old mice. We show that aging leads to increased transcriptional noise, indicating deregulated epigenetic control. We observe cell type-specific effects of aging, uncovering increased cholesterol biosynthesis in type-2 pneumocytes and lipofibroblasts and altered relative frequency of airway epithelial cells as hallmarks of lung aging. Proteomic profiling reveals extracellular matrix remodeling in old mice, including increased collagen IV and XVI and decreased Fraser syndrome complex proteins and collagen XIV. Computational integration of the aging proteome with the single cell transcriptomes predicts the cellular source of regulated proteins and creates an unbiased r
  27. 27A Review and Database of Snake Venom Proteomes.Advances in the last decade combining transcriptomics with established proteomics methods have made possible rapid identification and quantification of protein families in snake venoms. Although over 100 studies have been published, the value of this information is increased when it is collated, allowing rapid assimilation and evaluation of evolutionary trends, geographical variation, and possible medical implications. This review brings together all compositional studies of snake venom proteomes published in the last decade. Compositional studies were identified for 132 snake species: 42 from 360 (12%) Elapidae (elapids), 20 from 101 (20%) Viperinae (true vipers), 65 from 239 (27%) Crotalinae (pit vipers), and five species of non-front-fanged snakes. Approximately 90% of their total venom composition consisted of eight protein families for elapids, 11 protein families for viperines and ten protein families for crotalines. There were four dominant protein families: phospholipase A₂s (t
  28. 28A Comprehensive Subcellular Atlas of the Toxoplasma Proteome via hyperLOPIT Provides Spatial Context for Protein Functions.Apicomplexan parasites cause major human disease and food insecurity. They owe their considerable success to highly specialized cell compartments and structures. These adaptations drive their recognition, nondestructive penetration, and elaborate reengineering of the host's cells to promote their growth, dissemination, and the countering of host defenses. The evolution of unique apicomplexan cellular compartments is concomitant with vast proteomic novelty. Consequently, half of apicomplexan proteins are unique and uncharacterized. Here, we determine the steady-state subcellular location of thousands of proteins simultaneously within the globally prevalent apicomplexan parasite Toxoplasma gondii. This provides unprecedented comprehensive molecular definition of these unicellular eukaryotes and their specialized compartments, and these data reveal the spatial organizations of protein expression and function, adaptation to hosts, and the underlying evolutionary trajectories of these patho
  29. 29FungiDB: An Integrated Bioinformatic Resource for Fungi and Oomycetes.FungiDB (fungidb.org) is a free online resource for data mining and functional genomics analysis for fungal and oomycete species. FungiDB is part of the Eukaryotic Pathogen Genomics Database Resource (EuPathDB, eupathdb.org) platform that integrates genomic, transcriptomic, proteomic, and phenotypic datasets, and other types of data for pathogenic and nonpathogenic, free-living and parasitic organisms. FungiDB is one of the largest EuPathDB databases containing nearly 100 genomes obtained from GenBank, Aspergillus Genome Database (AspGD), The Broad Institute, Joint Genome Institute (JGI), Ensembl, and other sources. FungiDB offers a user-friendly web interface with embedded bioinformatics tools that support custom in silico experiments that leverage FungiDB-integrated data. In addition, a Galaxy-based workspace enables users to generate custom pipelines for large-scale data analysis (e.g., RNA-Seq, variant calling, etc.). This review provides an introduction to the FungiDB resources an
  30. 30A proteomics sample metadata representation for multiomics integration and big data analysis.The amount of public proteomics data is rapidly increasing but there is no standardized format to describe the sample metadata and their relationship with the dataset files in a way that fully supports their understanding or reanalysis. Here we propose to develop the transcriptomics data format MAGE-TAB into a standard representation for proteomics sample metadata. We implement MAGE-TAB-Proteomics in a crowdsourcing project to manually curate over 200 public datasets. We also describe tools and libraries to validate and submit sample metadata-related information to the PRIDE repository. We expect that these developments will improve the reproducibility and facilitate the reanalysis and integration of public proteomics datasets.
  31. 31Approaches to Integrating Metabolomics and Multi-Omics Data: A Primer.Metabolomics deals with multiple and complex chemical reactions within living organisms and how these are influenced by external or internal perturbations. It lies at the heart of omics profiling technologies not only as the underlying biochemical layer that reflects information expressed by the genome, the transcriptome and the proteome, but also as the closest layer to the phenome. The combination of metabolomics data with the information available from genomics, transcriptomics, and proteomics offers unprecedented possibilities to enhance current understanding of biological functions, elucidate their underlying mechanisms and uncover hidden associations between omics variables. As a result, a vast array of computational tools have been developed to assist with integrative analysis of metabolomics data with different omics. Here, we review and propose five criteria-hypothesis, data types, strategies, study design and study focus- to classify statistical multi-omics data integration a
  32. 32Multi-Omics Pipeline and Omics-Integration Approach to Decipher Plant's Abiotic Stress Tolerance Responses.The present day's ongoing global warming and climate change adversely affect plants through imposing environmental (abiotic) stresses and disease pressure. The major abiotic factors such as drought, heat, cold, salinity, etc., hamper a plant's innate growth and development, resulting in reduced yield and quality, with the possibility of undesired traits. In the 21st century, the advent of high-throughput sequencing tools, state-of-the-art biotechnological techniques and bioinformatic analyzing pipelines led to the easy characterization of plant traits for abiotic stress response and tolerance mechanisms by applying the 'omics' toolbox. Panomics pipeline including genomics, transcriptomics, proteomics, metabolomics, epigenomics, proteogenomics, interactomics, ionomics, phenomics, etc., have become very handy nowadays. This is important to produce climate-smart future crops with a proper understanding of the molecular mechanisms of abiotic stress responses by the plant's genes, transcrip
  33. 33Current state and future prospects of spatial biology in colorectal cancer.Over the past century, colorectal cancer (CRC) has become one of the most devastating cancers impacting the human population. To gain a deeper understanding of the molecular mechanisms driving this solid tumor, researchers have increasingly turned their attention to the tumor microenvironment (TME). Spatial transcriptomics and proteomics have emerged as a particularly powerful technology for deciphering the complexity of CRC tumors, given that the TME and its spatial organization are critical determinants of disease progression and treatment response. Spatial transcriptomics enables high-resolution mapping of the whole transcriptome. While spatial proteomics maps protein expression and function across tissue sections. Together, they provide a detailed view of the molecular landscape and cellular interactions within the TME. In this review, we delve into recent advances in spatial biology technologies applied to CRC research, highlighting both the methodologies and the challenges associ
  34. 34UniProt: the universal protein knowledgebase in 2021.The aim of the UniProt Knowledgebase is to provide users with a comprehensive, high-quality and freely accessible set of protein sequences annotated with functional information. In this article, we describe significant updates that we have made over the last two years to the resource. The number of sequences in UniProtKB has risen to approximately 190 million, despite continued work to reduce sequence redundancy at the proteome level. We have adopted new methods of assessing proteome completeness and quality. We continue to extract detailed annotations from the literature to add to reviewed entries and supplement these in unreviewed entries with annotations provided by automated systems such as the newly implemented Association-Rule-Based Annotator (ARBA). We have developed a credit-based publication submission interface to allow the community to contribute publications and annotations to UniProt entries. We describe how UniProtKB responded to the COVID-19 pandemic through expert curat
  35. 35Highly accurate protein structure prediction for the human proteome.Protein structures can provide invaluable information, both for reasoning about biological processes and for enabling interventions such as structure-based drug development or targeted mutagenesis. After decades of effort, 17% of the total residues in human protein sequences are covered by an experimentally determined structure 1 . Here we markedly expand the structural coverage of the proteome by applying the state-of-the-art machine learning method, AlphaFold 2 , at a scale that covers almost the entire human proteome (98.5% of human proteins). The resulting dataset covers 58% of residues with a confident prediction, of which a subset (36% of all residues) have very high confidence. We introduce several metrics developed by building on the AlphaFold model and use them to interpret the dataset, identifying strong multi-domain predictions as well as regions that are likely to be disordered. Finally, we provide some case studies to illustrate how high-quality predictions could be used t
  36. 36Plasma proteomic associations with genetics and health in the UK Biobank.The Pharma Proteomics Project is a precompetitive biopharmaceutical consortium characterizing the plasma proteomic profiles of 54,219 UK Biobank participants. Here we provide a detailed summary of this initiative, including technical and biological validations, insights into proteomic disease signatures, and prediction modelling for various demographic and health indicators. We present comprehensive protein quantitative trait locus (pQTL) mapping of 2,923 proteins that identifies 14,287 primary genetic associations, of which 81% are previously undescribed, alongside ancestry-specific pQTL mapping in non-European individuals. The study provides an updated characterization of the genetic architecture of the plasma proteome, contextualized with projected pQTL discovery rates as sample sizes and proteomic assay coverages increase over time. We offer extensive insights into trans pQTLs across multiple biological domains, highlight genetic influences on ligand-receptor interactions and pathw
  37. 37Metal nanoparticles: understanding the mechanisms behind antibacterial activity.As the field of nanomedicine emerges, there is a lag in research surrounding the topic of nanoparticle (NP) toxicity, particularly concerned with mechanisms of action. The continuous emergence of bacterial resistance has challenged the research community to develop novel antibiotic agents. Metal NPs are among the most promising of these because show strong antibacterial activity. This review summarizes and discusses proposed mechanisms of antibacterial action of different metal NPs. These mechanisms of bacterial killing include the production of reactive oxygen species, cation release, biomolecule damages, ATP depletion, and membrane interaction. Finally, a comprehensive analysis of the effects of NPs on the regulation of genes and proteins (transcriptomic and proteomic) profiles is discussed.
  38. 38Data-independent acquisition-based SWATH-MS for quantitative proteomics: a tutorial.Many research questions in fields such as personalized medicine, drug screens or systems biology depend on obtaining consistent and quantitatively accurate proteomics data from many samples. SWATH-MS is a specific variant of data-independent acquisition (DIA) methods and is emerging as a technology that combines deep proteome coverage capabilities with quantitative consistency and accuracy. In a SWATH-MS measurement, all ionized peptides of a given sample that fall within a specified mass range are fragmented in a systematic and unbiased fashion using rather large precursor isolation windows. To analyse SWATH-MS data, a strategy based on peptide-centric scoring has been established, which typically requires prior knowledge about the chromatographic and mass spectrometric behaviour of peptides of interest in the form of spectral libraries and peptide query parameters. This tutorial provides guidelines on how to set up and plan a SWATH-MS experiment, how to perform the mass spectrometric
  39. 39Revisiting biomarker discovery by plasma proteomics.Clinical analysis of blood is the most widespread diagnostic procedure in medicine, and blood biomarkers are used to categorize patients and to support treatment decisions. However, existing biomarkers are far from comprehensive and often lack specificity and new ones are being developed at a very slow rate. As described in this review, mass spectrometry (MS)-based proteomics has become a powerful technology in biological research and it is now poised to allow the characterization of the plasma proteome in great depth. Previous "triangular strategies" aimed at discovering single biomarker candidates in small cohorts, followed by classical immunoassays in much larger validation cohorts. We propose a "rectangular" plasma proteome profiling strategy, in which the proteome patterns of large cohorts are correlated with their phenotypes in health and disease. Translating such concepts into clinical practice will require restructuring several aspects of diagnostic decision-making, and we disc
  40. 40NMR Spectroscopy for Metabolomics Research.Over the past two decades, nuclear magnetic resonance (NMR) has emerged as one of the three principal analytical techniques used in metabolomics (the other two being gas chromatography coupled to mass spectrometry (GC-MS) and liquid chromatography coupled with single-stage mass spectrometry (LC-MS)). The relative ease of sample preparation, the ability to quantify metabolite levels, the high level of experimental reproducibility, and the inherently nondestructive nature of NMR spectroscopy have made it the preferred platform for long-term or large-scale clinical metabolomic studies. These advantages, however, are often outweighed by the fact that most other analytical techniques, including both LC-MS and GC-MS, are inherently more sensitive than NMR, with lower limits of detection typically being 10 to 100 times better. This review is intended to introduce readers to the field of NMR-based metabolomics and to highlight both the advantages and disadvantages of NMR spectroscopy for metab
  41. 41The ProteomeXchange consortium at 10 years: 2023 update.Mass spectrometry (MS) is by far the most used experimental approach in high-throughput proteomics. The ProteomeXchange (PX) consortium of proteomics resources (http://www.proteomexchange.org) was originally set up to standardize data submission and dissemination of public MS proteomics data. It is now 10 years since the initial data workflow was implemented. In this manuscript, we describe the main developments in PX since the previous update manuscript in Nucleic Acids Research was published in 2020. The six members of the Consortium are PRIDE, PeptideAtlas (including PASSEL), MassIVE, jPOST, iProX and Panorama Public. We report the current data submission statistics, showcasing that the number of datasets submitted to PX resources has continued to increase every year. As of June 2022, more than 34 233 datasets had been submitted to PX resources, and from those, 20 062 (58.6%) just in the last three years. We also report the development of the Universal Spectrum Identifiers and the i
  42. 42metaX: a flexible and comprehensive software for processing metabolomics data.Background Non-targeted metabolomics based on mass spectrometry enables high-throughput profiling of the metabolites in a biological sample. The large amount of data generated from mass spectrometry requires intensive computational processing for annotation of mass spectra and identification of metabolites. Computational analysis tools that are fully integrated with multiple functions and are easily operated by users who lack extensive knowledge in programing are needed in this research field. Results We herein developed an R package, metaX, that is capable of end-to-end metabolomics data analysis through a set of interchangeable modules. Specifically, metaX provides several functions, such as peak picking and annotation, data quality assessment, missing value imputation, data normalization, univariate and multivariate statistics, power analysis and sample size estimation, receiver operating characteristic analysis, biomarker selection, pathway annotation, correlation network analysis,
  43. 43Large-scale deep multi-layer analysis of Alzheimer's disease brain reveals strong proteomic disease-related changes not observed at the RNA level.The biological processes that are disrupted in the Alzheimer's disease (AD) brain remain incompletely understood. In this study, we analyzed the proteomes of more than 1,000 brain tissues to reveal new AD-related protein co-expression modules that were highly preserved across cohorts and brain regions. Nearly half of the protein co-expression modules, including modules significantly altered in AD, were not observed in RNA networks from the same cohorts and brain regions, highlighting the proteopathic nature of AD. Two such AD-associated modules unique to the proteomic network included a module related to MAPK signaling and metabolism and a module related to the matrisome. The matrisome module was influenced by the APOE ε4 allele but was not related to the rate of cognitive decline after adjustment for neuropathology. By contrast, the MAPK/metabolism module was strongly associated with the rate of cognitive decline. Disease-associated modules unique to the proteome are sources of promis
  44. 44Organ aging signatures in the plasma proteome track health and disease.Animal studies show aging varies between individuals as well as between organs within an individual 1-4 , but whether this is true in humans and its effect on age-related diseases is unknown. We utilized levels of human blood plasma proteins originating from specific organs to measure organ-specific aging differences in living individuals. Using machine learning models, we analysed aging in 11 major organs and estimated organ age reproducibly in five independent cohorts encompassing 5,676 adults across the human lifespan. We discovered nearly 20% of the population show strongly accelerated age in one organ and 1.7% are multi-organ agers. Accelerated organ aging confers 20-50% higher mortality risk, and organ-specific diseases relate to faster aging of those organs. We find individuals with accelerated heart aging have a 250% increased heart failure risk and accelerated brain and vascular aging predict Alzheimer's disease (AD) progression independently from and as strongly as plasma pTa
  45. 45IonQuant Enables Accurate and Sensitive Label-Free Quantification With FDR-Controlled Match-Between-Runs.Missing values weaken the power of label-free quantitative proteomic experiments to uncover true quantitative differences between biological samples or experimental conditions. Match-between-runs (MBR) has become a common approach to mitigate the missing value problem, where peptides identified by tandem mass spectra in one run are transferred to another by inference based on m/z, charge state, retention time, and ion mobility when applicable. Though tolerances are used to ensure such transferred identifications are reasonably located and meet certain quality thresholds, little work has been done to evaluate the statistical confidence of MBR. Here, we present a mixture model-based approach to estimate the false discovery rate (FDR) of peptide and protein identification transfer, which we implement in the label-free quantification tool IonQuant. Using several benchmarking datasets generated on both Orbitrap and timsTOF mass spectrometers, we demonstrate superior performance of IonQuant
  46. 46Software Tools and Approaches for Compound Identification of LC-MS/MS Data in Metabolomics.The annotation of small molecules remains a major challenge in untargeted mass spectrometry-based metabolomics. We here critically discuss structured elucidation approaches and software that are designed to help during the annotation of unknown compounds. Only by elucidating unknown metabolites first is it possible to biologically interpret complex systems, to map compounds to pathways and to create reliable predictive metabolic models for translational and clinical research. These strategies include the construction and quality of tandem mass spectral databases such as the coalition of MassBank repositories and investigations of MS/MS matching confidence. We present in silico fragmentation tools such as MS-FINDER, CFM-ID, MetFrag, ChemDistiller and CSI:FingerID that can annotate compounds from existing structure databases and that have been used in the CASMI (critical assessment of small molecule identification) contests. Furthermore, the use of retention time models from liquid chrom
  47. 47Large-scale plasma proteomics comparisons through genetics and disease associations.High-throughput proteomics platforms measuring thousands of proteins in plasma combined with genomic and phenotypic information have the power to bridge the gap between the genome and diseases. Here we performed association studies of Olink Explore 3072 data generated by the UK Biobank Pharma Proteomics Project 1 on plasma samples from more than 50,000 UK Biobank participants with phenotypic and genotypic data, stratifying on British or Irish, African and South Asian ancestries. We compared the results with those of a SomaScan v4 study on plasma from 36,000 Icelandic people 2 , for 1,514 of whom Olink data were also available. We found modest correlation between the two platforms. Although cis protein quantitative trait loci were detected for a similar absolute number of assays on the two platforms (2,101 on Olink versus 2,120 on SomaScan), the proportion of assays with such supporting evidence for assay performance was higher on the Olink platform (72% versus 43%). A considerable numb
  48. 48Ultra-high sensitivity mass spectrometry quantifies single-cell proteome changes upon perturbation.Single-cell technologies are revolutionizing biology but are today mainly limited to imaging and deep sequencing. However, proteins are the main drivers of cellular function and in-depth characterization of individual cells by mass spectrometry (MS)-based proteomics would thus be highly valuable and complementary. Here, we develop a robust workflow combining miniaturized sample preparation, very low flow-rate chromatography, and a novel trapped ion mobility mass spectrometer, resulting in a more than 10-fold improved sensitivity. We precisely and robustly quantify proteomes and their changes in single, FACS-isolated cells. Arresting cells at defined stages of the cell cycle by drug treatment retrieves expected key regulators. Furthermore, it highlights potential novel ones and allows cell phase prediction. Comparing the variability in more than 430 single-cell proteomes to transcriptome data revealed a stable-core proteome despite perturbation, while the transcriptome appears stochasti
  49. 49Genome, transcriptome and proteome: the rise of omics data and their integration in biomedical sciences.Advances in the technologies and informatics used to generate and process large biological data sets (omics data) are promoting a critical shift in the study of biomedical sciences. While genomics, transcriptomics and proteinomics, coupled with bioinformatics and biostatistics, are gaining momentum, they are still, for the most part, assessed individually with distinct approaches generating monothematic rather than integrated knowledge. As other areas of biomedical sciences, including metabolomics, epigenomics and pharmacogenomics, are moving towards the omics scale, we are witnessing the rise of inter-disciplinary data integration strategies to support a better understanding of biological systems and eventually the development of successful precision medicine. This review cuts across the boundaries between genomics, transcriptomics and proteomics, summarizing how omics data are generated, analysed and shared, and provides an overview of the current strengths and weaknesses of this glo
  50. 50Deep Visual Proteomics defines single-cell identity and heterogeneity.Despite the availabilty of imaging-based and mass-spectrometry-based methods for spatial proteomics, a key challenge remains connecting images with single-cell-resolution protein abundance measurements. Here, we introduce Deep Visual Proteomics (DVP), which combines artificial-intelligence-driven image analysis of cellular phenotypes with automated single-cell or single-nucleus laser microdissection and ultra-high-sensitivity mass spectrometry. DVP links protein abundance to complex cellular or subcellular phenotypes while preserving spatial context. By individually excising nuclei from cell culture, we classified distinct cell states with proteomic profiles defined by known and uncharacterized proteins. In an archived primary melanoma tissue, DVP identified spatially resolved proteome changes as normal melanocytes transition to fully invasive melanoma, revealing pathways that change in a spatial manner as cancer progresses, such as mRNA splicing dysregulation in metastatic vertical gr
  51. 51Inflammation-Related Mechanisms in Chronic Kidney Disease Prediction, Progression, and Outcome.Persistent, low-grade inflammation is now considered a hallmark feature of chronic kidney disease (CKD), being involved in the development of all-cause mortality of these patients. Although substantial improvements have been made in clinical care, CKD remains a major public health burden, affecting 10-15% of the population, and its prevalence is constantly growing. Due to its insidious nature, CKD is rarely diagnosed in early stages, and once developed, its progression is unfortunately irreversible. There are many factors that contribute to the setting of the inflammatory status in CKD, including increased production of proinflammatory cytokines, oxidative stress and acidosis, chronic and recurrent infections, altered metabolism of adipose tissue, and last but not least, gut microbiota dysbiosis, an underestimated source of microinflammation. In this scenario, a huge step forward was made by the increasing progression of omics approaches, specially designed for identification of biomar
  52. 52Single-cell proteomic and transcriptomic analysis of macrophage heterogeneity using SCoPE2.Background Macrophages are innate immune cells with diverse functional and molecular phenotypes. This diversity is largely unexplored at the level of single-cell proteomes because of the limitations of quantitative single-cell protein analysis. Results To overcome this limitation, we develop SCoPE2, which substantially increases quantitative accuracy and throughput while lowering cost and hands-on time by introducing automated and miniaturized sample preparation. These advances enable us to analyze the emergence of cellular heterogeneity as homogeneous monocytes differentiate into macrophage-like cells in the absence of polarizing cytokines. SCoPE2 quantifies over 3042 proteins in 1490 single monocytes and macrophages in 10 days of instrument time, and the quantified proteins allow us to discern single cells by cell type. Furthermore, the data uncover a continuous gradient of proteome states for the macrophages, suggesting that macrophage heterogeneity may emerge in the absence of pola
  53. 53Persistent clotting protein pathology in Long COVID/Post-Acute Sequelae of COVID-19 (PASC) is accompanied by increased levels of antiplasmin.Background Severe acute respiratory syndrome coronavirus 2 (SARS-Cov-2)-induced infection, the cause of coronavirus disease 2019 (COVID-19), is characterized by acute clinical pathologies, including various coagulopathies that may be accompanied by hypercoagulation and platelet hyperactivation. Recently, a new COVID-19 phenotype has been noted in patients after they have ostensibly recovered from acute COVID-19 symptoms. This new syndrome is commonly termed Long COVID/Post-Acute Sequelae of COVID-19 (PASC). Here we refer to it as Long COVID/PASC. Lingering symptoms persist for as much as 6 months (or longer) after acute infection, where COVID-19 survivors complain of recurring fatigue or muscle weakness, being out of breath, sleep difficulties, and anxiety or depression. Given that blood clots can block microcapillaries and thereby inhibit oxygen exchange, we here investigate if the lingering symptoms that individuals with Long COVID/PASC manifest might be due to the presence of persis
  54. 54Metallodrugs are unique: opportunities and challenges of discovery and development.Metals play vital roles in nutrients and medicines and provide chemical functionalities that are not accessible to purely organic compounds. At least 10 metals are essential for human life and about 46 other non-essential metals (including radionuclides) are also used in drug therapies and diagnostic agents. These include platinum drugs (in 50% of cancer chemotherapies), lithium (bipolar disorders), silver (antimicrobials), and bismuth (broad-spectrum antibiotics). While the quest for novel and better drugs is now as urgent as ever, drug discovery and development pipelines established for organic drugs and based on target identification and high-throughput screening of compound libraries are less effective when applied to metallodrugs. Metallodrugs are often prodrugs which undergo activation by ligand substitution or redox reactions, and are multi-targeting, all of which need to be considered when establishing structure-activity relationships. We focus on early-stage in vitro drug disc
  55. 55Evolving concepts in bone infection: redefining "biofilm", "acute vs. chronic osteomyelitis", "the immune proteome" and "local antibiotic therapy".Osteomyelitis is a devastating disease caused by microbial infection of bone. While the frequency of infection following elective orthopedic surgery is low, rates of reinfection are disturbingly high. Staphylococcus aureus is responsible for the majority of chronic osteomyelitis cases and is often considered to be incurable due to bacterial persistence deep within bone. Unfortunately, there is no consensus on clinical classifications of osteomyelitis and the ensuing treatment algorithm. Given the high patient morbidity, mortality, and economic burden caused by osteomyelitis, it is important to elucidate mechanisms of bone infection to inform novel strategies for prevention and curative treatment. Recent discoveries in this field have identified three distinct reservoirs of bacterial biofilm including: Staphylococcal abscess communities in the local soft tissue and bone marrow, glycocalyx formation on implant hardware and necrotic tissue, and colonization of the osteocyte-lacuno canalic
  56. 56Profiling the human intestinal environment under physiological conditions.The spatiotemporal structure of the human microbiome 1,2 , proteome 3 and metabolome 4,5 reflects and determines regional intestinal physiology and may have implications for disease 6 . Yet, little is known about the distribution of microorganisms, their environment and their biochemical activity in the gut because of reliance on stool samples and limited access to only some regions of the gut using endoscopy in fasting or sedated individuals 7 . To address these deficiencies, we developed an ingestible device that collects samples from multiple regions of the human intestinal tract during normal digestion. Collection of 240 intestinal samples from 15 healthy individuals using the device and subsequent multi-omics analyses identified significant differences between bacteria, phages, host proteins and metabolites in the intestines versus stool. Certain microbial taxa were differentially enriched and prophage induction was more prevalent in the intestines than in stool. The host proteome
  57. 57A novel affinity-based method for the isolation of highly purified extracellular vesicles.Extracellular vesicles (EVs) such as exosomes and microvesicles serve as messengers of intercellular network, allowing exchange of cellular components between cells. EVs carry lipids, proteins, and RNAs derived from their producing cells, and have potential as biomarkers specific to cell types and even cellular states. However, conventional methods (such as ultracentrifugation or polymeric precipitation) for isolating EVs have disadvantages regarding purity and feasibility. Here, we have developed a novel method for EV purification by using Tim4 protein, which specifically binds the phosphatidylserine displayed on the surface of EVs. Because the binding is Ca 2+ -dependent, intact EVs can be easily released from Tim4 by adding Ca 2+ chelators. Tim4 purification, which we have applied to cell conditioned media and biofluids, is capable of yielding EVs of a higher purity than those obtained using conventional methods. The lower contamination found in Tim4-purified EV preparations allows
  58. 58The Need for Multi-Omics Biomarker Signatures in Precision Medicine.Recent advances in omics technologies have led to unprecedented efforts characterizing the molecular changes that underlie the development and progression of a wide array of complex human diseases, including cancer. As a result, multi-omics analyses-which take advantage of these technologies in genomics, transcriptomics, epigenomics, proteomics, metabolomics, and other omics areas-have been proposed and heralded as the key to advancing precision medicine in the clinic. In the field of precision oncology, genomics approaches, and, more recently, other omics analyses have helped reveal several key mechanisms in cancer development, treatment resistance, and recurrence risk, and several of these findings have been implemented in clinical oncology to help guide treatment decisions. However, truly integrated multi-omics analyses have not been applied widely, preventing further advances in precision medicine. Additional efforts are needed to develop the analytical infrastructure necessary to
  59. 59dia-PASEF data analysis using FragPipe and DIA-NN for deep proteomics of low sample amounts.The dia-PASEF technology uses ion mobility separation to reduce signal interferences and increase sensitivity in proteomic experiments. Here we present a two-dimensional peak-picking algorithm and generation of optimized spectral libraries, as well as take advantage of neural network-based processing of dia-PASEF data. Our computational platform boosts proteomic depth by up to 83% compared to previous work, and is specifically beneficial for fast proteomic experiments and those with low sample amounts. It quantifies over 5300 proteins in single injections recorded at 200 samples per day throughput using Evosep One chromatography system on a timsTOF Pro mass spectrometer and almost 9000 proteins in single injections recorded with a 93-min nanoflow gradient on timsTOF Pro 2, from 200 ng of HeLa peptides. A user-friendly implementation is provided through the incorporation of the algorithms in the DIA-NN software and by the FragPipe workflow for spectral library generation.
  60. 60Single-Cell Multiomics: Multiple Measurements from Single Cells.Single-cell sequencing provides information that is not confounded by genotypic or phenotypic heterogeneity of bulk samples. Sequencing of one molecular type (RNA, methylated DNA or open chromatin) in a single cell, furthermore, provides insights into the cell's phenotype and links to its genotype. Nevertheless, only by taking measurements of these phenotypes and genotypes from the same single cells can such inferences be made unambiguously. In this review, we survey the first experimental approaches that assay, in parallel, multiple molecular types from the same single cell, before considering the challenges and opportunities afforded by these and future technologies.
  61. 61Cancer-associated fibroblast classification in single-cell and spatial proteomics data.Cancer-associated fibroblasts (CAFs) are a diverse cell population within the tumour microenvironment, where they have critical effects on tumour evolution and patient prognosis. To define CAF phenotypes, we analyse a single-cell RNA sequencing (scRNA-seq) dataset of over 16,000 stromal cells from tumours of 14 breast cancer patients, based on which we define and functionally annotate nine CAF phenotypes and one class of pericytes. We validate this classification system in four additional cancer types and use highly multiplexed imaging mass cytometry on matched breast cancer samples to confirm our defined CAF phenotypes at the protein level and to analyse their spatial distribution within tumours. This general CAF classification scheme will allow comparison of CAF phenotypes across studies, facilitate analysis of their functional roles, and potentially guide development of new treatment strategies in the future.
  62. 62Quantitative single-cell proteomics as a tool to characterize cellular hierarchies.Large-scale single-cell analyses are of fundamental importance in order to capture biological heterogeneity within complex cell systems, but have largely been limited to RNA-based technologies. Here we present a comprehensive benchmarked experimental and computational workflow, which establishes global single-cell mass spectrometry-based proteomics as a tool for large-scale single-cell analyses. By exploiting a primary leukemia model system, we demonstrate both through pre-enrichment of cell populations and through a non-enriched unbiased approach that our workflow enables the exploration of cellular heterogeneity within this aberrant developmental hierarchy. Our approach is capable of consistently quantifying ~1000 proteins per cell across thousands of individual cells using limited instrument time. Furthermore, we develop a computational workflow (SCeptre) that effectively normalizes the data, integrates available FACS data and facilitates downstream analysis. The approach presented
  63. 63The SARS-CoV-2 RNA-protein interactome in infected human cells.Characterizing the interactions that SARS-CoV-2 viral RNAs make with host cell proteins during infection can improve our understanding of viral RNA functions and the host innate immune response. Using RNA antisense purification and mass spectrometry, we identified up to 104 human proteins that directly and specifically bind to SARS-CoV-2 RNAs in infected human cells. We integrated the SARS-CoV-2 RNA interactome with changes in proteome abundance induced by viral infection and linked interactome proteins to cellular pathways relevant to SARS-CoV-2 infections. We demonstrated by genetic perturbation that cellular nucleic acid-binding protein (CNBP) and La-related protein 1 (LARP1), two of the most strongly enriched viral RNA binders, restrict SARS-CoV-2 replication in infected cells and provide a global map of their direct RNA contact sites. Pharmacological inhibition of three other RNA interactome members, PPIA, ATP1A1, and the ARP2/3 complex, reduced viral replication in two human cell
  64. 64DAPAR & ProStaR: software to perform statistical analyses in quantitative discovery proteomics.DAPAR and ProStaR are software tools to perform the statistical analysis of label-free XIC-based quantitative discovery proteomics experiments. DAPAR contains procedures to filter, normalize, impute missing value, aggregate peptide intensities, perform null hypothesis significance tests and select the most likely differentially abundant proteins with a corresponding false discovery rate. ProStaR is a graphical user interface that allows friendly access to the DAPAR functionalities through a web browser. Availability and implementation DAPAR and ProStaR are implemented in the R language and are available on the website of the Bioconductor project (http://www.bioconductor.org/). A complete tutorial and a toy dataset are accompanying the packages. Contact samuel.wieczorek@cea.fr, florence.combes@cea.fr, thomas.burger@cea.fr.
  65. 65Gut Microbiota Profiling: Metabolomics Based Approach to Unravel Compounds Affecting Human Health.The gut microbiota is composed of a huge number of different bacteria, that produce a large amount of compounds playing a key role in microbe selection and in the construction of a metabolic signaling network. The microbial activities are affected by environmental stimuli leading to the generation of a wide number of compounds, that influence the host metabolome and human health. Indeed, metabolite profiles related to the gut microbiota can offer deep insights on the impact of lifestyle and dietary factors on chronic and acute diseases. Metagenomics, metaproteomics and metabolomics are some of the meta-omics approaches to study the modulation of the gut microbiota. Metabolomic research applied to biofluids allows to: define the metabolic profile; identify and quantify classes and compounds of interest; characterize small molecules produced by intestinal microbes; and define the biochemical pathways of metabolites. Mass spectrometry and nuclear magnetic resonance spectroscopy are the pr
  66. 66Chemical proteomics approaches for identifying the cellular targets of natural products.Covering: 2010 up to 2016Deconvoluting the mode of action of natural products and drugs remains one of the biggest challenges in chemistry and biology today. Chemical proteomics is a growing area of chemical biology that seeks to design small molecule probes to understand protein function. In the context of natural products, chemical proteomics can be used to identify the protein binding partners or targets of small molecules in live cells. Here, we highlight recent examples of chemical probes based on natural products and their application for target identification. The review focuses on probes that can be covalently linked to their target proteins (either via intrinsic chemical reactivity or via the introduction of photocrosslinkers), and can be applied "in situ" - in living systems rather than cell lysates. We also focus here on strategies that employ a click reaction, the copper-catalysed azide-alkyne cycloaddition reaction (CuAAC), to allow minimal functionalisation of natural pro
  67. 67Rapid and In-Depth Coverage of the (Phospho-)Proteome With Deep Libraries and Optimal Window Design for dia-PASEF.Data-independent acquisition (DIA) methods have become increasingly attractive in mass spectrometry-based proteomics because they enable high data completeness and a wide dynamic range. Recently, we combined DIA with parallel accumulation-serial fragmentation (dia-PASEF) on a Bruker trapped ion mobility (IM) separated quadrupole time-of-flight mass spectrometer. This requires alignment of the IM separation with the downstream mass selective quadrupole, leading to a more complex scheme for dia-PASEF window placement compared with DIA. To achieve high data completeness and deep proteome coverage, here we employ variable isolation windows that are placed optimally depending on precursor density in the m/z and IM plane. This is implemented in the freely available py_diAID (Python package for DIA with an automated isolation design) package. In combination with in-depth project-specific proteomics libraries and the Evosep liquid chromatography system, we reproducibly identified over 7700 pro
  68. 68SARS-CoV-2 simulations go exascale to predict dramatic spike opening and cryptic pockets across the proteome.SARS-CoV-2 has intricate mechanisms for initiating infection, immune evasion/suppression and replication that depend on the structure and dynamics of its constituent proteins. Many protein structures have been solved, but far less is known about their relevant conformational changes. To address this challenge, over a million citizen scientists banded together through the Folding@home distributed computing project to create the first exascale computer and simulate 0.1 seconds of the viral proteome. Our adaptive sampling simulations predict dramatic opening of the apo spike complex, far beyond that seen experimentally, explaining and predicting the existence of 'cryptic' epitopes. Different spike variants modulate the probabilities of open versus closed structures, balancing receptor binding and immune evasion. We also discover dramatic conformational changes across the proteome, which reveal over 50 'cryptic' pockets that expand targeting options for the design of antivirals. All data a
  69. 69DEqMS: A Method for Accurate Variance Estimation in Differential Protein Expression Analysis.Quantitative proteomics by mass spectrometry is widely used in biomarker research and basic biology research for investigation of phenotype level cellular events. Despite the wide application, the methodology for statistical analysis of differentially expressed proteins has not been unified. Various methods such as t test, linear model and mixed effect models are used to define changes in proteomics experiments. However, none of these methods consider the specific structure of MS-data. Choices between methods, often originally developed for other types of data, are based on compromises between features such as statistical power, general applicability and user friendliness. Furthermore, whether to include proteins identified with one peptide in statistical analysis of differential protein expression varies between studies. Here we present DEqMS, a robust statistical method developed specifically for differential protein expression analysis in mass spectrometry data. In all data sets inv
  70. 70Tear fluid biomarkers in ocular and systemic disease: potential use for predictive, preventive and personalised medicine.In the field of predictive, preventive and personalised medicine, researchers are keen to identify novel and reliable ways to predict and diagnose disease, as well as to monitor patient response to therapeutic agents. In the last decade alone, the sensitivity of profiling technologies has undergone huge improvements in detection sensitivity, thus allowing quantification of minute samples, for example body fluids that were previously difficult to assay. As a consequence, there has been a huge increase in tear fluid investigation, predominantly in the field of ocular surface disease. As tears are a more accessible and less complex body fluid (than serum or plasma) and sampling is much less invasive, research is starting to focus on how disease processes affect the proteomic, lipidomic and metabolomic composition of the tear film. By determining compositional changes to tear profiles, crucial pathways in disease progression may be identified, allowing for more predictive and personalised
  71. 71A Critical Review of Bottom-Up Proteomics: The Good, the Bad, and the Future of this Field.Proteomics is the field of study that includes the analysis of proteins, from either a basic science prospective or a clinical one. Proteins can be investigated for their abundance, variety of proteoforms due to post-translational modifications (PTMs), and their stable or transient protein-protein interactions. This can be especially beneficial in the clinical setting when studying proteins involved in different diseases and conditions. Here, we aim to describe a bottom-up proteomics workflow from sample preparation to data analysis, including all of its benefits and pitfalls. We also describe potential improvements in this type of proteomics workflow for the future.
  72. 72The gut microbiome as a target for prevention and treatment of hyperglycaemia in type 2 diabetes: from current human evidence to future possibilities.The totality of microbial genomes in the gut exceeds the size of the human genome, having around 500-fold more genes that importantly complement our coding potential. Microbial genes are essential for key metabolic processes, such as the breakdown of indigestible dietary fibres to short-chain fatty acids, biosynthesis of amino acids and vitamins, and production of neurotransmitters and hormones. During the last decade, evidence has accumulated to support a role for gut microbiota (analysed from faecal samples) in glycaemic control and type 2 diabetes. Mechanistic studies in mice support a causal role for gut microbiota in metabolic diseases, although human data favouring causality is insufficient. As it may be challenging to sort the human evidence from the large number of animal studies in the field, there is a need to provide a review of human studies. Thus, the aim of this review is to cover the current and future possibilities and challenges of using the gut microbiota, with its ca
  73. 73Recent advances in mass spectrometry based clinical proteomics: applications to cancer research.Cancer biomarkers have transformed current practices in the oncology clinic. Continued discovery and validation are crucial for improving early diagnosis, risk stratification, and monitoring patient response to treatment. Profiling of the tumour genome and transcriptome are now established tools for the discovery of novel biomarkers, but alterations in proteome expression are more likely to reflect changes in tumour pathophysiology. In the past, clinical diagnostics have strongly relied on antibody-based detection strategies, but these methods carry certain limitations. Mass spectrometry (MS) is a powerful method that enables increasingly comprehensive insights into changes of the proteome to advance personalized medicine. In this review, recent improvements in MS-based clinical proteomics are highlighted with a focus on oncology. We will provide a detailed overview of clinically relevant samples types, as well as, consideration for sample preparation methods, protein quantitation stra
  74. 74Machine Learning Applications for Mass Spectrometry-Based Metabolomics.The metabolome of an organism depends on environmental factors and intracellular regulation and provides information about the physiological conditions. Metabolomics helps to understand disease progression in clinical settings or estimate metabolite overproduction for metabolic engineering. The most popular analytical metabolomics platform is mass spectrometry (MS). However, MS metabolome data analysis is complicated, since metabolites interact nonlinearly, and the data structures themselves are complex. Machine learning methods have become immensely popular for statistical analysis due to the inherent nonlinear data representation and the ability to process large and heterogeneous data rapidly. In this review, we address recent developments in using machine learning for processing MS spectra and show how machine learning generates new biological insights. In particular, supervised machine learning has great potential in metabolomics research because of the ability to supply quantitati
  75. 75MSBooster: improving peptide identification rates using deep learning-based features.Peptide identification in liquid chromatography-tandem mass spectrometry (LC-MS/MS) experiments relies on computational algorithms for matching acquired MS/MS spectra against sequences of candidate peptides using database search tools, such as MSFragger. Here, we present a new tool, MSBooster, for rescoring peptide-to-spectrum matches using additional features incorporating deep learning-based predictions of peptide properties, such as LC retention time, ion mobility, and MS/MS spectra. We demonstrate the utility of MSBooster, in tandem with MSFragger and Percolator, in several different workflows, including nonspecific searches (immunopeptidomics), direct identification of peptides from data independent acquisition data, single-cell proteomics, and data generated on an ion mobility separation-enabled timsTOF MS platform. MSBooster is fast, robust, and fully integrated into the widely used FragPipe computational platform.
  76. 76Ion identity molecular networking for mass spectrometry-based metabolomics in the GNPS environment.Molecular networking connects mass spectra of molecules based on the similarity of their fragmentation patterns. However, during ionization, molecules commonly form multiple ion species with different fragmentation behavior. As a result, the fragmentation spectra of these ion species often remain unconnected in tandem mass spectrometry-based molecular networks, leading to redundant and disconnected sub-networks of the same compound classes. To overcome this bottleneck, we develop Ion Identity Molecular Networking (IIMN) that integrates chromatographic peak shape correlation analysis into molecular networks to connect and collapse different ion species of the same molecule. The new feature relationships improve network connectivity for structurally related molecules, can be used to reveal unknown ion-ligand complexes, enhance annotation within molecular networks, and facilitate the expansion of spectral reference libraries. IIMN is integrated into various open source feature finding too
  77. 77Beyond mass spectrometry, the next step in proteomics.Proteins can be the root cause of a disease, and they can be used to cure it. The need to identify these critical actors was recognized early (1951) by Sanger; the first biopolymer sequenced was a peptide, insulin. With the advent of scalable, single-molecule DNA sequencing, genomics and transcriptomics have since propelled medicine through improved sensitivity and lower costs, but proteomics has lagged behind. Currently, proteomics relies mainly on mass spectrometry (MS), but instead of truly sequencing, it classifies a protein and typically requires about a billion copies of a protein to do it. Here, we offer a survey that illuminates a few alternatives with the brightest prospects for identifying whole proteins and displacing MS for sequencing them. These alternatives all boast sensitivity superior to MS and promise to be scalable and seem to be adaptable to bioinformatics tools for calling the sequence of amino acids that constitute a protein.
  78. 78MaxDIA enables library-based and library-free data-independent acquisition proteomics.MaxDIA is a software platform for analyzing data-independent acquisition (DIA) proteomics data within the MaxQuant software environment. Using spectral libraries, MaxDIA achieves deep proteome coverage with substantially better coefficients of variation in protein quantification than other software. MaxDIA is equipped with accurate false discovery rate (FDR) estimates on both library-to-DIA match and protein levels, including when using whole-proteome predicted spectral libraries. This is the foundation of discovery DIA-hypothesis-free analysis of DIA samples without library and with reliable FDR control. MaxDIA performs three- or four-dimensional feature detection of fragment data, and scoring of matches is augmented by machine learning on the features of an identification. MaxDIA's bootstrap DIA workflow performs multiple rounds of matching with increasing quality of recalibration and stringency of matching to the library. Combining MaxDIA with two new technologies-BoxCar acquisition
  79. 79Onco-Multi-OMICS Approach: A New Frontier in Cancer Research.The acquisition of cancer hallmarks requires molecular alterations at multiple levels including genome, epigenome, transcriptome, proteome, and metabolome. In the past decade, numerous attempts have been made to untangle the molecular mechanisms of carcinogenesis involving single OMICS approaches such as scanning the genome for cancer-specific mutations and identifying altered epigenetic-landscapes within cancer cells or by exploring the differential expression of mRNA and protein through transcriptomics and proteomics techniques, respectively. While these single-level OMICS approaches have contributed towards the identification of cancer-specific mutations, epigenetic alterations, and molecular subtyping of tumors based on gene/protein-expression, they lack the resolving-power to establish the casual relationship between molecular signatures and the phenotypic manifestation of cancer hallmarks. In contrast, the multi-OMICS approaches involving the interrogation of the cancer cells/tis
  80. 80Multi-omics analysis identifies therapeutic vulnerabilities in triple-negative breast cancer subtypes.Triple-negative breast cancer (TNBC) is a collection of biologically diverse cancers characterized by distinct transcriptional patterns, biology, and immune composition. TNBCs subtypes include two basal-like (BL1, BL2), a mesenchymal (M) and a luminal androgen receptor (LAR) subtype. Through a comprehensive analysis of mutation, copy number, transcriptomic, epigenetic, proteomic, and phospho-proteomic patterns we describe the genomic landscape of TNBC subtypes. Mesenchymal subtype tumors display high mutation loads, genomic instability, absence of immune cells, low PD-L1 expression, decreased global DNA methylation, and transcriptional repression of antigen presentation genes. We demonstrate that major histocompatibility complex I (MHC-I) is transcriptionally suppressed by H3K27me3 modifications by the polycomb repressor complex 2 (PRC2). Pharmacological inhibition of PRC2 subunits EZH2 or EED restores MHC-I expression and enhances chemotherapy efficacy in murine tumor models, providin
  81. 81Quantitative metabolomics by H-NMR and LC-MS/MS confirms altered metabolic pathways in diabetes.Insulin is as a major postprandial hormone with profound effects on carbohydrate, fat, and protein metabolism. In the absence of exogenous insulin, patients with type 1 diabetes exhibit a variety of metabolic abnormalities including hyperglycemia, glycosurea, accelerated ketogenesis, and muscle wasting due to increased proteolysis. We analyzed plasma from type 1 diabetic (T1D) humans during insulin treatment (I+) and acute insulin deprivation (I-) and non-diabetic participants (ND) by (1)H nuclear magnetic resonance spectroscopy and liquid chromatography-tandem mass spectrometry. The aim was to determine if this combination of analytical methods could provide information on metabolic pathways known to be altered by insulin deficiency. Multivariate statistics differentiated proton spectra from I- and I+ based on several derived plasma metabolites that were elevated during insulin deprivation (lactate, acetate, allantoin, ketones). Mass spectrometry revealed significant perturbations in
  82. 82Depression and anxiety in patients with active ulcerative colitis: crosstalk of gut microbiota, metabolomics and proteomics.Patients with ulcerative colitis (UC) have a high prevalence of mental disorders, such as depression and anxiety. Gut microbiota imbalance and disturbed metabolism have been suggested to play an important role in either UC or mental disorders. However, little is known about their detailed multi-omics characteristics in patients with UC and depression/anxiety. In this prospective observational study, 240 Chinese patients were enrolled, including 129 patients with active UC (69 in Phase 1 and 60 in Phase 2; divided into depression/non-depression or anxiety/non-anxiety groups), 49 patients with depression and anxiety (non-UC), and 62 healthy people. The gut microbiota of all subjects was analyzed using 16S rRNA sequencing. The serum metabolome and proteome of patients with UC in Phase 2 were analyzed using liquid chromatography/mass spectrometry. Associations between multi-omics were evaluated by correlation analysis. The prophylactic effect of candidate metabolites on the depressive-like
  83. 83Development, characterization and comparisons of targeted and non-targeted metabolomics methods.The potential of a metabolomics method to detect statistically significant perturbations in the metabolome of an organism is enhanced by excellent analytical precision, unequivocal identification, and broad metabolomic coverage. While the former two metrics are usually associated with targeted metabolomics and the latter with non-targeted metabolomics, a systematic comparison of the performance of both approaches has not yet been carried out. The present work reports on the development and performance evaluation of separate targeted and non-targeted metabolomics methods. The targeted approach facilitated determination of 181 metabolites (quantitative analysis of 18 amino acids, 11 biogenic amines, 5 neurotransmitters, 5 nucleobases and semi-quantitative analysis of 50 carnitines, 83 phosphatidylcholines, and 9 sphingomyelins) using ultra-performance liquid chromatography-tandem mass spectrometry (UPLC-MS/MS) and flow injection-tandem mass spectrometry (FI-MS/MS). Method accuracy and/or
  84. 84Multiomics study of nonalcoholic fatty liver disease.Nonalcoholic fatty liver (NAFL) and its sequelae are growing health problems. We performed a genome-wide association study of NAFL, cirrhosis and hepatocellular carcinoma, and integrated the findings with expression and proteomic data. For NAFL, we utilized 9,491 clinical cases and proton density fat fraction extracted from 36,116 liver magnetic resonance images. We identified 18 sequence variants associated with NAFL and 4 with cirrhosis, and found rare, protective, predicted loss-of-function variants in MTARC1 and GPAM, underscoring them as potential drug targets. We leveraged messenger RNA expression, splicing and predicted coding effects to identify 16 putative causal genes, of which many are implicated in lipid metabolism. We analyzed levels of 4,907 plasma proteins in 35,559 Icelanders and 1,459 proteins in 47,151 UK Biobank participants, identifying multiple proteins involved in disease pathogenesis. We show that proteomics can discriminate between NAFL and cirrhosis. The presen
  85. 85Comprehensive metabolic profiling of Parkinson's disease by liquid chromatography-mass spectrometry.Background Parkinson's disease (PD) is a prevalent neurological disease in the elderly with increasing morbidity and mortality. Despite enormous efforts, rapid and accurate diagnosis of PD is still compromised. Metabolomics defines the final readout of genome-environment interactions through the analysis of the entire metabolic profile in biological matrices. Recently, unbiased metabolic profiling of human sample has been initiated to identify novel PD metabolic biomarkers and dysfunctional metabolic pathways, however, it remains a challenge to define reliable biomarker(s) for clinical use. Methods We presented a comprehensive metabolic evaluation for identifying crucial metabolic disturbances in PD using liquid chromatography-high resolution mass spectrometry-based metabolomics approach. Plasma samples from 3 independent cohorts (n = 460, 223 PD, 169 healthy controls (HCs) and 68 PD-unrelated neurological disease controls) were collected for the characterization of metabolic changes r
  86. 86Novel Biomarkers in the Diagnosis of Chronic Kidney Disease and the Prediction of Its Outcome.In its early stages, symptoms of chronic kidney disease (CKD) are usually not apparent. Significant reduction of the kidney function is the first obvious sign of disease. If diagnosed early (stages 1 to 3), the progression of CKD can be altered and complications reduced. In stages 4 and 5 extensive kidney damage is observed, which usually results in end-stage renal failure. Currently, the diagnosis of CKD is made usually on the levels of blood urea and serum creatinine (sCr), however, sCr has been shown to be lacking high predictive value. Due to the development of genomics, epigenetics, transcriptomics, proteomics, and metabolomics, the introduction of novel techniques will allow for the identification of novel biomarkers in renal diseases. This review presents some new possible biomarkers in the diagnosis of CKD and in the prediction of outcome, including asymmetric dimethylarginine (ADMA), symmetric dimethylarginine (SDMA), uromodulin, kidney injury molecule-1 (KIM-1), neutrophil ge
  87. 87Mass spectrometry imaging for plant biology: a review.Mass spectrometry imaging (MSI) is a developing technique to measure the spatio-temporal distribution of many biomolecules in tissues. Over the preceding decade, MSI has been adopted by plant biologists and applied in a broad range of areas, including primary metabolism, natural products, plant defense, plant responses to abiotic and biotic stress, plant lipids and the developing field of spatial metabolomics. This review covers recent advances in plant-based MSI, general aspects of instrumentation, analytical approaches, sample preparation and the current trends in respective plant research.
  88. 88Arabidopsis plasmodesmal proteome.The multicellular nature of plants requires that cells should communicate in order to coordinate essential functions. This is achieved in part by molecular flux through pores in the cell wall, called plasmodesmata. We describe the proteomic analysis of plasmodesmata purified from the walls of Arabidopsis suspension cells. Isolated plasmodesmata were seen as membrane-rich structures largely devoid of immunoreactive markers for the plasma membrane, endoplasmic reticulum and cytoplasmic components. Using nano-liquid chromatography and an Orbitrap ion-trap tandem mass spectrometer, 1341 proteins were identified. We refer to this list as the plasmodesmata- or PD-proteome. Relative to other cell wall proteomes, the PD-proteome is depleted in wall proteins and enriched for membrane proteins, but still has a significant number (35%) of putative cytoplasmic contaminants, probably reflecting the sensitivity of the proteomic detection system. To validate the PD-proteome we searched for known plas
  89. 89Standardized multi-omics of Earth's microbiomes reveals microbial and metabolite diversity.Despite advances in sequencing, lack of standardization makes comparisons across studies challenging and hampers insights into the structure and function of microbial communities across multiple habitats on a planetary scale. Here we present a multi-omics analysis of a diverse set of 880 microbial community samples collected for the Earth Microbiome Project. We include amplicon (16S, 18S, ITS) and shotgun metagenomic sequence data, and untargeted metabolomics data (liquid chromatography-tandem mass spectrometry and gas chromatography mass spectrometry). We used standardized protocols and analytical methods to characterize microbial communities, focusing on relationships and co-occurrences of microbially related metabolites and microbial taxa across environments, thus allowing us to explore diversity at extraordinary scale. In addition to a reference database for metagenomic and metabolomic data, we provide a framework for incorporating additional studies, enabling the expansion of exis
  90. 90Causal integration of multi-omics data with prior knowledge to generate mechanistic hypotheses.Multi-omics datasets can provide molecular insights beyond the sum of individual omics. Various tools have been recently developed to integrate such datasets, but there are limited strategies to systematically extract mechanistic hypotheses from them. Here, we present COSMOS (Causal Oriented Search of Multi-Omics Space), a method that integrates phosphoproteomics, transcriptomics, and metabolomics datasets. COSMOS combines extensive prior knowledge of signaling, metabolic, and gene regulatory networks with computational methods to estimate activities of transcription factors and kinases as well as network-level causal reasoning. COSMOS provides mechanistic hypotheses for experimental observations across multi-omics datasets. We applied COSMOS to a dataset comprising transcriptomics, phosphoproteomics, and metabolomics data from healthy and cancerous tissue from eleven clear cell renal cell carcinoma (ccRCC) patients. COSMOS was able to capture relevant crosstalks within and between mul
  91. 91Omics-Based Strategies in Precision Medicine: Toward a Paradigm Shift in Inborn Errors of Metabolism Investigations.The rise of technologies that simultaneously measure thousands of data points represents the heart of systems biology. These technologies have had a huge impact on the discovery of next-generation diagnostics, biomarkers, and drugs in the precision medicine era. Systems biology aims to achieve systemic exploration of complex interactions in biological systems. Driven by high-throughput omics technologies and the computational surge, it enables multi-scale and insightful overviews of cells, organisms, and populations. Precision medicine capitalizes on these conceptual and technological advancements and stands on two main pillars: data generation and data modeling. High-throughput omics technologies allow the retrieval of comprehensive and holistic biological information, whereas computational capabilities enable high-dimensional data modeling and, therefore, accessible and user-friendly visualization. Furthermore, bioinformatics has enabled comprehensive multi-omics and clinical data in
  92. 92Integrative analysis of metabolomics and proteomics reveals amino acid metabolism disorder in sepsis.Background Sepsis is defined as a systemic inflammatory response to microbial infections with multiple organ dysfunction. This study analysed untargeted metabolomics combined with proteomics of serum from patients with sepsis to reveal the underlying pathological mechanisms involved in sepsis. Methods A total of 63 patients with sepsis and 43 normal controls were enrolled from a prospective multicentre cohort. The biological functions of the metabolome were assessed by coexpression network analysis. A molecular network based on metabolomics and proteomics data was constructed to investigate the key molecules. Results Untargeted metabolomics analysis revealed widespread dysregulation of amino acid metabolism, which regulates inflammation and immunity, in patients with sepsis. Seventy-three differentially expressed metabolites (|log 2 fold change| > 1.5, adjusted P value 1.5) that could predict sepsis were identified. External validation of the hub metabolites was consistent with the der
  93. 93Single Cell Multi-Omics Technology: Methodology and Application.In the era of precision medicine, multi-omics approaches enable the integration of data from diverse omics platforms, providing multi-faceted insight into the interrelation of these omics layers on disease processes. Single cell sequencing technology can dissect the genotypic and phenotypic heterogeneity of bulk tissue and promises to deepen our understanding of the underlying mechanisms governing both health and disease. Through modification and combination of single cell assays available for transcriptome, genome, epigenome, and proteome profiling, single cell multi-omics approaches have been developed to simultaneously and comprehensively study not only the unique genotypic and phenotypic characteristics of single cells, but also the combined regulatory mechanisms evident only at single cell resolution. In this review, we summarize the state-of-the-art single cell multi-omics methods and discuss their applications, challenges, and future directions.
  94. 94Multi-Omics of Single Cells: Strategies and Applications.Most genome-wide assays provide averages across large numbers of cells, but recent technological advances promise to overcome this limitation. Pioneering single-cell assays are now available for genome, epigenome, transcriptome, proteome, and metabolome profiling. Here, we describe how these different dimensions can be combined into multi-omics assays that provide comprehensive profiles of the same cell.
  95. 95Understanding and Designing the Strategies for the Microbe-Mediated Remediation of Environmental Contaminants Using Omics Approaches.Rapid industrialization and population explosion has resulted in the generation and dumping of various contaminants into the environment. These harmful compounds deteriorate the human health as well as the surrounding environments. Current research aims to harness and enhance the natural ability of different microbes to metabolize these toxic compounds. Microbial-mediated bioremediation offers great potential to reinstate the contaminated environments in an ecologically acceptable approach. However, the lack of the knowledge regarding the factors controlling and regulating the growth, metabolism, and dynamics of diverse microbial communities in the contaminated environments often limits its execution. In recent years the importance of advanced tools such as genomics, proteomics, transcriptomics, metabolomics, and fluxomics has increased to design the strategies to treat these contaminants in ecofriendly manner. Previously researchers has largely focused on the environmental remediation
  96. 96Subcellular Transcriptomics and Proteomics: A Comparative Methods Review.The internal environment of cells is molecularly crowded, which requires spatial organization via subcellular compartmentalization. These compartments harbor specific conditions for molecules to perform their biological functions, such as coordination of the cell cycle, cell survival, and growth. This compartmentalization is also not static, with molecules trafficking between these subcellular neighborhoods to carry out their functions. For example, some biomolecules are multifunctional, requiring an environment with differing conditions or interacting partners, and others traffic to export such molecules. Aberrant localization of proteins or RNA species has been linked to many pathological conditions, such as neurological, cancer, and pulmonary diseases. Differential expression studies in transcriptomics and proteomics are relatively common, but the majority have overlooked the importance of subcellular information. In addition, subcellular transcriptomics and proteomics data do not a
  97. 97Enablers and challenges of spatial omics, a melting pot of technologies.Spatial omics has emerged as a rapidly growing and fruitful field with hundreds of publications presenting novel methods for obtaining spatially resolved information for any omics data type on spatial scales ranging from subcellular to organismal. From a technology development perspective, spatial omics is a highly interdisciplinary field that integrates imaging and omics, spatial and molecular analyses, sequencing and mass spectrometry, and image analysis and bioinformatics. The emergence of this field has not only opened a window into spatial biology, but also created multiple novel opportunities, questions, and challenges for method developers. Here, we provide the perspective of technology developers on what makes the spatial omics field unique. After providing a brief overview of the state of the art, we discuss technological enablers and challenges and present our vision about the future applications and impact of this melting pot.
  98. 98A review of spatial profiling technologies for characterizing the tumor microenvironment in immuno-oncology.Interpreting the mechanisms and principles that govern gene activity and how these genes work according to -their cellular distribution in organisms has profound implications for cancer research. The latest technological advancements, such as imaging-based approaches and next-generation single-cell sequencing technologies, have established a platform for spatial transcriptomics to systematically quantify the expression of all or most genes in the entire tumor microenvironment and explore an array of disease milieus, particularly in tumors. Spatial profiling technologies permit the study of transcriptional activity at the spatial or single-cell level. This multidimensional classification of the transcriptomic and proteomic signatures of tumors, especially the associated immune and stromal cells, facilitates evaluation of tumor heterogeneity, details of the evolutionary trajectory of each tumor, and multifaceted interactions between each tumor cell and its microenvironment. Therefore, sp
  99. 99Quantitative proteomic analysis reveals a simple strategy of global resource allocation in bacteria.A central aim of cell biology was to understand the strategy of gene expression in response to the environment. Here, we study gene expression response to metabolic challenges in exponentially growing Escherichia coli using mass spectrometry. Despite enormous complexity in the details of the underlying regulatory network, we find that the proteome partitions into several coarse-grained sectors, with each sector's total mass abundance exhibiting positive or negative linear relations with the growth rate. The growth rate-dependent components of the proteome fractions comprise about half of the proteome by mass, and their mutual dependencies can be characterized by a simple flux model involving only two effective parameters. The success and apparent generality of this model arises from tight coordination between proteome partition and metabolism, suggesting a principle for resource allocation in proteome economy of the cell. This strategy of global gene regulation should serve as a basis
  100. 100SwissPalm: Protein Palmitoylation database.Protein S-palmitoylation is a reversible post-translational modification that regulates many key biological processes, although the full extent and functions of protein S-palmitoylation remain largely unexplored. Recent developments of new chemical methods have allowed the establishment of palmitoyl-proteomes of a variety of cell lines and tissues from different species. As the amount of information generated by these high-throughput studies is increasing, the field requires centralization and comparison of this information. Here we present SwissPalm ( http://swisspalm.epfl.ch), our open, comprehensive, manually curated resource to study protein S-palmitoylation. It currently encompasses more than 5000 S-palmitoylated protein hits from seven species, and contains more than 500 specific sites of S-palmitoylation. SwissPalm also provides curated information and filters that increase the confidence in true positive hits, and integrates predictions of S-palmitoylated cysteine scores, ortho
  101. 101Effects of diet on resource utilization by a model human gut microbiota containing Bacteroides cellulosilyticus WH2, a symbiont with an extensive glycobiome.The human gut microbiota is an important metabolic organ, yet little is known about how its individual species interact, establish dominant positions, and respond to changes in environmental factors such as diet. In this study, gnotobiotic mice were colonized with an artificial microbiota comprising 12 sequenced human gut bacterial species and fed oscillating diets of disparate composition. Rapid, reproducible, and reversible changes in the structure of this assemblage were observed. Time-series microbial RNA-Seq analyses revealed staggered functional responses to diet shifts throughout the assemblage that were heavily focused on carbohydrate and amino acid metabolism. High-resolution shotgun metaproteomics confirmed many of these responses at a protein level. One member, Bacteroides cellulosilyticus WH2, proved exceptionally fit regardless of diet. Its genome encoded more carbohydrate active enzymes than any previously sequenced member of the Bacteroidetes. Transcriptional profiling i
  102. 102Integration of transcriptomics, proteomics, and metabolomics data to reveal HER2-associated metabolic heterogeneity in gastric cancer with response to immunotherapy and neoadjuvant chemotherapy.Background Currently available prognostic tools and focused therapeutic methods result in unsatisfactory treatment of gastric cancer (GC). A deeper understanding of human epidermal growth factor receptor 2 (HER2)-coexpressed metabolic pathways may offer novel insights into tumour-intrinsic precision medicine. Methods The integrated multi-omics strategies (including transcriptomics, proteomics and metabolomics) were applied to develop a novel metabolic classifier for gastric cancer. We integrated TCGA-STAD cohort (375 GC samples and 56753 genes) and TCPA-STAD cohort (392 GC samples and 218 proteins), and rated them as transcriptomics and proteomics data, resepectively. 224 matched blood samples of GC patients and healthy individuals were collected to carry out untargeted metabolomics analysis. Results In this study, pan-cancer analysis highlighted the crucial role of ERBB2 in the immune microenvironment and metabolic remodelling. In addition, the metabolic landscape of GC indicated that
  103. 103Multi-omics of human plasma reveals molecular features of dysregulated inflammation and accelerated aging in schizophrenia.Schizophrenia is a devastating psychiatric illness that detrimentally affects a significant portion of the worldwide population. Aging of schizophrenia patients is associated with reduced longevity, but the potential biological factors associated with aging in this population have not yet been investigated in a global manner. To address this gap in knowledge, the present study assesses proteomics and metabolomics profiles in the plasma of subjects afflicted with schizophrenia compared to non-psychiatric control patients over six decades of life. Global, unbiased analyses of circulating blood plasma can provide knowledge of prominently dysregulated molecular pathways and their association with schizophrenia, as well as features of aging and gender in this disease. The resulting data compiled in this study represent a compendium of molecular changes associated with schizophrenia over the human lifetime. Supporting the clinical finding of schizophrenia's association with more rapid aging,
  104. 104Integrated proteogenomic characterization of urothelial carcinoma of the bladder.Background Urothelial carcinoma (UC) is the most common pathological type of bladder cancer, a malignant tumor. However, an integrated multi-omics analysis of the Chinese UC patient cohort is lacking. Methods We performed an integrated multi-omics analysis, including whole-exome sequencing, RNA-seq, proteomic, and phosphoproteomic analysis of 116 Chinese UC patients, comprising 45 non-muscle-invasive bladder cancer patients (NMIBCs) and 71 muscle-invasive bladder cancer patients (MIBCs). Result Proteogenomic integration analysis indicated that SND1 and CDK5 amplifications on chromosome 7q were associated with the activation of STAT3, which was relevant to tumor proliferation. Chromosome 5p gain in NMIBC patients was a high-risk factor, through modulating actin cytoskeleton implicating in tumor cells invasion. Phosphoproteomic analysis of tumors and morphologically normal human urothelium produced UC-associated activated kinases, including CDK1 and PRKDC. Proteomic analysis identified t
  105. 105UniProt: a worldwide hub of protein knowledge.The UniProt Knowledgebase is a collection of sequences and annotations for over 120 million proteins across all branches of life. Detailed annotations extracted from the literature by expert curators have been collected for over half a million of these proteins. These annotations are supplemented by annotations provided by rule based automated systems, and those imported from other resources. In this article we describe significant updates that we have made over the last 2 years to the resource. We have greatly expanded the number of Reference Proteomes that we provide and in particular we have focussed on improving the number of viral Reference Proteomes. The UniProt website has been augmented with new data visualizations for the subcellular localization of proteins as well as their structure and interactions. UniProt resources are available under a CC-BY (4.0) license via the web at https://www.uniprot.org/.
  106. 106UniProt: the universal protein knowledgebase.The UniProt knowledgebase is a large resource of protein sequences and associated detailed annotation. The database contains over 60 million sequences, of which over half a million sequences have been curated by experts who critically review experimental and predicted data for each protein. The remainder are automatically annotated based on rule systems that rely on the expert curated knowledge. Since our last update in 2014, we have more than doubled the number of reference proteomes to 5631, giving a greater coverage of taxonomic diversity. We implemented a pipeline to remove redundant highly similar proteomes that were causing excessive redundancy in UniProt. The initial run of this pipeline reduced the number of sequences in UniProt by 47 million. For our users interested in the accessory proteomes, we have made available sets of pan proteome sequences that cover the diversity of sequences for each species that is found in its strains and sub-strains. To help interpretation of geno
  107. 107MZmine 2: modular framework for processing, visualizing, and analyzing mass spectrometry-based molecular profile data.Background Mass spectrometry (MS) coupled with online separation methods is commonly applied for differential and quantitative profiling of biological samples in metabolomic as well as proteomic research. Such approaches are used for systems biology, functional genomics, and biomarker discovery, among others. An ongoing challenge of these molecular profiling approaches, however, is the development of better data processing methods. Here we introduce a new generation of a popular open-source data processing toolbox, MZmine 2. Results A key concept of the MZmine 2 software design is the strict separation of core functionality and data processing modules, with emphasis on easy usability and support for high-resolution spectra processing. Data processing modules take advantage of embedded visualization tools, allowing for immediate previews of parameter settings. Newly introduced functionality includes the identification of peaks using online databases, MSn data support, improved isotope p
  108. 108MitoCarta2.0: an updated inventory of mammalian mitochondrial proteins.Mitochondria are complex organelles that house essential pathways involved in energy metabolism, ion homeostasis, signalling and apoptosis. To understand mitochondrial pathways in health and disease, it is crucial to have an accurate inventory of the organelle's protein components. In 2008, we made substantial progress toward this goal by performing in-depth mass spectrometry of mitochondria from 14 organs, epitope tagging/microscopy and Bayesian integration to assemble MitoCarta (www.broadinstitute.org/pubs/MitoCarta): an inventory of genes encoding mitochondrial-localized proteins and their expression across 14 mouse tissues. Using the same strategy we have now reconstructed this inventory separately for human and for mouse based on (i) improved gene transcript models, (ii) updated literature curation, including results from proteomic analyses of mitochondrial sub-compartments, (iii) improved homology mapping and (iv) updated versions of all seven original data sets. The updated huma
  109. 109Online Parallel Accumulation-Serial Fragmentation (PASEF) with a Novel Trapped Ion Mobility Mass Spectrometer.In bottom-up proteomics, peptides are separated by liquid chromatography with elution peak widths in the range of seconds, whereas mass spectra are acquired in about 100 microseconds with time-of-flight (TOF) instruments. This allows adding ion mobility as a third dimension of separation. Among several formats, trapped ion mobility spectrometry (TIMS) is attractive because of its small size, low voltage requirements and high efficiency of ion utilization. We have recently demonstrated a scan mode termed parallel accumulation - serial fragmentation (PASEF), which multiplies the sequencing speed without any loss in sensitivity (Meier et al. , PMID: 26538118). Here we introduce the timsTOF Pro instrument, which optimally implements online PASEF. It features an orthogonal ion path into the ion mobility device, limiting the amount of debris entering the instrument and making it very robust in daily operation. We investigate different precursor selection schemes for shotgun proteomics to opt
  110. 110The ProteomeXchange consortium in 2017: supporting the cultural change in proteomics public data deposition.The ProteomeXchange (PX) Consortium of proteomics resources (http://www.proteomexchange.org) was formally started in 2011 to standardize data submission and dissemination of mass spectrometry proteomics data worldwide. We give an overview of the current consortium activities and describe the advances of the past few years. Augmenting the PX founding members (PRIDE and PeptideAtlas, including the PASSEL resource), two new members have joined the consortium: MassIVE and jPOST. ProteomeCentral remains as the common data access portal, providing the ability to search for data sets in all participating PX resources, now with enhanced data visualization components.We describe the updated submission guidelines, now expanded to include four members instead of two. As demonstrated by data submission statistics, PX is supporting a change in culture of the proteomics field: public data sharing is now an accepted standard, supported by requirements for journal submissions resulting in public data
  111. 111SCoPE-MS: mass spectrometry of single mammalian cells quantifies proteome heterogeneity during cell differentiation.Some exciting biological questions require quantifying thousands of proteins in single cells. To achieve this goal, we develop Single Cell ProtEomics by Mass Spectrometry (SCoPE-MS) and validate its ability to identify distinct human cancer cell types based on their proteomes. We use SCoPE-MS to quantify over a thousand proteins in differentiating mouse embryonic stem cells. The single-cell proteomes enable us to deconstruct cell populations and infer protein abundance relationships. Comparison between single-cell proteomes and transcriptomes indicates coordinated mRNA and protein covariation, yet many genes exhibit functionally concerted and distinct regulatory patterns at the mRNA and the protein level.
  112. 112Connecting genetic risk to disease end points through the human blood plasma proteome.Genome-wide association studies (GWAS) with intermediate phenotypes, like changes in metabolite and protein levels, provide functional evidence to map disease associations and translate them into clinical applications. However, although hundreds of genetic variants have been associated with complex disorders, the underlying molecular pathways often remain elusive. Associations with intermediate traits are key in establishing functional links between GWAS-identified risk-variants and disease end points. Here we describe a GWAS using a highly multiplexed aptamer-based affinity proteomics platform. We quantify 539 associations between protein levels and gene variants (pQTLs) in a German cohort and replicate over half of them in an Arab and Asian cohort. Fifty-five of the replicated pQTLs are located in trans. Our associations overlap with 57 genetic risk loci for 42 unique disease end points. We integrate this information into a genome-proteome network and provide an interactive web-tool
  113. 113The ProteomeXchange consortium in 2020: enabling 'big data' approaches in proteomics.The ProteomeXchange (PX) consortium of proteomics resources (http://www.proteomexchange.org) has standardized data submission and dissemination of mass spectrometry proteomics data worldwide since 2012. In this paper, we describe the main developments since the previous update manuscript was published in Nucleic Acids Research in 2017. Since then, in addition to the four PX existing members at the time (PRIDE, PeptideAtlas including the PASSEL resource, MassIVE and jPOST), two new resources have joined PX: iProX (China) and Panorama Public (USA). We first describe the updated submission guidelines, now expanded to include six members. Next, with current data submission statistics, we demonstrate that the proteomics field is now actively embracing public open data policies. At the end of June 2019, more than 14 100 datasets had been submitted to PX resources since 2012, and from those, more than 9 500 in just the last three years. In parallel, an unprecedented increase of data re-use ac
  114. 114VolcaNoseR is a web app for creating, exploring, labeling and sharing volcano plots.Comparative genome- and proteome-wide screens yield large amounts of data. To efficiently present such datasets and to simplify the identification of hits, the results are often presented in a type of scatterplot known as a volcano plot, which shows a measure of effect size versus a measure of significance. The data points with the largest effect size and a statistical significance beyond a user-defined threshold are considered as hits. Such hits are usually annotated in the plot by a label with their name. Volcano plots can represent ten thousands of data points, of which typically only a handful is annotated. The information of data that is not annotated is hardly or not accessible. To simplify access to the data and enable its re-use, we have developed an open source and online web tool with R/Shiny. The web app is named VolcaNoseR and it can be used to create, explore, label and share volcano plots ( https://huygens.science.uva.nl/VolcaNoseR ). When the data is stored in an online
  115. 115Integrated Proteogenomic Characterization of Clear Cell Renal Cell Carcinoma.To elucidate the deregulated functional modules that drive clear cell renal cell carcinoma (ccRCC), we performed comprehensive genomic, epigenomic, transcriptomic, proteomic, and phosphoproteomic characterization of treatment-naive ccRCC and paired normal adjacent tissue samples. Genomic analyses identified a distinct molecular subgroup associated with genomic instability. Integration of proteogenomic measurements uniquely identified protein dysregulation of cellular mechanisms impacted by genomic alterations, including oxidative phosphorylation-related metabolism, protein translation processes, and phospho-signaling modules. To assess the degree of immune infiltration in individual tumors, we identified microenvironment cell signatures that delineated four immune-based ccRCC subtypes characterized by distinct cellular pathways. This study reports a large-scale proteogenomic analysis of ccRCC to discern the functional impact of genomic alterations and provides evidence for rational tre
  116. 1161D and 2D annotation enrichment: a statistical method integrating quantitative proteomics with complementary high-throughput data.Quantitative proteomics now provides abundance ratios for thousands of proteins upon perturbations. These need to be functionally interpreted and correlated to other types of quantitative genome-wide data such as the corresponding transcriptome changes. We describe a new method, 2D annotation enrichment, which compares quantitative data from any two 'omics' types in the context of categorical annotation of the proteins or genes. Suitable genome-wide categories are membership of proteins in biochemical pathways, their annotation with gene ontology terms, sub-cellular localization, presence of protein domains or membership in protein complexes. 2D annotation enrichment detects annotation terms whose members show consistent behavior in one or both of the data dimensions. This consistent behavior can be a correlation between the two data types, such as simultaneous up- or down-regulation in both data dimensions, or a lack thereof, such as regulation in one dimension but no change in the ot
  117. 117Global, quantitative and dynamic mapping of protein subcellular localization.Subcellular localization critically influences protein function, and cells control protein localization to regulate biological processes. We have developed and applied Dynamic Organellar Maps, a proteomic method that allows global mapping of protein translocation events. We initially used maps statically to generate a database with localization and absolute copy number information for over 8700 proteins from HeLa cells, approaching comprehensive coverage. All major organelles were resolved, with exceptional prediction accuracy (estimated at >92%). Combining spatial and abundance information yielded an unprecedented quantitative view of HeLa cell anatomy and organellar composition, at the protein level. We subsequently demonstrated the dynamic capabilities of the approach by capturing translocation events following EGF stimulation, which we integrated into a quantitative model. Dynamic Organellar Maps enable the proteome-wide analysis of physiological protein movements, without requirin
  118. 118Chromatogram libraries improve peptide detection and quantification by data independent acquisition mass spectrometry.Data independent acquisition (DIA) mass spectrometry is a powerful technique that is improving the reproducibility and throughput of proteomics studies. Here, we introduce an experimental workflow that uses this technique to construct chromatogram libraries that capture fragment ion chromatographic peak shape and retention time for every detectable peptide in a proteomics experiment. These coordinates calibrate protein databases or spectrum libraries to a specific mass spectrometer and chromatography setup, facilitating DIA-only pipelines and the reuse of global resource libraries. We also present EncyclopeDIA, a software tool for generating and searching chromatogram libraries, and demonstrate the performance of our workflow by quantifying proteins in human and yeast cells. We find that by exploiting calibrated retention time and fragmentation specificity in chromatogram libraries, EncyclopeDIA can detect 20-25% more peptides from DIA experiments than with data dependent acquisition-b
  119. 119Quantification of microenvironmental metabolites in murine cancers reveals determinants of tumor nutrient availability.Cancer cell metabolism is heavily influenced by microenvironmental factors, including nutrient availability. Therefore, knowledge of microenvironmental nutrient levels is essential to understand tumor metabolism. To measure the extracellular nutrient levels available to tumors, we utilized quantitative metabolomics methods to measure the absolute concentrations of >118 metabolites in plasma and tumor interstitial fluid, the extracellular fluid that perfuses tumors. Comparison of nutrient levels in tumor interstitial fluid and plasma revealed that the nutrients available to tumors differ from those present in circulation. Further, by comparing interstitial fluid nutrient levels between autochthonous and transplant models of murine pancreatic and lung adenocarcinoma, we found that tumor type, anatomical location and animal diet affect local nutrient availability. These data provide a comprehensive characterization of the nutrients present in the tumor microenvironment of widely used mode
  120. 120Ultra-High-Throughput Clinical Proteomics Reveals Classifiers of COVID-19 Infection.The COVID-19 pandemic is an unprecedented global challenge, and point-of-care diagnostic classifiers are urgently required. Here, we present a platform for ultra-high-throughput serum and plasma proteomics that builds on ISO13485 standardization to facilitate simple implementation in regulated clinical laboratories. Our low-cost workflow handles up to 180 samples per day, enables high precision quantification, and reduces batch effects for large-scale and longitudinal studies. We use our platform on samples collected from a cohort of early hospitalized cases of the SARS-CoV-2 pandemic and identify 27 potential biomarkers that are differentially expressed depending on the WHO severity grade of COVID-19. They include complement factors, the coagulation system, inflammation modulators, and pro-inflammatory factors upstream and downstream of interleukin 6. All protocols and software for implementing our approach are freely available. In total, this work supports the development of routine
  121. 121Multi-laboratory assessment of reproducibility, qualitative and quantitative performance of SWATH-mass spectrometry.Quantitative proteomics employing mass spectrometry is an indispensable tool in life science research. Targeted proteomics has emerged as a powerful approach for reproducible quantification but is limited in the number of proteins quantified. SWATH-mass spectrometry consists of data-independent acquisition and a targeted data analysis strategy that aims to maintain the favorable quantitative characteristics (accuracy, sensitivity, and selectivity) of targeted proteomics at large scale. While previous SWATH-mass spectrometry studies have shown high intra-lab reproducibility, this has not been evaluated between labs. In this multi-laboratory evaluation study including 11 sites worldwide, we demonstrate that using SWATH-mass spectrometry data acquisition we can consistently detect and reproducibly quantify >4000 proteins from HEK293 cells. Using synthetic peptide dilution series, we show that the sensitivity, dynamic range and reproducibility established with SWATH-mass spectrometry are u
  122. 122Heavy Metal Tolerance in Plants: Role of Transcriptomics, Proteomics, Metabolomics, and Ionomics.Heavy metal contamination of soil and water causing toxicity/stress has become one important constraint to crop productivity and quality. This situation has further worsened by the increasing population growth and inherent food demand. It has been reported in several studies that counterbalancing toxicity due to heavy metal requires complex mechanisms at molecular, biochemical, physiological, cellular, tissue, and whole plant level, which might manifest in terms of improved crop productivity. Recent advances in various disciplines of biological sciences such as metabolomics, transcriptomics, proteomics, etc., have assisted in the characterization of metabolites, transcription factors, and stress-inducible proteins involved in heavy metal tolerance, which in turn can be utilized for generating heavy metal-tolerant crops. This review summarizes various tolerance strategies of plants under heavy metal toxicity covering the role of metabolites (metabolomics), trace elements (ionomics), tra
  123. 123Detailed analysis of the plasma extracellular vesicle proteome after separation from lipoproteins.The isolation of extracellular vesicles (EVs) from blood is of great importance to understand the biological role of circulating EVs and to develop EVs as biomarkers of disease. Due to the concurrent presence of lipoprotein particles, however, blood is one of the most difficult body fluids to isolate EVs from. The aim of this study was to develop a robust method to isolate and characterise EVs from blood with minimal contamination by plasma proteins and lipoprotein particles. Plasma and serum were collected from healthy subjects, and EVs were isolated by size-exclusion chromatography (SEC), with most particles being present in fractions 8-12, while the bulk of the plasma proteins was present in fractions 11-28. Vesicle markers peaked in fractions 7-11; however, the same fractions also contained lipoprotein particles. The purity of EVs was improved by combining a density cushion with SEC to further separate lipoprotein particles from the vesicles, which reduced the contamination of lipo
  124. 124Nanodroplet processing platform for deep and quantitative proteome profiling of 10-100 mammalian cells.Nanoscale or single-cell technologies are critical for biomedical applications. However, current mass spectrometry (MS)-based proteomic approaches require samples comprising a minimum of thousands of cells to provide in-depth profiling. Here, we report the development of a nanoPOTS (nanodroplet processing in one pot for trace samples) platform for small cell population proteomics analysis. NanoPOTS enhances the efficiency and recovery of sample processing by downscaling processing volumes to 3000 proteins are consistently identified from as few as 10 cells. Furthermore, we demonstrate quantification of ~2400 proteins from single human pancreatic islet thin sections from type 1 diabetic and control donors, illustrating the application of nanoPOTS for spatially resolved proteome measurements from clinical tissues.
  125. 125Plasma proteomic signature of age in healthy humans.To characterize the proteomic signature of chronological age, 1,301 proteins were measured in plasma using the SOMAscan assay (SomaLogic, Boulder, CO, USA) in a population of 240 healthy men and women, 22-93 years old, who were disease- and treatment-free and had no physical and cognitive impairment. Using a p ≤ 3.83 × 10 -5 significance threshold, 197 proteins were positively associated, and 20 proteins were negatively associated with age. Growth differentiation factor 15 (GDF15) had the strongest, positive association with age (GDF15; 0.018 ± 0.001, p = 7.49 × 10 -56 ). In our sample, GDF15 was not associated with other cardiovascular risk factors such as cholesterol or inflammatory markers. The functional pathways enriched in the 217 age-associated proteins included blood coagulation, chemokine and inflammatory pathways, axon guidance, peptidase activity, and apoptosis. Using elastic net regression models, we created a proteomic signature of age based on relative concentrations of 7
  126. 126Status of large-scale analysis of post-translational modifications by mass spectrometry.Cellular function can be controlled through the gene expression program, but often protein post-translational modifications (PTMs) provide a more precise and elegant mechanism. Key functional roles of specific modification events--for instance, during the cell cycle--have been known for decades, but only in the past 10 years has mass-spectrometry-(MS)-based proteomics begun to reveal the true extent of the PTM universe. In this overview for the special PTM issue of Molecular and Cellular Proteomics, we take stock of where MS-based proteomics stands in the large-scale analysis of protein modifications. For many PTMs, including phosphorylation, ubiquitination, glycosylation, and acetylation, tens of thousands of sites can now be confidently identified and localized in the sequence of the protein. The quantification of PTM levels between different cellular states is likewise established, with label-free methods showing particular promise. It is also becoming possible to determine the abso
  127. 127Missing Value Imputation Approach for Mass Spectrometry-based Metabolomics Data.Missing values exist widely in mass-spectrometry (MS) based metabolomics data. Various methods have been applied for handling missing values, but the selection can significantly affect following data analyses. Typically, there are three types of missing values, missing not at random (MNAR), missing at random (MAR), and missing completely at random (MCAR). Our study comprehensively compared eight imputation methods (zero, half minimum (HM), mean, median, random forest (RF), singular value decomposition (SVD), k-nearest neighbors (kNN), and quantile regression imputation of left-censored data (QRILC)) for different types of missing values using four metabolomics datasets. Normalized root mean squared error (NRMSE) and NRMSE-based sum of ranks (SOR) were applied to evaluate imputation accuracy. Principal component analysis (PCA)/partial least squares (PLS)-Procrustes analysis were used to evaluate the overall sample distribution. Student's t-test followed by correlation analysis was condu
  128. 128Modeling aspects of the language of life through transfer-learning protein sequences.Background Predicting protein function and structure from sequence is one important challenge for computational biology. For 26 years, most state-of-the-art approaches combined machine learning and evolutionary information. However, for some applications retrieving related proteins is becoming too time-consuming. Additionally, evolutionary information is less powerful for small families, e.g. for proteins from the Dark Proteome. Both these problems are addressed by the new methodology introduced here. Results We introduced a novel way to represent protein sequences as continuous vectors (embeddings) by using the language model ELMo taken from natural language processing. By modeling protein sequences, ELMo effectively captured the biophysical properties of the language of life from unlabeled big data (UniRef50). We refer to these new embeddings as SeqVec (Sequence-to-Vector) and demonstrate their effectiveness by training simple neural networks for two different tasks. At the per-resid
  129. 129An Optimized Shotgun Strategy for the Rapid Generation of Comprehensive Human Proteomes.This study investigates the challenge of comprehensively cataloging the complete human proteome from a single-cell type using mass spectrometry (MS)-based shotgun proteomics. We modify a classical two-dimensional high-resolution reversed-phase peptide fractionation scheme and optimize a protocol that provides sufficient peak capacity to saturate the sequencing speed of modern MS instruments. This strategy enables the deepest proteome of a human single-cell type to date, with the HeLa proteome sequenced to a depth of ∼584,000 unique peptide sequences and ∼14,200 protein isoforms (∼12,200 protein-coding genes). This depth is comparable with next-generation RNA sequencing and enables the identification of post-translational modifications, including ∼7,000 N-acetylation sites and ∼10,000 phosphorylation sites, without the need for enrichment. We further demonstrate the general applicability and clinical potential of this proteomics strategy by comprehensively quantifying global proteome ex
  130. 130Pan-cancer molecular subtypes revealed by mass-spectrometry-based proteomic characterization of more than 500 human cancers.Mass-spectrometry-based proteomic profiling of human cancers has the potential for pan-cancer analyses to identify molecular subtypes and associated pathway features that might be otherwise missed using transcriptomics. Here, we classify 532 cancers, representing six tissue-based types (breast, colon, ovarian, renal, uterine), into ten proteome-based, pan-cancer subtypes that cut across tumor lineages. The proteome-based subtypes are observable in external cancer proteomic datasets surveyed. Gene signatures of oncogenic or metabolic pathways can further distinguish between the subtypes. Two distinct subtypes both involve the immune system, one associated with the adaptive immune response and T-cell activation, and the other associated with the humoral immune response. Two additional subtypes each involve the tumor stroma, one of these including the collagen VI interacting network. Three additional proteome-based subtypes-respectively involving proteins related to Golgi apparatus, hemog
  131. 131The Mount Sinai cohort of large-scale genomic, transcriptomic and proteomic data in Alzheimer's disease.Alzheimer's disease (AD) affects half the US population over the age of 85 and is universally fatal following an average course of 10 years of progressive cognitive disability. Genetic and genome-wide association studies (GWAS) have identified about 33 risk factor genes for common, late-onset AD (LOAD), but these risk loci fail to account for the majority of affected cases and can neither provide clinically meaningful prediction of development of AD nor offer actionable mechanisms. This cohort study generated large-scale matched multi-Omics data in AD and control brains for exploring novel molecular underpinnings of AD. Specifically, we generated whole genome sequencing, whole exome sequencing, transcriptome sequencing and proteome profiling data from multiple regions of 364 postmortem control, mild cognitive impaired (MCI) and AD brains with rich clinical and pathophysiological data. All the data went through rigorous quality control. Both the raw and processed data are publicly avail
  132. 132Analytical methods in untargeted metabolomics: state of the art in 2015.Metabolomics comprises the methods and techniques that are used to measure the small molecule composition of biofluids and tissues, and is actually one of the most rapidly evolving research fields. The determination of the metabolomic profile - the metabolome - has multiple applications in many biological sciences, including the developing of new diagnostic tools in medicine. Recent technological advances in nuclear magnetic resonance and mass spectrometry are significantly improving our capacity to obtain more data from each biological sample. Consequently, there is a need for fast and accurate statistical and bioinformatic tools that can deal with the complexity and volume of the data generated in metabolomic studies. In this review, we provide an update of the most commonly used analytical methods in metabolomics, starting from raw data processing and ending with pathway analysis and biomarker identification. Finally, the integration of metabolomic profiles with molecular data from
  133. 133Systematic Comparison of Two Animal-to-Human Transmitted Human Coronaviruses: SARS-CoV-2 and SARS-CoV.After the outbreak of the severe acute respiratory syndrome (SARS) in the world in 2003, human coronaviruses (HCoVs) have been reported as pathogens that cause severe symptoms in respiratory tract infections. Recently, a new emerged HCoV isolated from the respiratory epithelium of unexplained pneumonia patients in the Wuhan seafood market caused a major disease outbreak and has been named the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). This virus causes acute lung symptoms, leading to a condition that has been named as "coronavirus disease 2019" (COVID-19). The emergence of SARS-CoV-2 and of SARS-CoV caused widespread fear and concern and has threatened global health security. There are some similarities and differences in the epidemiology and clinical features between these two viruses and diseases that are caused by these viruses. The goal of this work is to systematically review and compare between SARS-CoV and SARS-CoV-2 in the context of their virus incubation, o
  134. 134Optimization of Experimental Parameters in Data-Independent Mass Spectrometry Significantly Increases Depth and Reproducibility of Results.Comprehensive, reproducible and precise analysis of large sample cohorts is one of the key objectives of quantitative proteomics. Here, we present an implementation of data-independent acquisition using its parallel acquisition nature that surpasses the limitation of serial MS2 acquisition of data-dependent acquisition on a quadrupole ultra-high field Orbitrap mass spectrometer. In deep single shot data-independent acquisition, we identified and quantified 6,383 proteins in human cell lines using 2-or-more peptides/protein and over 7100 proteins when including the 717 proteins that were identified on the basis of a single peptide sequence. 7739 proteins were identified in mouse tissues using 2-or-more peptides/protein and 8121 when including the 382 proteins that were identified based on a single peptide sequence. Missing values for proteins were within 0.3 to 2.1% and median coefficients of variation of 4.7 to 6.2% among technical triplicates. In very complex mixtures, we could quanti
  135. 135Characterisation of the transcriptome and proteome of SARS-CoV-2 reveals a cell passage induced in-frame deletion of the furin-like cleavage site from the spike glycoprotein.Background SARS-CoV-2 is a recently emerged respiratory pathogen that has significantly impacted global human health. We wanted to rapidly characterise the transcriptomic, proteomic and phosphoproteomic landscape of this novel coronavirus to provide a fundamental description of the virus's genomic and proteomic potential. Methods We used direct RNA sequencing to determine the transcriptome of SARS-CoV-2 grown in Vero E6 cells which is widely used to propagate the novel coronavirus. The viral transcriptome was analysed using a recently developed ORF-centric pipeline. Allied to this, we used tandem mass spectrometry to investigate the proteome and phosphoproteome of the same virally infected cells. Results Our integrated analysis revealed that the viral transcripts (i.e. subgenomic mRNAs) generally fitted the expected transcription model for coronaviruses. Importantly, a 24 nt in-frame deletion was detected in over half of the subgenomic mRNAs encoding the spike (S) glycoprotein and was
蛋白质组循证手册 · GeniOmics