Malatang without sesame paste
1 Introduction
Genomic analysis of cancer has greatly promoted the identification of driver gene mutations and the development of targeted therapies. However, the functions of most somatic mutations and copy number variants in tumors are still unclear, and there is a lack of clear understanding of the causes of resistance to targeted therapies and methods to solve drug resistance. Now mass spectrometry-based proteomics can directly examine the consequences of genomic abnormalities, providing deeper insights into oncology. The integration of proteins and their post-translational modifications with genomic, epigenomic and transcriptome data constitutes a new field of proteomics. The following review describes recent developments in proteomics and how proteomics is used to study cancer.
2 Background knowledge
Cancer proteomics focuses on combining mass spectrometer (MS)-based protein abundance measurements and post-translational modifications (PTMs) with genomic, epigenomic and transcriptomic data from clinical cancer models and tumor samples. Analysis of multiple sets of data in cancer proteomics has brought new knowledge about biology and improved our understanding of malignant tumors and treatment. Genomics and epigenomics provide the cellular blueprint for what may happen, while proteomics provides decisions about what has happened, as proteins and their modifications are the ultimate arbiters of biological phenotypes .
Over the years, numerous studies have identified drivers (EGFR, ERBB2 (also known as HER2), BCR-ABL, ALK, and BRAF) that influence targeted therapies in cancers such as metastatic breast cancer, chronic myelogenous leukemia, non-small cell lung cancer, and melanoma. However, the vast majority of somatic mutations are passenger mutations with no specific oncogenic function. Distinguishing driver and passenger mutations is still a difficult problem, and researchers generally rely on recurring statistics to determine whether a gene is a driver. However, the number of recurring point mutations, small insertions, or small deletions in a single tumor is small, and those mutated genes that appear infrequently are not easy to find or treat with drugs, which limits the development of targeted therapies. Amplification or loss of chromosomal copy number can affect multiple genes, but the consequences of these changes are poorly understood and difficult to model. Because cancer cells have the ability to find alternative pathways for cell survival and growth, researchers frequently study resistance to targeted therapies. Recently, immuno-oncology approaches have considered the interaction between the cancer genome and the immune system to be critical. Proteogenomics also provides new perspectives in this regard.
- Typical proteomics data
Proteomics technology is mainly based on liquid-mass spectrometry (LC-MS/MS) and is widely used in cancer research. The figure below shows the expected number of features (copy number variation (CNA), DNA methylation ) measured by current technologies (clinical data; WGS, whole-genome sequencing ; WXS, whole-exome sequencing; methylation array; RNA-seq, RNA sequencing; sequencing). Typically, data from tumor tissue samples are compared to normal adjacent tissue (NAT) samples, using a variety of data types for analysis. Post-translational modifications (PTMs) data are generated by in-depth LC-MS analysis of enriched samples, which allows for higher detection sensitivity. In-depth analysis of tumor samples is now commonly performed using multiplexed isobaric chemical labeling reagents (TMTPro, capable of up to 18 samples).
Recently, a minimally invasive method has emerged: DNA, RNA and proteome analysis can be performed from a single needle biopsy.Researchers use a new sample preparation method, that is, complete proteome analysis in FFPE samples (paraffin slices or paraffin block samples formed by tissue samples fixed in formalin and embedded in paraffin, which are common biological materials in the medical field). However, due to the influence of the embedding process itself, PTMs analysis still has limitations. In addition, researchers continue to develop various methods, such as FAIMS, SPS-MS3 and real-time data analysis, to improve the depth, accuracy and reproducibility of proteome measurements. Finally, Data Independent Acquisition (DAI), a technique that can reduce missing values in proteomic data across samples and datasets, has begun to be applied to multi-omics and proteomics studies (Figure 1).
Figure 1 Typical proteomics data
- Using cloud computing mode to analyze proteome data
Cloud computing (Cloud computing) mode is where people access shared computing resources according to their own needs. These resources can be quickly configured for use and can be published after data analysis. As scientists delve deeper into the research, the size of the data sets becomes larger and larger, and more and more algorithms are used to analyze these data. Cloud computing is a type of distributed computing, which refers to decomposing huge data computing processing programs into countless small programs through the network "cloud", and then processing and analyzing these small programs through a system composed of multiple servers to obtain results and return them to the user. In the early days of cloud computing, to put it simply, it was simple distributed computing, solving task distribution and merging calculation results. Therefore, cloud computing is also called grid computing . Through this technology, tens of thousands of data can be processed in a very short time (a few seconds), thereby achieving powerful network services. National Institutes of Health (NIH) Data Sharing Initiative is a "cloud"-based initiative to share genomic and proteomic data. The National Cancer Institute (NCI) Genomics Cloud Pilot has created a shared cloud-based computing environment. The Terra cloud platform can quickly analyze massive amounts of medical data. Leveraging Terra's infrastructure, scientists have developed platforms such as PANOPLY to automate the analysis of proteogenomic data. Using these platforms we can rapidly analyze proteomes to propose relevant disease hypotheses, which can be further explored through additional analyzes and laboratory experiments (Figure 2).
Figure 2 Proteomic data analysis through cloud computing
- Proteomic landscape in cancer
Over the years, researchers have published a series of proteome studies for different cancer types. These studies systematically integrate and analyze genome, transcriptome, group, whole proteome, and PTMs data to further gain biological insights into disease pathogenesis and identify therapeutic targets for each cancer type. In short, scientific researchers have established commonly used statistical methods to refine cancer subtypes, explore the immune microenvironment, infer new antigen profiles, explore the coordinated regulation of DNA, RNA, protein and PTM levels and the specific regulation of proteome or PTM, and ultimately determine cancer drivers and therapeutic targets. The following are these published proteogenomic studies on cancer (Table 1).
Table 1 Conceptual insights into proteogenomic studies
- Computational methods and tools for proteomics
General processing of raw genomic data aims at identifying mutations in somatic and germline cells, quantifying copy number deviations and transcriptome profiling. So computing tools generally need to have the following functions: reading various components, quality control, variant identification and RNA identification.For example, the Genome Analysis Toolkit (GATK) is a tool set developed by the Broad Institute that can discover diversity sites. It is mainly used to analyze SNP (Single Nucleotide Polymorphisms) and indel (insertdelete) in DNA sequencing and RNA sequencing data. Nowadays, scientific researchers most commonly use LC-MS/MS, which is liquid phase chromatography tandem mass spectrometry technology, for the separation, purification and qualitative and quantitative identification of mixtures. Analysis of raw LC-MS/MS proteome data can produce qualitative and quantitative proteins, phosphorylation sites, acetyl sites and ubiquitination sites using MS-GF+, Spectrum Mill, MaxQuant, MSFragger, Philosopher, and CDAP. A special problem exists in proteomics: a greater degree of missing values in the data matrix due to the stochastic nature of peptide sampling and the limited dynamic range of RNA-seq measurements. This phenomenon is especially evident in the PTMs data set. Researchers have proposed various imputation strategies for proteomic data, hoping to solve this problem computationally. DIA, for example, can alleviate the problem of losing data. Commonly used proteomics computational tools are selected in the figure below (Figure 3).
Figure 3 Proteomics analysis tools
- Scientific research ideas related to proteomics
Proteomics (Proteomics) is a combination of the words protein and genomics, which means "a complete set of proteins expressed by a genome". It is a science based on comprehensive research on protein properties and exploring disease mechanisms, cellular patterns, functional connections, etc. at the protein level. In recent years, with the rapid development of proteomics technology, it has been applied to all aspects of life science research. So how to use proteomics technology for scientific research.
① Proteomics identifies new potential therapeutic targets
Scientists have discovered that the incidence of early-onset cancer has increased sharply around the world since 1990, and there is an urgent need to develop therapies to treat cancer. Cancer has become a common disease in our daily lives. With the development of medical technology and science and technology, clinical treatment of malignant tumors is no longer limited to surgery, radiotherapy, and chemotherapy. More and more new methods such as molecular targeted therapy, immunotherapy, , etc. are used in clinical practice, and the results are very good, which greatly improves the success rate of clinical treatment of tumors and prolongs the life and quality of life of tumor patients. proteomic analysis can determine transcriptomic and proteomic changes at high temporal resolution, accelerating the identification of pathogenicity-related biological pathways and the search for potential drug targets.
Article 1 below evaluates the expression of 384 total and post-translationally modified proteins in 871 patients with CLL (chronic lymphocytic leukemia) and MSBL (diffuse large B-cell lymphoma) and combines it with clinical data with the goal of improving diagnostic and therapeutic strategies. This study is the largest proteomics study of CLL to date. This study identifies protein expression signatures (SGs) in relapsed cancers that classify CLL and identify a new subset. These characteristics have a strong predictive effect on time to first treatment (TTFT) and overall survival (OS). The authors also identified a novel SG with hairy cell leukemia-like proteomic signatures but poor treatment response. Protein expression signatures can surrogate other prognosis (such as Rai stage, IGHV status) and predict response to modern treatments (such as the BTK inhibitor ). The addition of SGs and PFGs provides new drug targets for the disease and identifies the best candidate markers for "Watch Wait" (a strategy in which patients avoid or delay more aggressive treatments by receiving close follow-up and monitoring, and watch waiting in a narrow sense is limited to whether to perform surgery after neoadjuvant treatment) and early intervention. In summary, proteomics holds great promise in improving cancer classification, treatment strategy selection, and identifying new therapeutic targets (Figure 4).
Figure 4 Potential therapeutic targets for CLL
② Selection of biomarkers for treatment and prognosis
Biomarkers are crucial for accurate clinical assessment, which includes disease diagnosis, prognosis and selection of optimal therapeutic intervention. Cancer is the second leading cause of death from non-communicable diseases globally, and most biomarker research currently focuses on cancer. Biomarkers can be not only molecular signatures, but also histological , radiological or physiological features. Common biomarkers are based on clinical symptoms, signs, physiological and laboratory test results, among which histology, radiology and magnetic resonance imaging biomarkers rely on doctors' subjective interpretation. The detection accuracy of the marker group is much higher than that of one or several proteins . More and more research focus is shifting from single biomarkers to biomarker groups, and even integrated with existing clinical indicators. high-throughput protein mass spectrometry technology generates large data sets based on clinical samples, which can serve as a rich resource for re-mining biomarkers.
Breast cancer is one of the most common cancers in women. Both men and women are likely to suffer from breast cancer, but the incidence rate in women is much higher than that in men. Despite advances in genomic classification of breast cancer that have significantly improved diagnosis and treatment, current clinical testing and treatment decisions are often based on protein-level information. The authors below 2 performed a comprehensive proteomic analysis of clinically archived FFPE breast cancer surgical specimens from 300 confirmed patients (75 of each PAM50 subtype), quantifying a total of 4214 proteins and characterizing breast cancer heterogeneity in a clinically applicable manner. Researchers identified immune differences, ECM, and lipid metabolism pathways that may be clinically relevant. Among aggressive PAM50-classified basaloid carcinomas, proteomic analysis identified a group with an immune "hot" expression signature and optimal survival. In 88 triple-negative breast cancers, four protein groups showed basal immune "hot," basal immune "cold," mesenchymal-like, and luminal-like characteristics, with different survival outcomes. Proteomic analysis identified potential biomarkers and therapeutic targets: the combined expression of the antigen-presenting proteins TAP1 and HLADQA1 was associated with patient survival, and our study also provides a resource for clinical breast cancer classification (Figure 5).
Figure 5 Proteomic analysis of FFPE breast cancer tissue specimens
③ Using multi-omics methods to redefine and characterize cancer subtypes
Modern oncology believes that tumors are no longer a disease, but a type of disease. For the same cancer type (named after the cancer organ), due to the complexity of tumor pathogenesis, there is a high degree of heterogeneity in both histopathology and molecular biology .
Tumors are usually complex diseases involving multiple genes. Different stages have different gene expression profiles. The heredity, individual differences and complexity of molecular mechanisms of cancer require its characteristics to be described by gene groups or gene clusters. Cancer subtypes are primarily defined by clinical, genomic, or transcriptional characteristics. Multi-omics approaches enable further refinement or redefinition of subtypes based on underlying biology and/or outcomes. This more refined subtyping of cancers helps us tailor treatment strategies and evaluate efficacy. For example, breast cancer has subtypes such as lumA, lumB, basal, and HER2. Among them, TNBC (triple negative breast cancer) can be further subdivided into 3 to 7 subtypes. Of course, now has the support of proteomics, and the cell subtypes will become increasingly clear . Classification will also be more complex if multi-omics data are to be integrated.
Genomic studies of lung adenocarcinoma (LUAD) improve our understanding of the biology of the disease and accelerate targeted therapies. However, the proteomic properties of LUAD remain poorly understood. The following article3 conducted a comprehensive proteomic analysis of LUAD in 103 Chinese patients.Integrated analysis of proteomics, phosphoproteomics, transcriptome, and whole-exome sequencing data revealed cancer-related signatures such as tumor-associated protein variants, unique proteomic signatures, and clinical outcomes in early-stage patients or those with EGFR and TP53 mutations. Proteomic-based stratification of LUAD revealed three subtypes (S-I, S-II and S-III) associated with distinct clinical and molecular features. Furthermore, the authors proposed potential drug targets and validated plasma protein levels of HSP90β as a potential prognostic biomarker for LUAD. The comprehensive proteomic analysis of this study could provide a more comprehensive understanding of the molecular structure of LUAD and provide opportunities for more precise diagnosis and treatment (Figure 6).
Figure 6 Multi-omics pattern of LUAD samples
Three Full text summary
Although some related studies on proteomics have been published, proteomics is still an emerging field. Similar to early studies in genomics, batch analysis of hundreds of tumors may yield important and actionable insights into tumor biology, but more data are needed to identify less common or less impactful oncogenic processes, as well as to use different methods to explore the heterogeneity of tumor proteomics.