Original Strange Point Cake Lung cancer is the second most common cancer worldwide and the leading cause of cancer-related death [1]. Although radiological methods such as low-dose computed tomography (LDCT) can reduce the risk of lung cancer-related death by 20%, there are many

Original Strange Point Cake

Lung cancer is the second most common cancer worldwide and the leading cause of cancer-related death [1].

Although radiological methods such as low-dose computed tomography (LDCT) can reduce the risk of lung cancer-related death by 20%, there are many factors that limit its use [2]. Therefore, developing a reliable non-invasive method to detect early-stage lung cancer accurately and cost-effectively is an urgent problem to be solved.

In recent years, liquid biopsy based on cell-free DNA (cfDNA) has shown advantages in early tumor screening, but the prediction sensitivity of a single feature of cfDNA is low. Using a stacked integration method to integrate cfDNA genomic features from whole-genome sequencing (WGS) and create a highly sensitive model has achieved initial results in the detection of early colorectal adenocarcinoma [3]. Whether this method is suitable for early screening of lung cancer is currently little known.

Recently, a research team led by Xu Lin and Yin Rong from Jiangsu Provincial Cancer Hospital (Nanjing Medical University Affiliated Cancer Hospital) developed an accurate and economical early lung cancer detection method by integrating the omics characteristics of cfDNA fragments. The research results were published in the top respiratory journal "American Journal of Respiratory and Critical Care Medicine" [4].

The researchers found that a stacked ensemble model integrating five cfDNA features and five machine learning algorithms outperformed all models based on a single feature-algorithm combination. The sensitivity and specificity of the ensemble model in predicting early-stage non-small cell lung cancer (NSCLC) was over 90%.

It is worth mentioning that this model can still maintain high sensitivity and specificity even when the sequencing depth is reduced to 0.5×. Wang Siwei, Meng Fanchen and Li Ming from Jiangsu Cancer Hospital are the co-first authors of the paper.

Screenshot of the paper's homepage

Next, let's take a look at how this research was carried out.

The research team first randomly assigned 354 subjects to the training set and validation set I. The training set included 113 untreated NSCLC patients (adenocarcinoma ADC: 96; squamous cell carcinoma SCC: 17; stage I: 66; large tumor Small 1cm: 15) and 113 non-cancer healthy volunteers; validation set I included 81 NSCLC patients (ADC: 66; SCC: 15; Phase I: 46; tumor size 1cm: 16) and 47 healthy volunteers. The training set and validation set I are used to build the model and conduct internal verification.

Subsequently, they assigned another 188 subjects (70 healthy volunteers, 118 untreated ADC) to validation set II for external validation. In addition, they also designed an independent validation cohort, including 240 people from other retrospective studies, including 120 healthy people and 120 untreated NSCLC patients. Construction and verification of the

model

The researchers collected plasma samples from all subjects, extracted cfDNA, and then constructed a WGS library. They uniformly conducted model construction and evaluation based on the sequencing depth of 5×, and used WGS data of the original sequencing depth (5.28×-27.85×), or WGS data that reduced the sequencing depth to 4×, 3×, 2×, 1×, and 0.5×, to further evaluate the selected model.

They extracted five different fragment features from WGS data for feature selection and model building. The five fragment characteristics include: copy number variation (CNV), fragment size coverage (FSC), fragment size distribution (FSD), end sequence (EDM) and breakpoint sequence (BPM) .

They then used each cfDNA fragment group feature to build their base model and used five base algorithms: generalized linear model (GLM), gradient boosting machine (GBM), random forest, deep learning, and XGBoost.

Schematic diagram of building a stacked ensemble model and determining cancer probability score

The researchers tested the area under the curve (AUC) of the above five fragment features in five basic models to evaluate the predictive performance of the model. The results showed that the AUC values ​​of EDM, BPM, FSC, FSD and CNV features in the stacked ensemble model were higher than in the single algorithm model. Therefore, they established a stacked ensemble model that integrated the omics characteristics of plasma cfDNA fragments and five machine learning algorithms, with an AUC value of 0.985.

For each cancer or non-cancer sample in this study, a cancer probability score will be generated by the algorithm, ranging from 0 to 1. The higher the score output by the model, the higher the probability of developing cancer. The researchers found that cancer patients had significantly higher cancer probability scores than healthy subjects, and that the distribution of scores increased from stage I to stage IV cancer patients.

To assess the performance of the stacked ensemble model, the researchers used validation set I to determine a cutoff for 95% specificity (there were 46 healthy individuals in validation set I, so the calculated specificity was 44/46 = 95.7%, with a corresponding cancer score cutoff of 0.66) and then applied the cutoff to validation set II and an independent validation cohort for external evaluation.

They found that in the validation set I and validation set II, the AUC values ​​were relatively high, 0.984 and 0.987 respectively. Based on the 95.7% specificity in validation set I and applying 0.66 as the cancer score cutoff, the specificity of validation set II was 98.6%. The resulting sensitivities of validation set I and validation set II were 91.4% and 84.7% respectively.

Development and evaluation of the prediction model in the validation cohort

To further evaluate the generalizability of the stacked ensemble model, the researchers conducted independent experiments It was tested in the validation cohort, and it was found that the prediction model had an AUC value of 0.974 in the independent validation cohort. Using 0.66 as the cancer score cutoff, the prediction model could well distinguish cancer and non-cancer samples, with sensitivity and specificity of 92.5% and 94.2% respectively.. Moreover, in the independent validation cohort, the cancer scores of all patients also showed an upward trend from stage I to stage IV.

They also evaluated the stability and robustness of the model under different WGS sequencing depths and found that the model remained stable when using raw or 5× sequencing depth WGS data, Even after the sequencing depth of is reduced to 4×, 3×, 2×, 1× and 0.5×, their AUC values are still high in validation set I (≥0.966) and validation set II (≥0.971), suggesting good robustness . Moreover, even with the lowest variant allele frequency (VAF) (0.05%) and sequencing depth (0.5×), the model still had a sensitivity of 75.0% in identifying cancer.

Finally, they used the validation set to further evaluate the prediction performance of the model in different lung cancer subgroups. The results showed that the model can reliably distinguish SCC and ADC with sensitivities of 93.3% and 87.0% respectively, and can be used to detect early pathological characteristics such as stage I (sensitivity 83.2%) or tumors <1cm>

Diagnostic sensitivity of the prediction model in different subgroups of lung cancer patients and their combinations in validation sets I and II

In summary, this study established a stacked ensemble machine learning model integrating the omics features of five cfDNA fragments, which can distinguish early-stage NSCLC from non-cancer subjects with high sensitivity, high stability and robustness, and is helpful for the early detection of NSCLC.