| Sumario: | Simple Summary: This study tested a previously developed machine learning model that uses routine [18F]FDG-PET/CT scans and clinical data to judge whether lung cancer has spread to central chest lymph nodes before treatment. Researchers applied the model, without changes, to two separate patient groups: 87 patients from one hospital and 124 patients who had primary surgery from a public dataset. They compared the model with a common PET/CT rule based on lymph node activity and size, using surgical tissue results as the standard. The rate of advanced lymph node disease differed between groups. In the hospital group, accuracy for ruling out disease was similar between methods. In the surgical group, the model produced fewer false positives than the common rule for PET/CT image interpretation. Ability to detect disease was comparable in both groups. Overall, the machine learning model's performance was confirmed, with potential advantages in reducing false positives. In non-small cell lung cancer (NSCLC), [18F]FDG-PET/CT is limited in pretherapeutic lymph node (LN) staging by false-positives. We previously demonstrated that a machine learning (ML) classifier using routine [18F]FDG-PET/CT and clinical variables can improve diagnostic accuracy compared to visual assessment. The present study aimed at independent validation. Cohort 1 (Charité) included 87 NSCLC patients (surgical and non-surgical), prospectively enrolled at our institution. Cohort 2 (TCIA) comprised 124 patients with primary surgery from the multi-institution NSCLC Radiogenomics dataset. Our ML classifier for differentiating N0/1 vs. N2/3 status was applied without modification. As comparator, the combined standard PET/CT criterion of "mediastinal LN uptake > mediastinum and/or short-axis > 10 mm" was used. Histology of N2/3 LNs served as reference standard. Prevalence of pN2/3 differed significantly between cohorts (Charité: 40%, TCIA: 12%; p < 0.001). Specificity was similar between ML and the standard PET/CT criterion in the Charité cohort (65% vs. 60%; p = 0.5) but significantly higher with ML in TCIA (90% vs. 70%; p < 0.001). Sensitivity for pN2/3 was comparable between the two comparators in both the Charité cohort (97% each; p = 1.0) and TCIA (27% vs. 33%; p = 1.0). Lower sensitivity in TCIA patients reflects the preselection of surgical patients who had already been clinically staged and deemed suitable for surgery. The diagnostic performance of the ML classifier and its (potentially) superior specificity were thus successfully validated in two independent cohorts.
|