
by Christos Evangelou, MSc, PhD – Medical Writer and Editor
Researchers at Rutgers Health and the University of Pittsburgh Medical Center developed a multiresolution deep learning tool to address interobserver variability in risk stratification for potentially malignant oral lesions. The model demonstrated 80% accuracy and outperformed both traditional pathology grading systems and human pathologists in identifying oral lesions likely to progress into malignant lesions.
The study was published in npj Digital Medicine.
Study Rationale
Potentially malignant oral lesions are visible mucosal changes that could progress to squamous cell carcinoma. Yingci Liu-Swetz, DDS, MS, the first author of the study, explained that the current World Health Organization (WHO) grading system, which categorizes oral dysplasia as mild, moderate, or severe, provides limited accuracy in distinguishing which patients will actually develop cancer.
“Potentially malignant oral lesions are common findings in dental and medical clinics, but our current grading system is subjective and often inaccurate in identifying which patients will go on to develop oral cancer,” Liu-Swetz said. “Our goal was to use artificial intelligence to bring more objectivity and precision to this process, specifically for histopathologic evaluation.”
Liu-Swetz noted that 20%-35% of severe dysplasia cases progress to carcinoma without treatment, and 5%-15% of mild and moderate dysplasias also progress to malignant lesions. The WHO grading system is prone to inter- and intra-rater variability (κ: 0.41-0.5), with pathologists frequently disagreeing on how to classify the same lesion. This variability can influence patient care, as severe dysplasia typically prompts complete excision, whereas mild and moderate cases are often managed with watchful waiting.
Methodology
The research team developed a vision transformer (ViT) model trained on 221 digitized whole-slide images from patients with known outcomes; lesions in 111 patients progressed to squamous cell carcinoma, and those in 110 did not. Patients were followed up for a minimum of five years.
The ViT model integrated a multiresolution framework that first scans at low magnification to assess overall architecture, and then zooms in to evaluate cellular details. The model analyzed tissue patches at three magnifications (10×, 20×, and 40×), extracting features at each resolution level before integrating them for lesion classification. Ablation experiments demonstrated that models with three resolution backbones outperformed those with two, and models with two resolution backbones outperformed single-resolution models.
The research team implemented vision transformer technology rather than convolutional neural networks (CNNs), which are commonly used in AI models for digital pathology. Liu-Swetz explained that although CNNs can identify local features with high accuracy using their convolutional filters, they can fall short at capturing long-range spatial relationships, which can reveal tissue architectural patterns common in high-risk premalignant disease. Vision transformers, on the other hand, can identify distant spatial patterns, as they process images as sequences of patches.
Model Performance
When tested on 50 whole-slide images from three institutions, the ViT model outperformed CNN-based architectures (VGG16, InceptionV3, and ResNet50) on most performance metrics, including specificity, which remained relatively low for the CNN models despite hyperparameter tuning.
The ViT model demonstrated 80% accuracy in lesion classification, with an F1-score of 0.773 and an area under the receiver operating characteristic curve of 0.798. In addition, the ViT model demonstrated a positive predictive value of 73.9%, exceeding the maximum of 50% reported for severe dysplasia under the WHO grading system.
The model also outperformed binary classification by three blinded pathologists who categorized cases as high-risk (moderate to severe dysplasia) or low-risk (mild dysplasia). The AI model showed higher precision (73.9% vs. 69.6%) and sensitivity (81.0% vs. 76.2%) than expert pathologists.
“One of the most exciting findings was that the model’s predictions aligned with well-established histopathologic features of dysplasia, such as abnormal keratinization, increased apoptosis, and N:C ratio,” Liu-Swetz stated. “This tells us the AI is making predictions based on known histopathologic features of malignancy, rather than relying on random image features.”
Three pathologists independently scored 24 established histopathologic features of dysplasia for each test case. Eight features appeared significantly more frequently and severely in progressing lesions, and all eight were also identified by the AI as predictors of progression. Four histopathologic features showed particularly strong correlations with progression to cancer: karyorrhectic/apoptotic cells, single-cell keratinization, premature keratinization, and increased nuclear-to-cytoplasmic ratio (all p-values <0.0001).
The model’s predictions also correlated with overall dysplasia severity. AI-classified progressors had higher average dysplasia scores than non-progressors (2.17 vs. 1.24, p<0.001). Decision curve analysis demonstrated potential clinical utility across a broad range of risk thresholds, with the model demonstrating superior net benefit compared to “treat all” or “treat none” strategies.
Interpretation
Liu-Swetz explained that the superior performance of the ViT model compared to CNN-based architectures and expert human assessment likely reflects consistency rather than fundamentally different diagnostic capabilities.
“This does not necessarily mean the AI is diagnostically superior to a human pathologist,” Liu-Swetz said. “It likely reflects the model’s ability to apply the same criteria with better consistency, without the variability that can occur in human interpretation. That consistency may translate into better predictive performance in this specific task.”
Visual analysis of the model’s predictions on whole-slide images showed that regions with pronounced cytologic and architectural atypia, such as nuclear pleomorphism, hyperchromasia, severe dyskeratosis, and proliferative basal layers, were frequently classified as high-risk. Conversely, areas with milder atypia correlated with predictions of non-progression.
False positive cases typically showed projecting rete ridges, dyskeratosis, and inflammatory infiltrates, which could mimic malignant potential. In one case, the atypia appeared reactive rather than neoplastic, associated with an adjacent ulcer. False negative cases often exhibited minimal architectural distortion and low inflammation.
Limitations and Future Directions
The authors acknowledge that the dataset of 221 cases is relatively small, limiting conclusions about the generalizability of the predictions. The study was retrospective and did not incorporate clinical variables such as tobacco and alcohol use, which contribute to oral cancer risk. Additionally, the model relies solely on histopathologic images, excluding other potentially informative data sources, such as clinical characteristics and demographic variables.
“If validated in larger, multi-institutional cohorts, this approach could support clinicians by identifying truly high-risk lesions earlier, guiding decisions around surveillance versus surgical management, and ultimately helping prevent cases of oral cancer before they develop,” Liu-Swetz stated.
The study received financial support from the NIH National Center for Advancing Translational Sciences.
References
- Liu-Swetz Y, Niksic S, Seethala RR, Shasteen A, Foran DJ, Bilodeau EA. AI-driven prediction of progression to oral squamous cell carcinoma using a multiresolution pathology model. NPJ Digit Med. 2025;8(1):657. Published 2025 Nov 13. doi:10.1038/s41746-025-02014-1
No audio available for this article yet.
No quiz available for this article yet.









