
by Christos Evangelou, MSc, PhD – Medical Writer and Editor
In a comparative study of 100 oral histopathology slides, three multimodal large language models (ChatGPT, Grok, and MANUS) achieved diagnostic accuracies between 94% and 97%, compared with 98% accuracy achieved by two board-certified oral pathologists. Grok was the most accurate, ChatGPT was the most reproducible across runs, and MANUS showed the closest alignment with expert diagnostic reasoning.
The study was published in Scientific Reports.
Study Aim
The research team examined whether multimodal large language models can interpret oral histopathology images with enough accuracy and consistency to be clinically useful. The authors used 100 high-resolution slides representing a mix of oral lesions, including epithelial dysplasia, odontogenic tumors, salivary gland neoplasms, cysts, and reactive lesions. The images were obtained from the textbook Oral and Maxillofacial Pathology and independently confirmed by two oral pathology consultants with at least eight years of diagnostic experience.
Each slide was provided as input to ChatGPT (GPT-4-turbo), Grok, and MANUS using the same prompt. To test repeatability, the researchers resubmitted the same 100 slides to each model two weeks later. The main outcomes were diagnostic accuracy, intra-model consistency over time, agreement between the AI systems, and agreement with the expert pathologists.
“Artificial intelligence is rapidly entering healthcare, yet its ability to interpret complex histopathology images remains relatively underexplored,” said corresponding author Abdullah F. Alshammari, PhD. In oral pathology, he said, “diagnosis relies heavily on microscopic interpretation and expert judgment,” making the field a useful proving ground for testing whether modern multimodal systems can handle specialist visual reasoning.
Each Model Shows Different Strengths
All three AI models performed well, and the accuracy of all improved slightly on the second round of testing. Diagnostic accuracy in the first round was 95% with Grok, 94% with MANUS, and 93% with ChatGPT. In the second round, the accuracy of Grok increased to 97%, MANUS to 96%, and ChatGPT to 94%. The human experts achieved 98% accuracy.
“Our findings show that modern AI models can achieve remarkably high diagnostic accuracy when interpreting histopathology images, approaching expert-level performance,” Alshammari said in an interview with Pathology News. “However, human specialists still achieved the highest overall accuracy, reinforcing that AI should be viewed as a supportive tool rather than a replacement for clinicians.”
Although Grok provided the highest diagnostic accuracy, ChatGPT was the most consistent when the slides were reviewed a second time. The intra-model agreement of ChatGPT across the two rounds was near-perfect, with a Cohen’s kappa of 0.918. Intra-model agreement was lower with MANUS (Cohen’s kappa = 0.790) and Grok (Cohen’s kappa = 0.740). McNemar testing showed no statistically significant shifts between rounds for any of the models.
Where AI Agreed With Pathologists, and Where It Did Not
The authors also looked at the agreement between the large language models and the two oral pathologists. MANUS had the highest concordance with the experts, matching them in 94 cases and showing a kappa of 0.485 (p = 0.125). ChatGPT aligned on 93 cases, with a kappa of 0.427 (p = 0.063), and Grok matched 94 cases, with a kappa of 0.265 (p = 0.375).
Alshammari explained that this distinction is clinically relevant, as a model can arrive at the right answer often, but if it does so in ways that diverge from how specialists classify cases, trust and adoption may be limited.
The study showed that most misclassifications clustered in histologically ambiguous cases. Alshammari noted that many oral lesions have overlapping microscopic features, and pathologists do not work from slides alone in clinical practice. Age, lesion site, and clinical history are also taken into account when interpreting histology findings. In this study, the models worked from standardized slide images with limited contextual information, and the authors acknowledge that the lack of patient metadata is a limitation.
Why This Matters to Pathologists
Alshammari said these results suggest that AI could play an important role as a diagnostic support tool in oral pathology.
“Large language models may assist clinicians by providing second opinions, supporting triage of complex cases, and improving diagnostic consistency.” He added that “the goal is not to replace clinicians, but to augment their expertise and improve diagnostic accuracy, efficiency, and access to care.”
Alshammari also pointed to a global access problem that digital tools could help address:
“Such systems could be particularly valuable in regions where access to specialized pathologists is limited.”
He further explained that in many settings, oral pathology capacity is limited, turnaround times are long, and a second expert review may be difficult to obtain quickly. A well-validated multimodal model could help flag cases that need closer review, provide quality assurance, and support teaching in residency or dental training programs.
The authors also argue that comparing several models side by side adds value.
“One of the novel aspects of this study was the direct comparison of multiple multimodal large language models in interpreting oral histopathology images, while simultaneously benchmarking their performance against experienced oral pathologists,” Alshammari said.
Limitations and Future Work
The dataset included 100 textbook-derived images rather than real-world, multi-institutional clinical images. The cases were selected to represent a range of oral pathologies, but the sample may not capture the full range of oral histologies seen in routine practice. Poor staining, artifacts, partial sampling, rare entities, and cases where clinicopathologic correlation changes the likely diagnosis are also common in practice.
The authors also noted that the models were not tested with integrated clinical metadata, and their internal training data are proprietary. Interpretability was not examined in a formal way either, despite the fact that explainability is likely to matter for regulation and clinician trust.
“Further research is needed using larger and more diverse clinical datasets, including more diagnostically challenging cases,” Alshammari said. He also noted that future work should combine patient data with slide interpretation and “investigate explainable AI methods to better understand how these systems reach their conclusions.”
Alshammari concluded by emphasizing that “artificial intelligence has significant potential to enhance diagnostic workflows in pathology, but responsible integration is essential.”
References
- Alshammari AF, Madfa AA, Anazi BA. Human versus artificial intelligence in oral pathology diagnosis: a comparative study of ChatGPT, Grok, and MANUS. Sci Rep. Published online February 25, 2026. doi:10.1038/s41598-026-40792-0









