
Frederico Gaia, MD, a research fellow specializing in digital and computational pathology at Nagasaki University, Japan, challenged conventional thinking about diagnostic disagreement in pathology. Rather than viewing it as a problem to eliminate, he argued that diagnostic diversity reflects genuine scientific uncertainty and that AI can help illuminate this complexity rather than suppress it.
Interobserver variability can be particularly pronounced in complex cases. Although pathologists typically agree on straightforward benign conditions, agreement drops for rare diseases and challenging diagnoses.
“If you ask the same diagnostic question to two pathologists, they will probably say the same words,” Gaia explained during his presentation. “But sometimes they are going to generate different reports, especially for complex and rare conditions. This is diagnostic diversity.”
Gaia discussed their previous research on interobserver variability in lung adenocarcinoma subtyping. Eighteen expert pulmonary pathologists independently evaluated over 5,000 images and labeled them differently, with some identifying lepidic patterns while others noted micropapillary or papillary features in the same areas.
Despite these diagnostic differences in subtype, all expert interpretations maintained clinical relevance, Gaia noted, pointing out that survival curves remained significant across different expert classifications. The explanation lies in how pathologists apply diagnostic criteria, according to Gaia.
“We have different criteria for a subtype; we have some essential and some desirable criteria,” Gaia said.
Through hierarchical clustering analysis, Gaia’s team identified two main diagnostic subgroups among the experts, comprising ten and five pathologists, respectively, with three who did not fit either cluster due to individual diagnostic features. Using each subgroup’s consensus as ground truth, they trained two separate AI models, both demonstrating similarly strong correlations with patient survival.
Gaia’s team used a collaborative platform called NEDO that allows pathologists to view digital slides, make annotations, and see where both AI models identified different histologic patterns. Five pathologists (one beginner, two junior, and two senior) evaluated 99 cases before and after AI assistance. Agreement between pathologists and AI improved significantly, particularly for the beginner (Cohen’s kappa increased from 0.547 to 0.923, p < .0001) and junior pathologists (p = 0.0265). Inter-rater agreement among junior pathologists also increased (kappa from 0.593 to 0.810, p = 0.0463). Overall mean Cohen’s kappa increased from 0.316 to 0.489 after AI assistance.
Gaia emphasized that providing outputs from two distinct AI models encouraged active diagnostic reflection.
“This AI-augmented approach can promote the self-reflection of pathology,” Gaia said. “Pathologists are going to rethink what they are doing and reinforce their knowledge.”
Implementation of AI-augmented interpretation also led to a better separation of survival curves between grade 2 and grade 3 tumors compared to pathologist assessment alone, and this improvement held across all experience levels.
Gaia emphasized that this collaborative model creates a virtuous cycle:
“Pathologists can input data to our diagnostic platform, then make adjustments based on the AI-enhanced interpretation, and check the pathology significance. Then we can use this to improve our AI model.”
Gaia concluded by emphasizing that defining the ground truth is challenging, and that combining pathologist expertise, AI interpretation, and clinical outcomes together may represent the best approach to defining diagnostic standards.

Frederico Gaia, MD, Research Fellow specializing in digital and computational pathology at Nagasaki University, Japan.
No audio available for this article yet.
No quiz available for this article yet.








