
Several studies have tried to address the ‘black box’ question in AI tools for pathology: are these models getting the right answers for the right reasons? A new study presented at Pathology Visions 2025 provides insights into how LLMs approach diagnostic reasoning, and the findings reveal both promise and significant limitations.
Ghulam Rasool, PhD, and Ehsan Ullah, PhD, from Moffitt Cancer Center and Health New Zealand, presented findings from a study evaluating four LLMs with reasoning capabilities on pathology board-style questions. The responses of LLMs were assessed by 11 expert pathologists worldwide. The researchers went beyond accuracy metrics to examine whether these models employ the reasoning strategies that experienced pathologists use daily.
“When a pathologist reviews a case, they are not just analyzing images through pattern recognition; they go through an intricate diagnostic decision making process by adding up visual cues from the slides and then weighing up the probabilities in their mind based on patient demographics, clinical presentation, ancillary studies and relevant guidelines,” Rasool explained in an interview with Pathology News.
Gemini outperformed DeepSeek and OpenAI’s ChatGPT-o1 and o3-mini. Gemini achieved the highest scores across all metrics, including relevance (571 out of 600), coherence (548), analytical depth (542), accuracy (546), and conciseness (448).
“What stood out in our study was that Gemini and DeepSeek could demonstrate a structured, step-by-step reasoning process, ‘their chain of thought’ in reaching from visual and other clinical information available in board-style pathology questions to decipher a diagnosis,” Rasool stated in the interview. “These models laid out logical decision trees and explained how they moved from multiple sources of data to diagnosis — that’s what we refer to as beyond the black box.”
However, the study also showed that all models struggled with heuristic reasoning and probabilistic thinking, the kind of fast, experience-based judgments that define expert practice.
“Heuristic reasoning and pattern recognition are the ‘gut instincts’ of pathology, the fast, experience-based judgments that come from seeing thousands of slides,” Ullah explained in an interview with Pathology News. “Those are incredibly hard for language models to replicate, because they depend on visual intuition and contextual memory, not just text-based logic.”
The speakers provided more insights into the evaluation process. Each pathologist spent roughly 24 hours reviewing 720 responses. In their review, they assessed both language quality (accuracy, relevance, coherence, depth, and conciseness) and seven reasoning strategies that are commonly used in pathology practice. Inter-rater agreement hovered around 50%, with the highest concordance observed for Gemini’s outputs.
The study also demonstrated that models that reasoned most thoroughly were the least concise.
As Rasool explained, “Our results showed that the models that reasoned most like experts were also the least concise, that is, they took the time to explain why a diagnosis made sense.” He suggested context-dependent use of LLMs as a solution: “For routine or high-volume work, concise summaries are best, but for complex or ambiguous cases, you want the AI to show its reasoning trail.”
The practical applications of LLMs in pathology were also discussed in a different session by Harshwardhan Thaker, MD, PhD, from the University of Texas Medical Branch, who presented work on custom LLMs for integrated diagnostic reporting. His team developed a customized GPT that analyzes patient data from multiple sources, including clinical notes, lab values, radiology, and pathology reports, to generate comprehensive reports with risk stratification and management recommendations grounded in NCCN guidelines. The system even produces patient-friendly versions tailored to different reading levels.
All the presenters emphasized that rigorous validation and human oversight of LLMs remain essential.
“Two things need to happen before clinical use: first, prospective, multi-site validation on representative case mixes, and not just ‘right/wrong’ grading, but whether the model’s reasoning is clear, auditable, and consistent across raters,” Rasool stated in the interview. “Second, deployment should keep a human firmly in the loop with guardrails: show the model’s reasoning steps, surface uncertainty, log versions, and enable easy override with feedback for monitoring.”
The message is clear: while newer LLMs can approximate aspects of expert pathology reasoning, they work best as augmentation tools rather than replacements. As Rasool and Ullah put it, these systems should serve as “partners to pathologists to augment their efficiency,” leaving final diagnostic judgment with human experts.

Ghulam Rasool, PhD, assistant member of the Department of Machine Learning at the H. Lee Moffitt Cancer Center & Research Institute, Tampa, FL, USA.
No audio available for this article yet.
No quiz available for this article yet.









