
Nadieh Khalili, PhD, and colleagues at Radboud University Medical Center developed an Uncertainty-Guided Annotation (UGA) framework, which allows AI systems to effectively communicate their uncertainties to pathologists, enabling targeted improvements in model performance.[1] The work represents a significant step forward in making AI systems more transparent and applicable to real-world clinical settings.
“Even the most modern foundational models nowadays are making mistakes,” Khalili noted during her presentation. “For example, ChatGPT, which we are using every day, always answers our questions and barely says ‘I don’t know.’ That is more tricky in the medical domain, as it destroys the trust of pathologists and radiologists.”
A particular challenge in pathology AI has been the variation in staining protocols between different laboratories. These differences can influence the performance of AI models when used in different institutions. The UGA framework addresses this issue by identifying areas in which the model is uncertain about its predictions.
In an interview following her presentation, Khalili emphasized the importance of finding the right balance between automation and clinical oversight. “Clinicians emphasized the need for a balance between automation and control; they want AI to assist, not dictate, their workflow,” she said. This insight led to the development of an interactive interface that allows pathologists to efficiently refine the model outputs while maintaining control over the diagnostic process.
The research team evaluated their framework using the CAMELYON dataset, which includes WSIs from five medical centers. The model’s Dice coefficient, a measure of segmentation accuracy, improved from 0.66 to 0.76 after incorporating five uncertain patches, and further increased to 0.84 after manual corrections.
As Khalili explained, the ability to quantify uncertainty at the pixel level makes the UGA framework innovative. The system uses an ensemble of neural networks trained with different random initializations to identify the areas where the model is most likely to make mistakes. This information is then presented to pathologists as a heat map, highlighting the regions that require human attention.
“We quantified uncertainty per pixel based on the model disagreement and then aggregated the uncertainty by 1024 x 1024 patches,” Khalili said during her talk. This approach is particularly useful when dealing with slides from centers with different staining characteristics, as the framework helps mitigate domain shifts caused by staining differences by actively identifying regions in which the deep learning model has high uncertainty.
“Instead of requiring exhaustive manual re-annotation, UGA strategically suggests areas where corrections would have the greatest impact,” Khalili noted. “This ensures that models generalize better across different centers, reducing the burden on clinicians while improving reliability in clinical workflows.”
The UGA framework was designed with clinical workflows in mind. “The framework is designed to be intuitive for clinicians, requiring minimal training,” Khalili said in the interview. “Users can review the errors AI makes, adjust, and approve segmentations without extensive technical knowledge.”
The team is currently working on improving the computational efficiency of the system. “The ensemble network is computationally very expensive,” Khalili acknowledged. “Therefore, we are exploring Bayesian neural networks with Kullback–Leibler divergence for uncertainty estimation.”
The research team is also expanding their work beyond the CAMELYON dataset. Khalili announced an upcoming challenge at MICCAI 2025 that will combine pathology with other modalities, including MRI and genomic data, for tasks such as recurrence prediction and survival analysis in bladder cancer.
To promote collaboration and further development, the research team made their code publicly available on GitHub. “Open-sourcing the UGA framework encourages transparency and community-driven improvements,” Khalili stated. “We hope this will lead to further research on active learning strategies in pathology.”Nadieh Khalili, PhD, workgroup leader at Radboud University Medical Center, Netherlands.
[1] Nadieh Khalili, A human-in-the-loop framework for refining deep learning models in pathology segmentation. Presented at SPIE 2025 Digital and Computational Pathology conference, February 19, 2025; San Diego, CA.
No audio available for this article yet.
No quiz available for this article yet.









