
Novel AI Framework Enhances Image-Text Retrieval in Digital Pathology
by Christos Evangelou, MSc, PhD – Medical Writer and Editor
In a recent study, researchers at the University of Waterloo (Ontario, Canada), Brock University (Ontario, Canada), and Mayo Clinic (Minnesota, USA) developed a novel artificial intelligence (AI) framework that may change the way medical professionals analyze and retrieve histopathology data.
The proposed approach, which combines self-supervised learning with a tailored patching strategy for whole slide images (WSIs), significantly enhances cross-modal search capabilities in histopathology archives.1 According to the authors, implementation of this method could lead to more accurate diagnoses, improved research outcomes, and, ultimately, better patient care.
“Cross-modal image-text analysis integrates visual and textual data to improve diagnostic accuracy, treatment planning, and clinical decision support. It enhances research, education, automated reporting, and patient monitoring. This approach leverages comprehensive data for more effective healthcare solutions,”
said Hamid R. Tizhoosh, PhD, professor at Mayo Clinic and the corresponding author of the study.
The report was published in Scientific Reports.
The Unmet Need: Bridging the Gap in Multi-Modal Medical Data Analysis
As Dr. Tizhoosh explained, the rapid increase in data volume and variety, along with advancements in deep learning, has led to the integration of multi-modal data in real-world applications.
“These applications require comprehensive data from various sources, but most current models usually handle only one data type, limiting their effectiveness,”
he added.
As the volume and variety of medical data continue to grow exponentially, there is an increasing demand for sophisticated techniques to analyze and interpret multi-modal big data, particularly in fields like computational pathology.
“Multi-modal learning combines data from various types like text, images, and molecular data, aiming to link and understand diverse data for comprehensive insights,”
noted Dr. Tizhoosh.
Cross-modal retrieval, a specialized form of multi-modal learning that aims to find a common latent space where different modalities such as image-text pairs align closely, has emerged as a promising solution. However, the representation of tissue features in visual models has remained a significant challenge, primarily due to the scarcity of labeled data in the medical domain.1,2
“Cross-modal retrieval retrieves entities from one modality using queries from another. This presents challenges in bridging semantic gaps and managing inconsistencies,”
said Dr. Tizhoosh.
A Novel Approach to Cross-Modal Retrieval
To address these challenges, the researchers developed a framework that builds upon the self-supervised learning scheme known as DINO (Distillation without supervision). The LILE (Look In-depth Before Looking Elsewhere) architecture, designed for cross-modality tasks, was used as a backbone.1
“Both multi-modal learning and cross-modal retrieval are important for integrating diverse data, enhancing decision-making in areas like medical diagnosis, recommendations, and multi-modal search engines. The proposed methodology integrates several key components to address cross-modality retrieval challenges in pathology,”
noted Dr. Tizhoosh.
The key innovation in their approach is the introduction of a novel patching technique called ‘harmonizing DINO’ or H-DINO, specifically designed for histopathology WSIs. Unlike traditional methods, H-DINO extracts larger patches from the WSI and then crops views from these larger patches rather than using different magnifications.1 This strategy maintains consistent magnification across views, which is crucial for preserving the contextual integrity and structural consistency of tissue samples.
“The harmonization of scale refines the DINO paradigm through a novel WSI patching approach, overcoming the complexities posed by gigapixel WSIs in digital pathology,”
Dr. Tizhoosh emphasized.
Two types of views are extracted using this approach: 1) global views, cropped to 140-224 pixels; and 2) local views, cropped to 50-140 pixels. These patches undergo different augmentation pipelines, including a ‘hyper’ augmentation pipeline with stronger transformations and a ‘mild’ augmentation pipeline with less intensive transformations.1 The final processed patches are resized to 224 ´ 224 pixels as input to the network. This approach allows the model to capture both global and local perspectives while maintaining the unique characteristics of histopathology images.1
Network Architecture and Training
The researchers integrated the H-DINO approach into a larger framework that includes a vision transformer (ViT) architecture for image encoding, BioBERT for text encoding, and multi-head self-attention modules to highlight key features in both images and text.1 The framework also includes cross-attention modules to extract important segments from images and text, considering their relevance to the other modality, and a gated memory to refine extracted features in an iterative scheme.
The entire model is trained end-to-end using a novel loss function that combines self-supervised learning objectives with cross-modal retrieval objectives. This integrated approach allows the model to learn rich, aligned representations of both image and text data simultaneously.1
Improved Performance Across Multiple Datasets
The researchers evaluated their framework on three datasets: the GRH dataset (a private breast cancer dataset involving 22 primary diagnoses), PatchGastricADC22 (a dataset of pathologist-composed images and descriptions), and LC25000 (a simpler dataset with five primary diagnoses).1
Across all datasets, the proposed LILE + H-DINO method consistently outperformed existing approaches in both patch-based and WSI-based retrieval tasks. On the GRH dataset, LILE + H-DINO achieved superior performance in both patch-based and WSI-based tasks, with significant improvements in R@sum metrics.1 For the PatchGastricADC22 dataset, the method demonstrated robust performance in handling real-world pathologist-written descriptions, outperforming other approaches including zero-shot models like BioMedCLIP. In the LC25000 dataset, LILE + H-DINO also outperformed other methods, particularly in the R@1 metric, despite the relative simplicity of this dataset.1
Commenting on the implications of these findings, Dr. Tizhoosh said:
“Self-supervised learning in cross-modal retrieval is quite effective, and the integration of tailored patching for WSIs enhances performance. The novel ‘harmonizing’ patching preserves contextual integrity and structural consistency, contributing to richer tissue interpretation.”
Future Work
The development of this novel cross-modal search framework represents a significant advancement in digital pathology. By effectively bridging the gap between image and text data in histopathology archives, this approach opens new avenues for more efficient and accurate diagnoses. Although the model shows promise in controlled research settings, extensive validation in real-world clinical environments is necessary to ensure its practical applicability and reliability.
“A promising direction for future research involves incorporating diverse clinical information or additional data modalities into cross- and multi-modality tasks, enhancing the comprehensiveness and applicability of the analysis,”
noted Dr. Tizhoosh.
Furthermore, future work should evaluate retrieved images and texts based on content similarity rather than just category matching to provide a more nuanced understanding of the model’s performance.
No audio available for this article yet.
No quiz available for this article yet.








