
by Christos Evangelou, MSc, PhD – Medical Writer and Editor
In a significant advancement for digital pathology and artificial intelligence (AI), researchers at Fondazione Bruno Kessler, Santa Chiara Hospital, and University and Hospital Trust of Verona in Italy developed a framework for generating and evaluating synthetic medical images that are so realistic that even expert pathologists struggle to distinguish them from real tissue samples.
The research team demonstrated that their approach could generate high-quality synthetic images of five different tissue types, with pathologists achieving only slightly better than chance performance (56% accuracy) in distinguishing real from synthetic images. According to the authors, this pipeline could address current challenges in medical data sharing and privacy protection.
The report was published in Scientific Reports.
The Challenge of Medical Data Access
Digital pathology has transformed how medical professionals analyze tissue samples, replacing traditional glass slides with high-resolution digital images. However, the field faces significant challenges in developing AI-assisted diagnostic tools due to data scarcity and privacy concerns surrounding medical images. Synthetic data generation offers a promising solution, but its implementation requires careful validation to ensure that the generated images maintain accuracy and clinical relevance.
A Three-Pronged Evaluation Approach
Researchers developed a novel pipeline using a state-of-the-art AI technique called denoising diffusion probabilistic models (DDPM).1 They worked with 650 whole slide images (WSIs) from the Genotype-Tissue Expression (GTEx) dataset, focusing on the following five tissue types: brain, kidney, lung, pancreas, and uterus. This diverse dataset allowed them to test their framework across tissues with varying morphological complexities.
“We chose diffusion models over traditional methods like generative adversarial networks because they produce higher-quality and more diverse images, which is critical in medical imaging,” explained Dr. Giuseppe Jurman, one of the study’s lead investigators. “Diffusion models are less prone to issues like mode collapse, which can limit the variety of images generative adversarial networks generate.”
The evaluation framework consisted of three distinct components. First, the team employed established metrics to measure the quality and similarity between real and synthetic images, including Inception Score, Fréchet Inception Distance (FID), improved precision-recall metrics, density-coverage measurements, and IL-NIQE image quality score.1
The second evaluation component involved training deep learning models to classify tissue types using both real and synthetic images. The researchers also utilized explainable AI techniques to understand whether the models learned similar features from both types of images. This practical assessment provided insights into the usability of synthetic images for machine learning applications.
The third evaluation component involved validation from expert pathologists. Three pathologists participated in two questionnaires: one focused on identifying tissue types from sets of three images, and another aimed at distinguishing between real and synthetic images.1 This human expert evaluation served as a benchmark for assessing the biological relevance of the AI-generated images.
Matteo Pozzi, the co-first author of the study, emphasized the importance of their comprehensive evaluation strategy:
“Our evaluation strategy was designed to provide a comprehensive understanding of the quality and usability of the synthetic pathology data by assessing it from multiple angles.”
Pozzi further explained that each metric complemented the others by offering unique insights, which, together, provide a comprehensive understanding of the quality and usability of synthetic data for medical research and clinical applications.
“Quantitative measures provide an objective baseline for image quality, practical usability testing ensures the synthetic data is functionally robust and interpretable, and qualitative evaluation validates the biological credibility of the synthetic images,” he added.
Image Quality and Similarity
The synthetic images achieved high quantitative scores, with an Inception Score of 2.9115 (±0.0339).1 The generated images showed strong similarity to real images across multiple metrics. However, performance and generation quality varied among different tissue types, with brain and pancreas tissues showing high FID scores (indicating high dissimilarity between synthetic and real images) and uterus and lung tissues showing low FID scores.1
Clinical Validation
In the clinical validation phase, pathologists showed nearly identical performance in identifying tissue types from both real and synthetic images. Their ability to distinguish between real and synthetic data was only slightly better than chance, achieving 56% accuracy in distinguishing real from synthetic images.1 When pathologists could identify synthetic images, it was primarily due to visual attributes like blurriness rather than biological inconsistencies.
“One of the most interesting findings from the pathologists’ evaluations was that they relied almost entirely on visual features, like color differences or blurriness, to tell real and synthetic images apart,” noted Dr. Shahryar Noei, co-first author of the study. “They did not notice any significant biological or structural flaws in the synthetic images, which shows how convincing these generated images were in replicating real tissue appearances.”
Testing of the deep learning performance showed that models trained on synthetic images achieved 93% accuracy when classifying real images, compared to 98% accuracy for models trained on real images. The explainability analysis confirmed that models learned similar relevant features from both real and synthetic images.
Implications, Challenges, and Potential Solutions
According to Dr. Noei, integrating synthetic data into clinical practice has the potential to make a real difference, especially in addressing data scarcity and privacy concerns.
“Synthetic data can work alongside real datasets, helping train machine learning models when it’s hard to collect enough annotated medical data, like for rare conditions or underrepresented tissue types,” Dr. Noei said. He added that synthetic data replicate the patterns of real data without containing any actual patient information, which “makes it much easier for institutions to share data and collaborate without the risk of breaching privacy rules.”
However, the team faced several technical hurdles during the development of the evaluation framework.
“In the initial stages, while experimenting with various diffusion models, we encountered significant GPU vRAM constraints that hampered progress,” explained Dr. Jurman. “To address this, we pivoted to a simpler architecture based on U-Net, which yielded impressive results without the need for more computationally demanding models.”
Pozzi added, “Another challenge arose from the lack of diversity in tissue tiles, often caused by staining or processing artifacts. We utilized pixel gradient information to compute a complexity metric, enabling us to filter out low-information tiles. This ensured that the generative model trained on a dataset rich in meaningful morphological features.”
Moreover, the authors acknowledge that models trained solely on synthetic data do not yet perform as well as those trained on real data, which shows that generative algorithms still have room to improve in capturing the full complexity of medical images.
“Even for generating rare datasets, generative models still need a good amount of real data to learn from, which can be tricky to get in the first place,” Dr. Jurman noted.
Looking Ahead
Researchers envision that with close collaboration between AI researchers, pathologists, and medical experts, synthetic data could become a powerful tool for improving research, education, and clinical workflows.
“The next big step for synthetic data in digital pathology is to focus on generating WSIs with detailed annotations and exploring ways to create multimodal datasets,” explained Pozzi. “Generating WSIs would be a game-changer, as it would allow synthetic data to mirror the scale and complexity of the images pathologists use in their work.”
Dr. Noei added, “Another exciting direction is multimodal data generation, where synthetic datasets combine pathology images with other types of data, like genetic information, radiology scans, or clinical records. This could open up new possibilities for training AI models to analyze data in a way that better reflects how clinicians make decisions.”
The study received financial support from the Italian Ministry of University and Research.
References
- Pozzi M, Noei S, Robbi E, et al. Generating and evaluating synthetic data in digital pathology through diffusion models. Sci Rep. 2024;14(1):28435. Published 2024 Nov 18. doi:10.1038/s41598-024-79602-w
No audio available for this article yet.
No quiz available for this article yet.







