
by Christos Evangelou, MSc, PhD – Medical Writer and Editor
The combination of digital pathology with artificial intelligence (AI) shows promise in improving cancer diagnosis by enhancing analysis of histopathology images. However, recent studies revealed concerning biases in these systems that could impact their reliability in clinical settings.
In a new study, researchers from Ontario Tech University, Brock University, and Wilfrid Laurier University in Canada identified factors that contribute to the ability of deep learning models designed to learn cancerous patterns to classify data centers. According to the authors, these findings suggest that deep learning models may be learning institution-specific patterns rather than true cancer features.
The report was published in Scientific Reports.
Uncovering the Sources of Center-Specific Bias in AI Cancer Detection
The research team investigated why deep neural networks trained to detect cancer types could also predict which institution acquired the histopathology image with nearly 70% accuracy. This unexpected capability raised serious questions about what these models are actually learning.
“If a trained model relies on center-specific artifacts like staining procedures, it may perform well only on data from familiar centers but fail in unseen clinical settings,” noted Farnaz Kheiri, PhD researcher at Ontario Tech University and the first author of the study. “High reported accuracy might not result from true learning of cancer-related patterns, but rather from shortcut learning based on cues unique to specific hospitals.”
In histopathology, these shortcuts can take various forms, from staining variations to image acquisition differences between medical centers, Kheiri explained.
The research team conducted four case studies using two deep learning models: KimiaNet (a pre-trained model) and EfficientNet (which they trained themselves) to determine the sources of bias at different stages of the AI development pipeline.
Imbalances in Data Distribution
In the first case study, the researchers used mutual information analysis to quantify the dependency between cancer type and data centers. The results showed a significant correlation, with a mutual information value of 1.47 for the original dataset compared to zero for an artificially uncorrelated dataset. These results suggest that certain cancer types are predominantly sourced from specific centers.
“When a cancer type is disproportionately represented by data from a single institution, models may inadvertently learn center-specific artifacts rather than true biological patterns,” Kheiri explained. “This is especially problematic for cancers with subtle or non-distinctive histopathological features, which provide weaker biological signals, increasing the likelihood that models rely on superficial, non-biological cues.”
The Impact of Patching
The second case study examined how the patching process affects model performance. When patches from the same slide were excluded from the analysis, accuracy dropped from 96% to 48% for center identification and from 99% to 66% for cancer classification, suggesting that models were recognizing slide-specific features rather than cancer characteristics.
“Slide-specific bias can involve patient-specific signatures unique to individual slides, such as tissue characteristics or staining patterns,” noted Kheiri. “This can harm generalization and reduce clinical reliability. Addressing this requires strategies like slide-level cross-validation and regularization to promote more robust, generalizable learning.”
Ensuring cancer types are evenly distributed across data centers can reduce correlation-based biases, according to Kheiri. She noted, however, that this alone does not guarantee that the model will not become biased toward specific data centers.
“I recommend intervening in the training process to ensure that the model focuses on cancer-specific patterns, rather than center-specific features,” she added.
Site-Specific Patterns in Cancer Features
The third case study demonstrated that patches from the same cancer type and data center share more similarities than those from different centers, indicating the presence of site-specific patterns in supposedly cancer-specific features.
“This center-specific bias could result in suboptimal performance when deployed in diverse clinical environments, leading to misdiagnosis, or even health inequities for patients from underrepresented centers,” Kheiri emphasized.
The researchers found that certain cancer types showed higher contribution rates to this bias. Renal clear cell carcinoma, uterine carcinosarcoma, and sarcoma had the highest rates, while mesothelioma, pheochromocytoma, and cervical squamous cell carcinoma had the lowest.
Kheiri provided an explanation and a potential solution for this variation:
“In some cases, institution-specific characteristics may correlate with biological features, further complicating model interpretation and leading to confounding effects. To address data center-specific bias, researchers should ensure balanced and stratified sampling across institutions, perform cross-center validation, and transparently report sample distributions per cancer type and center.”
Staining Variations and Color Bias
In the fourth case study, the researchers assessed the impact of color variations by converting images to grayscale and introducing controlled noise. Noise-based grayscale normalization (NBGN) reduced the model’s ability to identify data centers while maintaining reasonable cancer detection accuracy, suggesting that staining patterns contribute to bias.
The team compared their NBGN method with established techniques like Reinhard normalization and Optical Density transformation. Although all methods reduced center-specific bias to some degree, NBGN and Reinhard normalization showed the best balance between bias reduction and maintaining cancer detection accuracy.
“To reduce color variation bias, pathology labs should implement stain normalization techniques like Macenko or Reinhard methods,” Kheiri suggested. “Standardizing staining protocols across data centers will also help ensure consistency across sites. Regular calibration of equipment and collaboration across centers can further minimize staining discrepancies.”
The researchers found that while converting images to grayscale and adding controlled noise reduced the model’s ability to identify data centers (from 75% to 65% accuracy), it maintained strong cancer detection performance (dropping only from 84% to 82%). This suggests that color normalization can effectively reduce bias without significantly compromising diagnostic capability.
Future Directions
According to Kheiri, machine unlearning could help reduce site-specific biases.
“One method of implementing machine unlearning in histopathology AI systems is to retrain the model using an unlearning algorithm that adjusts the model’s weights, specifically focusing on retaining task-relevant features, such as cancer-specific patterns, while forgetting irrelevant patterns related to data centers, like staining or imaging artifacts,” Kheiri said.
The researchers also suggest that innovative strategies are needed to detect and correct biases during model training without compromising the ability to detect clinically relevant features. For pre-trained models like KimiaNet, where retraining is not feasible, post-training methods like feature selection may effectively eliminate bias signatures from learned features.
As Kheiri concluded, addressing these biases is essential not just for technical accuracy but for ensuring equitable healthcare.
“Techniques such as domain-adversarial training and data harmonization can further help reduce reliance on non-biological, institution-specific artifacts and promote more generalizable models that serve all patients equally well, regardless of where they receive care,” she said.
References
- Kheiri F, Rahnamayan S, Makrehchi M, Asilian Bidgoli A. Investigation on potential bias factors in histopathology datasets. Sci Rep. 2025;15(1):11349. Published 2025 Apr 2. doi:10.1038/s41598-025-89210-x
No audio available for this article yet.
No quiz available for this article yet.







