
by Christos Evangelou, MSc, PhD – Medical Writer and Editor
Although HER2-targeted treatments have been approved for the treatment of previously treated, advanced biliary tract cancer (BTC), determining eligibility requires reliable quantification of HER2 expression. However, interpretation of HER2 staining can be challenging, especially in tumors expressing low levels of HER2. In a recent study, researchers at the CHA University School of Medicine in the Republic of Korea found that AI could be used to help clinicians identify patients with BTC who might benefit from HER2-directed therapies.
The study was published in Laboratory Investigation.
Study Rationale
BTC is a rare but aggressive malignancy often diagnosed too late for curative intervention. Hong Jae Chon, MD, PhD, one of the corresponding authors of the study, noted that recent clinical trials have shown meaningful benefits from drugs like trastuzumab, zanidatamab, and trastuzumab deruxtecan in patients with HER2-positive BTC.
“HER2-targeted therapies are emerging as promising treatment options in BTC,” Chon said. “However, in real-world practice, we often encounter substantial variability in how HER2 is interpreted by different pathologists. This challenge becomes even more important as the therapeutic focus expands to include HER2-low patients, who may also derive clinical benefit.”
A Multi-Modal Assessment Strategy
To assess variability in HER2 interpretation, the researchers examined 309 immunohistochemistry slides from patients with advanced BTC treated at CHA Bundang Medical Center between 2019 and 2022. They evaluated inter-observer variability (between different pathologists), intra-observer variability (within the same pathologist’s assessments), and between human and AI interpretation.
“BTC is highly heterogeneous, making HER2 interpretation difficult, even within a single slide,” Chon explained. “To understand the variability more comprehensively, we assessed each case using three modalities: light microscopy, digital pathology, and AI-powered analysis.”
Three experienced, board-certified pathologists first evaluated physical slides using conventional light microscopy according to guidelines established for gastroesophageal adenocarcinoma. After a washout period exceeding four weeks to minimize recall bias, they re-evaluated the same specimens as digitized whole slide images. The researchers used Lunit SCOPE HER2, an AI system for HER2 biomarker analysis, to assess the digital images independently. The system uses algorithms trained to detect tumor cells and segment invasive cancer areas before scoring HER2 expression.
Ground truth for each case was established through majority consensus among the pathologists, with the final determination made collaboratively when light microscopy and digital pathology assessments diverged.
Quantifying Diagnostic Disagreement
Pathologists achieved complete three-way agreement in 62.1% of light microscopy evaluations and 63.4% of digital pathology assessments—meaning more than one-third of cases prompted at least some disagreement.
Individual pathologists showed excellent internal consistency, with weighted kappa values ranging from 0.979 to 0.984 when comparing their own light microscopy and digital pathology interpretations. However, inter-observer variability was substantial, with kappa values ranging from 0.819 to 0.876.
HER2 3+ and 0 cases showed less variability than intermediate cases. Pathologists were 87% less likely to disagree on HER2 3+ cases compared to HER2 1+ cases in light microscopy, and 60% less likely to disagree on HER2 0 cases. Surgical specimens generated more consistent assessments than biopsy samples.
The AI system demonstrated 83.5% overall concordance with the consensus ground truth, with slightly better performance in digital pathology (weighted kappa 0.878) than in light microscopy (0.867). In 108 cases in which pathologists did not reach unanimous agreement using either modality, the AI matched the consensus determination 72.2% of the time.
“We were struck by the high level of concordance between AI scoring and pathologist consensus, including in many borderline cases,” said Gwangil Kim, MD, one of the corresponding authors of the study. “AI showed notable strength in providing consistent scores where human interpretation tends to vary.”
The authors also reported one instance in which AI outperformed expert pathologists, identifying a small but definite HER2 3+ tumor cluster of approximately 20 cells that pathologists had missed at standard magnification.
AI Limitations
In 25 cases (8.1%), the AI system completely disagreed with pathologist evaluations. Review of these cases revealed two primary failure modes.
Tissue segmentation errors occurred in 10 cases, where the AI either incorrectly classified non-cancerous areas as invasive cancer (false positives in seven cases) or failed to identify actual cancer regions (false negative in one case). These errors predominantly led to over-grading. The remaining 15 discordant cases involved cell model errors, primarily misclassification of membrane staining intensity, which led to under-grading.
Kim explained that the clinical factors influencing AI performance differed from those affecting the performance of pathologists. Although pathologists struggled more with low HER2 expression levels and biopsy specimens, the AI system showed difficulty with poorly differentiated tumors, where the odds of discordance increased 2.5-fold compared to moderately differentiated cancers.
“Discrepancies were observed in samples with marked inflammation or fibrosis, situations in which human experts rely on nuanced histologic context that AI cannot yet fully incorporate,” Kim said. “This reinforces the idea that AI should serve as an adjunctive tool that complements, rather than replaces, expert pathologist judgment.”
Potential Implications and Future Directions
“Minimizing variability in HER2 interpretation can directly enhance our ability to identify patients who may benefit from HER2-targeted therapies, particularly those with HER2-low expression,” Chon said. “As such therapies continue to expand in BTC, AI-supported assessment can meaningfully improve standardization and reproducibility, ultimately strengthening the accuracy of patient selection and clinical decision-making.”
Study limitations include the single-institution design and the participation of only three pathologists, which limit the generalizability of findings. In addition, the study did not examine whether AI assistance would change pathologist behavior in real-world practice, nor did it correlate HER2 assessments with treatment outcomes or compare immunohistochemistry findings with HER2 gene amplification status by fluorescence in situ hybridization.
“Our next step is to validate these AI algorithms in larger, multi-institutional cohorts and to examine how AI-based HER2 scoring correlates with treatment response and other clinical outcomes,” Kim noted. “We are also interested in studying how AI can be integrated into real-world pathology workflows to enhance efficiency and support pathologists in high-volume clinical environments.”
Chon and Kim concluded by emphasizing that AI should be regarded as a partner to enhance accuracy and consistency in settings with high diagnostic complexity.
“We hope that our work contributes to establishing a more reliable and standardized approach to HER2 assessment for patients with BTC,” they said.
The study received financial support from the National Research Foundation of Korea (NRF) grant 296 funded by the Korean government (MSIT).
References
- Kim H, Heo J, Cho SI, et al. Pathologist-Artificial Intelligence (AI) Concordance in HER2 Interpretation for Advanced Biliary Tract Cancer: Intra-observer, Inter-observer, and Human-AI Variability. Lab Invest. Published online November 11, 2025. doi:10.1016/j.labinv.2025.104259
No audio available for this article yet.
No quiz available for this article yet.







