
by Christos Evangelou, MSc, PhD – Medical Writer and Editor
In a recent study, researchers in the Republic of Korea evaluated the ability of deep learning-based image analysis (DLIA) to enhance the accuracy of Gleason grading and tumor quantification in radical prostatectomy specimens in collaboration with Deep Bio Inc. (Republic of Korea). The study showed that DLIA not only matched pathologists’ diagnostic capabilities but also surpassed them in certain aspects of tumor assessment.
“In areas with pathologist shortages, diagnostic gaps in prostate cancer management can become significant due to pathologist overload, which can lead to delayed or inaccurate diagnoses,” noted Hong Koo Ha, MD, PhD, Professor at Pusan National University and the lead investigator of the study. “By efficiently providing highly sensitive cancer assessments, AI algorithms like the one used in the research can serve as valuable assistants to help pathologists maintain diagnostic accuracy and throughput while preventing errors in diagnosis and grading.”
The report was published in Scientific Reports.
Study Rationale
Prostate cancer diagnosis and management face dual challenges: the complex pathological characteristics of the disease and a global shortage of specialized pathologists. The traditional examination of radical prostatectomy specimens is time-consuming and subjective, often leading to inter- and intra-observer variability. This is particularly problematic given that prostate cancer is frequently multifocal, with significant heterogeneity in each tumor focus.
Additionally, while biochemical recurrence (BCR) after radical prostatectomy is a crucial marker for determining the need for salvage treatment, the role of tumor volume (TV) and percent tumor volume (PTV) in predicting biochemical progression-free survival (BPFS) remains controversial.
Dr. Ha explained that prostate cancer’s more frequent multifocality compared to other cancers makes manual volume assessment technically more challenging and subjective, which could contribute to inconsistent findings regarding the prognostic value of TV. “The AI algorithms can overcome this by systematically and objectively analyzing the entire digitized specimen, consistently quantifying malignant lesions, and reducing inter-observer variability,” he added.
Methodology
The researchers analyzed 29,646 digitized H&E-stained slides from 992 patients who underwent radical prostatectomy. The median follow-up duration was 72.7 months. The study compared case-level algorithm results with pathologist assessments for International Society of Urological Pathology grade groups (GG), TV, and PTV.
The DLIA algorithm employed a deep learning-based semantic segmentation convolutional neural network model with a proprietary encoder-decoder structure. This was applied to tissue-containing image patches extracted from whole-slide images at 5× magnification. The model generated pixel-level likelihood values for each category: benign tissue, Gleason patterns 3, 4, and 5.
For each pixel classified as malignant, the algorithm computed a pixel-wise Gleason index (GI) as the sum of grade values weighted by their relative likelihood values. The case-level GI was defined as the average of pixel-level GIs across all tumor lesions, with ISUP GG determined by discretization of this case-level GI.
To calculate TV and PTV, the algorithm measured the areas of each tissue fragment and tumor lesion, accounting for serial sections and recut slides to avoid redundancy. Specimen volumes and TVs were determined by summing all tissue and tumor areas and multiplying by the average specimen slice thickness (4.5 mm). PTV was calculated by dividing TV by specimen volume.
Cancer Detection and Grading
While pathologists identified cancer in 986 cases and assigned GG in 980, the DLIA algorithm identified cancer and assigned GG to all 992 cases without omission. This included 12 cases where pathologists could not detect residual tumors (six cases) or assign GGs (six cases). Four of these patients without valid pathologist-assigned GGs experienced BCR, with one presenting BCR within one year.
The algorithm-assigned GG (aGG) showed fair concordance with pathologist assessments (pGG), with a linear-weighted Cohen’s kappa of 0.374. Compared with pathologists, the algorithm assigned the same GGs for 44.9% of patients, lower GGs for 26.6%, and higher GGs for 28.5%.
Despite these differences, both aGG and pGG demonstrated similar efficacy in predicting BPFS, with c-index values of 0.644 and 0.654, respectively (P=0.52). When stratifying patients by GG, BPFS rates significantly differed between GG1 and GG2, and between GG2 and GG3 for both pGG and aGG (all P<0.001), but differences were not significant between GG3 and GG4, or between GG4 and GG5.
Regarding the fair rather than strong concordance between pathologist and AI grading, Dr. Ha envisions a human-in-the-loop approach:
“The ideal collaboration can involve AI as a supportive tool, identifying areas of interest with malignant tumors, analyzing histological grades, and providing objective measurements, such as TV, to enhance pathologist efficiency and consistency. Pathologists remain essential for the final interpretation and integration of findings into clinical decision-making.”
Tumor Quantification
The DLIA-measured TV and PTV showed strong correlations with pathologist-based measurements, with Pearson’s correlation coefficients of 0.830 and 0.846, respectively. However, the algorithm consistently estimated smaller TV and PTV compared to pathologists, likely because it focused on measuring only epithelial cancer cells, excluding intra- and peritumoral empty spaces or stromal areas often included in pathologist assessments.
In addition, algorithm-based tumor quantifications (aTV and aPTV) demonstrated stronger efficacy in BPFS prediction than pathologist-based measurements (pTV and pPTV), with c-index values of 0.657 and 0.672 compared to 0.622 and 0.641, respectively. When patients were divided into quartile subgroups, significant differences in BPFS were observed between groups for all tumor measurements, except between the lowest two quartiles as measured by pathologists.
“The AI algorithm identified and quantified only malignant tumor cells, resulting in a more precise TV compared to pathologists’ quantifications, which were affected by subjectivity and variability,” Dr. Ha said. “Clinically, this AI-measured TV could allow us to better tailor post-operative care by more accurately stratifying patient risk.”
Both pathologist-based and algorithm-based tumor quantifications showed significant correlations with adverse pathological features such as extraprostatic extension, seminal vesicle invasion, perineural invasion, lymphovascular invasion, and surgical margin status.
Enhanced Risk Assessment
The researchers extended the Cancer of the Prostate Risk Assessment post-surgical (CAPRA-S) score by incorporating PTV. Incorporating algorithm-derived PTV (aPTV) into the CAPRA-S score significantly improved its predictive accuracy for BCR (P=0.006), increasing the c-index from 0.704 to 0.715. In contrast, incorporating pathologist-measured PTV (pPTV) did not yield comparable improvements.
When asked about the timeline for implementing such AI-enhanced risk assessment tools in clinical practice, Dr. Ha cautioned that
“While AI can enhance both diagnostics and prognostics, clinical implementation of AI-enhanced tools within routine practice is likely still years away because it requires overcoming hurdles like digital pathology system deployment, broad validation studies, and regulatory approvals.”
Future Directions
Dr. Ha noted that the study focused on radical prostatectomy specimens and that it is important to test the technology on biopsy specimens, which are often the first diagnostic step.
“Prostate biopsy specimens have different characteristics, including separate diagnostic requirements, which necessitate independent studies on the development and validation of such AI algorithms,” Dr. Ha explained. “There are already several commercially available AI algorithms for prostate biopsies, which are ready to affect early detection and treatment planning of prostate cancer.”
The researchers indicate that nationwide studies are being planned to evaluate the algorithm’s impact on workflow efficiency and diagnostic accuracy in real-world clinical settings, particularly in resource-constrained environments where expertise in uropathology is limited.
References
- Kwak TY, Lee CH, Park WY, et al. Clinical implications of deep learning based image analysis of whole radical prostatectomy specimens. Sci Rep. 2025;15(1):11006. Published 2025 Mar 31. doi:10.1038/s41598-025-95267-5
No audio available for this article yet.
No quiz available for this article yet.









