
The following abstract is drawn from a recently published paper in Journal of Pathology Informatics. We invite you to read the full paper and join the conversation, become a member of the Pathology News community to share your thoughts, ask questions, and engage with others around this work.
Authors: Joshua W. Tashman, Chris Gorman, Emilio Madrigal
Abstract
Medical imaging data are an essential resource for research and teaching; however, regulations such as the Health Insurance Portability and Accountability Act require careful consideration of protected health information (PHI) contained within images and their metadata. De-identification, the removal of PHI to prevent subject re-identification, is often essential for regulatory compliance. In pathology, whole-slide imaging (WSI) files—high-resolution scans of entire glass slides—are essential for research and machine learning model training. Currently, many scanner vendors use proprietary WSI formats; some vendors offer de-identification tools, which are often manual and tedious to use. Alternatively, physical slides may be re-scanned with obfuscated labels, though this is resource intensive. Here, we describe an institutional-scale automated WSI de-identification pipeline, which utilizes an informatics-based approach to convert clinical WSI from 11 proprietary formats into de-identified WSI in multiple open formats. This approach minimizes manual processes and additional slide-scanning hardware, storage, and personnel. De-identification requests are submitted through a zero-footprint web portal following the Fast Healthcare Interoperability Resources ServiceRequest data model. De-identification is verified via a human-in-the-loop review, and archives are distributed using cloud-based storage. Since deployment in November 2024, our pipeline has de-identified 819 cases comprising 4322 WSI files generated by 7 scanner models in 4 formats. We demonstrate that the conversion process is predictable, linearly scalable, and reliable across de-identification request sizes (238× variation), image sizes (120× variation), and capture format. This pipeline eliminates the need for manual de-identification or re-scanning while achieving high throughput and reliability at an institutional scale.
Read the full article: An automated end-to-end pipeline for the management, de-identification, and distribution of whole-slide images using DICOM: An institutional implementation – ScienceDirect
Affiliations
Department of Pathology, Mass General Brigham, Boston, MA, United States of America
No audio available for this article yet.
No quiz available for this article yet.









