
Several of the talks presented at Pathology Visions 2025 demonstrated how next-generation AI models and large language models (LLMs) can enhance diagnostic capabilities and improve pathology workflows.
Hematopathology Foundation Model Demonstrates Strong Performance on Blood and Bone Marrow Analysis
Hematopathology requires pathologists to evaluate blood and bone marrow specimens for a wide range of conditions, from anemia to leukemia. Unlike solid tissue pathology, where pathologists examine architectural patterns and tissue structure, hematopathologists must assess cellular morphology, maturation sequences, and cell differentials, which often requires counting hundreds of individual cells per case. Adoption of AI models in hematopathology has lagged behind digital transformation in solid tumor pathology, although companies like Grundium, Scopio Labs, and Techcyte have developed specialized scanners and analysis tools for blood smears and bone marrow specimens in recent years.
Now, a new AI model specifically designed for hematopathology is on par or outperforming general-purpose pathology models on tasks involving blood and bone marrow specimens, according to research presented by Brendan O’Fallon, PhD, Bioinformatics Director at ARUP Labs. The foundation model was trained on 27,735 whole slide images from 9,544 cases and represents one of the first AI systems built exclusively for hematopathology applications. Although numerous foundation models have emerged for solid tumor pathology in recent years, blood and bone marrow specimens have largely been overlooked, according to O’Fallon.
“There hasn’t really been a model dedicated to hematopathology,” O’Fallon explained during his presentation. The diversity of hematopathology specimens, including peripheral blood smears, bone marrow core biopsies, bone marrow aspirate smears, clots, and touch preps, requires specialized approaches that differ from solid tumor analysis.
The team’s first challenge involved developing specialized techniques to extract meaningful data from various specimen types.
“The first major obstacle for working with hematopathology data is defining custom models for the various specimen types,” O’Fallon said in an interview with Pathology News. “The simple tissue-detection algorithms used in solid tumor models aren’t sufficient here.”
For bone marrow cores, the team focused on identifying high-quality tissue regions while differentiating core samples from control tissue. Blood and aspirate smears required a different strategy. Rather than randomly sampling tiles from tissue, the researchers trained object detection models to identify individual white blood cells in areas where cells formed an even monolayer without clumping. This approach generated over 100 million individual cell detections, which, according to O’Fallon, is “probably the biggest fundamental difference from solid tissue models.”
The training dataset included approximately 165 million tiles at multiple magnifications (5x, 10x, 20x, and 40x) with a 60/40 split between bone marrow core tiles and individual cell images. All slides were scanned at 40x resolution. The Vision Transformer-Large model was trained using the DINO v2 algorithm over approximately two epochs on an 8xH100 system.
The model was evaluated against three established pathology foundation models: Prov-GigaPath, Uni, and H-Optimus-0. In terms of cellularity estimation, the hematopathology-specific model achieved the lowest root mean square error at 7.71. For fibrosis quantification on reticulin-stained bone marrow cores, it narrowly outperformed competitors with an RMSE of 0.52.
The strongest improvement appeared in white blood cell classification tasks. On a public dataset containing 21 different cell types from peripheral blood, the model achieved an F1 score of 0.79, which was substantially higher than that achieved by general-purpose models.
“Our model has much lower error than other models overall when it comes to cell differential estimations,” O’Fallon noted.
The research team also explored advanced training techniques. Multi-task training, where the model learned to predict multiple outcomes simultaneously, improved performance compared to single-task approaches. Contrastive training using additional data sources, such as next-generation sequencing results and flow cytometry, further boosted accuracy. However, this benefit applied only to transformer-based aggregation models and not to attention-based multiple instance learning approaches.
The attention mechanisms of the model revealed disease-specific patterns that aligned with clinical expectations, O’Fallon emphasized in his presentation. For acute myeloid leukemia cases, the model focused heavily on blast cells while largely ignoring mature neutrophils. Conversely, in plasma cell neoplasms, the model learned to prioritize plasma cells, without being explicitly programmed to do so.
Why this matters?
“Cellularity estimation, IHC quantification, and fibrosis grading are routine tasks that hematopathologists perform multiple times daily,” O’Fallon explained in the interview. “By automating this process, we can save a few minutes of precious time and also provide reproducible, standardized values.”
He also emphasized that hematopathologists dedicate substantial time to investigating cell morphology and counting cells.
“Most solid tumor models have very few, if any, representatives of these important cell types. That’s why our models perform better than others out there. I would imagine other sub-disciplines may also benefit from inclusion of domain-specific data like this.”

Brendan O'Fallon, PhD, Bioinformatics Director at ARUP Labs, USA.
No audio available for this article yet.
No quiz available for this article yet.









