At PathAI, we are constantly pushing the boundaries of how foundation models can transform digital pathology. Today, we’re excited to share our latest work on Multimodal Pathology: a new approach that combines the visual power of our PLUTO-4 foundation model with rich, descriptive language to improve disease classification.
How it works:
Vision + Language with Contrastive Learning: We combine image embeddings from PLUTO with histological descriptions using contrastive learning. The model learns to align image features directly with descriptive text.
Better Performance:
This approach outperformed image-only MIL models, showing a ~4-6% improvement in Dermatopathology and ~8-10% improvement in GI pathology with similar inference cost.
Why it matters:
By bringing language into the loop, we’re moving from fixed-label classification toward more flexible, expressive systems that better reflect how pathology is practiced, with the potential to handle rare and unseen conditions through open-vocabulary prediction without retraining, natural language search over slide datasets, and more adaptable AI that can evolve with new knowledge and emerging diagnostic categories. This is a step towards AI that is not just accurate, but versatile, scalable, and grounded in clinical reasoning.








