
by Sidney Ocanagil-Tunstall
“AI is now a reality.” — Dr. Miere Crispin Ortuzar, University of Cambridge, School of Clinical Medicine, UK
For digital pathology enthusiasts, those five words opened one of the most anticipated sessions at this year’s ESMO congress. Until recently, AI in imaging at ESMO had taken a back seat, a topic for hallway debates and Q&A sessions rather than a main-stage appointment. But this year, the focus was clearly different. For the first time, ESMO introduced a dedicated AI track, underscoring how integral artificial intelligence has become to diagnostic approaches. The hall, close to bursting, illustrated the audience’s appetite for this long-overdue conversation.
Presenting the first session in the track, Dr. Crispin Ortuzar took to the stage.
“We’ve been talking about it [AI adoption] for a long time, but now it is really happening. It is happening for healthcare!”
From 2010 to 2024, the number of FDA approvals for medical devices using AI has shown an exponential increase. Studies have observed the same trend in CE-IVD-approved devices here in Europe.
“The foundations have been built,” she said, “and the ground is more fertile than ever.”
With over a decade of groundwork laid through regulatory approvals, academic research, and hospital readiness, healthcare systems are now poised to integrate AI tools at scale.
So, how are we using AI in healthcare and more specifically, diagnostics?
Foundation Models
First in line, and central to Dr. Crispin Ortuzar’s talk, were foundation models. Large scale algorithms trained on vast imaging datasets that learn the underlying structure of data rather than being engineered for a single diagnostic task. They are the “Swiss Army knives” of AI: versatile tools that can be used for segmentation, detection, or report generation. While abundantly useful across a wide range of tasks, you should not consider them as a ‘one size fits all’ model. Dr. Crispin reflected,
“If you have a specific clinical research question, you’ll need to look beyond foundation models.”
Dr. Crispin’s group is particularly interested in spatial biomarkers, with her team aiming to “classify tissue in a spatial manner.” With this task in mind the opportunity to put a foundation model to the test presented itself. Together with her team, Dr. Crispin developed SMILe (soon to be published in Nature Cancer), an AI method for mapping and classifying tissue regions within whole-slide images. This model demonstrated that encoding slides with a foundation model improved spatial classification accuracy by almost ten percent in some cases, relative to standard, non–foundation-based approaches.
The takeaway was clear: foundation models show fantastic promise, but it’s not just about applying them indiscriminately. It’s about understanding the context in which you want to use them and re-engineering them thoughtfully to meet the realities of the problem at hand. Only then can you truly harness their full potential.
Visual Language Models
Next under the spotlight were Visual Language Models (VLMs) — algorithms capable of simultaneously understanding both medical images and text, allowing for the interpretation and generation of language about or from visual data. In pathology and radiology, VLMs can analyse whole-slide images or scans alongside associated reports or annotations.
An example includes PathChat, a pathology-focused multimodal model that connects visual understanding of whole-slide images with language-based reasoning.
With VLMs, you can now “generate a report by giving your model an image and asking: what are the radiology findings related to the image?” The downside to the current use of VLMs is that they are unable to work on the bigger picture to classify whole-slide images. Instead, they “break whole-slide images into tiles.” If you’re probing for spatial descriptors of tissue, you need to understand each descriptor in its relevant context — that is, relative to the rest of the image and its microenvironment.
With that in mind, Dr. Crispin’s team set out to re-engineer a VLM that would allow them to interpret spatial descriptors in their biological context. They named this model ALPaCA.¹
ALPaCA utilises a chatbot (Llama 3.1) to answer questions about an entire whole-slide image by first compressing multi-magnification patch features from a pathology encoder (CONCH) into a small set of slide tokens. Thanks to this architecture, the model captures both fine and broad morphological details before generating an answer.
How did the team evaluate this model? With closed-ended questions where answers were, for instance, A, B, or C — the model achieved an accuracy of 92%. For open-ended questions, “we saw performances of 60–70%, dependent on the task.”
This work underscores AI evolving from a descriptive tool to a system capable of reasoning, bridging the gap between human interpretation and AI-driven insight while proving itself to be remarkably capable.
Agentic AI
Looking forward, the session explored agentic AI, autonomous systems capable of performing multi-step analytical tasks.
“These agents are able to interact with the environment and the data presented to them, evolving autonomously.”
New studies showcase the use of multiple agents; each trained in specific areas. “The agents have discussions between themselves and do research.”
“We are seeing EU calls for substantial funding for this [agentic AI].”Could we one day harness these tools to simulate multidisciplinary discussions, integrating radiology and pathology data?
“There’s a lot of excitement around multi-omic modelling… the agents are designed for it.”
Instead of having radiologists and pathologists, could the future instead feature diagnosticians empowered by AI agents, capable of utilising multimodal data to reach their final diagnosis?
This session was just the beginning. As ESMO’s new AI-based image biomarkers track unfolds, the dialogue between oncology, radiology, and pathology will only grow more connected and much more consequential.
References
- https://www.medrxiv.org/content/10.1101/2025.04.22.25326190v1.full
No audio available for this article yet.
No quiz available for this article yet.









