paper-with-me

Papers

Insights into a radiology-specialised multimodal large language model with sparse autoencoders

2025-07-17 · Kenza Bouzid, Shruthi Bannur, Felix Meissen, Daniel Coelho de Castro, Anton Schwaighofer, Javier Alvarez-Valle, Stephanie L. Hyland arxiv

Interpretability can improve the safety, transparency and trust of AI models, which is especially important in healthcare applications where decisions often carry significant consequences. Mechanistic interpretability, particularly through the use of sparse autoencoders (SAEs), offers a promising approach for uncovering human-interpretable features within large transformer-based models. In this study, we apply Matryoshka-SAE to the radiology-specialised multimodal large language model, MAIRA-2, to interpret its internal representations. Using large-scale automated interpretability of the SAE features, we identify a range of clinically relevant concepts - including medical devices (e.g., line and tube placements, pacemaker presence), pathologies such as pleural effusion and cardiomegaly, longitudinal changes and textual features. We further examine the influence of these features on model behaviour through steering, demonstrating directional control over generations with mixed success. Our results reveal practical and methodological challenges, yet they offer initial insights into the internal concepts learned by MAIRA-2 - marking a step toward deeper mechanistic understanding and interpretability of a radiology-adapted multimodal large language model, and paving the way for improved model transparency. We release the trained SAEs and interpretations: https://huggingface.co/microsoft/maira-2-sae.

📄 PDF Abstract BibTeX arXiv:2507.12950

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MAIRA-Seg: Enhancing Radiology Report Generation with Segmentation-Aware Multimodal Large Language Models

2024-11-18 · Harshita Sharma, Valentina Salvatelli, Shaury Srivastav, Kenza Bouzid 외

There is growing interest in applying AI to radiology report generation, particularly for chest X-rays (CXRs). This paper investigates whether incorporating pixel-level information through segmentation masks can improve …

SegmentationSemantic Segmentation

MAIRA-1: A specialised large multimodal model for radiology report generation

2023-11-22 · Stephanie L. Hyland, Shruthi Bannur, Kenza Bouzid, Daniel C. Castro 외

We present a radiology-specific multimodal model for the task for generating radiological reports from chest X-rays (CXRs). Our work builds on the idea that large language model(s) can be equipped with multimodal capabil…

Data AugmentationLanguage ModelingLanguage ModellingLarge Language Model

Opportunities and challenges in the application of large artificial intelligence models in radiology

2024-03-24 · Liangrui Pan, Zhenyu Zhao, Ying Lu, Kewei Tang 외

Influenced by ChatGPT, artificial intelligence (AI) large models have witnessed a global upsurge in large model research and development. As people enjoy the convenience by this AI large model, more and more large models…

Video Generation

Building RadiologyNET: Unsupervised annotation of a large-scale multimodal medical database

2023-07-27 · Mateja Napravnik, Franko Hržić, Sebastian Tschauner, Ivan Štajduhar

Background and objective: The usage of machine learning in medical diagnosis and treatment has witnessed significant growth in recent years through the development of computer-aided diagnosis systems that are often relyi…

ClusteringMedical DiagnosisSemantic SimilaritySemantic Textual Similarity

GRILLBot In Practice: Lessons and Tradeoffs Deploying Large Language Models for Adaptable Conversational Task Assistants

2024-02-12 · Sophie Fischer, Carlos Gemmell, Niklas Tecklenburg, Iain Mackie 외

We tackle the challenge of building real-world multimodal assistants for complex real-world tasks. We describe the practicalities and challenges of developing and deploying GRILLBot, a leading (first and second prize win…

Code GenerationManagementQuestion AnsweringWorld Knowledge