paper-with-me

Papers

FunnyNodules: A Customizable Medical Dataset Tailored for Evaluating Explainable AI

2025-11-19 · Luisa Gallée, Yiheng Xiong, Meinrad Beer, Michael Götz arxiv

Densely annotated medical image datasets that capture not only diagnostic labels but also the underlying reasoning behind these diagnoses are scarce. Such reasoning-related annotations are essential for developing and evaluating explainable AI (xAI) models that reason similarly to radiologists: making correct predictions for the right reasons. To address this gap, we introduce FunnyNodules, a fully parameterized synthetic dataset designed for systematic analysis of attribute-based reasoning in medical AI models. The dataset generates abstract, lung nodule-like shapes with controllable visual attributes such as roundness, margin sharpness, and spiculation. The target class is derived from a predefined attribute combination, allowing full control over the decision rule that links attributes to the diagnostic class. We demonstrate how FunnyNodules can be used in model-agnostic evaluations to assess whether models learn correct attribute-target relations, to interpret over- or underperformance in attribute prediction, and to analyze attention alignment with attribute-specific regions of interest. The framework is fully customizable, supporting variations in dataset complexity, target definitions, class balance, and beyond. With complete ground truth information, FunnyNodules provides a versatile foundation for developing, benchmarking, and conducting in-depth analyses of explainable AI methods in medical image analysis.

📄 PDF Abstract BibTeX arXiv:2511.15481

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models

2024-09-20 · Junfeng Jiang, Jiahao Huang, Akiko Aizawa

Recent developments in Japanese large language models (LLMs) primarily focus on general domains, with fewer advancements in Japanese biomedical LLMs. One obstacle is the absence of a comprehensive, large-scale benchmark …

Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI

2024-01-25 · Elron Bandel, Yotam Perlitz, Elad Venezian, Roni Friedman-Melamed 외

In the dynamic landscape of generative NLP, traditional text processing pipelines limit research flexibility and reproducibility, as they are tailored to specific dataset, task, and model combinations. The escalating com…

SUMMPILOT: Bridging Efficiency and Customization for Interactive Summarization System

2026-01-13 · JungMin Yun, Juhwan Choi, Kyohoon Jin, Soojin Jang 외 arxiv

This paper incorporates the efficiency of automatic summarization and addresses the challenge of generating personalized summaries tailored to individual users' interests and requirements. To tackle this challenge, we in…

Less Is More: A Comparison of Active Learning Strategies for 3D Medical Image Segmentation

2022-07-02 · Josafat-Mattias Burmeister, Marcel Fernandez Rosas, Johannes Hagemann, Jonas Kordt 외

Since labeling medical image data is a costly and labor-intensive process, active learning has gained much popularity in the medical image segmentation domain in recent years. A variety of active learning strategies have…

Active LearningBenchmarkingImage SegmentationMedical Image Segmentation+2

Medical artificial intelligence toolbox (MAIT): an explainable machine learning framework for binary classification, survival modelling, and regression analyses

2025-01-08 · Ramtin Zargari Marandi, Anne Svane Frahm, Jens Lundgren, Daniel Dawson Murray 외

While machine learning offers diverse techniques suitable for exploring various medical research questions, a cohesive synergistic framework can facilitate the integration and understanding of new approaches within unifi…

Binary ClassificationFeature Importance