paper-with-me

Papers

Enhancing Visual Inspection Capability of Multi-Modal Large Language Models on Medical Time Series with Supportive Conformalized and Interpretable Small Specialized Models

2025-01-27 · Huayu Li, Xiwen Chen, Ci Zhang, Stuart F. Quan, William D. S. Killgore, Shu-Fen Wung, Chen X. Chen, Geng Yuan, Jin Lu, Ao Li

Large language models (LLMs) exhibit remarkable capabilities in visual inspection of medical time-series data, achieving proficiency comparable to human clinicians. However, their broad scope limits domain-specific precision, and proprietary weights hinder fine-tuning for specialized datasets. In contrast, small specialized models (SSMs) excel in targeted tasks but lack the contextual reasoning required for complex clinical decision-making. To address these challenges, we propose ConMIL (Conformalized Multiple Instance Learning), a decision-support SSM that integrates seamlessly with LLMs. By using Multiple Instance Learning (MIL) to identify clinically significant signal segments and conformal prediction for calibrated set-valued outputs, ConMIL enhances LLMs' interpretative capabilities for medical time-series analysis. Experimental results demonstrate that ConMIL significantly improves the performance of state-of-the-art LLMs, such as ChatGPT4.0 and Qwen2-VL-7B. Specifically, \ConMIL{}-supported Qwen2-VL-7B achieves 94.92% and 96.82% precision for confident samples in arrhythmia detection and sleep staging, compared to standalone LLM accuracy of 46.13% and 13.16%. These findings highlight the potential of ConMIL to bridge task-specific precision and broader contextual reasoning, enabling more reliable and interpretable AI-driven clinical decision support.

📄 PDF Abstract BibTeX arXiv:2501.16215

Code (1)

HuayuLiArizona/Conformalized-Multiple-Instance-Learning-For-MedTS 공식 구현 pytorch

Tasks

Arrhythmia DetectionConformal PredictionDecision MakingMultiple Instance LearningSleep StagingTime SeriesTime Series Analysis

Similar Papers 제목 키워드 기반

MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding

2026-05-29 · Qian Kou, Xiaofeng Shi, Yulin Li, Xiaosong Qiu 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks. However, they remain brittle on mechanical engineering drawings, where high annotation…

Visual Question Answering

A Study on Unsupervised Anomaly Detection and Defect Localization using Generative Model in Ultrasonic Non-Destructive Testing

2024-05-26 · Yusaku Ando, Miya Nakajima, Takahiro Saitoh, Tsuyoshi Kato

In recent years, the deterioration of artificial materials used in structures has become a serious social issue, increasing the importance of inspections. Non-destructive testing is gaining increased demand due to its ca…

Anomaly DetectionDefect Detectionobject-detectionObject Detection+1

ActFER: Agentic Facial Expression Recognition via Active Tool-Augmented Visual Reasoning

2026-04-10 · Shifeng Liu, Zhengye Zhang, Sirui Zhao, Xinglong Mao 외 arxiv

Recent advances in Multimodal Large Language Models (MLLMs) have created new opportunities for facial expression recognition (FER), moving it beyond pure label prediction toward reasoning-based affect understanding. Howe…

Facial Expression RecognitionReinforcement LearningMultimodal ReasoningVisual Reasoning

Advancing Precision in Multi-Point Cloud Fusion Environments

2025-08-05 · Ulugbek Alibekov, Vanessa Staderini, Philipp Schneider, Doris Antensteiner arxiv

This research focuses on visual industrial inspection by evaluating point clouds and multi-point cloud matching methods. We also introduce a synthetic dataset for quantitative evaluation of registration method and variou…

Point Clouds

DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes

2025-05-29 · Sungjune Park, Hyunjun Kim, Junho Kim, Seongho Kim 외

Multimodal Large Language Models (MLLMs) have demonstrated significant visual understanding capabilities, yet their fine-grained visual perception in complex real-world scenarios, such as densely crowded public areas, re…

Decision MakingReinforcement Learning (RL)