paper-with-me

홈 › Papers

Cyclic Vision-Language Manipulator: Towards Reliable and Fine-Grained Image Interpretation for Automated Report Generation

2024-11-08 · Yingying Fang, Zihao Jin, Shaojie Guo, Jinda Liu, Zhiling Yue, Yijian Gao, Junzhi Ning, Zhi Li, Simon Walsh, Guang Yang

Despite significant advancements in automated report generation, the opaqueness of text interpretability continues to cast doubt on the reliability of the content produced. This paper introduces a novel approach to identify specific image features in X-ray images that influence the outputs of report generation models. Specifically, we propose Cyclic Vision-Language Manipulator CVLM, a module to generate a manipulated X-ray from an original X-ray and its report from a designated report generator. The essence of CVLM is that cycling manipulated X-rays to the report generator produces altered reports aligned with the alterations pre-injected into the reports for X-ray generation, achieving the term "cyclic manipulation". This process allows direct comparison between original and manipulated X-rays, clarifying the critical image features driving changes in reports and enabling model users to assess the reliability of the generated texts. Empirical evaluations demonstrate that CVLM can identify more precise and reliable features compared to existing explanation methods, significantly enhancing the transparency and applicability of AI-generated reports.

📄 PDF Abstract BibTeX arXiv:2411.05261

Code (0)

등록된 구현이 없습니다.

Tasks

counterfactualDecision MakingText Generation

Similar Papers 제목 키워드 기반

Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather Removal

2025-03-12 · Rongxin Liao, Feng Li, Yanyan Wei, Zenglin Shi 외

Universal adverse weather removal (UAWR) seeks to address various weather degradations within a unified framework. Recent methods are inspired by prompt learning using pre-trained vision-language models (e.g., CLIP), lev…

Image RestorationPrompt Learning

HyReach: Vision-Guided Hybrid Manipulator Reaching in Unseen Cluttered Environments

2026-03-22 · Shivani Kamtikar, Kendall Koe, Justin Wasserman, Samhita Marri 외 arxiv

As robotic systems increasingly operate in unstructured, cluttered, and previously unseen environments, there is a growing need for manipulators that combine compliance, adaptability, and precise control. This work prese…

Motion Planning

AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models

2025-11-03 · Sarthak Mishra, Rishabh Dev Yadav, Avirup Das, Saksham Gupta 외 arxiv

The rapid progress of vision--language models (VLMs) has sparked growing interest in robotic control, where natural language can express the operation goals while visual feedback links perception to action. However, dire…

Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots

2025-10-20 · Haochen Su, Cristian Meo, Francesco Stella, Andrea Peirone 외 arxiv

Robotic systems are increasingly expected to operate in human-centered, unstructured environments where safety, adaptability, and generalization are essential. Vision-Language-Action (VLA) models have been proposed as a …

Per-Group Error, Not Total MSE: Fine-Tuning Vision-Language-Action Models for 11-DoF Mobile Manipulation

2026-05-29 · Pau Montagut Bofi, Mario García Blasco, Tessa Pulli, Markus Vincze arxiv

Fine-tuning Vision-Language-Action (VLA) models for mobile manipulators with heterogeneous joint spaces can produce a counterintuitive result: the checkpoint with the lowest aggregate MSE is not the one that performs bes…