paper-with-me

홈 › Papers

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

2024-12-17 · Zihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei, Xiaocheng Feng, Wanxiang Che, Min Li, Libo Qin

Large Vision-Language Models (LVLMs) have recently demonstrated amazing success in multi-modal tasks, including advancements in Multi-modal Chain-of-Thought (MCoT) reasoning. Despite these successes, current benchmarks still follow a traditional paradigm with multi-modal input and text-modal output, which leads to significant drawbacks such as missing visual operations and vague expressions. Motivated by this, we introduce a novel Chain of Multi-modal Thought (CoMT) benchmark to address these limitations. Different from the traditional MCoT benchmark, CoMT requires both multi-modal input and multi-modal reasoning output, aiming to mimic human-like reasoning that inherently integrates visual operation. Specifically, CoMT consists of four categories: (1) Visual Creation, (2) Visual Deletion, (3) Visual Update, and (4) Visual Selection to comprehensively explore complex visual operations and concise expression in real scenarios. We evaluate various LVLMs and strategies on CoMT, revealing some key insights into the capabilities and limitations of the current approaches. We hope that CoMT can inspire more research on introducing multi-modal generation into the reasoning process.

📄 PDF Abstract BibTeX arXiv:2412.12932

Code (1)

czhhzc/CoMT 공식 구현

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation

2024-06-17 · Yue Jiang, Jiawei Chen, Dingkang Yang, Mingcheng Li 외

Automatic medical report generation (MRG), which possesses significant research value as it can aid radiologists in clinical diagnosis and report composition, has garnered increasing attention. Despite recent progress, g…

DiagnosticHallucinationMedical Report Generation

TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis

2025-10-07 · Austin Feng, Andreas Varvarigos, Ioannis Panitsas, Daniela Fernandez 외 arxiv

Modern enterprises generate vast streams of time series metrics when monitoring complex systems, known as observability data. Unlike conventional time series from domains such as climate, observability data are zero-infl…

Anomaly Detection

M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

2024-05-26 · Qiguang Chen, Libo Qin, Jin Zhang, Zhi Chen 외

Multi-modal Chain-of-Thought (MCoT) requires models to leverage knowledge from both textual and visual modalities for step-by-step reasoning, which gains increasing attention. Nevertheless, the current MCoT benchmark sti…

Multimodal Transformer With a Low-Computational-Cost Guarantee

2024-02-23 · Sungjin Park, Edward Choi

Transformer-based models have significantly improved performance across a range of multimodal understanding tasks, such as visual question answering and action recognition. However, multimodal Transformers significantly …

Action RecognitionQuestion AnsweringVisual Question Answering

Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning

2026-01-04 · Wenting Lu, Didi Zhu, Tao Shen, Donglin Zhu 외 arxiv

Multi-modal reasoning requires the seamless integration of visual and linguistic cues, yet existing Chain-of-Thought methods suffer from two critical limitations in cross-modal scenarios: (1) over-reliance on single coar…

Visual Reasoning