paper-with-me

홈 › Papers

The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

2023-11-15 · Yifan Wu, Pengchuan Zhang, Wenhan Xiong, Barlas Oguz, James C. Gee, Yixin Nie

The study explores the effectiveness of the Chain-of-Thought approach, known for its proficiency in language tasks by breaking them down into sub-tasks and intermediate steps, in improving vision-language tasks that demand sophisticated perception and reasoning. We present the "Description then Decision" strategy, which is inspired by how humans process signals. This strategy significantly improves probing task performance by 50%, establishing the groundwork for future research on reasoning paradigms in complex vision-language tasks.

📄 PDF Abstract BibTeX arXiv:2311.09193

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

Sample Complexity of Autoregressive Reasoning: Chain-of-Thought vs. End-to-End

2026-04-13 · Steve Hanneke, Idan Mehalel, Shay Moran arxiv

Modern large language models generate text autoregressively, producing tokens one at a time. To study the learnability of such systems, Joshi et al. (COLT 2025) introduced a PAC-learning framework for next-token generato…

Natural Questions

Chain of Thought Prompt Tuning in Vision Language Models

2023-04-16 · Jiaxin Ge, Hongyin Luo, Siyuan Qian, Yulu Gan 외

Language-Image Pre-training has demonstrated promising results on zero-shot and few-shot downstream tasks by prompting visual models with natural language prompts. However, most recent studies only use a single prompt fo…

Domain Generalizationimage-classificationImage ClassificationLanguage Modeling+4

X-Ray-CoT: Interpretable Chest X-ray Diagnosis with Vision-Language Models via Chain-of-Thought Reasoning

2025-08-17 · Chee Ng, Liliang Sun, Shaoqing Tang arxiv

Chest X-ray imaging is crucial for diagnosing pulmonary and cardiac diseases, yet its interpretation demands extensive clinical experience and suffers from inter-observer variability. While deep learning models offer hig…

Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents

2025-03-11 · Rui Xu, Mingyu Wang, Xintao Wang, Dakuan Lu 외

Recent advances in LLM-based role-playing language agents (RPLAs) have attracted broad attention in various applications. While chain-of-thought reasoning has shown importance in many tasks for LLMs, the internal thinkin…

Multi-modal Latent Space Learning for Chain-of-Thought Reasoning in Language Models

2023-12-14 · Liqi He, Zuchao Li, Xiantao Cai, Ping Wang

Chain-of-thought (CoT) reasoning has exhibited impressive performance in language models for solving complex tasks and answering questions. However, many real-world questions require multi-modal information, such as text…

Machine Translation