paper-with-me

홈 › Papers

CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation

2026-01-12 · Chaoyu Li, Deeparghya Dutta Barua, Fei Tao, Pooyan Fazli arxiv

Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasoning remains unstable. Repeated sampling over the same input often produces divergent reasoning trajectories and inconsistent final predictions. To address this, we introduce two complementary approaches inspired by test-time scaling: (1) CASHEW, an inference-time framework that stabilizes reasoning by iteratively aggregating multiple candidate trajectories into higher-quality reasoning traces, with explicit visual verification filtering hallucinated steps and grounding reasoning in visual evidence, and (2) CASHEW-RL, a learned variant that internalizes this aggregation behavior within a single model. CASHEW-RL is trained using Group Sequence Policy Optimization (GSPO) with a composite reward that encourages correct answers grounded in minimal yet sufficient visual evidence, while adaptively allocating reasoning effort based on task difficulty. This training objective enables robust self-aggregation at inference. Extensive experiments on 13 image understanding, video understanding, and video reasoning benchmarks show significant performance improvements, including gains of up to +26.2 percentage points on ScienceQA and +9.1 percentage points on EgoSchema.

📄 PDF Abstract BibTeX arXiv:2601.08010

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Reasoning

Similar Papers 제목 키워드 기반

Mapping smallholder cashew plantations to inform sustainable tree crop expansion in Benin

2023-01-01 · Leikun Yin, Rahul Ghosh, Chenxi Lin, David Hale 외

Cashews are grown by over 3 million smallholders in more than 40 countries worldwide as a principal source of income. As the third largest cashew producer in Africa, Benin has nearly 200,000 smallholder cashew growers co…

ClusteringDecision Making

A Remote Approach to Cashew Orchard Detection: Leveraging Active Learning with Satellite Imagery in Guinea-Bissau

2026-08-12 · Miguel Pereira, Sofia C. Pereira, Maria J. P. Vasconcelos, Patrícia Guedes 외 arxiv

Cashew production is a widespread economic activity in Guinea-Bissau, as well as other countries in West Africa. However, unregulated cashew production can be directly associated with increasing regionwide deforestation …

Active Learning

IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning

2026-05-18 · Chenghao Li, Fusheng Hao, Xikai Zhang, Likang Xiao 외 arxiv

Multimodal large language models via reinforcement learning (RL) have demonstrated remarkable capabilities in complex visual reasoning tasks, yet they remain limited in long-horizon multimodal scenarios, often suffering …

Reinforcement LearningMultimodal ReasoningVisual GroundingVisual Reasoning

Cashew dataset generation using augmentation and RaLSGAN and a transfer learning based tinyML approach towards disease detection

2023-04-18 · Varsha Jayaprakash, Akilesh K, Ajay Kumar, Balamurugan M. S 외

Cashew is one of the most extensively consumed nuts in the world, and it is also known as a cash crop. A tree may generate a substantial yield in a few months and has a lifetime of around 70 to 80 years. Yet, in addition…

Dataset GenerationTransfer Learning

Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering

2025-08-31 · Changin Choi, Wonseok Lee, Jungmin Ko, Wonjong Rhee arxiv

Knowledge-intensive visual question answering (VQA) requires external knowledge beyond image content, demanding precise visual grounding and coherent integration of visual and textual information. Although multimodal ret…

Visual Question AnsweringVisual Grounding