paper-with-me

홈 › Papers

Limits and Gains of Test-Time Scaling in Vision-Language Reasoning

2025-12-11 · Mohammadjavad Ahmadpour, Amirmahdi Meighani, Payam Taebi, Omid Ghahroodi, Amirmohammad Izadi, Mahdieh Soleymani Baghshah arxiv

Test-time scaling (TTS) has emerged as a powerful paradigm for improving the reasoning ability of Large Language Models (LLMs) by allocating additional computation at inference, yet its application to multimodal systems such as Vision-Language Models (VLMs) remains underexplored. In this work, we present a systematic empirical study of inference time reasoning methods applied across both open-source and closed-source VLMs on different benchmarks. Our results reveal that while closed-source models consistently benefit from structured reasoning and iterative Self-Refinement, open-source VLMs show inconsistent behavior: external verification provides the most reliable gains, whereas iterative refinement often degrades performance. We further find that the effectiveness of TTS is dataset-dependent, yielding clear improvements on multi-step reasoning tasks but offering only limited gains on perception-focused benchmarks. These findings demonstrate that TTS is not a universal solution and must be tailored to both model capabilities and task characteristics, motivating future work on adaptive TTS strategies and multimodal reward models.

📄 PDF Abstract BibTeX arXiv:2512.11109

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment

2026-02-12 · Jacky Kwok, Xilun Zhang, Mengdi Xu, Yuejiang Liu 외 arxiv

The long-standing vision of general-purpose robots hinges on their ability to understand and act upon natural language instructions. Vision-Language-Action (VLA) models have made remarkable progress toward this goal, yet…

Instruction Following

Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling

2026-01-05 · Falcon LLM Team, Iheb Chaabane, Puneesh Khanna, Suhail Mohmad 외 arxiv

This work introduces Falcon-H1R, a 7B-parameter reasoning-optimized model that establishes the feasibility of achieving competitive reasoning performance with small language models (SLMs). Falcon-H1R stands out for its p…

The Compression Gap: Why Discrete Tokenization Limits Vision-Language-Action Model Scaling

2026-04-03 · Takuya Shiba arxiv

Scaling Vision-Language-Action (VLA) models by upgrading the vision encoder is expected to improve downstream manipulation performance--as it does in vision-language modeling. We show that this expectation fails when act…

Calibrating Self-supervised Monocular Depth Estimation

2020-09-16 · Robert McCraith, Lukas Neumann, Andrea Vedaldi

In the recent years, many methods demonstrated the ability of neural networks to learn depth and pose changes in a sequence of images, using only self-supervision as the training signal. Whilst the networks achieve good …

Depth EstimationMonocular Depth Estimation

Beyond Many-Shot Translation: Scaling In-Context Demonstrations For Low-Resource Machine Translation

2026-02-04 · Luis Frentzen Salim, Esteban Carlin, Alexandre Morinvil, Xi Ai 외 arxiv

Building machine translation (MT) systems for low-resource languages is notably difficult due to the scarcity of high-quality data. Although Large Language Models (LLMs) have improved MT system performance, adapting them…

Machine Translation