paper-with-me

홈 › Papers

Test-Time Scaling for Small VLMs on Multilingual Visual MCQ

2026-07-10 · Spiros Baxevanakis, Peng-Jian Yang arxiv

Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models remains unclear. We examine this on EXAMS-V, a multilingual visual multiple-choice benchmark, comparing self-consistency, describe-then-reason with PRM-guided beam search, and two post-hoc selectors across Qwen2.5-VL-7B-Instruct and Qwen3.5-4B. What matters is the conditions under which TTS runs, not the search or verification machinery. The largest factor is parseability: an early prompt format left many chains reasoning correctly yet never committing to an answer letter, which a standard answer cue and a guided repair step largely remove. A larger decoding budget removes the rest: raising the per-chain token limit from 1k to 2k recovers 3.7 pp, whereas sampling more chains (8 to 16) adds only 0.15 pp. Once chains have room to finish, elaborate methods contribute little: PRM-guided beam search trails plain self-consistency by 0.39 pp at over eight times the cost, and neither a training-free generative critic nor a trained multimodal PRM beats majority vote across both policies. The largest gain comes instead from the policy model itself (+11.4 pp). Our best configuration reaches 84.1% on the held-out ImageCLEF 2026 test split, ranking first on the Visual MCQ leaderboard.

📄 PDF Abstract BibTeX arXiv:2607.09438

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

On Test-Time Scaling for Vision-Language Models

2026-06-27 · Fawaz Sammani, Tzoulio Chamiti, Nikos Deligiannis arxiv

Test-time scaling is a paradigm where large models use additional compute at inference to achieve better performance, without changing model weights. While it has been widely studied for Large Language Models (LLMs), its…

Efficient Test-Time Scaling for Small Vision-Language Models

2025-10-03 · Mehmet Onurcan Kaya, Desmond Elliott, Dim P. Papadopoulos arxiv

Small Vision-Language Models (VLMs) provide a computationally efficient alternative to larger models, at the cost of weaker generalization abilities and downstream task performance. These shortcomings could be addressed …

Computational EfficiencyTest-time Adaptation

Florenz: Scaling Laws for Systematic Generalization in Vision-Language Models

2025-03-12 · Julian Spravil, Sebastian Houben, Sven Behnke

Cross-lingual transfer enables vision-language models (VLMs) to perform vision tasks in various languages with training data only in one language. Current approaches rely on large pre-trained multilingual language models…

Cross-Lingual TransferImage CaptioningLarge Language ModelMachine Translation+3

Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned

2025-09-27 · Brandon Ong, Tej Deep Pala, Vernon Toh, William Chandra Tjhi 외 arxiv

Process Reward Models (PRMs) provide step-level supervision that improves the reliability of reasoning in large language models. While PRMs have been extensively studied in text-based domains, their extension to Vision L…

Multimodal ReasoningVisual Grounding

Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning

2025-02-24 · Guijin Son, Jiwoo Hong, Hyunwoo Ko, James Thorne

Scaling pre-training compute has proven effective for achieving mulitlinguality, but does the same hold for test-time scaling? In this work, we introduce MCLM, a multilingual math benchmark featuring competition-level pr…

MathMathematical Reasoning