paper-with-me

홈 › Papers

Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models

2026-05-29 · Yijie Tong, Yifan Hou, Shaobo Cui, Antoine Bosselut, Mrinmaya Sachan arxiv

Test-time compute (TTC) strategies have emerged as a lightweight approach to boost reasoning in large language models (LLMs). However, their application and benefits for vision-language models (VLMs) remain underexplored. We present a systematic study of TTC across seven VLMs and six benchmarks, specifically analyzing feature-based scoring and majority voting methods. We find that feature heuristics fail and voting yields only modest gains in single-model settings. We theoretically show that this limitation stems from a lack of prediction diversity: when outputs are highly correlated, voting provides little benefit. In contrast, multi-model ensembles offer richer diversity, yet standard majority voting fails to account for varying model capabilities. To address this, we propose Entropy-based TTC (ETTC), which selects the most confident prediction based on predictive entropy. Our method reduces to majority voting in the single-model case, but in model ensembles, it leverages confidence disparities to prioritize stronger models. We prove that ETTC outperforms majority voting under mild assumptions and empirically demonstrate that it consistently surpasses both voting and the best individual model. Crucially, our results show that smaller models can synergistically enhance larger ones, unlocking ensembling gains not achievable with standard strategies.

📄 PDF Abstract BibTeX arXiv:2605.30713

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning

2025-06-05 · Ho-Lam Chung, Teng-Yun Hsiao, Hsiao-Ying Huang, Chunerh Cho 외

Test-Time Scaling (TTS) improves the reasoning performance of Large Language Models (LLMs) by allocating additional compute during inference. We conduct a structured survey of TTS methods and categorize them into samplin…

DiversityMathematical Reasoning

Where You Inject Diversity Matters: A Unified Framework for Diverse Generation

2026-06-09 · Cheng Zhang, Rui Xin, Chudi Zhong arxiv

Open-ended generation tasks often require a set of meaningfully different outputs, yet large language models often produce similar generations. Existing test-time diversity methods operate at different stages of generati…

Diversity Matters: Dataset Diversification and Dual-Branch Network for Generalized AI-Generated Image Detection

2026-03-29 · Nusrat Tasnim, Kutub Uddin, Khalid Malik arxiv

The rapid proliferation of AI-generated images, powered by generative adversarial networks (GANs), diffusion models, and other synthesis techniques, has raised serious concerns about misinformation, copyright violations,…

Boosting Ensemble Accuracy by Revisiting Ensemble Diversity Metrics

2021-06-19 · CVPR 2021 1 · Yanzhao Wu, Ling Liu, Zhongwei Xie, Ka-Ho Chow 외

Neural network ensembles are gaining popularity by harnessing the complementary wisdom of multiple base models. Ensemble teams with high diversity promote high failure independence, which is effective for boosting th…

DiversityEnsemble LearningEnsemble PruningImage Classification

Graph Exploration Matters: Improving both individual-level and system-level diversity in WeChat Feed Recommender

2023-05-29 · Shuai Yang, Lixin Zhang, Feng Xia, Leyu Lin

There are roughly three stages in real industrial recommendation systems, candidates generation (retrieval), ranking and reranking. Individual-level diversity and system-level diversity are both important for industrial …

DiversityRecommendation SystemsRerankingRetrieval