paper-with-me

홈 › Papers

Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference

2026-08-19 · Zishan Ahmad, Vishal Vaddina arxiv

Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking penalties; especially in document tasks where visual layouts drive complexity. To address this, we introduce BudgetDoc, the first multimodal benchmark providing explicit supervision for model-budget-performance trade-offs across three document tasks. Using BudgetDoc, we train DRB (Document-Reasoning Balancer), an approx. 1B-parameter pre-flight estimator (SigLIP-2 + Qwen3-0.6B) that predicts ordinal model performance across budget levels, achieving a 0.753 weighted F1. When dynamically allocating reasoning budgets across five frontier models and three datasets, DRB matches or improves F1 scores compared to always-maximum-budget baselines in 9 of 15 configurations while drastically reducing cost. Finally, preliminary evaluations demonstrate DRB's potential to generalize to cross-model selection.

📄 PDF Abstract BibTeX arXiv:2608.18591

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Conformalized Multimodal Uncertainty Regression and Reasoning

2023-09-20 · Domenico Parente, Nastaran Darabi, Alex C. Stutts, Theja Tulabandhula 외

This paper introduces a lightweight uncertainty estimator capable of predicting multimodal (disjoint) uncertainty bounds by integrating conformal prediction with a deep-learning regressor. We specifically discuss its app…

Conformal PredictionOptical Flow EstimationPredictionregression+1

Multimodal Video Emotion Recognition with Reliable Reasoning Priors

2025-07-29 · Zhepeng Wang, Yingjian Zhu, Guanghao Dong, Hongzhu Yi 외 arxiv

This study investigates the integration of trustworthy prior reasoning knowledge from MLLMs into multimodal emotion recognition. We employ Gemini to generate fine-grained, modality-separable reasoning traces, which are i…

Multimodal Emotion RecognitionVideo Emotion RecognitionContrastive Learning

Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation

2025-08-07 · Haonan Shangguan, Xiaocui Yang, Shi Feng, Daling Wang 외 arxiv

The surge in rich multimodal content on social media platforms has greatly advanced Multimodal Sentiment Analysis (MSA), with Large Language Models (LLMs) further accelerating progress in this field. Current approaches p…

Multimodal Sentiment AnalysisMulti-Task Learning

CORA: Analyzing and bridging thinking-answer gap in Multimodal RLVR via Consistency-Oriented Reasoning Alignment

2026-06-12 · Jiayue Cao, Zhicong Lu, Xuehan Sun, Wei Jia 외 arxiv

Reinforcement learning with verifiable rewards (RLVR) has successfully elicited the reasoning capabilities of large language models, motivating its extension to multimodal scenarios. Existing methods primarily focus on i…

Reinforcement LearningMultimodal Reasoning

Towards Efficient Multimodal Unified Reasoning Model via Model Merging

2025-10-10 · Qixiang Yin, Huanjin Yao, Jianghao Chen, Jiaxing Huang 외 arxiv

Although Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across diverse tasks, they encounter challenges in terms of reasoning efficiency, large model size and overthinking. However, ex…

Reinforcement LearningMathematical ReasoningMultimodal Reasoning