paper-with-me

홈 › Papers

Ranked from Within: Ranking Large Multimodal Models for Visual Question Answering Without Labels

2024-12-09 · Weijie Tu, Weijian Deng, Dylan Campbell, Yu Yao, Jiyang Zheng, Tom Gedeon, Tongliang Liu

As large multimodal models (LMMs) are increasingly deployed across diverse applications, the need for adaptable, real-world model ranking has become paramount. Traditional evaluation methods are largely dataset-centric, relying on fixed, labeled datasets and supervised metrics, which are resource-intensive and may lack generalizability to novel scenarios, highlighting the importance of unsupervised ranking. In this work, we explore unsupervised model ranking for LMMs by leveraging their uncertainty signals, such as softmax probabilities. We evaluate state-of-the-art LMMs (e.g., LLaVA) across visual question answering benchmarks, analyzing how uncertainty-based metrics can reflect model performance. Our findings show that uncertainty scores derived from softmax distributions provide a robust, consistent basis for ranking models across varied tasks. This finding enables the ranking of LMMs on real-world, unlabeled data for visual question answering, providing a practical approach for selecting models across diverse domains without requiring manual annotation.

📄 PDF Abstract BibTeX arXiv:2412.06461

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Question Answering

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Images Don't Lie: Transferring Deep Visual Semantic Features to Large-Scale Multimodal Learning to Rank

2015-11-20 · Corey Lynch, Kamelia Aryafar, Josh Attenberg

Search is at the heart of modern e-commerce. As a result, the task of ranking search results automatically (learning to rank) is a multibillion dollar machine learning problem. Traditional models optimize over a few hand…

Learning-To-Rank

Comment Ranking Diversification in Forum Discussions

2020-02-27 · Curtis G. Northcutt, Kimberly A. Leon, Naichun Chen

Viewing consumption of discussion forums with hundreds or more comments depends on ranking because most users only view top-ranked comments. When comments are ranked by an ordered score (e.g. number of replies or up-vote…

Re-RankingSemantic SimilaritySemantic Textual Similarity

Mutimodal Ranking Optimization for Heterogeneous Face Re-identification

2022-12-11 · Hui Hu, Jiawei Zhang, Zhen Han

Heterogeneous face re-identification, namely matching heterogeneous faces across disjoint visible light (VIS) and near-infrared (NIR) cameras, has become an important problem in video surveillance application. However, t…

MMCIG: Multimodal Cover Image Generation for Text-only Documents and Its Dataset Construction via Pseudo-labeling

2025-08-24 · Hyeyeon Kim, Sungwoo Han, Jingun Kwon, Hidetaka Kamigaito 외 arxiv

In this study, we introduce a novel cover image generation task that produces both a concise summary and a visually corresponding image from a given text-only document. Because no existing datasets are available for this…

Image Generation

SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval

2025-09-30 · Ren-Di Wu, Yu-Yen Lin, Huei-Fang Yang arxiv

Composed Image Retrieval (CIR) aims to retrieve target images that preserve the visual content of a reference image while incorporating user-specified textual modifications. Training-free zero-shot CIR (ZS-CIR) approache…

Image Retrieval