paper-with-me

홈 › Papers

Scale, Don't Fine-tune: Guiding Multimodal LLMs for Efficient Visual Place Recognition at Test-Time

2025-09-02 · Jintao Cheng, Weibin Li, Jiehao Luo, Xiaoyu Tang, Zhijian He, Jin Wu, Yao Zou, Wei Zhang arxiv

Visual Place Recognition (VPR) has evolved from handcrafted descriptors to deep learning approaches, yet significant challenges remain. Current approaches, including Vision Foundation Models (VFMs) and Multimodal Large Language Models (MLLMs), enhance semantic understanding but suffer from high computational overhead and limited cross-domain transferability when fine-tuned. To address these limitations, we propose a novel zero-shot framework employing Test-Time Scaling (TTS) that leverages MLLMs' vision-language alignment capabilities through Guidance-based methods for direct similarity scoring. Our approach eliminates two-stage processing by employing structured prompts that generate length-controllable JSON outputs. The TTS framework with Uncertainty-Aware Self-Consistency (UASC) enables real-time adaptation without additional training costs, achieving superior generalization across diverse environments. Experimental results demonstrate significant improvements in cross-domain VPR performance with up to 210$\times$ computational efficiency gains.

📄 PDF Abstract BibTeX arXiv:2509.02129

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Place RecognitionComputational Efficiency

Similar Papers 제목 키워드 기반

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation

2025-07-11 · Liu He, Xiao Zeng, Yizhi Song, Albert Y. C. Chen 외 arxiv

Multimodal Large Language Models (MLLMs) struggle with accurately capturing camera-object relations, especially for object orientation, camera viewpoint, and camera shots. This stems from the fact that existing MLLMs are…

Image Generation

MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications

2025-11-17 · Anshul Kumar, Gagan Raj Gupta, Manish Rai, Apu Chakraborty 외 arxiv

Large Language Models (LLMs) have emerged as powerful tools for automating complex reasoning and decision-making tasks. In telecommunications, they hold the potential to transform network optimization, automate troublesh…

Explainable Multimodal Aspect-Based Sentiment Analysis with Dependency-guided Large Language Model

2026-01-11 · Zhongzheng Wang, Yuanhe Tian, Hongzhi Wang, Yan Song arxiv

Multimodal aspect-based sentiment analysis (MABSA) aims to identify aspect-level sentiments by jointly modeling textual and visual information, which is essential for fine-grained opinion understanding in social media. E…

Sentiment Analysis

On the Performance of Multimodal Language Models

2023-10-04 · Utsav Garg, Erhan Bas

Instruction-tuned large language models (LLMs) have demonstrated promising zero-shot generalization capabilities across various downstream tasks. Recent research has introduced multimodal capabilities to LLMs by integrat…

BenchmarkingBinary ClassificationImage CaptioningImage Comprehension+2

Guiding Large Language Models to Post-Edit Machine Translation with Error Annotations

2024-04-11 · Dayeon Ki, Marine Carpuat

Machine Translation (MT) remains one of the last NLP tasks where large language models (LLMs) have not yet replaced dedicated supervised systems. This work exploits the complementary strengths of LLMs and supervised MT b…

Machine TranslationTranslation