paper-with-me

홈 › Papers

See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model

2025-09-19 · Pengteng Li, Pinhao Song, Wuyang Li, Weiyu Guo, Huizai Yao, Yijie Xu, Dugang Liu, Hui Xiong arxiv

We introduce SEE&TREK, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMS) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely visualspatial understanding remains underexplored. SEE&TREK addresses this gap by focusing on two core principles: increasing visual diversity and motion reconstruction. For visual diversity, we conduct Maximum Semantic Richness Sampling, which employs an off-the-shell perception model to extract semantically rich keyframes that capture scene structure. For motion reconstruction, we simulate visual trajectories and encode relative spatial positions into keyframes to preserve both spatial relations and temporal coherence. Our method is training&GPU-free, requiring only a single forward pass, and can be seamlessly integrated into existing MLLM'S. Extensive experiments on the VSI-B ENCH and STI-B ENCH show that S EE &T REK consistently boosts various MLLM S performance across diverse spatial reasoning tasks with the most +3.5% improvement, offering a promising path toward stronger spatial intelligence.

📄 PDF Abstract BibTeX arXiv:2509.16087

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial ReasoningPoint Clouds

Similar Papers 제목 키워드 기반

ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way

2025-07-11 · Rajarshi Roy, Devleena Das, Ankesh Banerjee, Arjya Bhattacharjee 외

We introduce ByDeWay, a training-free framework designed to enhance the performance of Multimodal Large Language Models (MLLMs). ByDeWay uses a novel prompting strategy called Layered-Depth-Based Prompting (LDP), which i…

Depth EstimationHallucinationLanguage ModelingLanguage Modelling+2

Federated Personal Knowledge Graph Completion with Lightweight Large Language Models for Personalized Recommendations

2026-02-25 · Fernando Spadea, Oshani Seneviratne arxiv

Personalized recommendation increasingly relies on private user data, motivating approaches that can adapt to individuals without centralizing their information. We present Federated Targeted Recommendations with Evolvin…

Knowledge Graph CompletionFederated LearningKnowledge Graphs

ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting

2025-04-29 · Yu Zhang, Wenxiang Guo, Changhao Pan, Zhiyuan Zhu 외

Multimodal immersive spatial drama generation focuses on creating continuous multi-speaker binaural speech with dramatic prosody based on multimodal prompts, with potential applications in AR, VR, and others. This task r…

Contrastive LearningMamba

Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting

2026-03-02 · Giacomo Frisoni, Lorenzo Molfetta, Mattia Buzzoni, Gianluca Moro arxiv

Recent advances in training-free visual prompting, such as Set-of-Mark, have emerged as a promising direction for enhancing the grounding capabilities of multimodal language models (MLMs). These techniques operate by par…

Visual Question AnsweringSpatial Reasoning

Enterprise to Computer: Star Trek chatbot

2017-08-02 · Grishma Jena, Mansi Vashisht, Abheek Basu, Lyle Ungar 외

Human interactions and human-computer interactions are strongly influenced by style as well as content. Adding a persona to a chatbot makes it more human-like and contributes to a better and more engaging user experience…

ChatbotDecoder