paper-with-me

홈 › Papers

HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding

2026-03-02 · Lei Yao, Yong Chen, Yuejiao Su, Yi Wang, Moyun Liu, Lap-Pui Chau arxiv

Humans commonly identify 3D object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. Inspired by this principle, we advocate for a novel framework that leverages emerging multimodal large language models (MLLMs) for interaction intention-driven 3D affordance grounding, namely HAMMER. Instead of generating explicit object attribute descriptions or relying on off-the-shelf 2D segmenters, we alternatively aggregate the interaction intention depicted in the image into a contact-aware embedding and guide the model to infer textual affordance labels, ensuring it thoroughly excavates object semantics and contextual cues. We further devise a hierarchical cross-modal integration mechanism to fully exploit the complementary information from the MLLM for 3D representation refinement and introduce a multi-granular geometry lifting module that infuses spatial characteristics into the extracted intention embedding, thus facilitating accurate 3D affordance localization. Extensive experiments on public datasets and our newly constructed corrupted benchmark demonstrate the superiority and robustness of HAMMER compared to existing approaches. All code and weights are publicly available.

📄 PDF Abstract BibTeX arXiv:2603.02329

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Hammer: Robust Function-Calling for On-Device Language Models via Function Masking

2024-10-06 · Qiqiang Lin, Muning Wen, Qiuying Peng, Guanyu Nie 외

Large language models have demonstrated impressive value in performing as autonomous agents when equipped with external tools and API calls. Nonetheless, effectively harnessing their potential for executing complex tasks…

The Future of MLLM Prompting is Adaptive: A Comprehensive Experimental Evaluation of Prompt Engineering Methods for Robust Multimodal Performance

2025-04-14 · Anwesha Mohanty, Venkatesh Balavadhani Parthasarathy, Arsalan Shahid

Multimodal Large Language Models (MLLMs) are set to transform how machines process and generate human-like responses by integrating diverse modalities such as text, images, and code. Yet, effectively harnessing their cap…

Code GenerationHallucinationPrompt EngineeringRetrieval

Vela: Scalable Embeddings with Voice Large Language Models for Multimodal Retrieval

2025-06-17 · Ruofan Hu, Yan Xia, Minjie Hong, Jieming Zhu 외

Multimodal large language models (MLLMs) have seen substantial progress in recent years. However, their ability to represent multimodal information in the acoustic domain remains underexplored. In this work, we introduce…

In-Context LearningRetrieval

Can Multimodal Large Language Model Think Analogically?

2024-11-02 · Diandian Guo, Cong Cao, Fangfang Yuan, Dakui Wang 외

Analogical reasoning, particularly in multimodal contexts, is the foundation of human perception and creativity. Multimodal Large Language Model (MLLM) has recently sparked considerable discussion due to its emergent cap…

Language ModelingLanguage ModellingLarge Language Modelmodel+1

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce

2026-04-22 · Biao Zhang, Lixin Chen, Bin Zhang, Zongwei Wang 외 arxiv

Multimodal representation is crucial for E-commerce tasks such as identical product retrieval. Large representation models (e.g., VLM2Vec) demonstrate strong multimodal understanding capabilities, yet they struggle with …

Representation LearningContrastive Learning