paper-with-me

홈 › Papers

Multi-Task Domain Adaptation for Language Grounding with 3D Objects

2024-07-03 · Penglei Sun, Yaoxian Song, Xinglin Pan, Peijie Dong, Xiaofei Yang, Qiang Wang, Zhixu Li, Tiefeng Li, Xiaowen Chu

The existing works on object-level language grounding with 3D objects mostly focus on improving performance by utilizing the off-the-shelf pre-trained models to capture features, such as viewpoint selection or geometric priors. However, they have failed to consider exploring the cross-modal representation of language-vision alignment in the cross-domain field. To answer this problem, we propose a novel method called Domain Adaptation for Language Grounding (DA4LG) with 3D objects. Specifically, the proposed DA4LG consists of a visual adapter module with multi-task learning to realize vision-language alignment by comprehensive multimodal feature representation. Experimental results demonstrate that DA4LG competitively performs across visual and non-visual language descriptions, independent of the completeness of observation. DA4LG achieves state-of-the-art performance in the single-view setting and multi-view setting with the accuracy of 83.8% and 86.8% respectively in the language grounding benchmark SNARE. The simulation experiments show the well-practical and generalized performance of DA4LG compared to the existing methods. Our project is available at https://sites.google.com/view/da4lg.

📄 PDF Abstract BibTeX arXiv:2407.02846

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationMulti-Task Learning

Methods 이 논문이 사용한 방법론

Adapter 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding

2025-08-08 · Jian Hu, Zixu Cheng, Shaogang Gong, Isabel Guan 외 arxiv

Video Temporal Grounding (TG) aims to temporally locate video segments matching a natural language description (a query) in a long video. While Vision-Language Models (VLMs) are effective at holistic semantic matching, t…

Reinforcement Learning

Multi-modal Domain Adaptation for REG via Relation Transfer

2023-09-23 · Yifan Ding, Liqiang Wang, Boqing Gong

Domain adaptation, which aims to transfer knowledge between domains, has been well studied in many areas such as image classification and object detection. However, for multi-modal tasks, conventional approaches rely on …

Domain Adaptationimage-classificationImage Classificationobject-detection+3

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

2025-07-01 · Bob Zhang, Haoran Li, Tao Zhang, Jianan Li 외 arxiv

Multimodal Large Language Models (MLLMs) perform well in single-image visual grounding but struggle with real-world tasks that demand cross-image reasoning and multi-modal instructions. To address this, we adopt a reinfo…

Reinforcement LearningVisual Grounding

ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding

2024-08-13 · Yubin Wang, Xinyang Jiang, De Cheng, Dongsheng Li 외

Video temporal grounding is an emerging topic aiming to identify specific clips within videos. In addition to pre-trained video models, contemporary methods utilize pre-trained vision-language models (VLM) to capture det…

Prompt Learning

Sim-To-Real Transfer of Visual Grounding for Human-Aided Ambiguity Resolution

2022-05-24 · Georgios Tziafas, Hamidreza Kasaei

Service robots should be able to interact naturally with non-expert human users, not only to help them in various tasks but also to receive guidance in order to resolve ambiguities that might be present in the instructio…

Domain AdaptationVisual Grounding