paper-with-me

홈 › Papers

VGNMN: Video-grounded Neural Module Networks for Video-Grounded Dialogue Systems

2022-07-01 · NAACL 2022 7 · Hung Le, Nancy Chen, Steven Hoi

Neural module networks (NMN) have achieved success in image-grounded tasks such as Visual Question Answering (VQA) on synthetic images. However, very limited work on NMN has been studied in the video-grounded dialogue tasks. These tasks extend the complexity of traditional visual tasks with the additional visual temporal variance and language cross-turn dependencies. Motivated by recent NMN approaches on image-grounded tasks, we introduce Video-grounded Neural Module Network (VGNMN) to model the information retrieval process in video-grounded language tasks as a pipeline of neural modules. VGNMN first decomposes all language components in dialogues to explicitly resolve any entity references and detect corresponding action-based inputs from the question. The detected entities and actions are used as parameters to instantiate neural module networks and extract visual cues from the video. Our experiments show that VGNMN can achieve promising performance on a challenging video-grounded dialogue benchmark as well as a video QA benchmark.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalQuestion AnsweringRetrievalVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

VGNMN: Video-grounded Neural Module Network to Video-Grounded Language Tasks

2021-04-16 · Hung Le, Nancy F. Chen, Steven C. H. Hoi

Neural module networks (NMN) have achieved success in image-grounded tasks such as Visual Question Answering (VQA) on synthetic images. However, very limited work on NMN has been studied in the video-grounded dialogue ta…

Information RetrievalQuestion AnsweringRetrievalVisual Question Answering+1

VilNMN: A Neural Module Network approach to Video-Grounded Language Tasks

2021-01-01 · Hung Le, Nancy F. Chen, Steven Hoi

Neural module networks (NMN) have achieved success in image-grounded tasks such as question answering (QA) on synthetic images. However, very limited work on NMN has been studied in the video-grounded language tasks. The…

Information RetrievalQuestion AnsweringRetrieval

Action Images: End-to-End Policy Learning via Multiview Video Generation

2026-04-07 · Haoyu Zhen, Zixian Gao, Qiao Sun, Yilin Zhao 외 arxiv

World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the future states. However, existing approaches often rely on separate ac…

Video Generation

Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars

2026-02-02 · Youliang Zhang, Zhengguang Zhou, Zhentao Yu, Ziyao Huang 외 arxiv

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (G…

Video Generation

Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

2024-10-04 · Haibo Wang, Zhiyang Xu, Yu Cheng, Shizhe Diao 외

Video Large Language Models (Video-LLMs) have demonstrated remarkable capabilities in coarse-grained video understanding, however, they struggle with fine-grained temporal grounding. In this paper, we introduce Grounded-…

Dense Video CaptioningSentenceTemporal Sentence GroundingVideo Captioning+1