paper-with-me

홈 › Papers

VilNMN: A Neural Module Network approach to Video-Grounded Language Tasks

2021-01-01 · Hung Le, Nancy F. Chen, Steven Hoi

Neural module networks (NMN) have achieved success in image-grounded tasks such as question answering (QA) on synthetic images. However, very limited work on NMN has been studied in the video-grounded language tasks. These tasks extend the complexity of traditional visual tasks with the additional visual temporal variance. Motivated by recent NMN approaches on image-grounded tasks, we introduce Visio-Linguistic Neural Module Network (VilNMN) to model the information retrieval process in video-grounded language tasks as a pipeline of neural modules. VilNMN first decomposes all language components to explicitly resolves entity references and detect corresponding action-based inputs from the question. Detected entities and actions are used as parameters to instantiate neural module networks and extract visual cues from the video. Our experiments show that VilNMN can achieve promising performance on two video-grounded language tasks: video QA and video-grounded dialogues.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalQuestion AnsweringRetrieval

Similar Papers 제목 키워드 기반

VGNMN: Video-grounded Neural Module Network to Video-Grounded Language Tasks

2021-04-16 · Hung Le, Nancy F. Chen, Steven C. H. Hoi

Neural module networks (NMN) have achieved success in image-grounded tasks such as Visual Question Answering (VQA) on synthetic images. However, very limited work on NMN has been studied in the video-grounded dialogue ta…

Information RetrievalQuestion AnsweringRetrievalVisual Question Answering+1

VGNMN: Video-grounded Neural Module Networks for Video-Grounded Dialogue Systems

2022-07-01 · NAACL 2022 7 · Hung Le, Nancy Chen, Steven Hoi

Neural module networks (NMN) have achieved success in image-grounded tasks such as Visual Question Answering (VQA) on synthetic images. However, very limited work on NMN has been studied in the video-grounded dialogue ta…

Information RetrievalQuestion AnsweringRetrievalVisual Question Answering+1

Attend What You Need: Motion-Appearance Synergistic Networks for Video Question Answering

2021-06-19 · ACL 2021 5 · Ahjeong Seo, Gi-Cheon Kang, Joonhan Park, Byoung-Tak Zhang

Video Question Answering is a task which requires an AI agent to answer questions grounded in video. This task entails three key challenges: (1) understand the intention of various questions, (2) capturing various elemen…

AI AgentQuestion AnsweringVideo Question Answering

Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

2024-10-04 · Haibo Wang, Zhiyang Xu, Yu Cheng, Shizhe Diao 외

Video Large Language Models (Video-LLMs) have demonstrated remarkable capabilities in coarse-grained video understanding, however, they struggle with fine-grained temporal grounding. In this paper, we introduce Grounded-…

Dense Video CaptioningSentenceTemporal Sentence GroundingVideo Captioning+1

Harnessing Object Grounding for Time-Sensitive Video Understanding

2025-09-08 · Tz-Ying Wu, Sharath Nittur Sridhar, Subarna Tripathi arxiv

We propose to improve the time-sensitive video understanding (TSV) capability of video large language models (Video-LLMs) with grounded objects (GO). We hypothesize that TSV tasks can benefit from GO within frames, which…

Dense Captioning