paper-with-me

홈 › Papers

You Need to Read Again: Multi-granularity Perception Network for Moment Retrieval in Videos

2022-05-25 · Xin Sun, Xuan Wang, Jialin Gao, Qiong Liu, Xi Zhou

Moment retrieval in videos is a challenging task that aims to retrieve the most relevant video moment in an untrimmed video given a sentence description. Previous methods tend to perform self-modal learning and cross-modal interaction in a coarse manner, which neglect fine-grained clues contained in video content, query context, and their alignment. To this end, we propose a novel Multi-Granularity Perception Network (MGPN) that perceives intra-modality and inter-modality information at a multi-granularity level. Specifically, we formulate moment retrieval as a multi-choice reading comprehension task and integrate human reading strategies into our framework. A coarse-grained feature encoder and a co-attention mechanism are utilized to obtain a preliminary perception of intra-modality and inter-modality information. Then a fine-grained feature encoder and a conditioned interaction module are introduced to enhance the initial perception inspired by how humans address reading comprehension problems. Moreover, to alleviate the huge computation burden of some existing methods, we further design an efficient choice comparison module and reduce the hidden size with imperceptible quality loss. Extensive experiments on Charades-STA, TACoS, and ActivityNet Captions datasets demonstrate that our solution outperforms existing state-of-the-art methods. Codes are available at github.com/Huntersxsx/MGPN.

📄 PDF Abstract BibTeX arXiv:2205.12886

Code (1)

Huntersxsx/MGPN 공식 구현 pytorch

Tasks

Moment RetrievalReading ComprehensionRetrievalSentence

Similar Papers 제목 키워드 기반

AI-Generated Image Quality Assessment Based on Task-Specific Prompt and Multi-Granularity Similarity

2024-11-25 · Jili Xia, Lihuo He, Fei Gao, Kaifan Zhang 외

Recently, AI-generated images (AIGIs) created by given prompts (initial prompts) have garnered widespread attention. Nevertheless, due to technical nonproficiency, they often suffer from poor perception quality and Text-…

Image Quality Assessment

Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

2026-07-17 · Wei Feng, Xin Wang, Yu-Wei Zhan, Yuwei Zhou 외 arxiv

Video Large Language Models (Video LLMs) have made significant advancements in various video understanding tasks. However, long-video scenarios remain challenging due to the tension between limited visual token budgets a…

Reinforcement Learning

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring

2026-05-23 · Xinge Peng, Yiting Lu, Xin Li, Zhibo Chen arxiv

We present IQA-Spider, the first image quality assessment (IQA) framework that unifies reasoning, grounding, and referring into a single LMM-based framework for multi-granularity quality understanding. Existing LMM-based…

Image Quality AssessmentQuestion Answering

Document-Level Event Role Filler Extraction using Multi-Granularity Contextualized Encoding

2020-05-13 · ACL 2020 6 · Xinya Du, Claire Cardie

Few works in the literature of event extraction have gone beyond individual sentences to make extraction decisions. This is problematic when the information needed to recognize an event argument is spread across multiple…

Document-level Event ExtractionEvent ExtractionLanguage ModelingLanguage Modelling+1

Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors

2024-01-29 · Shiyin Dong, Mingrui Zhu, Kun Cheng, Nannan Wang 외

The remarkable prowess of diffusion models in image generation has spurred efforts to extend their application beyond generative tasks. However, a persistent challenge exists in lacking a unified approach to apply diffus…

DecoderImage GenerationImage RetrievalOpen Vocabulary Semantic Segmentation+3