paper-with-me

Papers

DVCFlow: Modeling Information Flow Towards Human-like Video Captioning

2021-11-19 · Xu Yan, Zhengcong Fei, Shuhui Wang, Qingming Huang, Qi Tian

Dense video captioning (DVC) aims to generate multi-sentence descriptions to elucidate the multiple events in the video, which is challenging and demands visual consistency, discoursal coherence, and linguistic diversity. Existing methods mainly generate captions from individual video segments, lacking adaptation to the global visual context and progressive alignment between the fast-evolved visual content and textual descriptions, which results in redundant and spliced descriptions. In this paper, we introduce the concept of information flow to model the progressive information changing across video sequence and captions. By designing a Cross-modal Information Flow Alignment mechanism, the visual and textual information flows are captured and aligned, which endows the captioning process with richer context and dynamics on event/topic evolution. Based on the Cross-modal Information Flow Alignment module, we further put forward DVCFlow framework, which consists of a Global-local Visual Encoder to capture both global features and local features for each video segment, and a pre-trained Caption Generator to produce captions. Extensive experiments on the popular ActivityNet Captions and YouCookII datasets demonstrate that our method significantly outperforms competitive baselines, and generates more human-like text according to subject and objective tests.

📄 PDF Abstract BibTeX arXiv:2111.10146

Code (0)

등록된 구현이 없습니다.

Tasks

Dense Video CaptioningDiversitySentenceVideo Captioning

Similar Papers 제목 키워드 기반

Normalizing Flows on the Product Space of SO(3) Manifolds for Probabilistic Human Pose Modeling

2024-04-08 · CVPR 2024 1 · Olaf Dünkel, Tim Salzmann, Florian Pfaff

Normalizing flows have proven their efficacy for density estimation in Euclidean space, but their application to rotational representations, crucial in various domains such as robotics or human pose modeling, remains und…

Density Estimation

Spatial Shortcut Network for Human Pose Estimation

2019-04-05 · Te Qi, Bayram Bayramli, Usman Ali, Qinchuan Zhang 외

Like many computer vision problems, human pose estimation is a challenging problem in that recognizing a body part requires not only information from local area but also from areas with large spatial distance. In order t…

Pose Estimation

The languages of actions, formal grammars and qualitive modeling of companies

2016-08-19 · Vladislav B Kovchegov

In this paper we discuss methods of using the language of actions, formal languages, and grammars for qualitative conceptual linguistic modeling of companies as technological and human institutions. The main problem foll…

Modeling Fairness in Recruitment AI via Information Flow

2025-11-16 · Mattias Brännström, Themis Dimitra Xanthopoulou, Lili Jiang arxiv

Avoiding bias and understanding the real-world consequences of AI-supported decision-making are critical to address fairness and assign accountability. Existing approaches often focus either on technical aspects, such as…

Human mobility is well described by closed-form gravity-like models learned automatically from data

2023-12-18 · Oriol Cabanas-Tirapu, Lluís Danús, Esteban Moro, Marta Sales-Pardo 외

Modeling of human mobility is critical to address questions in urban planning and transportation, as well as global challenges in sustainability, public health, and economic development. However, our understanding and ab…

Form