paper-with-me

Papers

Live Video Captioning

2024-06-20 · Eduardo Blanco-Fernández, Carlos Gutiérrez-Álvarez, Nadia Nasri, Saturnino Maldonado-Bascón, Roberto J. López-Sastre

Dense video captioning is the task that involves the detection and description of events within video sequences. While traditional approaches focus on offline solutions where the entire video of analysis is available for the captioning model, in this work we introduce a paradigm shift towards Live Video Captioning (LVC). In LVC, dense video captioning models must generate captions for video streams in an online manner, facing important constraints such as having to work with partial observations of the video, the need for temporal anticipation and, of course, ensuring ideally a real-time response. In this work we formally introduce the novel problem of LVC and propose new evaluation metrics tailored for the online scenario, demonstrating their superiority over traditional metrics. We also propose an LVC model integrating deformable transformers and temporal filtering to address the LVC new challenges. Experimental evaluations on the ActivityNet Captions dataset validate the effectiveness of our approach, highlighting its performance in LVC compared to state-of-the-art offline methods. Results of our model as well as an evaluation kit with the novel metrics integrated are made publicly available to encourage further research on LVC.

📄 PDF Abstract BibTeX arXiv:2406.14206

Code (1)

gramuah/lvc 공식 구현

Tasks

Dense Video CaptioningLive Video CaptioningVideo Captioning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark

2024-01-25 · Zeyu Xi, Ge Shi, Xuefen Li, Junchi Yan 외

Despite the recent emergence of video captioning models, how to generate the text description with specific entity names and fine-grained actions is far from being solved, which however has great applications such as bas…

DecoderVideo Captioning

The Use of Video Captioning for Fostering Physical Activity

2021-04-07 · Soheyla Amirian, Abolfazl Farahani, Hamid R. Arabnia, Khaled Rasheed 외

Video Captioning is considered to be one of the most challenging problems in the field of computer vision. Video Captioning involves the combination of different deep learning models to perform object detection, action d…

Action Detectionobject-detectionObject DetectionVideo Captioning

Response to LiveBot: Generating Live Video Comments Based on Visual and Textual Contexts

2020-06-04 · Hao Wu, Gareth J. F. Jones, Francois Pitie

Live video commenting systems are an emerging feature of online video sites. Recently the Chinese video sharing platform Bilibili, has popularised a novel captioning system where user comments are displayed as streams of…

SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries

2023-04-10 · Hassan Mkhallati, Anthony Cioppa, Silvio Giancola, Bernard Ghanem 외

Soccer is more than just a game - it is a passion that transcends borders and unites people worldwide. From the roar of the crowds to the excitement of the commentators, every moment of a soccer match is a thrill. Yet, w…

Dense Video CaptioningVideo Captioning

ALIVE: Animate Your World with Lifelike Audio-Video Generation

2026-02-09 · Ying Guo, Qijun Gan, Yifu Zhang, Jinlai Liu 외 arxiv

Video generation is rapidly evolving towards unified audio-video generation. In this paper, we present ALIVE, a generation model that adapts a pretrained Text-to-Video (T2V) model to Sora-style audio-video generation and…

Video CaptioningVideo Generation