paper-with-me

Papers Video Description

“Video Description” 태그가 달린 논문 104편 · 필터 해제

DANTE-AD: Dual-Vision Attention Network for Long-Term Audio Description

2025-03-31 · Adrienne Deganutti, Simon Hadfield, Andrew Gilbert

Audio Description is a narrated commentary designed to aid vision-impaired audiences in perceiving key visual elements in a video. While short-form video understanding has advanced rapidly, a solution for maintaining coh…

Video DescriptionVideo UnderstandingVisual Storytelling

HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation

2025-03-31 · CVPR 2025 1 · Kun Liu, Qi Liu, Xinchen Liu, Jie Li 외

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the la…

HallucinationHuman-Object Interaction DetectionVideo DescriptionVideo Generation

Cross-Modal Learning for Music-to-Music-Video Description Generation

2025-03-14 · Zhuoyuan Mao, Mengjie Zhao, Qiyu Wu, Zhi Zhong 외

Music-to-music-video generation is a challenging task due to the intrinsic differences between the music and video modalities. The advent of powerful text-to-video diffusion models has opened a promising pathway for musi…

Video DescriptionVideo Generation

VideoA11y: Method and Dataset for Accessible Video Description

2025-02-27 · Chaoyu Li, Sid Padmanabhuni, Maryam Cheema, Hasti Seifi 외

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality…

Video Description

AVD2: Accident Video Diffusion for Accident Video Description

2025-02-20 · Cheng Li, Keyuan Zhou, Tong Liu, Yu Wang 외

Traffic accidents present complex challenges for autonomous driving, often featuring unpredictable scenarios that hinder accurate system interpretation and responses.Nonetheless, prevailing methodologies fall short in el…

Autonomous DrivingScene UnderstandingVideo DescriptionVideo Understanding

Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis

2025-02-11 · Amir Hosein Fadaei, Mohammad-Reza A. Dehaqani

It's no secret that video has become the primary way we share information online. That's why there's been a surge in demand for algorithms that can analyze and understand video content. It's a trend going to continue as …

Action RecognitionVideo DescriptionVideo Understanding

Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

2025-01-14 · Liping Yuan, Jiawei Wang, Haomiao Sun, Yuchen Zhang 외

We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior general video understanding capabilities. Tars…

Embodied Question AnsweringHallucinationLanguage ModelingLanguage Modelling+5

Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time

2025-01-14 · Mihai Masala, Marius Leordeanu

In the current era of Machine Learning, Transformers have become the de facto approach across a variety of domains, such as computer vision and natural language processing. Transformer-based solutions are the backbone of…

Object RecognitionText GenerationVideo ClassificationVideo Description

Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning

2024-12-17 · Shiping Ge, Qiang Chen, Zhiwei Jiang, Yafeng Yin 외

Weakly-Supervised Dense Video Captioning (WSDVC) aims to localize and describe all events of interest in a video without requiring annotations of event boundaries. This setting poses a great challenge in accurately locat…

Dense Video CaptioningDescriptiveVideo CaptioningVideo Description

StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification

2024-11-11 · Yichen He, Yuan Lin, Jianchao Wu, Hanchong Zhang 외

Existing large vision-language models (LVLMs) are largely limited to processing short, seconds-long videos and struggle with generating coherent descriptions for extended video spanning minutes or more. Long video descri…

Large Language ModelMultimodal Large Language ModelMultiple-choiceVideo Description

PV-VTT: A Privacy-Centric Dataset for Mission-Specific Anomaly Detection and Natural Language Interpretation

2024-10-30 · Ryozo Masukawa, Sanggeon Yun, Yoshiki Yamaguchi, Mohsen Imani

Video crime detection is a significant application of computer vision and artificial intelligence. However, existing datasets primarily focus on detecting severe crimes by analyzing entire video clips, often neglecting t…

Anomaly DetectionDescriptiveGraph Neural NetworkLanguage Modelling+2

FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning

2024-10-20 · Shiyu Hu, Xuchen Li, Xuzhao Li, Jing Zhang 외

Despite rapid progress in large vision-language models (LVLMs), existing video caption benchmarks remain limited in evaluating their alignment with human understanding. Most rely on a single annotation per video and lexi…

DiagnosticVideo CaptioningVideo DescriptionVideo Understanding

VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models

2024-10-01 · Jiapeng Wang, Chengyu Wang, Kunzhe Huang, Jun Huang 외

Contrastive Language-Image Pre-training (CLIP) has been widely studied and applied in numerous applications. However, the emphasis on brief summary texts during pre-training prevents CLIP from understanding long descript…

Hallucinationtext similarityVideo DescriptionVideo Retrieval

Technical Report: Competition Solution For Modelscope-Sora

2024-09-24 · Shengfu Chen, Hailong Liu, Wenzhao Wei

This report presents the approach adopted in the Modelscope-Sora challenge, which focuses on fine-tuning data for video generation models. The challenge evaluates participants' ability to analyze, clean, and generate hig…

Text-to-Video GenerationVideo DescriptionVideo Generation

Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation

2024-08-19 · Liu He, Yizhi Song, Hejun Huang, Pinxin Liu 외

Text-to-video generation has been dominated by diffusion-based or autoregressive models. These novel models provide plausible versatility, but are criticized for improper physical motion, shading and illumination, camera…

Instruction FollowingLarge Language ModelText-to-Video GenerationVideo Description+1

SUSTechGAN: Image Generation for Object Detection in Adverse Conditions of Autonomous Driving

2024-07-18 · Gongjin Lan, Yang Peng, Qi Hao, Chengzhong Xu

Autonomous driving significantly benefits from data-driven deep neural networks. However, the data in autonomous driving typically fits the long-tailed distribution, in which the critical driving data in adverse conditio…

Autonomous DrivingImage Generationobject-detectionObject Detection+2

https://arxiv.org/abs/2407.00634

2024-07-02 · Jiawei Wang, Liping Yuan, Yuchen Zhang

Generating fine-grained video descriptions is a fundamental challenge in video understanding. In this work, we introduce Tarsier, a family of large-scale video-language models designed to generate high-quality video desc…

Video CaptioningVideo DescriptionVideo UnderstandingVisual Question Answering (VQA)

Tarsier: Recipes for Training and Evaluating Large Video Description Models

2024-06-30 · arXiv 2024 7 · Jiawei Wang, Liping Yuan, Yuchen Zhang, Haomiao Sun

Generating fine-grained video descriptions is a fundamental challenge in video understanding. In this work, we introduce Tarsier, a family of large-scale video-language models designed to generate high-quality video desc…

Video CaptioningVideo DescriptionVideo Question AnsweringVideo Understanding+2

LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living

2024-06-13 · CVPR 2025 1 · Dominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind 외

Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant represent…

BenchmarkingHuman-Object Interaction DetectionRepresentation LearningVideo Description+1

A Labelled Dataset for Sentiment Analysis of Videos on YouTube, TikTok, and Other Sources about the 2024 Outbreak of Measles

2024-06-11 · Nirmalya Thakur, Vanessa Su, Mingchen Shao, Kesha A. Patel 외

The work of this paper presents a dataset that contains the data of 4011 videos about the ongoing outbreak of measles published on 264 websites on the internet between January 1, 2024, and May 31, 2024. The dataset is av…

Sentiment AnalysisSubjectivity AnalysisVideo Description
1–20 / 104 다음 →