Papers Video Description
“Video Description” 태그가 달린 논문 104편 · 필터 해제
DANTE-AD: Dual-Vision Attention Network for Long-Term Audio Description
Audio Description is a narrated commentary designed to aid vision-impaired audiences in perceiving key visual elements in a video. While short-form video understanding has advanced rapidly, a solution for maintaining coh…
Video DescriptionVideo UnderstandingVisual StorytellingHOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation
Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the la…
HallucinationHuman-Object Interaction DetectionVideo DescriptionVideo GenerationCross-Modal Learning for Music-to-Music-Video Description Generation
Music-to-music-video generation is a challenging task due to the intrinsic differences between the music and video modalities. The advent of powerful text-to-video diffusion models has opened a promising pathway for musi…
Video DescriptionVideo GenerationVideoA11y: Method and Dataset for Accessible Video Description
Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality…
Video DescriptionAVD2: Accident Video Diffusion for Accident Video Description
Traffic accidents present complex challenges for autonomous driving, often featuring unpredictable scenarios that hinder accurate system interpretation and responses.Nonetheless, prevailing methodologies fall short in el…
Autonomous DrivingScene UnderstandingVideo DescriptionVideo UnderstandingEnhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis
It's no secret that video has become the primary way we share information online. That's why there's been a surge in demand for algorithms that can analyze and understand video content. It's a trend going to continue as …
Action RecognitionVideo DescriptionVideo UnderstandingTarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding
We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior general video understanding capabilities. Tars…
Embodied Question AnsweringHallucinationLanguage ModelingLanguage Modelling+5Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
In the current era of Machine Learning, Transformers have become the de facto approach across a variety of domains, such as computer vision and natural language processing. Transformer-based solutions are the backbone of…
Object RecognitionText GenerationVideo ClassificationVideo DescriptionImplicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
Weakly-Supervised Dense Video Captioning (WSDVC) aims to localize and describe all events of interest in a video without requiring annotations of event boundaries. This setting poses a great challenge in accurately locat…
Dense Video CaptioningDescriptiveVideo CaptioningVideo DescriptionStoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification
Existing large vision-language models (LVLMs) are largely limited to processing short, seconds-long videos and struggle with generating coherent descriptions for extended video spanning minutes or more. Long video descri…
Large Language ModelMultimodal Large Language ModelMultiple-choiceVideo DescriptionPV-VTT: A Privacy-Centric Dataset for Mission-Specific Anomaly Detection and Natural Language Interpretation
Video crime detection is a significant application of computer vision and artificial intelligence. However, existing datasets primarily focus on detecting severe crimes by analyzing entire video clips, often neglecting t…
Anomaly DetectionDescriptiveGraph Neural NetworkLanguage Modelling+2FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
Despite rapid progress in large vision-language models (LVLMs), existing video caption benchmarks remain limited in evaluating their alignment with human understanding. Most rely on a single annotation per video and lexi…
DiagnosticVideo CaptioningVideo DescriptionVideo UnderstandingVideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
Contrastive Language-Image Pre-training (CLIP) has been widely studied and applied in numerous applications. However, the emphasis on brief summary texts during pre-training prevents CLIP from understanding long descript…
Hallucinationtext similarityVideo DescriptionVideo RetrievalTechnical Report: Competition Solution For Modelscope-Sora
This report presents the approach adopted in the Modelscope-Sora challenge, which focuses on fine-tuning data for video generation models. The challenge evaluates participants' ability to analyze, clean, and generate hig…
Text-to-Video GenerationVideo DescriptionVideo GenerationKubrick: Multimodal Agent Collaborations for Synthetic Video Generation
Text-to-video generation has been dominated by diffusion-based or autoregressive models. These novel models provide plausible versatility, but are criticized for improper physical motion, shading and illumination, camera…
Instruction FollowingLarge Language ModelText-to-Video GenerationVideo Description+1SUSTechGAN: Image Generation for Object Detection in Adverse Conditions of Autonomous Driving
Autonomous driving significantly benefits from data-driven deep neural networks. However, the data in autonomous driving typically fits the long-tailed distribution, in which the critical driving data in adverse conditio…
Autonomous DrivingImage Generationobject-detectionObject Detection+2https://arxiv.org/abs/2407.00634
Generating fine-grained video descriptions is a fundamental challenge in video understanding. In this work, we introduce Tarsier, a family of large-scale video-language models designed to generate high-quality video desc…
Video CaptioningVideo DescriptionVideo UnderstandingVisual Question Answering (VQA)Tarsier: Recipes for Training and Evaluating Large Video Description Models
Generating fine-grained video descriptions is a fundamental challenge in video understanding. In this work, we introduce Tarsier, a family of large-scale video-language models designed to generate high-quality video desc…
Video CaptioningVideo DescriptionVideo Question AnsweringVideo Understanding+2LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant represent…
BenchmarkingHuman-Object Interaction DetectionRepresentation LearningVideo Description+1A Labelled Dataset for Sentiment Analysis of Videos on YouTube, TikTok, and Other Sources about the 2024 Outbreak of Measles
The work of this paper presents a dataset that contains the data of 4011 videos about the ongoing outbreak of measles published on 264 websites on the internet between January 1, 2024, and May 31, 2024. The dataset is av…
Sentiment AnalysisSubjectivity AnalysisVideo Description