Video Caption Dataset for Describing Human Actions in Japanese
In recent years, automatic video caption generation has attracted considerable attention. This paper focuses on the generation of Japanese captions for describing human actions. While most currently available video caption datasets have been constructed for English, there is no equivalent Japanese dataset. To address this, we constructed a large-scale Japanese video caption dataset consisting of 79,822 videos and 399,233 captions. Each caption in our dataset describes a video in the form of "who does what and where." To describe human actions, it is important to identify the details of a person, place, and action. Indeed, when we describe human actions, we usually mention the scene, person, and action. In our experiments, we evaluated two caption generation methods to obtain benchmark results. Further, we investigated whether those generation methods could specify "who does what and where."
Code (0)
등록된 구현이 없습니다.
Tasks
Caption GenerationSimilar Papers 제목 키워드 기반
Towards Fine-Grained Human Motion Video Captioning
Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motion details, resulting in vague or semanti…
Human Mesh RecoveryVideo CaptioningOSVidCap: A Framework for the Simultaneous Recognition and Description of Concurrent Actions in Videos in an Open-Set Scenario
Automatically understanding and describing the visual content of videos in natural language is a challenging task in computer vision. Existing approaches are often designed to describe single events in a closed-set setti…
DecoderOpen Set Video CaptioningVideo CaptioningHuman Action Adverb Recognition: ADHA Dataset and A Three-Stream Hybrid Model
We introduce the first benchmark for a new problem --- recognizing human action adverbs (HAA): "Adverbs Describing Human Actions" (ADHA). This is the first step for computer vision to change over from pattern recognition…
Action RecognitionImage CaptioningTemporal Action LocalizationVideo ReCap: Recursive Captioning of Hour-Long Videos
Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g., objects, scenes, atomic actions). However, most real-world videos last for…
EgoSchemaVideo CaptioningVideo UnderstandingZero-Shot Video Question AnswerHuman-centric Behavior Description in Videos: New Benchmark and Model
In the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because desc…
Video Captioning