paper-with-me

Papers

Video Caption Dataset for Describing Human Actions in Japanese

2020-03-10 · LREC 2020 5 · Yutaro Shigeto, Yuya Yoshikawa, Jiaqing Lin, Akikazu Takeuchi

In recent years, automatic video caption generation has attracted considerable attention. This paper focuses on the generation of Japanese captions for describing human actions. While most currently available video caption datasets have been constructed for English, there is no equivalent Japanese dataset. To address this, we constructed a large-scale Japanese video caption dataset consisting of 79,822 videos and 399,233 captions. Each caption in our dataset describes a video in the form of "who does what and where." To describe human actions, it is important to identify the details of a person, place, and action. Indeed, when we describe human actions, we usually mention the scene, person, and action. In our experiments, we evaluated two caption generation methods to obtain benchmark results. Further, we investigated whether those generation methods could specify "who does what and where."

📄 PDF Abstract BibTeX arXiv:2003.04865

Code (0)

등록된 구현이 없습니다.

Tasks

Caption Generation

Similar Papers 제목 키워드 기반

Towards Fine-Grained Human Motion Video Captioning

2025-10-24 · Guorui Song, Guocun Wang, Zhe Huang, Jing Lin 외 arxiv

Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motion details, resulting in vague or semanti…

Human Mesh RecoveryVideo Captioning

OSVidCap: A Framework for the Simultaneous Recognition and Description of Concurrent Actions in Videos in an Open-Set Scenario

2021-09-29 · IEEE Access 2021 9 · Andrei De Souza Inácio, Matheus Gutoski, André Eugênio Lazzaretti, Heitor Silvério Lopes

Automatically understanding and describing the visual content of videos in natural language is a challenging task in computer vision. Existing approaches are often designed to describe single events in a closed-set setti…

DecoderOpen Set Video CaptioningVideo Captioning

Human Action Adverb Recognition: ADHA Dataset and A Three-Stream Hybrid Model

2018-02-04 · Bo Pang, Kaiwen Zha, Cewu Lu

We introduce the first benchmark for a new problem --- recognizing human action adverbs (HAA): "Adverbs Describing Human Actions" (ADHA). This is the first step for computer vision to change over from pattern recognition…

Action RecognitionImage CaptioningTemporal Action Localization

Video ReCap: Recursive Captioning of Hour-Long Videos

2024-02-20 · CVPR 2024 1 · Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan 외

Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g., objects, scenes, atomic actions). However, most real-world videos last for…

EgoSchemaVideo CaptioningVideo UnderstandingZero-Shot Video Question Answer

Human-centric Behavior Description in Videos: New Benchmark and Model

2023-10-04 · Lingru Zhou, Yiqi Gao, Manqing Zhang, Peng Wu 외

In the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because desc…

Video Captioning