YouMakeup VQA Challenge: Towards Fine-grained Action Understanding in Domain-Specific Videos
The goal of the YouMakeup VQA Challenge 2020 is to provide a common benchmark for fine-grained action understanding in domain-specific videos e.g. makeup instructional videos. We propose two novel question-answering tasks to evaluate models' fine-grained action understanding abilities. The first task is \textbf{Facial Image Ordering}, which aims to understand visual effects of different actions expressed in natural language to the facial object. The second task is \textbf{Step Ordering}, which aims to measure cross-modal semantic alignments between untrimmed videos and multi-sentence texts. In this paper, we present the challenge guidelines, the dataset used, and performances of baseline models on the two proposed tasks. The baseline codes and models are released at \url{https://github.com/AIM3-RUC/YouMakeup_Baseline}.
Code (1)
Tasks
Action UnderstandingQuestion AnsweringSentenceVisual Question Answering (VQA)Similar Papers 제목 키워드 기반
YouMakeup: A Large-Scale Domain-Specific Multimodal Dataset for Fine-Grained Semantic Comprehension
Multimodal semantic comprehension has attracted increasing research interests recently such as visual question answering and caption generation. However, due to the data limitation, fine-grained semantic comprehension ha…
Caption GenerationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Dual-Path Temporal Map Optimization for Make-up Temporal Video Grounding
Make-up temporal video grounding (MTVG) aims to localize the target video segment which is semantically related to a sentence describing a make-up activity, given a long video. Compared with the general video grounding t…
Sentencetext similarityVideo GroundingLearning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos
Joint video-language learning has received increasing attention in recent years. However, existing works mainly focus on single or multiple trimmed video clips (events), which makes human-annotated event boundaries neces…
Dense Video CaptioningNatural Language Moment RetrievalSentenceText Generation+1MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding
Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communication. However, their short duration, weak motion patterns, and fine-grain…
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing benchmarks for video understanding often tr…
AttributeVideo Understanding