paper-with-me

Papers

YouMakeup VQA Challenge: Towards Fine-grained Action Understanding in Domain-Specific Videos

2020-04-12 · Shizhe Chen, Weiying Wang, Ludan Ruan, Linli Yao, Qin Jin

The goal of the YouMakeup VQA Challenge 2020 is to provide a common benchmark for fine-grained action understanding in domain-specific videos e.g. makeup instructional videos. We propose two novel question-answering tasks to evaluate models' fine-grained action understanding abilities. The first task is \textbf{Facial Image Ordering}, which aims to understand visual effects of different actions expressed in natural language to the facial object. The second task is \textbf{Step Ordering}, which aims to measure cross-modal semantic alignments between untrimmed videos and multi-sentence texts. In this paper, we present the challenge guidelines, the dataset used, and performances of baseline models on the two proposed tasks. The baseline codes and models are released at \url{https://github.com/AIM3-RUC/YouMakeup_Baseline}.

📄 PDF Abstract BibTeX arXiv:2004.05573

Code (1)

AIM3-RUC/YouMakeup_Baseline 공식 구현 pytorch

Tasks

Action UnderstandingQuestion AnsweringSentenceVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

YouMakeup: A Large-Scale Domain-Specific Multimodal Dataset for Fine-Grained Semantic Comprehension

2019-11-01 · IJCNLP 2019 11 · Weiying Wang, Yongcheng Wang, Shi-Zhe Chen, Qin Jin

Multimodal semantic comprehension has attracted increasing research interests recently such as visual question answering and caption generation. However, due to the data limitation, fine-grained semantic comprehension ha…

Caption GenerationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Dual-Path Temporal Map Optimization for Make-up Temporal Video Grounding

2023-09-12 · Jiaxiu Li, Kun Li, Jia Li, Guoliang Chen 외

Make-up temporal video grounding (MTVG) aims to localize the target video segment which is semantically related to a sentence describing a make-up activity, given a long video. Compared with the general video grounding t…

Sentencetext similarityVideo Grounding

Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos

2023-03-11 · Teng Wang, Jinrui Zhang, Feng Zheng, Wenhao Jiang 외

Joint video-language learning has received increasing attention in recent years. However, existing works mainly focus on single or multiple trimmed video clips (events), which makes human-annotated event boundaries neces…

Dense Video CaptioningNatural Language Moment RetrievalSentenceText Generation+1

MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding

2026-07-10 · Kun Li, Dan Guo, Jihao Gu, Pengyu Liu 외 arxiv

Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communication. However, their short duration, weak motion patterns, and fine-grain…

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

2025-05-26 · Fanheng Kong, Jingyuan Zhang, Hongzhi Zhang, Shi Feng 외

Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing benchmarks for video understanding often tr…

AttributeVideo Understanding