paper-with-me

Papers

RareAct: A video dataset of unusual interactions

2020-08-03 · Antoine Miech, Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, Andrew Zisserman

This paper introduces a manually annotated video dataset of unusual actions, namely RareAct, including actions such as "blend phone", "cut keyboard" and "microwave shoes". RareAct aims at evaluating the zero-shot and few-shot compositionality of action recognition models for unlikely compositions of common action verbs and object nouns. It contains 122 different actions which were obtained by combining verbs and nouns rarely co-occurring together in the large-scale textual corpus from HowTo100M, but that frequently appear separately. We provide benchmarks using a state-of-the-art HowTo100M pretrained video and text model and show that zero-shot and few-shot compositionality of actions remains a challenging and unsolved task.

📄 PDF Abstract BibTeX arXiv:2008.01018

Code (1)

antoine77340/RareAct 공식 구현

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

What is usual in unusual videos? Trajectory snippet histograms for discovering unusualness

2014-01-03 · Ahmet Iscen, Anil Armagan, Pinar Duygulu

Unusual events are important as being possible indicators of undesired consequences. Moreover, unusualness in everyday life activities may also be amusing to watch as proven by the popularity of such videos shared in soc…

UAL-Bench: The First Comprehensive Unusual Activity Localization Benchmark

2024-10-02 · Hasnat Md Abdullah, Tian Liu, Kangda Wei, Shu Kong 외

Localizing unusual activities, such as human errors or surveillance incidents, in videos holds practical significance. However, current video understanding models struggle with localizing these unusual events likely beca…

Unusual Activity LocalizationVideo Understanding

Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition

2026-06-21 · Prajwal Gatti, Simon Jenni, Fabian Caba Heilbron, Dima Damen arxiv

We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video generative model, conditioned on diverse text prompts grounded in action p…

Action Recognition

Reformulating Zero-shot Action Recognition for Multi-label Actions

2021-12-01 · NeurIPS 2021 12 · Alec Kerrigan, Kevin Duarte, Yogesh Rawat, Mubarak Shah

The goal of zero-shot action recognition (ZSAR) is to classify action classes which were not previously seen during training. Traditionally, this is achieved by training a network to map, or regress, visual inputs to a s…

Action ClassificationAction DetectionAction RecognitionActivity Detection+1

MoCha:End-to-End Video Character Replacement without Structural Guidance

2026-01-13 · Zhengbo Xu, Jie Ma, Ziheng Wang, Zhan Peng 외 arxiv

Controllable video character replacement with a user-provided identity remains a challenging problem due to the lack of paired video data. Prior works have predominantly relied on a reconstruction-based paradigm that req…