paper-with-me

홈 › Papers

AViD Dataset: Anonymized Videos from Diverse Countries

2020-07-10 · NeurIPS 2020 12 · AJ Piergiovanni, Michael S. Ryoo

We introduce a new public video dataset for action recognition: Anonymized Videos from Diverse countries (AViD). Unlike existing public video datasets, AViD is a collection of action videos from many different countries. The motivation is to create a public dataset that would benefit training and pretraining of action recognition models for everybody, rather than making it useful for limited countries. Further, all the face identities in the AViD videos are properly anonymized to protect their privacy. It also is a static dataset where each video is licensed with the creative commons license. We confirm that most of the existing video datasets are statistically biased to only capture action videos from a limited number of countries. We experimentally illustrate that models trained with such biased datasets do not transfer perfectly to action videos from the other countries, and show that AViD addresses such problem. We also confirm that the new AViD dataset could serve as a good dataset for pretraining the models, performing comparably or better than prior datasets.

📄 PDF Abstract BibTeX arXiv:2007.05515

Code (1)

piergiaj/AViD 공식 구현

Tasks

Action ClassificationAction DetectionAction Recognition

Similar Papers 제목 키워드 기반

DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data

2026-04-02 · Wonjoon Jin, Jiyun Won, Janghyeok Han, Qi Dai 외 arxiv

Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A central limitation lies in the scarcity o…

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

2025-06-16 · Zhucun Xue, Jiangning Zhang, Teng Hu, Haoyang He 외

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for video applications sets higher requirements …

4k8kVideo Generation

AVID: Adapting Video Diffusion Models to World Models

2024-10-01 · Marc Rigter, Tarun Gupta, Agrin Hilmkil, Chao Ma

Large-scale generative models have achieved remarkable success in a number of domains. However, for sequential decision-making problems, such as robotics, action-labelled data is often scarce and therefore scaling-up fou…

Decision MakingSequential Decision Making

LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living

2024-06-13 · CVPR 2025 1 · Dominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind 외

Current Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant represent…

BenchmarkingHuman-Object Interaction DetectionRepresentation LearningVideo Description+1

GAViD: A Large-Scale Multimodal Dataset for Context-Aware Group Affect Recognition from Videos

2026-04-17 · Deepak Kumar, Abhishek Pratap Singh, Puneet Kumar, Xiaobai Li 외 arxiv

Understanding affective dynamics in real-world social systems is fundamental to modeling and analyzing human-human interactions in complex environments. Group affect emerges from intertwined human-human interactions, con…