paper-with-me

Papers

A Large-scale Study of Spatiotemporal Representation Learning with a New Benchmark on Action Recognition

2023-03-23 · ICCV 2023 1 · Andong Deng, Taojiannan Yang, Chen Chen

The goal of building a benchmark (suite of datasets) is to provide a unified protocol for fair evaluation and thus facilitate the evolution of a specific area. Nonetheless, we point out that existing protocols of action recognition could yield partial evaluations due to several limitations. To comprehensively probe the effectiveness of spatiotemporal representation learning, we introduce BEAR, a new BEnchmark on video Action Recognition. BEAR is a collection of 18 video datasets grouped into 5 categories (anomaly, gesture, daily, sports, and instructional), which covers a diverse set of real-world applications. With BEAR, we thoroughly evaluate 6 common spatiotemporal models pre-trained by both supervised and self-supervised learning. We also report transfer performance via standard finetuning, few-shot finetuning, and unsupervised domain adaptation. Our observation suggests that current state-of-the-art cannot solidly guarantee high performance on datasets close to real-world applications, and we hope BEAR can serve as a fair and challenging evaluation benchmark to gain insights on building next-generation spatiotemporal learners. Our dataset, code, and models are released at: https://github.com/AndongDeng/BEAR

📄 PDF Abstract BibTeX arXiv:2303.13505

Code (1)

andongdeng/bear 공식 구현

Tasks

Action RecognitionDomain AdaptationRepresentation LearningSelf-Supervised LearningTemporal Action LocalizationUnsupervised Domain Adaptation

Similar Papers 제목 키워드 기반

A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning

2021-04-29 · CVPR 2021 1 · Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick 외

We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simple objective that can easily generalize …

Representation LearningSelf-Supervised Action RecognitionUnsupervised Pre-training

Learning Transferable Spatiotemporal Representations from Natural Script Knowledge

2022-09-30 · CVPR 2023 1 · Ziyun Zeng, Yuying Ge, Xihui Liu, Bin Chen 외

Pre-training on large-scale video data has become a common recipe for learning transferable spatiotemporal representations in recent years. Despite some progress, existing methods are mostly limited to highly curated dat…

DescriptiveRepresentation LearningVideo Understanding

Local Spatiotemporal Representation Learning for Longitudinally-consistent Neuroimage Analysis

2022-06-09 · Mengwei Ren, Neel Dey, Martin A. Styner, Kelly Botteron 외

Recent self-supervised advances in medical computer vision exploit global and local anatomical self-similarity for pretraining prior to downstream tasks such as segmentation. However, current methods assume i.i.d. image …

One-Shot SegmentationRepresentation LearningSegmentation

Would Mega-scale Datasets Further Enhance Spatiotemporal 3D CNNs?

2020-04-10 · Hirokatsu Kataoka, Tenga Wakamiya, Kensho Hara, Yutaka Satoh

How can we collect and use a video dataset to further improve spatiotemporal 3D Convolutional Neural Networks (3D CNNs)? In order to positively answer this open question in video recognition, we have conducted an explora…

General ClassificationOpen-Ended Question AnsweringVideo ClassificationVideo Recognition

ChronoEarth-492K: A Large Scale and Long Horizon Spatiotemporal Hyperspectral Earth Observation Dataset and Benchmark

2026-05-15 · Haozhe Si, Yuxuan Wan, Yuqing Wang, Minh Do 외 arxiv

Hyperspectral imaging (HSI) provides dense spectral information for the Earth's surface, enabling material-level understanding of land cover and ecosystem dynamics. Despite recent progress in hyperspectral self-supervise…

Self-Supervised LearningRepresentation LearningTemporal Sequences