Oops! Predicting Unintentional Action in Video
From just a short glance at a video, we can often tell whether a person's action is intentional or not. Can we train a model to recognize this? We introduce a dataset of in-the-wild videos of unintentional action, as well as a suite of tasks for recognizing, localizing, and anticipating its onset. We train a supervised neural network as a baseline and analyze its performance compared to human consistency on the tasks. We also investigate self-supervised representations that leverage natural signals in our dataset, and show the effectiveness of an approach that uses the intrinsic speed of video to perform competitively with highly-supervised pretraining. However, a significant gap between machine and human performance remains. The project website is available at https://oops.cs.columbia.edu
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Tragedy Plus Time: Capturing Unintended Human Activities from Weakly-labeled Videos
In videos that contain actions performed unintentionally, agents do not achieve their desired goals. In such videos, it is challenging for computer vision systems to understand high-level concepts such as goal-directed b…
Action UnderstandingVideo CaptioningPLSM: A Parallelized Liquid State Machine for Unintentional Action Detection
Reservoir Computing (RC) offers a viable option to deploy AI algorithms on low-end embedded system platforms. Liquid State Machine (LSM) is a bio-inspired RC model that mimics the cortical microcircuits and uses spiking …
Action DetectionGPUNavigating Hallucinations for Reasoning of Unintentional Activities
In this work we present a novel task of understanding unintentional human activities in videos. We formalize this problem as a reasoning task under zero-shot scenario, where given a video of an unintentional activity we …
HallucinationNavigateLearning Goals from Failure
We introduce a framework that predicts the goals behind observable human action in video. Motivated by evidence in developmental psychology, we leverage video of unintentional action to learn video representations of goa…
Representation LearningSelf-supervised Learning for Unintentional Action Prediction
Distinguishing if an action is performed as intended or if an intended action fails is an important skill that not only humans have, but that is also important for intelligent systems that operate in human environments. …
Action ClassificationPredictionRepresentation LearningSelf-Supervised Learning