paper-with-me

홈 › Papers

Event-based Vision for Early Prediction of Manipulation Actions

2023-07-26 · Daniel Deniz, Cornelia Fermuller, Eduardo Ros, Manuel Rodriguez-Alvarez, Francisco Barranco

Neuromorphic visual sensors are artificial retinas that output sequences of asynchronous events when brightness changes occur in the scene. These sensors offer many advantages including very high temporal resolution, no motion blur and smart data compression ideal for real-time processing. In this study, we introduce an event-based dataset on fine-grained manipulation actions and perform an experimental study on the use of transformers for action prediction with events. There is enormous interest in the fields of cognitive robotics and human-robot interaction on understanding and predicting human actions as early as possible. Early prediction allows anticipating complex stages for planning, enabling effective and real-time interaction. Our Transformer network uses events to predict manipulation actions as they occur, using online inference. The model succeeds at predicting actions early on, building up confidence over time and achieving state-of-the-art classification. Moreover, the attention-based transformer architecture allows us to study the role of the spatio-temporal patterns selected by the model. Our experiments show that the Transformer network captures action dynamic features outperforming video-based approaches and succeeding with scenarios where the differences between actions lie in very subtle cues. Finally, we release the new event dataset, which is the first in the literature for manipulation action recognition. Code will be available at https://github.com/DaniDeniz/EventVisionTransformer.

📄 PDF Abstract BibTeX arXiv:2307.14332

Code (1)

danideniz/davishanddataset-events 공식 구현

Tasks

Action RecognitionData CompressionEvent-based visionPrediction

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

Action Prediction in Humans and Robots

2019-07-03 · Florentin Wörgötter, Fatemeh Ziaeetabar, Stefan Pfeiffer, Osman Kaya 외

Efficient action prediction is of central importance for the fluent workflow between humans and equally so for human-robot interaction. To achieve prediction, actions can be encoded by a series of events, where every eve…

PredictionTime SeriesTime Series Analysis

Early Prediction of Sepsis From Clinical Datavia Heterogeneous Event Aggregation

2019-10-14 · Lu-chen Liu, Haoxian Wu, Zichang Wang, Zequn Liu 외

Sepsis is a life-threatening condition that seriously endangers millions of people over the world. Hopefully, with the widespread availability of electronic health records (EHR), predictive models that can effectively de…

Predicting decision-making in the future: Human versus Machine

2021-10-09 · Hoe Sung Ryu, Uijong Ju, Christian Wallraven

Deep neural networks (DNNs) have become remarkably successful in data prediction, and have even been used to predict future actions based on limited input. This raises the question: do these systems actually "understand"…

Decision Making

Neuro-Symbolic Manipulation Understanding with Enriched Semantic Event Chains

2026-04-22 · Fatemeh Ziaeetabar arxiv

Robotic systems operating in human environments must reason about how object interactions evolve over time, which actions are currently being performed, and what manipulation step is likely to follow. Classical enriched …

Action RecognitionDecision Making

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model

2026-06-28 · Jiaxin Liu, Xun Xu, Zhenhao Zhang, Hanqing Wang 외 arxiv

Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable indoor settings, while real-world embodied manipulation may involve …