paper-with-me

Papers Action Detection

“Action Detection” 태그가 달린 논문 850편 · 필터 해제

TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

2026-08-18 · Yearang Lee, Ho-Joong Kim, Seong-Whan Lee arxiv

Zero-Shot Temporal Action Detection (ZSTAD) aims to lo- calize and recognize action instances from unseen action categories in untrimmed videos. Although existing meth- ods have shown effectiveness by advancing architect…

Action DetectionVideo Alignment

TubeLite: Lightweight Multi-Actor Spatio-Temporal Action Detection

2026-07-06 · Ali Soltaninezhad, Melissa Cote, Alejandro Rico Espinosa, Tunai Porto Marques 외 arxiv

Spatio-temporal action detection in videos requires jointly localizing actors in space and identifying action boundaries over time. A common challenge is constructing temporally stable action tubes, as frame-level detect…

Action Detection

SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos

2026-06-30 · Björn Braun, Christian Holz arxiv

To enable personalized, real-time coaching using Augmented Reality glasses or fixed camera setups in domains such as sports, cooking, or music, a system must understand not just what a person does, but how well they exec…

Action Detection

A New Multi-Domain Benchmark for Micro-Action Recognition and Detection

2026-06-12 · Yanbin Hao, Pengyu Liu, Xing Wei, Xun Yang 외 arxiv

Micro-actions are short-duration, low-amplitude subtle body movements at the whole-body level that can reveal latent intentions, involuntary reactions, and fine-grained affective changes. Our previous MA-52 benchmark has…

Micro-Action RecognitionAction UnderstandingEmotion RecognitionAction Detection

Motion Reinforces Appearance: RGB-Skeleton Gated Residual Fusion for Micro-Gesture Online Recognition

2026-06-10 · Jialin Liu, Xinwen He, Pengyu Liu, Jiale Shi 외 arxiv

Micro-gesture analysis attracts increasing attention for inferring spontaneous emotion from subtle body movements. Micro-gesture online recognition, which localizes and classifies each gesture instance in untrimmed video…

Action Detection

SpikeTAD: Spiking Neural Networks for End-to-End Temporal Action Detection

2026-06-10 · Min Yang, Mi Zhou, Limin Wang arxiv

Video understanding is a crucial part of computer vision, with numerous application scenarios. With the increasing popularity of mobile devices, an increasing number of efforts are trying to deploy video understanding mo…

Action Detection

EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026

2026-05-23 · Zhiheng Fu, Zixu Li, Zhiwei Chen, Fangxu Liu 외 arxiv

The EPIC-KITCHENS-100 Action Detection challenge evaluates whether a model can localize the start and end of each action in long untrimmed egocentric videos and assign the corresponding verb--noun action label. In this r…

Action Detection

Improving Viewpoint-Invariance and Temporal Consistency for Action Detection

2026-05-21 · Yannick Porto, Renato Martins, Thomas Chalumeau, Cedric Demonceaux arxiv

Viewpoint change invariance and action temporal consistency are critical aspects for the effective deployment of human action detection of untrimmed videos. Existing appearance-based video detection methods often struggl…

Action Detection

EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms

2026-04-23 · Brian VanVoorst, Nicholas Walczak, Christopher Gilleo, Charles Meissner 외 arxiv

This paper introduces EgoMAGIC (Medical Assistance, Guidance, Instruction, and Correction), an egocentric medical activity dataset collected as part of DARPA's Perceptually-enabled Task Guidance (PTG) program. This datas…

Action RecognitionAction Detection

Explicit Dropout: Deterministic Regularization for Transformer Architectures

2026-04-22 · Vidhi Agrawal, Illia Oleksiienko, Alexandros Iosifidis arxiv

Dropout is a widely used regularization technique in deep learning, but its effects are typically realized through stochastic masking rather than explicit optimization objectives. We propose a deterministic formulation t…

Audio ClassificationImage ClassificationAction Detection

LiquidTAD: Efficient Temporal Action Detection via Parallel Liquid-Inspired Temporal Relaxation

2026-04-20 · Zepeng Sun, Naichuan Zheng, Hailun Xia, Junjie Wu 외 arxiv

Temporal Action Detection (TAD) requires precise localization of action boundaries within long, untrimmed video sequences. While current high-performing methods achieve strong accuracy, they are often characterized by ex…

Action Detection

Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection

2026-04-20 · Sa Zhu, Wanqian Zhang, Lin Wang, Jinchao Zhang 외 arxiv

Open-Vocabulary Temporal Action Detection (OV-TAD) aims to localize and classify action segments of unseen categories in untrimmed videos, where effective alignment between action semantics and video representations is c…

Action DetectionVideo Alignment

Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection

2026-04-10 · Yicheng Qiu, Keiji Yanai arxiv

Temporal human action detection aims to identify and localize action segments within untrimmed videos, serving as a pivotal task in video understanding. Despite the progress achieved by prior architectures like CNN and T…

Action Detection

Single-agent vs. Multi-agents for Automated Video Analysis of On-Screen Collaborative Learning Behaviors

2026-04-04 · Likai Peng, Shihui Feng arxiv

On-screen learning behavior provides valuable insights into how students seek, use, and create information during learning. Analyzing on-screen behavioral engagement is essential for capturing students' cognitive and col…

Action Detection

From Skeletons to Semantics: Design and Deployment of a Hybrid Edge-Based Action Detection System for Public Safety

2026-03-31 · Ganen Sethupathy, Lalit Dumka, Jan Schagen arxiv

Public spaces such as transport hubs, city centres, and event venues require timely and reliable detection of potentially violent behaviour to support public safety. While automated video analysis has made significant pr…

Action Detection

Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection

2026-03-25 · Sa Zhu, Wanqian Zhang, Lin Wang, Xiaohua Chen 외 arxiv

Open-Vocabulary Temporal Action Detection (OV-TAD) aims to classify and localize action segments in untrimmed videos for unseen categories. Previous methods rely solely on global alignment between label-level semantics a…

Action Detection

SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion

2026-03-13 · Xiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu 외 arxiv

Error detection is crucial in industrial training, healthcare, and assembly quality control. Most existing work assumes a single-view setting and cannot handle the practical case where a third-person (exo) demonstration …

Dense Video CaptioningAction Detection

HomeSafe-Bench: Evaluating Vision-Language Models on Unsafe Action Detection for Embodied Agents in Household Scenarios

2026-03-12 · Jiayue Pu, Zhongxiang Sun, Zilu Zhang, Xiao Zhang 외 arxiv

The rapid evolution of embodied agents has accelerated the deployment of household robots in real-world environments. However, unlike structured industrial settings, household spaces introduce unpredictable safety risks,…

Multimodal ReasoningAction DetectionVideo Generation

When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use Agents

2026-02-09 · Yuting Ning, Jaylen Jones, Zhehao Zhang, Chentao Ye 외 arxiv

Computer-use agents (CUAs) have made tremendous progress in the past year, yet they still frequently produce misaligned actions that deviate from the user's original intent. Such misaligned actions may arise from externa…

Action Detection

Towards Adaptive Fusion of Multimodal Deep Networks for Human Action Recognition

2025-12-04 · Novanto Yudistira arxiv

This study introduces a pioneering methodology for human action recognition by harnessing deep neural network techniques and adaptive fusion strategies across multiple modalities, including RGB, optical flows, audio, and…

Self-Supervised LearningAction RecognitionAction Detection
1–20 / 850 다음 →