paper-with-me

Papers

Unified Framework with Consistency across Modalities for Human Activity Recognition

2024-09-04 · Tuyen Tran, Thao Minh Le, Hung Tran, Truyen Tran

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their ability to exploit the complementary advantages across modalities. Recent studies focus on combining these two modalities using simple feature fusion techniques. However, due to the inherent disparities in representation between these input modalities, designing a unified neural network architecture to effectively leverage their complementary information remains a significant challenge. To address this, we propose a comprehensive multimodal framework for robust video-based human activity recognition. Our key contribution is the introduction of a novel compositional query machine, called COMPUTER ($\textbf{COMP}ositional h\textbf{U}man-cen\textbf{T}ric qu\textbf{ER}y$ machine), a generic neural architecture that models the interactions between a human of interest and its surroundings in both space and time. Thanks to its versatile design, COMPUTER can be leveraged to distill distinctive representations for various input modalities. Additionally, we introduce a consistency loss that enforces agreement in prediction between modalities, exploiting the complementary information from multimodal inputs for robust human movement recognition. Through extensive experiments on action localization and group activity recognition tasks, our approach demonstrates superior performance when compared with state-of-the-art methods. Our code is available at: https://github.com/tranxuantuyen/COMPUTER.

📄 PDF Abstract BibTeX arXiv:2409.02385

Code (1)

tranxuantuyen/computer 공식 구현 pytorch

Tasks

Action LocalizationActivity RecognitionGroup Activity RecognitionHuman Activity Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?

2025-05-30 · Jiwan Chung, Janghan Yoon, Junhyeong Park, Sangeyl Lee 외

Any-to-any generative models aim to enable seamless interpretation and generation across multiple modalities within a unified framework, yet their ability to preserve relationships across modalities remains uncertain. Do…

LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV

2026-05-25 · Tengfei Liu, Yang Shi, Xuanyu Zhu, Jiafu Tang 외 arxiv

Audio-visual generation is rapidly advancing from short clips to minute-long content, while existing evaluation protocols remain largely confined to short-form settings. Existing benchmarks primarily focus on 5--10 secon…

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

2026-04-27 · Weixing Wang, Liudvikas Zekas, Anton Hackl, Constantin Alexander Auga 외 arxiv

Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these two capabilities independently and do no…

Towards Customized Multimodal Role-Play

2026-05-01 · Chao Tang, Jianzong Wu, Qingyu Shi, Ye Tian 외 arxiv

Unified multimodal understanding and generation models enable richer human-AI interaction. Yet jointly customizing a character's persona, dialogue style, and visual identity while maintaining output consistency across mo…

UniSOT: A Unified Framework for Multi-Modality Single Object Tracking

2025-11-03 · Yinchao Ma, Yuyang Tang, Wenfei Yang, Tianzhu Zhang 외 arxiv

Single object tracking aims to localize target object with specific reference modalities (bounding box, natural language or both) in a sequence of specific video modalities (RGB, RGB+Depth, RGB+Thermal or RGB+Event.). Di…

Object TrackingVisual Tracking