paper-with-me

홈 › Papers

MOMA: Multi-Object Multi-Actor Activity Parsing

2021-12-01 · NeurIPS 2021 12 · Zelun Luo, Wanze Xie, Siddharth Kapoor, Yiyun Liang, Michael Cooper, Juan Carlos Niebles, Ehsan Adeli, Fei-Fei Li

Complex activities often involve multiple humans utilizing different objects to complete actions (e.g., in healthcare settings, physicians, nurses, and patients interact with each other and various medical devices). Recognizing activities poses a challenge that requires a detailed understanding of actors' roles, objects' affordances, and their associated relationships. Furthermore, these purposeful activities are composed of multiple achievable steps, including sub-activities and atomic actions, which jointly define a hierarchy of action parts. This paper introduces Activity Parsing as the overarching task of temporal segmentation and classification of activities, sub-activities, atomic actions, along with an instance-level understanding of actors, objects, and their relationships in videos. Involving multiple entities (actors and objects), we argue that traditional pair-wise relationships, often used in scene or action graphs, do not appropriately represent the dynamics between them. Hence, we introduce Action Hypergraph, a spatial-temporal graph containing hyperedges (i.e., edges with higher-order relationships), as a new representation. In addition, we introduce Multi-Object Multi-Actor (MOMA), the first benchmark and dataset dedicated to activity parsing. Lastly, to parse a video, we propose the HyperGraph Activity Parsing (HGAP) network, which outperforms several baselines, including those based on regular graphs and raw video data.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Object

Similar Papers 제목 키워드 기반

MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing

2022-11-28 · NeurIPS 2022 11 · Zelun Luo, Zane Durante, Linden Li, Wanze Xie 외

Video-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabi…

Activity RecognitionFew Shot Action RecognitionGraph GenerationVideo Understanding

MOMA-AC: A preference-driven actor-critic framework for continuous multi-objective multi-agent reinforcement learning

2025-11-22 · Adam Callaghan, Karl Mason, Patrick Mannion arxiv

This paper addresses a critical gap in Multi-Objective Multi-Agent Reinforcement Learning (MOMARL) by introducing the first dedicated inner-loop actor-critic framework for continuous state and action spaces: Multi-Object…

Multi-agent Reinforcement Learning

MOMAland: A Set of Benchmarks for Multi-Objective Multi-Agent Reinforcement Learning

2024-07-23 · Florian Felten, Umut Ucak, Hicham Azmani, Gao Peng 외

Many challenging tasks such as managing traffic systems, electricity grids, or supply chains involve complex decision-making processes that must balance multiple conflicting objectives and coordinate the actions of vario…

BenchmarkingDecision MakingMulti-agent Reinforcement LearningMulti-Objective Multi-Agent Reinforcement Learning+3

MoMA: Multimodal LLM Adapter for Fast Personalized Image Generation

2024-04-08 · Kunpeng Song, Yizhe Zhu, Bingchen Liu, Qing Yan 외

In this paper, we present MoMA: an open-vocabulary, training-free personalized image model that boasts flexible zero-shot capabilities. As foundational text-to-image models rapidly evolve, the demand for robust image-to-…

Image GenerationImage-to-Image TranslationLanguage ModelingLanguage Modelling+3

Towards Fine-Grained Video Question Answering

2025-03-10 · Wei Dai, Alan Luo, Zane Durante, Debadutta Dash 외

In the rapidly evolving domain of video understanding, Video Question Answering (VideoQA) remains a focal point. However, existing datasets exhibit gaps in temporal and spatial granularity, which consequently limits the …

Language ModelingLanguage ModellingLarge Language ModelQuestion Answering+3