paper-with-me

홈 › Papers

Semantic Tracklets: An Object-Centric Representation for Visual Multi-Agent Reinforcement Learning

2021-08-06 · Iou-Jen Liu, Zhongzheng Ren, Raymond A. Yeh, Alexander G. Schwing

Solving complex real-world tasks, e.g., autonomous fleet control, often involves a coordinated team of multiple agents which learn strategies from visual inputs via reinforcement learning. Many existing multi-agent reinforcement learning (MARL) algorithms however don't scale to environments where agents operate on visual inputs. To address this issue, algorithmically, recent works have focused on non-stationarity and exploration. In contrast, we study whether scalability can also be achieved via a disentangled representation. For this, we explicitly construct an object-centric intermediate representation to characterize the states of an environment, which we refer to as semantic tracklets.' We evaluate semantic tracklets' on the visual multi-agent particle environment (VMPE) and on the challenging visual multi-agent GFootball environment. `Semantic tracklets' consistently outperform baselines on VMPE, and achieve a +2.4 higher score difference than baselines on GFootball. Notably, this method is the first to successfully learn a strategy for five players in the GFootball environment using only visual data.

📄 PDF Abstract BibTeX arXiv:2108.03319

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-agent Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Multi-Face Tracking by Extended Bag-of-Tracklets in Egocentric Videos

2015-07-16 · Maedeh Aghaei, Mariella Dimiccoli, Petia Radeva

Wearable cameras offer a hands-free way to record egocentric images of daily experiences, where social events are of special interest. The first step towards detection of social events is to track the appearance of multi…

STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation

2026-01-28 · Alexandre Chapin, Emmanuel Dellandréa, Liming Chen arxiv

Visual foundation models provide strong perceptual features for robotics, but their dense representations lack explicit object-level structure, limiting robustness and controllability in manipulation tasks. We propose ST…

CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs

2025-11-18 · Jingyu Lei, Gaoang Wang, Der-Horng Lee arxiv

Large Vision-Language Models (LVLMs) usually suffer from prohibitive computational and memory costs due to the quadratic growth of visual tokens with image resolution. Existing token compression methods, while varied, of…

Self-Supervised Visual Representation Learning with Semantic Grouping

2022-05-30 · Xin Wen, Bingchen Zhao, Anlin Zheng, Xiangyu Zhang 외

In this paper, we tackle the problem of learning visual representations from unlabeled scene-centric data. Existing works have demonstrated the potential of utilizing the underlying complex structure within scene-centric…

Contrastive LearningInstance SegmentationObject DetectionObject Discovery+5

GSLAMOT: A Tracklet and Query Graph-based Simultaneous Locating, Mapping, and Multiple Object Tracking System

2024-08-17 · Shuo Wang, Yongcai Wang, Zhimin Xu, Yongyu Guo 외

For interacting with mobile objects in unfamiliar environments, simultaneously locating, mapping, and tracking the 3D poses of multiple objects are crucially required. This paper proposes a Tracklet Graph and Query Graph…

Multiple Object TrackingObjectobject-detectionObject Detection+1