paper-with-me

홈 › Papers

ReMoT: Reinforcement Learning with Motion Contrast Triplets

2026-02-28 · Cong Wan, Zeyu Guo, Jiangyang Li, SongLin Dong, Yifan Bai, Lin Peng, Zhiheng Ma, Yihong Gong arxiv

We present ReMoT, a unified training paradigm to systematically address the fundamental shortcomings of VLMs in spatio-temporal consistency -- a critical failure point in navigation, robotics, and autonomous driving. ReMoT integrates two core components: (1) A rule-based automatic framework that generates ReMoT-16K, a large-scale (16.5K triplets) motion-contrast dataset derived from video meta-annotations, surpassing costly manual or model-based generation. (2) Group Relative Policy Optimization, which we empirically validate yields optimal performance and data efficiency for learning this contrastive reasoning, far exceeding standard Supervised Fine-Tuning. We also construct the first benchmark for fine-grained motion contrast triplets to measure a VLM's discrimination of subtle motion attributes (e.g., opposing directions). The resulting model achieves state-of-the-art performance on our new benchmark and multiple standard VLM benchmarks, culminating in a remarkable 25.1% performance leap on spatio-temporal reasoning tasks.

📄 PDF Abstract BibTeX arXiv:2603.00461

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningAutonomous Driving

Similar Papers 제목 키워드 기반

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation

2024-12-10 · Thong Thanh Nguyen, Xiaobao Wu, Yi Bin, Cong-Duy T Nguyen 외

To equip artificial intelligence with a comprehensive understanding towards a temporal world, video and 4D panoptic scene graph generation abstracts visual data into nodes to represent entities and edges to capture tempo…

Contrastive LearningGraph GenerationPanoptic Scene Graph GenerationRelation+2

Representation Learning for Remote Sensing: An Unsupervised Sensor Fusion Approach

2021-08-11 · Aidan M. Swope, Xander H. Rudelis, Kyle T. Story

In the application of machine learning to remote sensing, labeled data is often scarce or expensive, which impedes the training of powerful models like deep convolutional neural networks. Although unlabeled data is abund…

Representation LearningSelf-Supervised LearningSensor Fusion

Informative and Representative Triplet Selection for Multilabel Remote Sensing Image Retrieval

2021-05-08 · Gencer Sumbul, Mahdyar Ravanbakhsh, Begüm Demir

Learning the similarity between remote sensing (RS) images forms the foundation for content-based RS image retrieval (CBIR). Recently, deep metric learning approaches that map the semantic similarity of images into an em…

Image RetrievalMetric LearningRetrievalSemantic Similarity+2

REMOTE: A Unified Multimodal Relation Extraction Framework with Multilevel Optimal Transport and Mixture-of-Experts

2025-09-05 · Xinkui Lin, Yongxiu Xu, Minghao Tang, Shilong Zhang 외 arxiv

Multimodal relation extraction (MRE) is a crucial task in the fields of Knowledge Graph and Multimedia, playing a pivotal role in multimodal knowledge graph construction. However, existing methods are typically limited t…

Relation Extraction

Gaze on the Prize: Shaping Visual Attention with Return-Guided Contrastive Learning

2025-10-09 · Andrew Lee, Ian Chuang, Dechen Gao, Kai Fukazawa 외 arxiv

Visual Reinforcement Learning (RL) agents must learn to act based on high-dimensional image data where only a small fraction of the pixels is task-relevant. This forces agents to waste exploration and computational resou…

Reinforcement LearningContrastive Learning