paper-with-me

홈 › Papers

Contrastive Learning for Enhancing Robust Scene Transfer in Vision-based Agile Flight

2023-09-18 · Jiaxu Xing, Leonard Bauersfeld, Yunlong Song, Chunwei Xing, Davide Scaramuzza

Scene transfer for vision-based mobile robotics applications is a highly relevant and challenging problem. The utility of a robot greatly depends on its ability to perform a task in the real world, outside of a well-controlled lab environment. Existing scene transfer end-to-end policy learning approaches often suffer from poor sample efficiency or limited generalization capabilities, making them unsuitable for mobile robotics applications. This work proposes an adaptive multi-pair contrastive learning strategy for visual representation learning that enables zero-shot scene transfer and real-world deployment. Control policies relying on the embedding are able to operate in unseen environments without the need for finetuning in the deployment environment. We demonstrate the performance of our approach on the task of agile, vision-based quadrotor flight. Extensive simulation and real-world experiments demonstrate that our approach successfully generalizes beyond the training domain and outperforms all baselines.

📄 PDF Abstract BibTeX arXiv:2309.09865

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningRepresentation Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation

2026-02-15 · Xinhua Wang, Kun Wu, Zhen Zhao, Hu Cao 외 arxiv

Enhancing the generalization capability of robotic learning to enable robots to operate effectively in diverse, unseen scenes is a fundamental and challenging problem. Existing approaches often depend on pretraining with…

Data AugmentationObject Detection

Multimodal Fusion for Sim2real Transfer in Visual Reinforcement Learning

2025-07-12 · Zichun Xu, Jingdong Zhao, Chenyu Guo, Qianxue Zhang 외 arxiv

Depth information is robust to scene appearance variations and inherently carries 3D spatial details. Thus, a visual backbone based on the vision transformer is proposed to fuse RGB and depth modalities for enhancing gen…

Reinforcement LearningContrastive Learning

Vision-Language Pre-training with Object Contrastive Learning for 3D Scene Understanding

2023-05-18 · Taolin Zhang, Sunan He, Dai Tao, Bin Chen 외

In recent years, vision language pre-training frameworks have made significant progress in natural language processing and computer vision, achieving remarkable performance improvement on various downstream tasks. Howeve…

Contrastive LearningObjectScene UnderstandingVisual Grounding

Human Centric General Physical Intelligence for Agile Manufacturing Automation

2025-08-16 · Sandeep Kanta, Mehrdad Tavassoli, Varun Teja Chirkuri, Venkata Akhil Kumar 외 arxiv

Agile human-centric manufacturing increasingly requires resilient robotic solutions that are capable of safe and productive interactions within unstructured environments of modern factories. While multi-modal sensor fusi…

Representation Learning

Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models

2025-10-01 · Yu Zeng, Wenxuan Huang, Shiting Huang, Xikun Bao 외 arxiv

Although current large Vision-Language Models (VLMs) have advanced in multimodal understanding and reasoning, their fundamental perceptual and reasoning abilities remain limited. Specifically, even on simple jigsaw tasks…

Reinforcement Learning