paper-with-me

홈 › Papers

CLEVER: Stream-based Active Learning for Robust Semantic Perception from Human Instructions

2025-07-21 · Jongseok Lee, Timo Birr, Rudolph Triebel, Tamim Asfour arxiv

We propose CLEVER, an active learning system for robust semantic perception with Deep Neural Networks (DNNs). For data arriving in streams, our system seeks human support when encountering failures and adapts DNNs online based on human instructions. In this way, CLEVER can eventually accomplish the given semantic perception tasks. Our main contribution is the design of a system that meets several desiderata of realizing the aforementioned capabilities. The key enabler herein is our Bayesian formulation that encodes domain knowledge through priors. Empirically, we not only motivate CLEVER's design but further demonstrate its capabilities with a user validation study as well as experiments on humanoid and deformable objects. To our knowledge, we are the first to realize stream-based active learning on a real robot, providing evidence that the robustness of the DNN-based semantic perception can be improved in practice. The project website can be accessed at https://sites.google.com/view/thecleversystem.

📄 PDF Abstract BibTeX arXiv:2507.15499

Code (0)

등록된 구현이 없습니다.

Tasks

Active Learning

Similar Papers 제목 키워드 기반

CleverDistiller: Simple and Spatially Consistent Cross-modal Distillation

2025-03-12 · Hariprasath Govindarajan, Maciej K. Wozniak, Marvin Klingner, Camille Maurice 외

Vision foundation models (VFMs) such as DINO have led to a paradigm shift in 2D camera-based perception towards extracting generalized features to support many downstream tasks. Recent works introduce self-supervised cro…

3D Object DetectionAutonomous DrivingKnowledge Distillationobject-detection+4

Visually Grounded Commonsense Knowledge Acquisition

2022-11-22 · Yuan YAO, Tianyu Yu, Ao Zhang, Mengdi Li 외

Large-scale commonsense knowledge bases empower a broad range of AI applications, where the automatic extraction of commonsense knowledge (CKE) is a fundamental and challenging problem. CKE from text is known for sufferi…

Language Modelling

StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding

2025-08-03 · Haolin Yang, Feilong Tang, Lingxiao Zhao, Xinlin Zhuang 외 arxiv

Real-time streaming video understanding in domains such as autonomous driving and intelligent surveillance poses challenges beyond conventional offline video processing, requiring continuous perception, proactive decisio…

Semantic RetrievalAutonomous DrivingDecision Making

Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction

2025-01-06 · CVPR 2025 1 · Rui Qian, Shuangrui Ding, Xiaoyi Dong, Pan Zhang 외

Active Real-time interaction with video LLMs introduces a new paradigm for human-computer interaction, where the model not only understands user intent but also responds while continuously processing streaming video on t…

Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs

2024-08-16 · Jinming Liu, Yuntao Wei, Junyan Lin, Shengyang Zhao 외

We present a new image compression paradigm to achieve ``intelligently coding for machine'' by cleverly leveraging the common sense of Large Multimodal Models (LMMs). We are motivated by the evidence that large language/…

Common Sense Reasoningimage-classificationImage ClassificationImage Compression+4