paper-with-me

Papers

ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation

2025-11-25 · Yuhan Wu, Tiantian Wei, Shuo Wang, ZhiChao Wang, Yanyong Zhang, Daniel Cremers, Yan Xia arxiv

Interactive articulated manipulation requires long-horizon, multi-step interactions with appliances while maintaining physical consistency. Existing vision-language and diffusion-based policies struggle to generalize across parts, instances, and categories. We first introduce ArtiBench, a five-level benchmark covering kitchen, storage, office, and tool environments. ArtiBench enables structured evaluation from cross-part and cross-instance variation to long-horizon multi-object tasks, revealing the core generalization challenges of articulated object manipulation. Building on this benchmark, we propose ArtiBrain, a modular framework that unifies high-level reasoning with adaptive low-level control. ArtiBrain uses a VLM-based Task Reasoner (GPT-4.1) to decompose and validate subgoals, and employs a Hybrid Controller that combines geometry-aware keyframe execution with affordance-guided diffusion for precise and interpretable manipulation. An Affordance Memory Bank continually accumulates successful execution episodes and propagates part-level actionable affordances to unseen articulated parts and configurations. Extensive experiments on ArtiBench show that our ArtiBrain significantly outperforms state-of-the-art multimodal and diffusion-based methods in robustness and generalization. Code and dataset will be released upon acceptance.

📄 PDF Abstract BibTeX arXiv:2511.20330

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MarsRetrieval: Benchmarking Vision-Language Models for Planetary-Scale Geospatial Retrieval on Mars

2026-02-15 · Shuoyuan Wang, Yiran Wang, Hongxin Wei arxiv

Data-driven approaches like deep learning are rapidly advancing planetary science, particularly in Mars exploration. Despite recent progress, most existing benchmarks remain confined to closed-set supervised visual tasks…

Text Retrieval

Benchmarking Vision Foundation Models for Domain-Generalizable Face Anti-Spoofing

2026-04-21 · Mika Feng, Pierre Gallin-Martel, Koichi Ito, Takafumi Aoki arxiv

Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision-Language Models (VLMs) for semantic supervision, these …

Computational EfficiencyDomain GeneralizationFace Anti-SpoofingData Augmentation

BenchHAR: Benchmarking Self-Supervised Learning for Generalizable Sensor-based Activity Recognition

2026-05-08 · Yize Cai, Rui Feng, Anlan Yu, Baoshen Guo 외 arxiv

Human Activity Recognition (HAR) from wearable sensors supports broad healthcare and behavior science applications. However, data heterogeneity and the scarcity of labeled data limit its real-world generalization. Recent…

Human Activity RecognitionSelf-Supervised Learning

HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning

2026-04-17 · Yanbin Wei, Chun Kang, Siwei Li, Haoxuan Che 외 arxiv

Large Vision-Language Models (LVLMs) consistently require new arenas to guide their expanding boundaries, yet their capabilities with hypergraphs remain unexplored. In the real world, hypergraphs have significant practic…

FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents

2025-06-02 · Bobo Li, Yuheng Wang, Hao Fei, Juncheng Li 외

Online form filling is a common yet labor-intensive task involving extensive keyboard and mouse interactions. Despite the long-standing vision of automating this process with "one click", existing tools remain largely ru…

BenchmarkingForm