paper-with-me

Papers

ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs

2026-02-09 · Xin Wu, Zhixuan Liang, Yue Ma, Mengkang Hu, Zhiyuan Qin, Xiu Li arxiv

Multimodal Large Language Models (MLLMs) have significantly advanced the landscape of embodied AI, yet transitioning to synchronized bimanual coordination introduces formidable challenges in multi-stream multimodal integration. We introduce ST-BiBench, a comprehensive multi-tier framework for evaluating spatio-temporal multimodal coordination. Our approach centers on Strategic Coordination Planning, assessing high-level cross-modal reasoning over multiple action and perception streams. To investigate the "proximity paradox"-where semantically coherent plans fail to align with spatially grounded visual inputs-we incorporate Foundational Spatial Grounding to verify workspace awareness and arm-selection logic. Furthermore, we probe model frontiers through Fine-Grained Action Control, investigating whether MLLMs can directly synthesize high-dimensional continuous action modalities (16-Dim) from complex multimodal metadata. Evaluating 30+ state-of-the-art MLLMs, we uncover a persistent and pervasive "coordination paradox"-a significant gap between high-level strategic reasoning and fine-grained physical execution. Results reveal that while frontier MLLMs excel at logic-driven strategy, they frequently suffer from perception-logic disconnection and multi-stream interference during multimodal fusion. ST-BiBench provides a platform for identifying critical bottlenecks in multi-stream multimodal fusion and cross-modal alignment for complex embodied tasks.

📄 PDF Abstract BibTeX arXiv:2602.08392

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BiBench: Benchmarking and Analyzing Network Binarization

2023-01-26 · Haotong Qin, Mingyuan Zhang, Yifu Ding, Aoyu Li 외

Network binarization emerges as one of the most promising compression approaches offering extraordinary computation and memory savings by minimizing the bit-width. However, recent research has shown that applying existin…

BenchmarkingBinarization

CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics

2025-05-06 · Junqi Liu, Xiaohan Lin, Jonas Bayer, Yael Dillies 외

Neurosymbolic approaches integrating large language models with formal reasoning have recently achieved human-level performance on mathematics competition problems in algebra, geometry and number theory. In comparison, c…

Benchmarking

Benchmark for Antibody Binding Affinity Maturation and Design

2025-05-23 · Xinyan Zhao, Yi-Ching Tang, Akshita Singh, Victor J Cantu 외

We introduce AbBiBench (Antibody Binding Benchmarking), a benchmarking framework for antibody binding affinity maturation and design. Unlike existing antibody evaluation strategies that rely on antibody alone and its sim…

Benchmarking

MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents

2025-12-14 · Youngmin Im, Byeongung Jo, Jaeyoung Wi, Seungwoo Baek 외 arxiv

Mobile GUI Agents, AI agents capable of interacting with mobile applications on behalf of users, have the potential to transform human computer interaction. However, current evaluation practices for GUI agents face two f…

ARBiBench: Benchmarking Adversarial Robustness of Binarized Neural Networks

2023-12-21 · Peng Zhao, Jiehua Zhang, Bowen Peng, Longguang Wang 외

Network binarization exhibits great potential for deployment on resource-constrained devices due to its low computational cost. Despite the critical importance, the security of binarized neural networks (BNNs) is rarely …

Adversarial RobustnessBenchmarkingBinarization