paper-with-me

홈 › Papers

T2MBench: A Benchmark for Out-of-Distribution Text-to-Motion Generation

2026-02-14 · Bin Yang, Rong Ou, Weisheng Xu, Jiaqi Xiong, Xintao Li, Taowen Wang, Luyu Zhu, Xu Jiang, Jing Tan, Renjing Xu arxiv

Most existing evaluations of text-to-motion generation focus on in-distribution textual inputs and a limited set of evaluation criteria, which restricts their ability to systematically assess model generalization and motion generation capabilities under complex out-of-distribution (OOD) textual conditions. To address this limitation, we propose a benchmark specifically designed for OOD text-to-motion evaluation, which includes a comprehensive analysis of 14 representative baseline models and the two datasets derived from evaluation results. Specifically, we construct an OOD prompt dataset consisting of 1,025 textual descriptions. Based on this prompt dataset, we introduce a unified evaluation framework that integrates LLM-based Evaluation, Multi-factor Motion evaluation, and Fine-grained Accuracy Evaluation. Our experimental results reveal that while different baseline models demonstrate strengths in areas such as text-to-motion semantic alignment, motion generalizability, and physical quality, most models struggle to achieve strong performance with Fine-grained Accuracy Evaluation. These findings highlight the limitations of existing methods in OOD scenarios and offer practical guidance for the design and evaluation of future production-level text-to-motion models.

📄 PDF Abstract BibTeX arXiv:2602.13751

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VMBench: A Benchmark for Perception-Aligned Video Motion Generation

2025-03-13 · Xinrang Ling, Chen Zhu, Meiqi Wu, Hangyu Li 외

Video generation has advanced rapidly, improving evaluation methods, yet assessing video's motion remains a major challenge. Specifically, there are two key issues: 1) current motion metrics do not fully align with human…

Motion GenerationVideo Generation

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

2026-07-02 · Yuanzhi Liu, Shousheng Zhao, Bo Zhou, Kongming Liang 외 arxiv

Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance. We pre…

Latent Temporal Discrepancy as Motion Prior: A Loss-Weighting Strategy for Dynamic Fidelity in T2V

2026-01-28 · Meiqi Wu, Bingze Song, Ruimin Lin, Chen Zhu 외 arxiv

Video generation models have achieved notable progress in static scenarios, yet their performance in motion video generation remains limited, with quality degrading under drastic dynamic changes. This is due to noise dis…

Video Generation

The Quest for Generalizable Motion Generation: Data, Model, and Evaluation

2025-10-30 · Jing Lin, Ruisi Wang, Junzhe Lu, Ziqi Huang 외 arxiv

Despite recent advances in 3D human motion generation (MoGen) on standard benchmarks, existing text-to-motion models still face a fundamental bottleneck in their generalization capability. In contrast, adjacent generativ…

Video Generation

Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy

2024-10-02 · Ricardo Garcia, ShiZhe Chen, Cordelia Schmid

Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address this gap by introducing GemBench, a novel…

Motion PlanningRobot ManipulationRobot Manipulation GeneralizationTask Planning