paper-with-me

홈 › Papers

FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning

2025-12-14 · Yue Jiang, Dingkang Yang, Minghao Han, Jinghang Han, Zizhi Chen, Yizhou Liu, Mingcheng Li, Peng Zhai, Lihua Zhang arxiv

Despite rapid progress in multimodal large language models (MLLMs) and emerging omni-modal architectures, current benchmarks remain limited in scope and integration, suffering from incomplete modality coverage, restricted interaction to text-centric outputs, and weak interdependence and complementarity among modalities. To bridge these gaps, we introduce FysicsWorld, the first unified full-modality benchmark that supports bidirectional input-output across image, video, audio, and text, enabling comprehensive any-to-any evaluation across understanding, generation, and reasoning. FysicsWorld encompasses 16 primary tasks and 3,268 curated samples, aggregated from over 40 high-quality sources and covering a rich set of open-domain categories with diverse question types. We also propose the Cross-Modal Complementarity Screening (CMCS) strategy integrated in a systematic data construction framework that produces omni-modal data for spoken interaction and fusion-dependent cross-modal reasoning. Through a comprehensive evaluation of over 30 state-of-the-art baselines, spanning MLLMs, modality-specific models, unified understanding-generation models, and omni-modal language models, FysicsWorld exposes the performance disparities and limitations across models in understanding, generation, and reasoning. Our benchmark establishes a unified foundation and strong baselines for evaluating and advancing next-generation full-modality architectures.

📄 PDF Abstract BibTeX arXiv:2512.12756

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation

2025-06-20 · Teng Li, Quanfeng Lu, Lirui Zhao, Hao Li 외

Unified image understanding and generation has emerged as a promising paradigm in multimodal artificial intelligence. Despite recent progress, the optimal architectural design for such unified models remains an open chal…

Representation Learning

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

2024-08-22 · Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang 외

We present a unified transformer, i.e., Show-o, that unifies multimodal understanding and generation. Unlike fully autoregressive models, Show-o unifies autoregressive and (discrete) diffusion modeling to adaptively hand…

10-shot image generationImage GenerationQuestion Answering+3

Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

2026-03-09 · Jaeik Kim, Woojin Kim, Jihwan Hong, Yejoon Lee 외 arxiv

We present Dynin-Omni, the first masked-diffusion-based omnimodal foundation model that unifies text, image, and speech understanding and generation, together with video understanding, within a single architecture. Unlik…

Cross-Modal RetrievalSpeech RecognitionImage Generation

See or Say Graphs: Agent-Driven Scalable Graph Structure Understanding with Vision-Language Models

2025-10-19 · Shuo Han, Yukun Cao, Zezhong Ding, Zengyi Gao 외 arxiv

Vision-language models (VLMs) have shown promise in graph structure understanding, but remain limited by input-token constraints, facing scalability bottlenecks and lacking effective mechanisms to coordinate textual and …

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

2026-08-05 · Yue Zhang, Yingzhao Jian, Yunqiu Xu, Xiaoxiao Sun 외 hf

Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these mod…

Multimodal ReasoningScene Understanding