paper-with-me

Papers

UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models

2025-10-21 · Chen Chen, ZeYang Hu, Fengjiao Chen, Liya Ma, Jiaxing Liu, Xiaoyu Li, Ziwen Wang, Xuezhi Cao, Xunliang Cai arxiv

Multimodal Large Languages models have been progressing from uni-modal understanding toward unifying visual, audio and language modalities, collectively termed omni models. However, the correlation between uni-modal and omni-modal remains unclear, which requires comprehensive evaluation to drive omni model's intelligence evolution. In this work, we introduce a novel, high-quality, and UNified Omni model benchmark, UNO-Bench. This benchmark is designed to effectively evaluate both UNi-modal and Omni-modal capabilities under a unified ability taxonomy, spanning 44 task types and 5 modality combinations. It includes 1250 human curated samples for omni-modal with 98% cross-modality solvability, and 2480 enhanced uni-modal samples. The human-generated dataset is well-suited to real-world scenarios, particularly within the Chinese context, whereas the automatically compressed dataset offers a 90% increase in speed and maintains 98% consistency across 18 public benchmarks. In addition to traditional multi-choice questions, we propose an innovative multi-step open-ended question format to assess complex reasoning. A general scoring model is incorporated, supporting 6 question types for automated evaluation with 95% accuracy. Experimental result shows the Compositional Law between omni-modal and uni-modal performance and the omni-modal capability manifests as a bottleneck effect on weak models, while exhibiting synergistic promotion on strong models.

📄 PDF Abstract BibTeX arXiv:2510.18915

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring the Spectrum of Visio-Linguistic Compositionality and Recognition

2024-06-13 · Youngtaek Oh, Pyunghwan Ahn, Jinhyung Kim, Gwangmo Song 외

Vision and language models (VLMs) such as CLIP have showcased remarkable zero-shot recognition abilities yet face challenges in visio-linguistic compositionality, particularly in linguistic comprehension and fine-grained…

Retrievalzero-shot-classificationZero-Shot Learning

Seen to Unseen: Exploring Compositional Generalization of Multi-Attribute Controllable Dialogue Generation

2023-06-17 · Weihao Zeng, Lulu Zhao, Keqing He, Ruotong Geng 외

Existing controllable dialogue generation work focuses on the single-attribute control and lacks generalization capability to out-of-distribution multiple attribute combinations. In this paper, we explore the composition…

AttributeDialogue GenerationDisentanglement

CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration

2025-09-22 · Seyed Amir Kasaei, Ali Aghayari, Arash Marioriyad, Niki Sepasian 외 arxiv

Text-to-image diffusion models, such as Stable Diffusion, can produce high-quality and diverse images but often fail to achieve compositional alignment, particularly when prompts describe complex object relationships, at…

Infinity and Beyond: Compositional Alignment in VAR and Diffusion T2I Models

2025-12-12 · Hossein Shahabadi, Niki Sepasian, Arash Marioriyad, Ali Sharifi-Zarchi 외 arxiv

Achieving compositional alignment between textual descriptions and generated images - covering objects, attributes, and spatial relationships - remains a core challenge for modern text-to-image (T2I) models. Although dif…

Exploring Continual Learning of Compositional Generalization in NLI

2024-03-07 · Xiyan Fu, Anette Frank

Compositional Natural Language Inference has been explored to assess the true abilities of neural models to perform NLI. Yet, current evaluations assume models to have full access to all primitive inferences in advance, …

Continual LearningNatural Language Inference