paper-with-me

홈 › Papers

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

2026-09-17 · Danning Zhang, Yijing Lin, Shuhan Zhuang, Mengqi Huang, Shaojin Wu, Shancheng Fang, Zhendong Mao hf

Multi-modal image generation, particularly subject-driven customization, has garnered growing attention in recent years. Despite the rapid advancement of generative models, their evaluation remains largely lagging. Existing methods, whether embedding-based or Multi-modal Large Language Model (MLLM)-based, evaluate alignment with each modal condition in isolation, which contradicts the simultaneous condition alignment objective of multi-modal image generation, leading to poor consistency with human judgments. To address this challenge, we propose UFO, the first unified framework for omni-condition alignment simultaneous evaluation. Specifically, UFO introduces a novel Atomized Chain-of-Evaluation paradigm, i.e., it first decomposes omni-condition alignment into a sequential chain of fine-grained, disentangled Atomic Evaluation Units (AEUs), categorizes them into distinct modality-relevance classes, and then employs general or dedicated functional calls for accurate verification of different AEU types. Experimental results demonstrate that UFO achieves the highest correlation with human evaluation preferences, delivering an average improvement of 15.25%. Furthermore, we present UFO-Bench, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.

📄 PDF Abstract BibTeX arXiv:2609.12397

Code (2)

Tavish9/awesome-daily-AI-arxiv ★ 120
Valiant-Cat/hfpaper

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?

2026-02-02 · Susan Liang, Chao Huang, Filippos Bellos, Yolo Yunlong Tang 외 arxiv

State-of-the-art text-to-video generation models such as Sora 2 and Veo 3 can now produce high-fidelity videos with synchronized audio directly from a textual prompt, marking a new milestone in multi-modal generation. Ho…

Text-to-Video Generation

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs

2026-04-16 · Ziyang Luo, Nian Liu, Junwei Han arxiv

Omni-modal Large Language Models (Omni-MLLMs) promise a unified integration of diverse sensory streams. However, recent evaluations reveal a critical performance paradox: unimodal baselines frequently outperform joint mu…

OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment

2025-10-12 · Yiting Lu, Fengbin Guan, Yixin Gao, Yan Zhong 외 arxiv

Current visual evaluation approaches are typically constrained to a single task. To address this, we propose OmniQuality-R, a unified reward modeling framework that transforms multi-task quality reasoning into continuous…

Reinforcement Learning

Omni-Modal Dissonance Benchmark: Systematically Breaking Modality Consensus to Probe Robustness and Calibrated Abstention

2026-03-28 · Zabir Al Nazi, Shubhashis Roy Dipta, Md Rizwan Parvez arxiv

Existing omni-modal benchmarks attempt to measure modality-specific contributions, but their measurements are confounded: naturally co-occurring modalities carry correlated yet unequal information, making it unclear whet…

OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs

2025-03-27 · John Murzaku, Owen Rambow

The use of omni-LLMs (large language models that accept any modality as input), particularly for multimodal cognitive state tasks involving speech, is understudied. We present OmniVox, the first systematic evaluation of …

Emotion Recognition