paper-with-me

홈 › Papers

Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching

2026-02-12 · Onkar Susladkar, Tushar Prakash, Gayatri Deshmukh, Kiet A. Nguyen, Jiaxun Zhang, Adheesh Juvekar, Tianshu Bao, Lin Chai, Sparsh Mittal, Inderjit S Dhillon, Ismini Lourentzou arxiv

We propose UniDFlow, a unified discrete flow-matching framework for multimodal understanding, generation, and editing. It decouples understanding and generation via task-specific low-rank adapters, avoiding objective interference and representation entanglement, while a novel reference-based multimodal preference alignment optimizes relative outcomes under identical conditioning, improving faithfulness and controllability without large-scale retraining. UniDFlpw achieves SOTA performance across eight benchmarks and exhibits strong zero-shot generalization to tasks including inpainting, in-context image generation, reference-based editing, and compositional generation, despite no explicit task-specific training.

📄 PDF Abstract BibTeX arXiv:2602.12221

Code (0)

등록된 구현이 없습니다.

Tasks

Zero-shot GeneralizationMultimodal ReasoningImage Generation

Similar Papers 제목 키워드 기반

TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning

2024-02-29 · Kate Sanders, Nathaniel Weir, Benjamin Van Durme

It is challenging for models to understand complex, multimodal content such as television clips, and this is in part because video-language models often rely on single-modality reasoning and lack interpretability. To com…

Question AnsweringVideo Understanding

Kling-Omni Technical Report

2025-12-18 · Kling Team, Jialu Chen, Yuanzheng Ci, Xiangyu Du 외 arxiv

We present Kling-Omni, a generalist generative framework designed to synthesize high-fidelity videos directly from multimodal visual language inputs. Adopting an end-to-end perspective, Kling-Omni bridges the functional …

Instruction FollowingVideo Generation

UEval: A Benchmark for Unified Multimodal Generation

2026-01-29 · Bo Li, Yida Yin, Wenhao Chai, Xingyu Fu 외 arxiv

We introduce UEval, a benchmark to evaluate unified models, i.e., models capable of generating both images and text. UEval comprises 1,000 expert-curated questions that require both images and text in the model output, s…

multimodal generation

MemoryDocDataSet: A Benchmark for Joint Conversational Memory and Long Document Reasoning

2026-06-03 · Qiyang Xie, Jialun Wu, Xinjie He, Su Liu 외 arxiv

AI systems increasingly need to combine two demanding capabilities: navigating multi-session conversation history and performing deep reading comprehension within long documents. Yet no existing benchmark evaluates both …

Reading Comprehension

Best of Both Worlds: A Hybrid Approach for Multi-Hop Explanation with Declarative Facts

2021-12-17 · AAAI Workshop CLeaR 2022 2 · Shane Storks, Qiaozi Gao, Aishwarya Reganti, Govind Thattai

Language-enabled AI systems can answer complex, multi-hop questions to high accuracy, but supporting answers with evidence is a more challenging task which is important for the transparency and trustworthiness to users. …

Explanation GenerationRetrieval