paper-with-me

홈 › Papers

Context Unrolling in Omni Models

2026-04-23 · Ceyuan Yang, Zhijie Lin, Yang Zhao, Fei Xiao, Hao He, Qi Zhao, Chaorui Deng, Kunchang Li, Zihan Ding, Yuwei Guo, Fuyun Wang, Fangqi Zhu, Xiaonan Nie, Shenhan Zhu, Shanchuan Lin, Hongsheng Li, Weilin Huang, Guang Shi, Haoqi Fan arxiv

We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such training enables Context Unrolling, where the model explicitly reasons across multiple modal representations before producing predictions. This process enables the model to aggregate complementary information across heterogeneous modalities, facilitating a more faithful approximation of the shared multimodal knowledge manifold and improving downstream reasoning fidelity. As a result, Omni achieves strong performance on both multimodal generation and understanding benchmarks, while demonstrating advanced multimodal reasoning capabilities, including in-context generation of text, image, video, and 3D geometry.

📄 PDF Abstract BibTeX arXiv:2604.21921

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generationMultimodal Reasoning

Similar Papers 제목 키워드 기반

Deep Algorithm Unrolling for Biomedical Imaging

2021-08-15 · Yuelong Li, Or Bar-Shira, Vishal Monga, Yonina C. Eldar

In this chapter, we review biomedical applications and breakthroughs via leveraging algorithm unrolling, an important technique that bridges between traditional iterative algorithms and modern deep learning techniques. T…

Image GenerationRolling Shutter Correction

A statistical perspective on algorithm unrolling models for inverse problems

2023-11-10 · Yves Atchade, Xinru Liu, Qiuyun Zhu

We consider inverse problems where the conditional distribution of the observation ${\bf y}$ given the latent variable of interest ${\bf x}$ (also known as the forward model) is known, and we have access to a data set in…

Learning Variational Models with Unrolling and Bilevel Optimization

2022-09-26 · Christoph Brauer, Niklas Breustedt, Timo de Wolff, Dirk A. Lorenz

In this paper we consider the problem of learning variational models in the context of supervised learning via risk minimization. Our goal is to provide a deeper understanding of the two approaches of learning of variati…

Bilevel OptimizationRolling Shutter Correction

OmniGen2: Exploration to Advanced Multimodal Generation

2025-06-23 · Chenyuan Wu, Pengfei Zheng, Ruiran Yan, Shitao Xiao 외

In this work, we introduce OmniGen2, a versatile and open-source generative model designed to provide a unified solution for diverse generation tasks, including text-to-image, image editing, and in-context generation. Un…

Image Generationmultimodal generationText Generation

OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs

2025-03-27 · John Murzaku, Owen Rambow

The use of omni-LLMs (large language models that accept any modality as input), particularly for multimodal cognitive state tasks involving speech, is understudied. We present OmniVox, the first systematic evaluation of …

Emotion Recognition