paper-with-me

홈 › Papers

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

2026-06-07 · Jiahao Wang, An Ping, Yanghai Wang, Yuanxing Zhang, Shihao Li, Hanyan Bian, Yichi Ren, Yize Zhang, Han Wang, Haowen Chen, Junze Li, Jiaqi Wang, Yiyang Hu, Zhuze Xu, Zijie Zhang, Jiaheng Liu arxiv

While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex, multi-faceted user instructions remains largely unexplored. Existing benchmarks primarily focus on holistic video understanding or text-only instruction following, failing to capture the intricate interplay between modalities and user constraints. To bridge this gap, we introduce OmniCap-IF, the first comprehensive benchmark specifically designed to evaluate instruction-following capabilities in omni-modal captioning. OmniCap-IF incorporates a systematic framework that assesses captions on two dimensions: format correctness and content correctness. Our benchmark encompasses 50 distinct constraint types across pure visual, pure audio, and audio-visual modalities, while integrating Temporal Grounding to assess spatio-temporal precision. Extensive evaluations of prominent models on 1,920 high-quality samples reveal significant performance disparities. Furthermore, our analysis uncovers a critical "format-content tradeoff", demonstrating that increasing formatting complexity directly degrades models' omni-modal reasoning abilities. Finally, to advance the field, we curate a 54K instruction-tuning dataset, OmniCap-IF-54K and present OmniCaptioner-IF, which achieves notable improvements in both complex instruction adherence and general omni-modal captioning performance.

📄 PDF Abstract BibTeX arXiv:2606.08572

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingVideo Captioning

Similar Papers 제목 키워드 기반

OmniCaptioner: One Captioner to Rule Them All

2025-04-09 · Yiting Lu, Jiakang Yuan, Zhen Li, Shitian Zhao 외

We propose OmniCaptioner, a versatile visual captioning framework for generating fine-grained textual descriptions across a wide variety of visual domains. Unlike prior methods limited to specific image types (e.g., natu…

AllImage CaptioningImage GenerationText to Image Generation+2

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks

2025-05-24 · Jiayu Wang, Yang Jiao, Yue Yu, Tianwen Qian 외

Recent breakthroughs in large multimodal models (LMMs), such as the impressive GPT-4o-Native, have demonstrated remarkable proficiency in following general-purpose instructions for image generation. However, current benc…

Image GenerationInstruction Followingmultimodal generation

Omni-Interactive Universal Embedder

2026-08-27 · Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon 외 arxiv

Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, e…

Representation Learning

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization

2026-02-05 · Minghao Han, Dingkang Yang, Yue Jiang, Yizhou Liu 외 arxiv

The autonomous evolution of networked AI systems relies heavily on robust environmental perception. However, physical understanding remains brittle in current models because key physical signals are visually ambiguous an…

Omnibenchmark (alpha) for continuous and open benchmarking in bioinformatics

2024-09-25 · Izaskun Mallona, Almut Luetge, Ben Carrillo, Daniel Incicau 외

Benchmarking in bioinformatics is a process of designing, running and disseminating rigorous performance evaluations of methods (software). Benchmarking systems facilitate the benchmarking process by providing an entrypo…

Benchmarking