paper-with-me

홈 › Papers

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models

2026-06-18 · Yifan Shen, Pei Tian, Xinzhuo Li, Bowen Fang, Shujun Xia, Bingxuan Li, Ana Jojic, Wenming Ye, Xu Cao, James Matthew Rehg, Ismini Lourentzou arxiv

Omni-modal models can ingest video, audio, and text, but unified access to multiple modalities does not guarantee that a model uses the right evidence. This gap is especially pronounced in social video question answering, where the answer may hinge on a gesture, vocal tone, temporal cue, or mismatch between what is said and what is visually expressed. We introduce CogniRoute, a schema-guided Mixture-of-Experts framework for social omni reasoning. CogniRoute uses a training-only cognitive schema that factorizes each example by cross-modal relation, reasoning demand, and temporal scope, and aligns global routing signatures with this structure during supervised fine-tuning. We further introduce route-aware reinforcement learning, which jointly optimizes token generation and expert allocation using rewards for answer correctness, modality-consistent reasoning, and cognitive temporal grounding. To support training and evaluation, we construct OmniSocialBench, a diagnostic social video QA resource with 118K structured training examples, grounded reasoning traces, schema labels, temporal evidence spans, and a manually verified evaluation split. CogniRoute achieves 59.38\% average accuracy on OmniSocialBench, improving over the strongest proprietary baseline by 15.33 percentage points and the strongest open-source omni baseline by 26.77 points, with the largest gains on questions requiring audio-visual coordination, conflict resolution, and temporally grounded social inference.

📄 PDF Abstract BibTeX arXiv:2606.20970

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question AnsweringReinforcement Learning

Similar Papers 제목 키워드 기반

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering

2026-05-25 · Ming Xie, Zizheng Huang, Xudong Tan, Chao Wang 외 arxiv

While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal methods are inherently designed for offlin…

Question AnsweringVisual Reasoning

OmniEgo-R$^2$: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026

2026-05-23 · Zixu Li, Zhiwei Chen, Zhiheng Fu, Wenbo Wang 외 arxiv

The 1st Cross-Domain EgoCross Challenge at EgoVis, CVPR 2026 evaluates whether multimodal large language models can reason over egocentric videos across surgery, industry, extreme sports, and animal perspective. We achie…

Visual Question AnsweringMultimodal Reasoning

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection

2026-05-02 · Tianxiao Li, Zhenglin Huang, Haiquan Wen, Yiwei He 외 arxiv

Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-modality scope, simplified manipulations…

DeepFake Detection

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

2026-05-12 · Che Liu, Lichao Ma, Xiangyu Tony Zhang, Yuxin Zhang 외 arxiv

Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evidence alone is enough to answer a query. We study whether current omni-…

OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models

2026-07-03 · Shijie Cao, Qingyu Zhang, Boxi Yu, Yuzhong Zhang 외 arxiv

Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they generate large token sequences under audio-visual inputs, leading to substantial in…