paper-with-me

Papers

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

2026-05-12 · Chaeyoung Jung, Kyeongha Rho, Joon Son Chung arxiv

Omnimodal Large Language Models (Omni-LLMs) incur substantial computational overhead due to the large number of multimodal input tokens they process, making token reduction essential for real-world deployment. Existing Omni-LLM pruning methods typically reduce this cost by selecting tokens that are important for the current query or strongly aligned with cross-modal cues. However, such strategies can discard evidence that falls outside these criteria, even when needed for different questions or for understanding context beyond aligned audio-visual cues. To address this limitation, we reframe Omni-LLM token reduction as preserving broad audio-visual context while removing cross-modal redundancy. We propose ContextGuard, an inference-time token pruning framework built on this principle. ContextGuard predicts coarse visual semantics from audio and prunes video tokens whose coarse semantics are likely recoverable from audio, while retaining additional video tokens to preserve localized visual details that audio alone cannot specify. For further compression, our method merges temporally similar video tokens. The framework requires no downstream LLM fine-tuning and uses only an independently trained lightweight predictor. On Qwen2.5-Omni and Video-SALMONN2+ at 3B and 7B scales across six audio-visual benchmarks, ContextGuard outperforms prior inference-time pruning methods while pruning more tokens. Notably, on Qwen2.5-Omni 7B, ContextGuard achieves full-token-level performance on five of six benchmarks while pruning 55% of input tokens.

📄 PDF Abstract BibTeX arXiv:2605.11605

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Wan-Streamer v0.2: Higher Resolution, Same Latency

2026-07-05 · Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang 외 hf

We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps the v0.1 modeling formulation, but raises the interactive output stream from 192x3…

PReM: Learning What to Preserve and When to Refresh for Context Compression

2026-07-15 · Bohan Yu, Lei Shen, Chenxi Zhou, Chen Han 외 arxiv

Efficient long-context inference is not only about reducing memory cost, but also about keeping useful contextual evidence accessible as generation proceeds. However, existing compression-oriented approaches, such as key…

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

2026-08-12 · Yuxuan Zhang, Haozhong Xiong, Jiayi Song, Jinpeng Yu 외 arxiv

Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized m…

ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA

2026-03-10 · Aviad Dahan, Moran Yanuka, Noa Kraicer, Lior Wolf 외 arxiv

Existing video personalization methods preserve visual likeness but treat video and audio separately. Without access to the visual scene, audio models cannot synchronize sounds with on-screen actions; and because classic…

Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory

2023-10-27 · Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov 외

The interactive use of large language models (LLMs) in AI assistants (at work, home, etc.) introduces a new set of inference-time privacy risks: LLMs are fed different types of information from multiple sources in their …

Privacy Preserving