paper-with-me

홈 › Papers

Attend to Anything: Foundation Model for Unified Human Attention Modeling

2026-06-02 · Wenzhuo Zhao, Ronghao Xian, Keren Fu, Qijun Zhao arxiv

Existing human attention (saliency) modeling methods persist as highly fragmented across modalities, scenes, and task formulations. Consequently, even with increasing model capacity and data scale, current models predominantly remain scene-dependent and task-specific, failing to practically generalize in real-world applications. To address the fundamental limitations, we present the Attend to Anything Model (AAM), a multi-modal foundation model that unifies attention modeling across various image, video, and audio-visual tasks and scenes. AAM reformulates attention as a cognitive entailment relationship organized in a general-to-specific hierarchy, implemented through language prompts with hierarchical embeddings in hyperbolic space. Furthermore, to unify static image and dynamic video attention, we adopt a fluid-dynamics perspective, formulating video-frame attention as a diffusive temporal evolution governed by the Fokker--Planck equation. Extensive experiments on 16 benchmarks demonstrate that AAM consistently outperforms state-of-the-art methods by an average of 6\% across various scenarios, while achieving approximately a 4$\times$ speedup in video inference. Overall, these results demonstrate that AAM provides a principled foundation for future research on attention and saliency-related tasks. The dataset and code will be available at https://github.com/wz-zhao/Attend-to-Anything.

📄 PDF Abstract BibTeX arXiv:2606.03540

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MATCHA:Towards Matching Anything

2025-01-24 · Fei Xue, Sven Elflein, Laura Leal-Taixé, Qunjie Zhou

Establishing correspondences across images is a fundamental challenge in computer vision, underpinning tasks like Structure-from-Motion, image editing, and point tracking. Traditional methods are often specialized for sp…

Point Tracking

MATCHA: Towards Matching Anything

2025-01-01 · CVPR 2025 1 · Fei Xue, Sven Elflein, Laura Leal-Taixé, Qunjie Zhou

Establishing correspondences across images is a fundamental challenge in computer vision, underpinning tasks like Structure-from-Motion, image editing, and point tracking. Traditional methods are often specialized fo…

Point Tracking

Judge Anything: MLLM as a Judge Across Any Modality

2025-03-21 · Shu Pu, Yaochen Wang, Dongping Chen, Yuhang Chen 외

Evaluating generative foundation models on open-ended multimodal understanding (MMU) and generation (MMG) tasks across diverse modalities (e.g., images, audio, video) poses significant challenges due to the complexity of…

Hallucination

SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object Detection

2025-11-13 · Jia Lin, Xiaofei Zhou, Jiyuan Liu, Runmin Cong 외 arxiv

Recently segment anything model (SAM) has attracted widespread concerns, and it is often treated as a vision foundation model for universal segmentation. Some researchers have attempted to directly apply the foundation m…

Video Salient Object Detection

A Unified Model for Extractive and Abstractive Summarization using Inconsistency Loss

2018-05-16 · ACL 2018 7 · Wan-Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min 외

We propose a unified model combining the strength of extractive and abstractive summarization. On the one hand, a simple extractive model can obtain sentence-level attention with high ROUGE scores but less readable. On t…

Abstractive Text SummarizationSentence