paper-with-me

Papers

PRE-MAP: Personalized Reinforced Eye-tracking Multimodal LLM for High-Resolution Multi-Attribute Point Prediction

2025-07-25 · Hanbing Wu, Ping Jiang, Anyang Su, Chenxu Zhao, Tianyu Fu, Minghui Wu, Beiping Tan, Huiying Li arxiv

Visual selective attention, driven by individual preferences, regulates human prioritization of visual stimuli by bridging subjective cognitive mechanisms with objective visual elements, thereby steering the semantic interpretation and hierarchical processing of dynamic visual scenes. However, existing models and datasets predominantly neglect the influence of subjective cognitive diversity on fixation behavior. Conventional saliency prediction models, typically employing segmentation approaches, rely on low-resolution imagery to generate saliency heatmaps, subsequently upscaled to native resolutions, which limiting their capacity to capture personalized attention patterns. Furthermore, MLLMs are constrained by factors such as hallucinations, making it very costly to strictly adhere to the expected format in tasks involving multiple point predictions, and achieving precise point positioning is challenging. To address these limitations, we present Subjective Personalized Attention for Advertisement Videos, namely SPA-ADV, a large-scale multimodal dataset capturing gaze behaviors from over 4,500 participants varying in age and gender with 486 videos. Furthermore, we propose PRE-MAP, a novel eye-tracking saliency model that characterizes Personalized visual disparities through Reinforcement learning-optimized Eye-tracking, built upon MLLMs and guided by Multi-Attribute user profiles to predict Points. To ensure MLLMs produce prediction points that are both format-correct and spatially accurate, we introduce Consistency Group Relative Policy Optimization (C-GRPO), inspired by the variability in eye movement points and Multi-Attribute profiles. Extensive experiments on SPA-ADV and other benchmarks demonstrate the effectiveness of our approach. The code and dataset are available at \href{https://github.com/mininglamp-MLLM/PRE-MAP}{this URL}.

📄 PDF Abstract BibTeX arXiv:2507.19213

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningSaliency Prediction

Similar Papers 제목 키워드 기반

Personalizing MLLMs via Reinforced Multimodal Reference Game

2026-06-27 · Deepayan Das, Davide Talon, Yiming Wang, Massimiliano Mancini 외 arxiv

Personalizing Multimodal Large Language Models (MLLMs) aims to recognize users' unique concepts from visual data and provide personalized responses. Although prior work has shown the benefit of concept descriptions and r…

Text Reinforcement for Multimodal Time Series Forecasting

2025-08-31 · Chen Su, Yuanhe Tian, Yan Song, Yongdong Zhang arxiv

Recent studies in time series forecasting (TSF) use multimodal inputs, such as text and historical time series data, to predict future values. These studies mainly focus on developing advanced techniques to integrate tex…

Time Series ForecastingReinforcement Learning

Progressive Modality Reinforcement for Human Multimodal Emotion Recognition From Unaligned Multimodal Sequences

2021-06-19 · CVPR 2021 1 · Fengmao Lv, Xiang Chen, Yanyong Huang, Lixin Duan 외

Human multimodal emotion recognition involves time-series data of different modalities, such as natural language, visual motions, and acoustic behaviors. Due to the variable sampling rates for sequences from differen…

Emotion RecognitionMultimodal Emotion RecognitionTime SeriesTime Series Analysis

ScrapMem: A Bio-inspired Framework for On-device Personalized Agent Memory via Optical Forgetting

2026-05-05 · Jiale Chang, Yuxiang Ren arxiv

Long-term personalized memory for LLM agents is challenging on resource-limited edge devices due to high storage costs and multimodal complexity. To address this, we propose ScrapMem, a framework that integrates multimod…

Towards General Multimodal Visual Tracking

2025-03-14 · Andong Lu, Mai Wen, Jinhu Wang, Yuanzhi Guo 외

Existing multimodal tracking studies focus on bi-modal scenarios such as RGB-Thermal, RGB-Event, and RGB-Language. Although promising tracking performance is achieved through leveraging complementary cues from different …

MambaVisual Tracking