paper-with-me

홈 › Papers

To See or To Read: User Behavior Reasoning in Multimodal LLMs

2025-11-05 · Tianning Dong, Luyi Ma, Varun Vasudevan, Jason Cho, Sushant Kumar, Kannan Achan arxiv

Multimodal Large Language Models (MLLMs) are reshaping how modern agentic systems reason over sequential user-behavior data. However, whether textual or image representations of user behavior data are more effective for maximizing MLLM performance remains underexplored. We present \texttt{BehaviorLens}, a systematic benchmarking framework for assessing modality trade-offs in user-behavior reasoning across six MLLMs by representing transaction data as (1) a text paragraph, (2) a scatter plot, and (3) a flowchart. Using a real-world purchase-sequence dataset, we find that when data is represented as images, MLLMs next-purchase prediction accuracy is improved by 87.5% compared with an equivalent textual representation without any additional computational cost.

📄 PDF Abstract BibTeX arXiv:2511.03845

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

From Rows to Reasoning: A Retrieval-Augmented Multimodal Framework for Spreadsheet Understanding

2026-01-13 · Anmol Gulati, Sahil Sen, Waqar Sarguroh, Kevin Paul arxiv

Large Language Models (LLMs) struggle to reason over large-scale enterprise spreadsheets containing thousands of numeric rows, multiple linked sheets, and embedded visual content such as charts and receipts. Prior state-…

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

2026-08-13 · Xinming Wang, Weinong Wang, Hongming Yang, Yansong Lin 외 arxiv

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, thei…

Reinforcement Learning

Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models

2025-09-19 · Renjie Pi, Kehao Miao, Li Peihang, Runtao Liu 외 arxiv

Multimodal large language models (MLLMs) have demonstrated extraordinary capabilities in conducting conversations based on image inputs. However, we observe that MLLMs exhibit a pronounced form of visual sycophantic beha…

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?

2025-05-17 · Zihao Dongfang, Xu Zheng, Ziqiao Weng, Yuanhuiyi Lyu 외

The 180x360 omnidirectional field of view captured by 360-degree cameras enables their use in a wide range of applications such as embodied AI and virtual reality. Although recent advances in multimodal large language mo…

HallucinationObject CountingSpatial Reasoning

Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning

2025-07-07 · Yana Wei, Liang Zhao, Jianjian Sun, Kangheng Lin 외

The remarkable reasoning capability of large language models (LLMs) stems from cognitive behaviors that emerge through reinforcement with verifiable rewards. This work investigates how to transfer this principle to Multi…

Reinforcement Learning (RL)Visual Reasoning