paper-with-me

Papers

Tokenization of Gaze Data

2025-03-28 · Tim Rolff, Jurik Karimian, Niklas Hypki, Susanne Schmidt, Markus Lappe, Frank Steinicke

A considerable part of the performance of today's large language models (LLM's) and multimodal large language models (MLLM's) depends on their tokenization strategies. While tokenizers are extensively researched for textual and visual input, there is no research on tokenization strategies for gaze data due to its nature. However, a corresponding tokenization strategy would allow using the vision capabilities of pre-trained MLLM's for gaze data, for example, through fine-tuning. In this paper, we aim to close this research gap by analyzing five different tokenizers for gaze data on three different datasets for the forecasting and generation of gaze data through LLMs (cf.~\cref{fig:teaser}). We evaluate the tokenizers regarding their reconstruction and compression abilities. Further, we train an LLM for each tokenization strategy, measuring its generative and predictive performance. Overall, we found that a quantile tokenizer outperforms all others in predicting the gaze positions and k-means is best when predicting gaze velocities.

📄 PDF Abstract BibTeX arXiv:2503.22145

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers

2025-07-21 · Ian Chuang, Jinyu Zou, Andrew Lee, Dechen Gao 외 arxiv

Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing. In contrast, robot learning systems typically rely on p…

Robot ManipulationImage Segmentation

Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling

2026-08-24 · Virmarie Maquiling, Zhuojiang Cai, Enkelejda Kasneci arxiv

Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw traces preserve detail but are noisy and device-dependent, while coarse e…

On the Proper Treatment of Tokenization in Psycholinguistics

2024-10-03 · Mario Giulianelli, Luca Malagutti, Juan Luis Gastaldi, Brian DuSell 외

Language models are widely used in computational psycholinguistics to test theories that relate the negative log probability (the surprisal) of a region of interest (a substring of characters) under a language model to i…

Language ModelingLanguage Modelling

STARE: Predicting Decision Making Based on Spatio-Temporal Eye Movements

2025-08-06 · Moshe Unger, Alexander Tuzhilin, Michel Wedel arxiv

The present work proposes a Deep Learning architecture for the prediction of various consumer choice behaviors from time series of raw gaze or eye fixations on images of the decision environment, for which currently no f…

Decision Making

GazeGen: Gaze-Driven User Interaction for Visual Content Generation

2024-11-07 · He-Yen Hsieh, Ziyun Li, Sai Qian Zhang, Wei-Te Mark Ting 외

We present GazeGen, a user interaction system that generates visual content (images and videos) for locations indicated by the user's eye gaze. GazeGen allows intuitive manipulation of visual content by targeting regions…

Gaze EstimationKnowledge DistillationRaspberry Pi 4