paper-with-me

Papers

Seeing the Unseen: Visual Similarity for Pixel Language Model Adaptation

2026-08-31 · Ran Zhang, Miryam de Lhoneux, Wessel Poelman arxiv

Pixel-based language models (LMs) replace traditional tokenizers by processing rendered images of text, making cross-lingual transfer heavily dependent on the visual and structural properties of writing systems. However, the dynamics of adapting these models to low-resource languages with complex morphology and written in unique scripts are not yet explored. Using Tibetan as a case study, we analyze how continued pre-training of pixel-based LMs is influenced by data scale, initial script exposure, and cross-lingual transfer from languages written in other Brahmic scripts. We introduce four rendering-level metrics to quantify visual script similarity. We evaluate downstream performance across three tasks. Our results show that higher orthographic proximity enhances semantic transfer, even under severe data constraints. Additionally, we find a performance asymmetry based on the pre-training starting point: while multilingual pre-training PIXEL-M4 has stronger initial performance, its capacity for subsequent adaptation seems to be constrained, whereas adapting a monolingual model PIXEL with mixed scripts yields more gains on sentence-level tasks. Our metrics and case study offer empirical observations that could help inform data selection and script adaptation choices when working with pixel-based models in similar low-resource settings.

📄 PDF Abstract BibTeX arXiv:2608.30541

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Lingual Transfer

Similar Papers 제목 키워드 기반

Multimodal One-Shot Learning of Speech and Images

2018-11-09 · Ryan Eloff, Herman A. Engelbrecht, Herman Kamper

Imagine a robot is shown new concepts visually together with spoken tags, e.g. "milk", "eggs", "butter". After seeing one paired audio-visual example per class, it is shown a new set of unseen instances of these objects,…

Dynamic Time WarpingOne-Shot Learning

Seeing Through the Clouds: Cloud Gap Imputation with Prithvi Foundation Model

2024-04-30 · Denys Godwin, Hanxi Li, Michael Cecil, Hamed Alemohammad

Filling cloudy pixels in multispectral satellite imagery is essential for accurate data analysis and downstream applications, especially for tasks which require time series data. To address this issue, we compare the per…

Generative Adversarial NetworkImputationTime Series

Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy

2025-12-12 · Kechun Xu, Zhenjie Zhu, Anzhe Chen, Shuqi Zhao 외 arxiv

The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone during fine-tuning. While co-training with…

Instruction Following

Zero-Shot Object Recognition by Semantic Manifold Distance

2015-06-01 · CVPR 2015 6 · Zhenyong Fu, Tao Xiang, Elyor Kodirov, Shaogang Gong

Object recognition by zero-shot learning (ZSL) aims to recognise objects without seeing any visual examples by learning knowledge transfer between seen and unseen object classes. This is typically achieved by exploring a…

AttributeObjectObject RecognitionTransfer Learning+1

A Deep Visual Correspondence Embedding Model for Stereo Matching Costs

2015-12-01 · ICCV 2015 12 · Zhuoyuan Chen, Xun Sun, Liang Wang, Yinan Yu 외

This paper presents a data-driven matching cost for stereo matching. A novel deep visual correspondence embedding model is trained via Convolutional Neural Network on a large set of stereo images with ground truth dispar…

Stereo MatchingStereo Matching Hand