paper-with-me

Papers

Rate-Utility Frontiers for Language Encodings: Comparing Tokens, Bytes, and Pixels Under Controlled Linguistic Content

2026-07-17 · Ingo Ziegler, Martin Krebs, Desmond Elliott arxiv

Language models encode text as subword tokens, raw bytes, or rendered pixels, but these encodings are usually compared under modeling constraints that expose different amounts of linguistic content to models across different languages. We instead ask what each encoding preserves when both the content and the downstream capacity are controlled. Using verified parallel sentences across thirteen languages and five scripts, we compare tokens, bytes, and pixels through a shared bottleneck whose width is swept to trace rate-utility frontiers. This separates three quantities that are often conflated: the number of input positions an encoding creates, the latent capacity available after encoding, and the task-relevant information that survives compression. We evaluate three utilities: surface form preservation, cross-lingual sentence alignment, and topic classification. No encoding dominates across tasks or capacity regimes. Pixels preserve surface form best, bytes preserve cross-lingual alignment best, especially in same-script multilingual settings, and tokens support topic prediction best. These performances are not explained by sequence length alone. Short inputs can discard useful meaning, while long inputs can preserve information that compresses well. Choosing an encoding is therefore not a fixed preference for tokens, bytes, or pixels, but a rate-utility tradeoff that depends on the task, language mix, capacity regime, and compute budget.

📄 PDF Abstract BibTeX arXiv:2607.16117

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

The efficient frontiers of mean-variance portfolio rules under distribution misspecification

2021-06-19 · Andrew Paskaramoorthy, Tim Gebbie, Terence van Zyl

Mean-variance portfolio decisions that combine prediction and optimisation have been shown to have poor empirical performance. Here, we consider the performance of various shrinkage methods by their efficient frontiers u…

A Framework for Benchmarking Fairness-Utility Trade-offs in Text-to-Image Models via Pareto Frontiers

2025-08-22 · Marco N. Bochernitsan, Rodrigo C. Barros, Lucas S. Kupssinskü arxiv

Achieving fairness in text-to-image generation demands mitigating social biases without compromising visual fidelity, a challenge critical to responsible AI. Current fairness evaluation procedures for text-to-image model…

Text-to-Image Generation

Classification and Clustering of arXiv Documents, Sections, and Abstracts, Comparing Encodings of Natural and Mathematical Language

2020-05-22 · Philipp Scharpf, Moritz Schubotz, Abdou Youssef, Felix Hamborg 외

In this paper, we show how selecting and combining encodings of natural and mathematical language affect classification and clustering of documents with mathematical content. We demonstrate this by using sets of document…

ClassificationClusteringGeneral ClassificationMath+2

Semantic-Aware Guided Drone Exploration for Language-Conditioned 3D Indoor Mapping

2026-05-22 · Nitin Vegesna, Avideh Zakhor arxiv

We present Semantic-Aware Guided Exploration, SAGE, a system for open-vocabulary exploration in unknown 3D indoor environments that preserves coverage-oriented behavior while allowing semantic cues to reprioritize fronti…

Comparing Neural Network Encodings for Logic-based Explainability

2025-05-26 · Levi Cordeiro Carvalho, Saulo A. F. Oliveira, Thiago Alves Rocha

Providing explanations for the outputs of artificial neural networks (ANNs) is crucial in many contexts, such as critical systems, data protection laws and handling adversarial examples. Logic-based methods can offer exp…