paper-with-me

홈 › Papers

LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings

2025-12-08 · Sebastian Sztwiertnia, Felix Friedrich, Kristian Kersting, Patrick Schramowski, Björn Deiseroth arxiv

Pre-training decoder-only language models relies on vast amounts of high-quality data, yet the availability of such data is increasingly reaching its limits. While metadata is commonly used to create and curate these datasets, its potential as a direct training signal remains under-explored. We challenge this status quo and propose LIME (Linguistic Metadata Embeddings), a method that enriches token embeddings with metadata capturing syntax, semantics, and contextual properties. LIME substantially improves pre-training efficiency. Specifically, it adapts up to 56% faster to the training data distribution, while introducing only 0.01% additional parameters at negligible compute overhead. Beyond efficiency, LIME improves tokenization, leading to remarkably stronger language modeling capabilities and generative task performance. These benefits persist across model scales (500M to 2B). In addition, we develop a variant with shifted metadata, LIME+1, that can guide token generation. Given prior metadata for the next token, LIME+1 improves reasoning performance by up to 38% and arithmetic accuracy by up to 35%.

📄 PDF Abstract BibTeX arXiv:2512.07522

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LIME: Towards a Metadata Module for Ontolex

2013-09-01 · WS 2013 9 · Manuel Fiorelli, Maria Teresa Pazienza, Arm Stellato, o

Standardizing a Component Metadata Infrastructure

2012-05-01 · LREC 2012 5 · Daan Broeder, Dieter van Uytvanck, Maria Gavrilidou, Thorsten Trippel 외

This paper describes the status of the standardization efforts of a Component Metadata approach for describing Language Resources with metadata. Different linguistic and Language {\&} Technology communities as CLARIN, ME…

A Psycho-linguistic Analysis of BitChute

2022-04-17 · Benjamin D. Horne

In order to better support researchers, journalist, and practitioners in their use of the MeLa-BitChute dataset for exploration and investigative reporting, we provide new psycho-linguistic metadata for the videos, comme…

GraphDerm: Fusing Imaging, Physical Scale, and Metadata in a Population-Graph Classifier for Dermoscopic Lesions

2025-09-14 · Mehdi Yousefzadeh, Parsa Esfahanian, Sara Rashidifar, Hossein Salahshoor Gavalan 외 arxiv

Introduction. Dermoscopy aids melanoma triage, yet image-only AI often ignores patient metadata (age, sex, site) and the physical scale needed for geometric analysis. We present GraphDerm, a population-graph framework th…

Lesion SegmentationNode Classification

Mobilizing Metadata: Open Data Kit (ODK) for Language Resource Development in East Africa

2020-05-01 · LREC 2020 5 · Richard Griscom

Linguistic fieldworkers collect and archive metadata as part of the language resources (LRs) that they create, but they often work in resource-constrained environments that prevent them from using computers for data entr…