paper-with-me

홈 › Papers

Fractal Language Modelling by Universal Sequence Maps (USM)

2025-08-08 · Jonas S Almeida, Daniel E Russ, Susana Vinga, Ines Duarte, Lee Mason, Praphulla Bhawsar, Aaron Ge, Arlindo Oliveira, Jeya Balaji Balasubramanian arxiv

Motivation: With the advent of Language Models using Transformers, popularized by ChatGPT, there is a renewed interest in exploring encoding procedures that numerically represent symbolic sequences at multiple scales and embedding dimensions. The challenge that encoding addresses is the need for mechanisms that uniquely retain contextual information about the succession of individual symbols, which can then be modeled by nonlinear formulations such as neural networks. Context: Universal Sequence Maps(USM) are iterated functions that bijectively encode symbolic sequences onto embedded numerical spaces. USM is composed of two Chaos Game Representations (CGR), iterated forwardly and backwardly, that can be projected into the frequency domain (FCGR). The corresponding USM coordinates can be used to compute a Chebyshev distance metric as well as k-mer frequencies, without having to recompute the embedded numeric coordinates, and, paradoxically, allowing for non-integers values of k. Results: This report advances the bijective fractal encoding by Universal Sequence Maps (USM) by resolving seeding biases affecting the iterated process. The resolution had two results, the first expected, the second an intriguing outcome: 1) full reconciliation of numeric positioning with sequence identity; and 2) uncovering the nature of USM as an efficient numeric process converging towards a steady state sequence embedding solution. We illustrate these results for genomic sequences because of the convenience of a planar representation defined by an alphabet with only 4 tokens (the 4 nucleotides). Nevertheless, the application to alphabet of arbitrary cardinality was found to be straightforward.

📄 PDF Abstract BibTeX arXiv:2508.06641

Code (0)

등록된 구현이 없습니다.

Tasks

Language Modelling

Similar Papers 제목 키워드 기반

Correlation Dimension of Natural Language in a Statistical Manifold

2024-05-10 · Xin Du, Kumiko Tanaka-Ishii

The correlation dimension of natural language is measured by applying the Grassberger-Procaccia algorithm to high-dimensional sequences produced by a large-scale language model. This method, previously studied only in a …

Language ModelingLanguage Modelling

Language Independent Emotion Quantification using Non linear Modelling of Speech

2021-02-11 · Uddalok Sarkar, Sayan Nag, Chirayata Bhattacharya, Shankha Sanyal 외

At present emotion extraction from speech is a very important issue due to its diverse applications. Hence, it becomes absolutely necessary to obtain models that take into consideration the speaking styles of a person, v…

Clustering

Bifractal nature of chromosome contact maps

2019-06-28 · Simone Pigolotti, Mogens H. Jensen, Yinxiu Zhan, Guido Tiana

Modern biological techniques such as Hi-C permit to measure probabilities that different chromosomal regions are close in space. These probabilities can be visualised as matrices called contact maps. In this paper, we in…

The new face of multifractality: Multi-branchedness and the phase transitions in time series of mean inter-event times

2020-04-25

Empirical time series of inter-event or waiting times are investigated using a modified Multifractal Detrended Fluctuation Analysis operating on fluctuations of mean detrended dynamics. The core of the extended multifrac…

Time SeriesTime Series Analysis

Long-range memory and multifractality in gold markets

2015-05-17

Long-range correlation and fluctuation in the gold market time series of world's two leading gold consuming countries, namely China and India, are studied. For both the market series during the period 1985-2013 we observ…

Time SeriesTime Series Analysis