paper-with-me

Papers

A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?

2025-02-19 · Ibrahim Alabdulmohsin, Andreas Steiner

Language exhibits a fractal structure in its information-theoretic complexity (i.e. bits per token), with self-similarity across scales and long-range dependence (LRD). In this work, we investigate whether large language models (LLMs) can replicate such fractal characteristics and identify conditions-such as temperature setting and prompting method-under which they may fail. Moreover, we find that the fractal parameters observed in natural language are contained within a narrow range, whereas those of LLMs' output vary widely, suggesting that fractal parameters might prove helpful in detecting a non-trivial portion of LLM-generated texts. Notably, these findings, and many others reported in this work, are robust to the choice of the architecture; e.g. Gemini 1.0 Pro, Mistral-7B and Gemma-2B. We also release a dataset comprising of over 240,000 articles generated by various LLMs (both pretrained and instruction-tuned) with different decoding temperatures and prompting methods, along with their corresponding human-generated texts. We hope that this work highlights the complex interplay between fractal properties, prompting, and statistical mimicry in LLMs, offering insights for generating, evaluating and detecting synthetic texts.

📄 PDF Abstract BibTeX arXiv:2502.14924

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

A Systems Biology view of Breast cancer via Fractal Geometry and Fractional Calculus

2025-05-12 · Abhijeet Das, Ramray Bhat, Mohit Kumar Jolly

Breast cancer (BC) is the most widespread cancer globally, yet current diagnostic and prognostic methods inadequately capture its biological complexity, despite the benefits of early detection. Cancer systems biology (SB…

Diagnostic

A multifractal-based masked auto-encoder: an application to medical images

2026-05-25 · Joao Batista Florindo, Viviane de Moura arxiv

Masked autoencoders (MAE) have shown great promise in medical image classification. However, the random masking strategy employed by traditional MAEs may overlook critical areas in medical images, where even subtle chang…

Medical Image Classification

Shape-aware Sampling Matters in the Modeling of Multi-Class Tubular Structures

2025-06-14 · Minghui Zhang, Yaoyu Liu, Xin You, Hanxiao Zhang 외

Accurate multi-class tubular modeling is critical for precise lesion localization and optimal treatment planning. Deep learning methods enable automated shape modeling by prioritizing volumetric overlap accuracy. However…

Recursive Self-Similarity in Deep Weight Spaces of Neural Architectures: A Fractal and Coarse Geometry Perspective

2025-03-18 · Ambarish Moharil, Indika Kumara, Damian Andrew Tamburri, Majid Mohammadi 외

This paper conceptualizes the Deep Weight Spaces (DWS) of neural architectures as hierarchical, fractal-like, coarse geometric structures observable at discrete integer scales through recursive dilation. We introduce a c…

A Language and Its Dimensions: Intrinsic Dimensions of Language Fractal Structures

2023-11-16 · Vasilii A. Gromov, Nikita S. Borodin, Asel S. Yerbolova

The present paper introduces a novel object of study - a language fractal structure. We hypothesize that a set of embeddings of all $n$-grams of a natural language constitutes a representative sample of this fractal set.…

AllTopological Data Analysis