paper-with-me

Papers

ConlangBench: Exploring Language Knowledge and Learning in LLMs through Diverse Constructed Languages

2026-08-04 · Jinhong Jeong, Seungyeop Yi, Sangah Lee, Youngjae Yu arxiv

Constructed languages (conlangs) are intentionally created human languages with a rich tradition of linguistic creativity. Despite their potential for studying language learning in large language models (LLMs), existing conlangs remain largely underexplored in LLM research. We present ConlangBench, the first large-scale benchmark for evaluating and training LLMs on 21 existing conlangs. We collect over 21M conlang-English parallel sentence pairs (including 430K pairs across the 20 non-Esperanto conlangs) and 321K vocabulary entries. In bidirectional translation experiments, we find that models perform better on a posteriori conlangs, whose vocabularies are derived from natural languages, reflecting the design characteristics of conlangs. Training on ConlangBench also shows that models can learn all eight conlangs for which sufficient parallel corpora are available, while their learning curves vary depending on how the conlangs were created. Our findings suggest that conlangs provide a unique testbed for investigating how LLMs acquire low-resource languages.

📄 PDF Abstract BibTeX arXiv:2608.03505

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Research Trends for the Interplay between Large Language Models and Knowledge Graphs

2024-06-12 · Hanieh Khorashadizadeh, Fatima Zahra Amara, Morteza Ezzabady, Frédéric Ieng 외

This survey investigates the synergistic relationship between Large Language Models (LLMs) and Knowledge Graphs (KGs), which is crucial for advancing AI's capabilities in understanding, reasoning, and language processing…

DescriptiveKnowledge GraphsNatural Language QueriesQuestion Answering

Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models

2025-03-03 · Boyu Jia, Junzhe Zhang, Huixuan Zhang, Xiaojun Wan

In recent years, multimodal large language models (MLLMs) have achieved significant breakthroughs, enhancing understanding across text and vision. However, current MLLMs still face challenges in effectively integrating k…

Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs

2024-08-20 · Maxim Ifergan, Leshem Choshen, Roee Aharoni, Idan Szpektor 외

The veracity of a factoid is largely independent of the language it is written in. However, language models are inconsistent in their ability to answer the same factual question across languages. This raises questions ab…

knowledge editing

LLMs Are Globally Multilingual Yet Locally Monolingual: Exploring Knowledge Transfer via Language and Thought Theory

2025-05-30 · Eojin Kang, Juae Kim

Multilingual large language models (LLMs) open up new possibilities for leveraging information across languages, but their factual knowledge recall remains inconsistent depending on the input language. While previous stu…

Cross-Lingual TransferTransfer Learning

ConceptViz: A Visual Analytics Approach for Exploring Concepts in Large Language Models

2025-09-20 · Haoxuan Li, Zhen Wen, Qiqi Jiang, Chenxiao Li 외 arxiv

Large language models (LLMs) have achieved remarkable performance across a wide range of natural language tasks. Understanding how LLMs internally represent knowledge remains a significant challenge. Despite Sparse Autoe…