paper-with-me

홈 › Papers

Why are some word orders more common than others? A uniform information density account

2010-12-01 · NeurIPS 2010 12 · Luke Maurits, Dan Navarro, Amy Perfors

Languages vary widely in many ways, including their canonical word order. A basic aspect of the observed variation is the fact that some word orders are much more common than others. Although this regularity has been recognized for some time, it has not been well-explained. In this paper we offer an information-theoretic explanation for the observed word-order distribution across languages, based on the concept of Uniform Information Density (UID). We suggest that object-first languages are particularly disfavored because they are highly non-optimal if the goal is to distribute information content approximately evenly throughout a sentence, and that the rest of the observed word-order distribution is at least partially explainable in terms of UID. We support our theoretical analysis with data from child-directed speech and experimental work.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Khmer Word Search: Challenges, Solutions, and Semantic-Aware Search

2021-12-16 · Rina Buoy, Nguonly Taing, Sovisal Chenda

Search is one of the key functionalities in digital platforms and applications such as an electronic dictionary, a search engine, and an e-commerce platform. While the search function in some languages is trivial, Khmer …

A Cross-Linguistic Pressure for Uniform Information Density in Word Order

2023-06-06 · Thomas Hikaru Clark, Clara Meister, Tiago Pimentel, Michael Hahn 외

While natural languages differ widely in both canonical word order and word order flexibility, their word orders still follow shared cross-linguistic statistical patterns, often attributed to functional pressures. In the…

counterfactual

CharBench: Evaluating the Role of Tokenization in Character-Level Tasks

2025-08-04 · Omri Uzan, Yuval Pinter arxiv

Tasks that require character-level reasoning, such as counting or locating characters within words, remain challenging for contemporary language models. A common conjecture is that language models' reliance on subword un…

Neural language modeling of free word order argument structure

2019-11-30 · Charlotte Rochereau, Benoît Sagot, Emmanuel Dupoux

Neural language models trained with a predictive or masked objective have proven successful at capturing short and long distance syntactic dependencies. Here, we focus on verb argument structure in German, which has the …

Language ModelingLanguage Modelling

Fast and Fuzzy Private Set Intersection

2014-05-13 · Nicholas Kersting

Private Set Intersection (PSI) is usually implemented as a sequence of encryption rounds between pairs of users, whereas the present work implements PSI in a simpler fashion: each set only needs to be encrypted once, aft…