paper-with-me

Papers

The distribution of information content in English sentences

2016-09-24 · Shuiyuan Yu, Jin Cong, Junying Liang, Haitao Liu

Sentence is a basic linguistic unit, however, little is known about how information content is distributed across different positions of a sentence. Based on authentic language data of English, the present study calculated the entropy and other entropy-related statistics for different sentence positions. The statistics indicate a three-step staircase-shaped distribution pattern, with entropy in the initial position lower than the medial positions (positions other than the initial and final), the medial positions lower than the final position and the medial positions showing no significant difference. The results suggest that: (1) the hypotheses of Constant Entropy Rate and Uniform Information Density do not hold for the sentence-medial positions; (2) the context of a word in a sentence should not be simply defined as all the words preceding it in the same sentence; and (3) the contextual information content in a sentence does not accumulate incrementally but follows a pattern of "the whole is greater than the sum of parts".

📄 PDF Abstract BibTeX arXiv:1609.07681

Code (0)

등록된 구현이 없습니다.

Tasks

PositionSentence

Similar Papers 제목 키워드 기반

Cross-lingual Alignment of Knowledge Graph Triples with Sentences

2021-12-01 · ICON 2021 12 · Swayatta Daw, Shivprasad Sagare, Tushar Abhishek, Vikram Pudi 외

The pairing of natural language sentences with knowledge graph triples is essential for many downstream tasks like data-to-text generation, facts extraction from sentences (semantic parsing), knowledge graph completion, …

Data-to-Text GenerationKnowledge Graph CompletionNERSemantic Parsing+4

Neural Machine Translation for Sinhala-English Code-Mixed Text

2021-09-01 · RANLP 2021 9 · Archchana Kugathasan, Sagara Sumathipala

Code-mixing has become a moving method of communication among multilingual speakers. Most of the social media content of the multilingual societies are written in code-mixed text. However, most of the current translation…

DecoderMachine TranslationNMTTranslation

Neural Machine Translation System using a Content-equivalently Translated Parallel Corpus for the Newswire Translation Tasks at WAT 2019

2019-11-01 · WS 2019 11 · Hideya Mino, Hitoshi Ito, Isao Goto, Ichiro Yamada 외

This paper describes NHK and NHK Engineering System (NHK-ES){'}s submission to the newswire translation tasks of WAT 2019 in both directions of Japanese→English and English→Japanese. In addition to the JIJI Corpus that w…

Machine TranslationSentenceTranslation

Analysing the Correlation between Lexical Ambiguity and Translation Quality in a Multimodal Setting using WordNet

2022-07-01 · NAACL (ACL) 2022 7 · Ali Hatami, Paul Buitelaar, Mihael Arcan

Multimodal Neural Machine Translation is focusing on using visual information to translate sentences in the source language into the target language. The main idea is to utilise information from visual modalities to prom…

Machine TranslationSentenceTranslation

WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia

2019-07-10 · EACL 2021 2 · Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong 외

We present an approach based on multilingual sentence embeddings to automatically extract parallel sentences from the content of Wikipedia articles in 85 languages, including several dialects or low-resource languages. W…

ArticlesSentenceSentence Embeddings