paper-with-me

홈 › Papers

Examining Language Modeling Assumptions Using an Annotated Literary Dialect Corpus

2024-10-03 · Craig Messner, Tom Lippincott

We present a dataset of 19th century American literary orthovariant tokens with a novel layer of human-annotated dialect group tags designed to serve as the basis for computational experiments exploring literarily meaningful orthographic variation. We perform an initial broad set of experiments over this dataset using both token (BERT) and character (CANINE)-level contextual language models. We find indications that the "dialect effect" produced by intentional orthographic variation employs multiple linguistic channels, and that these channels are able to be surfaced to varied degrees given particular language modelling assumptions. Specifically, we find evidence showing that choice of tokenization scheme meaningfully impact the type of orthographic information a model is able to surface.

📄 PDF Abstract BibTeX arXiv:2410.02674

Code (1)

comp-int-hum/orthography-embedding-clustering 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

American 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Jointly Learning Author and Annotated Character N-gram Embeddings: A Case Study in Literary Text

2019-09-01 · RANLP 2019 9 · Suraj Maharjan, Deepthi Mave, Prasha Shrestha, Manuel Montes 외

An author{'}s way of presenting a story through his/her writing style has a great impact on whether the story will be liked by readers or not. In this paper, we learn representations for authors of literary texts togethe…

Authorship AttributionGenre classificationLanguage ModelingLanguage Modelling

An annotated dataset of literary entities

2019-06-01 · NAACL 2019 6 · David Bamman, Sejal Popat, Sheng Shen

We present a new dataset comprised of 210,532 tokens evenly drawn from 100 different English-language literary texts annotated for ACE entity categories (person, location, geo-political entity, facility, organization, an…

Annotating Characters in Literary Corpora: A Scheme, the CHARLES Tool, and an Annotated Novel

2016-05-01 · LREC 2016 5 · Hardik Vala, Stefan Dimitrov, David Jurgens, Andrew Piper 외

Characters form the focus of various studies of literary works, including social network analysis, archetype induction, and plot comparison. The recent rise in the computational modelling of literary works has produced a…

The Project Dialogism Novel Corpus: A Dataset for Quotation Attribution in Literary Texts

2022-04-12 · LREC 2022 6 · Krishnapriya Vishnubhotla, Adam Hammond, Graeme Hirst

We present the Project Dialogism Novel Corpus, or PDNC, an annotated dataset of quotations for English literary texts. PDNC contains annotations for 35,978 quotations across 22 full-length novels, and is by an order of m…

Referring Expression

LiteraryTaste: A Preference Dataset for Creative Writing Personalization

2025-11-12 · John Joon Young Chung, Vishakh Padmakumar, Melissa Roemmele, Yi Wang 외 arxiv

People have different creative writing preferences, and large language models (LLMs) for these tasks can benefit from adapting to each user's preferences. However, these models are often trained over a dataset that consi…