paper-with-me

홈 › Papers

Higher Criticism for Discriminating Word-Frequency Tables and Testing Authorship

2019-10-30 · Alon Kipnis

We adapt the Higher Criticism (HC) goodness-of-fit test to measure the closeness between word-frequency tables. We apply this measure to authorship attribution challenges, where the goal is to identify the author of a document using other documents whose authorship is known. The method is simple yet performs well without handcrafting and tuning; reporting accuracy at the state of the art level in various current challenges. As an inherent side effect, the HC calculation identifies a subset of discriminating words. In practice, the identified words have low variance across documents belonging to a corpus of homogeneous authorship. We conclude that in comparing the similarity of a new document and a corpus of a single author, HC is mostly affected by words characteristic of the author and is relatively unaffected by topic structure.

📄 PDF Abstract BibTeX arXiv:1911.01208

Code (2)

alonkipnis/AuthorshipAttribution 공식 구현
alonkipnis/HCAuthorship 공식 구현

Tasks

Authorship Attribution

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

Discriminating between Similar Languages Using a Combination of Typed and Untyped Character N-grams and Words

2017-04-01 · WS 2017 4 · Helena Gomez, Ilia Markov, Jorge Baptista, Grigori Sidorov 외

This paper presents the cic{\_}ualg{'}s system that took part in the Discriminating between Similar Languages (DSL) shared task, held at the VarDial 2017 Workshop. This year{'}s task aims at identifying 14 languages acro…

General ClassificationInformation RetrievalMachine Translation

Building a language evolution tree based on word vector combination model

2018-10-04 · Zhu Gao, Yanhui Jiang, Junhui Gao

In this paper, we try to explore the evolution of language through case calculations. First, we chose the novels of eleven British writers from 1400 to 2005 and found the corresponding works; Then, we use the natural lan…

Clustering

Scientific Table Search Using Keyword Queries

2017-07-11 · Gao Kyle Yingkai, Callan Jamie

Tables are common and important in scientific documents, yet most text-based document search systems do not capture structures and semantics specific to tables. How to bridge different types of mismatch between keywords …

Table Search

AnlamVer: Semantic Model Evaluation Dataset for Turkish - Word Similarity and Relatedness

2018-08-01 · COLING 2018 8 · G{\"o}khan Ercan, Olcay Taner Y{\i}ld{\i}z

In this paper, we present AnlamVer, which is a semantic model evaluation dataset for Turkish designed to evaluate word similarity and word relatedness tasks while discriminating those two relations from each other. Our d…

Word EmbeddingsWord Similarity

Modeling Image Quantization Tradeoffs for Optimal Compression

2021-12-14 · Johnathan Chiu

All Lossy compression algorithms employ similar compression schemes -- frequency domain transform followed by quantization and lossless encoding schemes. They target tradeoffs by quantizating high frequency data to incre…

Quantization