paper-with-me

Papers

Document Author Classification Using Parsed Language Structure

2024-03-20 · Todd K Moon, Jacob H. Gunther

Over the years there has been ongoing interest in detecting authorship of a text based on statistical properties of the text, such as by using occurrence rates of noncontextual words. In previous work, these techniques have been used, for example, to determine authorship of all of \emph{The Federalist Papers}. Such methods may be useful in more modern times to detect fake or AI authorship. Progress in statistical natural language parsers introduces the possibility of using grammatical structure to detect authorship. In this paper we explore a new possibility for detecting authorship using grammatical structural information extracted using a statistical natural language parser. This paper provides a proof of concept, testing author classification based on grammatical structure on a set of "proof texts," The Federalist Papers and Sanditon which have been as test cases in previous authorship detection studies. Several features extracted from the statistical natural language parser were explored: all subtrees of some depth from any level; rooted subtrees of some depth, part of speech, and part of speech by level in the parse tree. It was found to be helpful to project the features into a lower dimensional space. Statistical experiments on these documents demonstrate that information from a statistical parser can, in fact, assist in distinguishing authors.

📄 PDF Abstract BibTeX arXiv:2403.13253

Code (0)

등록된 구현이 없습니다.

Tasks

Classification

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

The BDCam\~oes Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology

2020-05-01 · LREC 2020 5 · Sara Grilo, M{\'a}rcia Bolrinha, Jo{\~a}o Silva, Rui Vaz 외

This paper presents the BDCam{\~o}es Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese that in its inaugural version includes close to 4 million words from over 200 complet…

Genre classification

HERMES: a multi-agent framework for structured knowledge extraction from ultra-long documents in geoscience

2026-08-14 · Ziqi Song, Zongyuan Xiang, James G. Ogg, Bruce S. Lieberman 외 arxiv

Authoritative scientific knowledge in geoscience remains largely trapped in legacy monographs and historical literature, where unstructured text and complex layouts hinder computational access. We introduce HERMES, a sca…

The Semantic Scholar Open Data Platform

2023-01-24 · Rodney Kinney, Chloe Anastasiades, Russell Authur, Iz Beltagy 외

The volume of scientific output is creating an urgent need for automated tools to help scientists keep up with developments in their field. Semantic Scholar (S2) is an open data platform and website aimed at accelerating…

graph construction

Dolphin-v2: Universal Document Parsing via Scalable Anchor Prompting

2026-02-05 · Hao Feng, Wei Shi, Ke Zhang, Xiang Fei 외 arxiv

Document parsing has garnered widespread attention as vision-language models (VLMs) advance OCR capabilities. However, the field remains fragmented across dozens of specialized models with varying strengths, forcing user…

Attribute Extraction

Stylometry Analysis of Multi-authored Documents for Authorship and Author Style Change Detection

2024-01-12 · Muhammad Tayyab Zamir, Muhammad Asif Ayub, Asma Gul, Nasir Ahmad 외

In recent years, the increasing use of Artificial Intelligence based text generation tools has posed new challenges in document provenance, authentication, and authorship detection. However, advancements in stylometry ha…

Change DetectionStyle change detectionText Generation