paper-with-me

Papers

A stylometric analysis of speaker attribution from speech transcripts

2025-12-15 · Cristina Aggazzotti, Elizabeth Allyn Smith arxiv

Forensic scientists often need to identify an unknown speaker or writer in cases such as ransom calls, covert recordings, alleged suicide notes, or anonymous online communications, among many others. Speaker recognition in the speech domain usually examines phonetic or acoustic properties of a voice, and these methods can be accurate and robust under certain conditions. However, if a speaker disguises their voice or employs text-to-speech software, vocal properties may no longer be reliable, leaving only their linguistic content available for analysis. Authorship attribution methods traditionally use syntactic, semantic, and related linguistic information to identify writers of written text (authorship attribution). In this paper, we apply a content-based authorship approach to speech that has been transcribed into text, using what a speaker says to attribute speech to individuals (speaker attribution). We introduce a stylometric method, StyloSpeaker, which incorporates character, word, token, sentence, and style features from the stylometric literature on authorship, to assess whether two transcripts were produced by the same speaker. We evaluate this method on two types of transcript formatting: one approximating prescriptive written text with capitalization and punctuation and another normalized style that removes these conventions. The transcripts' conversation topics are also controlled to varying degrees. We find generally higher attribution performance on normalized transcripts, except under the strongest topic control condition, in which overall performance is highest. Finally, we compare this more explainable stylometric model to black-box neural approaches on the same data and investigate which stylistic features most effectively distinguish speakers.

📄 PDF Abstract BibTeX arXiv:2512.13667

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Recognition

Similar Papers 제목 키워드 기반

The Impact of Automatic Speech Transcription on Speaker Attribution

2025-07-11 · Cristina Aggazzotti, Matthew Wiesner, Elizabeth Allyn Smith, Nicholas Andrews arxiv

Speaker attribution from speech transcripts is the task of identifying a speaker from the transcript of their speech based on patterns in their language use. This task is especially useful when the audio is unavailable (…

Speech Recognition

Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?

2023-11-13 · Cristina Aggazzotti, Nicholas Andrews, Elizabeth Allyn Smith

Authorship verification is the task of determining if two distinct writing samples share the same author and is typically concerned with the attribution of written text. In this paper, we explore the attribution of trans…

Authorship AttributionAuthorship Verification

MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models

2024-11-27 · Thai-Binh Nguyen, Alexander Waibel

Speaker-attributed automatic speech recognition (SA-ASR) aims to transcribe speech while assigning transcripts to the corresponding speakers accurately. Existing methods often rely on complex modular systems or require e…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Stylometric Analysis of Parliamentary Speeches: Gender Dimension

2017-04-01 · WS 2017 4 · M, Justina ravickait{\.e}, Tomas Krilavi{\v{c}}ius

Relation between gender and language has been studied by many authors, however, there is still some uncertainty left regarding gender influence on language usage in the professional environment. Often, the studied data s…

Improving Quotation Attribution with Fictional Character Embeddings

2024-06-17 · Gaspard Michel, Elena V. Epure, Romain Hennequin, Christophe Cerisara

Humans naturally attribute utterances of direct speech to their speaker in literary works. When attributing quotes, we process contextual information but also access mental representations of characters that we build and…

AttributeAuthorship Verification