paper-with-me

홈 › Papers

negativas: a prototype for searching and classifying sentential negation in speech data

2025-04-05 · Túlio Sousa de Gois, Paloma Batista Cardoso

Negation is a universal feature of natural languages. In Brazilian Portuguese, the most commonly used negation particle is n\~ao, which can scope over nouns or verbs. When it scopes over a verb, n\~ao can occur in three positions: pre-verbal (NEG1), double negation (NEG2), or post-verbal (NEG3), e.g., n\~ao gosto, n\~ao gosto n\~ao, gosto n\~ao ("I do not like it"). From a variationist perspective, these structures are different forms of expressing negation. Pragmatically, they serve distinct communicative functions, such as politeness and modal evaluation. Despite their grammatical acceptability, these forms differ in frequency. NEG1 dominates across Brazilian regions, while NEG2 and NEG3 appear more rarely, suggesting its use is contextually restricted. This low-frequency challenges research, often resulting in subjective, non-generalizable interpretations of verbal negation with n\~ao. To address this, we developed negativas, a tool for automatically identifying NEG1, NEG2, and NEG3 in transcribed data. The tool's development involved four stages: i) analyzing a dataset of 22 interviews from the Falares Sergipanos database, annotated by three linguists, ii) creating a code using natural language processing (NLP) techniques, iii) running the tool, iv) evaluating accuracy. Inter-annotator consistency, measured using Fleiss' Kappa, was moderate (0.57). The tool identified 3,338 instances of n\~ao, classifying 2,085 as NEG1, NEG2, or NEG3, achieving a 93% success rate. However, negativas has limitations. NEG1 accounted for 91.5% of identified structures, while NEG2 and NEG3 represented 7.2% and 1.2%, respectively. The tool struggled with NEG2, sometimes misclassifying instances as overlapping structures (NEG1/NEG2/NEG3).

📄 PDF Abstract BibTeX arXiv:2504.04275

Code (1)

tuliosg/negativas 공식 구현

Tasks

Negation

Similar Papers 제목 키워드 기반

Exploiting weak-supervision for classifying Non-Sentential Utterances in Mandarin Conversations

2020-10-01 · PACLIC 2020 10 · Xin-Yi Chen, Laurent Prévot

When a sentence does not introduce a discourse entity, Transformer-based models still sometimes refer to it

2022-05-06 · NAACL 2022 7 · Sebastian Schuster, Tal Linzen

Understanding longer narratives or participating in conversations requires tracking of discourse entities that have been mentioned. Indefinite noun phrases (NPs), such as 'a dog', frequently introduce discourse entities …

NegationSentence

When a sentence does not introduce a discourse entity, Transformer-based models still often refer to it

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Understanding longer narratives or participating in conversations requires tracking of discourse entities that have been mentioned. Indefinite noun phrases, such as 'a dog', frequently introduce discourse entities but th…

NegationSentence

Saying What You're Looking For: Linguistics Meets Video Search

2013-09-20 · Andrei Barbu, N. Siddharth, Jeffrey Mark Siskind

We present an approach to searching large video corpora for video clips which depict a natural-language query in the form of a sentence. This approach uses compositional semantics to encode subtle meaning that is lost in…

object-detectionObject DetectionSentence

Analyzing and Mitigating Negation Artifacts using Data Augmentation for Improving ELECTRA-Small Model Accuracy

2025-11-09 · Mojtaba Noghabaei arxiv

Pre-trained models for natural language inference (NLI) often achieve high performance on benchmark datasets by using spurious correlations, or dataset artifacts, rather than understanding language touches such as negati…

Natural Language InferenceData Augmentation