paper-with-me

Papers

MIST: a Large-Scale Annotated Resource and Neural Models for Functions of Modal Verbs in English Scientific Text

2022-12-14 · Sophie Henning, Nicole Macher, Stefan Grünewald, Annemarie Friedrich

Modal verbs (e.g., "can", "should", or "must") occur highly frequently in scientific articles. Decoding their function is not straightforward: they are often used for hedging, but they may also denote abilities and restrictions. Understanding their meaning is important for various NLP tasks such as writing assistance or accurate information extraction from scientific text. To foster research on the usage of modals in this genre, we introduce the MIST (Modals In Scientific Text) dataset, which contains 3737 modal instances in five scientific domains annotated for their semantic, pragmatic, or rhetorical function. We systematically evaluate a set of competitive neural architectures on MIST. Transfer experiments reveal that leveraging non-scientific data is of limited benefit for modeling the distinctions in MIST. Our corpus analysis provides evidence that scientific communities differ in their usage of modal verbs, yet, classifiers trained on scientific data generalize to some extent to unseen scientific domains.

📄 PDF Abstract BibTeX arXiv:2212.07156

Code (1)

boschresearch/mist_emnlp_findings2022 공식 구현 pytorch

Tasks

Articles

Similar Papers 제목 키워드 기반

The N2 corpus: A semantically annotated collection of Islamist extremist stories

2014-05-01 · LREC 2014 5 · Mark Finlayson, Jeffry Halverson, Steven Corman

We describe the N2 (Narrative Networks) Corpus, a new language resource. The corpus is unique in three important ways. First, every text in the corpus is a story, which is in contrast to other language resources that may…

Translation

ArPoMeme: An Annotated Arabic Multimodal Dataset for Political Ideology and Polarization

2026-05-20 · Wajdi Zaghouani, Kais Attia, Md. Rafiul Biswas, Fadhl Eryani arxiv

Memes have become a prominent medium of political communication in the Arab world, reflecting how humor, imagery, and text interact to express ideological and cultural positions. Despite the centrality of memes to online…

EnzChemRED, a rich enzyme chemistry relation extraction dataset

2024-04-22 · Po-Ting Lai, Elisabeth Coudert, Lucila Aimo, Kristian Axelsen 외

Expert curation is essential to capture knowledge of enzyme functions from the scientific literature in FAIR open knowledgebases but cannot keep pace with the rate of new discoveries and new publications. In this work we…

Benchmarkingnamed-entity-recognitionNamed Entity RecognitionNER+2

Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection

2025-10-19 · Akif Islam, Mohd Ruhul Ameen arxiv

Bengali social media platforms have witnessed a sharp increase in hate speech, disproportionately affecting women and adolescents. While datasets such as BD-SHS provide a basis for structured evaluation, most prior appro…

parameter-efficient fine-tuningHate Speech Detection

The Index Thomisticus Treebank as Linked Data in the LiLa Knowledge Base

2022-06-01 · LREC 2022 6 · Francesco Mambrini, Marco Passarotti, Giovanni Moretti, Matteo Pellegrini

Although the Universal Dependencies initiative today allows for cross-linguistically consistent annotation of morphology and syntax in treebanks for several languages, syntactically annotated corpora are not yet interope…