paper-with-me

Papers

ColLex.en: Automatically Generating and Evaluating a Full-form Lexicon for English

2014-05-01 · LREC 2014 5 · Tim vor der Br{\"u}ck, Alex Mehler, er, Zahurul Islam

The paper describes a procedure for the automatic generation of a large full-form lexicon of English. We put emphasis on two statistical methods to lexicon extension and adjustment: in terms of a letter-based HMM and in terms of a detector of spelling variants and misspellings. The resulting resource, {\textbackslash}collexen, is evaluated with respect to two tasks: text categorization and lexical coverage by example of the SUSANNE corpus and the {\textbackslash}openanc.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

FormText CategorizationText Classification

Similar Papers 제목 키워드 기반

CollEX -- A Multimodal Agentic RAG System Enabling Interactive Exploration of Scientific Collections

2025-04-10 · Florian Schneider, Narges Baba Ahmadi, Niloufar Baba Ahmadi, Iris Vogel 외

In this paper, we introduce CollEx, an innovative multimodal agentic Retrieval-Augmented Generation (RAG) system designed to enhance interactive exploration of extensive scientific collections. Given the overwhelming vol…

RAGRetrieval-augmented Generation

IndoCollex: A Testbed for Morphological Transformation of Indonesian Word Colloquialism

2021-08-01 · Findings (ACL) 2021 8 · Haryo Akbarianto Wibowo, Made Nindyatama Nityasya, Afra Feyza Akyürek, Suci Fitriany 외

Detecting Optional Arguments of Verbs

2016-05-01 · LREC 2016 5 · Andr{\'a}s Kornai, D{\'a}vid M{\'a}rk Nemeskey, G{\'a}bor Recski

We propose a novel method for detecting optional arguments of Hungarian verbs using only positive data. We introduce a custom variant of collexeme analysis that explicitly models the noise in verb frames. Our method is, …

Clustering

OARelatedWork: A Large-Scale Dataset of Related Work Sections with Full-texts from Open Access Sources

2024-05-03 · Martin Docekal, Martin Fajcik, Pavel Smrz

This paper introduces OARelatedWork, the first large-scale multi-document summarization dataset for related work generation containing whole related work sections and full-texts of cited papers. The dataset includes 94 4…

Document SummarizationExtractive SummarizationMulti-Document Summarization

Fully Convolutional Networks for Automatically Generating Image Masks to Train Mask R-CNN

2020-03-03 · Hao Wu, Jan Paul Siebert, Xiangrong Xu

This paper proposes a novel automatically generating image masks method for the state-of-the-art Mask R-CNN deep learning method. The Mask R-CNN method achieves the best results in object detection until now, however, it…

Objectobject-detectionObject DetectionSegmentation