paper-with-me

Papers

Building Monolingual Comparable and Annotated Corpora: An experimental study from a pos tagged corpus (Construire un corpus monolingue annot\'e comparable Exp\'erience \`a partir d'un corpus annot\'e morpho-syntaxiquement) [in French]

2014-07-01 · JEPTALNRECITAL 2014 7 · Hern, Nicolas ez
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Part-Of-Speech TaggingPOS

Similar Papers 제목 키워드 기반

Building Comparable Corpora for Assessing Multi-Word Term Alignment

2022-06-01 · LREC 2022 6 · Omar Adjali, Emmanuel Morin, Pierre Zweigenbaum

Recent work has demonstrated the importance of dealing with Multi-Word Terms (MWTs) in several Natural Language Processing applications. In particular, MWTs pose serious challenges for alignment and machine translation s…

Machine Translation

Reducing the Search Space for Parallel Sentences in Comparable Corpora

2020-05-01 · LREC 2020 5 · R{\'e}mi Cardon, Natalia Grabar

This paper describes and evaluates simple techniques for reducing the research space for parallel sentences in monolingual comparable corpora. Initially, when searching for parallel sentences between two comparable docum…

Sentence

Towards Unsupervised Speech-to-Text Translation

2018-11-04 · Yu-An Chung, Wei-Hung Weng, Schrasing Tong, James Glass

We present a framework for building speech-to-text translation (ST) systems using only monolingual speech and text corpora, in other words, speech utterances from a source language and independent text from a target lang…

DenoisingLanguage ModelingLanguage ModellingSpeech-to-Text+3

The CONCISUS Corpus of Event Summaries

2012-05-01 · LREC 2012 5 · Horacio Saggion, S Szasz, ra

Text summarization and information extraction systems require adaptation to new domains and languages. This adaptation usually depends on the availability of language resources such as corpora. In this paper we present a…

Text GenerationText Summarization

The MARCELL Legislative Corpus

2020-05-01 · LREC 2020 5 · Tam{\'a}s V{\'a}radi, Svetla Koeva, Martin Yamalov, Marko Tadi{\'c} 외

This article presents the current outcomes of the MARCELL CEF Telecom project aiming to collect and deeply annotate a large comparable corpus of legal documents. The MARCELL corpus includes 7 monolingual sub-corpora (Bul…

Sentence