Parallel Sentence Retrieval From Comparable Corpora for Biomedical Text Simplification
Parallel sentences provide semantically similar information which can vary on a given dimension, such as language or register. Parallel sentences with register variation (like expert and non-expert documents) can be exploited for the automatic text simplification. The aim of automatic text simplification is to better access and understand a given information. In the biomedical field, simplification may permit patients to understand medical and health texts. Yet, there is currently no such available resources. We propose to exploit comparable corpora which are distinguished by their registers (specialized and simplified versions) to detect and align parallel sentences. These corpora are in French and are related to the biomedical area. Manually created reference data show 0.76 inter-annotator agreement. Our purpose is to state whether a given pair of specialized and simplified sentences is parallel and can be aligned or not. We treat this task as binary classification (alignment/non-alignment). We perform experiments with a controlled ratio of imbalance and on the highly unbalanced real data. Our results show that the method we present here can be used to automatically generate a corpus of parallel sentences from our comparable corpus.
Code (0)
등록된 구현이 없습니다.
Tasks
Binary ClassificationRetrievalSentenceSentence RetrievalText SimplificationSimilar Papers 제목 키워드 기반
French Biomedical Text Simplification: When Small and Precise Helps
We present experiments on biomedical text simplification in French. We use two kinds of corpora {--} parallel sentences extracted from existing health comparable corpora in French and WikiLarge corpus translated from Eng…
Text SimplificationBuilding Subject-aligned Comparable Corpora and Mining it for Truly Parallel Sentence Pairs
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our methodology for mining such…
ArticlesMachine TranslationRetrievalSentence+1Harvesting comparable corpora and mining them for equivalent bilingual sentences using statistical classification and analogy- based heuristics
Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for minin…
General ClassificationMachine TranslationRetrievalTranslationIdentification of Parallel Sentences in Comparable Monolingual Corpora from Different Registers
Parallel aligned sentences provide useful information for different NLP applications. Yet, this kind of data is seldom available, especially for languages other than English. We propose to exploit comparable corpora in F…
Information RetrievalMachine TranslationSTSText SimplificationPEXACC: A Parallel Sentence Mining Algorithm from Comparable Corpora
Extracting parallel data from comparable corpora in order to enrich existing statistical translation models is an avenue that attracted a lot of research in recent years. There are experiments that convincingly show how …
Information RetrievalMachine TranslationSentenceTranslation