paper-with-me

Papers

STACC, OOV Density and N-gram Saturation: Vicomtech's Participation in the WMT 2018 Shared Task on Parallel Corpus Filtering

2018-10-01 · WS 2018 10 · Andoni Azpeitia, Thierry Etchegoyhen, Eva Mart{\'\i}nez Garcia

We describe Vicomtech{'}s participation in the WMT 2018 Shared Task on parallel corpus filtering. We aimed to evaluate a simple approach to the task, which can efficiently process large volumes of data and can be easily deployed for new datasets in different language pairs and domains. We based our approach on STACC, an efficient and portable method for parallel sentence identification in comparable corpora. To address the specifics of the corpus filtering task, which features significant volumes of noisy data, the core method was expanded with a penalty based on the amount of unknown words in sentence pairs. Additionally, we experimented with a complementary data saturation method based on source sentence n-grams, with the goal of demoting parallel sentence pairs that do not contribute significant amounts of yet unobserved n-grams. Our approach requires no prior training and is highly efficient on the type of large datasets featured in the corpus filtering task. We achieved competitive results with this simple and portable method, ranking in the top half among competing systems overall.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationOutlier DetectionSentence

Similar Papers 제목 키워드 기반

DOCAL - Vicomtech's Participation in the WMT16 Shared Task on Bilingual Document Alignment

2016-08-01 · WS 2016 8 · Andoni Azpeitia, Thierry Etchegoyhen
Machine TranslationSemantic Textual Similarity

Supervised and Unsupervised Minimalist Quality Estimators: Vicomtech's Participation in the WMT 2018 Quality Estimation Task

2018-10-01 · WS 2018 10 · Thierry Etchegoyhen, Eva Mart{\'\i}nez Garcia, Andoni Azpeitia

We describe Vicomtech{'}s participation in the WMT 2018 shared task on quality estimation, for which we submitted minimalist quality estimators. The core of our approach is based on two simple features: lexical translati…

Language ModelingLanguage ModellingMachine TranslationSentence+1

Vicomtech at eHealth-KD Challenge 2020: Deep End-to-End Model for Entity and Relation Extraction in Medical Text

2020-09-20 · Aitor García-Pablos, Naiara Perez, Montse Cuadros and Elena Zotova

This paper describes the participation of the Vicomtech NLP team in the eHealth-KD 2020 shared task about detecting and classifying entities and relations in health-related texts written in Spanish. The proposed system …

Medical DiagnosisMedical ProcedureMulti-Label Classification Of Biomedical TextsRelation Extraction

Text Adaptation to Plain Language and Easy Read via Automatic Post-Editing Cycles

2025-09-15 · Jesús Calleja, David Ponce, Thierry Etchegoyhen arxiv

We describe Vicomtech's participation in the CLEARS challenge on text adaptation to Plain Language and Easy Read in Spanish. Our approach features automatic post-editing of different types of initial Large Language Model…

ASASVIcomtech: The Vicomtech-UGR Speech Deepfake Detection and SASV Systems for the ASVspoof5 Challenge

2024-08-19 · Juan M. Martín-Doñas, Eros Roselló, Angel M. Gomez, Aitor Álvarez 외

This paper presents the work carried out by the ASASVIcomtech team, made up of researchers from Vicomtech and University of Granada, for the ASVspoof5 Challenge. The team has participated in both Track 1 (speech deepfake…

DeepFake DetectionFace SwappingSpeaker Verification