paper-with-me

홈 › Papers

The constitution of an Arabic semantic resource from a multilingual aligned corpus (Constitution d'une ressource s\'emantique arabe \`a partir de corpus multilingue align\'e) [in French]

2013-06-01 · JEPTALNRECITAL 2013 6 · Authoul Abdul Hay, Olivier Kraif
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

IdiomX A Multilingual Benchmark for Idiom Understanding, Retrieval, and Interpretation

2026-04-25 · Ayman Ali Sharara arxiv

Idiomatic expressions remain a persistent challenge for natural language processing because their meanings are often non-compositional, context-dependent, and difficult to align across languages. Existing idiom resources…

Semantic Retrieval

Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus

2026-04-14 · John E. Ortega, Rodolfo Zevallos, Fabricio Carraro arxiv

We present a unified pipeline for synthesizing high-quality Quechua and Spanish speech for the Peruvian Constitution using three state-of-the-art text-to-speech (TTS) architectures: XTTS v2, F5-TTS, and DiFlow-TTS. Our m…

Cross-Lingual Transfer

The Inception Team at NSURL-2019 Task 8: Semantic Question Similarity in Arabic

2020-04-24 · NSURL 2019 9 · Hana Al-Theiabat, Aisha Al-Sadi

This paper describes our method for the task of Semantic Question Similarity in Arabic in the workshop on NLP Solutions for Under-Resourced Languages (NSURL). The aim is to build a model that is able to detect similar se…

Question Similarity

Can Multilingual Language Models Transfer to an Unseen Dialect? A Case Study on North African Arabizi

2020-05-01 · Benjamin Muller, Benoit Sagot, Djamé Seddah

Building natural language processing systems for non standardized and low resource languages is a difficult challenge. The recent success of large-scale multilingual pretrained language models provides new modeling tools…

Dependency ParsingPart-Of-Speech TaggingTransliteration

The AMARA Corpus: Building Parallel Language Resources for the Educational Domain

2014-05-01 · LREC 2014 5 · Ahmed Abdelali, Francisco Guzman, Hassan Sajjad, Stephan Vogel

This paper presents the AMARA corpus of on-line educational content: a new parallel corpus of educational video subtitles, multilingually aligned for 20 languages, i.e. 20 monolingual corpora and 190 parallel corpora. Th…

Machine TranslationTranslation