paper-with-me

Papers

Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering

2020-10-23 · EMNLP 2021 11 · Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sagot, Djamé Seddah, Jacopo Staiano

Coupled with the availability of large scale datasets, deep learning architectures have enabled rapid progress on the Question Answering task. However, most of those datasets are in English, and the performances of state-of-the-art multilingual models are significantly lower when evaluated on non-English data. Due to high data collection costs, it is not realistic to obtain annotated data for each language one desires to support. We propose a method to improve the Cross-lingual Question Answering performance without requiring additional annotated data, leveraging Question Generation models to produce synthetic samples in a cross-lingual fashion. We show that the proposed method allows to significantly outperform the baselines trained on English data only. We report a new state-of-the-art on four multilingual datasets: MLQA, XQuAD, SQuAD-it and PIAF (fr).

📄 PDF Abstract BibTeX arXiv:2010.12643

Code (1)

microsoft/unilm 공식 구현 pytorch

Tasks

Cross-Lingual Question AnsweringData AugmentationQuestion AnsweringQuestion GenerationQuestion-Generation

Similar Papers 제목 키워드 기반

Text-guided Synthetic Geometric Augmentation for Zero-shot 3D Understanding

2025-01-16 · Kohei Torimi, Ryosuke Yamada, Daichi Otsuka, Kensho Hara 외

Zero-shot recognition models require extensive training data for generalization. However, in zero-shot 3D classification, collecting 3D data and captions is costly and laborintensive, posing a significant barrier compare…

3D ClassificationZero-shot 3D classificationZero-Shot Learning

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

2026-06-18 · Satwinder Singh, Qianli Wang, Zihan Zhong, Clarion Mendes 외 arxiv

Automatic speech recognition remains unreliable for dysarthric speech due to data scarcity and high inter-speaker variability. While synthetic data can address these gaps, traditional methods often require extensive spea…

Speech RecognitionData Augmentation

ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis

2026-03-04 · Youngwon Choi, Jinwoo Oh, Hwayeon Kim, Hyeonyu Kim arxiv

We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. While synthetic augmentation can provide linguistically rich and phonetically dive…

Data AugmentationSpeech Synthesis

Aslema at NADI 2026: Augmentation through Fewshot for SLU

2026-08-19 · Tajwaar Shafiq, Hunzalah Hassan Bhatti, Shammur Absar Chowdhury, Firoj Alam arxiv

We present Aslema, our system for NADI 2026 Shared Task 5, which consists of two subtasks: intent recognition and slot filling. We evaluate four omni LLMs in a zero-shot setting and compare them with fine-tuned models. O…

Intent RecognitionData AugmentationSlot Filling

Accelerating New Product Introduction for Visual Quality Inspection via Few-Shot Diffusion-Based Defect Synthesis

2026-04-22 · Serkan Hamdi Güğül, Kemal Levi, Burak Acar arxiv

Industrial visual inspection systems often suffer from a severe scarcity of labeled defect data, particularly during the early stages of New Product Introduction (NPI). This limitation hinders the deployment of robust su…

Representation LearningData AugmentationDomain Adaptation