paper-with-me

홈 › Papers

Efficient Multilingual Dialogue Processing via Translation Pipelines and Distilled Language Models

2026-01-14 · Santiago Martínez Novoa, Nicolás Rozo Fajardo, Diego Alejandro González Vargas, Nicolás Bedoya Figueroa arxiv

This paper presents team Kl33n3x's multilingual dialogue summarization and question answering system developed for the NLPAI4Health 2025 shared task. The approach employs a three-stage pipeline: forward translation from Indic languages to English, multitask text generation using a 2.55B parameter distilled language model, and reverse translation back to source languages. By leveraging knowledge distillation techniques, this work demonstrates that compact models can achieve highly competitive performance across nine languages. The system achieved strong win rates across the competition's tasks, with particularly robust performance on Marathi (86.7% QnA), Tamil (86.7% QnA), and Hindi (80.0% QnA), demonstrating the effectiveness of translation-based approaches for low-resource language processing without task-specific fine-tuning.

📄 PDF Abstract BibTeX arXiv:2601.09059

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge DistillationQuestion AnsweringText Generation

Similar Papers 제목 키워드 기반

Comparative Evaluation of Machine Translation Systems on Images with Text

2026-05-28 · Blai Puchol, Sergio Gómez González, Miguel Domingo, Francisco Casacuberta arxiv

This work presents a comparative evaluation of machine translation systems applied to images containing textual information, a task that lies at the intersection of computer vision and natural language processing. The st…

Machine TranslationText Detection

Towards Multilingual Automatic Dialogue Evaluation

2023-08-31 · John Mendonça, Alon Lavie, Isabel Trancoso

The main limiting factor in the development of robust multilingual dialogue evaluation metrics is the lack of multilingual data and the limited availability of open sourced multilingual dialogue systems. In this work, we…

Dialogue EvaluationMachine TranslationTranslation

mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset

2021-08-31 · Luiz Bonifacio, Vitor Jeronymo, Hugo Queiroz Abonizio, Israel Campiotti 외

The MS MARCO ranking dataset has been widely used for training deep learning models for IR tasks, achieving considerable effectiveness on diverse zero-shot scenarios. However, this type of resource is scarce in languages…

Information RetrievalMachine TranslationPassage RankingReranking+3

Zero-shot hashtag segmentation for multilingual sentiment analysis

2021-12-06 · Ruan Chaves Rodrigues, Marcelo Akira Inuzuka, Juliana Resplande Sant'Anna Gomes, Acquila Santos Rocha 외

Hashtag segmentation, also known as hashtag decomposition, is a common step in preprocessing pipelines for social media datasets. It usually precedes tasks such as sentiment analysis and hate speech detection. For sentim…

Feature EngineeringHate Speech DetectionMachine TranslationSegmentation+2

DeepCon: An End-to-End Multilingual Toolkit for Automatic Minuting of Multi-Party Dialogues

2022-09-01 · SIGDIAL (ACL) 2022 9 · Aakash Bhatnagar, Nidhir Bhavsar, Muskaan Singh

In this paper, we present our minuting tool DeepCon, an end-to-end toolkit for minuting the multiparty dialogues of meetings. It provides technological support for (multilingual) communication and collaboration, with a s…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationnamed-entity-recognition+6