Sentence Compression for Arbitrary Languages via Multilingual Pivoting
In this paper we advocate the use of bilingual corpora which are abundantly available for training sentence compression models. Our approach borrows much of its machinery from neural machine translation and leverages bilingual pivoting: compressions are obtained by translating a source string into a foreign language and then back-translating it into the source while controlling the translation length. Our model can be trained for any language as long as a bilingual corpus is available and performs arbitrary rewrites without access to compression specific data. We release. Moss, a new parallel Multilingual Compression dataset for English, German, and French which can be used to evaluate compression models across languages and genres.
Code (1)
Tasks
Machine TranslationSentenceSentence CompressionText GenerationText SummarizationTranslationSimilar Papers 제목 키워드 기반
Image Pivoting for Learning Multilingual Multimodal Representations
In this paper we propose a model to learn multimodal multilingual representations for matching images and sentences in different languages, with the aim of advancing multilingual versions of image search and image unders…
Image DescriptionImage RetrievalSemantic Textual SimilarityHow effective is Multi-source pivoting for Translation of Low Resource Indian Languages?
Machine Translation (MT) between linguistically dissimilar languages is challenging, especially due to the scarcity of parallel corpora. Prior works suggest that pivoting through a high-resource language can help transla…
Machine TranslationSentenceTranslationMultilingual Models for Compositional Distributed Semantics
We present a novel technique for learning semantic representations, which extends the distributional hypothesis to multilingual data and joint-space embeddings. Our models leverage parallel data and learn to strongly ali…
Cross-Lingual Document ClassificationDocument ClassificationGeneral ClassificationLearning Semantic RepresentationsZero-Shot Paraphrase Generation with Multilingual Language Models
Leveraging multilingual parallel texts to automatically generate paraphrases has drawn much attention as size of high-quality paraphrase corpus is limited. Round-trip translation, also known as the pivoting method, is a …
DenoisingDiversityMachine TranslationParaphrase Generation+2IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages
Natural Language Generation (NLG) for non-English languages is hampered by the scarcity of datasets in these languages. In this paper, we present the IndicNLG Benchmark, a collection of datasets for benchmarking NLG for …
ArticlesBenchmarkingHeadline GenerationMachine Translation+6