paper-with-me

Papers

RuSentEval: Linguistic Source, Encoder Force!

2021-02-28 · EACL (BSNLP) 2021 4 · Vladislav Mikhailov, Ekaterina Taktasheva, Elina Sigdel, Ekaterina Artemova

The success of pre-trained transformer language models has brought a great deal of interest on how these models work, and what they learn about language. However, prior research in the field is mainly devoted to English, and little is known regarding other languages. To this end, we introduce RuSentEval, an enhanced set of 14 probing tasks for Russian, including ones that have not been explored yet. We apply a combination of complementary probing methods to explore the distribution of various linguistic properties in five multilingual transformers for two typologically contrasting languages -- Russian and English. Our results provide intriguing findings that contradict the common understanding of how linguistic knowledge is represented, and demonstrate that some properties are learned in a similar manner despite the language differences.

📄 PDF Abstract BibTeX arXiv:2103.00573

Code (1)

RussianNLP/rusenteval 공식 구현 pytorch

Similar Papers 제목 키워드 기반

BERnaT: Basque Encoders for Representing Natural Textual Diversity

2025-12-03 · Ekhi Azurmendi, Joseba Fernandez de Landa, Jaione Bengoetxea, Maite Heredia 외 arxiv

Language models depend on massive text corpora that are often filtered for quality, a process that can unintentionally exclude non-standard linguistic varieties, reduce model robustness and reinforce representational bia…

Natural Language Understanding

Linguistic Information in Neural Semantic Parsing with Multiple Encoders

2019-05-01 · WS 2019 5 · Rik van Noord, Antonio Toral, Johan Bos

Recently, sequence-to-sequence models have achieved impressive performance on a number of semantic parsing tasks. However, they often do not exploit available linguistic resources, while these, when employed correctly, a…

DRS ParsingMachine TranslationSemantic ParsingTranslation

Efficiently Fusing Pretrained Acoustic and Linguistic Encoders for Low-resource Speech Recognition

2021-01-17 · Cheng Yi, Shiyu Zhou, Bo Xu

End-to-end models have achieved impressive results on the task of automatic speech recognition (ASR). For low-resource ASR tasks, however, labeled data can hardly satisfy the demand of end-to-end models. Self-supervised …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Improving Multilingual Neural Machine Translation by Utilizing Semantic and Linguistic Features

2024-08-02 · Mengyu Bu, Shuhao Gu, Yang Feng

The many-to-many multilingual neural machine translation can be regarded as the process of integrating semantic features from the source sentences and linguistic features from the target sentences. To enhance zero-shot t…

DecoderMachine TranslationText GenerationTransfer Learning+1

Many-to-Many Voice Conversion based Feature Disentanglement using Variational Autoencoder

2021-07-11 · Manh Luong, Viet Anh Tran

Voice conversion is a challenging task which transforms the voice characteristics of a source speaker to a target speaker without changing linguistic content. Recently, there have been many works on many-to-many Voice Co…

DisentanglementVoice Conversion