paper-with-me

Papers

Cross-linguistic Comparison of Linguistic Feature Encoding in BERT Models for Typologically Different Languages

2022-07-01 · NAACL (SIGTYP) 2022 7 · Yulia Otmakhova, Karin Verspoor, Jey Han Lau

Though recently there have been an increased interest in how pre-trained language models encode different linguistic features, there is still a lack of systematic comparison between languages with different morphology and syntax. In this paper, using BERT as an example of a pre-trained model, we compare how three typologically different languages (English, Korean, and Russian) encode morphology and syntax features across different layers. In particular, we contrast languages which differ in a particular aspect, such as flexibility of word order, head directionality, morphological type, presence of grammatical gender, and morphological richness, across four different tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Metric-Learning Encoding Models Identify Processing Profiles of Linguistic Features in BERT's Representations

2024-02-18 · Louis Jalouzot, Robin Sobczyk, Bastien Lhopitallier, Jeanne Salle 외

We introduce Metric-Learning Encoding Models (MLEMs) as a new approach to understand how neural systems represent the theoretical features of the objects they process. As a proof-of-concept, we apply MLEMs to neural repr…

Metric Learning

What Makes Two Language Models Think Alike?

2024-06-18 · Jeanne Salle, Louis Jalouzot, Nur Lan, Emmanuel Chemla 외

Do architectural differences significantly affect the way models represent and process language? We propose a new approach, based on metric-learning encoding models (MLEMs), as a first step to answer this question. The a…

MambaMetric Learning

Probing LLMs for Joint Encoding of Linguistic Categories

2023-10-28 · Giulio Starace, Konstantinos Papakostas, Rochelle Choenni, Apostolos Panagiotopoulos 외

Large Language Models (LLMs) exhibit impressive performance on a range of NLP tasks, due to the general-purpose linguistic knowledge acquired during pretraining. Existing model interpretability research (Tenney et al., 2…

POS

Document-aware Positional Encoding and Linguistic-guided Encoding for Abstractive Multi-document Summarization

2022-09-13 · Congbo Ma, Wei Emma Zhang, Pitawelayalage Dasun Dileepa Pitawela, Yutong Qu 외

One key challenge in multi-document summarization is to capture the relations among input documents that distinguish between single document summarization (SDS) and multi-document summarization (MDS). Few existing MDS wo…

Document SummarizationMulti-Document Summarization

Subword-Based Comparative Linguistics across 242 Languages Using Wikipedia Glottosets

2026-01-26 · Iaroslav Chelombitko, Mika Hämäläinen, Aleksey Komissarov arxiv

We present a large-scale comparative study of 242 Latin and Cyrillic-script languages using subword-based methodologies. By constructing 'glottosets' from Wikipedia lexicons, we introduce a framework for simultaneous cro…