paper-with-me

Papers

A Computational Approach to Language Contact -- A Case Study of Persian

2026-01-28 · Ali Basirat, Danial Namazifard, Navid Baradaran Hemmati arxiv

We investigate structural traces of language contact in the intermediate representations of a monolingual language model. Focusing on Persian (Farsi) as a historically contact-rich language, we probe the representations of a Persian-trained model when exposed to languages with varying degrees and types of contact with Persian. Our methodology quantifies the amount of linguistic information encoded in intermediate representations and assesses how this information is distributed across model components for different morphosyntactic features. The results show that universal syntactic information is largely insensitive to historical contact, whereas morphological features such as Case and Gender are strongly shaped by language-specific structure, suggesting that contact effects in monolingual language models are selective and structurally constrained.

📄 PDF Abstract BibTeX arXiv:2601.20592

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Opportunities for Persian Digital Humanities Research with Artificial Intelligence Language Models; Case Study: Forough Farrokhzad

2024-05-10 · Arash Rasti Meymandi, Zahra Hosseini, Sina Davari, Abolfazl Moshiri 외

This study explores the integration of advanced Natural Language Processing (NLP) and Artificial Intelligence (AI) techniques to analyze and interpret Persian literature, focusing on the poetry of Forough Farrokhzad. Uti…

It is not all downhill from here: Syllable Contact Law in Persian

2015-10-03 · Afshin Rahimi, Moharram Eslami, Bahram Vazirnezhad

Syllable contact pairs crosslinguistically tend to have a falling sonority slope a constraint which is called the Syllable Contact Law SCL In this study the phonotactics of syllable contacts in 4202 CVCCVC words of Persi…

All

Old wine in old glasses: Comparing computational and qualitative methods in identifying incivility on Persian Twitter during the #MahsaAmini movement

2026-02-09 · Hossein Kermani, Fatemeh Oudlajani, Pardis Yarahmadi, Hamideh Mahdi Soltani 외 arxiv

This paper compares three approaches to detecting incivility in Persian tweets: human qualitative coding, supervised learning with ParsBERT, and large language models (ChatGPT). Using 47,278 tweets from the #MahsaAmini m…

Dialectal Layers in West Iranian: a Hierarchical Dirichlet Process Approach to Linguistic Relationships

2020-01-13 · Chundra Aroor Cathcart

This paper addresses a series of complex and unresolved issues in the historical phonology of West Iranian languages. The West Iranian languages (Persian, Kurdish, Balochi, and other languages) display a high degree of n…

Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation

2024-12-17 · Samin Mahdizadeh Sani, Pouya Sadeghi, Thuy-Trang Vu, Yadollah Yaghoobzadeh 외

Large language models (LLMs) have made great progress in classification and text generation tasks. However, they are mainly trained on English data and often struggle with low-resource languages. In this study, we explor…

Classificationparameter-efficient fine-tuningText GenerationTransfer Learning