paper-with-me

Papers

Turkish Native Language Identification

2023-07-27 · Ahmet Yavuz Uluslu, Gerold Schneider

In this paper, we present the first application of Native Language Identification (NLI) for the Turkish language. NLI involves predicting the writer's first language by analysing their writing in different languages. While most NLI research has focused on English, our study extends its scope to Turkish. We used the recently constructed Turkish Learner Corpus and employed a combination of three syntactic features (CFG production rules, part-of-speech n-grams, and function words) with L2 texts to demonstrate their effectiveness in this task.

📄 PDF Abstract BibTeX arXiv:2307.14850

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationNative Language Identification

Similar Papers 제목 키워드 기반

Enhancing the PARSEME Turkish Corpus of Verbal Multiword Expressions

2022-06-01 · LREC (MWE) 2022 6 · Yagmur Ozturk, Najet Hadj Mohamed, Adam Lion-Bouton, Agata Savary

The PARSEME (Parsing and Multiword Expressions) project proposes multilingual corpora annotated for multiword expressions (MWEs). In this case study, we focus on the Turkish corpus of PARSEME. Turkish is an agglutinative…

SU-NLP at SemEval-2020 Task 12: Offensive Language IdentifiCation in Turkish Tweets

2020-12-01 · SEMEVAL 2020 · Anil Ozdemir, Reyyan Yeniterzi

This paper summarizes our group{'}s efforts in the offensive language identification shared task, which is organized as part of the International Workshop on Semantic Evaluation (Sem-Eval2020). Our final submission syste…

Language IdentificationWord Embeddings

Introducing cosmosGPT: Monolingual Training for Turkish Language Models

2024-04-26 · H. Toprak Kesgin, M. Kaan Yuce, Eren Dogan, M. Egemen Uzun 외

The number of open source language models that can produce Turkish is increasing day by day, as in other languages. In order to create the basic versions of such models, the training of multilingual models is usually con…

Pin\_cod\_ at SemEval-2020 Task 12: Injecting Lexicons into Bidirectional Long Short-Term Memory Networks to Detect Turkish Offensive Tweets

2020-12-01 · SEMEVAL 2020 · Pinar Arslan

This paper describes a system (pin{\_}cod{\_}) built for SemEval 2020 Task 12: OffensEval: Multilingual Offensive Language Identification in Social Media (Zampieri et al., 2020). I present the system based on the archite…

Language Identification

MSVD-Turkish: A Comprehensive Multimodal Dataset for Integrated Vision and Language Research in Turkish

2020-12-13 · Begum Citamak, Ozan Caglayan, Menekse Kuyu, Erkut Erdem 외

Automatic generation of video descriptions in natural language, also called video captioning, aims to understand the visual content of the video and produce a natural language sentence depicting the objects and actions i…

Machine TranslationMultimodal Machine TranslationSentenceTranslation+2