paper-with-me

홈 › Papers

Korean L2 Vocabulary Prediction: Can a Large Annotated Corpus be Used to Train Better Models for Predicting Unknown Words?

2018-05-01 · LREC 2018 5 · Kevin Yancey, Yves Lepage
📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language AcquisitionLexical SimplificationReading ComprehensionText Simplification

Similar Papers 제목 키워드 기반

Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts

2025-10-28 · Seyoung Song, Nawon Kim, Songeun Chae, Kiwoong Park 외 arxiv

The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remaine…

GECKO: Generative Language Model for English, Code and Korean

2024-05-24 · Sungwoo Oh, Donggyu Kim

We introduce GECKO, a bilingual large language model (LLM) optimized for Korean and English, along with programming languages. GECKO is pretrained on the balanced, high-quality corpus of Korean and English employing LLaM…

kmmluLanguage ModelingLanguage ModellingLarge Language Model+1

KR-BERT: A Small-Scale Korean-Specific Language Model

2020-08-10 · Sangah Lee, Hansol Jang, Yunmee Baik, Suzi Park 외

Since the appearance of BERT, recent works including XLNet and RoBERTa utilize sentence embedding models pre-trained by large corpora and a large number of parameters. Because such models have large hardware and a huge a…

Language ModelingLanguage ModellingSentenceSentence Embedding+1

Automated Pronunciation Evaluation for Korean Toddler Speech using Speech Diarization and Self-Supervised Learning

2026-06-08 · Diane Myung-kyung Woodbridge, Jee Hyun Suh arxiv

Speech sound disorders affect approximately 44% of Korean pediatric communication disorder cases, yet automated assessment tools for Korean toddler speech remain underdeveloped. This paper presents an end-to-end pipeline…

Self-Supervised LearningRepresentation LearningSpeaker Diarization

K-Wav2vec 2.0: Automatic Speech Recognition based on Joint Decoding of Graphemes and Syllables

2021-10-11 · Jounghee Kim, Pilsung Kang

Wav2vec 2.0 is an end-to-end framework of self-supervised learning for speech representation that is successful in automatic speech recognition (ASR), but most of the work on the topic has been developed with a single la…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Cross-Lingual TransferDecoder+3