paper-with-me

Papers

Optimizing Korean-Centric LLMs via Token Pruning

2026-04-17 · Hoyeol Kim, Hyeonwoo Kim arxiv

This paper presents a systematic benchmark of state-of-the-art multilingual large language models (LLMs) adapted via token pruning - a compression technique that eliminates tokens and embedding parameters corresponding to languages irrelevant to the target application. Focusing on Korean-centric natural language processing (NLP) tasks, we evaluate architectures including Qwen3, Gemma-3, Llama-3, and Aya across three vocabulary configurations: Original, English-Korean (EnKo), and English-Korean-Chinese (EnKoZh). Performance is assessed using established benchmarks for general aptitude, cultural literacy, instruction following, and machine translation. Our findings indicate that token pruning significantly improves generation stability by eliminating language confusion, and in the case of machine translation, frequently enhances performance on Korean-specific tasks. While instruction-following capabilities display architecture-dependent variance linked to latent cross-lingual representations, the significant reduction in vocabulary size validates token pruning as a highly effective optimization strategy for memory-constrained, domain-specific deployments, despite modest gains in inference latency.

📄 PDF Abstract BibTeX arXiv:2604.16235

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction FollowingMachine Translation

Similar Papers 제목 키워드 기반

Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language Models

2024-02-22 · Seungduk Kim, Seungtaek Choi, Myeongho Jeong

This report introduces \texttt{EEVE-Korean-v1.0}, a Korean adaptation of large language models that exhibit remarkable capabilities across English and Korean text understanding. Building on recent highly capable but Engl…

Mi:dm 2.0 Korea-centric Bilingual Language Models

2026-01-14 · Donghoon Shin, Sejung Lee, Soonmin Bae, Hwijung Ryu 외 arxiv

We introduce Mi:dm 2.0, a bilingual large language model (LLM) specifically engineered to advance Korea-centric AI. This model goes beyond Korean text processing by integrating the values, reasoning patterns, and commons…

Synthetic Data Generation

Beyond Intermediate States: Explaining Visual Redundancy through Language

2025-03-26 · Dingchen Yang, Bowen Cao, Anran Zhang, Weibo Gu 외

Multi-modal Large Langue Models (MLLMs) often process thousands of visual tokens, which consume a significant portion of the context window and impose a substantial computational burden. Prior work has empirically explor…

OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models

2026-05-18 · Morunliu Yang, Ruotao Xu, Le Li, Yue Wang 외 arxiv

Omnimodal large language models (OmniLLMs) have recently gained increasing attention for unified audio-video understanding. However, processing long multimodal token sequences introduces substantial computational overhea…

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

2025-07-12 · Wencan Huang, Daizong Liu, Wei Hu arxiv

While 3D Multi-modal Large Language Models (MLLMs) demonstrate remarkable scene understanding capabilities, their practical deployment faces critical challenges due to computational inefficiency. The key bottleneck stems…

Scene Understanding