paper-with-me

홈 › Papers

Building Resource-Constrained Language Agents: A Korean Case Study on Chemical Toxicity Information

2025-03-22 · Hojun Cho, Donghu Kim, Soyoung Yang, Chan Lee, Hunjoo Lee, Jaegul Choo

Language agents powered by large language models (LLMs) face significant deployment challenges in resource-constrained environments, particularly for specialized domains and less-common languages. This paper presents Tox-chat, a Korean chemical toxicity information agent devised within these limitations. We propose two key innovations: a context-efficient architecture that reduces token consumption through hierarchical section search, and a scenario-based dialogue generation methodology that effectively distills tool-using capabilities from larger models. Experimental evaluations demonstrate that our fine-tuned 8B parameter model substantially outperforms both untuned models and baseline approaches, in terms of DB faithfulness and preference. Our work offers valuable insights for researchers developing domain-specific language agents under practical constraints.

📄 PDF Abstract BibTeX arXiv:2503.17753

Code (0)

등록된 구현이 없습니다.

Tasks

Dialogue Generation

Similar Papers 제목 키워드 기반

Korean Language Modeling via Syntactic Guide

2022-06-01 · LREC 2022 6 · Hyeondey Kim, Seonhoon Kim, Inho Kang, Nojun Kwak 외

While pre-trained language models play a vital role in modern language processing tasks, but not every language can benefit from them. Most existing research on pre-trained language models focuses primarily on widely-use…

Language ModelingLanguage ModellingPOS

Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

2026-06-30 · Utsav Garg, Sungjin Hong, Jason Jung, Justin Lee 외 arxiv

We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterprise agents under practical memory and serving constraints. The mode…

Reinforcement LearningMathematical Reasoning

Discovering Lexical Gaps Using Embeddings from Multilingual LLMs

2026-05-23 · Yoonwon Jung, Aaron S. Cohen, Benjamin K. Bergen arxiv

Lexical gaps are words that do not exist in certain languages. They pose challenges for building multilingual lexical resources, for machine translation, and for cross-lingual transfer. Existing lexical gap detection rel…

Cross-Lingual TransferSemantic SimilarityMachine Translation

A Dog Is Passing Over The Jet? A Text-Generation Dataset for Korean Commonsense Reasoning and Evaluation

2022-07-01 · Findings (NAACL) 2022 7 · Jaehyung Seo, Seounghoon Lee, Chanjun Park, Yoonna Jang 외

Recent natural language understanding (NLU) research on the Korean language has been vigorously maturing with the advancements of pretrained language models and datasets. However, Korean pretrained language models still …

Language Model EvaluationLanguage ModelingLanguage ModellingNatural Language Understanding+2

Building Korean Linguistic Resource for NLU Data Generation of Banking App CS Dialog System

2022-10-01 · PANDL (COLING) 2022 10 · Jeongwoo Yoon, Onyu Park, Changhoe Hwang, Gwanghoon Yoo 외

Natural language understanding (NLU) is integral to task-oriented dialog systems, but demands a considerable amount of annotated training data to increase the coverage of diverse utterances. In this study, we report the …

Natural Language Understanding