paper-with-me

Papers

Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models

2024-10-07 · Dahyun Kim, Sukyung Lee, Yungi Kim, Attapol Rutherford, Chanjun Park

The rapid advancement of large language models (LLMs) has highlighted the need for robust evaluation frameworks that assess their core capabilities, such as reasoning, knowledge, and commonsense, leading to the inception of certain widely-used benchmark suites such as the H6 benchmark. However, these benchmark suites are primarily built for the English language, and there exists a lack thereof for under-represented languages, in terms of LLM development, such as Thai. On the other hand, developing LLMs for Thai should also include enhancing the cultural understanding as well as core capabilities. To address these dual challenge in Thai LLM research, we propose two key benchmarks: Thai-H6 and Thai Cultural and Linguistic Intelligence Benchmark (ThaiCLI). Through a thorough evaluation of various LLMs with multi-lingual capabilities, we provide a comprehensive analysis of the proposed benchmarks and how they contribute to Thai LLM development. Furthermore, we will make both the datasets and evaluation code publicly available to encourage further research and development for Thai LLMs.

📄 PDF Abstract BibTeX arXiv:2410.04795

Code (1)

UpstageAI/ThaiCLI_H6 공식 구현

Similar Papers 제목 키워드 기반

LLM for Everyone: Representing the Underrepresented in Large Language Models

2024-09-20 · Samuel Cahyawijaya

Natural language processing (NLP) has witnessed a profound impact of large language models (LLMs) that excel in a multitude of tasks. However, the limitation of LLMs in multilingual settings, particularly in underreprese…

In-Context Learning

Representing Interlingual Meaning in Lexical Databases

2023-01-22 · Fausto Giunchiglia, Gabor Bella, Nandu Chandran Nair, Yang Chi 외

In today's multilingual lexical databases, the majority of the world's languages are under-represented. Beyond a mere issue of resource incompleteness, we show that existing lexical databases have structural limitations …

Diversity

Cross-Lingual Mental Health Ontologies for Indian Languages: Bridging Patient Expression and Clinical Understanding through Explainable AI and Human-in-the-Loop Validation

2025-10-06 · Ananth Kandala, Ratna Kandala, Akshata Kishore Moharir, Niva Manchanda 외 arxiv

Mental health communication in India is linguistically fragmented, culturally diverse, and often underrepresented in clinical NLP. Current health ontologies and mental health resources are dominated by diagnostic framewo…

Do Large Language Models Truly Understand Cross-cultural Differences?

2025-12-08 · Shiwei Guo, Sihang Jiang, Qianxi He, Yanghua Xiao 외 arxiv

In recent years, large language models (LLMs) have demonstrated strong performance on multilingual tasks. Given its wide range of applications, cross-cultural understanding capability is a crucial competency. However, ex…

CultureScore: Evaluating Cultural Faithfulness in Video Generation Models

2026-06-05 · Anku Rani, Wei Dai, Shravan Nayak, Pattie Maes 외 arxiv

As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metrics, such as VideoScore, only measure vis…

Video Generation