paper-with-me

Papers

VisTai: Benchmarking Vision-Language Models for Traditional Chinese in Taiwan

2025-03-13 · Zhi Rui Tam, Ya-Ting Pai, Yen-Wei Lee

In this paper, we propose a comprehensive evaluation benchmark for Visual Language Models (VLM) in Traditional Chinese. Our evaluation suite, the first of its kind, contains two complementary components: (1) VisTai-MCQ, a collection of manually curated exam multi-choice questions from 21 academic subjects designed to test the broad knowledge and reasoning capabilities of VLMs; and (2) VisTai-Dialogue, an open dialogue benchmark comprising 131 image-question pairs manually created to evaluate VLMs' ability in free-form dialogue generation within Taiwanese cultural contexts. These benchmarks address a critical gap in the evaluation landscape, where existing benchmarks predominantly focus on English or Simplified Chinese, neglecting the unique linguistic and cultural aspects of Traditional Chinese used in regions like Taiwan and Hong Kong. Our analysis reveals significant performance differences across various VLMs and highlights specific challenges in processing Traditional Chinese visual content.

📄 PDF Abstract BibTeX arXiv:2503.10427

Code (1)

TMMMU-Benchmark/evaluation 공식 구현 pytorch

Tasks

BenchmarkingDialogue Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning

2025-09-10 · Haiyang Yu, Yuchuan Wu, Fan Shi, Lei Liao 외 arxiv

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only …

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan

2025-08-02 · Jui-Ming Yao, Bing-Cheng Xie, Sheng-Wei Peng, Hao-Yuan Chen 외 arxiv

Multimodal Large Language Models (MLLMs) process visual, acoustic, and textual inputs, addressing the limitations of single-modality LLMs. However, existing benchmarks often overlook tri-modal evaluation in Traditional C…

Question Answering

Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese

2025-05-28 · Hanjia Lyu, Jiebo Luo, Jian Kang, Allison Koenecke

While the capabilities of Large Language Models (LLMs) have been studied in both Simplified and Traditional Chinese, it is yet unclear whether LLMs exhibit differential performance when prompted in these two variants of …

Benchmarking

The Breeze 2 Herd of Models: Traditional Chinese LLMs Based on Llama with Vision-Aware and Function-Calling Capabilities

2025-01-23 · MediaTek Research, :, Chan-Jan Hsu, Chia-Sheng Liu 외

Llama-Breeze2 (hereinafter referred to as Breeze2) is a suite of advanced multi-modal language models, available in 3B and 8B parameter configurations, specifically designed to enhance Traditional Chinese language repres…

General KnowledgeInstruction FollowingLanguage ModelingLanguage Modelling

Benchmarking Machine Translation on Chinese Social Media Texts

2026-01-30 · Kaiyan Zhao, Zheyong Xie, Zhongtao Miao, Xinze Lyu 외 arxiv

The prevalence of rapidly evolving slang, neologisms, and highly stylized expressions in informal user-generated text, particularly on Chinese social media, poses significant challenges for Machine Translation (MT) bench…

Machine Translation