paper-with-me

홈 › Papers

CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation

2024-07-01 · Yuxuan Wang, Yijun Liu, Fei Yu, Chen Huang, Kexin Li, Zhiguo Wan, Wanxiang Che

Despite the rapid development of Chinese vision-language models (VLMs), most existing Chinese vision-language (VL) datasets are constructed on Western-centric images from existing English VL datasets. The cultural bias in the images makes these datasets unsuitable for evaluating VLMs in Chinese culture. To remedy this issue, we present a new Chinese Vision- Language Understanding Evaluation (CVLUE) benchmark dataset, where the selection of object categories and images is entirely driven by Chinese native speakers, ensuring that the source images are representative of Chinese culture. The benchmark contains four distinct VL tasks ranging from image-text retrieval to visual question answering, visual grounding and visual dialogue. We present a detailed statistical analysis of CVLUE and provide a baseline performance analysis with several open-source multilingual VLMs on CVLUE and its English counterparts to reveal their performance gap between English and Chinese. Our in-depth category-level analysis reveals a lack of Chinese cultural knowledge in existing VLMs. We also find that fine-tuning on Chinese culture-related VL datasets effectively enhances VLMs' understanding of Chinese culture.

📄 PDF Abstract BibTeX arXiv:2407.01081

Code (1)

WangYuxuan93/CVLUE 공식 구현 pytorch

Tasks

Image-text RetrievalQuestion AnsweringText RetrievalVisual GroundingVisual Question Answering

Similar Papers 제목 키워드 기반

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

2022-11-02 · An Yang, Junshu Pan, Junyang Lin, Rui Men 외

The tremendous success of CLIP (Radford et al., 2021) has promoted the research and application of contrastive learning for vision-language pretraining. In this work, we construct a large-scale dataset of image-text pair…

Contrastive Learningimage-classificationImage ClassificationImage Retrieval+7

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning

2025-09-10 · Haiyang Yu, Yuchuan Wu, Fan Shi, Lei Liao 외 arxiv

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only …

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model

2025-10-13 · Chunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 외 arxiv

Fine-grained vision-language understanding requires precise alignment between visual content and linguistic descriptions, a capability that remains limited in current models, particularly in non-English settings. While m…

CCMB: A Large-scale Chinese Cross-modal Benchmark

2022-05-08 · Chunyu Xie, Heng Cai, Jincheng Li, Fanjing Kong 외

Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets a…

image-classificationImage ClassificationImage GenerationImage Retrieval+9

Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch

2025-02-24 · Xueru Wen, Jie Lou, Zichao Li, Yaojie Lu 외

Reward models (RMs) are crucial for aligning large language models (LLMs) with human preferences. However, most RM research is centered on English and relies heavily on synthetic resources, which leads to limited and les…