paper-with-me

Papers

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

2025-05-30 · Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky, Henok Biadglign Ademtew, Alham Fikri Aji, Israel Abebe Azime, Jinheon Baek, Frederico Belcavello, Fermin Cristobal, Jan Christian Blaise Cruz, Mary Dabre, Raj Dabre, Toqeer Ehsan, Naome A Etori, Fauzan Farooqui, Jiahui Geng, Guido Ivetta, Thanmay Jayakumar, Soyeong Jeong, Zheng Wei Lim, Aishik Mandal, Sofia Martinelli, Mihail Minkov Mihaylov, Daniil Orel, Aniket Pramanick, Sukannya Purkayastha, Israfel Salazar, Haiyue Song, Tiago Timponi Torrent, Debela Desalegn Yadeta, Injy Hamed, Atnafu Lambebo Tonja, Thamar Solorio

Cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey sufficient context to capture region-specific meanings. In this work, we investigate whether images can act as cultural context in multimodal translation. We introduce CaMMT, a human-curated benchmark of over 5,800 triples of images along with parallel captions in English and regional languages. Using this dataset, we evaluate five Vision Language Models (VLMs) in text-only and text+image settings. Through automatic and human evaluations, we find that visual context generally improves translation quality, especially in handling Culturally-Specific Items (CSIs), disambiguation, and correct gender usage. By releasing CaMMT, we aim to support broader efforts in building and evaluating multimodal translation systems that are better aligned with cultural nuance and regional variation.

📄 PDF Abstract BibTeX arXiv:2505.24456

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingMachine TranslationMultimodal Machine TranslationTranslation

Similar Papers 제목 키워드 기반

TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs

2025-05-16 · Pengju Xu, Yan Wang, Shuyuan Zhang, Xuan Zhou 외

Recent progress in Multimodal Large Language Models (MLLMs) have significantly enhanced the ability of artificial intelligence systems to understand and generate multimodal content. However, these models often exhibit li…

BenchmarkingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs

2025-11-06 · Ali Faraz, Akash, Shaharukh Khan, Raja Kolla 외 arxiv

Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions about their performance in culturally diver…

Multimodal Machine TranslationVisual Question Answering

Survey of Cultural Awareness in Language Models: Text and Beyond

2024-10-30 · Siddhesh Pawar, Junyeong Park, Jiho Jin, Arnav Arora 외

Large-scale deployment of large language models (LLMs) in various applications, such as chatbots and virtual assistants, requires LLMs to be culturally sensitive to the user to ensure inclusivity. Culture has been widely…

Benchmarking

Benchmarking Machine Translation with Cultural Awareness

2023-05-23 · Binwei Yao, Ming Jiang, Tara Bobinac, Diyi Yang 외

Translating culture-related content is vital for effective cross-cultural communication. However, many culture-specific items (CSIs) often lack viable translations across languages, making it challenging to collect high-…

BenchmarkingIn-Context LearningMachine TranslationNMT+1

Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs

2025-05-21 · Hao Wang, Pinzhi Huang, Jihan Yang, Saining Xie 외

The rapid evolution of multimodal large language models (MLLMs) has significantly enhanced their real-world applications. However, achieving consistent performance across languages, especially when integrating cultural k…

BenchmarkingQuestion AnsweringVisual Question Answering