paper-with-me

홈 › Papers

Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains

2024-11-22 · Yurii Paniv, Artur Kiulian, Dmytro Chaplynskyi, Mykola Khandoga, Anton Polishko, Tetiana Bas, Guillermo Gabrielli

While the evaluation of multimodal English-centric models is an active area of research with numerous benchmarks, there is a profound lack of benchmarks or evaluation suites for low- and mid-resource languages. We introduce ZNO-Vision, a comprehensive multimodal Ukrainian-centric benchmark derived from standardized university entrance examination (ZNO). The benchmark consists of over 4,300 expert-crafted questions spanning 12 academic disciplines, including mathematics, physics, chemistry, and humanities. We evaluated the performance of both open-source models and API providers, finding that only a handful of models performed above baseline. Alongside the new benchmark, we performed the first evaluation study of multimodal text generation for the Ukrainian language: we measured caption generation quality on the Multi30K-UK dataset, translated the VQA benchmark into Ukrainian, and measured performance degradation relative to original English versions. Lastly, we tested a few models from a cultural perspective on knowledge of national cuisine. We believe our work will advance multimodal generation capabilities for the Ukrainian language and our approach could be useful for other low-resource languages.

📄 PDF Abstract BibTeX arXiv:2411.14647

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingCaption Generationmultimodal generationText GenerationVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian

2025-01-12 · Mykyta Syromiatnikov, Victoria Ruvinskaya, Anastasiya Troynina

As the usage of large language models for problems outside of simple text understanding or generation increases, assessing their abilities and limitations becomes crucial. While significant progress has been made in this…

BenchmarkingMathMultiple-choice

From Bytes to Borsch: Fine-Tuning Gemma and Mistral for the Ukrainian Language Representation

2024-04-14 · Artur Kiulian, Anton Polishko, Mykola Khandoga, Oryna Chubych 외

In the rapidly advancing field of AI and NLP, generative large language models (LLMs) stand at the forefront of innovation, showcasing unparalleled abilities in text understanding and generation. However, the limited rep…

BenchmarkingDiversityLanguage ModelingLanguage Modelling

Ukrainian Visual Word Sense Disambiguation Benchmark

2026-03-24 · Yurii Laba, Yaryna Mohytych, Ivanna Rohulia, Halyna Kyryleyza 외 arxiv

This study presents a benchmark for evaluating the Visual Word Sense Disambiguation (Visual-WSD) task in Ukrainian. The main goal of the Visual-WSD task is to identify, with minimal contextual information, the most appro…

Word Sense Disambiguation

The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs

2025-09-16 · Asif Azad, Mohammad Sadat Hossain, MD Sadik Hossain Shanto, M Saifur Rahman 외 arxiv

Vision-Language Models (VLMs) have achieved remarkable progress in complex visual understanding across scientific and reasoning tasks. While performance benchmarking has advanced our understanding of these capabilities, …

Ukrainian-to-English folktale corpus: Parallel corpus creation and augmentation for machine translation in low-resource languages

2024-10-14 · AMTA 2022 9 · Olena Burda-Lassen

Folktales are linguistically very rich and culturally significant in understanding the source language. Historically, only human translation has been used for translating folklore. Therefore, the number of translated tex…

Machine TranslationSentenceTranslation