paper-with-me

Papers

"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models

2026-05-29 · Mihai Masala, Marius Leordeanu, Mihai Dascalu, Traian Rebedea arxiv

Vision-Language Models (VLMs) largely follow the text-only LLM trajectory, excelling on English benchmarks but sharply degrading on low-resource languages, where neither large-scale image-text corpora nor culturally grounded evaluations exist. We present a systematic study of building a language-specific VLM for Romanian, covering the full pipeline from data construction to architectural choices. We translate established English VLM training and evaluation corpora into Romanian, applying machine translation to textual annotations and to in-image text, preserving visual grounding while adapting the textual content. Using this data, we train and ablate a series of VLMs to isolate the contribution of (i) vision backbones of varying scale and pretraining, (ii) language backbones from multilingual to Romanian-adapted LLMs, and (iii) OCR-style image-text data. We further curate HoraVQA, a culturally native evaluation set grounded in Romanian everyday scenes. Romanian-adapted VLMs consistently outperform their same-sized counterparts and, across all evaluated benchmarks, even surpass models from the next larger size category.

📄 PDF Abstract BibTeX arXiv:2605.31401

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationVisual Grounding

Similar Papers 제목 키워드 기반

"Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions

2024-06-26 · Mihai Masala, Denis C. Ilie-Ablachim, Alexandru Dima, Dragos Corlatescu 외

In recent years, Large Language Models (LLMs) have achieved almost human-like performance on various tasks. While some LLMs have been trained on multilingual data, most of the training data is in English; hence, their pe…

Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models

2025-12-16 · George-Andrei Dima, Răzvan-Alexandru Smădu, Dumitru-Clementin Cercel arxiv

Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. We translate the widely known Flickr30K da…

Visual Question Answering

The Enemy from Within: A Study of Political Delegitimization Discourse in Israeli Political Speech

2025-08-21 · Naama Rivlin-Angert, Guy Mor-Lan arxiv

We present the first large-scale computational study of political delegitimization discourse (PDD), defined as symbolic attacks on the normative validity of political entities. We curate and manually annotate a novel Heb…

A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models

2026-05-29 · Iosif Tsangko, Andreas Triantafyllopoulos, George Margetis, Ioana Crihana 외 arxiv

Blind and low-vision (BLV) audiences remain underserved by visual art descriptions, particularly across languages and in museum settings where privacy and intellectual-property constraints may favour small on-premise vis…

Short Video Uprising: How #BlackLivesMatter Content on TikTok Challenges the Protest Paradigm

2022-06-20 · Yanru Jiang, Xin Jin, Qinhao Deng

This study uses TikTok (N = 8,173) to examine how short-form video platforms challenge the protest paradigm in the recent Black Lives Matter movement. A computer-mediated visual analysis, computer vision, is employed to …

Descriptive