paper-with-me

홈 › Papers

An image speaks a thousand words, but can everyone listen? On image transcreation for cultural relevance

2024-04-01 · Simran Khanuja, Sathyanarayanan Ramamoorthy, Yueqi Song, Graham Neubig

Given the rise of multimedia content, human translators increasingly focus on culturally adapting not only words but also other modalities such as images to convey the same meaning. While several applications stand to benefit from this, machine translation systems remain confined to dealing with language in speech and text. In this work, we take a first step towards translating images to make them culturally relevant. First, we build three pipelines comprising state-of-the-art generative models to do the task. Next, we build a two-part evaluation dataset: i) concept: comprising 600 images that are cross-culturally coherent, focusing on a single concept per image, and ii) application: comprising 100 images curated from real-world applications. We conduct a multi-faceted human evaluation of translated images to assess for cultural relevance and meaning preservation. We find that as of today, image-editing models fail at this task, but can be improved by leveraging LLMs and retrievers in the loop. Best pipelines can only translate 5% of images for some countries in the easier concept dataset and no translation is successful for some countries in the application dataset, highlighting the challenging nature of the task. Our code and data is released here: https://github.com/simran-khanuja/image-transcreation.

📄 PDF Abstract BibTeX arXiv:2404.01247

Code (1)

simran-khanuja/image-transcreation 공식 구현 pytorch

Tasks

Machine TranslationTranslation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Requirements for Mass Adoption of Assistive Listening Technology by the General Public

2023-03-04 · Thomas B. Kaufmann, Mehdi Foroogozar, Julie Liss, Visar Berisha

Assistive listening systems (ALSs) dramatically increase speech intelligibility and reduce listening effort. It is very likely that essentially everyone, not only individuals with hearing loss, would benefit from the inc…

An Alternate Approach for Designing a Domain Specific Image Search Prototype Using Histogram

2013-11-28 · Sukanta Sinha, Rana Dattagupta, Debajyoti Mukhopadhyay

Everyone knows that thousand of words are represented by a single image. As a result image search has become a very popular mechanism for the Web searchers. Image search means, the search results are produced by the sear…

Image Retrieval

On measuring linguistic intelligence

2015-03-20 · Maxim Litvak

This work addresses the problem of measuring how many languages a person "effectively" speaks given that some of the languages are close to each other. In other words, to assign a meaningful number to her language portfo…

Can Language Models Learn to Listen?

2023-08-21 · ICCV 2023 1 · Evonne Ng, Sanjay Subramanian, Dan Klein, Angjoo Kanazawa 외

We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the speaker's words with their timestamps, ou…

Language ModelingLanguage ModellingLarge Language Model

TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar

2025-10-16 · Yinxi Li, Yuntian Deng, Pengyu Nie arxiv

Large language models (LLMs) for code rely on subword tokenizers, such as byte-pair encoding (BPE), learned from mixed natural language text and programming language code but driven by statistics rather than grammar. As …