paper-with-me

홈 › Papers

Language-Grounded Multi-Domain Image Translation via Semantic Difference Guidance

2026-01-12 · Jongwon Ryu, Joonhyung Park, Jaeho Han, Yeong-Seok Kim, Hye-rin Kim, Sunjae Yoon, Junyeong Kim arxiv

Multi-domain image-to-image translation re quires grounding semantic differences ex pressed in natural language prompts into corresponding visual transformations, while preserving unrelated structural and seman tic content. Existing methods struggle to maintain structural integrity and provide fine grained, attribute-specific control, especially when multiple domains are involved. We propose LACE (Language-grounded Attribute Controllable Translation), built on two compo nents: (1) a GLIP-Adapter that fuses global semantics with local structural features to pre serve consistency, and (2) a Multi-Domain Control Guidance mechanism that explicitly grounds the semantic delta between source and target prompts into per-attribute translation vec tors, aligning linguistic semantics with domain level visual changes. Together, these modules enable compositional multi-domain control with independent strength modulation for each attribute. Experiments on CelebA(Dialog) and BDD100K demonstrate that LACE achieves high visual fidelity, structural preservation, and interpretable domain-specific control, surpass ing prior baselines. This positions LACE as a cross-modal content generation framework bridging language semantics and controllable visual translation.

📄 PDF Abstract BibTeX arXiv:2601.07221

Code (0)

등록된 구현이 없습니다.

Tasks

Image-to-Image Translation

Similar Papers 제목 키워드 기반

Grounded Word Sense Translation

2019-06-01 · WS 2019 6 · Chiraag Lala, Pranava Madhyastha, Lucia Specia

Recent work on visually grounded language learning has focused on broader applications of grounded representations, such as visual question answering and multimodal machine translation. In this paper we consider grounded…

Grounded language learningMachine TranslationMultimodal Machine TranslationQuestion Answering+3

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

2025-08-07 · Yufei Gao, Jiaying Fei, Nuo Chen, Ruirui Chen 외 arxiv

Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings. We argue that this failure is not merely a lingu…

Machine Translation

Lessons learned in multilingual grounded language learning

2018-09-20 · CONLL 2018 10 · Ákos Kádár, Desmond Elliott, Marc-Alexandre Côté, Grzegorz Chrupała 외

Recent work has shown how to learn better visual-semantic embeddings by leveraging image descriptions in more than one language. Here, we investigate in detail which conditions affect the performance of this type of grou…

Grounded language learningSentence

Imagination improves Multimodal Translation

2017-05-11 · IJCNLP 2017 11 · Desmond Elliott, Ákos Kádár

We decompose multimodal translation into two sub-tasks: learning to translate and learning visually grounded representations. In a multitask learning framework, translations are learned in an attention-based encoder-deco…

DecoderPredictionTranslation

A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking

2020-05-01 · LREC 2020 5 · Hideki Nakayama, Akihiro Tamura, Takashi Ninomiya

Visually-grounded natural language processing has become an important research direction in the past few years. However, majorities of the available cross-modal resources (e.g., image-caption datasets) are built in Engli…

Image CaptioningMachine TranslationMultimodal Machine TranslationTranslation