paper-with-me

Papers

DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance

2025-12-11 · Peiying Zhang, Nanxuan Zhao, Matthew Fisher, Yiran Xu, Jing Liao, Difan Liu arxiv

Recent vision-language model (VLM)-based approaches have achieved impressive results on SVG generation. However, because they generate only text and lack visual signals during decoding, they often struggle with complex semantics and fail to produce visually appealing or geometrically coherent SVGs. We introduce DuetSVG, a unified multimodal model that jointly generates image tokens and corresponding SVG tokens in an end-to-end manner. DuetSVG is trained on both image and SVG datasets. At inference, we apply a novel test-time scaling strategy that leverages the model's native visual predictions as guidance to improve SVG decoding quality. Extensive experiments show that our method outperforms existing methods, producing visually faithful, semantically aligned, and syntactically clean SVGs across a wide range of applications.

📄 PDF Abstract BibTeX arXiv:2512.10894

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding

2026-04-15 · Yibo Jiang, Tao Wu, Rui Jiang, Yehao Lu 외 arxiv

Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability significa…

Visual Reasoning

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

2026-06-25 · Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar, Abdelrahman Shaker 외 arxiv

Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward mo…

Image Generation

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models

2026-06-22 · Hongxiang Li, Hongxu Chen, Chenyang Zhu, Xiaoshuang Huang 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable success in visual understanding but remain constrained in visual generation due to the fundamental feature discrepancy between semantic perception and pix…

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models

2026-04-27 · Weixing Wang, Liudvikas Zekas, Anton Hackl, Constantin Alexander Auga 외 arxiv

Unified Multimodal Models (uMMs) aim to support both visual understanding and visual generation within a shared representation. However, existing evaluation protocols assess these two capabilities independently and do no…

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

2026-02-27 · Xiang Deng, Feng Gao, Yong Zhang, Youxin Pang 외 arxiv

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or s…

Instruction FollowingQuestion Answering