Unsupervised Image to Sequence Translation with Canvas-Drawer Networks
Encoding images as a series of high-level constructs, such as brush strokes or discrete shapes, can often be key to both human and machine understanding. In many cases, however, data is only available in pixel form. We present a method for generating images directly in a high-level domain (e.g. brush strokes), without the need for real pairwise data. Specifically, we train a "canvas" network to imitate the mapping of high-level constructs to pixels, followed by a high-level "drawing" network which is optimized through this mapping towards solving a desired image recreation or translation task. We successfully discover sequential vector representations of symbols, large sketches, and 3D objects, utilizing only pixel data. We display applications of our method in image segmentation, and present several ablation studies comparing various configurations.
Code (1)
Tasks
Image SegmentationSemantic SegmentationTranslationSimilar Papers 제목 키워드 기반
CoDraw: Collaborative Drawing as a Testbed for Grounded Goal-driven Communication
In this work, we propose a goal-driven collaborative task that combines language, perception, and action. Specifically, we develop a Collaborative image-Drawing game between two agents, called CoDraw. Our game is grounde…
Imitation LearningTranslation Canvas: An Explainable Interface to Pinpoint and Analyze Translation Systems
With the rapid advancement of machine translation research, evaluation toolkits have become essential for benchmarking system progress. Tools like COMET and SacreBLEU offer single quality score assessments that are effec…
BenchmarkingMachine TranslationTranslationCanvasVAE: Learning to Generate Vector Graphic Documents
Vector graphic documents present visual elements in a resolution free, compact format and are often seen in creative applications. In this work, we attempt to learn a generative model of vector graphic documents. We defi…
Length-Adaptive Decoding for Masked Diffusion Machine Translation
Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising. Existing masked diffusion…
Machine TranslationMulti-task Sequence to Sequence Learning
Sequence to sequence learning has recently emerged as a new paradigm in supervised learning. To date, most of its applications focused on only one task and not much work explored this framework for multiple tasks. This p…
Caption GenerationDecoderMachine TranslationMulti-Task Learning+1