paper-with-me

홈 › Papers

Semimage: HSV-Based Semantic Image Encoding for Disentangled Text Representation

2025-11-26 · Mohammad Zare arxiv

We propose SemImage, a novel method for representing a text document as a two-dimensional semantic image to be processed by convolutional neural networks (CNNs). In a SemImage, each word is represented as a pixel in a 2D image: rows correspond to sentences and an additional boundary row is inserted between sentences to mark semantic transitions. Each pixel is not a typical RGB value but a vector in a disentangled HSV color space, encoding different linguistic features: the Hue with two components H_cos and H_sin to account for circularity encodes the topic, Saturation encodes the sentiment, and Value encodes intensity or certainty. We enforce this disentanglement via a multi-task learning framework: a ColorMapper network maps each word embedding to the HSV space, and auxiliary supervision is applied to the Hue and Saturation channels to predict topic and sentiment labels, alongside the main task objective. The insertion of dynamically computed boundary rows between sentences yields sharp visual boundaries in the image when consecutive sentences are semantically dissimilar, effectively making paragraph breaks salient. We integrate SemImage with standard 2D CNNs (e.g., ResNet) for document classification. Experiments on multi-label datasets (with both topic and sentiment annotations) and single-label benchmarks demonstrate that SemImage can achieve competitive or better accuracy than strong text classification baselines (including BERT and hierarchical attention networks) while offering enhanced interpretability. An ablation study confirms the importance of the multi-channel HSV representation and the dynamic boundary rows. Finally, we present visualizations of SemImage that qualitatively reveal clear patterns corresponding to topic shifts and sentiment changes in the generated image, suggesting that our representation makes these linguistic features visible to both humans and machines.

📄 PDF Abstract BibTeX arXiv:2512.00088

Code (0)

등록된 구현이 없습니다.

Tasks

Document ClassificationText ClassificationMulti-Task Learning

Similar Papers 제목 키워드 기반

Towards Disentangled Representations for Human Retargeting by Multi-view Learning

2019-12-12 · Chao Yang, Xiaofeng Liu, Qingming Tang, C. -C. Jay Kuo

We study the problem of learning disentangled representations for data across multiple domains and its applications in human retargeting. Our goal is to map an input image to an identity-invariant latent representation t…

MULTI-VIEW LEARNING

Explaining Word Embeddings via Disentangled Representation

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Keng-Te Liao, Cheng-Syuan Lee, Zhong-Yu Huang, Shou-De Lin

Disentangled representations have attracted increasing attention recently. However, how to transfer the desired properties of disentanglement to word representations is unclear. In this work, we propose to transform typi…

DisentanglementWord Embeddings

Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders

2026-05-28 · Pierre-Antoine Lequeu, Camille Barboule, Benjamin Piwowarski arxiv

Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and stored remains poorly understood. Modern PE methods such as RoPE still s…

Long-Context Understanding

ALADIN-NST: Self-supervised disentangled representation learning of artistic style through Neural Style Transfer

2023-04-12 · Dan Ruta, Gemma Canet Tarres, Alexander Black, Andrew Gilbert 외

Representation learning aims to discover individual salient features of a domain in a compact and descriptive form that strongly identifies the unique characteristics of a given sample respective to its domain. Existing …

DescriptiveDisentanglementRepresentation LearningStyle Transfer

Representation Disentanglement for Multi-task Learning with application to Fetal Ultrasound

2019-08-21 · Qingjie Meng, Nick Pawlowski, Daniel Rueckert, Bernhard Kainz

One of the biggest challenges for deep learning algorithms in medical image analysis is the indiscriminate mixing of image properties, e.g. artifacts and anatomy. These entangled image properties lead to a semantically r…

AnatomyDisentanglementMedical Image AnalysisMulti-Task Learning