Learning Distributional Token Representations from Visual Features
In this study, we compare token representations constructed from visual features (i.e., pixels) with standard lookup-based embeddings. Our goal is to gain insight about the challenges of encoding a text representation from low-level features, e.g. from characters or pixels. We focus on Chinese, which{---}as a logographic language{---}has properties that make a representation via visual features challenging and interesting. To train and evaluate different models for the token representation, we chose the task of character-based neural machine translation (NMT) from Chinese to English. We found that a token representation computed only from visual features can achieve competitive results to lookup embeddings. However, we also show different strengths and weaknesses in the models{'} performance in a part-of-speech tagging task and also a semantic similarity task. In summary, we show that it is possible to achieve a \textit{text representation} only from pixels. We hope that this is a useful stepping stone for future studies that exclusively rely on visual input, or aim at exploiting visual features of written language.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationNMTPart-Of-Speech TaggingRepresentation LearningSemantic SimilaritySemantic Textual SimilarityTranslationSimilar Papers 제목 키워드 기반
OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
Multi-modal large language models (MLLMs) achieve strong visual-language reasoning but suffer from high inference cost due to redundant visual tokens. Recent work explores visual token pruning to accelerate inference, wh…
Building Static Embeddings from Contextual Ones: Is It Useful for Building Distributional Thesauri?
While contextual language models are now dominant in the field of Natural Language Processing, the representations they build at the token level are not always suitable for all uses. In this article, we propose a new met…
Semantic SimilaritySemantic Textual SimilarityVocal Bursts Type PredictionHigh-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
Recent large-scale vision-language models (VLMs) have shown remarkable text-to-image generation capabilities, yet their visual fidelity remains constrained by the discrete image tokenization, which poses a major challeng…
Text-to-Image GenerationGKR: Bridging the Gap between Symbolic/structural and Distributional Meaning Representations
Three broad approaches have been attempted to combine distributional and structural/symbolic aspects to construct meaning representations: a) injecting linguistic features into distributional representations, b) injectin…
Neural Text Classification by Jointly Learning to Cluster and Align
Distributional text clustering delivers semantically informative representations and captures the relevance between each word and semantic clustering centroids. We extend the neural text clustering approach to text class…
ClassificationClusteringGeneral Classificationtext-classification+3