paper-with-me

Papers

Modality Conversion of Handwritten Patterns by Cross Variational Autoencoders

2019-06-14 · Taichi Sumi, Brian Kenji Iwana, Hideaki Hayashi, Seiichi Uchida

This research attempts to construct a network that can convert online and offline handwritten characters to each other. The proposed network consists of two Variational Auto-Encoders (VAEs) with a shared latent space. The VAEs are trained to generate online and offline handwritten Latin characters simultaneously. In this way, we create a cross-modal VAE (Cross-VAE). During training, the proposed Cross-VAE is trained to minimize the reconstruction loss of the two modalities, the distribution loss of the two VAEs, and a novel third loss called the space sharing loss. This third, space sharing loss is used to encourage the modalities to share the same latent space by calculating the distance between the latent variables. Through the proposed method mutual conversion of online and offline handwritten characters is possible. In this paper, we demonstrate the performance of the Cross-VAE through qualitative and quantitative analysis.

📄 PDF Abstract BibTeX arXiv:1906.06142

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

USD Coin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

Cross-Domain Image Conversion by CycleDM

2024-03-05 · Sho Shimotsumagari, Shumpei Takezaki, Daichi Haraguchi, Seiichi Uchida

The purpose of this paper is to enable the conversion between machine-printed character images (i.e., font images) and handwritten character images through machine learning. For this purpose, we propose a novel unpaired …

Denoising

Indic Handwritten Script Identification using Offline-Online Multimodal Deep Network

2018-02-23 · Ayan Kumar Bhunia, Subham Mukherjee, Aneeshan Sain, Ankan Kumar Bhunia 외

In this paper, we propose a novel approach of word-level Indic script identification using only character-level data in training stage. The advantages of using character level data for training have been outlined in sect…

Learning Modality-Invariant Representations for Speech and Images

2017-12-11 · Kenneth Leidal, David Harwath, James Glass

In this paper, we explore the unsupervised learning of a semantic embedding space for co-occurring sensory inputs. Specifically, we focus on the task of learning a semantic vector space for both spoken and handwritten di…

Information RetrievalRetrievalSemantic SimilaritySemantic Textual Similarity+1

Survey: Transformer-based Models in Data Modality Conversion

2024-08-08 · Elyas Rashno, Amir Eskandari, Aman Anand, Farhana Zulkernine

Transformers have made significant strides across various artificial intelligence domains, including natural language processing, computer vision, and audio processing. This success has naturally garnered considerable in…

Survey

A Change of Heart: Improving Speech Emotion Recognition through Speech-to-Text Modality Conversion

2023-07-21 · Zeinab Sadat Taghavi, Ali Satvaty, Hossein Sameti

Speech Emotion Recognition (SER) is a challenging task. In this paper, we introduce a modality conversion concept aimed at enhancing emotion recognition performance on the MELD dataset. We assess our approach through two…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion RecognitionSpeech Emotion Recognition+3