Japanese Sentence Compression with a Large Training Dataset
In English, high-quality sentence compression models by deleting words have been trained on automatically created large training datasets. We work on Japanese sentence compression by a similar approach. To create a large Japanese training dataset, a method of creating English training dataset is modified based on the characteristics of the Japanese language. The created dataset is used to train Japanese sentence compression models based on the recurrent neural network.
Code (0)
등록된 구현이 없습니다.
Tasks
SentenceSentence CompressionSimilar Papers 제목 키워드 기반
Flexible Japanese Sentence Compression by Relaxing Unit Constraints
Japanese SimCSE Technical Report
We report the development of Japanese SimCSE, Japanese sentence embedding models fine-tuned with SimCSE. Since there is a lack of sentence embedding models for Japanese that can be used as a baseline in sentence embeddin…
SentenceSentence EmbeddingSentence-EmbeddingSentence EmbeddingsJCSE: Contrastive Learning of Japanese Sentence Embeddings and Its Applications
Contrastive learning is widely used for sentence representation learning. Despite this prevalence, most studies have focused exclusively on English and few concern domain adaptation for domain-specific downstream tasks, …
Contrastive LearningDomain AdaptationInformation RetrievalLanguage Modelling+7Domain Adaptation for Japanese Sentence Embeddings with Contrastive Learning based on Synthetic Sentence Generation
Several backbone models pre-trained on general domain datasets can encode a sentence into a widely useful embedding. Such sentence embeddings can be further enhanced by domain adaptation that adapts a backbone model to a…
Contrastive LearningDomain AdaptationSemantic Textual SimilaritySentence+2A Corpus for English-Japanese Multimodal Neural Machine Translation with Comparable Sentences
Multimodal neural machine translation (NMT) has become an increasingly important area of research over the years because additional modalities, such as image data, can provide more context to textual data. Furthermore, t…
Image CaptioningMachine TranslationNMTSentence+1