paper-with-me

홈 › Papers

Extending the Vocabulary of Fictional Languages using Neural Networks

2022-01-18 · Thomas Zacharias, Ashutosh Taklikar, Raja Giryes

Fictional languages have become increasingly popular over the recent years appearing in novels, movies, TV shows, comics, and video games. While some of these fictional languages have a complete vocabulary, most do not. We propose a deep learning solution to the problem. Using style transfer and machine translation tools, we generate new words for a given target fictional language, while maintaining the style of its creator, hence extending this language vocabulary.

📄 PDF Abstract BibTeX arXiv:2201.07288

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationStyle TransferTranslation

Similar Papers 제목 키워드 기반

Statistical laws and linguistics differ in naturalistic video and fictional conversations

2025-12-19 · Ashley M. A. Fehr, Calla G. Beauregard, Julia Witte Zimmerman, Katie Ekström 외 arxiv

Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations…

FAME: Fictional Actors for Multilingual Erasure

2025-12-17 · Claudio Savelli, Moreno La Quatra, Alkis Koudounas, Flavio Giobergia arxiv

LLMs trained on web-scale data raise concerns about privacy and the right to be forgotten. To address these issues, Machine Unlearning provides techniques to remove specific information from trained models without retrai…

Extending the Subwording Model of Multilingual Pretrained Models for New Languages

2022-11-29 · Kenji Imamura, Eiichiro Sumita

Multilingual pretrained models are effective for machine translation and cross-lingual processing because they contain multiple languages in one model. However, they are pretrained after their tokenizers are fixed; there…

Machine TranslationTranslation

LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language

2024-05-13 · Cagri Toraman

Despite advancements in English-dominant generative large language models, further development is needed for low-resource languages to enhance global accessibility. The primary methods for representing these languages ar…

Constructing Cross-lingual Consumer Health Vocabulary with Word-Embedding from Comparable User Generated Content

2022-06-23 · Chia-Hsuan Chang, Lei Wang, Christopher C. Yang

The online health community (OHC) is the primary channel for laypeople to share health information. To analyze the health consumer-generated content (HCGC) from the OHCs, identifying the colloquial medical expressions us…