paper-with-me

홈 › Papers

Examining the Limits of Word2Vec with Toki Pona

2026-06-15 · Daniel Zhenhan Huang, Hongchen Wu arxiv

Word2Vec's effectiveness at generating semantic embeddings has been widely validated, yet it has been tested almost exclusively on languages with large vocabulary inventories. This study examines whether Word2Vec can successfully capture semantic relationships within an extremely reduced vocabulary using data from Toki Pona, a constructed language with approximately 130 words. We sourced 1.4 million sentences (7.95 million tokens) from the Toki Pona community for training. Approximately 23% of sentences in the corpus contain non-Toki Pona tokens such as named entities, loanwords, and neologisms. To investigate whether this linguistic noise enhances or hinders performance -- a topic rarely addressed in word embedding literature -- we trained two distinct models: one retaining these incidental tokens and another filtering them out completely. Evaluation was conducted using quantitative methods measuring word proximity to semantic category centroids, automated silhouette scores via agglomerative clustering, and qualitative analysis utilizing representational similarity matrices compared against English. The results indicate that while sparse, non-core tokens do not affect the relative structure of the learned embeddings, they actually draw similar words closer together in the vector space. Importantly, Word2Vec's effectiveness depends more on distributional patterns than lexicon size even at this extreme lower bound.

📄 PDF Abstract BibTeX arXiv:2606.17299

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Computational Approach to Analyzing Language Change and Variation in the Constructed Language Toki Pona

2025-08-14 · Daniel Huang, Hyoun-A Joo arxiv

This study explores language change and variation in Toki Pona, a constructed language with approximately 120 core words. Taking a computational and corpus-based approach, the study examines features including fluid word…

Basic concepts and tools for the Toki Pona minimal and constructed language: description of the language and main issues; analysis of the vocabulary; text synthesis and syntax highlighting; Wordnet synsets

2017-12-26 · Renato Fabbri

A minimal constructed language (conlang) is useful for experiments and comfortable for making tools. The Toki Pona (TP) conlang is minimal both in the vocabulary (with only 14 letters and 124 lemmas) and in the (about) 1…

Sentence

Automated Trustworthiness Oracle Generation for Machine Learning Text Classifiers

2024-10-30 · Lam Nguyen Tung, Steven Cho, Xiaoning Du, Neelofar Neelofar 외

Machine learning (ML) for text classification has been widely used in various domains. These applications can significantly impact ethics, economics, and human behavior, raising serious concerns about trusting ML decisio…

Adversarial AttackChatbotEthicstext-classification+2

EponaV2: Driving World Model with Comprehensive Future Reasoning

2026-05-14 · Jiawei Xu, Zhizhou Zhong, Zhijian Shu, Mingkai Jia 외 arxiv

Data scaling plays a pivotal role in the pursuit of general intelligence. However, the prevailing perception-planning paradigm in autonomous driving relies heavily on expensive manual annotations to supervise trajectory …

Scene UnderstandingTrajectory PlanningAutonomous Driving

Structured Exploration and Exploitation of Label Functions for Automated Data Annotation

2026-03-28 · Phong Lam, Ha-Linh Nguyen, Thu-Trang Nguyen, Son Nguyen 외 arxiv

High-quality labeled data is critical for training reliable machine learning and deep learning models, yet manual annotation remains costly and error-prone. Programmatic labeling addresses this challenge by using label f…