paper-with-me

홈 › Papers

Afro-MNIST: Synthetic generation of MNIST-style datasets for low-resource languages

2020-09-28 · Daniel J Wu, Andrew C Yang, Vinay U Prabhu

We present Afro-MNIST, a set of synthetic MNIST-style datasets for four orthographies used in Afro-Asiatic and Niger-Congo languages: Ge`ez (Ethiopic), Vai, Osmanya, and N'Ko. These datasets serve as "drop-in" replacements for MNIST. We also describe and open-source a method for synthetic MNIST-style dataset generation from single examples of each digit. These datasets can be found at https://github.com/Daniel-Wu/AfroMNIST. We hope that MNIST-style datasets will be developed for other numeral systems, and that these datasets vitalize machine learning education in underrepresented nations in the research community.

📄 PDF Abstract BibTeX arXiv:2009.13509

Code (1)

Daniel-Wu/AfroMNIST 공식 구현

Tasks

BIG-bench Machine LearningDataset Generation

Similar Papers 제목 키워드 기반

Typography-MNIST (TMNIST): an MNIST-Style Image Dataset to Categorize Glyphs and Font-Styles

2022-02-12 · Nimish Magre, Nicholas Brown

We present Typography-MNIST (TMNIST), a dataset comprising of 565,292 MNIST-style grayscale images representing 1,812 unique glyphs in varied styles of 1,355 Google-fonts. The glyph-list contains common characters from o…

OmniStyle: Filtering High Quality Style Transfer Data at Scale

2025-05-20 · CVPR 2025 1 · Ye Wang, Ruiqi Liu, Jiang Lin, Fei Liu 외

In this paper, we introduce OmniStyle-1M, a large-scale paired style transfer dataset comprising over one million content-style-stylized image triplets across 1,000 diverse style categories, each enhanced with textual de…

Style Transfer

MNIST-Gen: A Modular MNIST-Style Dataset Generation Using Hierarchical Semantics, Reinforcement Learning, and Category Theory

2025-07-16 · Pouya Shaeri, Arash Karimi, Ariane Middel arxiv

Neural networks are often benchmarked using standard datasets such as MNIST, FashionMNIST, or other variants of MNIST, which, while accessible, are limited to generic classes such as digits or clothing items. For researc…

Reinforcement Learning

OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas

2025-11-23 · James Y. Huang, Wenxuan Zhou, Nan Xu, Fei Wang 외 arxiv

The ability of Large Language Models (LLMs) to generate structured outputs that follow arbitrary schemas is crucial to a wide range of downstream tasks that require diverse structured representations of results such as i…

Information Extraction

Correlated discrete data generation using adversarial training

2018-04-03 · Shreyas Patel, Ashutosh Kakadiya, Maitrey Mehta, Raj Derasari 외

Generative Adversarial Networks (GAN) have shown great promise in tasks like synthetic image generation, image inpainting, style transfer, and anomaly detection. However, generating discrete data is a challenge. This wor…

Anomaly DetectionImage GenerationImage InpaintingStyle Transfer