paper-with-me

홈 › Papers

Parameterized Synthetic Text Generation with SimpleStories

2025-04-12 · Lennart Finke, Chandan Sreedhara, Thomas Dooms, Mat Allen, Emerald Zhang, Juan Diego Rodriguez, Noa Nabeshima, Thomas Marshall, Dan Braun

We present SimpleStories, a large synthetic story dataset in simple language, consisting of 2 million samples each in English and Japanese. Through parameterizing prompts at multiple levels of abstraction, we achieve control over story characteristics at scale, inducing syntactic and semantic diversity. Ablations on a newly trained model suite show improved sample efficiency and model interpretability compared to the TinyStories dataset. We open-source all constituent parts of model creation, hoping to enable novel ways to study the end-to-end training process. As a byproduct, we move the frontier regarding the fewest-parameter language model that outputs grammatical natural language.

📄 PDF Abstract BibTeX arXiv:2504.09184

Code (1)

lennart-finke/simple_stories_generate 공식 구현

Tasks

DiversityLanguage ModelingLanguage ModellingText Generation

Similar Papers 제목 키워드 기반

TAGPerson: A Target-Aware Generation Pipeline for Person Re-identification

2021-12-28 · Kai Chen, Weihua Chen, Tao He, Rong Du 외

Nowadays, real data in person re-identification (ReID) task is facing privacy issues, e.g., the banned dataset DukeMTMC-ReID. Thus it becomes much harder to collect real data for ReID task. Meanwhile, the labor cost of l…

Person Re-Identification

Control3D: Towards Controllable Text-to-3D Generation

2023-11-09 · Yang Chen, Yingwei Pan, Yehao Li, Ting Yao 외

Recent remarkable advances in large-scale text-to-image diffusion models have inspired a significant breakthrough in text-to-3D generation, pursuing 3D content creation solely from a given text prompt. However, existing …

3D GenerationNeRFText to 3D

No more hard prompts: SoftSRV prompting for synthetic data generation

2024-10-21 · Giulia Desalvo, Jean-Fracois Kagy, Lazaros Karydas, Afshin Rostamizadeh 외

We present a novel soft prompt based framework, SoftSRV, that leverages a frozen pre-trained large language model (LLM) to generate targeted synthetic text sequences. Given a sample from the target distribution, our prop…

Language ModelingLanguage ModellingLarge Language ModelMath+1

FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models

2025-10-02 · Karan Dua, Hitesh Laxmichand Patel, Puneet Mittal, Ranjeet Gupta 외 arxiv

Developing document understanding models at enterprise scale requires large, diverse, and well-annotated datasets spanning a wide range of document types. However, collecting such data is prohibitively expensive due to p…

Key Information ExtractionSynthetic Data Generation

A Reparameterized Discrete Diffusion Model for Text Generation

2023-02-11 · Lin Zheng, Jianbo Yuan, Lei Yu, Lingpeng Kong

This work studies discrete diffusion probabilistic models with applications to natural language generation. We derive an alternative yet equivalent formulation of the sampling from discrete diffusion processes and levera…

modelText Generation