Parameterized Synthetic Text Generation with SimpleStories

Juan Rodriguez (Mila - Quebec Artificial Intelligence Institute) · Lennart Finke (Harvard University ETH Zürich) · Chandan Sreedhara (RespiQ) · Thomas Dooms (Universiteit Antwerpen) · Mat Allen (Independent) · Noa Nabeshima · Thomas Marshall (EleutherAI) · Dan Braun (Goodfire AI)
end-to-end training processfewest-parameter language modelgrammatical englishmodel creationmodel interpretabilitymultiple levels of abstractionopen-source.parameterizing promptssample efficiencysemantic diversitysimplestoriesstory characteristicssyntactic diversitysynthetic story datasettinystories dataset

We present SimpleStories, a large synthetic story dataset in simple language, consisting of 2 million samples each in English and Japanese. Through parameterizing prompts at multiple levels of abstraction, we achieve control over story characteristics at scale, inducing syntactic and semantic diversity. Ablations on a newly trained tiny model suite then show improved sample efficiency and model interpretability in comparison with the TinyStories dataset. We open-source all constituent parts of model creation, hoping to enable novel ways to study the end-to-end training process. As a byproduct, we move the frontier with regards to the fewest-parameter language model that outputs grammatical English.