236 GB
505 files
Updated 4 days ago
Name
Size
figures
.gitattributes2.46 kB
xet
README.md11.9 kB
xet
synth_001.parquet472 MB
xet
synth_002.parquet474 MB
xet
synth_003.parquet474 MB
xet
synth_004.parquet473 MB
xet
synth_005.parquet473 MB
xet
synth_006.parquet473 MB
xet
synth_007.parquet474 MB
xet
synth_008.parquet473 MB
xet
synth_009.parquet474 MB
xet
synth_010.parquet473 MB
xet
synth_011.parquet473 MB
xet
synth_012.parquet473 MB
xet
synth_013.parquet473 MB
xet
synth_014.parquet473 MB
xet
synth_015.parquet473 MB
xet
synth_016.parquet473 MB
xet
synth_017.parquet474 MB
xet
synth_018.parquet473 MB
xet
synth_019.parquet473 MB
xet
synth_020.parquet474 MB
xet
synth_021.parquet474 MB
xet
synth_022.parquet472 MB
xet
synth_023.parquet473 MB
xet
synth_024.parquet472 MB
xet
synth_025.parquet473 MB
xet
synth_026.parquet473 MB
xet
synth_027.parquet473 MB
xet
synth_028.parquet474 MB
xet
synth_029.parquet473 MB
xet
synth_030.parquet472 MB
xet
synth_031.parquet474 MB
xet
synth_032.parquet474 MB
xet
synth_033.parquet474 MB
xet
synth_034.parquet473 MB
xet
synth_035.parquet474 MB
xet
synth_036.parquet472 MB
xet
synth_037.parquet474 MB
xet
synth_038.parquet472 MB
xet
synth_039.parquet473 MB
xet
synth_040.parquet473 MB
xet
synth_041.parquet472 MB
xet
synth_042.parquet473 MB
xet
synth_043.parquet472 MB
xet
synth_044.parquet472 MB
xet
synth_045.parquet473 MB
xet
synth_046.parquet474 MB
xet
synth_047.parquet474 MB
xet
synth_048.parquet472 MB
xet
synth_049.parquet473 MB
xet
synth_050.parquet473 MB
xet
synth_051.parquet471 MB
xet
synth_052.parquet475 MB
xet
synth_053.parquet475 MB
xet
synth_054.parquet475 MB
xet
synth_055.parquet474 MB
xet
synth_056.parquet473 MB
xet
synth_057.parquet474 MB
xet
synth_058.parquet473 MB
xet
synth_059.parquet472 MB
xet
synth_060.parquet473 MB
xet
synth_061.parquet473 MB
xet
synth_062.parquet471 MB
xet
synth_063.parquet473 MB
xet
synth_064.parquet472 MB
xet
synth_065.parquet473 MB
xet
synth_066.parquet473 MB
xet
synth_067.parquet473 MB
xet
synth_068.parquet472 MB
xet
synth_069.parquet472 MB
xet
synth_070.parquet473 MB
xet
synth_071.parquet473 MB
xet
synth_072.parquet474 MB
xet
synth_073.parquet473 MB
xet
synth_074.parquet473 MB
xet
synth_075.parquet474 MB
xet
synth_076.parquet472 MB
xet
synth_077.parquet473 MB
xet
synth_078.parquet473 MB
xet
synth_079.parquet473 MB
xet
synth_080.parquet474 MB
xet
synth_081.parquet473 MB
xet
synth_082.parquet474 MB
xet
synth_083.parquet473 MB
xet
synth_084.parquet472 MB
xet
synth_085.parquet474 MB
xet
synth_086.parquet473 MB
xet
synth_087.parquet473 MB
xet
synth_088.parquet472 MB
xet
synth_089.parquet472 MB
xet
synth_090.parquet473 MB
xet
synth_091.parquet473 MB
xet
synth_092.parquet472 MB
xet
synth_093.parquet472 MB
xet
synth_094.parquet473 MB
xet
synth_095.parquet473 MB
xet
synth_096.parquet473 MB
xet
synth_097.parquet472 MB
xet
README.md

SYNTH

Pleias

Blog announcement

SYNTH is the first open generalist synthetic dataset for training small reasoning model end-to-end, jointly released by Pleias and the AI Alliance.

SYNTH includes 79,648,272 individual text samples, comprising over 41 billion words (about 75 billion tokens with Pleias tokenizer). It is based on the amplification of 58,698 articles from Wikipedia and made possible thanks to the Structured Wikipedia dataset from Wikimedia Enterprise.

SYNTH differs from existing open synthetic dataset in being:

  • fully open based on seed text under open license (CC-By-SA) and generated with models allowing for output reuse. This means that SYNTH can be universally release and serve as a basis for further reproducible synthetic pipelines.
  • state of the art for small models below 350 million parameters. We release two models train on SYNTH achieving current best results for size range on MMLU and other standard evaluation metrics.
  • data efficient with best results attained with only 100-200 billions tokens trained on SYNTH.
  • reasoning by design with all generated answers being accompanied with intermediary reasoning traces in an entirely new syntax.
  • diverse comprising a wide range of exercises that cover many use cases of small models: retrieval-augmented generation, creative writing, arithmetics, information extraction, etc.
  • multilingual with about 20% of all texts in other languages than English, for now limited on European languages (German, French, Spanish, Italian, Polish, Dutch, Latin).

SYNTH is not only the name of a dataset but an initiative for open synthetic data and open environment led by AI Alliance and Pleias that aims to address the critical gap in open-source AI development by creating a cutting-edge, open-source data corpus for training sovereign AI models and advanced AI agents.

Dataset Design

Amplified knowledge

At its core, SYNTH is a fully synthetic and engineered corpus derived from a sample of 50,000 pages curated by the Wikipedia community. Throughout the past two decades, thousands of contributors selected a collection of core topics that every encyclopedia should have, Wikipedia:Vital articles. It’s a concentric selection starting at level 1 (10 articles) up to level 5 (50,000 articles). SYNTH includes as its starting point all articles featured in level 5.

SYNTH further expands on this core nucleus with three additional seed collections:

  • specialized articles: following on intermediary evaluation, we added 8,698 articles to reinforce coverage of specific fields like law, medicine, chemistry. Selection was based on category tree search analysis and aimed to fill remaining holes in knowledge coverage from Wikipedia:Vital articles.
  • textbooks: wikipedia articles are focused on encyclopedic knowledge but lag on practical knowledge and how to, which happens to be the focus of another Wikimedia project, Wikibooks. For now we included 3,727 pages on cooking from Wikibooks but looking forward to expand on additional forms of experential knowledge (gardening, language acquisition, etc.)
  • recent/self knowledge: we incorporated a small sample of 130 texts hand-crafted internally to expand model familiarity with recent events, self-awareness about training condition and general research information on AI. This collection has been highly amplified.

This content act as the SYNTH memory base and has been amplified at a minimum 100 times (about 10000 times for recent/self knowledge). Our amplification strategy relies on a new synthetic pipeline, partly inspired by RAG applications:

  • Selection of individual consistent sections from the original articles (about 250,000 for the core sample of 50,000 pages).
  • Generation of queries with randomized constraints for style variation, query outcomes. It proved especially determining to have enough negative queries to reinforce world knowledge and limit hallucinations.

Synthetic exercises

The approach has been originally explored by Pleias for retrieval-augmented generation. It has been extended to virtually most of the expected use case of small reasoning models:

  • arithmetics
  • creative writing We injected randomized constraints

Dataset Details

Dataset Description

  • Curated by: Wikipedia community (Wikipedia:Vital Articles) and Pleias.
  • Funded by [optional]: Pleias
  • Shared by [optional]: Pleias
  • Language(s) (NLP): English (80%), French, German, Italian, Spanish, Polish, Dutch and Latin.
  • License:

Dataset Sources [optional]

While the final training data is fully synthetic, it relied on seeds collected from three data sources:

  • Structured Wikipedia: We used directly the dumps made available by the Wikimedia Foundation.
  • Wikibooks: extracted through the official Wikimedia API.
  • Internal documents from Pleias: mostly model-self documentation and few updated information.

Uses

The dataset aims to support data efficient training of small reasoning model. It provide a generalist, self-sufficient collection of multilingual amplified encyclopedic texts along with synthetic reasoning traces, as well as synthetic tasks that reinforce most of the expected capacities of small model.

In contrast with organic pretraining dataset, SYNTH allows for fast convergence to the existing SOTA (about 100 billion tokens). Furthermore, SYNTH is fully releasable, only use sourced text under free license.

Overall, SYNTH aims to support an emerging ecosystem of small training model by providing a reusable generalist foundational dataset.

Direct Use

Direct use include:

  • Pretraining of small reasoning models: the dataset is sufficient to elicit most expected capacities in small models.
  • Mid-training/fine-tuning of existing models: we already led successful experiments with Pleias-350m.
  • Research/explainability experiment: with its openness and data efficiency, SYNTH should be an ideal resource for research on model memorization or skill acquisition.

Out-of-Scope Use

Current out-of-scope use include:

  • Code generation: we intently excluded code data from SYNTH as this would require the development of specific synthetic pipeline.
  • Global multilingual support: SYNTH only claims support from our current list of eight languages.
  • Training of large models: the difficulty of synthetic exercises has been calibrated for models smaller than a few billion parameters.

Yet, SYNTH is a live resources and we intend to cover some of these use cases in future releases.

Dataset Structure

Field Type Description
synth_id string Unique synthetic identifier for each generated sample.
language string Language of the text sample (e.g., "en", "fr", "it", "es", "de", "pl", "nl", "la").
exercise string Type of synthetic exercise (e.g., reasoning, writing, retrieval, arithmetic). Describes the synthetic task context.
model string Finetuned model used to generate the synthetic sample
query string Backtranslated query.
query_seed_url string URL of the Wikipedia or Wikibooks section that served as the seed for query generation.
query_seed_text string Extend text used as seed for query generation.
additional_seed_url string Optional additional URL(s) used as supplementary seed
seed_license string License of the seed text (most of the time "CC-BY-SA 4.0").
constraints string Generation constraints applied to answer generation. Varies depending on the exercise
script string Internal template or script identifier defining the structure of the synthetic exercise.
synthetic_reasoning string Generated reasoning draft.
synthetic_answer string Final generated answer or output corresponding to the query.
words int64 Word count of the full generated text sample (query + draft + answer)

Dataset Creation

Curation Rationale

SYNTH is structured around a “memory core”, the Wikipedia vital articles.. Throughout the past two decades, thousands of contributors selected a collection of core topics that every encyclopedia should have: it’s a concentric selection starting at level 1 (10 articles) up to level 5 (50,000 articles). SYNTH includes as its starting point all articles featured in level 5. It further expands on this selection by increasing coverage of more specialized domains (physics, chemistry, law…) through targeted expansion of wikidata knowledge graphs.

Source Data

The 58,698 Wikipedia articles were collected thanks to ''Structured Wikipedia'', a project from Wikimedia Enterprise that parsed directly rendered Wikipedia articles in html. Structured Wikipedia fixed most of the formatting issues linked with the mediawiki syntax and provides a clean, section-based version of all Wikipedia pages.

We additionally extracted 3,000 cooking recipes from Wikibooks using the standard API method from Wikimedia.

Data Collection and Processing

Who are the source data producers?

The main sourced dataset used for synthetic amplification was curated by the English Wikipedia communities throughout nearly 2 decades. Rationale for selection are available on the relevant talk pages of Wikipedia:Vital articles.

The selection reflect similar bias for "canon" general knowledge in English-speaking countries than major LLM benchmarks like MMLU (drawn from high school exams).

Personal and Sensitive Information

The dataset only contain encyclopedic information on highly well-known historical people. No PII curation was needed.

Bias, Risks, and Limitations

The dataset was created from a collection of 50,000 Wikipedia articles curated by the community (Wikipedia:Vital Articles).

On top of the well documented structural bias in Wikipedia contribution and editing, the selection has been intently made from the perspective of western US/European culture.

Due to systematic Wikipedia grounding, the data presents a very low risk of toxic or problematic content, as well as poor or highly hallucinated information.

Total size
236 GB
Files
505
Last updated
Oct 2
Pre-warmed CDN
US EU US EU

Contributors