Alphabell.
Dataset · Training data

FineWeb-Edu

A 1.3-trillion-token English web corpus filtered out of FineWeb by an educational-quality classifier, released together with the classifier and its training code. A looser 5.4-trillion-token variant (threshold 2) is also published.

The loop

Llama-3-70B-Instruct scores sampled FineWeb pages for educational value on a 0 to 5 scale, and a small embedding-based regressor trained on those LLM labels then scores the full 15-trillion-token corpus. Pages scoring below 3 are dropped (92% of the data) and the rest is used to pretrain new language models, so an existing LLM's judgment sets the training diet of the next ones.

The loop
FineWeb-Edu
Dataset · Hugging Face
  1. Llama-3-70B scores pages for educational value
  2. Regressor trained on labels filters corpus
  3. Filtered corpus pretrains new language models
  4. LLM judgment sets next training diet
  1. Llama-3-70B scores pages for educational value
  2. Regressor trained on labels filters corpus
  3. Filtered corpus pretrains new language models
  4. LLM judgment sets next training diet
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

Each stronger model can relabel the web for the next one (the dataset card cites Meta using Llama 2 to build the quality classifiers behind Llama 3), so the curator improves with every generation it helps train.

Evidence

In the paper's ablations, moving from FineWeb to FineWeb-Edu raised MMLU from 33% to 37% and ARC from 46% to 57%, and FineWeb-Edu matched the final performance of the Matrix dataset with almost 10x fewer tokens.

pretraining-datallm-as-annotatordata-filteringsynthetic-labelsopen-dataset