FineWeb-Edu
A 1.3-trillion-token English web corpus filtered out of FineWeb by an educational-quality classifier, released together with the classifier and its training code. A looser 5.4-trillion-token variant (threshold 2) is also published.
Llama-3-70B-Instruct scores sampled FineWeb pages for educational value on a 0 to 5 scale, and a small embedding-based regressor trained on those LLM labels then scores the full 15-trillion-token corpus. Pages scoring below 3 are dropped (92% of the data) and the rest is used to pretrain new language models, so an existing LLM's judgment sets the training diet of the next ones.
- Llama-3-70B scores pages for educational value
- Regressor trained on labels filters corpus
- Filtered corpus pretrains new language models
- LLM judgment sets next training diet
- Llama-3-70B scores pages for educational value
- Regressor trained on labels filters corpus
- Filtered corpus pretrains new language models
- LLM judgment sets next training diet
Why it is a road to recursion
Each stronger model can relabel the web for the next one (the dataset card cites Meta using Llama 2 to build the quality classifiers behind Llama 3), so the curator improves with every generation it helps train.
Evidence
In the paper's ablations, moving from FineWeb to FineWeb-Edu raised MMLU from 33% to 37% and ARC from 46% to 57%, and FineWeb-Edu matched the final performance of the Matrix dataset with almost 10x fewer tokens.
Related loops
More training data →- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat
- Agent synthesizes agentic pretraining data
- Pipeline adjusts training set in real time
- Data feeds back into model training
- repeat
- Language model generates a self-edit
- Edit is applied as weight update
- RL rewards self-edit by downstream performance
- repeat