STaR
STaR: Bootstrapping Reasoning With ReasoningSTaR (Self-Taught Reasoner) turns a few rationale examples and a dataset of question-answer pairs into a large set of model-written chain-of-thought rationales, then fine-tunes the model on them. Repeating this lets a language model improve its reasoning from its own generated explanations.
The language model writes step-by-step rationales for many questions; rationales that reach the correct answer are kept, and for questions it got wrong the model is shown the correct answer and asked to justify it (rationalization). The model is fine-tuned on this self-generated set, and the improved model generates the rationale dataset for the next iteration.
- Language model writes step-by-step rationales
- Correct rationales are kept as dataset
- Model is fine-tuned on self-generated set
- Improved model generates next rationale dataset
- Language model writes step-by-step rationales
- Correct rationales are kept as dataset
- Model is fine-tuned on self-generated set
- Improved model generates next rationale dataset
Why it is a road to recursion
It shows a language model can produce the training data for its own next version and improve each round, the core step of recursive self-improvement through data.
Evidence
On CommonsenseQA, STaR applied to GPT-J (6B) reached 72.5% accuracy, close to the 73.0% of a fine-tuned GPT-3 that is 30x larger, and 12.5 points above GPT-J fine-tuned to predict answers directly.
Related loops
More training data →- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- repeat
- Agent synthesizes agentic pretraining data
- Pipeline adjusts training set in real time
- Data feeds back into model training
- repeat
- Language model generates a self-edit
- Edit is applied as weight update
- RL rewards self-edit by downstream performance
- repeat