Nemotron-4 340B
Nemotron-4 340B Technical ReportNVIDIA's open release of Nemotron-4 340B Base, Instruct and Reward models, whose report documents the synthetic data pipeline that produced over 98% of the alignment data. The pipeline and models are licensed for generating training data for other LLMs.
An aligned model generates prompts, dialogues and responses, and Nemotron-4-340B-Reward judges, ranks and filters them into supervised and preference training data. In the report's iterative weak-to-strong alignment, Mixtral-8x7B-Instruct generates the first round of data, and the resulting 340B intermediate instruct model then becomes the generator for the next round, so each aligned model produces the data that aligns a better one.
- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
- Aligned model generates training data
- Reward model judges and filters it
- Data trains next intermediate instruct model
- Next model generates better alignment data
Why it is a road to recursion
The report runs the generator-to-successor handoff more than once, replacing the weaker generator with the newly aligned model, which is the basic shape of models producing the data for their own successors.
Evidence
Only about 20K human-annotated examples were used across the whole alignment process, while the generation pipeline synthesized over 98% of the data used for supervised and preference fine-tuning.
Related loops
More training data →- Agent synthesizes agentic pretraining data
- Pipeline adjusts training set in real time
- Data feeds back into model training
- repeat
- Language model generates a self-edit
- Edit is applied as weight update
- RL rewards self-edit by downstream performance
- repeat
- o3-mini injects bugs and writes issues
- SWE-agent solves tasks for expert trajectories
- Trajectories fine-tune Qwen 2.5 Coder
- Open coding agent trained on agent data
- repeat