Alphabell.
Paper · Training data

Nemotron-4 340B

Nemotron-4 340B Technical ReportNVIDIA's open release of Nemotron-4 340B Base, Instruct and Reward models, whose report documents the synthetic data pipeline that produced over 98% of the alignment data. The pipeline and models are licensed for generating training data for other LLMs.

The loop

An aligned model generates prompts, dialogues and responses, and Nemotron-4-340B-Reward judges, ranks and filters them into supervised and preference training data. In the report's iterative weak-to-strong alignment, Mixtral-8x7B-Instruct generates the first round of data, and the resulting 340B intermediate instruct model then becomes the generator for the next round, so each aligned model produces the data that aligns a better one.

The loop
Nemotron-4 340B
Paper · arXiv
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
  1. Aligned model generates training data
  2. Reward model judges and filters it
  3. Data trains next intermediate instruct model
  4. Next model generates better alignment data
↻ The improved system does the next round, and the loop turns again.

Why it is a road to recursion

The report runs the generator-to-successor handoff more than once, replacing the weaker generator with the newly aligned model, which is the basic shape of models producing the data for their own successors.

Evidence

Only about 20K human-annotated examples were used across the whole alignment process, while the generation pipeline synthesized over 98% of the data used for supervised and preference fine-tuning.

synthetic-datareward-modelalignmentweak-to-strongopen-weights