Post-training is everything done to a model after pre-training to shape its behavior: instruction tuning, preference optimization (RLHF and successors), safety
Everything done to a model after pre-training to shape its behavior: instruction tuning, preference optimization like RLHF, safety training, reasoning training, and tool-use training. Pre-training builds raw capability; post-training decides what kind of assistant it becomes.
Because by 2026 labs compete as much on data quality and alignment technique as on scale. Post-training is where reasoning models were born and where much frontier behavioral differentiation now lives.