In more detail
Pretraining is a model’s childhood: the long, costly first phase where it reads an enormous slice of the internet and learns how language works by endlessly predicting missing words. No manners, no helpfulness yet — just raw pattern-learning at staggering scale.
It matters because this is where most of a model’s knowledge, and most of its cost, comes from. Everything afterwards — fine-tuning, RLHF — is finishing school by comparison. When a lab announces a new model generation, months of pretraining are what’s really being announced.
Goes with
