AI for Everybody - Lesson 10
How a Language Model Actually Works: What "Training" Means
The model produces a probability over the next token, samples from it, appends, repeats. Fine. But who taught it those probabilities? Nobody handed them down. The probabilities were learned, in a procedure called training, and training is the engine behind every machine-learning system worth knowing about. This week we open it up.
The short version of training, in one sentence: the system is shown many examples, measured on how wrong it is, and nudged in the direction of being slightly less wrong, trillions of times in a row. That’s the whole of it. No magic. No understanding. No teacher in a tweed jacket explaining anything. The principle is over a hundred years old and the modern form of it is over forty years old (Rumelhart, Hinton, & Williams, 1986). Everything that looks miraculous about large language models comes out of this one procedure run at scale.




