Grace Ann Hansen's Substack

Grace Ann Hansen's Substack

AI for Everybody - Lesson 10

How a Language Model Actually Works: What "Training" Means

Grace Ann Hansen's avatar
Grace Ann Hansen
Jul 21, 2026
∙ Paid
Image by Grace Ann Hansen using NANO BANANA 2

The model produces a probability over the next token, samples from it, appends, repeats. Fine. But who taught it those probabilities? Nobody handed them down. The probabilities were learned, in a procedure called training, and training is the engine behind every machine-learning system worth knowing about. This week we open it up.

The short version of training, in one sentence: the system is shown many examples, measured on how wrong it is, and nudged in the direction of being slightly less wrong, trillions of times in a row. That’s the whole of it. No magic. No understanding. No teacher in a tweed jacket explaining anything. The principle is over a hundred years old and the modern form of it is over forty years old (Rumelhart, Hinton, & Williams, 1986). Everything that looks miraculous about large language models comes out of this one procedure run at scale.


Grace Ann Hansen's Substack is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

User's avatar

Continue reading this post for free, courtesy of Grace Ann Hansen.

Or purchase a paid subscription.
© 2026 Grace Ann Hansen LLC · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture