How does a model “roll the dice” to pick the next word?
A short test to skip this level
Solve these 5 questions on your own. Answer all of them correctly and the level counts as cleared with three stars, and every part of the page opens. Showing an answer doesn’t count.
Number Scores [2, 0] at temperature T = 2, with e¹ ≈ 2.72. What probability does the first word get? (2 decimals)
Number The model gave the correct word p = 0.1. What is the loss −ln p? Use ln 10 ≈ 2.30. (2 decimals)
CodeReturn the gradient of the loss with respect to the scores: p minus the one-hot target.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Number A model always splits its guess evenly over 4 words, p = 0.25 each, and the right word is always one of them. What is its perplexity?
CodeWrite mean_loss for a batch: for each row of scores, divide by T, take a safe softmax (subtract the max before np.exp), and take −ln of the probability of that row’s right word. Return the mean of these losses over the rows. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Warm-up2 questions from earlier levels
A quick review before you start. Optional. Nothing here locks the level.
A language model never writes a word directly. Its last step gives one score for each possible next word.
A higher score means “more likely”. (These raw scores are also called logits. Later levels use both words for the same thing.)
Then two things happen:
softmax turns the scores into probabilities that add up to 1.
The model picks one word at random, like rolling uneven dice: a word with a higher probability is picked more often.
1. Picking a word at random
Here are four possible next words. Sample a few times and watch the counts.
Leave the temperature slider at 1 for now. Section 3 explains it.
Temperature changes the probabilities; sampling picks words with them
Move the temperature, then sample words and compare the counts with the probabilities.
The cat sat on the ___
wordscorescore / Tprobabilitysampled
mat2.02.00?—
sofa1.01.00?—
roof0.50.50?—
moon-1.0-1.00?—
The probabilities appear once you answer the question below.
No samples yet. Changing T clears the counts.
probabilityfraction of the samples so farthe word drawn last
Each word’s probability is its part of the total: its escore divided by the sum over all four words.
(Section 2 explains why we use escore. If e is new to you, the math page has a short table.)
Number The scores are mat 2.0, sofa 1.0, roof 0.5, moon −1.0. Rounded, e^score is 7.4, 2.7, 1.6 and 0.4, which add up to 12.1. What probability does “mat” get? (2 decimals)
🔒 Answer the question above to unlock
2. How softmax works
Two steps. First, take escore of every score. This makes every number positive,
and it makes big scores much bigger. Then divide each one by the total, so they add up to 1.
pi=∑jesjesi
Read ∑j as “add up over every word j”. So the bottom is es0+es1+…, the total (words count from 0).
Try it with just two words, scores [2, 0]. You need e2≈7.39 and e0=1.
Number Scores [2, 0], with e² ≈ 7.39 and e⁰ = 1. What probability does the first word get? (2 decimals)
ChooseScores [2, 0] give the probabilities 0.88 and 0.12. Add 10 to both scores: [12, 10]. What happens to the probabilities?
I got stuck here Why not just divide each score by the sum of the scores?
Because scores can be negative or zero. Scores [2, -2] add up to 0, and you can’t divide by 0.
Scores [1, -3] would give a probability of −0.5 for one word, which makes no sense.
ex is positive for every x, so after taking e to the power of each score, every number is a valid weight.
It also only cares about differences: es+10=es⋅e10, and the e10 is in the top and in the
bottom of the fraction, so it divides away. That is why adding 10 to every score changed nothing.
🔒 Answer the question above to unlock
3. Temperature
Before softmax, the model can divide every score by a number T, the temperature.
T<1 makes the differences bigger. The top word wins almost every time.
T>1 makes the differences smaller. Rare words get picked more often.
ChooseThe scores are mat 2.0, sofa 1.0, roof 0.5, moon −1.0. At T = 1, sampling 1000 words gives about 600 mat. Temperature divides every score by T. At temperature 0.1, what will sampling 1000 words give?
Now go back to the lab and drag the temperature from 0.1 to 3.
Number Scores [2, 0] at temperature T = 2, with e¹ ≈ 2.72. What probability does the first word get? (2 decimals)
Now write it. Divide the scores by T, then do softmax.
CodeWrite the line that applies the temperature. scores arrives as a Python list, and a list can’t be divided by a number, so make it a float array first.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
The template subtracts the largest score before np.exp, so exp never overflows; softmax gives the same result.
Optional side trip: level U2 shows why this matters, by looking at how a float is stored.
Go deeper Greedy decoding, and why chat models are not greedy
At T→0 the biggest score gets probability 1. Always picking the top word is called greedy decoding.
It is predictable, but it tends to repeat itself, because the same context always leads to the same word.
Sampling with T around 0.7–1 keeps some variety. A common extra trick is top-k: keep only the
k highest scores, set the rest to −∞ (so e−∞=0), then sample. The rare, nonsense words
can never be picked, but there is still some choice among the good ones.
In level 17 you will run a real model step by step and see these scores for every word it writes.
🔒 Answer the question above to unlock
4. Scoring a guess: cross-entropy
During training we know the correct next word. We need one number that says how wrong the model was.
Look at the probability the model gave the correct word, and take −ln of it.
loss=−lnpcorrect
This is the loss you met in level 3, where there were only two answers (yes or no). There you took −lnp when the
answer was yes and −ln(1−p) when it was no; both are “−ln of the probability given to the right answer”.
Now there can be any number of answers, and softmax supplies their probabilities.
In this course log always means the natural log, ln (base e≈2.718), never base 10. np.log in code is ln too.
Here is −lnp for a few values of p:
p
1
0.5
0.25
0.1
0.01
−lnp
0
0.69
1.39
2.30
4.61
Two rules make these easy. −ln(1/x)=lnx: for example −ln0.1=−ln(1/10)=ln10≈2.30.
And a smaller p always gives a bigger loss. If ln is new to you or you have forgotten it, the math page has the rules with examples.
Choose the correct word with the round button, then drag the scores.
Cross-entropy: the surprise at the correct answer
Pick the correct answer with the round buttons. Drag the sliders to change the scores.
correct?scorep
2.00.659
1.00.242
0.10.099
The loss number appears once you answer the two questions below.
correct answer: cat · p = 0.659 · loss = −log(0.659) = ?
p of the correct answerp of the other answersloss = −log pwhere you are now
Number The model gave the correct word p = 0.1. What is the loss −ln p? Use ln 10 ≈ 2.30. (2 decimals)
Number The model gave the correct word p = 0.5. What is the loss? Use ln 2 ≈ 0.69. (2 decimals)
I got stuck here Why the log? Why not just use 1 − p?
Look at the curve. With 1−p, a model that gives the right word p=0.01 gets a loss of 0.99,
and one that gives it p=0.0001 also gets about 0.99. They look equally bad.
With −lnp, the first gets 4.6 and the second gets 9.2. The closer the model comes to giving the right word
zero probability, the larger the loss grows, without limit. So training strongly pushes the model to give every
possible right word at least some probability.
There is a second reason. The probability of a whole sentence is the product of the probabilities of
its words. The log of a product is the sum of the logs. So the total loss is just the sum of the
per-word losses.
Write the loss for one example: softmax the scores, pick the correct one, take −ln.
CodeReturn the cross-entropy loss: softmax the scores, take the correct one, then −log.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
I got stuck here In PyTorch, should I apply softmax before F.cross_entropy?
No. F.cross_entropy(logits, target) does the softmax and the −ln in one step, so you give it the raw scores.
If you apply softmax first, it is applied twice.
ChooseScores [2, 0]. Softmax gives [0.88, 0.12]. If you pass these probabilities to F.cross_entropy as if they were scores, it applies softmax again. What happens to training?
🔒 Answer the question above to unlock
5. Which way to push: the gradient
Training needs to know how to change each score to lower the loss. For softmax followed by cross-entropy,
the answer is short enough to remember:
∂si∂L=pi−yi
Here y is the one-hot target: 1 for the correct word, 0 for the others. Take three words with
probabilities p=[0.7,0.2,0.1] for cat, dog and car, and say the right answer is dog, so y=[0,1,0]:
cat
dog (correct)
car
p
0.7
0.2
0.1
y
0
1
0
p−y
0.7
?
0.1
Number p = [0.7, 0.2, 0.1] for cat, dog, car, and the correct word is dog, so y = [0, 1, 0]. What is ∂L/∂s for dog, p − y?
🔒 Answer the question above to unlock
The correct word is the only one with a negative entry: it is the only score that should go up.
ChooseGradient descent moves every score against its gradient: s ← s − lr × (p − y). With p − y = [0.7, −0.8, 0.1] for cat, dog, car, which score goes up?
Go deeper Where p − y comes from
The loss is L=−lnpc for the correct word c, and pc=esc/∑jesj. Take the ln of that fraction:
L=−sc+lnj∑esj
Now change one score si a little. The first term only depends on sc: it gives −1 when i=c and 0 otherwise,
which is −yi. The second term changes by esi/∑jesj, which is exactly pi. Add them: pi−yi.
With two answers this is the p−y that level 6 uses for the yes/no neuron (sigmoid + its cross-entropy). Same rule, more answers.
Write it in one line: subtract 1 from the probability of the correct word.
CodeReturn the gradient of the loss with respect to the scores: p minus the one-hot target.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
🔒 Answer the question above to unlock
6. How surprised is the model? Perplexity
Cross-entropy scores one guess. To judge a model on a whole text, take the average loss over every word.
Each −lnp is the model’s surprise at the word that actually came, so the average is its
average surprise.
That number is hard to picture. So we turn it back with ex, which undoes the ln:
perplexity=eaverage of (−lnp)
Perplexity reads like a count: the model is as unsure as if it were choosing evenly among this many words.
Number A model always splits its guess evenly over 4 words, p = 0.25 each, and the right word is always one of them. What is its perplexity?
🔒 Answer the question above to unlock
Now drag the bars. Each one is how much the model believes in that word. The lab shows the perplexity you would
get if the words really came with these probabilities: “as unsure as choosing evenly among this many words”.
Entropy and perplexity: how unsure is the model?
Drag how much the model believes in each word, or press a preset.
wordbelief (drag)p−ln p
mat0.2501.386
sofa0.2501.386
roof0.2501.386
moon0.2501.386
perplexity4.00
entropy = Σ p · (−ln p) = 1.386 nats · perplexity = e1.386 = 4.00: as unsure as an even choice among 4.00 words
p of each wordone equally likely word; the filled part is the perplexity
Go deeper Entropy: the lab’s other number
The lab averages the surprise over the four words, weighted by how likely each one is. That weighted average is
the entropy: how uncertain the model is before it sees the answer. The perplexity in the lab is eentropy.
On real text the two can differ. Perplexity uses the surprise at the words that actually came. Entropy weights by
the model’s own probabilities. They are equal only when the words really come with the model’s probabilities.
You don’t need entropy for the rest of the course.
Try it
Press “even” and you get perplexity 4: four equally likely words. Press “sure” and you get 1: no doubt at all.
Now make two words equal and the others zero. Perplexity is 2. The count is the number of words
the model is really choosing between.
ChooseTen words. Model A gives the right word p = 0.9 nine times and p = 0.001 once. Model B gives it p = 0.5 every time. Which has the lower perplexity?
I got stuck here Why is perplexity like a number of words?
If a model splits its guess evenly over k words, every right word gets p=1/k. The surprise is
−ln(1/k)=lnk every time, so the average is lnk, and elnk=k.
A real model is never exactly even. Perplexity says which even split would be just as surprised on average.
A perplexity of 2.3 means “about as unsure as a choice between two or three words”.
Here is a model reading the sentence “the cat sat on the mat”. After each word it guesses the next one.
next word
cat
sat
on
the
mat
p the model gave it
0.5
0.25
0.125
0.25
0.25
surprise −lnp
0.693
1.386
2.079
1.386
1.386
Number For “the cat sat on the mat” the model gave the right next words p = 0.5, 0.25, 0.125, 0.25, 0.25. The surprises −ln p add up to 6.93, and ln 4 ≈ 1.386. What is its perplexity?
Write it for any list of probabilities.
CodeReturn the perplexity of a list of probabilities (the p the model gave each right word).
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Go deeper Nats, bits, and why perplexity doesn’t care
With the natural log, surprise is measured in nats. With log2 it is in bits: a surprise of 1 bit
means “as surprising as one coin flip”. Both are fine, as long as you undo the log with the matching base:
ex for nats, 2x for bits. The perplexity is the same either way.
Perplexity is also the geometric mean of 1/p. For the sentence above: the 1/p values are 2, 4, 8, 4, 4,
their product is 1024 = 4⁵, and the fifth root of 1024 is 4.
When you compare language models in level 11, this is the number you will use. Lower is better.
🔒 Answer the question above to unlock
7. Put it together
The core of this level in one function: turn scores into probabilities with a temperature, safely, and score the
right word. “Safely” means the trick from the probs code in section 3: subtract the largest score before np.exp,
so a score like 1000 does not overflow. Training never scores one example at a time. It takes a batch of B
examples, each with its own row of scores and its own right word. The loss of the batch is the mean of the
B losses. Write the whole function; a for loop over the rows is fine.
To loop over two lists side by side, Python has zip: list(zip([1, 2], "ab")) is [(1, 'a'), (2, 'b')].
So for scores, c in zip(rows, correct): gives you each row together with its right word.
CodeWrite mean_loss for a batch: for each row of scores, divide by T, take a safe softmax (subtract the max before np.exp), and take −ln of the probability of that row’s right word. Return the mean of these losses over the rows. Several lines.
Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs
Optional side trip: the Classic networks branch starts here. It builds convolutions, recurrent networks, LSTMs
and the first attention (level N1 to N6). Level 11 continues the main line: the rest of this page (section 8) is optional.
8. Optional: Gaussian noise (for the Diffusion branch)
This section is not needed for the main line: level 11 does not use it, and the level counts as cleared without it.
The Diffusion branch (level D1) uses it, so read it before you go there.
Sampling a word picks from a short list. Sometimes we need random numbers instead, like a little noise
added to a point. The usual choice is the Gaussian (normal) distribution.
First, one word you need: the standard deviation (std) says how far numbers typically sit from their mean.
[1, −1, 1, −1] has mean 0 and std 1. [10, −10, 10, −10] has mean 0 and std 10. (Exactly: the variance is the
average squared distance from the mean, and the std is its square root. Level 7 already used std for starting weights.)
To make one noisy number:
x=mean+std×z
Here z comes from a standard normal: mean 0, std 1. Most z values are between −1 and 1, and values
beyond ±3 are rare. The diffusion branch of this course is built entirely on this one line.
Gaussian noise: x = mean + std × z
Move the mean and the std. Every point keeps its own z, so the whole cloud shifts and stretches together.
300 samples (x, y). The ring is one std from the mean.The same points, counted by x only. The curve is the Gaussian’s bell shape for this mean and std.
highlighted point: z = (0.97, 0.02) → x = 0.0 + 1.0 × 0.97 = 0.97· measured over 300 points: mean of x = 0.06, std of x = 1.01, within one std of the mean (in x): 66%
the highlighted point, computed in the line abovethe other samplesone std from the meanfraction of points per x
Number One noisy number is x = mean + std × z, where z comes from a standard normal. mean = 3, std = 2, and z = −1.5. What is x?
Try it
Set std to 0. Every point lands on the mean. Now set std to 2. The cloud is wider, but the fraction of points
within one std of the mean stays near 68%. The highlighted point keeps the same z, so it slides
outward along the same direction.
Go deeper Why the Gaussian appears everywhere
Add up many small, independent random effects and the total is close to a Gaussian, whatever the small
effects looked like. That is why measurement errors, heights, and the starting weights of a network
are all modeled this way.
In NumPy, rng.standard_normal((n, d)) (or np.random.randn(n, d)) gives an (n, d) array of standard
normal z values: the shape you pass in is the shape you get. Multiply by std and add the mean, and you get any
Gaussian you like.
Shape What is the shape of rng.standard_normal((1000, 2)) * 2 + 3?
Recap
a summary for when you finish the level
The key formulas and common mistakes appear here once you clear the level.
You can now
Turn scores into probabilities with softmax, at any temperature.
Compute the cross-entropy loss and its gradient p−y for one word.
Read perplexity as “choosing evenly among this many words”, and compute it from the p of each right word.
Keep in mind
pi=∑jesjesi; with temperature, softmax of s/T
loss=−lnpcorrect
∂L/∂si=pi−yi, with y one-hot
perplexity=eaverage of (−lnp)
Optional (section 8, for the Diffusion branch): Gaussian noise, x=mean+std×z
Common mistakes
Dividing a score by the sum of the scores: take e to the power of each score first.
Applying softmax before F.cross_entropy: it wants the raw scores and applies softmax itself.