Level 10 · Foundations · runs in your browser

Probability and sampling

How does a model “roll the dice” to pick the next word?

A language model never writes a word directly. Its last step gives one score for each possible next word. A higher score means “more likely”. (These raw scores are also called logits. Later levels use both words for the same thing.) Then two things happen:

  1. softmax turns the scores into probabilities that add up to 1.
  2. The model picks one word at random, like rolling uneven dice: a word with a higher probability is picked more often.

1. Picking a word at random

Here are four possible next words. Sample a few times and watch the counts. Leave the temperature slider at 1 for now. Section 3 explains it.

Temperature changes the probabilities; sampling picks words with them

Move the temperature, then sample words and compare the counts with the probabilities.

The cat sat on the ___

wordscorescore / Tprobabilitysampled
mat 2.0 2.00 ? —
sofa 1.0 1.00 ? —
roof 0.5 0.50 ? —
moon -1.0 -1.00 ? —

The probabilities appear once you answer the question below.

No samples yet. Changing T clears the counts.
probability fraction of the samples so far the word drawn last

Each word’s probability is its part of the total: its escoree^{\text{score}} divided by the sum over all four words. (Section 2 explains why we use escoree^{\text{score}}. If ee is new to you, the math page has a short table.)

Number The scores are mat 2.0, sofa 1.0, roof 0.5, moon −1.0. Rounded, e^score is 7.4, 2.7, 1.6 and 0.4, which add up to 12.1. What probability does “mat” get? (2 decimals)
🔒 Answer the question above to unlock

2. How softmax works

Two steps. First, take escoree^{\text{score}} of every score. This makes every number positive, and it makes big scores much bigger. Then divide each one by the total, so they add up to 1.

pi=esi∑jesjp_i = \frac{e^{s_i}}{\sum_j e^{s_j}}

Read ∑j\sum_j as “add up over every word jj”. So the bottom is es0+es1+…e^{s_0} + e^{s_1} + \dots, the total (words count from 0).

Try it with just two words, scores [2, 0]. You need e2≈7.39e^2 \approx 7.39 and e0=1e^0 = 1.

Number Scores [2, 0], with e² ≈ 7.39 and e⁰ = 1. What probability does the first word get? (2 decimals)
ChooseScores [2, 0] give the probabilities 0.88 and 0.12. Add 10 to both scores: [12, 10]. What happens to the probabilities?
I got stuck here Why not just divide each score by the sum of the scores?

Because scores can be negative or zero. Scores [2, -2] add up to 0, and you can’t divide by 0. Scores [1, -3] would give a probability of −0.5 for one word, which makes no sense.

exe^x is positive for every xx, so after taking ee to the power of each score, every number is a valid weight. It also only cares about differences: es+10=es⋅e10e^{s+10} = e^{s} \cdot e^{10}, and the e10e^{10} is in the top and in the bottom of the fraction, so it divides away. That is why adding 10 to every score changed nothing.

🔒 Answer the question above to unlock

3. Temperature

Before softmax, the model can divide every score by a number TT, the temperature.

  • T<1T < 1 makes the differences bigger. The top word wins almost every time.
  • T>1T > 1 makes the differences smaller. Rare words get picked more often.
ChooseThe scores are mat 2.0, sofa 1.0, roof 0.5, moon −1.0. At T = 1, sampling 1000 words gives about 600 mat. Temperature divides every score by T. At temperature 0.1, what will sampling 1000 words give?

Now go back to the lab and drag the temperature from 0.1 to 3.

Number Scores [2, 0] at temperature T = 2, with e¹ ≈ 2.72. What probability does the first word get? (2 decimals)

Now write it. Divide the scores by T, then do softmax.

CodeWrite the line that applies the temperature. scores arrives as a Python list, and a list can’t be divided by a number, so make it a float array first.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

The template subtracts the largest score before np.exp, so exp never overflows; softmax gives the same result. Optional side trip: level U2 shows why this matters, by looking at how a float is stored.

Go deeper Greedy decoding, and why chat models are not greedy

At T→0T \to 0 the biggest score gets probability 1. Always picking the top word is called greedy decoding. It is predictable, but it tends to repeat itself, because the same context always leads to the same word.

Sampling with TT around 0.7–1 keeps some variety. A common extra trick is top-k: keep only the k highest scores, set the rest to −∞-\infty (so e−∞=0e^{-\infty} = 0), then sample. The rare, nonsense words can never be picked, but there is still some choice among the good ones.

In level 17 you will run a real model step by step and see these scores for every word it writes.

🔒 Answer the question above to unlock

4. Scoring a guess: cross-entropy

During training we know the correct next word. We need one number that says how wrong the model was. Look at the probability the model gave the correct word, and take −ln⁡-\ln of it.

loss=−ln⁡pcorrect\text{loss} = -\ln p_{\text{correct}}

This is the loss you met in level 3, where there were only two answers (yes or no). There you took −ln⁡p-\ln p when the answer was yes and −ln⁡(1−p)-\ln(1 - p) when it was no; both are “−ln of the probability given to the right answer”. Now there can be any number of answers, and softmax supplies their probabilities.

In this course log always means the natural log, ln (base e≈2.718e \approx 2.718), never base 10. np.log in code is ln too. Here is −ln⁡p-\ln p for a few values of pp:

pp10.50.250.10.01
−ln⁡p-\ln p00.691.392.304.61

Two rules make these easy. −ln⁡(1/x)=ln⁡x-\ln(1/x) = \ln x: for example −ln⁡0.1=−ln⁡(1/10)=ln⁡10≈2.30-\ln 0.1 = -\ln(1/10) = \ln 10 \approx 2.30. And a smaller pp always gives a bigger loss. If ln is new to you or you have forgotten it, the math page has the rules with examples.

Choose the correct word with the round button, then drag the scores.

Cross-entropy: the surprise at the correct answer

Pick the correct answer with the round buttons. Drag the sliders to change the scores.

correct?scorep
2.0 0.659
1.0 0.242
0.1 0.099

The loss number appears once you answer the two questions below.

loss00.51p of the correct answer
correct answer: cat · p = 0.659 · loss = −log(0.659) = ?
p of the correct answer p of the other answers loss = −log p where you are now
Number The model gave the correct word p = 0.1. What is the loss −ln p? Use ln 10 ≈ 2.30. (2 decimals)
Number The model gave the correct word p = 0.5. What is the loss? Use ln 2 ≈ 0.69. (2 decimals)
I got stuck here Why the log? Why not just use 1 − p?

Look at the curve. With 1−p1 - p, a model that gives the right word p=0.01p = 0.01 gets a loss of 0.99, and one that gives it p=0.0001p = 0.0001 also gets about 0.99. They look equally bad.

With −ln⁡p-\ln p, the first gets 4.6 and the second gets 9.2. The closer the model comes to giving the right word zero probability, the larger the loss grows, without limit. So training strongly pushes the model to give every possible right word at least some probability.

There is a second reason. The probability of a whole sentence is the product of the probabilities of its words. The log of a product is the sum of the logs. So the total loss is just the sum of the per-word losses.

Write the loss for one example: softmax the scores, pick the correct one, take −ln⁡-\ln.

CodeReturn the cross-entropy loss: softmax the scores, take the correct one, then −log.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

I got stuck here In PyTorch, should I apply softmax before F.cross_entropy?

No. F.cross_entropy(logits, target) does the softmax and the −ln⁡-\ln in one step, so you give it the raw scores. If you apply softmax first, it is applied twice.

ChooseScores [2, 0]. Softmax gives [0.88, 0.12]. If you pass these probabilities to F.cross_entropy as if they were scores, it applies softmax again. What happens to training?
🔒 Answer the question above to unlock

5. Which way to push: the gradient

Training needs to know how to change each score to lower the loss. For softmax followed by cross-entropy, the answer is short enough to remember:

∂L∂si=pi−yi\frac{\partial L}{\partial s_i} = p_i - y_i

Here yy is the one-hot target: 1 for the correct word, 0 for the others. Take three words with probabilities p=[0.7,0.2,0.1]p = [0.7, 0.2, 0.1] for cat, dog and car, and say the right answer is dog, so y=[0,1,0]y = [0, 1, 0]:

catdog (correct)car
pp0.70.20.1
yy010
p−yp - y0.7?0.1
Number p = [0.7, 0.2, 0.1] for cat, dog, car, and the correct word is dog, so y = [0, 1, 0]. What is ∂L/∂s for dog, p − y?
🔒 Answer the question above to unlock

The correct word is the only one with a negative entry: it is the only score that should go up.

ChooseGradient descent moves every score against its gradient: s ← s − lr × (p − y). With p − y = [0.7, −0.8, 0.1] for cat, dog, car, which score goes up?
Go deeper Where p − y comes from

The loss is L=−ln⁡pcL = -\ln p_c for the correct word cc, and pc=esc/∑jesjp_c = e^{s_c} / \sum_j e^{s_j}. Take the ln of that fraction:

L=−sc+ln⁡∑jesjL = -s_c + \ln \sum_j e^{s_j}

Now change one score sis_i a little. The first term only depends on scs_c: it gives −1-1 when i=ci = c and 00 otherwise, which is −yi-y_i. The second term changes by esi/∑jesje^{s_i} / \sum_j e^{s_j}, which is exactly pip_i. Add them: pi−yip_i - y_i.

With two answers this is the p−yp - y that level 6 uses for the yes/no neuron (sigmoid + its cross-entropy). Same rule, more answers.

Write it in one line: subtract 1 from the probability of the correct word.

CodeReturn the gradient of the loss with respect to the scores: p minus the one-hot target.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

6. How surprised is the model? Perplexity

Cross-entropy scores one guess. To judge a model on a whole text, take the average loss over every word. Each −ln⁡p-\ln p is the model’s surprise at the word that actually came, so the average is its average surprise.

That number is hard to picture. So we turn it back with exe^x, which undoes the ln:

perplexity=eaverage of (−ln⁡p)\text{perplexity} = e^{\text{average of } (-\ln p)}

Perplexity reads like a count: the model is as unsure as if it were choosing evenly among this many words.

Number A model always splits its guess evenly over 4 words, p = 0.25 each, and the right word is always one of them. What is its perplexity?
🔒 Answer the question above to unlock

Now drag the bars. Each one is how much the model believes in that word. The lab shows the perplexity you would get if the words really came with these probabilities: “as unsure as choosing evenly among this many words”.

Entropy and perplexity: how unsure is the model?

Drag how much the model believes in each word, or press a preset.

wordbelief (drag)p−ln p
mat 0.250 1.386
sofa 0.250 1.386
roof 0.250 1.386
moon 0.250 1.386
perplexity 4.00
entropy = Σ p · (−ln p) = 1.386 nats · perplexity = e1.386 = 4.00 : as unsure as an even choice among 4.00 words
p of each word one equally likely word; the filled part is the perplexity
Go deeper Entropy: the lab’s other number

The lab averages the surprise over the four words, weighted by how likely each one is. That weighted average is the entropy: how uncertain the model is before it sees the answer. The perplexity in the lab is eentropye^{\text{entropy}}.

On real text the two can differ. Perplexity uses the surprise at the words that actually came. Entropy weights by the model’s own probabilities. They are equal only when the words really come with the model’s probabilities. You don’t need entropy for the rest of the course.

Try it

Press “even” and you get perplexity 4: four equally likely words. Press “sure” and you get 1: no doubt at all. Now make two words equal and the others zero. Perplexity is 2. The count is the number of words the model is really choosing between.

ChooseTen words. Model A gives the right word p = 0.9 nine times and p = 0.001 once. Model B gives it p = 0.5 every time. Which has the lower perplexity?
I got stuck here Why is perplexity like a number of words?

If a model splits its guess evenly over kk words, every right word gets p=1/kp = 1/k. The surprise is −ln⁡(1/k)=ln⁡k-\ln(1/k) = \ln k every time, so the average is ln⁡k\ln k, and eln⁡k=ke^{\ln k} = k.

A real model is never exactly even. Perplexity says which even split would be just as surprised on average. A perplexity of 2.3 means “about as unsure as a choice between two or three words”.

Here is a model reading the sentence “the cat sat on the mat”. After each word it guesses the next one.

next wordcatsatonthemat
p the model gave it0.50.250.1250.250.25
surprise −ln⁡p-\ln p0.6931.3862.0791.3861.386
Number For “the cat sat on the mat” the model gave the right next words p = 0.5, 0.25, 0.125, 0.25, 0.25. The surprises −ln p add up to 6.93, and ln 4 ≈ 1.386. What is its perplexity?

Write it for any list of probabilities.

CodeReturn the perplexity of a list of probabilities (the p the model gave each right word).

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Go deeper Nats, bits, and why perplexity doesn’t care

With the natural log, surprise is measured in nats. With log⁡2\log_2 it is in bits: a surprise of 1 bit means “as surprising as one coin flip”. Both are fine, as long as you undo the log with the matching base: exe^x for nats, 2x2^x for bits. The perplexity is the same either way.

Perplexity is also the geometric mean of 1/p1/p. For the sentence above: the 1/p1/p values are 2, 4, 8, 4, 4, their product is 1024 = 4⁵, and the fifth root of 1024 is 4.

When you compare language models in level 11, this is the number you will use. Lower is better.

🔒 Answer the question above to unlock

7. Put it together

The core of this level in one function: turn scores into probabilities with a temperature, safely, and score the right word. “Safely” means the trick from the probs code in section 3: subtract the largest score before np.exp, so a score like 1000 does not overflow. Training never scores one example at a time. It takes a batch of B examples, each with its own row of scores and its own right word. The loss of the batch is the mean of the B losses. Write the whole function; a for loop over the rows is fine.

To loop over two lists side by side, Python has zip: list(zip([1, 2], "ab")) is [(1, 'a'), (2, 'b')]. So for scores, c in zip(rows, correct): gives you each row together with its right word.

CodeWrite mean_loss for a batch: for each row of scores, divide by T, take a safe softmax (subtract the max before np.exp), and take −ln of the probability of that row’s right word. Return the mean of these losses over the rows. Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Optional side trip: the Classic networks branch starts here. It builds convolutions, recurrent networks, LSTMs and the first attention (level N1 to N6). Level 11 continues the main line: the rest of this page (section 8) is optional.

8. Optional: Gaussian noise (for the Diffusion branch)

This section is not needed for the main line: level 11 does not use it, and the level counts as cleared without it. The Diffusion branch (level D1) uses it, so read it before you go there.

Sampling a word picks from a short list. Sometimes we need random numbers instead, like a little noise added to a point. The usual choice is the Gaussian (normal) distribution.

First, one word you need: the standard deviation (std) says how far numbers typically sit from their mean. [1, −1, 1, −1] has mean 0 and std 1. [10, −10, 10, −10] has mean 0 and std 10. (Exactly: the variance is the average squared distance from the mean, and the std is its square root. Level 7 already used std for starting weights.)

To make one noisy number:

x=mean+std×zx = \text{mean} + \text{std} \times z

Here zz comes from a standard normal: mean 0, std 1. Most zz values are between −1 and 1, and values beyond ±3 are rare. The diffusion branch of this course is built entirely on this one line.

Gaussian noise: x = mean + std × z

Move the mean and the std. Every point keeps its own z, so the whole cloud shifts and stretches together.

−4−2024
300 samples (x, y). The ring is one std from the mean.
−4−2024how many land at each x66% within one std
The same points, counted by x only. The curve is the Gaussian’s bell shape for this mean and std.
highlighted point: z = (0.97, 0.02) → x = 0.0 + 1.0 × 0.97 = 0.97 · measured over 300 points: mean of x = 0.06, std of x = 1.01, within one std of the mean (in x): 66%
the highlighted point, computed in the line above the other samples one std from the mean fraction of points per x
Number One noisy number is x = mean + std × z, where z comes from a standard normal. mean = 3, std = 2, and z = −1.5. What is x?
Try it

Set std to 0. Every point lands on the mean. Now set std to 2. The cloud is wider, but the fraction of points within one std of the mean stays near 68%. The highlighted point keeps the same zz, so it slides outward along the same direction.

Go deeper Why the Gaussian appears everywhere

Add up many small, independent random effects and the total is close to a Gaussian, whatever the small effects looked like. That is why measurement errors, heights, and the starting weights of a network are all modeled this way.

In NumPy, rng.standard_normal((n, d)) (or np.random.randn(n, d)) gives an (n, d) array of standard normal zz values: the shape you pass in is the shape you get. Multiply by std and add the mean, and you get any Gaussian you like.

Shape What is the shape of rng.standard_normal((1000, 2)) * 2 + 3?

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Turn scores into probabilities with softmax, at any temperature.
  • Compute the cross-entropy loss and its gradient for one word.
  • Read perplexity as “choosing evenly among this many words”, and compute it from the p of each right word.

Keep in mind

  • ; with temperature, softmax of
  • , with one-hot
  • Optional (section 8, for the Diffusion branch): Gaussian noise,

Common mistakes

  • Dividing a score by the sum of the scores: take to the power of each score first.
  • Applying softmax before F.cross_entropy: it wants the raw scores and applies softmax itself.
Side trips after this levelOptional; the next level does not need them.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars