# 11. What is a language model?

> How can predicting the next word turn into writing?

LLM by Hand · Theory · runs in your browser · interactive page: https://llm.liko.page/learn/language-models/

A language model does one thing. Given the text so far, it gives a probability to every word that could come next.

To write, you use it in a loop: get the probabilities, pick a word (level 10 showed how to sample), append it,
and ask again with the longer text. Everything a chatbot writes is made by that loop.

This level builds the smallest possible language model, first by counting, then as a neural network.
The rest of Theory makes it look further back than one word.

## 1. Count the neighbors

The simplest model only looks at the **previous word**. To learn it, read some text and count how often
each word is directly followed by each other word. A pair of neighbors is called a **bigram**.

Here is a tiny piece of text. It has 12 words, so 11 pairs of neighbors:

```
the cat sat . the dog sat . a cat ran .
```

**Question.** In “the cat sat . the dog sat . a cat ran .”, how many times is “sat” directly followed by “.”?

*Answer it on the page to check your work.*

A count turns into a probability when you divide by the row total: of all the times you saw "cat",
how often was the next word "sat"?

**Question.** In “the cat sat . the dog sat . a cat ran .”, “cat” appears twice: once before “sat”, once before “ran”. What is P(sat | cat)?

*Answer it on the page to check your work.*

Here is the whole count table for that text. Row = this word, column = the next word.

|        | the | a | cat | dog | sat | ran | . |
|--------|----:|--:|----:|----:|----:|----:|--:|
| **the** | 0 | 0 | 1 | 1 | 0 | 0 | 0 |
| **a**   | 0 | 0 | 1 | 0 | 0 | 0 | 0 |
| **cat** | 0 | 0 | 0 | 0 | 1 | 1 | 0 |
| **dog** | 0 | 0 | 0 | 0 | 1 | 0 | 0 |
| **sat** | 0 | 0 | 0 | 0 | 0 | 0 | 2 |
| **ran** | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| **.**   | 1 | 1 | 0 | 0 | 0 | 0 | 0 |

**Predict.** A counting model looks one word back. It counted only the text “the cat sat . the dog sat . a cat ran .”. What probability does it give the sentence “a dog sat .”?

A. 0: it can never write it
B. Small but not zero
C. The same as “a cat sat .”

*Answer it on the page to check your work.*

## 2. A bigger story collection

Three sentences are too few. The lab uses 200 tiny stories made by one rule of our own:

```
(the | a) (cat | dog) (sat | ran) [on the (mat | rug)] .
```

Each choice is random: "the" 70% of the time, "cat" 60%, "sat" 50%, and half the sentences get "on the mat" or "on the rug".
That gives a vocabulary of 10 words (the period “.” counts as a word). 160 stories are used for counting; 40 are held out (not used for counting) to test the model later.

Before you open the lab, one question about what the counting model can write.

**Predict.** A counting model looks one word back. Its training stories contain “… on the mat .” many times, but no story contains “the mat .” without “on” before it. Can the model write “the mat .”?

A. Yes
B. No, it never saw it

*Answer it on the page to check your work.*

*[Interactive lab: Bigram — open the page to use it]*

**Try it**

Press "Write a sentence" until you get one marked **never in the stories**. Find the step where it went wrong,
and look at the probability under that word. Every single step was allowed. Only the whole sentence is strange.

**If you are stuck: It only ever picks one word. Where does the writing come from?**

From the loop. Each pick becomes part of the text, and the next pick is made from the new text.
A sentence is just the record of many single picks.

That is also why one bad pick can make the rest of the sentence wrong: the model never goes back to change a word.
Level 18 looks at smarter ways to pick.

## 3. Scoring a model: perplexity

How good is a language model? Show it text it has not seen, and look at the probability it gave to each
word that actually came next. Average the $-\ln p$ values and undo the log:

$$
\text{perplexity} = \exp\Big(\frac{1}{n}\sum_{k=1}^{n} -\ln p_k\Big)
$$

Level 10 introduced the formula and wrote it as code; here you only need it by hand. A perplexity of 4 means the model was, on average,
as unsure as if it were choosing among 4 equally likely words. Lower is better. Guessing uniformly among 10 words gives exactly 10.

**Question.** After a “.”, a model reads “the cat ran” and gives each next word P(the | .) = 0.5, P(cat | the) = 0.5, P(ran | cat) = 0.5. What is the perplexity?

*Answer it on the page to check your work.*

**If you are stuck: What if the model gave a real next word probability 0?**

Then $-\ln 0$ is infinite, and so is the perplexity. One impossible-looking word ruins the whole score.
The counting model does this whenever the test text has a pair it never counted, like "a dog" in the tiny example.

The usual fix for counting is **smoothing**: add a small number, say 1, to every cell before dividing, so nothing is
exactly impossible. Neural models never have this problem, because softmax never outputs exactly 0.

## 4. The same model as a neural network

The count table can also be **learned**. Give the network a table of weights W, one row per previous word.
The row is a list of scores, and softmax turns them into probabilities:

$$
P(\text{next} \mid \text{prev}) = \text{softmax}(\text{onehot}(\text{prev})\, W)
$$

Multiplying by a one-hot row just picks row "prev" of W. (Level 13 uses the same trick to give every word a list of numbers.)
Training is the loop from level 2: compute the cross-entropy of the real next words, take the gradient, step.

**Question.** The stories have a vocabulary of 10 words. W holds one row of scores per previous word, one score per possible next word. What is the shape of W?

*Answer it on the page to check your work.*

**Predict.** The counting model scores perplexity 2.016 on the 40 held-out stories. The neural version starts with W all zeros (perplexity 10) and trains on the same 160 stories. Where is its perplexity after training?

A. Much better than 2.016
B. About the same as 2.016
C. Clearly worse than 2.016

*Answer it on the page to check your work.*

*[Interactive lab: Neural bigram — open the page to use it]*

The weights start at zero, so every row starts uniform: perplexity 10. After 300 steps it is at 2.037 and still
slowly moving toward the counting model's 2.016.

The learning rate here is 5, much bigger than the 0.05 of level 2. That is safe because of the form of this loss surface:
softmax plus cross-entropy over one table changes slowly in every direction, so even a big step does not jump far past the lowest point.
(Level 2’s squared error on x values up to 3 changes much faster, which is why 0.2 already diverged there.)

**Code question.** Compute the mean cross-entropy of the real next words. W is (2, 2), prev and nxt are arrays of word ids.

Fill in the blank (`____`):

```python
# scores chosen so that softmax(W[0]) = [0.5, 0.5] and softmax(W[1]) = [0.9, 0.1]
W = np.log(np.array([[0.5, 0.5],
                     [0.9, 0.1]]))
prev = np.array([0, 0, 1])
nxt  = np.array([0, 1, 0])

def loss(W, prev, nxt):
    P = softmax(W[prev])                # (3, 2): one row of probabilities per pair
    p_real = P[np.arange(len(nxt)), nxt]
    print("probability of the real next word:", p_real)
    return ____

print(loss(W, prev, nxt))
```

*Answer it on the page to check your work.*

**Deeper: Why training finds the counts**

For one row, the loss is the cross-entropy between the real next words and softmax(W[row]).
With nothing else tying the rows together, the best you can do is to make softmax(W[row]) equal the fraction
of times each word followed, which is exactly the count table divided by its row total.

So for a one-word context, the neural network and the counting model are the same model, found two different ways.
The difference only appears when the context gets longer, which is the next section.

## 5. Why counting stops working

To look two words back, the count table needs one row for every **pair** of previous words.

**Question.** With a vocabulary of 10 words, how many rows does a count table need to look 2 words back (one row per possible pair of previous words)?

*Answer it on the page to check your work.*

**Predict.** A real vocabulary has 50,000 tokens. A count table that looks 3 tokens back needs 50,000³ = 125,000,000,000,000 rows. What goes wrong?

A. Almost every row is empty: those 3-token contexts never appear in any text
B. Counting all that text takes too long to finish
C. The rows no longer add up to 1

*Answer it on the page to check your work.*

A neural network does not need a row per context. It turns each word into a vector (level 13), and similar words
get similar vectors, so what it learns about "the cat sat" also helps with "the dog sat". Then attention (level 14)
lets it look at every earlier word at once, with weights it computes. That is the path from here to a GPT.

## 6. The whole counting model in one line

Before we stop using counts, build its probability table yourself. You have the count table `C` from section 1
(row = this word, column = the next word), this time with **smoothing**: add 1 to every cell first, so no pair is
exactly impossible (a pair counted 0 times now counts 1, a pair counted 2 times counts 3). Then each row has to become
probabilities that add up to 1, so each row is divided by its own total. Two NumPy tools do it:

```python
M = np.array([[1, 3], [2, 2]])
M.sum(axis=1)                  # [4, 4]       one total per row, shape (2,)
M.sum(axis=1, keepdims=True)   # [[4], [4]]   the same totals, shape (2, 1)
M / M.sum(axis=1, keepdims=True)   # [[0.25, 0.75], [0.5, 0.5]]   (2, 1) is stretched across the columns
```

**Code question.** C is the count table of “the cat sat . the dog sat . a cat ran .” (row = this word, column = the next word). Write probs: the next-word probabilities with add-one smoothing, so every row adds up to 1. A few lines.

Fill in the blank (`____`):

```python
words = ["the", "a", "cat", "dog", "sat", "ran", "."]
C = np.array([[0, 0, 1, 1, 0, 0, 0],
              [0, 0, 1, 0, 0, 0, 0],
              [0, 0, 0, 0, 1, 1, 0],
              [0, 0, 0, 0, 1, 0, 0],
              [0, 0, 0, 0, 0, 0, 2],
              [0, 0, 0, 0, 0, 0, 1],
              [1, 1, 0, 0, 0, 0, 0]], dtype=float)

def probs(C):
    ____
    return P

P = probs(C)
print("P(next | cat):", P[2])
```

*Answer it on the page to check your work.*

## You can now

- Count bigrams in a text and turn one row of counts into next-word probabilities.
- Compute the perplexity of a short text by hand from the probabilities of its real next words.
- Say why a count table that looks several words back grows too large to fill.
