# 7. Optimization

> How do you make training fast and stable on real data?

LLM by Hand · Foundations · runs in your browser · interactive page: https://llm.liko.page/learn/optimization/

You can already train a network: compute the loss, compute the gradient, step downhill.
Real training adds four habits that make it faster and more stable. Each one fixes a problem you can see with a few numbers:

1. **Mini-batches**: don't look at all the data before every step.
2. **Better optimizers**: don't let one steep direction set the speed for everything.
3. **Initialization**: start with weights that keep the numbers a sensible size.
4. **A learning-rate schedule**: take small steps at the start and at the end.

Level 8 is about the other half of training well: a model that also works on data it has never seen.

## 1. Mini-batches

Fit a line, $\hat y = w x + b$, to four points: (1, 3), (2, 3), (3, 7), (4, 7). They lie roughly on y = 2x, with some noise.
Start from w = 0.5, b = 0.

The errors (prediction − truth) are −2.5, −2, −5.5, −5. The gradient for w is the mean of 2 · error · x over the points.

**Question.** Points (1, 3), (2, 3), (3, 7), (4, 7); w = 0.5, b = 0, so the errors are −2.5, −2, −5.5, −5. Using all four points, what is dL/dw = mean of 2·error·x?

*Answer it on the page to check your work.*

That used all four points. A **mini-batch** uses only some of them. Take a batch of just one point, (3, 7):

**Question.** A line ŷ = w·x + b with w = 0.5, b = 0. Use a batch of just one point, (3, 7). Its error is −5.5. What is dL/dw = 2·error·x for this batch?

*Answer it on the page to check your work.*

One point pulls harder than the average: −33 instead of −21.5. Point (1, 3) alone gives only −5.
Each batch has its own opinion of which way is downhill. On average the opinions agree with the full gradient,
but any single step follows just one of them.

**Predict.** Batch size 1, learning rate 0.02, start w = 0.5, b = 0. After each step you measure the loss on all four points. Does it go down on every step?

A. Yes: every step lowers the loss on all four points
B. No: many steps raise it, and it never stays at the bottom
C. No: it rises on every step and never comes down

*Answer it on the page to check your work.*

Try it. Batch size 4 is the full data. Batch size 1 uses one point per step.

*[Interactive lab: Batch — open the page to use it]*

With batch size 4 the path is smooth. With batch size 1 it goes back and forth, and near the bottom it never settles.
Even at the best line, w = 1.6 and b = 1, each point alone still pulls in its own direction.

**If you are stuck: If the full batch is smoother, why not always use all the data?**

Because real datasets have millions of examples. One full-batch step would read all of them, just to move once.
A batch of 32 gives a noisy but useful direction, for a tiny fraction of the cost, so you can take thousands of steps
in the time of one full step.

The noise is not only a cost. It can move the weights away from flat places and small dips, which often helps.
A common fix for the small back-and-forth moves at the end is to lower the learning rate as training goes on.

One pass through all the training data, batch after batch, is called an **epoch**. With 4 points and batches of 1,
an epoch is 4 steps. Training usually runs for many epochs, shuffling the order each time.

**Question.** A training set has 1,000 examples and you use mini-batches of 50. How many steps make one epoch?

*Answer it on the page to check your work.*

**If you are stuck: Why shuffle the data again at the start of every epoch?**

Without shuffling, every epoch has the same batches in the same order. The weights then follow the same chain of
directions each time, and if the data is sorted (all 3s first, then all 4s), each part of the epoch sees only one kind
of example. The last, smaller batch is also always the same few examples. A new random order each epoch gives a new set
of batches, so the noise is different every time, and it helps the weights get out of flat places.

## 2. Optimizers

Here is a loss shaped like a narrow valley:

$$
f(w_1, w_2) = 0.5\,(w_1^2 + 25\,w_2^2)
$$

Its gradient is $(w_1,\ 25 w_2)$. The valley is 25 times steeper across it ($w_2$) than along it ($w_1$).
Start at (−4, 1). The gradient there is (−4, 25).

Plain gradient descent (SGD) does $w \leftarrow w - \text{lr} \cdot g$ with lr = 0.03.

**Question.** f(w1, w2) = 0.5·(w1² + 25·w2²), start at (−4, 1), gradient (−4, 25). One SGD step with lr = 0.03: what is the new w2?

*Answer it on the page to check your work.*

w2 jumped from 1 to 0.25, a move of 0.75. w1 moved only 0.12, from −4 to −3.88.
The steep direction moves fast and the gentle one moves very slowly. You want a bigger learning rate for w1, so raise it:

**Predict.** The loss is f(w1, w2) = 0.5·(w1² + 25·w2²), with gradient (w1, 25·w2). Start at (−4, 1). SGD with lr = 0.09 instead of 0.03, to make w1 move faster. What happens to w2?

A. It reaches the bottom faster
B. w2 jumps from side to side, further each time, and grows very large
C. It stops moving

*Answer it on the page to check your work.*

So with plain SGD the steepest direction sets the speed limit for every direction. Two fixes:

- **Momentum** keeps a velocity $u$ that remembers past steps: $u \leftarrow \beta u + g$, then $w \leftarrow w - \text{lr} \cdot u$.
  β (beta) is usually 0.9: each step keeps 90% of the old velocity and adds the new gradient.
  Steps that agree (along the valley) add up. Steps that flip sign (across the valley) cancel each other.
- **Adam** gives every weight its own step size. It keeps two **running averages**, each one mixing the newest value
  into the old average, for example $m \leftarrow 0.9\,m + 0.1\,g$:
  - $m$, the average gradient (with beta1 = 0.9),
  - $v$, the average **squared** gradient (with beta2 = 0.999).

  It steps by $\text{lr} \cdot \hat m / \sqrt{\hat v}$. Dividing by $\sqrt{\hat v}$ cancels the size of the gradient,
  so each weight moves about lr, whether its gradient is 4 or 25.

Here is Adam’s first step with lr = 0.3. The question asks how far w1 **moves**, not where it lands. $\hat m$ and $\hat v$ are m and v with a small correction for the first few steps. On step 1 they are simply $\hat m = g$ and $\hat v = g^2$.
The square root √ always means the positive root, so $\sqrt{(-4)^2} = \sqrt{16} = 4$.

**Question.** Adam with lr = 0.3 starts at (−4, 1), gradient (−4, 25). On step 1, m̂ = g and v̂ = g², so the step is 0.3·g/√(g²). By how much does w1 change? (Give the size of the move, not the new position.)

*Answer it on the page to check your work.*

Now run all three from (−4, 1) and compare them.

*[Interactive lab: Optimizer — open the page to use it]*

**Try it**

Set momentum β to 0 and press Play. Which path does momentum follow now? Then try β = 0.99. When does a velocity
that remembers more steps help, and when does it carry the weights past the bottom?

**Deeper: Why Adam divides by 1 − β^t**

m and v start at 0. After one step, m = 0.1·g: only a tenth of the gradient, just because it started at zero.
Dividing by 1 − 0.9¹ = 0.1 undoes that, so $\hat m = g$. After many steps, $0.9^t$ is almost 0 and the correction becomes too small to matter.
The same goes for v with beta2 = 0.999.

The ratio $\hat m / \sqrt{\hat v}$ is the key. If a weight's gradient keeps the same sign, the ratio is near ±1 and the weight moves about lr per step.
If the sign keeps flipping, $\hat m$ averages toward 0 while $\sqrt{\hat v}$ stays large, so the step shrinks. That is what stops the back-and-forth steps across the valley.

Write momentum's step yourself, as two lines: first the new velocity, then the move.

**Code question.** Write momentum’s step: first update the velocity u from the new gradient g, then move w by lr times u. Return both.

Fill in the blank (`____`):

```python
def momentum_step(w, u, g, lr=0.03, beta=0.9):
    ____
    return w, u

w, u = momentum_step(np.array([-4.0, 1.0]), np.zeros(2), np.array([-4.0, 25.0]))
print("w:", w, "u:", u)
```

*Answer it on the page to check your work.*

Now write Adam's update. In NumPy the square root is `np.sqrt`, for example `np.sqrt(16.0)` is `4.0`.
Real code adds a tiny number `eps` (1e-8) to the bottom of the fraction (the denominator), $\sqrt{\hat v} + \text{eps}$, so it never
divides by 0. The function receives it as the argument `eps`.

**Code question.** Write Adam’s update. mₕₐₜ and vₕₐₜ are already computed.

Fill in the blank (`____`):

```python
def adam_step(w, g, m, v, t, lr=0.1, beta1=0.9, beta2=0.999, eps=1e-8):
    m = beta1 * m + (1 - beta1) * g
    v = beta2 * v + (1 - beta2) * g**2
    m_hat = m / (1 - beta1**t)
    v_hat = v / (1 - beta2**t)
    w = w - ____
    return w, m, v

w, g = np.array([1.0, 1.0]), np.array([4.0, 0.01])
w, m, v = adam_step(w, g, np.zeros(2), np.zeros(2), t=1)
print("after one step:", w)
```

*Answer it on the page to check your work.*

## 3. Initialization

Before training starts, every weight needs a value. Random is the usual answer. But how big?

To talk about "how big", we need one word. The **standard deviation (std)** says how far numbers typically sit from
their mean. [1, −1, 1, −1] has std 1. [10, −10, 10, −10] has std 10. (Level 10 gives the exact formula. The
[math page](/math/#mean-std) has a short example too.)

Each output of a layer is a sum over its inputs: $z = x_1 w_1 + x_2 w_2 + \dots + x_n w_n$. If the inputs have standard deviation 1
and the weights have standard deviation s, the sum has standard deviation about √n · s. To keep it at 1, pick
**s = 1/√n**, where n is the number of inputs to the layer (its "fan-in").

**If you are stuck: Why does the std of a sum grow like √n, not n?**

Because the ups and downs partly cancel. What adds up exactly is the **variance** (std squared): the sum of n independent
terms, each with variance s², has variance n·s². Take the square root to get back to std: √n · s.
Example: 100 terms with std 0.1 each have variance 100 × 0.01 = 1, so std 1.

**Question.** A layer has 100 inputs (fan-in n = 100). Using the rule s = 1/√n, what standard deviation should its weights have?

*Answer it on the page to check your work.*

Now stack 10 layers and get the size wrong. Each layer multiplies the standard deviation by √n · s:

**Question.** Inputs have std 1. They pass through 10 layers with no activation, width 100, weights with std 0.2. Each layer multiplies the std by √100 · 0.2. What is the std after 10 layers?

*Answer it on the page to check your work.*

An activation function does not fix this. Try tanh:

**Predict.** A network has 10 layers of width 100 with tanh after every layer. The weights have std 0.01, and the inputs have std 1. Without tanh, each layer would multiply the std by √100 · 0.01. After 10 layers the outputs are…

A. About the same size as the inputs
B. Almost exactly 0
C. Huge

*Answer it on the page to check your work.*

Pass 40 random inputs through 10 layers of width 100. Change the weight size and the activation.

*[Interactive lab: Init — open the page to use it]*

**Deeper: Why ReLU wants √(2/n) instead of 1/√n**

ReLU sets every negative number to 0. About half of the values are negative, so ReLU removes about half the variance
at each layer. To correct for this, double the variance of the weights: s² = 2/n, so s = √(2/n). With n = 100 that is about 0.14.
Try it in the lab with ReLU. With 0.1 the values get smaller and smaller, to about 0.02 after 10 layers. With 0.14 they stay near 0.5.

tanh is close to a straight line near 0, so it behaves like the linear case: 1/√n keeps the values from vanishing,
although they still shrink, to about 0.2 after 10 layers.

## 4. Learning-rate schedules

Big models rarely keep one learning rate for the whole run. Two habits are common:

1. **Warmup**: start with a tiny lr and raise it over the first few hundred steps.
2. **Decay**: lower lr as training goes on.

**Deeper: Why use warmup, and why decay?**

At the start, $\hat m$ and $\hat v$ are averages of very few gradients, and the weights are random, so big early
steps can push the weights far in a bad direction. Warmup keeps those first steps small.

Near the end, the noisy mini-batch steps go back and forth around the bottom (the small back-and-forth moves you saw
in section 1). A smaller lr makes those moves smaller, so the weights end closer to the bottom.

The simplest schedule is two straight lines. Pick a highest learning rate `peak`, a number of warmup steps `warmup`,
and the number of steps in the whole run, `total`. At step $t$ (counting from 0):

$$
\text{lr}(t) =
\begin{cases}
\text{peak} \cdot t \,/\, \text{warmup} & \text{if } t < \text{warmup} \\
\text{peak} \cdot (\text{total} - t) \,/\, (\text{total} - \text{warmup}) & \text{otherwise}
\end{cases}
$$

The big bracket means "use one of these two lines". Read it as two rules: the top line while $t$ is less than
`warmup`, the bottom line after that. The rate rises from 0 to `peak` during warmup, then falls in a straight line to 0 at the last step.

**Question.** A schedule with peak = 0.004, warmup = 100 and total = 1000. Step 25 is still in warmup, so lr = peak · t / warmup. What is lr at step 25?

*Answer it on the page to check your work.*

After warmup, use the second line.

**Question.** A schedule has peak = 0.004, warmup = 100, total = 1000. Step 550 is after warmup, so lr = peak · (total − t) / (total − warmup). What is lr at step 550?

*Answer it on the page to check your work.*

Many models replace the falling straight line with a smooth curve (half a cosine wave), but the pattern is the same:
up quickly, then slowly down.

The GPT you write in level 21 uses these habits: it trains with Adam on mini-batches of 64.

Now put the pieces of this level together in one small training run. The loss is $(w - 3)^2$, lowest at $w = 3$, and the
schedule is given as `lr_at`. Write the body of the loop: each step takes the gradient, gets its learning rate from
the schedule, and moves w with momentum (section 2).

**Code question.** Write the body of a training loop for the loss (w − 3)², with momentum and a learning-rate schedule: each step gets its own learning rate from `lr_at`. Several lines.

Fill in the blank (`____`):

```python
def lr_at(step, peak, warmup, total):
    if step < warmup:
        return peak * step / warmup
    return peak * (total - step) / (total - warmup)

def grad(w):
    return 2 * (w - 3)                  # gradient of the loss (w − 3)²

def train(w, total, peak=0.1, warmup=10, beta=0.9):
    u = 0.0
    for t in range(total):
        ____
    return w

print("after 2 steps:  ", train(0.0, 2, warmup=2))
print("after 200 steps:", train(0.0, 200))
```

*Answer it on the page to check your work.*

## You can now

- Compute one step of SGD, momentum or Adam by hand on a two-weight loss.
- Pick a weight std for a layer from its fan-in, and say what 10 layers do to the wrong one.
- Compute a warmup-then-decay learning rate at any step.
