# D1. Adding and removing noise

> If you mix a picture with noise step by step, can you undo it?

LLM by Hand · Theory · side trip: Diffusion · runs in your browser · interactive page: https://llm.liko.page/learn/noise-and-denoise/

A Transformer writes one word at a time. Diffusion models make pictures in a different way.
They start from pure noise and remove a little of it, many times, until a picture is left.

To learn how to remove noise, the model first watches the opposite: clean data being mixed with noise.
That direction is easy. It takes one line of math.

> **What you need:** Foundations (levels 1–10). D1 and D2 use nothing else. D3 also uses attention and
> Transformer blocks, so take D3 after level 16.

Our "picture" is a cloud of 2D points shaped like a spiral. Every point is two numbers, so you can see everything.

## 1. Adding noise

Each step mixes a bit of random noise into every point. After `t` steps, a point `x0` has become:

$$x_t = \sqrt{\bar\alpha_t}\; x_0 + \sqrt{1-\bar\alpha_t}\; \varepsilon$$

- `ε` (eps) is noise drawn from the normal distribution you sampled in level 10: one new pair of numbers per point.
- `ᾱ_t` (alpha-bar) says how much of the original signal (the clean point x0) is left after `t` steps. It starts at 1 and falls toward 0.
- In this branch, `T` is the number of steps. It is not the temperature T from level 10.
  The bar over the α means “multiplied over all the steps so far”. In code it is called `ab`.

As `ᾱ_t` falls toward 0, the signal’s weight `√ᾱ_t` falls toward 0 and the noise’s weight `√(1 − ᾱ_t)` rises toward 1.
The lab below uses 50 steps (call that number T = 50). Before you drag anything, ask yourself: what does the last step look like?

**Predict.** Noise is mixed into a spiral of 2D points with xₜ = √ᾱₜ x0 + √(1 − ᾱₜ) eps. On a 50-step schedule, alpha-bar at t = 50 is 0.0045. What does the spiral look like there?

A. Still a spiral, just less sharp
B. A round cloud with no trace of the spiral
C. All the points at exactly [0, 0]

*Answer it on the page to check your work.*

Now drag `t` and watch the spiral turn into noise. The highlighted point is the one we follow.

*[Interactive lab: Noise — open the page to use it]*

Where does alpha-bar come from? Each step has a small number β (beta), the amount of noise it adds.
That step keeps α = 1 − β of the signal. After several steps, multiply what each step kept.

Try it on a small example with only 3 steps: β = 0.1, 0.2, 0.5. So α = 0.9, 0.8, 0.5.

**Question.** With β = 0.1, 0.2, 0.5, what is alpha-bar after 3 steps?

*Answer it on the page to check your work.*

### Why the square roots?

Alpha-bar is 0.36, yet the formula multiplies the signal by √0.36 = 0.6. Both are right, because they measure different things.
In level 10 you measured how spread out numbers are with the std. Squaring the std gives the **variance**, and variance is what
alpha-bar counts: 36% of the signal’s variance is left. Multiplying a number by 0.6 multiplies its variance by 0.6² = 0.36.

The noise gets the rest. Its weight is √(1 − 0.36) = √0.64 = 0.8, so its variance is 0.8² = 0.64.

**Question.** The signal is multiplied by 0.6 and the noise by 0.8. What is 0.6² + 0.8²?

*Answer it on the page to check your work.*

That sum is the reason for the square roots. The two weights, squared, always add up to 1, at every step.
So if the clean data has variance 1, the noisy data has variance 1 at every step too: the cloud only changes from
“spiral” into “noise”, and never grows or shrinks. Real models scale their data to variance 1 for this reason.
Our spiral has a variance of about 0.6, so its cloud grows a little as it becomes noise with variance 1.

Now compute it for one point. Take `x0 = [1, 2]` and the noise `eps = [0.5, -1]`.
At t = 3, √0.36 = 0.6 and √(1 − 0.36) = 0.8. After you answer, switch the lab to **Hand example · 3 steps** and drag to t = 3 to watch the point move.

**Question.** x0 = [1, 2], eps = [0.5, −1], alpha-bar = 0.36. What is the first number of x₃?

*Answer it on the page to check your work.*

**If you are stuck: Why jump straight to step t? Don’t you have to add the noise one step at a time?**

You can add it one step at a time and get the same kind of result.
Adding two independent normal noises gives one bigger normal noise, so all t steps combine into one formula.
That shortcut matters for training: to make a training example at step 37, you don't have to run 37 steps.
Compute alpha-bar once and jump there.

Write the shortcut yourself. It works on one point or on a whole array of points at once.

**Code question.** Write `add_noise`. ab is alpha-bar. It should work for one point or for an array of points.

Fill in the blank (`____`):

```python
def add_noise(x0, eps, ab):
    return ____

x0 = np.array([1.0, 2.0])
eps = np.array([0.5, -1.0])
print(add_noise(x0, eps, 0.36))
```

*Answer it on the page to check your work.*

Now put both parts together: start from the list of betas and return the noisy point at step t.
Three pieces of NumPy do it:

- `betas[:t]` takes the first t betas (positions 0 to t − 1), so step t uses steps 1 to t.
- `1 - betas[:t]` turns them into the α values, what each step keeps.
- `np.prod` multiplies all the numbers of an array: `np.prod(np.array([0.9, 0.8, 0.5]))` is 0.36.

**Code question.** Write `noisy_at`: from the clean point x0, the noise eps, the list of betas and a step t (counting from 1), compute alpha-bar for step t, then return the noisy point xₜ.

Fill in the blank (`____`):

```python
def noisy_at(x0, eps, betas, t):
    ____

betas = np.array([0.1, 0.2, 0.5])
print(noisy_at(np.array([1.0, 2.0]), np.array([0.5, -1.0]), betas, 3))
```

*Answer it on the page to check your work.*

Adding noise is the easy half. Next comes the half that makes generation possible: removing it.

## 2. Running it backward

Here is the key idea. If you knew exactly which noise `eps` was added, you could undo the formula and get `x0` back:

$$x_0 = \frac{x_t - \sqrt{1-\bar\alpha_t}\;\varepsilon}{\sqrt{\bar\alpha_t}}$$

Try it on a different point. At t = 3 (so 0.6 and 0.8 again), a point sits at `x_3 = [0.7, 1.0]`, and you are told its noise was `eps = [0.5, 0.5]`.

**Question.** At t = 3 (alpha-bar 0.36), x₃ = [0.7, 1.0] and the true noise was eps = [0.5, 0.5]. What is the second number of x0?

*Answer it on the page to check your work.*

When generating, nobody tells you the noise. So a network is trained to **guess** it.
The network sees a noisy point and the step number `t`, and outputs its guess of the noise.
The loss is the mean squared error between the guess and the true noise.

**Question.** The simplest network always guesses that the noise is [0, 0]. For one point the true noise is [0.5, −1]. What is the loss, mean((guess − eps)²)?

*Answer it on the page to check your work.*

**Question.** One training batch has 128 noisy points. Each input row is the point’s x and y, plus 12 numbers that encode the step t. What shape goes into the network?

*Answer it on the page to check your work.*

**Deeper: Why the network needs to know t**

Take the same noisy position at two different steps:

- At step 2 the point has hardly moved, so almost all of it is signal. The right noise guess is small.
- At step 48 the point is almost pure noise. The right guess is close to the point itself.

The same position needs a different answer at different steps.

So the network gets t as 12 extra numbers: sin and cos of t/T (T = 50, the number of steps) at 6 different frequencies (how fast each one repeats).
A single number would work too, but those 12 make it easy for the network to treat
nearby steps alike and far-apart steps differently. Transformers use the same idea for word positions (level 16).

## 3. Training the noise-guessing network

Next you train this network in your browser. Each training step:

1. takes 128 points from the spiral,
2. picks a random t for each and adds the matching noise,
3. asks the network to guess that noise and changes its weights a little to reduce the loss.

The simplest guess, always [0, 0], scored 0.625 on our one point. Averaged over many random noises it scores 1.0,
because each noise number has variance 1, so the average of ε² is 1. Training should do better than that. How much better?

**Predict.** Guess before you train: if training goes perfectly, where does the loss stop?

A. All the way down to 0
B. Below 1.0, but it stops falling somewhere above 0
C. It stays at 1.0, like the [0, 0] guess

*Answer it on the page to check your work.*

Now press **Train**. This is real gradient descent with Adam, the same loop you met in levels 2 and 7.

*[Interactive lab: Denoise train — open the page to use it]*

**Try it**

Train all 4,000 steps, then drag “look at step t”. At t = 5 the guessed points sit right on the spiral.
At t = 45 they no longer follow the spiral; they land in an unclear cloud. Why can't the network do better from so much noise?
Is it a bad network, or is the question impossible?

One piece is left: turning a noise guess into a guess of the clean point. It is the undo formula from section 2, with the network's guess in place of the true noise.
The guess is written `ε̂` (“eps-hat”, `eps_hat` in code): a hat on a letter marks an estimate of it.

**Code question.** Write `guess_clean`: given a noisy point, the network’s noise guess, and alpha-bar, return the guess of x0.

Fill in the blank (`____`):

```python
def guess_clean(xt, eps_hat, ab):
    return ____

print(guess_clean(np.array([1.0, 0.4]), np.array([0.5, -1.0]), 0.36))
```

*Answer it on the page to check your work.*

You now have both halves. Forward: one line that adds noise. Backward: a network that guesses the noise.
One guess from heavy noise gives an answer that is not sharp, as you saw. In level D2 you take many small steps instead of one big one,
and the spiral comes back sharp.

## You can now

- Compute alpha-bar from a list of betas: multiply what each step keeps, α = 1 − β.
- Compute the noisy point at any step t in one line, without running the steps in between.
- Undo the formula with a noise guess to get a guess of the clean point.
