Level D2 · Foundations · Diffusion · runs in your browser

Sampling and guidance

How do you make it draw the thing you asked for?

Side trip · best after level 10 · Probability and sampling

This level continues D1.

In level D1, one guess from heavy noise gave an answer that was not sharp. Generating works by taking many small steps instead:

  1. Start from pure random points at t = 50.
  2. Guess the noise and remove a little of it, which takes you one step back.
  3. Repeat: 49, 48, and so on down to 0.

1. One step back

Two letters from level D1 look almost the same, so here they are side by side for the 3-step example at t = 3:

meansat t = 3in code
α_t (alpha)what this one step keeps, 1 − β_t0.5alpha
ᾱ_t (alpha-bar)what all steps so far kept, multiplied together0.9 × 0.8 × 0.5 = 0.36ab

Removing “a little” noise has its own formula. It uses both letters. With the noise guess ε̂ (eps-hat):

mean=xt−βt1−αˉt  ε^αt\text{mean} = \frac{x_t - \dfrac{\beta_t}{\sqrt{1-\bar\alpha_t}}\;\hat\varepsilon}{\sqrt{\alpha_t}} xt−1=mean+βt  zx_{t-1} = \text{mean} + \sqrt{\beta_t}\; z

z is fresh random noise. Yes, a little noise goes back in. Most people do not expect that, and a box further down explains why.

Use the 3-step example from level D1 (β = 0.1, 0.2, 0.5) and the noisy point from there: x_3 = [1.0, 0.4]. Suppose the network guesses the noise was ε̂ = [0.5, -1]. At t = 3: β = 0.5, α = 0.5, alpha-bar = 0.36. Start with the number in front of the noise guess:

Number β = 0.5 and alpha-bar = 0.36. What is β / √(1 − alpha-bar)?

With that number, the rest of the step is subtraction and division.

🔒 Answer the question above to unlock
Number One reverse step: mean = (xₜ − 0.625 × ε̂) / √α. Here x₃ = [1.0, 0.4], the noise guess ε̂ = [0.5, −1] and α = 0.5. Use 1 / √0.5 ≈ 1.41. What is the first number of the mean? (2 decimals)

With fresh noise z = [0.2, -0.4] you get x_2 = [1.114, 1.167].

A step back does not go back along the same path. It picks a new point at random from the places it could have come from. The z here was chosen to land near where the same point would sit at t = 2 going forward, [1.113, 1.168]. With a different z, the step lands somewhere else nearby, and that is fine.

ChooseAt the very last step, from t = 1 to t = 0, do we add fresh noise z?
🔒 Answer the question above to unlock

Now write the step. The last step, from t = 1 to t = 0, passes zeros for z.

CodeWrite one reverse step. z is the fresh noise.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

One step works. Generation is that step, repeated 50 times.

🔒 Answer the question above to unlock

2. Watching 50 steps

Now a real trained network. It learned three shapes: a spiral, a ring, and two moons. Every point starts as random noise. Press Step a few times, then Play. The numbers on the right follow the highlighted point through each step.

Reverse sampling: from noise back to a shape

Pick a shape, then press Step. Each step the network guesses the noise and removes a little of it.

Loading the trained network…

t = 50 of 50

Every point starts as random noise. Press Step.

loading…
the target shape the points at step t the point we follow
I got stuck here Why add fresh noise back while removing noise? Isn’t that going the wrong way?

The noise guess is never exact. It is the network’s best average guess, and an average is less sharp than any single answer. In this sampler, if you only ever subtract, every point moves toward the “average” answer, and the samples gather in the middle of the shape or all land on a few spots.

The reason is in the formula: the mean above was derived for a step that adds fresh noise. Adding a small amount at each step keeps the points as spread as real data. The amount gets smaller as t goes down, and at the last step there is none. Try “To the end” on the spiral: the points spread along the whole arm instead of gathering in a few places.

Other samplers exist that add no noise at all. They use a different step formula, made for that purpose.

Go deeper Where the one-step formula comes from

Going forward, one step does x_t = √α_t · x_(t−1) + √β_t · noise. To go back you would like to know x_(t−1) given x_t.

If you also knew the clean point x_0, that question has an exact answer: x_(t−1) has a Gaussian (normal) distribution around a weighted mix of x_t and x_0. You don’t know x_0, but level D1 showed how to guess it from x_t and the noise guess. Put that guess into the mix, simplify, and the terms rearrange into the mean above.

The std around that mean should be a bit less than √β_t, but using exactly √β_t works just as well in practice and is simpler.

3. Asking for a shape

The network was trained with a label: “this point came from the spiral”, “this one from the ring”. In 20% of training examples the label was hidden. So the same network can make two guesses:

  • ε_free: the noise guess without the label. It only knows “some shape”.
  • ε_shape: the noise guess when told which shape.

Guidance mixes them:

ε^=εfree+w (εshape−εfree)\hat\varepsilon = \varepsilon_\text{free} + w\,(\varepsilon_\text{shape} - \varepsilon_\text{free})

  • w = 0 ignores the label.
  • w = 1 uses the labeled guess as it is.
  • w > 1 goes past it, pushing further in the direction that makes the point “more like this shape”.
Number Guidance mixes two noise guesses: ε̂ = ε_free + w (ε_shape − ε_free). ε_free = [0.2, 0.0], ε_shape = [0.6, −0.4], w = 3. What is the first number of the guided guess ε̂?
Predict firstGuess before you try it: what happens to the samples when w is very large, like 10?

Guidance: how strongly w moves points toward the shape

Drag w. The same starting noise is used every time, so only w changes.

Loading the trained network…

0: ignore the shape1: plain10: push hard
on the shape…
shape covered…
0%50%100%01510w
sampling…
the true shape samples at this w on the shape: samples within 0.15 of it shape covered: parts of it with a sample within 0.15 the curve is a quicker run with 120 points, so it differs a little from the boxes

At w = 1 the shapes are not sharp. That is normal for a network this small. Slide w up and watch them sharpen.

Try it

Pick the ring and slide w from 0 to 10. Watch both numbers: “on the shape” climbs and then falls, and “shape covered” falls once w passes about 3. Find the w that gives a good balance of the two. Is it the same for the spiral?

Number Sampling with guidance, ε̂ = ε_free + w (ε_shape − ε_free), takes 50 steps. How many times do we run the network in total? (One run handles all 300 points at once.)

The last piece is the guidance mix itself.

CodeWrite the guidance mix.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

Now put both halves together. One guided step back is: mix the two noise guesses, then take the reverse step from section 1 with the mixed guess. This is exactly what the lab does 50 times.

CodeWrite one guided step back, from x to the next x, with guidance weight w. z is the fresh noise (zeros on the last step).

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

You can now draw a chosen shape from pure noise. In level D3 the points become pictures, and the network that guesses the noise is a Transformer.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Take one reverse step by hand: compute the mean from the noise guess, then add fresh noise scaled by √β.
  • Explain why each step adds a little fresh noise back, and why the last step adds none.
  • Mix two noise guesses with a guidance weight w and use the mixed guess in the reverse step.

Keep in mind

  • , with on the last step
  • α_t = kept in this one step; ᾱ_t = kept by all steps so far
  • : w = 0 ignores the label, w > 1 goes past it
  • Guidance runs the network twice per step: 50 steps → 100 runs

Common mistakes

  • Dividing by instead of this step’s at the bottom of the mean.
  • Scaling the fresh noise by β instead of √β, or adding it on the last step.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars