Level N1 · Foundations · Classic networks · runs in your browser

Convolutions

How can one small filter find an edge anywhere in a picture?

Side trip · best after level 10 · Probability and sampling

A 28×28 picture is 784 numbers. An MLP from level 5 gives every one of them its own weight to every hidden unit. But an edge looks the same whether it is in the top-left corner or the middle. A convolution uses one small grid of weights, the filter, and slides it over the whole picture.

1. One filter, one window

The picture below is 5×5: a vertical stroke of 1s, like the digit 1. The filter is 3×3. To get one output number, lay the filter on a 3×3 window of the picture, multiply the 9 pairs, and add them up. Then slide one step and do it again.

A 3×3 filter sliding over a picture

Select a pixel to flip it. Point at, tap or focus an output cell to see the window it reads.

picture
∗
filter (editable)
=
output 3×3
?
?
?
?
?
?
?
?
?
Filter
output side = ⌊(5 − 3 + 2×0) / 1⌋ + 1 = 3
positive output negative output background = 0 the 3×3 window
Number The picture is 5×5 with 1s in column 2 and 0s everywhere else. The filter is [[−1, 0, 1], [−1, 0, 1], [−1, 0, 1]]. What is out[0][0], the output for the top-left 3×3 window?
🔒 Answer the question above to unlock

The left side of the filter is −1 and the right side is +1. So the output is large where the picture is brighter on the right than on the left. That is a vertical edge. Each output cell answers one question: “is there a dark-to-bright step here, going left to right?”

Number The picture is 5×5 with 1s in column 1 and 0s everywhere else. The filter is [[−1, 0, 1], [−1, 0, 1], [−1, 0, 1]], no padding, stride 1. What is out[0][1]?
ChooseA 5×5 picture has 1s in column 2 and 0s everywhere else. With the horizontal-edge filter [[−1, −1, −1], [0, 0, 0], [1, 1, 1]], no padding and stride 1, what does the 3×3 output look like?
I got stuck here Why does one stroke appear twice, as +3 and as −3?

The stroke has two edges. On its left side the picture goes from 0 to 1: dark to bright, so the filter says +3. On its right side it goes from 1 back to 0: bright to dark, so the filter says −3. A network learns to use both. Later layers can combine “a +3 here and a −3 two pixels to the right” into “a thin line”.

2. Output size: stride and padding

🔒 Answer the question above to unlock

Two settings change the output size:

  • Stride is how far the filter moves each step. Stride 2 uses only every second position (0, 2, 4, …) and roughly halves the output.
  • Padding adds a ring of zeros around the picture, so the filter can also sit on the border pixels.

For a picture of side nn, a filter of side kk, stride ss and padding pp:

output side=⌊n−k+2ps⌋+1\text{output side} = \left\lfloor \frac{n - k + 2p}{s} \right\rfloor + 1

The brackets ⌊  ⌋\lfloor \; \rfloor mean round down to a whole number: ⌊15.5⌋=15\lfloor 15.5 \rfloor = 15.

Try both settings in the lab above. With padding 1 and stride 1, a 5×5 picture stays 5×5.

Number A 28×28 picture, a 5×5 filter, no padding, stride 1. How many pixels wide is the output?
Number A 32×32 picture, a 3×3 filter, padding 1, stride 2. How many pixels wide is the output?

3. Why one shared filter

🔒 Answer the question above to unlock

The same 9 weights are used at every position. That is called weight sharing. Compare it with a dense layer (level 5) that produces the same output: it has one weight for every pair of (input pixel, output cell).

Number A 6×6 picture and a 3×3 filter give a 4×4 output. A dense layer that makes the same 4×4 output connects every one of the 36 pixels to every one of the 16 output cells. How many weights does it have (no biases)?
🔒 Answer the question above to unlock

For a real digit the gap is huge. A 28×28 picture has 784 pixels, a 26×26 output has 676 cells, and a dense layer needs 784 × 676 = 529,984 weights. The filter still needs 9: about 59,000 times fewer. The filter learns “what an edge looks like” once and finds it everywhere.

Predict firstTwo networks train on 60,000 handwritten digits for the same 1,500 steps. A convolutional network has 5,258 weights. An MLP has 50,890. Which scores higher on 10,000 digits it has never seen?
Go deeper Does a convolutional network still work on a shifted picture?

Not fully. Moving the picture moves the feature map by the same amount: the filters find the same edges, just elsewhere. But at the end, the network flattens the feature maps into one list and multiplies it by a dense layer. That layer learned which positions matter, and it does not treat every position the same.

The level’s demo.py measures it. Moving every test digit 3 pixels right and 3 pixels down drops the convolutional network from 98.0% to 31.2%, and the MLP from 96.2% to 16.2%. Both are hurt. The convolutional network is hurt less, because its pooling layers (section 5) keep only the largest number in each small block, so a small move changes less.

4. Channels: many filters at once

🔒 Answer the question above to unlock

One filter answers one question, such as “is there a vertical edge here?”. A real layer asks many questions at once: it has many filters, and each one makes its own output map. The maps are stacked, and each map is called a channel. A layer with 8 filters turns one picture into 8 channels.

The next layer then reads all 8 channels. So each of its filters is not 3×3 but 8×3×3: one 3×3 slice per input channel. At each position it multiplies all 8 windows by their slices and adds the 72 products into one number. A color picture works the same way from the start: it has 3 input channels (red, green, blue).

shape
input(C_in, H, W)
one filter(C_in, k, k)
the layerC_out filters, plus one bias each
output(C_out, H′, W′)

So a layer has k⋅k⋅Cin⋅Coutk \cdot k \cdot C_{in} \cdot C_{out} weights plus CoutC_{out} biases. Example: a 3×3 layer from 1 channel to 8 channels has 3 · 3 · 1 · 8 = 72 weights and 8 biases, 80 numbers in all.

Number A 3×3 convolution layer reads 8 input channels and makes 16 output channels. How many numbers does it learn, weights plus biases?
🔒 Answer the question above to unlock

That is how the convolutional network from the Predict above gets its 5,258 numbers: 80 for the first layer, 1,168 for the second, and 4,010 for the final dense layer (16 channels × 5 × 5 positions = 400 inputs, to 10 digit scores: 400 · 10 + 10).

Go deeper Same four steps, different connections

An MLP, a convolutional network, a recurrent network (side trip N3) and a Transformer (level 14 onward) all do the same four things: multiply inputs by weights, add a bias, apply a nonlinearity, and learn the weights with gradient descent. They differ only in which inputs each weight touches and which weights are shared.

networkeach output readsweights shared across
MLPevery inputnothing
convolutiona small windowevery position in the picture
recurrentthis step’s input and the last stateevery time step
Transformerevery token, mixed by attention weightsevery token position

In deep learning, “convolution” does not flip the filter first, as the math definition does. It is really a cross-correlation. Because the filter is learned, the difference does not matter.

5. Stacking layers: seeing more of the picture

🔒 Answer the question above to unlock

One 3×3 filter sees a 3×3 window. Put a second 3×3 layer on top of the first one. Each of its outputs reads a 3×3 window of first-layer outputs, and each of those reads 3×3 pixels. Together, one second-layer output depends on a bigger patch of the picture. That patch is its receptive field.

Number Two 3×3 layers are stacked, stride 1, no pooling. One output pixel of the second layer depends on a square patch of the picture. How many pixels wide is that patch?

Pooling shrinks a feature map. 2×2 max pooling keeps the largest number in each 2×2 block, so the map halves in each direction.

Number 2×2 max pooling of [[1, 3, 0, 2], [4, 2, 1, 1], [0, 0, 5, 6], [1, 2, 7, 0]] gives a 2×2 result. What is its bottom-right number?
🔒 Answer the question above to unlock

To find the receptive field of a deeper stack, track two numbers as you go up, one layer at a time:

  • width: how many pixels wide the patch is. Start at 1 (one pixel).
  • gap: how many pixels apart two neighboring cells of the current map are. Start at 1.

Then apply one rule per layer:

  • a k×k convolution adds (k − 1) × gap to the width;
  • a 2×2 pooling adds 1 × gap to the width, and then doubles the gap (its output cells sit twice as far apart).

Worked example, the first two layers of “conv, pool, …”:

layerwidthgap
start11
3×3 conv1 + 2 × 1 = 31
2×2 pool3 + 1 × 1 = 42

Without pooling the gap stays 1, which is why two 3×3 layers gave 5. Now continue the table for two more layers.

Number Track the patch width and the gap layer by layer. A k×k conv adds (k − 1) × gap to the width. A 2×2 pool adds 1 × gap to the width, then doubles the gap. After a 3×3 conv and a 2×2 pool, the width is 4 and the gap is 2. Add a second 3×3 conv, then a second 2×2 pool. How many pixels wide is the patch one final output sees?

6. What deeper layers see

🔒 Answer the question above to unlock

How much of the picture one output pixel sees

Add layers, switch the filter size, turn pooling on, and watch the field grow.

  1. conv 1 (3×3): sees 3×3
one pixel after the last layer sees 3×3 pixels of the picture
the output pixel darkest ring: what the first layer sees; each lighter ring: what one more layer adds

Early layers see small patches and find edges. Deeper layers see large patches and combine edges into loops, corners and whole digits.

Here is the whole digit network from section 4 in 3D. Each layer is a block: width and height are the picture, depth is the channels.

The picture gets smaller, the channels get deeper

Press Next to go one layer at a time. Turn the view to see the channels placed one behind another.

Drag to turn · click, then scroll to zoom
Press Next to start. · 5,258 parameters in all
Layers (6)

Look at the cone of lines: one output cell reads only a small window of the layer before it, through every channel.

forward (this step) Select a layer to see its numbers.

Try it

Switch the convolution lab to Real digits and try the vertical-edge filter on the 1 and on the 0. Then edit the filter so it gives large outputs only on strokes that lean like “/”. Which three cells did you make positive?

7. Write it yourself

First, how to cut out a window in NumPy. For a 2D array, x[0:3, 0:3] takes rows 0–2 and columns 0–2 in one step: a 3×3 block. The list habit x[0:3][0:3] does something else: x[0:3] takes rows 0–2, and the second [0:3] takes rows 0–2 of that again, so you get all columns. For a window starting at row i and column j, write x[i:i+3, j:j+3].

* multiplies two same-shape arrays cell by cell, and .sum() adds up every number in an array.

Write a convolution with two loops. For each output position, take the window under the filter and add up window × filter.

CodeWrite one line: the output at (i, j) is the 3×3 window starting at (i, j), times the filter, summed.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

Now let the filter side and the stride change. With filter side f and stride s, output cell (i, j) reads the window that starts at row i * s and column j * s, so the window is x[i*s:i*s+f, j*s:j*s+f]. With no padding, the output side is the formula from section 2 with p = 0. In Python, a // b divides and rounds down: 7 // 2 is 3.

CodeWrite a convolution for any square filter and any stride, no padding. Set the output side m, make out, and fill every out[i, j]. Replace ____ with as many lines as you need.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Compute one output cell by hand: lay the filter on its window, multiply cell by cell, add the products.
  • Compute a layer’s output size and its number of weights and biases before you build it.
  • Write a convolution with any filter size and stride in NumPy, with two loops and a window slice.

Keep in mind

  • Parameters: weights + biases
  • Window for output (i, j): x[i*s:i*s+f, j*s:j*s+f], filter side f, stride s
  • Receptive field: a k×k conv adds (k − 1) × gap; a 2×2 pool adds 1 × gap, then doubles the gap

Common mistakes

  • Forgetting that each filter reads every input channel: it is , not .
  • Cutting a window with x[0:3][0:3] (rows twice) instead of x[0:3, 0:3], or using @ instead of *.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars