Masks decide which words each word may attend to
Tap a word to turn it into padding. Point at or tap a weight to see why it is blocked.
| 0 | 0 | 0 | 1 | 1 |
| 0 | 0 | 0 | 1 | 1 |
| 0 | 0 | 0 | 1 | 1 |
| 0 | 0 | 0 | 1 | 1 |
| 0 | 0 | 0 | 1 | 1 |
| 0 | 0 | 0 | 1 | 1 |
| 0 | 1 | 1 | 1 | 1 |
| 0 | 0 | 1 | 1 | 1 |
| 0 | 0 | 0 | 1 | 1 |
| 0 | 0 | 0 | 0 | 1 |
| 0 | 0 | 0 | 0 | 0 |
| the | cat | sat | PAD | PAD | |
|---|---|---|---|---|---|
| the | 0 | 1 | 1 | 1 | 1 |
| cat | 0 | 0 | 1 | 1 | 1 |
| sat | 0 | 0 | 0 | 1 | 1 |
| PAD | 0 | 0 | 0 | 1 | 1 |
| PAD | 0 | 0 | 0 | 1 | 1 |
| the | cat | sat | PAD | PAD | |
|---|---|---|---|---|---|
| the | 1.00 | 0.00 | 0.00 | 0.00 | 0.00 |
| cat | 0.50 | 0.50 | 0.00 | 0.00 | 0.00 |
| sat | 0.33 | 0.33 | 0.33 | 0.00 | 0.00 |
| PAD | 0.33 | 0.33 | 0.33 | 0.00 | 0.00 |
| PAD | 0.33 | 0.33 | 0.33 | 0.00 | 0.00 |
In a batch the padding mask has shape (B, 1, 5); the 1 is the query axis it gets copied along.
In the mask lab, uncheck “causal” and tap the last two words to make them padding. Which columns are now empty? Turn “causal” back on. Which row always has exactly one open cell, whatever you pad?
I got stuck here Why is the causal mask a 6×6 table and not just 6 numbers?
Because it is a mask on the weights, and the weight table has one row per word that is looking and one column per word being looked at. For a causal mask, “may word i look at word j” depends on both i and j, so it needs a full 6×6 table.
The padding mask really is just 6 numbers, one per column: a PAD token is blocked for everyone. That is why it is stored as one row, with shape (1, 6), and copied down every row when it is used.
The size comes from the sequence that looks at itself: 6 tokens looking at the same 6 tokens. Side trip N7 has models with two sequences (a source and a target of different lengths). Its section 5 shows which mask each attention uses, and which length sets its size.