Level 3 · Foundations · runs in your browser

Neurons and classification loss

How does a model answer yes or no, and how do you measure a wrong answer?

In level 2 the model answered with a number. Now it answers yes or no. That takes one neuron:

z=w1x1+w2x2+bz = w_1 x_1 + w_2 x_2 + b p=sigmoid(z)=11+e−zp = \text{sigmoid}(z) = \frac{1}{1 + e^{-z}}

The first part is the line from level 2, with two inputs instead of one. Sigmoid then squeezes any number into the range 0 to 1, so pp reads as “how sure the neuron is that the answer is 1”. If p>0.5p > 0.5, it says 1. The lab also has a steepness kk: it computes p=sigmoid(k⋅z)p = \text{sigmoid}(k \cdot z), so a bigger kk makes the switch from 0 to 1 sharper. The lab also shows a loss for yes/no answers; section 3 explains how it is measured. (e≈2.72e \approx 2.72 and e−ze^{-z} are on the math page if you have forgotten them.)

1. One neuron draws one line

The data is four points, the four corners of a square. Start with AND: the answer is 1 only when both inputs are 1.

One neuron, four points

Drag on the board, or use the arrow keys, to slide the line. The sliders turn and sharpen it.

0001x1 →x2 ↑

p = sigmoid(k · (w1·x1 + w2·x2 + b))

x1, x2targetz = k·(…)psays
0, 00??0 ✓
0, 10??0 ✓
1, 00??0 ✓
1, 11??1 ✓
Data
4 / 4 correct · loss (cross-entropy) 0.4059
background: how sure the neuron is of 1 (stronger = surer) target 1 target 0 a wrong answer p = 0.5, the decision line
Number The AND neuron has w1 = 1, w2 = 1, b = −1.5, and z = w1·x1 + w2·x2 + b. What is z for the point (1, 1)?
🔒 Answer the question above to unlock

Now pass that z through the sigmoid (with k=1k = 1).

Number A neuron has z = 0.5. Its output is p = sigmoid(0.5) = 1 / (1 + e^(−0.5)). Use e^(−0.5) ≈ 0.6. What is p? (Three decimals.)

The neuron says 1 exactly where z>0z > 0, because sigmoid(0) = 0.5. So the boundary between “1” and “0” is the line w1x1+w2x2+b=0w_1 x_1 + w_2 x_2 + b = 0: the dashed line in the lab. This line is the decision boundary. One neuron draws one straight line.

Make the steepness larger. The colors get sharper, but the line doesn’t move. Now switch the data to OR and press the OR neuron preset. A line works for OR too. Now press XOR: the answer is 1 when the inputs differ.

ChooseXOR data: the points (0, 0) and (1, 1) have target 0, and (0, 1) and (1, 0) have target 1. One neuron draws one straight line. Can you find w1, w2, b so that all 4 points are correct?
🔒 Answer the question above to unlock
I got stuck here Sigmoid is a curve, not a line. So why does one neuron still always have a straight boundary?

The sigmoid comes after the line. It bends the output, turning a big negative zz into almost 0 and a big positive zz into almost 1. But it never changes where z=0z = 0, and that is where the answer flips. The boundary is decided by w1x1+w2x2+bw_1 x_1 + w_2 x_2 + b alone, and that is a straight line.

Try it

With the OR data, change only bb and watch the dashed line. Then change only w1w_1. Which one moves the line without turning it, and which one turns it? Then try to get more than 3 out of 4 right on XOR.

2. Write the neuron

Write the neuron in NumPy. x can be one point of shape (2,) or all four points as a (4, 2) matrix. x @ w handles both.

CodeWrite z for one neuron.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

3. How wrong is a yes/no answer?

The lab shows a loss. Level 2 measured a mistake with the squared error. For a yes/no answer there is a better measure, the cross-entropy loss. It uses ln⁡\ln, the natural logarithm (NumPy’s np.log). You only need a few values of −ln⁡p-\ln p (the math page explains ln⁡\ln):

pp10.750.50.250.1
−ln⁡p-\ln p00.290.691.392.30

The smaller pp is, the bigger the cost, and p=1p = 1 costs nothing. Take one point at a time:

  • target 1: the cost is −ln⁡p-\ln p. If the neuron is sure and right (p=1p = 1) the cost is 0. The closer pp gets to 0, the bigger the cost.
  • target 0: the cost is −ln⁡(1−p)-\ln(1 - p). This is the same idea with the neuron’s confidence in 0.

One formula covers both, because yy is 1 or 0, which makes one of the two terms 0:

loss=−mean(yln⁡p+(1−y)ln⁡(1−p))\text{loss} = -\text{mean}\big(y \ln p + (1 - y) \ln(1 - p)\big)

This loss is also called binary cross-entropy (“binary” because there are two answers). A neuron that answers p=0.5p = 0.5 for everything costs −ln⁡0.5≈0.69-\ln 0.5 \approx 0.69 per point. This is the cost of a coin flip. On XOR one neuron can do no better than that, so its loss stops at 0.6931. Level 10 extends the loss to more than two answers and explains where it comes from.

How much better is a model that is only partly sure, but right on every point?

Number On XOR, one neuron answers p = 0.5 for every point. Suppose a model answered p = 0.25 for the two points whose target is 0, and p = 0.75 for the two whose target is 1. What is its loss −mean(y·ln p + (1 − y)·ln(1 − p))? Use ln 0.75 ≈ −0.288 and ln 0.25 ≈ −1.386.
🔒 Answer the question above to unlock
Go deeper Why not the squared error for yes/no answers?

Squared error would cost (p−y)2(p - y)^2. For a target of 1, the worst answer p=0p = 0 costs only 1, and the cost hardly changes as pp moves from 0.01 to 0.001. A neuron that is confidently wrong gets almost no gradient, so it hardly changes.

Cross-entropy costs −ln⁡p-\ln p instead, which grows without limit as pp goes to 0: −ln⁡0.1=2.30-\ln 0.1 = 2.30, −ln⁡0.01=4.61-\ln 0.01 = 4.61. Being sure and wrong is expensive, so the gradient stays large exactly where the model needs to change most. Level 6 shows a second reason: with sigmoid and cross-entropy together, the gradient of the loss with respect to zz is simply p−yp - y.

Now write the loss for a whole array of answers. np.log works on every element of an array, and np.mean averages them (level 2):

p = np.array([0.5, 0.25])
np.log(p)            # array([-0.6931, -1.3863])
np.mean(np.log(p))   # -1.0397
CodeWrite the binary cross-entropy loss: −mean(y·ln p + (1 − y)·ln(1 − p)) for arrays p and y.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

🔒 Answer the question above to unlock

4. The neuron and its loss together

Training (level 6) needs one number for the whole dataset. Run the neuron on every point, then score all the answers at once. This combines section 2 and section 3:

  1. z = X @ w + b gives one z per point, shape (4,).
  2. The sigmoid turns each z into a p.
  3. The loss compares the four p values with the four targets.
CodeWrite the body of neuron_loss: return the binary cross-entropy of the neuron’s predictions against y. Several lines.

Enter keeps the indent · Tab indents · Esc then Tab leaves the editor · ⌘/Ctrl + Enter runs

One neuron draws one line, and XOR needs more than one line. Level 4 stacks neurons and shows what must sit between them so the stack can do more than a single neuron.

Recap

a summary for when you finish the level

The key formulas and common mistakes appear here once you clear the level.

You can now

  • Compute a neuron’s and for a point, and say which side of its line the point is on.
  • Explain why one neuron cannot solve XOR: its decision boundary is one straight line.
  • Compute the binary cross-entropy of a few yes/no answers, by hand and in NumPy.

Keep in mind

  • , in code X @ w + b
  • ; the neuron says 1 where
  • everywhere costs : a coin flip

Common mistakes

  • Using for a target-0 point: its cost is .
  • Forgetting the minus sign, so the loss is negative.
Side trips after this levelOptional; the next level does not need them.

Press ? for keyboard shortcuts

Reading mode · every part open, no stars