Back
propagation
The chain rule, one layer at a time
Preset
Dataset
Architecture
−
2-4-4-1
+
Or use the
−
/
+
under any layer on the canvas, and the
+
in the gaps between them. Weights survive the resize.
Hidden activation
tanh
ReLU
sigmoid
Learning rate α
0.500
Batch size
1
One sample — the numbers you see
are
the update.
Speed
2.0 /s
Phases per second.
Play
Step
Reset
Narrate each phase
Weight labels
Live decision boundary
positive weight
negative weight
forward signal
a
backward gradient
δ
Ready
—
Press play to run a training step.
iteration
0
epoch
0.00
loss (dataset)
—
accuracy
—
parameters
—
cross-entropy vs. iteration
—
what the network has learned
—
Chain rule for
Clear selection
Click a weight
to trace its chain rule ·
−
/
+
under a layer to resize ·
Space
play/pause ·
S
step ·
R
reset