Beginner3Blue1Brown
How a network learns
The follow-up. A cost function, then gradient descent: small changes to the weights that reduce error on examples you already labeled.
Gradient descent, how neural networks learn | Deep learning chapter 2
Key points
The network starts with weights that do not yet solve the task. It learns from labeled examples.
A cost function scores how far the outputs are from the labels, averaged over the training examples.
The gradient of that cost says which way to nudge each weight and bias so the cost falls fastest.
Gradient descent repeats the nudge: step downhill, then compute the gradient again. It reaches a local low point, not a guarantee about images it has not seen.
The efficient way to compute that gradient is backpropagation. This chapter says why you need it and leaves the calculus for the next one.
This video does not cover language models, pricing, or how a team should judge a product.
Sign in to earn points for it.
Ask while you watch
The coach stays on this lesson’s notes. If the notes do not cover it, the answer says so.
Sign in to ask and add it to the thread. Sign in
Questions on this lesson
Anyone can read. Posting needs a session.
No questions yet. Be the first to ask.