Beginner3Blue1Brown

How a network learns

The follow-up. A cost function, then gradient descent: small changes to the weights that reduce error on examples you already labeled.

Gradient descent, how neural networks learn | Deep learning chapter 2

Key points

The network starts with weights that do not yet solve the task. It learns from labeled examples.
A cost function scores how far the outputs are from the labels, averaged over the training examples.
The gradient of that cost says which way to nudge each weight and bias so the cost falls fastest.
Gradient descent repeats the nudge: step downhill, then compute the gradient again. It reaches a local low point, not a guarantee about images it has not seen.
The efficient way to compute that gradient is backpropagation. This chapter says why you need it and leaves the calculus for the next one.
This video does not cover language models, pricing, or how a team should judge a product.
Sign in to earn points for it.

Ask while you watch

The coach stays on this lesson’s notes. If the notes do not cover it, the answer says so.

Sign in to ask and add it to the thread. Sign in

Questions on this lesson

Anyone can read. Posting needs a session.

No questions yet. Be the first to ask.