ManagerAndrej Karpathy

Large language models, without training one

Andrej Karpathy’s one-hour talk for a general audience. What the model file is, what training cost looked like in 2023, and where the security problems sit. The shape of the explanation still holds.

[1hr Talk] Intro to Large Language Models

Key points

An LLM you run is two pieces: a parameters file and the code that runs it. It is not looking up a stored sentence.
Pretraining compresses a very large text corpus into those parameters. The compression is lossy, so the model produces a plausible continuation rather than a retrieved page.
The talk’s 2023 figures for Llama 2 70B are about 10TB of text, about 6,000 GPUs, about 12 days, and about $2 million. The talk says frontier systems were already roughly ten times past those figures.
Fine-tuning aims the model at assistant behavior. It does not stop hallucination. The talk says an answer is more trustworthy when the needed text is in the context, from browsing or retrieval, than when the model answers from parameters alone.
The talk’s “LLM OS” picture: the model is like a kernel, the context window is working memory, and tools are peripherals.
Security problems named in the talk are jailbreaks, prompt injection, and data poisoning.
Those figures and that security list are from this November 2023 talk. They are not a current bill or a full threat model.
Sign in to earn points for it.

Ask while you watch

The coach stays on this lesson’s notes. If the notes do not cover it, the answer says so.

Sign in to ask and add it to the thread. Sign in

Questions on this lesson

Anyone can read. Posting needs a session.

No questions yet. Be the first to ask.