Language Modeling from Scratch
Exhaustive lecture notes for Stanford's CS336, taught by Percy Liang & Tatsunori Hashimoto. The course's thesis: you understand language models by building them — tokenizer, architecture, optimizer, training, scaling, data, and alignment — end to end, with efficiency as the organizing obsession.
The course in one idea
Prompting frontier models is powerful but the abstraction is leaky, and frontier models are closed and cost ~$1B to train. So we can't reproduce the frontier — but we can learn what transfers: the mechanics of how things work and the mindset of squeezing every bit of efficiency from data and hardware. Five parts, five assignments, one recurring question: what is the best model you can build with a fixed data and compute budget?
tokenize · model · train
kernels · parallelism · inference
predict before you spend
eval · curate · clean
SFT · DPO · RL
Lectures
- 01 Overview & Tokenization done
- 02 Resource Accounting done
- 03 Architectures & Hyperparameters done
- 04 Attention Alternatives & Mixture of Experts done
- 05 GPUs & Making Them Go Fast done
- 06 Triton Kernels, Benchmarking & Profiling done
- 07 Multi-GPU Parallelism done
- 08 Parallelism Deep Dive (ZeRO & 4D) done
- 09 Scaling Laws (Basics) done
- 10 Inference done
- 11 Scaling Laws II — Recipes, Optimizers, μP done
- 12 Evaluation done
- 13 Data I — Sources, Copyright, Datasets done
- 14 Data II — Pipeline & Post-training done
- 15 Alignment — SFT & RLHF done
- 16 RLVR & Reasoning Models done
- 17 Multimodality done
- 18 Inference Systems done
Arc: tokenization → architecture → systems → scaling → data → alignment → reasoning → multimodality → inference systems.
Source: CS336 Spring 2026 playlist (Stanford Online) · cs336.stanford.edu. Notes are a rendering of each lecture for quick study.