Learn Up next
CS336: Language Modeling from Scratch
Stanford's build-everything-yourself LLM course. Tokenizer, transformer, systems, scaling laws, data, alignment. Five assignments, no shortcuts.
0/24 done cs336.stanford.edu ↗
Working through the Spring 2026 offering. Each lecture gets notes here; each assignment gets a write-up and a repo.
Lectures
- 1 · Overview, tokenization
- 2 · PyTorch (einops), resource accounting: FLOPs, memory, arithmetic intensity
- 3 · Architectures, hyperparameters
- 4 · Attention alternatives and mixture of experts
- 5 · GPUs, TPUs
- 6 · Kernels, Triton
- 7 · Parallelism I
- 8 · Parallelism II
- 9 · Scaling laws I
- 10 · Inference
- 11 · Scaling laws II
- 12 · Evaluation
- 13 · Data: sources, datasets
- 14 · Data: filtering, deduplication, mixing, synthetic data
- 15 · Mid/post-training: SFT, RLHF
- 16 · Post-training: RLVR
- 17 · Alignment, multimodality
- 18 · Guest lecture: Daniel Selsam
- 19 · Guest lecture: Dan Fu
Assignments
- A1 · Basics
- A2 · Systems
- A3 · Scaling
- A4 · Data
- A5 · Alignment and reasoning RL