Learn Up next

CS336: Language Modeling from Scratch

Stanford's build-everything-yourself LLM course. Tokenizer, transformer, systems, scaling laws, data, alignment. Five assignments, no shortcuts.

0/24 done cs336.stanford.edu ↗

Working through the Spring 2026 offering. Each lecture gets notes here; each assignment gets a write-up and a repo.

Lectures

  • 1 · Overview, tokenization
  • 2 · PyTorch (einops), resource accounting: FLOPs, memory, arithmetic intensity
  • 3 · Architectures, hyperparameters
  • 4 · Attention alternatives and mixture of experts
  • 5 · GPUs, TPUs
  • 6 · Kernels, Triton
  • 7 · Parallelism I
  • 8 · Parallelism II
  • 9 · Scaling laws I
  • 10 · Inference
  • 11 · Scaling laws II
  • 12 · Evaluation
  • 13 · Data: sources, datasets
  • 14 · Data: filtering, deduplication, mixing, synthetic data
  • 15 · Mid/post-training: SFT, RLHF
  • 16 · Post-training: RLVR
  • 17 · Alignment, multimodality
  • 18 · Guest lecture: Daniel Selsam
  • 19 · Guest lecture: Dan Fu

Assignments

  • A1 · Basics
  • A2 · Systems
  • A3 · Scaling
  • A4 · Data
  • A5 · Alignment and reasoning RL