Chapters
Chapter 1
What a language model is, probability over sequences, n-grams, and why counting fails.
Chapter 2
Perceptrons, activations, loss functions, backpropagation, and gradient descent.
Chapter 3
Dense vectors, Word2Vec, subword tokenization, and positional encoding.
Chapter 4
Self-attention, query-key-value, multi-head attention, and the math behind it.
Chapter 5
Layer norm, residual connections, feed-forward networks, and encoder/decoder variants.
Chapter 6
Tokenization at scale, next-token prediction, optimizers, and training dynamics.
Chapter 7
Greedy, beam, temperature, top-k, top-p, and interactive text generation.
Chapter 8
Scaling laws, GPT/LLaMA/Mistral families, MoE, and multimodal extensions.
Chapter 9
SFT, RLHF, LoRA, instruction tuning, and parameter-efficient methods.
Chapter 10
Prompt engineering, RAG, benchmarks, and failure modes.