Loss: measuring wrongness
Training starts with a loss function — a number measuring how wrong the model’s predictions are. For language models it is usually cross-entropy: how surprised the model was by the real next token. Lower loss means better predictions on the training task.
MAKE IT CONCRETE
If the true next word is “Paris” and the model gave it 90% probability, loss is low; at 1% it is high.
Learning means making a number go down.
CHECK YOURSELF
