In more detail
RLHF — reinforcement learning from human feedback — is how a raw model learns manners. People compare pairs of the model’s answers and pick the better one, thousands upon thousands of times, and the model is then trained to produce more of what humans preferred.
It matters because it’s the difference between a wild autocomplete engine and a helpful assistant. The politeness, the careful refusals, the reassuring tone — that whole familiar chatbot personality is largely RLHF’s handiwork, for better and occasionally for blander.
Goes with
