LLMs GuruLLMsGuru
Find your AI
AI ModelsAI ToolsLearnGlossaryAI NewsFree ToolsFind your AISpeakRightAstrology
Explore All Tools
← The AI Dictionary
Plain-English definition

RLHF

In one breath

Reinforcement learning from human feedback: training a model with human thumbs-up and thumbs-down until its answers feel helpful and polite. It's why chatbots sound the way they do.

In more detail

RLHF — reinforcement learning from human feedback — is how a raw model learns manners. People compare pairs of the model’s answers and pick the better one, thousands upon thousands of times, and the model is then trained to produce more of what humans preferred.

It matters because it’s the difference between a wild autocomplete engine and a helpful assistant. The politeness, the careful refusals, the reassuring tone — that whole familiar chatbot personality is largely RLHF’s handiwork, for better and occasionally for blander.

📌 See it in action

A model writes two answers to ‘Explain mortgages simply.’ One is accurate but dense; the other is clear and friendly. A human reviewer clicks the friendly one. Multiply that click by armies of reviewers across countless questions — that distilled preference is why chatbots sound the way they do.

Goes with
AlignmentFine-TuningPretraining

Now put the vocabulary to work.

Meet the models →Which AI is mine?