RLAIF (Reinforcement Learning from AI Feedback)

Appears in 1 paper · 1 tutorial

The stage of Constitutional AI where an AI (rather than a human) provides feedback on which response better follows the constitution.

As used in Paper 22 — Constitutional AI: Harmlessness from AI Feedback →

The stage of Constitutional AI where an AI (rather than a human) provides feedback on which response better follows the constitution. The feedback is used to train a reward model, which is then used for RL optimization.

As used in Fine-Tuning & Model Customization →

Preference tuning where the "better/worse" labels come from an AI rather than humans. (M12)