RLHF (Reinforcement Learning from Human Feedback)
The original preference-tuning approach: train a reward model, then optimize the LLM against it with RL.
The original preference-tuning approach: train a reward model, then optimize the LLM against it with RL. (M12)