ORPO
A method combining SFT and preference tuning into one step, no reference model needed.
A method combining SFT and preference tuning into one step, no reference model needed. (M12)
A method combining SFT and preference tuning into one step, no reference model needed.
A method combining SFT and preference tuning into one step, no reference model needed. (M12)