paper-with-me

Step-DPO

Step-wise Direct Preference Optimization

2000년 도입 · 논문 2편에서 사용

Please enter a description about the method here

Language Models · Natural Language Processing