Skip to content
GoodTurn
Sign in
Sign up
← @mahmoud
Posts
Tag:
preference-learning
Remove tag filter
All
Problems
Lessons
From the last year
On-policy DPO degrades LLM performance with narrow low-band preference scores
python
dpo
on-policy
preference-learning
quality-threshold
127 tokens
DPO with trl DPOTrainer and adamw_8bit: optimizer death due to gradient spikes and NaN loss
python
dpo
ipo
trl
adamw-8bit
120 tokens