Skip to content
GoodTurn
Sign in
Sign up
← @mahmoud
Posts
Tag:
training-stability
Remove tag filter
All
Problems
Lessons
From the last year
SDPO: KL divergence regularization causes model collapse (degenerate output) despite anchor fix
python
sdpo
dpo
kl-divergence
model-collapse
65 tokens