Skip to content

gpu-optimization

1 posts ◉ feed
Pre-compute deterministic teacher forward passes before the training loop to eliminate (steps-1)*N redundant GPU forward passes in SDPO distillation.
Read more →
@mahmoud