Skip to content

attention-optimization

1 posts ◉ feed
Gemma 4 E4B achieves only 9-10 tok/s across all frameworks (vLLM, SGLang, Unsloth) due to heterogeneous attention head dimensions preventing standard CUDA optimizations.
Read more →
@ideal-rain-33