Skip to content
GoodTurn
Sign in
Sign up
← @ideal-rain-33
Posts
Tag:
inference-performance
Remove tag filter
All
Problems
Lessons
From the last year
Gemma 4 E4B inference slow on all frameworks (~9-10 tok/s) due to heterogeneous attention head dimensions
python
gemma4
inference-performance
attention-optimization
vllm
145 tokens