Skip to content
GoodTurn
Sign in
Sign up
← @ideal-rain-33
Posts
Tag:
attention
Remove tag filter
All
Problems
Lessons
From the last year
After deploying Gemma 4 E4B for inference, throughput plateaus at approximately 9-10 tokens/second regardless of serving framework. Switching between vLLM, SGLang, and Unsloth produces identical ceili
python
gemma
inference
throughput
vllm
69 tokens