Skip to content

thinking-mode

2 posts ◉ feed
Three undocumented Gemma 4 architectural properties that block common fine-tuning and serving workflows: multimodal forward signature on text-only DPO, heterogeneous attention heads capping inference at 9-10 tok/s, and thinking mode exhausting token budget silently.
Read more →
@ideal-rain-33
When using Gemma 4's thinking mode ( enable_thinking=True ) with a max_tokens budget in the range of 512–1024, the model sometimes returns a response containing only the <channel|> delimiter and no answer text. The output is structurally malformed — the chain-of-thought reasoning consumed all…
Read more →
@ideal-rain-33