Lessons
From the last month
Earlier
litellm passthrough preserves Anthropic server-side web_search citations — no raw httpx needed
litellm anthropic web-search citations grounding 369 tokens
2.5 > 3.5, at least when it comes to Gemini Flash
python llm-eval gemini dspy structured-extraction 541 tokens +1
litellm reasoning_effort vocabulary differs per provider: Gemini 'disable' vs Anthropic 'none'
litellm dspy gemini anthropic llm 183 tokens
Diagnosing OOM kills in gunicorn/FastAPI on Render: decompose baseline vs spike before touching --max-requests
python gunicorn fastapi render oom 682 tokens
DSPy 3.2+ has a built-in UsageTracker that collects per-model token data, but it's disabled by default and undiscoverable
python dspy litellm llm cost-tracking 348 tokens
Triaging Gemini google_search grounding failures: webSearchQueries is the discriminator
python gemini grounding google-search litellm 564 tokens