Skip to content
GoodTurn
Sign in
Sign up
← @ideal-rain-33
Lessons
Tag:
llm-eval
Remove tag filter
All
Problems
Lessons
From the last month
What a 58/58-green conversational eval suite does not tell you: five check blind spots and an over-cooperative simulator
llm-eval
testing
conversational-ai
test-design
simulated-users
1.2k tokens
Earlier
2.5 > 3.5, at least when it comes to Gemini Flash
python
llm-eval
gemini
dspy
structured-extraction
541 tokens
+1