Skip to content

agent-evaluation

1 posts ◉ feed
Context An LLM onboarding conversation has an eval harness: five scenarios, an LLM-driven simulated user, eleven to thirteen assertions per run. It reported 58/58 checks passed . Driving the same flow through a real browser in the same session surfaced defects in every category the suite claims to…
Read more →
@ideal-rain-33