A 13-config eval of typed structured extraction: the thinking-disabled incumbent won on accuracy per dollar, and all six models hallucinated document dates the same way.
BLUF. Across 13 model/thinking configs on a typed structured-extraction task, the two-generation-old cheap model with reasoning turned off (
gemini-2.5-flash+reasoning_effort='disable') landed inside the accuracy noise band of every newer and pricier config while costing $0.0113 per document at 13.6s: about 1/7th the cost ofgemini-3.6-flashand 1/11th ofclaude-sonnet-4-6. Turning thinking on made the same model worse (2.9x latency, 2.3x cost, junk over-extractions). The second finding is the one that changed production code: all six models hallucinate the document's date the same way, inferring a month from content clues and fabricating the year from training priors. 80 of 88 date estimates fell outside a generous 90-day pre-publication window (median miss 272 days, worst 823 days), and no model choice fixes it; a deterministic clamp against a known publication date does.
Context
We extract structured financial data out of long user-submitted documents: free-form personal spending logs, a week or so of narrated purchases wrapped around a summary of the author's finances. One document becomes one typed object: demographics, the people whose money is described, accounts and debts, recurring monthly expenses, per-purchase rows, per-day totals, and a weekly total. (The code calls these documents "diaries", which is why the identifiers below say so.)
The extraction is a single typed LLM call. There is no agent loop, no tool use, no retrieval; a dspy.Predict over a signature whose output field is a Pydantic model, with litellm underneath talking to Gemini or Anthropic.
class DiaryExtractionSignature(dspy.Signature):
"""Extract ALL financial data from a personal money diary into structured JSON."""
diary_text: str = dspy.InputField(desc="Full money diary text including demographics and daily spending log")
extraction: DiaryExtraction = dspy.OutputField(desc="Structured extraction of every financial detail in the diary")The inputs are big and messy. Across our 10 labelled fixtures, raw text runs from 4.9 KB (the shortest, a terse forum post) to 63 KB (the longest, a two-earner household log), median 28 KB. Output is a nested object with up to seven list fields, so max_tokens has to be generous (32,768 in production; truncation shows up as an inscrutable JSON parse error rather than a clean length stop).
Two things make model choice a real decision rather than a vibe. First, this runs over a corpus, not one document: the local cache alone holds 2,044 of them, so a 7x cost difference per document is a 7x line item. Second, the same default_fast_model config key feeds 31 dspy.LM( callsites across 9 backend files (document import, view generation, rate scraping, feed scoring, CSV import), so "just upgrade to the new model" is not a local change. We had also drifted: production and staging were extracting with gemini-3.5-flash while dev used gemini-2.5-flash, and nobody could say which was better because the choice had been made on a single document.
Methodology
The harness is a CLI command (ff biz diary eval) that runs a model x reasoning matrix and writes a markdown report. Two evaluation modes:
Golden fixtures (labelled, n=10). Seven documents that share a publisher's house template and three freeform forum posts, hand-labelled with evidence-quoted assertions. A fixture "passes" only if every assertion holds, which makes the golden pass rate a harsh conjunctive metric: one bad field fails the whole document. Assertions are mechanical, not stylistic: occupation_contains, salary_range, a list of required holdings by substring and type, min_transaction_count, min_holding_count, max_holding_count, min_monthly_expense_count, min_holder_count.
Label-free batch (n=60). A deterministic sample of 60 documents from the local cache, fixtures excluded. No hand labels, so the metrics are proxies that scale: does it parse, is it internally consistent, how many of the dollar amounts in the text made it into the output.
The metrics, precisely:
- Golden pass (n/10). All assertions for that fixture hold, including the two label-free checks and the date-window check below.
- Assertion failures. Total individual failures across all 10 fixtures, the finer-grained signal. A config can fail 3 fixtures by 1 assertion each or by 5 each; the difference matters for judging noise.
- Parse rate. Fraction of the 60 batch documents where the model output validates as
DiaryExtractionafter sanitization. This is the reliability floor: a document that does not parse produces nothing. - Sum-consistency. An internal check with no labels needed: the day totals the model emitted must add up to the weekly total the model emitted, within $1. This catches arithmetic drift and dropped days without anyone hand-labelling a document.
- Dollar-amount recall. Distinct dollar figures of $1 or more in the source text (regex), divided into the numeric leaves of the extraction dict. It is a recall proxy for "did you actually transcribe the numbers", deliberately with no pass/fail threshold; it is only meaningful comparatively, since these documents mention amounts that are not transactions.
- Date-window. If the fixture has a known publication date and the model emitted
estimated_start_date, the week it describes must start within 90 days before publication and not after it. These documents are written days to weeks before they are published, so this is a loose sanity bound, not a tight accuracy test. - $/doc and latency. Latency is wall clock with the dspy cache off. Cost is computed from reported prompt/completion tokens against a price table (official Gemini pricing checked 2026-08-03; thinking tokens bill at output rates). The Anthropic rows in that table are unverified list prices and were display-only, never a decision input.
The matrix was six models crossed with thinking modes: {gemini-2.5-flash, gemini-3.5-flash, gemini-3.6-flash, gemini-3.5-flash-lite, claude-haiku-4-5, claude-sonnet-4-6} x {default, disable}, plus one middle point (budget:1024, a 1024-token thinking budget) that only gemini-2.5-flash supports. That is 13 configs, so 130 golden extractions, plus 240 batch extractions for the four batched configs. Everything ran at temperature=0, max_tokens=32768, cache=False, one attempt per document, extractions cached to disk per config so the report is reproducible without re-billing.
One provider wrinkle worth knowing: "turn thinking off" is not a portable string. litellm takes reasoning_effort='disable' for Gemini and 'none' for Anthropic.
# ff/src/ff/biz/_diary.py
def _eval_reasoning_kwargs(model: str, spec: str) -> dict | None:
"""dspy.LM kwargs for a reasoning spec. None means skip this (model, spec) pair."""
if spec == 'default':
return {}
if spec == 'disable':
return {'reasoning_effort': 'none' if model.startswith('anthropic/') else 'disable'}
if spec.startswith('budget:'):
if not model.startswith('gemini/gemini-2.5'):
return None
return {'thinking': {'type': 'enabled', 'budget_tokens': int(spec.split(':', 1)[1])}}
raise ValueError(f'unknown reasoning spec: {spec} (use default, disable, or budget:N)')Results
Golden fixtures, n=10, max_tokens=32768. Parse rate is blank for configs that were not run through the 60-document batch. Date-window denominators differ (7 fixtures have a known publication date; a 6 means the model returned no start date at all for one of them, so the check was skipped).
| gemini-2.5-flash | disable | 3/10 | 11 | 1/7 | 10/10 | 60/60 | 0.71 | $0.0113 | 13.6s |
| gemini-2.5-flash | default | 3/10 | 10 | 2/7 | 9/9 | 60/60 | 0.66 | $0.0257 | 38.9s |
| gemini-2.5-flash | budget:1024 | 2/10 | 14 | 0/7 | 7/9 | 0.70 | $0.0177 | 31.8s | |
| gemini-3.5-flash | disable | 2/10 | 16 | 0/7 | 7/10 | 0.67 | $0.0630 | 21.6s | |
| gemini-3.5-flash | default | 3/10 | 13 | 1/7 | 10/10 | 0.69 | $0.1045 | 41.0s | |
| gemini-3.5-flash-lite | disable | 4/10 | 14 | 1/6 | 6/9 | 60/60 | 0.69 | $0.0190 | 19.5s |
| gemini-3.5-flash-lite | default | 1/10 | 16 | 0/6 | 7/10 | 0.68 | $0.0175 | 20.3s | |
| gemini-3.6-flash | disable | 3/10 | 14 | 0/6 | 7/9 | 0.67 | $0.0583 | 21.4s | |
| gemini-3.6-flash | default | 4/10 | 10 | 1/7 | 10/10 | 60/60 | 0.68 | $0.0792 | 35.9s |
| claude-haiku-4-5 | disable | 1/10 | 15 | 0/7 | 10/10 | 0.70 | $0.0360 | 36.2s | |
| claude-haiku-4-5 | default | 1/10 | 15 | 0/7 | 10/10 | 0.71 | $0.0377 | 40.1s | |
| claude-sonnet-4-6 | disable | 4/10 | 11 | 1/7 | 10/10 | 0.68 | $0.1243 | 90.6s | |
| claude-sonnet-4-6 | default | 4/10 | 11 | 1/7 | 10/10 | 0.70 | $0.1262 | 91.5s |
The 60-document label-free batch, run for the top golden configs plus the shipped baseline:
| gemini-2.5-flash | disable | 60 | 60/60 | 49/52 (94%) | 0.78 | 12.6s | $0.0104 |
| gemini-2.5-flash | default | 60 | 60/60 | 49/52 (94%) | 0.74 | 49.1s | $0.0257 |
| gemini-3.5-flash-lite | disable | 60 | 60/60 | 40/53 (75%) | 0.76 | 17.7s | $0.0156 |
| gemini-3.6-flash | default | 60 | 60/60 | 52/54 (96%) | 0.75 | 36.2s | $0.0763 |
The decision rule was fixed before the run: rank by fixtures passed, then fewest assertion failures, then date-window rate, then dollar recall; among configs within 1 fixture-pass and 2 assertion failures of the leader, take the cheapest, breaking ties by latency. The leader was gemini-3.6-flash + default (4/10, 10 failures). Three deploy candidates landed in band (2.5-flash + disable at $0.0113, 2.5-flash + default at $0.0257, 3.6-flash + default at $0.0792; Sonnet was in band on numbers but had been declared a reference ceiling, not a deploy candidate, before the run). The cheapest of the three was the model we were already running in dev.
Deep-dive 1: the cheap incumbent won
gemini-2.5-flash with reasoning disabled passed 3 of 10 golden fixtures against the leader's 4, with 11 assertion failures against 10. That is one assertion apart on a 10-document sample; it is not a real difference. Everything else favored the cheap config: 10/10 sum-consistency on golden, 60/60 parses on the batch, the best dollar recall at scale (0.78, higher than 3.6-flash's 0.75 and flash-lite's 0.76), 12.6s per document and $0.0104. The leader cost $0.0763 and took 36.2s for that one extra fixture.
The instructive comparison is gemini-3.5-flash, which Google positions as its "most intelligent" flash-tier model for agentic and coding work, and which our production config was quietly using. With thinking disabled it was strictly worse than 2.5-flash on every extraction metric we measured: 2/10 vs 3/10 passes, 16 vs 11 assertion failures, 0/7 vs 1/7 date windows, 7/10 vs 10/10 sum-consistency, 0.67 vs 0.71 dollar recall. It cost $0.0630 against $0.0113, about 5.6x, because its input price is 5x and its output price 3.6x. With thinking on it reached parity on passes (3/10) at $0.1045, roughly 9x the baseline. Production moved down a model generation as a result.
The Anthropic rows are the other useful data point. claude-haiku-4-5 was the worst config in the sweep (1/10 in both modes) despite clean parses and perfect sum-consistency; it failed on content, missing required holdings and undercounting transactions. claude-sonnet-4-6 tied for the golden lead (4/10) with the best-in-class discipline you would expect, and took 90 seconds per document at $0.1243. It was included as a reference ceiling, never a deploy candidate, and it is worth saying plainly: the ceiling was one fixture above the floor-priced config. Whatever a frontier model is buying on this task, it is not transcription accuracy.
gemini-3.5-flash-lite deserves a note because it shows why the label-free batch existed. It led the golden table at 4/10 and looked like a cheap winner. At n=60 its sum-consistency collapsed to 40/53 (75%) against the winner's 94%. On a 10-document labelled set that failure mode was invisible; a 60-document unlabelled run priced at a few dollars caught it.
Deep-dive 2: thinking hurt
Same model, same prompt, same fixtures, three reasoning settings:
| Thinking | Golden | Failures | Sum-consistency | Latency | $/doc |
|---|---|---|---|---|---|
| disable | 3/10 | 11 | 10/10 | 13.6s | $0.0113 |
| budget:1024 | 2/10 | 14 | 7/9 | 31.8s | $0.0177 |
| default | 3/10 | 10 | 9/9 | 38.9s | $0.0257 |
Default thinking bought nothing measurable and charged 2.9x the latency and 2.3x the cost for it (thinking tokens bill at output rates, so a chatty reasoning pass is expensive twice over: once in tokens, once in wall clock). On the single-document probe that started this investigation, default thinking also actively degraded output: 19 holdings extracted where thinking-off found 12, and 8 of the 19 were junk ("Parents' EZ pass" and "Parents' vacations" filed as financial holdings). The model reasoned its way into inventing accounts out of narrative mentions of other people's spending. Every variant found the same 10 real accounts, so the "19 vs 11 holdings" gap we originally read as recall loss was over-extraction in the other direction.
The middle setting was the worst of the three. A 1024-token budget scored 2/10 with 14 assertion failures, the worst date accuracy in the entire sweep (0/7), and two sum-consistency failures the other two settings did not have, at 2.3x the latency of thinking-off. Our hypothesis had been that a small budget might recover multi-holder extraction (thinking-off sometimes collapses a two-earner household to a single person) without the verbosity tax. It recovered nothing. A truncated reasoning pass appears to be worse than either a full one or none at all.
The mechanism we believe, stated as a hypothesis rather than a finding: extraction is transcription, not deduction. The answer is already in the context window verbatim. Reasoning tokens give the model room to interpret, and interpretation on a transcription task is pure downside; it invents structure that the text merely implies. That comment now lives in the production call site:
# fsrv/src/fsrv/impex/service.py
# reasoning_effort='disable' keeps Gemini's thinking tokens from eating the
# output budget. Extraction is transcription, not deduction, and thinking
# roughly doubled latency for no parse-reliability gain.
lm = dspy.LM(model=model_name, api_key=api_key, max_tokens=DIARY_EXTRACT_MAX_TOKENS,
temperature=0, reasoning_effort='disable',
metadata={"subsystem": "diary_import"})Deep-dive 3: every model hallucinates dates the same way
This is the finding that outlived the model decision.
The extraction schema asks for estimated_start_date: the Monday of the week the document describes. The documents label their days ("Day One: Monday") and never give a calendar date, so the model has to infer one. The check is deliberately loose: the start must fall within 90 days before publication and not after it.
Across the full sweep, 80 of 88 date checks failed. Not 80 of 88 documents; 80 of 88 (config, fixture) pairs, where 7 fixtures have a known publication date and 13 configs each got a shot at them (3 of the 91 possible pairs returned no date at all, so the check was skipped). The best config in the sweep managed 2/7. Date failures were 80 of the 170 total assertion failures recorded across all configs, so nearly half of everything the eval flagged was this one field.
The misses are not small. Median 272 days (about 9 months), mean 338 days, worst 823 days (2.25 years). The minimum miss was 92 days, which means not a single wrong answer was merely a couple of days outside the bound.
The structure of the errors is the interesting part. Errors are correlated across models, not independent:
| A (2025-09-01) | 12 | 2023-06-01 .. 2024-06-17 (all 12 in June) |
| B (2026-03-04) | 13 | 2025-06-01 .. 2025-07-28 (11 in June) |
| C (2026-02-25) | 13 | 2025-06-01 .. 2025-09-01 (9 in June) |
| D (2026-02-04) | 13 | 2025-06-01 .. 2025-08-25 (8 in June) |
| E (2025-05-19) | 7 | 2024-03-17 .. 2024-03-18 (7 within one day of each other) |
| F (2025-08-27) | 11 | 2023-06-03 .. 2024-09-01 |
| G (2026-07-29) | 11 | 2026-01-01 .. 2026-04-28 |
Of the 80 wrong dates, 43 land in June. The years cluster too: 39 in 2025, 22 in 2024, 11 in 2026, 8 in 2023, for documents published between May 2025 and July 2026. Gemini, flash-lite, Haiku and Sonnet, thinking on and off, converge on the same wrong answers.
The mechanism is visible in the source text. Document A, published September 2025, contains this aside about the weather where its author lives:
They all leave to go back up north around April, when it becomes hot again. In June it is unbelievably hot and humid, and also stormy in the evenings, so people are less likely to come out to [...]
That sentence is not about the week being logged; it is background about the city. Every one of the 12 configs that produced a date for it put the week in June, then scattered the year across 2023 and 2024 for a document published in September 2025. That is the pattern in miniature: the month comes from a content clue in the text (weather, a holiday, a seasonal reference), and the year comes from nowhere in the text at all. It comes from the model's prior about what year it is, which is its training cutoff, which is not the year your document was published. Document E is the cleanest evidence for a shared prior: seven different configs, three different providers, landing on 2024-03-17 or 2024-03-18, exactly 14 months before publication.
The practical read: an LLM date estimate is a fluent guess with no error bar, and it is systematically biased toward the training era. Ensembling does not help, because the models agree. Picking a smarter model does not help, because Sonnet is on the same wrong dates as flash-lite. Prompting harder is a coin flip we did not want to keep flipping in production.
What helps is refusing to use the number when you have a better one. Publication date is ground truth we already store, so the fix is a deterministic clamp: an LLM start date more than 120 days before publication (or more than a week after it) is discarded in favor of the publication date, with the provenance recorded so the UI can tell the difference. Documents with no known publication date keep the LLM value, because there is nothing better to fall back on.
# fsrv/src/fsrv/impex/service.py, inside extract_diary()
# Date sanity guard: models sometimes hallucinate a diary week months away
# from publication (3.5/3.6-flash produced April dates for a July diary).
# The publication date is ground truth when we have one; an extracted start
# outside [pub - 120d, pub + 7d] is discarded in favor of it. Manual/casual
# diaries (no source_published_date) keep LLM dates untouched.
pub = diary.source_published_date
if pub is not None and llm_start is not None and not (
pub - datetime.timedelta(days=120) <= llm_start <= pub + datetime.timedelta(days=7)):
llm_start = None
if llm_end is not None and not (
pub - datetime.timedelta(days=120) <= llm_end <= pub + datetime.timedelta(days=7)):
llm_end = None
if llm_start is not None and diary.start_date_source != 'user':
diary.start_date = llm_start
diary.start_date_source = 'extraction'
if llm_end is not None and diary.end_date is None:
diary.end_date = llm_end
# Fallback: publication date when the LLM start date is missing or was discarded
if diary.start_date_source not in ('user', 'extraction') and pub:
diary.start_date = pub
diary.start_date_source = 'publication'The window is asymmetric on purpose (120 days back, 7 days forward) because these documents are written before they are published, never after. It is wider than the eval's 90-day check so that the guard clamps only clear hallucinations rather than plausible-but-old submissions. The start_date_source field ('user', 'extraction', 'publication') is the part worth copying: when you overwrite a model's output with a deterministic value, record which one you used, or you will not be able to measure the model again later.
Limitations
- n=10 labelled fixtures. One fixture is 10 percentage points. Every "beat" in this post that rests on a single fixture (3/10 vs 4/10) is inside the noise band, which is exactly why the decision rule treated a 1-fixture gap as a tie and fell through to cost. Do not read the golden column as a leaderboard.
- Conjunctive pass metric. A fixture passes only if all of its assertions hold, so absolute pass rates look low and are dominated by the date check. For the winner, 6 of its 7 failing fixtures failed on a date. Read assertion failures and the per-metric columns, not just passes.
- The batch is label-free. Parse rate, sum-consistency and dollar recall are proxies. A model can be perfectly self-consistent and wrong, and dollar recall counts amounts that were never transactions.
- One task class. Financial extraction from long English prose into a nested Pydantic model. The conclusion "thinking hurts" is a claim about transcription-shaped tasks, not about summarization, classification, or anything with a deduction step. Do not generalize it to your agent loop.
- Point-in-time. Model versions, prices and defaults as of 2026-08-03. The Anthropic prices in the cost column are unverified list prices used for display; the Gemini prices are from the official pricing page on that date. Costs are computed from reported token counts, not from billing invoices.
- One attempt per document, temperature=0. No retries, no self-consistency sampling, no variance estimate across repeated runs of the same config. Some of the 1-fixture gaps are probably run-to-run noise rather than config differences, and we did not measure that.
- Sample bias. Seven of ten fixtures come from one publisher and share a house template, so their structure is correlated. The three freeform forum posts are the only counterweight, and they have no publication date, so they contribute nothing to the date finding.
Conclusions
- Eval the cheap incumbent with thinking off before you pay for a newer tier. Our production config had drifted to a model that was 5.6x more expensive and measurably worse at the actual job. Nobody made that decision; it accreted from "newer is better".
- Intelligence tiers do not transfer to transcription accuracy. Sonnet-class discipline bought one fixture over a two-generation-old flash model at 11x the price and 6.7x the latency. If your task is "copy what is in the document into these fields", the frontier is not where the wins are.
- Reasoning is a cost multiplier on transcription tasks, and sometimes a correctness penalty. Thinking tokens bill at output rates and can push a model from transcribing into interpreting, which is how "Parents' EZ pass" becomes a financial holding. A partial thinking budget was worse than either extreme.
- Never trust an LLM date without a deterministic anchor. The failure is systematic, cross-vendor, and correlated: month from content clues, year from training priors. If you have any ground-truth timestamp (publication date, file mtime, ingest time, email header), clamp against it and record which source you used.
- Cheap label-free checks at scale catch what small labelled sets miss. Sum-consistency over 60 unlabelled documents eliminated a config that led the labelled table.
Next
- Re-eval on a cadence, not on release notes. The harness makes a re-decision a single command, so the trigger for reconsidering the model should be a scheduled run, not a vendor blog post. Flash-tier models do get retired (Google shut down all
gemini-2.0-flash*in June 2026), so this is technical debt with an unknown due date. - Generalize the guard. The date clamp is currently local to this one extraction path. Anywhere an LLM emits a date that we can bound with a known timestamp deserves the same treatment plus a provenance field.
- Grow the fixture set past the noise band. At n=10, one assertion decides the ranking. Doubling the labelled set, especially with source formats outside the dominant template, is worth more than adding another model column.
- Sweep the tiers we skipped.
gemini-3.1-flash-liteis reachable and cheaper than everything tested here; it was left out of this matrix. The open question is whether the cheapest tier available holds parse rate and sum-consistency at corpus scale, which is exactly where3.5-flash-litebroke. - Measure run-to-run variance. Repeating one config three times would tell us how much of a 1-fixture gap is real.
Appendix: reproducing this
Fixture shape. Fixtures are plain dicts in a module beside the eval; the corpus files they point at are local-only (gitignored). Source URL and cache path are redacted here; everything else is verbatim.
{
'source': 'article',
'cached_text_path': '<local cache path>',
'published_date': '2026-02-25',
'assertions': {
'demographics': {
'occupation_contains': 'Data Engineer',
'salary_range': (90000, 97000),
},
'holdings': [
{'name_contains': 'stipend', 'holding_type': 'income', 'required': True},
{'name_contains': 'car loan', 'holding_type': 'debt', 'required': True},
{'name_contains': 'student loan', 'holding_type': 'debt', 'required': True},
{'name_contains': '401k', 'holding_type': 'retirement', 'required': True},
{'name_contains': 'HSA', 'holding_type': 'hsa', 'required': False},
],
'min_transaction_count': 12,
'min_holding_count': 8,
'min_monthly_expense_count': 10,
'min_holder_count': 1,
},
},The label-free checks. These are the two functions worth stealing; neither needs a labelled corpus.
def _dollar_recall(diary_text: str, parsed: dict) -> float | None:
"""Fraction of distinct dollar amounts (>= $1) in the diary text found among the
numeric leaves of the extraction. Comparative metric only."""
amounts: set[float] = set()
for m in re.findall(r'\$([\d,]+(?:\.\d{1,2})?)', diary_text):
try:
val = round(float(m.replace(',', '')), 2)
except ValueError:
continue
if val >= 1.0:
amounts.add(val)
if not amounts:
return None
leaves: set = set()
_collect_numeric_leaves(parsed, leaves)
return sum(1 for a in amounts if a in leaves) / len(amounts)
def _consistency_check(parsed: dict) -> tuple[bool | None, str | None]:
"""Internal-consistency check: sum(day_totals) must match weekly_spending_total."""
weekly = parsed.get('weekly_spending_total')
day_totals = parsed.get('day_totals') or []
if weekly is None or not day_totals:
return None, None
try:
total = sum(float(d.get('total') or 0) for d in day_totals if isinstance(d, dict))
except (TypeError, ValueError):
return False, 'unparseable day_totals'
if abs(total - float(weekly)) <= 1.0:
return True, None
return False, f'day totals {total:.2f} != weekly total {float(weekly):.2f}'The date-window check.
def _date_window_check(published_date: str | None, parsed: dict) -> tuple[bool | None, str | None]:
"""The diary week must start within 90 days before publication, never after."""
start_s = parsed.get('estimated_start_date')
if not published_date or not start_s:
return None, None
...
return False, f'start_date {start} outside window ({published} - 90d .. {published})'Per-config caching. Extractions are cached to disk with the config baked into the filename, so re-running the report costs nothing and comparing two configs never reads the other's output.
def _eval_cache_path(cached_path: str, model: str, spec: str, max_tokens: int) -> str:
suffix = f"_eval_cache__{model.replace('/', '_')}__{spec.replace(':', '')}__{max_tokens}.json"
return cached_path.replace('.json', suffix)The LM under test. Note cache=False: dspy's cache would make the latency column a lie.
lm = dspy.LM(model=model, api_key=api_key, max_tokens=max_tokens, temperature=0,
cache=False, metadata={'subsystem': 'diary_eval'}, **reasoning_kwargs)
extractor = DiaryExtractorModule(lm=lm)Rerunning. The golden matrix and then the two batch-mode passes:
poetry -C fsrv run -- ff biz diary eval \
--models gemini/gemini-2.5-flash,gemini/gemini-3.5-flash,gemini/gemini-3.5-flash-lite,gemini/gemini-3.6-flash,anthropic/claude-haiku-4-5,anthropic/claude-sonnet-4-6 \
--reasoning disable,default --rerun \
--report-path biz_data/eval_report_2026-08.md
poetry -C fsrv run -- ff biz diary eval \
--models gemini/gemini-2.5-flash --reasoning budget:1024
poetry -C fsrv run -- ff biz diary eval \
--models gemini/gemini-2.5-flash,gemini/gemini-3.6-flash --reasoning disable,default --batch 60Total spend for the sweep was in the single-digit dollars, which is the real argument for running one: the eval cost less than a week of guessing wrong about the model.