DSPy: LLM evaluator incorrectly rejects segments as 'not_quantitative' when fed narrative summary instead of transcript text
LLM segment evaluator rejects quantitative content as 'not_quantitative' because it is fed the segmenter's narrative summary instead of transcript text.
A video-to-calculator pipeline segments a transcript, then a DSPy evaluator grades each segment for whether it holds numbers concrete enough to build a calculator from. not_quantitative was the single largest rejection reason across the corpus (98 of 164 rejections, 60%).
The evaluator never sees the transcript:
# _evaluation.py
prediction = predictor(
source='youtube',
title=title,
transcript_segment=segment.creator_context, # <-- a ~150-word prose SUMMARY
creator_name=creator_name,
)The InputField is documented as "Transcript segment text with quantitative content" and is fed the segmenter's first-person narrative summary. Whether any dollar figure reaches the evaluator is luck: the summary conveys what the creator discussed, not the numbers.
Confirmed case: a segment whose summary read "...from groceries and health insurance to my massive homeowners and flood insurance bills" (no figures) came back is_good_fit: false, rejection_reason: not_quantitative, financial_metrics: {}, confidence: 0.95 — against a 2,581-word cleaned transcript carrying 38 dollar figures, including a fully itemized monthly budget (rent $800, health insurance $175, utilities $45, groceries $400, $2,855/mo total) and the comparison case. A textbook sum-of-line-items calculator, rejected at 0.95.
Two things make this hard to spot:
- The rejection is locally correct. Given the input it received,
not_quantitativeat high confidence is the right answer. Nothing errors, nothing reads as miscalibrated, and prompt-tuning the evaluator cannot fix it. - The segment record looks like it should slice. Segments carry
start_time: "02:44"/end_time: "07:22"and no text field, which reads like a time-window reference into the transcript — but the stored transcript (raw AND cleaned) is flattened prose with no timestamps, so those offsets cannot resolve for anyone who wires them up later.
General detection heuristic: when an LLM judge's job is to detect the presence of a feature (numbers, citations, code) in a document, assert that the feature is present in the string you actually pass it. A judge whose most common verdict is "feature absent" may be reading a summary of the document rather than the document.
Pass the document, not a summary of it.
Minimal fix: send the cleaned transcript alongside creator_context (2,581 words fits comfortably inside the existing max_tokens=16000 LM), so financial_metrics is extracted from real spoken figures instead of narrative prose.
Structural fix: preserve timestamps through the transcript-cleanup stage so the start_time/end_time the segmenter already emits can slice the actual segment. Today they are dead metadata — cleanup strips timestamps, so the offsets reference a coordinate system no stored artifact has.
What NOT to do first: re-tune the evaluator prompt or swap models. The verdict is correct for the input it received; the bug is the argument. Before shipping a fix to a judge's calibration, print the exact string the judge received and grep it for the feature being judged.
Guardrail worth adding: when a judge returns feature-absent AND its extraction field comes back empty (financial_metrics: {}), assert the feature against the source document and log the mismatch. That converts a silent quality ceiling into a loud one.