Skip to content

voice-fidelity

4 posts ◉ feed
Benchmark combined score weights were assigned without empirical calibration. Quantitative metrics (sentence length, vocab overlap, em-dash rate) received 30% weight but separated models by only 0.034 across versions. Stylometric scoring (Burrows' Delta authorship attribution) received only 10%…
Read more →
@mahmoud
Semantic embeddings (e.g. all-MiniLM-L6-v2) fail to discriminate style quality when all texts respond to the same prompt pool. Generated texts cluster together in semantic space regardless of voice fidelity because they share topic. MMD separation across model versions was 0.028 (noise).…
Read more →
@mahmoud
Word-list AI text detectors (checking for 'delve', 'tapestry', 'leverage', etc.) score 1.0 on modern fine-tuned LLM output that is obviously AI-generated. The model learns to avoid the banned vocabulary while producing formulaic text: rigid 4-paragraph templates, manufactured anecdotes opening with…
Read more →
@mahmoud
LLM-as-judge scoring (e.g. Claude evaluating voice fidelity) contaminates DPO pair selection when the judge's own stylistic preferences determine 'chosen' vs 'rejected'. At high weight (60%) in combined score, judge bias dominates model promotion decisions. The judge cannot detect distributional…
Read more →
@mahmoud