ADR 0005 — Build evaluation in-repo; buy nothing
2026-08-29 · Status: accepted** · Full workings: strategy/32
Context
Seven research passes over the eval-vendor landscape (Cekura, 3CLogic, Hamming, Coval, Roark, Braintrust, W&B Weave) and the academic literature.
The governing finding: no benchmark, dataset or rubric exists anywhere for multi-turn spoken Hinglish conversation quality. Every mainstream speech-agent benchmark is English or English+Chinese. AI4Bharat's speech work is accuracy-only. No Indic negotiation corpus exists in any modality.
Two 2026 results contraindicate the obvious shortcut: LLM perplexity does not track native-speaker Hinglish preference, and LLM judges systematically overestimate quality on non-Latin scripts. Stanford's Talk Arena (7,500 interactions) found no static benchmark correlates with user preference above ρ ≈ 0.33.
Decision
Build. At ~10 calls/day the economics that force sampling in a contact centre do not apply — we score 100% of calls. Components taken rather than bought: EVA-Bench's accuracy/experience split (κ = 0.78–0.85) as the human rubric, PIER (ICASSP 2025) for scoring the recogniser only where the money is.
A human listening is the primary evaluator, not a stopgap. All 30+ closed feedback items came from a person; none from a test.
Consequences
rate_call.py,pier_wer.py,bakeoff.pyand the Pipecat regression suite are ours to maintain.- Any automated judge we build must be validated against human ratings and the correlation published in this repo. An uncalibrated Hinglish judge is worse than none.
- ⚠ Blocking dependency: 20+ human-rated calls. Currently zero.
Revisit when
A vendor demonstrates real Hindi negotiation calls, or call volume exceeds what one person can listen to.