ADR 0005 — Build evaluation in-repo; buy nothing

2026-08-29 · Status: accepted** · Full workings: strategy/32

Context

Seven research passes over the eval-vendor landscape (Cekura, 3CLogic, Hamming, Coval, Roark, Braintrust, W&B Weave) and the academic literature.

The governing finding: no benchmark, dataset or rubric exists anywhere for multi-turn spoken Hinglish conversation quality. Every mainstream speech-agent benchmark is English or English+Chinese. AI4Bharat's speech work is accuracy-only. No Indic negotiation corpus exists in any modality.

Two 2026 results contraindicate the obvious shortcut: LLM perplexity does not track native-speaker Hinglish preference, and LLM judges systematically overestimate quality on non-Latin scripts. Stanford's Talk Arena (7,500 interactions) found no static benchmark correlates with user preference above ρ ≈ 0.33.

Decision

Build. At ~10 calls/day the economics that force sampling in a contact centre do not apply — we score 100% of calls. Components taken rather than bought: EVA-Bench's accuracy/experience split (κ = 0.78–0.85) as the human rubric, PIER (ICASSP 2025) for scoring the recogniser only where the money is.

A human listening is the primary evaluator, not a stopgap. All 30+ closed feedback items came from a person; none from a test.

Consequences

  • rate_call.py, pier_wer.py, bakeoff.py and the Pipecat regression suite are ours to maintain.
  • Any automated judge we build must be validated against human ratings and the correlation published in this repo. An uncalibrated Hinglish judge is worse than none.
  • ⚠ Blocking dependency: 20+ human-rated calls. Currently zero.

Revisit when

A vendor demonstrates real Hindi negotiation calls, or call volume exceeds what one person can listen to.