ADR 0004 — Pipecat, not LiveKit, for the voice pipeline

2026-08-29 · Status: accepted** · Full workings: strategy/31

Context

Both frameworks were built as working POCs and driven with live calls. The choice needed to be made on evidence rather than reputation.

Decision

Pipecat. Reasons in order of weight:

  1. Licence. LiveKit's turn detector is Apache-2.0 with a §3(i) restriction that forbids standalone use outside their SDK. We had provisionally planned exactly that. Verified against the licence text.
  2. Frames are the right abstraction here. Money safety needs a processor sitting immediately before the voice, seeing final text after substitution and transliteration. Pipecat's pipeline makes that a ten-line processor.
  3. Native OpenTelemetry, conversation/turn spans, and a behavioural eval harness we now run as a regression suite.

Consequences

  • LiveKit's POC is archived at poc/_archive/lk-livekit-poc, not deleted.
  • We inherit Pipecat's defaults, and its defaults have bitten us three times (text-mode eval bypassing our processor, English TTS/STT in a Hindi suite, float32 Whisper on CPU). Treat every framework default as unset.
  • An earlier version of this decision cited a latency advantage that turned out to be a measurement artifact. Pipecat's real median (~1,184 ms) is close to LiveKit's 939 ms; latency is not a reason for this choice.

Revisit when

Telephony at 8 kHz is tested end to end, or concurrency requirements exceed what one process per call can serve.