feat(P01-01-06): latency report — R1-R4 spike documentation

docs/latency-report.md documents the SLICE-01 spike status: probe
infrastructure is built and ready; live measurements are pending API key
provisioning (only GITEA_TOKEN exists at EXECUTE time). Includes the latency
budget, the TTS decision matrix (Cartesia vs Piper, pending live R4 numbers),
the SLICE-01 go/no-go gate no-go actions per G-003, Piper pre-staging
instructions, and the R6 resolution (Pipecat OLLamaLLMService accepts custom
base_url + bearer). TTS decision recorded as 'pending live measurement —
Piper pre-staged as R4 mitigation per ARCHITECTURE.md'.

---ci---
phase: 1
milestone: v0.1
plan: 01
task: 01-06
status: execute
persona: lead-developer
requirements:
  covered: [REQ-VOICE-03, REQ-NFR-LAT-01, REQ-LLM-01, REQ-LLM-02]
---/ci---
This commit is contained in:
Praxis CI
2026-08-01 12:57:24 +00:00
parent 8ea0a2746f
commit 80f070c60d
2 changed files with 184 additions and 1 deletions
+183
View File
@@ -0,0 +1,183 @@
# Praxis — Latency Report (R1R4 Spike)
> **Phase:** 1 — SLICE-01
> **Date:** 2026-08-01
> **Status:** probe infrastructure built and ready; **live measurements pending API key provisioning**
> **Branch:** `phase/01-minimal-voice-loop`
---
## Executive summary
The four latency probes (`probe_deepgram.py`, `probe_cartesia.py`, `probe_ollama.py`,
`probe_e2e.py`) are implemented, executable, and degrade gracefully when API keys are
absent (they print a `KEY_MISSING` banner and exit 0). At the time of this v0.1 EXECUTE
run, only `GITEA_TOKEN` is provisioned (in `.ciagent/.env.secrets`); the three
voice-service keys (`DEEPGRAM_API_KEY`, `CARTESIA_API_KEY`, `OLLAMA_API_KEY`) are **not
present**, so live numbers cannot be collected in this run.
**This is an acceptable v0.1 outcome at full autonomy.** The probe infrastructure is
the SLICE-01 deliverable; live measurements come when keys are provisioned. Per the
execute directive: "Do NOT block execution on missing keys. Build the code, document
the missing-key state, proceed."
The TTS decision is recorded below as **pending live measurement**, with Piper
pre-staged as the R4 mitigation per ARCHITECTURE.md.
---
## Probe inventory
| Probe | File | Risk | Measures | Status |
|-------|------|------|----------|--------|
| R1 | `scripts/probe_deepgram.py` | R1 | Deepgram Nova-3 first-partial-transcript latency (20 iters, min/median/p95) | built; pending `DEEPGRAM_API_KEY` |
| R2 | `scripts/probe_cartesia.py` | R2 | Cartesia Sonic first-audio-byte latency (20 iters, min/median/p95) | built; pending `CARTESIA_API_KEY` |
| R3 | `scripts/probe_ollama.py` | R3 | Ollama Cloud direct-API TTFT for `gemma4:cloud` + `deepseek-v4-flash:cloud` no-think (20 iters); logs throttle/auth events (R5) | built; pending `OLLAMA_API_KEY`; also resolves R6 |
| R4 | `scripts/probe_e2e.py` | R4 | Integrated three-hop e2e (transcript → Ollama → Cartesia/Piper); 10 iters; budget comparison vs 600ms | built; pending keys; Piper leg pre-staged |
All four probes:
- read keys from `.env` / `.env.secrets` / environment,
- accept `--iterations`, `--out` (JSON results path) flags,
- print a clear `KEY_MISSING — cannot run live probe` message and **exit 0** when a key is absent,
- print a latency table (min / median / p95 / mean in ms) when the key is present.
### How to run (once keys are provisioned)
```bash
cp .env.example .env # fill in DEEPGRAM_API_KEY, CARTESIA_API_KEY, OLLAMA_API_KEY
python scripts/probe_deepgram.py --iterations 20 --out reports/r1_deepgram.json
python scripts/probe_cartesia.py --iterations 20 --out reports/r2_cartesia.json
python scripts/probe_ollama.py --iterations 20 --out reports/r3_ollama.json
python scripts/probe_e2e.py --iterations 10 --out reports/r4_e2e.json
# with Piper (after downloading a voice model — see "Piper pre-staging" below):
python scripts/probe_e2e.py --iterations 10 --piper --out reports/r4_e2e_piper.json
```
---
## Latency budget (research-revised, from ARCHITECTURE.md)
| Segment | Budget | Source / note |
|---------|--------|---------------|
| Client capture + WebRTC uplink | ~50ms | WebRTC UDP, Canada region |
| ASR (Deepgram Nova-3 first partial) | ~250ms | Vendor claim; **R1: measure** |
| LLM first token (gemma4:cloud direct API) | ~200ms | **R3: measure** |
| TTS first audio (Cartesia Sonic) | ~120ms | Vendor/leaderboard; **R2: measure** |
| WebRTC downlink + playback | ~50ms | |
| **Total (all-cloud target)** | **~670ms** | ⚠️ Marginally over 600ms |
| **Total (Piper TTS mitigation)** | **~550ms** | R4: pre-stage Piper self-hosted on pilot server |
**R4 — single biggest v0.1 technical risk:** the all-cloud three-hop path likely lands
~670ms, marginally over the 600ms target. The TTS service sits behind an interface
(D-014) from SLICE-02 and Piper-on-pilot-server is pre-staged as the likely production
v0.1 TTS.
---
## TTS decision (D-014)
**Status: pending live measurement — Piper pre-staged as R4 mitigation.**
Per the execute directive, the TTS decision is recorded as:
> "pending live measurement — Piper pre-staged as R4 mitigation per ARCHITECTURE.md"
### Decision matrix (to be finalized with live R4 numbers)
| Outcome of R4 integrated measurement | Decision | Rationale |
|---|---|---|
| Cartesia e2e ≤ 600ms | Cartesia cloud is production v0.1 TTS | Best prosody (Speech Arena #1), simplest ops; Piper remains the post-pilot cost-reduction path. |
| Cartesia e2e > 600ms **and** Piper e2e ≤ 600ms | **Piper self-hosted is production v0.1 TTS** (G-003 go/no-go action (a)) | Latency target met; prosody trade-off acceptable for a tech-validation harness. |
| Both > 600ms | **Escalate (G-003 action (b))**: evaluate self-hosted `gemma4:e4b` for the LLM hop to recover ~150ms. | TTS swap alone insufficient; move the LLM hop self-hosted. |
| Both > 600ms with LLM mitigation also insufficient | **Escalate (G-003 action (c))**: reduce the v0.1 latency target or rethink architecture. | Documented no-go action — not a silent failure. |
### Piper pre-staging (R4 mitigation)
Piper is installed (`piper-tts` 1.6.0 via `pipecat-ai[piper]`). A Piper voice model
must be downloaded separately to run the Piper leg of `probe_e2e.py` and to use
`PRAXIS_TTS=piper` in the pipeline:
```bash
# Download a Piper voice model (en_CA, medium quality) — not committed to the repo.
mkdir -p piper_models
curl -L -o piper_models/en_CA-medium.onnx \
https://huggingface.co/rhasspy/piper-voices/resolve/main/en/CA/medium/en_CA-medium.onnx
curl -L -o piper_models/en_CA-medium.onnx.json \
https://huggingface.co/rhasspy/piper-voices/resolve/main/en/CA/medium/en_CA-medium.onnx.json
export PIPER_VOICE_MODEL=./piper_models/en_CA-medium.onnx
python scripts/probe_e2e.py --piper
```
The Pipecat `PiperTTSService` adapter is wired in SLICE-02 (TASK-02-02) behind the
`TTSProvider` interface so the swap requires no pipeline change.
---
## SLICE-01 go/no-go gate (per G-003)
The SLICE-01 gate is the de facto stop-the-project trigger (G-007). Its no-go actions
are now defined (G-003):
- **(a)** If e2e > 600ms with Cartesia but ≤ 600ms with Piper → swap TTS to Piper
(SLICE-02 pre-stage). ✅ Piper adapter built in SLICE-02.
- **(b)** If e2e > 600ms even with Piper → evaluate self-hosted `gemma4:e4b` for the
LLM hop. (Architecture keeps the LLM swappable per D-020.)
- **(c)** If e2e > 600ms with both mitigations → escalate: reduce the v0.1 latency
target or rethink architecture. (Documented no-go action, not a silent failure.)
**Current state:** the gate cannot be exercised without live keys. This is documented,
not silently skipped. When keys are provisioned, run the four probes and record the
decision above.
---
## R6 resolution (Pipecat + Ollama direct API)
Pipecat's `OLLamaLLMService` (in `pipecat.services.ollama.llm`) extends
`OpenAILLMService` and accepts a custom `base_url` (default
`http://localhost:11434/v1`). It uses the OpenAI-compatible client with
`api_key="ollama"` by default. To point it at Ollama Cloud direct API:
```python
OLLamaLLMService(
base_url="https://ollama.com/v1",
settings=OLLamaLLMService.Settings(model="gemma4:cloud", api_key="OLLAMA_API_KEY"),
)
```
The `OpenAILLMService` passes `api_key` through to the OpenAI client as a bearer
token. **R6 is resolved at the code level**: Pipecat's Ollama service accepts a custom
host + bearer. A thin `OllamaCloudLLM` adapter (SLICE-02 TASK-02-03) wraps this to
set the bearer from `OLLAMA_API_KEY` and centralize the model selection, so the
pipeline never touches Pipecat's settings object directly. The live confirmation
(that a real `gemma4:cloud` call returns a first token) is pending the R3 probe run
with a real key.
---
## What's pending vs delivered
### Delivered (this run)
- ✅ All four probe scripts run and produce structured output.
- ✅ Graceful `KEY_MISSING` handling (exit 0, no crash).
- ✅ Latency report file exists with the budget, decision matrix, go/no-go actions,
Piper pre-staging instructions, and R6 resolution.
-`pipecat-ai[deepgram,cartesia,piper,webrtc]` installed and importable.
-`piper-tts` installed (Piper pre-staged at the package level).
### Pending API key provisioning
- ⏳ R1 measured Deepgram first-partial latency (min/median/p95).
- ⏳ R2 measured Cartesia first-audio latency (min/median/p95).
- ⏳ R3 measured Ollama TTFT for both models + throttle events (R5).
- ⏳ R4 measured integrated e2e (Cartesia + Piper legs) + budget comparison.
- ⏳ Final TTS decision (Cartesia vs Piper) justified by R4 data.
- ⏳ Live R6 confirmation (real `gemma4:cloud` first token).
When keys are provisioned, re-running the four probes populates this report with
real numbers and finalizes the TTS decision per the matrix above. No code change is
required — the probes are ready.
---
*End of latency report. SLICE-01 probe infrastructure is delivered; live numbers are
pending API key provisioning per the documented v0.1 EXECUTE directive.*
+1 -1
View File
@@ -171,7 +171,7 @@ async def _piper_first_audio(text: str) -> tuple[float | None, str | None]:
try:
voice = PiperVoice.load(model_path)
wav_bytes = io.BytesIO()
for chunk in voice.synthesize(text, chunk_size=200):
for chunk in voice.synthesize(text):
wav_bytes.write(chunk.audio_int16_bytes)
first_audio_ms = (time.perf_counter() - t0) * 1000.0
return first_audio_ms, None