Praxis CI 7b1b296430 feat(P01-02-05,P01-02-06): React client + latency readout
client/ — React + Vite + TypeScript scaffolded with the Pipecat client SDK
(@pipecat-ai/client-js) and SmallWebRTCTransport
(@pipecat-ai/small-webrtc-transport). useVoiceSession.ts hook manages mic
permission, WebRTC connect, audio playback, live transcript, and a latency
readout (captures the e2e_latency_ms metric the server emits). App.tsx is a
minimal one-page session UI: disclaimer, Start/End buttons, status badge,
latency readout (within/over 600ms budget), live transcript. vite.config.ts
proxies /pipecat + /health to the Python server (port 8789). npm run
typecheck + npm run build pass.

server/latency.py — LatencyObserver (a Pipecat FrameProcessor) timestamps
transcript-ready, LLM-first-token, TTS-first-audio, and playback-start per
turn, computes ASR→TTS-first-audio (the v0.1 latency target), and logs it
to console with a within/over-budget verdict. Wired into the pipeline
between STT/LLM/TTS so it observes without altering the frame stream. 5
unit tests pass (LatencyRecord e2e math + observer construction +
reset_turn). Full server suite: 18 passed.

---ci---
phase: 1
milestone: v0.1
plan: 02
task: 02-05,02-06
status: execute
persona: frontend-engineer,backend-engineer
requirements:
  covered: [REQ-VOICE-01, REQ-VOICE-02, REQ-VOICE-03, REQ-NFR-LAT-01]
---/ci---
2026-08-01 13:08:42 +00:00

Praxis — v0.1 Foundation

Voice-first AI apprenticeship platform. v0.1 is a tech-validation harness (per G-008) for the minimal viable voice loop: a single learner speaks to an AI tutor playing a Customer Service role-play scenario, hears a <600ms-latency response, receives an end-of-session coaching debrief, and has the session logged to SQLite.

Status

Phase 1 (minimal viable voice loop) — code-complete, pending live API keys for runtime verification.

Stack

  • Orchestration: Pipecat (D-017) with Silero VAD + interruptibility
  • ASR: Deepgram Nova-3 streaming (D-013)
  • LLM: Ollama Cloud direct API (D-020) — gemma4:cloud (role-play) + deepseek-v4-flash:cloud no-think (debrief)
  • TTS: Cartesia Sonic (primary, D-014) / Piper (self-hosted, R4 mitigation) — behind an interface
  • Client: React + Vite + WebRTC (Pipecat client SDK, D-015)
  • State: SQLite praxis.db (D-007, single hardcoded learner, no auth)

Layout

server/      Pipecat pipeline, services (TTS/LLM/Guardrail interfaces), scenario runtime, adapters
client/      React + Vite + WebRTC learner surface
scenarios/   YAML scenario definitions (D-018)
db/          SQLite schema, migrations, async store
scripts/     Latency probes (R1-R4), e2e smoke
tests/       Unit + e2e
docs/        Latency report, debrief templates

Quickstart

  1. Copy .env.example.env, fill in DEEPGRAM_API_KEY, CARTESIA_API_KEY, OLLAMA_API_KEY.
  2. Install server deps: pip install -e ".[dev]"
  3. Install client deps: cd client && npm install
  4. Run probes: python scripts/probe_deepgram.py (etc.)
  5. Run server: python -m server
  6. Run client: cd client && npm run dev

See docs/latency-report.md for the R1-R4 spike status and TTS decision.

S
Description
No description provided
Readme 1.7 MiB
2026-08-04 22:36:00 +00:00
Languages
Python 81.4%
Shell 13.8%
TypeScript 4.3%
CSS 0.3%
Dockerfile 0.2%