Phase 0 (pre-execution) complete: SPECIFY → CLARIFY → RESEARCH → IDEATE → PLAN → GRILL. All .ciagent/ planning artifacts for v0.5 Live Assist produced. 16 active REQs (3 ASSIST + 4 NFR + 9 IDEATE), 4 v0.6 backlog. 2 execution phases planned (P1 24 tasks, P2 9 tasks). Grill verdict: Proceed-with-conditions (0.70), 2 MUSTs (G-049, G-067), 1 escalation (PIPEDA consent-law review). ---ci--- project: praxis phase: 0 milestone: v0.5 status: complete requirements: covered: [REQ-ASSIST-01, REQ-ASSIST-02, REQ-ASSIST-03, REQ-NFR-ASSIST-01, REQ-NFR-ASSIST-02, REQ-NFR-ASSIST-03, REQ-NFR-ASSIST-04, REQ-IDEATE-01, REQ-IDEATE-02, REQ-IDEATE-03, REQ-IDEATE-04, REQ-IDEATE-05, REQ-IDEATE-06, REQ-IDEATE-07, REQ-IDEATE-08, REQ-IDEATE-09] partial: [] ---/ci---
58 KiB
Praxis — Voice-first AI Apprenticeship Platform
Milestone: v0.5 (Live Assist — on-the-job voice companion) Status: phase 0 — pre-execution (active milestone) Autonomy: full Previous milestone: v0.4 (Operator tier — cohort dashboard, auth, Postgres) — complete, tagged v0.1.9, release created, merged to main
Vision
Praxis is a voice-first, AI-tutored skill platform for learners in resource-constrained environments. Instead of courses, videos, and quizzes, learners practice real job scenarios through real-time spoken conversation with AI tutors. The platform treats every learner as an apprentice to a master craftsperson — open the app, talk, do the job, get better at it.
One-line pitch: Praxis turns every smartphone into a master craftsperson that talks to you, challenges you, and helps you get good at your job.
Objective
Build a voice-first AI apprenticeship platform where learners engage in spoken role-play scenarios with AI tutors, receive coaching debriefs, and progress via mastery gates — working on low-cost phones over constrained bandwidth.
v0.3 Scope (Mastery Scoring + Competency Rubrics — complete, retained for context)
v0.3 activated the mastery/assessment layer deferred from v0.1/v0.2 (per D-021, ROADMAP line 53). Learners progress via mastery gates — they move on only when they can do the thing across varied scenarios, scored against a competency rubric. v0.3 introduced a verifiable-credential issuer so mastery is portable. The operator tier (multi-tenant + auth + cohort dashboard) was deferred to v0.4 per GRILL-v0.3.md Axis 2.
v0.3 in scope (activated REQ groups — post-grill):
- Mastery core (REQ-MAST-01, REQ-MAST-02): competency rubric per skill; Mastery Score updated after each session, requiring varied-scenario success before a mastery gate opens
- Verifiable credentials (REQ-MAST-03): portable, tamper-evident credentials issued on week-final mastery gate (W3C VC Data Model 2.0, Ed25519, formative-tier, SQLite-backed issuer keys, public verification endpoint)
- Dynamic difficulty (REQ-SCEN-02): scenario difficulty adjusts to learner performance (item-response-theory-informed)
- Scenario library (REQ-SCEN-03, REQ-SCEN-04): library tagged by skill/difficulty/failure_mode; expert-authored format extended with rubric mappings + AI-generated variation hooks
- Path structure (REQ-PATH-02): path-as-job 6-week structure (PRD §6.4) — the progression container mastery gates live in
v0.3 out of scope (deferred to v0.4 per GRILL-v0.3.md Axis 2):
- REQ-DASH-01 (cohort dashboard) + REQ-AUTH-01 (operator auth) + REQ-MT-01/02 (operator Postgres + aggregation) + 4 NFRs — the operator tier was originally v0.8 on the ROADMAP; pulling it into v0.3 created a 2-milestone program. The grill's binding verdict splits it to v0.4. D-031 (override D-007) is deferred with the operator tier.
- REQ-PATH-01 (full multi-path launch) — v0.3 ships the Customer Service path only
- REQ-DASH-02 (full operator-suite dashboard) — later milestone
- REQ-ASSIST-01..03 (Live Assist) — later milestone
- REQ-LOWBW-01..03 (WhatsApp/USSD/offline) — later milestone
- REQ-VOICE-05/06 (multi-language, persona switching) — later milestone
- Active failure injection (D-009) — D-049 confirms stays off in v0.3
- Dynamic rubric weight re-weighting on branch outcome — static in v0.3 (grill Axis 9)
- Traefik proxy / public TLS — deferred from v0.2 (R-AUTH-01 deferred to v0.4 with the operator surface)
Carries forward from v0.2 (already in production):
- Docker-in-LXC deployment (
lxc-deploy.sh,praxis.service,/health:8789) - Voice loop (Deepgram Nova-3 + Cartesia + Pipecat + Ollama Cloud)
- v0.1 scenario (
cs_refund_ca_v01.yaml) + guardrails + debrief
v0.5 Scope (Live Assist — On-the-Job Voice Companion)
v0.5 activates the Live Assist surface deferred from v0.1 (per the original out-of-scope list: "Live Assist mode"). v0.1–v0.4 built and validated the practice surface — learners practice scenarios with AI tutors, scored against rubrics, progress via mastery gates, with a v0.4 operator tier observing cohort patterns. v0.5 adds the companion surface: a hands-free voice assistant a learner invokes while actually working on the job, context-aware of their current scenario/skill path, coaching in real time without doing the job for them.
v0.5 in scope (activated REQ groups — 3 REQs + NFRs TBD after RESEARCH/IDEATE):
- Hands-free voice companion (REQ-ASSIST-01): voice companion invocable while working — distinct from the practice voice loop (v0.1). Hands-free (earbuds/phone-in-pocket), always-listening or wake-word/hotkey-activated, short coaching turns interleaved with real work. Reuses the v0.1 voice pipeline (Pipecat + Deepgram + Cartesia + Ollama Cloud) but in a new "assist" mode, not the practice scenario loop.
- Context-aware (REQ-ASSIST-02): knows the learner's current scenario/skill path — binds to the learner's active path week (D-037) + scenario context, so coaching is relevant to the job they're actually doing, not generic. Carries forward learner state from SQLite (D-007 preserved).
- Guardrails (REQ-ASSIST-03): coaches, does not do the job; never lies to real customers — the safety-critical distinction from the practice surface. The AI is in the learner's ear during real customer interactions; it must never impersonate, never give answers the learner parrots, never claim authority it doesn't have. Extends D-019 guardrail layer with Live-Assist-specific ruleset. Safety-sensitive: real customers, real consequences.
v0.5 out of scope (still deferred):
- REQ-PATH-01 (full multi-path launch) — still Customer Service path only; Live Assist binds to that path
- REQ-LOWBW-01..03 (WhatsApp/USSD/offline) — v0.5 is voice; low-bandwidth surfaces later
- REQ-VOICE-05/06 (multi-language, persona switching) — Canadian English only in v0.5
- REQ-DASH-02 (full operator-suite dashboard) — v0.4's foundational cohort view is sufficient; Live Assist telemetry feeds the same aggregation pipeline
- Learner auth / multi-learner-per-device — still single-learner-per-device (D-007)
- Live Assist session recording/replay — v0.5 is live coaching, not recording; replay later
- Proactive intervention (AI speaks unprompted) — v0.5 is learner-invoked; proactive later
- Multi-modal (camera/screen context) — audio-only (C-4)
Carries forward from v0.4 (already in production):
- Operator-tier Postgres + cohort aggregation + operator auth + cohort dashboard (v0.4)
- Mastery scoring + competency rubrics + IRT + VC issuer (v0.3)
- Scenario library + Customer Service 6-week path (v0.3)
- Docker-in-LXC deployment (v0.2)
- Voice loop: Deepgram Nova-3 + Cartesia + Pipecat + Ollama Cloud (v0.1)
Open questions for CLARIFY/RESEARCH:
- ✅ RESOLVED (D-058, D-064): Invocation model = wake-word (Picovoice Porcupine on-device) + tap-to-talk fallback. Refined: built-in wake word for v0.5 pilot (MAU pricing has no recurring free tier — R-ASSIST-01); custom "Hey Praxis" post-pilot; Vosk fallback. NEW open: client architecture — React-Web (v0.1) can't do background wake-word; React-Native upgrade or defer wake-word to v0.6 (RESEARCH §7 Q1).
- ✅ RESOLVED (D-059): Context-binding = learner declares context at session start (path week + scenario tag); server reads
progress.current_weekfrom SQLite. Auto-detection impossible (C-4). - ✅ RESOLVED (D-060, D-068): "Coaches not does" enforced via 3-layer guardrail: (1) prompt rules (coaching-mode system prompt), (2) output filter (regex direct-answer + false-authority + impersonation patterns + one retry + canned fallback), (3) audit log (turns table guardrail_verdict + cohort guardrail_block_rate). NEW open: privacy/consent for ambient recording (R-ASSIST-08) — legal review of Canada PIPEDA.
- ⚠️ AT RISK (D-061, R-ASSIST-02): <600ms latency budget for assist turns estimated ~655-770ms (all-cloud) / ~655ms (Piper + lean prompt). Mitigations: D-065 (Piper TTS for assist), D-066 (≤150-token prompt). Flag for orchestrator: relax C-8 for assist or push hardening to v0.6. The same pipeline handles both modes (no second Pipecat instance) — confirmed. Wake-word → first-audio is a separate ~850-1150ms budget (warm WebRTC — D-067).
- ✅ RESOLVED (D-058, D-064, D-067): Hands-free UX = Porcupine on-device (offline, ~1MB RAM, <4% core — verified). Battery ~4-9% per 8h shift (estimated — R-ASSIST-14, needs Phase-1 measurement). Warm WebRTC per shift (D-067). Foreground service for background mic (Android 14+ requirement).
- ✅ RESOLVED (D-062, D-069): Session model = shift-bounded ("starting shift" / "ending shift"), with assist turns within. Auto-end after 8h (D-069). Aggregates as
session_type=assistin v0.4 cohort pipeline (no schema change). Does NOT update mastery (D-063).
NEW open questions from research (for orchestrator + PLAN): 7. Client architecture for v0.5 (RESEARCH §7 Q1): React-Web (v0.1, D-015) can't run a background foreground service on Android. Options: (a) upgrade to React Native, (b) separate native Android assist app, (c) defer wake-word to v0.6 and ship v0.5 assist as tap-to-talk only. Recommendation: (c) for v0.5 pilot. Scope decision. 8. Picovoice sales engagement timing (R-ASSIST-01): before PLAN or after v0.5 ships with tap-to-talk? If wake-word deferred to v0.6, sales engagement is v0.6. 9. Output filter regex corpus (R-ASSIST-06): how to build the tuning corpus before v0.5 ships? Synthetic corpus via LLM (prompt gemma4:cloud to produce coaching + direct-answer responses, label, tune). Phase-1 task. 10. Canada consent law review (R-ASSIST-08, D-070): PIPEDA + provincial one-party/two-party consent for ambient recording during coaching. Legal review recommended before v0.5 ship.
v0.4 Scope (Operator Tier — Cohort Dashboard + Auth + Postgres — complete)
v0.4 activates the operator tier deferred from v0.3 per GRILL-v0.3.md Axis 2 (the operator tier was originally v0.8 on this ROADMAP; pulling it into v0.3 created a 2-milestone program disguised as one). The v0.3 mastery/VC/scenario work carries forward unchanged; v0.4 layers the operator surface on top of it.
v0.4 in scope (activated REQ groups — 8 REQs total):
- Operator-tier Postgres (REQ-MT-01): second Docker service in the existing LXC CT (
docker-compose.ymladdspostgres), Postgres 16, persistent volume, internal Docker network only (D-040). Separate from learner-local SQLite (D-007 preserved for learner surface). Stores cohort aggregations, operator accounts, issued credentials, mastery-gate audit log. - Cohort aggregation pipeline (REQ-MT-02): on-session-end hook + nightly reconciliation job writes k-anonymized aggregates to Postgres from learner sessions (D-045). No raw learner PII in Postgres.
- Operator auth (REQ-AUTH-01): session-cookie, argon2id passwords, single
operatorrole, login rate-limited (5 attempts/min) (D-041). Cookie: httpOnly, secure, SameSite=Strict, 8h expiry. Protects cohort dashboard + credential issuance. - Cohort dashboard (REQ-DASH-01): anonymized cohort view (practice, mastery progression, failure patterns) for training operators — k-anonymity ≥ 10, 7-day aggregation window (D-034). React route under
/operator/*, served by the same FastAPI server (new/api/operator/*prefix), reuses v0.2 StaticFiles (D-044). No separate SPA build — sameclient/dist. - NFRs (4): REQ-NFR-AUTH-01 (argon2id + httpOnly + secure + rate-limited), REQ-NFR-MT-01 (Postgres-in-LXC without destabilizing learner service), REQ-NFR-DASH-01 (k-anonymity ≥ 10 enforced — cells < 10 suppressed), REQ-NFR-DASH-02 (freshness ≤ 24h stale).
v0.4 out of scope (still deferred):
- REQ-PATH-01 (full multi-path launch) — v0.3 ships Customer Service path only, multi-path later
- REQ-DASH-02 (full operator-suite dashboard) — later milestone (v0.4 ships the foundational cohort view only)
- REQ-ASSIST-01..03 (Live Assist) — later milestone
- REQ-LOWBW-01..03 (WhatsApp/USSD/offline) — later milestone
- REQ-VOICE-05/06 (multi-language, persona switching) — later milestone
- Learner auth / multi-learner-per-device — operator auth is v0.4; learner auth later
- RBAC (multiple operator roles) — single
operatorrole in v0.4; RBAC deferred - Third-party credential issuers — v0.9 credentialing milestone
- Differential privacy — k-anonymity ≥ 10 is sufficient for v0.4 scale (D-034)
Carries forward from v0.3 (already in production):
- Mastery scoring + competency rubrics + IRT dynamic difficulty (v0.3)
- Verifiable credential issuer (W3C VC 2.0, Ed25519, SQLite-backed) — v0.4 migrates the issuer key store to operator-tier Postgres + secrets (D-042)
- Scenario library + Customer Service 6-week path (v0.3)
- Docker-in-LXC deployment (v0.2)
- Voice loop (Deepgram Nova-3 + Cartesia + Pipecat + Ollama Cloud) (v0.1)
v0.3 Scope (Mastery Scoring + Competency Rubrics — complete)
v0.3 activated the mastery/assessment layer deferred from v0.1/v0.2 (per D-021). Learners progressed via mastery gates — they moved on only when they could do the thing across varied scenarios, scored against a competency rubric. v0.3 shipped competency rubric engine + Mastery Score + scenario library (≥6 CS scenarios) + dynamic difficulty (IRT) + Customer Service 6-week path + verifiable-credential issuer (W3C VC 2.0, Ed25519, SQLite-backed, formative-tier, public verification). All learner-facing. Released as v0.1.5.
v0.2 Scope (Proxmox LXC Deployment — complete)
v0.2 in scope:
- Docker image (multi-stage: Node builds
client/dist, Python runsserver+ serves dist via FastAPI StaticFiles) scripts/proxmox/adapted from coreci (api.sh, lxc-deploy, lxc-clone, lxc-config, lxc-start, health-check, rollback, stage-snippet, firstboot-hook, timing)scripts/install-service.sh(systemd unit fordocker compose up)- Secret wiring: PROXMOX_* sourced from coreci's
.env.secrets; GITEA_TOKEN + DEEPGRAM_API_KEY from praxis's secrets - Health-check adapted for
/health:8789 (praxis's endpoint, not coreci's/healthz:18080) - E2E deploy verification against the live Proxmox cluster
v0.2 out of scope (deferred):
- Mastery scoring, competency rubrics (deferred to v0.3)
- CARTESIA_API_KEY / OLLAMA_API_KEY provisioning (infrastructure-only; server degrades gracefully per v0.1 design)
- Traefik proxy / public TLS (pilot = direct bridge IP access)
- Multi-environment (dev/staging/prod) — single pilot CT
- vmbr1 private network (pilot uses vmbr0 DHCP)
Product Principles (non-negotiable)
- Voice is the primary interface. Text is fallback, not default.
- Doing > Knowing. Every session produces observable action, not passive consumption.
- One skill, one outcome. Each path is a job someone can get.
- Works on a cheap phone, on 2G. Engineering constraints are product features.
- The AI is a master, not a chatbot. Personality, standards, opinions.
- Mastery gates progression. Move on when you can do the thing.
- Failure is the curriculum. AI provokes mistakes, then coaches recovery.
Requirements (summary — see REQUIREMENTS.md for formal REQ-IDs)
- Voice conversation engine: real-time ASR + streaming TTS, <600ms round-trip, interruptible, persona switching
- Scenario engine: branching role-plays with failure-injection and dynamic difficulty (v0.1: one scenario)
- Learner state: progress, session history, mastery accumulation (v0.3: mastery scoring + competency rubrics + verifiable credentials)
- Scenario engine: branching role-plays with failure-injection and dynamic difficulty (v0.3: dynamic difficulty + scenario library + AI variations)
- Skill paths: path-as-job 6-week structure (v0.3: Customer Service path structured + mastery gates)
- Cohort dashboard: anonymized cohort view for training operators (v0.3: multi-tenant + auth + cohort view)
- LLM foundation: Ollama-hosted open-weights models
gemma4:cloudanddeepseek-v4-flash:cloud - Low-bandwidth surfaces (later milestones)
Constraints
- C-1 Voice is primary interface; text is fallback only
- C-2 Must work on $100 Android phone over 2G/3G
- C-3 Cost ≤ $3/active learner/month (target markets; v0.1 is Canada launch — relaxed for pilot)
- C-4 Audio-only in v1 (no large video assets)
- C-5 Open-weights LLM via Ollama catalog —
gemma4:cloud+deepseek-v4-flash:cloud - C-6 Domain safety guardrails + human-in-the-loop + disclaimers for safety-sensitive domains
- C-7 Scenarios authored by domain experts + learning designers; AI generates variations only
- C-8 Latency budget < 600ms end-to-end (ASR → LLM → TTS)
Key Decisions
| ID | Decision | Rationale | Confidence | Alternatives |
|---|---|---|---|---|
| D-001 | Launch market = Canada (path: Customer Service) | User-directed; Canada as initial market for v0.1 pilot. PRD named Kenya — overridden. | 0.70 | Kenya + Customer Service (PRD default) |
| D-002 | Milestone = v0.1 foundation (v1.0 reserved for working/tested product) | User-directed; v0.1 is the foundation slice (Phase 0 + Phase 1 minimal voice loop). v1.0 is a future milestone. | 0.90 | v1.0 = Phase 0 + Phase 1 (too ambitious for first milestone) |
| D-003 | LLM foundation = Ollama catalog — gemma4:cloud + deepseek-v4-flash:cloud |
User-directed; open-weights via Ollama, two base models for edge/cloud split. Research phase to verify exact catalog IDs. | 0.75 | Llama-family, Mistral-family, Qwen-family |
| D-004 | Defer monetization model decision to Phase 1 | PRD §11.5 explicitly lists this as a Phase 1 decision (B2C paid, B2B per-seat, donor-funded, government). | 0.85 | Decide now (insufficient data) |
| D-005 | Single-project mode | Fresh repo with one project; no multi-project need. | 1.00 | Multi-project mode |
| D-006 | "One persona" = one voice persona; scenario role-play uses the same TTS voice as mentor (no distinct character voice in v0.1) | Minimizes v0.1 surface area; PRD's full persona-switching (REQ-VOICE-06) is deferred. Same voice avoids a second TTS configuration to validate. | 0.70 | Two voices (mentor + character) — adds TTS config risk |
| D-007 | "Single learner state" = local single hardcoded profile, no auth, no multi-tenant; persisted via SQLite on-device (or local file fallback) | v0.1 is a pilot harness, not a production multi-user system. Auth/multi-tenant is a later-milestone concern. SQLite chosen as the default local store; research phase may refine. | 0.80 | In-memory only (no persistence), server-side Postgres (premature) |
| D-008 | Interruptibility = abort-and-yield (learner speech cuts AI TTS immediately, AI yields the floor, no pause/resume state machine in v0.1) | Matches real-conversation semantics per PRD §6.1; pause/resume adds state-machine complexity inappropriate for v0.1. | 0.75 | Pause/resume state machine |
| D-009 | Failure-injection hook = architecturally present (scenario declares a failure_mode field) but NOT actively provoked in v0.1 sessions |
v0.1 validates the data model and one scenario's success criteria; provoking failures is a coaching-debrief feature tied to mastery (deferred). Hook present so Phase 2+ can activate it without schema change. | 0.70 | Active failure injection in v0.1 (couples to deferred mastery engine) |
| D-010 | v0.1 Canada Customer Service scenario = "Angry customer requesting refund on a damaged product" (retail context, single branch point) | Concrete, universally recognizable, low safety-risk (non-medical/non-electrical). One branch point (customer escalates vs accepts resolution) keeps scenario runtime minimal while exercising branching. | 0.65 | "Customer with wrong booking" (hospitality — less universal for Canada pilot) |
| D-011 | Coaching debrief = included in v0.1 as a single end-of-session text+voice summary (not the full PRD §5.1 multi-moment replay) | The debrief is part of the core daily loop and cheap to include at a basic level. Full replay/multi-moment coaching is tied to mastery (deferred). | 0.70 | Exclude debrief entirely (loses core loop identity), full replay (over-scoped) |
| D-012 | v0.1 cost ceiling = no enforced ceiling (pilot); architecture must not bake in assumptions that would prevent meeting ≤$3/learner/month post-pilot | C-3 is a target-market constraint. Canada pilot is a foundation/tech-validation milestone, not a unit-economics milestone. Logging actual cost per session is a v0.1 NFR to inform later milestones. | 0.85 | Enforce $3 ceiling in v0.1 (premature optimization, wrong market) |
| D-013 | ASR = Deepgram Nova-3 streaming (cloud, WebSocket) | Research-verified: streaming-native, ~200-300ms first partial, accent-robust for Canadian English, first-class Pipecat integration, Canada data-residency available. Fallback: Groq-hosted Whisper. | 0.85 | whisper.cpp (breaks <600ms budget), OpenAI Whisper API (batch) |
| D-014 | TTS = Cartesia Sonic (cloud, ~120ms first audio) primary; Piper (self-hosted, ~80ms) fallback behind interface | Research-verified: Cartesia #1 on Speech Arena; Piper is open-weights post-pilot ≤$3/learner path. R4 risk: all-cloud path ~670ms — Piper local may be required for production v0.1 latency. | 0.80 | ElevenLabs (quality but higher latency/cost), Amazon Polly |
| D-015 | Client = React + WebRTC via Pipecat client SDK | Research-verified: Pipecat ships React/RN/Swift/Kotlin SDKs; web client = fastest v0.1 iteration, no app-store distribution, upgrades to React Native for Android later. | 0.85 | Python CLI harness (dev-integration only), native Android Kotlin (premature) |
| D-016 | Transport = WebRTC (UDP, sub-50ms audio); WebSocket dev fallback | Research-verified: WebRTC is Pipecat's production transport; adaptive bitrate, UDP. SSE/HTTP rejected (unidirectional/high overhead). | 0.85 | WebSocket-only (higher audio latency), custom raw HTTP/2 |
| D-017 | Orchestration = Pipecat (not custom, not Vocode) | Research-verified: 13.8k★, active, integrates Deepgram+Cartesia+Piper+Ollama natively, has VAD/interrupt/Flows for branching. Vocode stale since Nov 2024. Custom orchestration rebuilds solved problems. | 0.85 | Vocode (stale), custom from scratch |
| D-018 | Scenario format = YAML DSL → Pydantic → Pipecat Flows | Research-verified: YAML is human-authorable + diffable + supports comments (critical for learning-designer rationale per C-7); Pydantic gives typed runtime; Pipecat Flows consumes the schema for branching. JSON is wire format only. | 0.85 | JSON DSL (no comments), code-authored (couples authoring to engineering) |
| D-019 | v0.1 guardrail layer = pluggable interface with Customer Service ruleset implementation | Research: v0.1 is low-risk (Customer Service) but architecture must support pluggable guardrails for later high-risk domains (health/electrical). Ruleset: no legal/financial/medical advice, no real-company employee impersonation, stay-in-role, session-start disclaimer audio, no PII beyond hardcoded profile. | 0.80 | No guardrails (violates C-6), hardcoded non-pluggable rules (blocks future domains) |
| D-020 | LLM access = Ollama Cloud direct API (https://ollama.com/api/chat + OLLAMA_API_KEY) — no local daemon |
Research-verified: :cloud tags are real Ollama hosted-inference on NVIDIA cloud partners. Direct API eliminates local-daemon deployment dependency. gemma4:cloud (256K ctx) → role-play fast path; deepseek-v4-flash:cloud (1M ctx, no-think mode) → debrief. Self-host gemma4:e4b is the post-pilot cost-reduction path. |
0.85 | Local Ollama daemon proxy mode (adds deployment dependency) |
| D-021 | v0.2 scope = Proxmox LXC deployment (replaces roadmap's mastery-scoring v0.2) | User-directed: deploy praxis into an LXC container hosted on Proxmox, reusing ~/coreci/scripts/proxmox/ methods. Mastery scoring deferred to v0.3. |
0.95 | v0.2 = mastery scoring (original roadmap), v0.2 = LXC deploy + mastery (too large) |
| D-022 | Artifact = Docker image in LXC (nesting=1) | User-directed. Isolates Python/Pipecat deps; coreci's clone script already sets features=nesting=1. Avoids venv/pip first-boot fragility (Pipecat has many native deps). Multi-stage build: Node stage produces client/dist, Python stage runs the server. |
0.85 | Clone repo + venv + pip (fragile first-boot), sdist tarball (needs build/release step) |
| D-023 | Client serving = FastAPI serves client/dist as StaticFiles |
User-directed. Single port (8789), simplest pilot — no nginx/caddy. The Docker image bundles the pre-built dist. | 0.90 | Separate static server (nginx/caddy — more moving parts), client out of scope |
| D-024 | Voice-service keys = infrastructure-only for v0.2 | User-directed. Server starts and /health passes even without CARTESIA/OLLAMA keys (v0.1 graceful degradation). Keys provisioned in a later milestone. Only GITEA_TOKEN + DEEPGRAM_API_KEY are in .env.secrets. |
0.90 | Provision all keys in v0.2 (premature — deploy infra first) |
| D-025 | Image distribution = host-build → pct push tarball (research decision, see RESEARCH.md) |
The LXC CT may not route to the internet (coreci pattern: host-fetch → pct push). Build the Docker image on the PVE host (Docker available on Proxmox host) and `docker save | pct exec -- docker load, or pct push` a tarball. Avoids needing a container registry. |
0.75 |
| D-026 | Proxmox secrets sourced from ~/coreci/.ciagent/.env.secrets |
Same Proxmox cluster, same operator. PROXMOX_API_URL/TOKEN/NODE/STORAGE/TEMPLATE_VOLID already provisioned there. Praxis's .env.secrets adds GITEA_TOKEN + DEEPGRAM_API_KEY. The deploy script sources both. |
0.90 | Duplicate proxmox secrets in praxis (drift risk) |
| D-027 | VMID = auto (fresh allocation via pve_nextid) |
CLARIFY auto-decide (full autonomy). Don't reuse coreci's fixed PROXMOX_LXC_VMID — praxis gets its own CT on the same cluster. | 0.95 | Reuse coreci's VMID (collision), hardcode a new fixed VMID (manual allocation) |
| D-028 | Docker installed inside the CT via apt (CT has network via vmbr0 DHCP) | CLARIFY auto-decide. Avoids needing Docker on the PVE host. The debian-12 template + nesting=1 supports Docker-in-LXC. firstboot hook runs pct exec to install docker.io + docker-compose-v2. |
0.90 | Docker on PVE host (extra host dependency), pre-baked template (custom template maintenance) |
| D-029 | Image built inside the CT (clone repo from Gitea, docker build, docker compose up) |
CLARIFY auto-decide. Self-contained — CT fetches its own source + builds. No image transfer needed. Slower first-boot (~3-5 min for build) but simpler and reproducible. | 0.80 | Build on PVE host + pct push tarball (host Docker dependency), pre-built image from registry (external dependency) |
| D-030 | CT network = vmbr0 DHCP only (pilot, no vmbr1, no Traefik proxy) | CLARIFY auto-decide. v0.2 is infrastructure-only pilot. Direct bridge IP access for health-check. Proxy/TLS deferred to a later milestone. | 0.90 | vmbr1 + Traefik proxy (over-scoped for pilot) |
| D-031 | v0.3 introduces multi-tenant + auth — overrides D-007 for the cohort-dashboard surface | REQ-DASH-01 (anonymized cohort view for training operators) requires multi-tenant data. D-007's single-learner/no-auth stance was correct for v0.1/v0.2 pilot but blocks v0.3's cohort dashboard. Resolution: hybrid — learner-local state stays SQLite-on-device (D-007 preserved for learner surface); a new operator-tier Postgres stores cohort aggregations + operator accounts + issued credentials. Learner auth deferred (single-learner-per-device still valid for pilot). Operator auth = session-based, single operator role in v0.3. Research phase to validate Postgres-in-LXC + migration path. | 0.75 | Full Postgres migration (abandons SQLite pilot work), defer DASH-01 again (scope creep), no auth (insecure) |
| D-032 | Mastery gate = N-of-M varied-scenario success + rubric score ≥ threshold | Operationalizes PRD principle 6 ("move on when you can do the thing"). N=3 distinct scenarios, rubric mean ≥ 3.5/5.0 (configurable per path). Research phase to validate rubric model + threshold against competency-based-assessment literature. | 0.70 | Single-scenario pass (gaming risk), pure rubric score (no variety), pure time-on-task (invalid) |
| D-033 | Verifiable credentials = W3C VC Data Model 2.0, platform-issued (operator key), Ed25519 signatures | Research-anticipated: W3C VC 2.0 is the current standard; platform-issued is simplest viable issuer model (no DID method proliferation); Ed25519 is compact + widely supported. Self-issued (learner-side key) rejected — no tamper-evidence authority. Third-party issuer (university/agency) deferred to v0.9 credentialing milestone. Revocation = simple status list (VC Status List v2025). | 0.70 | Self-issued (no authority), third-party issuer (v0.9 scope), JWT-VC (less mature tooling) |
| D-034 | Cohort anonymization = k-anonymity ≥ 10 + aggregation window ≥ 7 days | REQ-DASH-01 operator view must not expose individual learners. k=10 is the conventional minimum for anonymized analytics; 7-day aggregation prevents re-identification via sparse windows. Research phase to validate against differential-privacy literature. Operator sees aggregate progression/failure-patterns only. | 0.70 | No anonymization (privacy violation), differential privacy (over-engineered for v0.3 scale), k=5 (too weak) |
| D-035 | Dynamic difficulty = IRT-informed (1-parameter Rasch), updated per session | REQ-SCEN-02. Item Response Theory (1PL/Rasch) is the simplest well-grounded model: learner ability θ, scenario difficulty b, P(success)=logistic(θ−b). Bayesian update of θ after each session. Avoids 2PL/3PL complexity (discrimination/guessing params — needs more data than v0.3 has). Research phase to validate. | 0.70 | ELO-like (less theoretically grounded), fixed difficulty steps (no adaptation), 2PL/3PL (data-hungry) |
| D-036 | Scenario library structure = YAML directory + index manifest, tagged by skill/difficulty/failure_mode/rubric | Extends D-018's YAML DSL. Library = scenarios/<path>/<scenario>.yaml + scenarios/index.yaml manifest (tagged, versioned). Expert-authored scenarios ship as YAML; AI-generated variations use the same schema with a generated_from backref. Rubric mapping added to scenario schema (each scenario declares which rubric criteria it exercises). |
0.80 | Database-backed library (premature — YAML is diffable + authorable per C-7), JSON (no comments per D-018), inline in code (couples authoring to engineering) |
| D-037 | Path structure = 6-week job-structured path, JSON + YAML, mastery gates between weeks | REQ-PATH-02 (PRD §6.4). Path = paths/<slug>.yaml defining 6 weeks, each week = a set of scenarios + a mastery gate. Gate opens when D-032 mastery condition met. v0.3 ships the Customer Service path fully (6 weeks) with ≥1 scenario per week (library REQ-SCEN-03 fills the rest). |
0.75 | Free-form progression (no structure), 12-week (too long for pilot), week-as-fixed-time (relax to mastery-paced) |
| D-038 | Rubric scoring path = rule-based final score, LLM-assisted criterion extraction only (REQ-NFR-MAST-01) | Final score must be deterministic. LLM (deepseek-v4-flash:cloud no_think) extracts criterion evidence from session turns (which utterance maps to which rubric criterion); a rule function computes the 1-5 score per criterion from the extracted evidence + branch outcome. No LLM in the numeric scoring step. Preserves REQ-NFR-MAST-01 determinism + keeps latency off the voice path. | 0.80 | Pure-LLM scoring (non-deterministic, violates NFR-MAST-01), pure-rule extraction (rigid — can't handle free-form speech) |
| D-039 | Rubric YAML format = rubrics/<skill>.yaml with criteria, 5-level anchors, per-skill weights |
Extends D-018's YAML-everywhere stance. One rubric file per skill (v0.3: rubrics/customer_service.yaml). Each criterion has id, name, 5 anchored levels (1=fail … 5=mastery), weight. Scenario YAML maps to rubric criteria via rubric_criteria field (D-036). |
0.80 | JSON (no comments per D-018), inline in scenario (couples rubric to scenario — rubric is per-skill not per-scenario), DB-backed (premature) |
| D-040 | Operator Postgres deployment = second Docker service in the existing LXC CT (docker-compose.yml adds postgres service) |
REQ-NFR-MT-01. Reuses v0.2's LXC + Docker-in-LXC. No new CT, no host Postgres. Postgres 16, persistent volume, internal Docker network only (not exposed to bridge). Operator auth + cohort API + VC issuer connect to it. | 0.80 | Separate CT (over-provisioned for v0.3 scale), host Postgres (PVE host dependency), SQLite for operator (cohort aggregation needs relational + k-anonymity queries — SQLite workable but Postgres is the safer default) |
| D-041 | Operator auth = session-cookie, argon2id passwords, single operator role, login rate-limited (5 attempts/min) |
REQ-NFR-AUTH-01. Simplest viable auth for v0.3's single operator role. No OAuth/JWT complexity for one role. Cookie: httpOnly, secure, SameSite=Strict, 8h expiry. Rate limit via in-memory counter (single-instance). RBAC deferred (one role). | 0.75 | JWT (over-engineered for server-side session), OAuth (no IdP yet), basic-auth (insecure), no rate-limit (brute-force risk) |
| D-042 | VC issuer key = Ed25519 keypair in operator-tier secrets (PRAXIS_VC_ISSUER_KEY), generated on first issuer init, not committed |
REQ-NFR-VC-01. Key generated at first boot if absent, stored in Postgres issuer_keys table encrypted at rest with a root key from secrets. Verification endpoint serves the public key. Rotation = new key + old key marked superseded (not revoked — old VCs still verify against archived public key). |
0.70 | RSA (larger, slower), KMS-managed (no KMS in LXC), self-signed cert chain (X.509 complexity unjustified for one issuer) |
| D-043 | VC verification endpoint = public, unauthenticated, GET /vc/verify/<credential_id> |
Third parties (employers/agencies) verify credentials without an account. Returns {valid: bool, status: "active"|"revoked", issuer: "praxis-v0.3", mastery: {...}}. No PII in the verification response beyond what the credential itself asserts. |
0.80 | Authenticated verification (friction for employers), no public endpoint (credentials not portable), returns full learner PII (privacy violation) |
| D-044 | Cohort dashboard UI = React route under /operator/*, served by the same FastAPI server (new prefix), reuses v0.2 StaticFiles |
REQ-DASH-01. Frontend-engineer reactivates (PERSONAS.md). Adds /operator React route + /api/operator/* FastAPI endpoints. Auth gate in React + server-side session check. No separate SPA build — same client/dist. |
0.75 | Separate operator SPA (extra build pipeline), server-rendered HTML (abandons React investment), no UI (operator reads JSON — not a product) |
| D-045 | Cohort aggregation trigger = on-session-end hook + nightly reconciliation job | REQ-MT-02. Hook fires after end_session() → writes k-anonymized aggregate to Postgres (incremental). Nightly job (cron in the praxis service) reconciles + recomputes 7-day windows. Hybrid: low-latency updates + correctness guarantee. |
0.70 | Pure real-time (race-prone), pure nightly (stale, violates NFR-DASH-02 if job lags), CDC/streaming (over-engineered) |
| D-046 | IRT θ persistence = in learner-local SQLite (learner_ability table: learner_id, path, theta, updated_at) |
REQ-NFR-IRT-01. θ is per-learner-per-path, computed in-process on session end, no LLM call. Stays in SQLite with the rest of learner state (D-007 preserved). Cohort dashboard sees only k-anonymized aggregates of θ, never raw θ. | 0.80 | Postgres (couples learner state to operator tier — violates D-031 hybrid), in-memory (lost on restart), file-based JSON (no queryability) |
| D-047 | Scenario library minimum for v0.3 = ≥6 expert-authored Customer Service scenarios (one per path week) + AI-generated variations gated by expert review | REQ-SCEN-03/04. 6 scenarios give the mastery gate's N=3 varied-scenario condition room (D-032) without being so few that mastery is gameable. AI variations: LLM generates a variation from an expert scenario's schema with generated_from backref; expert reviews + approves before it enters the library. |
0.70 | 3 scenarios (mastery gate N=3 = exactly the minimum — no room for failure-retry variety), 12 scenarios (over-scoped for one milestone), no AI variations (loses REQ-SCEN-04) |
| D-048 | Mastery gate open action = advance learner to next path week + issue VC if week-final gate | When D-032 condition met for a week's scenarios: learner progress.current_week advances. If the gate is the final week's gate, a VC is issued (REQ-MAST-03) asserting mastery of the path. Mid-path gates: no VC, just advancement. VCs are path-level, not week-level. |
0.75 | VC per week (credential spam — devalues the credential), no advancement (mastery gate is decorative), manual advancement (violates autonomy) |
| D-049 | v0.3 activation of D-009 failure-injection = NO — failure-injection stays architecturally present but not provoked in v0.3 | D-009 hook stays in the schema. v0.3 mastery scoring scores recovery from naturally-occurring failure branches (the escalate branch in cs_refund_ca_v01), not AI-provoked failures. Active failure injection couples to a "failure-recovery coaching" feature that's a later milestone. v0.3 RESEARCH confirms this — no new failure-injection scenarios authored. |
0.80 | Activate failure injection in v0.3 (couples mastery scoring to a new feature — scope creep), remove the hook (breaks forward compat) |
| D-050 | Postgres connection from praxis service = asyncpg pool over Docker internal network, service DNS name postgres |
CLARIFY auto-decide (full autonomy). docker-compose defines a postgres service on an internal bridge network; the praxis service reaches it via postgresql://praxis:${PRAXIS_PG_PASSWORD}@postgres:5432/praxis. asyncpg is the async Pg driver (matches FastAPI async). No external port exposure. Single connection pool (min 1, max 10 — v0.4 scale). |
0.85 | psycopg2 sync (blocks event loop), external port + host access (security surface), pgbouncer (over-provisioned for v0.4 scale) |
| D-051 | VC issuer key migration = fresh keypair on first v0.4 boot; v0.3 SQLite-issued VCs remain verifiable via archived public key | CLARIFY auto-decide. v0.3 stored the Ed25519 issuer key in SQLite (issuer_keys table). v0.4 generates a fresh keypair in Postgres issuer_keys (D-040), marks it active, and archives the v0.3 public key as superseded (not revoked — old VCs still verify against it). The verification endpoint tries the active key first, falls back to superseded keys for older credentials. No re-issuance of v0.3 VCs. |
0.80 | Re-issue all v0.3 VCs (unnecessary churn, learners hold old credentials), revoke v0.3 key (breaks old VCs), keep SQLite key store (defeats D-031 hybrid) |
| D-052 | Operator account bootstrap = first-run CLI script scripts/create-operator.py creates the initial operator from env-provided credentials |
CLARIFY auto-decide. No signup UI (operators are provisioned, not self-serve). Script reads PRAXIS_BOOTSTRAP_OPERATOR_USER + PRAXIS_BOOTSTRAP_OPERATOR_PASS from .env.secrets, hashes the password with argon2id, inserts into operators table. Idempotent (no-op if user exists). Subsequent operators added via the same script (run by the operator from the host). RBAC deferred (D-041 single role). |
0.80 | First-run web wizard (UI surface for a one-time action), hardcoded admin/admin (insecure), SQL insert (no password hashing) |
| D-053 | Cohort dashboard v0.4 scope = 3 views: practice-volume, mastery-progression, failure-patterns — all k-anonymized ≥10, 7-day rolling windows | CLARIFY auto-decide. REQ-DASH-01 names "practice, mastery progression, failure patterns" — v0.4 implements exactly those three views, no more. (1) Practice volume: sessions/day per path, anonymized. (2) Mastery progression: % learners at each week, gate-open rate. (3) Failure patterns: top failure modes by frequency, rubric criterion weak-spots. Each view = a /api/operator/<view> endpoint returning pre-aggregated rows from cohort_aggregates; React renders read-only tables + sparkline charts. No filters beyond path + window (no per-learner drill-down — k-anon). |
0.80 | Full BI dashboard (over-scoped for v0.4), single combined view (loses the three named aspects), per-learner drill-down (violates k-anon) |
| D-054 | Aggregation trigger = async fire-and-forget on session end (non-blocking); nightly reconciliation job at 03:00 CT | CLARIFY auto-decide. D-045 named the trigger; this clarifies the semantics. On end_session(), the server enqueues an aggregation task to an in-process asyncio.Task (no Celery/Redis for v0.4 scale) — non-blocking, the session-end response returns immediately. Failures log + the nightly job reconciles (idempotent upsert by window). Nightly job: cron-style asyncio.create_task loop, recomputes all 7-day windows. If the service restarts, the in-flight task is lost but nightly reconciliation covers it. |
0.80 | Sync on session-end (adds latency to learner path — violates C-8), Celery+Redis (over-provisioned), CDC streaming (over-engineered) |
| D-055 | Postgres backup = nightly pg_dump to a named Docker volume, 7-day retention |
CLARIFY auto-decide. Postgres data lives on a named Docker volume (pgdata) inside the LXC CT. Nightly cron job runs `pg_dump praxis |
gzip > /backups/praxis-$(date).sql.gz to a second named volume (pgbackups). 7-day retention (rotates oldest). Operator can pct pull` backups to the PVE host. No streaming replication (single CT, no replica target). This is pilot-tier backup; a later milestone adds off-CT replication. |
0.70 |
| D-056 | Auth session store = signed stateless cookies (HMAC-SHA256), no server-side session table | CLARIFY auto-decide. D-041 said "session-cookie" — clarifying: the cookie is a self-contained signed token (user_id, issued_at, expiry, HMAC). No sessions table in Postgres. Verification = recompute HMAC + check expiry. Logout = client clears cookie (stateless — no server revocation list in v0.4). Rate limit is in-memory (single-instance). This minimizes DB load + simplifies the auth surface. A later milestone adds a revocation list if multi-instance or forced-logout is needed. |
0.75 | Postgres sessions table (DB load + cleanup job), Redis sessions (extra service), JWT with claims (same idea, more complex tooling) |
| D-057 | Auth enforcement = server-side on every /api/operator/* request + React route guard for UX, never trust the client |
CLARIFY auto-decide. FastAPI middleware checks the signed cookie on every /api/operator/* request; 401 if missing/invalid/expired. React /operator/* routes check a /api/operator/me call on mount and redirect to /operator/login if 401 — this is UX only, the server is the authority. The cohort dashboard reads only k-anonymized aggregates (D-034) so even an auth bypass leaks no PII (defense in depth). VC issuance endpoints (/api/operator/credentials/*) are also auth-gated. |
0.85 | Server-only (poor UX — no redirect), React-only (insecure — bypassable), no auth on issuance (credential forgery risk) |
| D-058 | Live Assist invocation model = wake-word (Picovoice Porcupine on-device) + tap-to-talk fallback, NOT always-listening | CLARIFY auto-decide (full autonomy). Always-listening drains battery on a $100 Android phone the learner is actively using for work + raises privacy concerns (listening to real customers). Wake-word is the hands-free UX without always-on microphone. Picovoice Porcupine is on-device, offline, low-power, free-tier supports custom wake words. Tap-to-talk fallback covers wake-word failure or noisy environments. Research phase to validate Porcupine on Android + battery impact. | 0.65 | Always-listening (battery + privacy), pure tap-to-talk (not hands-free), cloud wake-word (latency + connectivity dependency) |
| D-059 | Live Assist context-binding source = learner declares context at session start (path + scenario tag), server reads active path week from SQLite for rubric/coaching alignment | CLARIFY auto-decide (full autonomy). Live Assist cannot auto-detect which real scenario the learner is in (no camera per C-4, no screen context). Learner taps their current path week / scenario tag when starting an assist session (or voice-declares it). Server reads the learner's progress.current_week from SQLite (D-007) for rubric alignment + coaching context. This keeps the learner in control + makes context explicit. Auto-detection from calendar/location is out of scope. |
0.70 | Full auto-detection (impossible without sensors), pure SQLite read without learner declaration (ambiguous which real scenario), no context (generic coaching — violates REQ-ASSIST-02) |
| D-060 | Live Assist "coaches not does" guardrail enforcement = (1) prompt-layer rules (system prompt forbids giving direct answers), (2) output filter (post-generation check for direct-answer patterns), (3) session audit log of all assist turns | CLARIFY auto-decide (full autonomy). REQ-ASSIST-03 is safety-critical. Three layers: (1) system prompt explicitly instructs the LLM to ask guiding questions, never give the answer, never speak on behalf of the learner. (2) Output filter scans the LLM response for direct-answer patterns (e.g., "you should say X to the customer") and rewrites/blocks. (3) All assist turns logged to SQLite for audit + the operator cohort dashboard (v0.4). Research phase to validate filter patterns + false-positive rate. | 0.70 | Prompt-only (single layer — bypassable), output-filter-only (inconsistent with prompt), no logging (no audit trail — unsafe for safety-critical surface) |
| D-061 | Live Assist latency budget = shares the v0.1 voice pipeline (Pipecat + Deepgram + Cartesia + Ollama) but assist turns are short (≤30s), and the <600ms round-trip (C-8) must hold for assist turns | CLARIFY auto-decide (full autonomy). Live Assist does NOT run concurrently with a practice session — it's a separate mode. The learner invokes assist, gets short coaching turns (≤30s each), dismisses. The same pipeline handles both modes (no second Pipecat instance). C-8's <600ms budget applies to assist turns too — coaching that arrives after the customer moment has passed is useless. Research phase to validate wake-word → first-audio latency + whether assist context adds LLM tokens that break the budget. | 0.75 | Separate pipeline (doubles infra cost + complexity), relaxed latency for assist (useless coaching), longer turns (loses the real-time moment) |
| D-062 | Live Assist session model = shift-bounded sessions (learner starts "I'm starting my shift", ends "ending shift"), with individual coaching turns within the shift; assist turns feed the v0.4 cohort aggregation as a new session_type=assist |
CLARIFY auto-decide (full autonomy). A shift-bounded session matches the real-world use case (a learner works a shift, invokes assist as needed). Within the shift, each assist turn is a discrete coaching exchange. Assist turns aggregate into the v0.4 cohort pipeline (D-045) as session_type=assist — operators see assist usage patterns alongside practice patterns. No double-counting with mastery: assist turns are coaching, not assessment, so they don't update θ (D-035) or count toward mastery gates (D-032). Continuous (no start/end) is ambiguous for aggregation. |
0.70 | Continuous (no aggregation boundary), per-turn sessions (too granular for cohort view), no aggregation (operators blind to assist usage) |
| D-063 | Live Assist does NOT update mastery score (D-035) or count toward mastery gates (D-032) — assist is coaching, not assessment | CLARIFY auto-decide (full autonomy). Mastery gates require demonstrated performance across varied scenarios (D-032). Live Assist is the AI helping during real work — it's coaching, not a performance demonstration. Counting assist turns toward mastery would be gaming (the AI did the work). Assist turns are logged for audit + cohort aggregation (D-062) but never update θ or open gates. A later milestone may add "assist-weaning" (track reducing assist reliance as a mastery signal) but v0.5 keeps them separate. | 0.85 | Assist counts toward mastery (gaming risk), assist updates θ (contaminates the ability estimate), no logging (no audit) |
| D-064 | Live Assist wake-word engine = Picovoice Porcupine (built-in wake word for v0.5 pilot; custom "Hey Praxis" post-pilot), with Vosk as the documented open-source fallback | RESEARCH-derived (RESEARCH-v0.5 §1.2). R-ASSIST-01: Porcupine MAU pricing has no recurring free tier (verified via Picovoice general FAQ). v0.5 ships with a built-in Porcupine wake word (e.g., "Bumblebee") to avoid custom-training costs during the pilot. Post-pilot, engage Picovoice sales for a custom "Hey Praxis" under a pilot/educational tier. Vosk (Apache 2.0, offline) is the fallback if Porcupine pricing is unsustainable. Snowboy rejected (deprecated). | 0.70 | Vosk for v0.5 (free but heavier), TFLite DIY (engineering effort), Snowboy (deprecated) |
| D-065 | Live Assist TTS = Piper (self-hosted on pilot server) as the default for assist turns, Cartesia as the quality fallback for practice mode | RESEARCH-derived (RESEARCH-v0.5 §3.3). R-ASSIST-02: assist turns are latency-critical (C-8). Piper ~80ms first audio vs Cartesia ~120ms. The v0.1 R4 mitigation pre-stages Piper; v0.5 assist mode defaults to Piper to claw back ~40ms toward the <600ms budget. Practice mode retains Cartesia (quality over latency for practice). | 0.75 | Cartesia for both (simpler, but +40ms on assist), Piper for both (lower quality for practice) |
| D-066 | Live Assist system prompt = ≤150 input tokens (coaching instruction ~80 tokens + context-binding ~50 tokens + voice-conciseness ~20 tokens) | RESEARCH-derived (RESEARCH-v0.5 §3.3). R-ASSIST-02: extra input tokens add prefill latency (~0.5ms/token). A lean prompt keeps the prefill delta under 50ms vs v0.1 practice. Avoid dumping the full rubric or scenario YAML into the prompt — context-binding is terse (path week, scenario tag, one-line coaching focus). | 0.78 | Verbose prompt (easier coaching quality, but +100-200ms latency) |
| D-067 | Live Assist WebRTC connection = warm for the entire shift (foreground service keepalive; not per-turn cold connect) | RESEARCH-derived (RESEARCH-v0.5 §3.4). R-ASSIST-03: cold WebRTC connect (~500-1000ms) is unacceptable for live assist. The assist foreground service opens a warm connection at shift start, keeps it alive (heartbeat every 30s), and reuses it for every assist turn. Closed at shift-end. Between turns, only keepalive flows (no audio streaming) to save battery. | 0.78 | Per-turn cold connect (too slow), always-streaming (battery + privacy) |
| D-068 | Live Assist guardrail output filter = regex-based direct-answer + false-authority + impersonation patterns, with one retry on block + canned coaching redirect fallback | RESEARCH-derived (RESEARCH-v0.5 §2.3). R-ASSIST-06/07: regex is the fast on-voice-path filter (matches the existing CustomerServiceGuardrail pattern). One retry gives the LLM a chance to self-correct; the canned fallback ensures a safe response if the retry also blocks. LLM-as-judge deferred to post-v0.5 (off-voice-path, more accurate, nightly). | 0.78 | LLM-as-judge on-voice-path (too slow for <600ms), no filter (unsafe) |
| D-069 | Live Assist shift = auto-end after 8 hours (configurable via PRAXIS_ASSIST_MAX_SHIFT_HOURS=8) |
RESEARCH-derived (RESEARCH-v0.5 §4.2). R-ASSIST-11: learners may forget "ending shift", leaving orphaned WebRTC connections + stale sessions. Auto-end after 8h (a typical shift length) closes the shift cleanly, fires the aggregation hook, and releases the foreground service. The learner can restart a new shift if needed. | 0.75 | No auto-end (orphan risk), shorter (4h — too short for some shifts), longer (12h — battery risk) |
| D-070 | Live Assist consent disclosure = foreground-service notification + learner-facing "Assist is on — those around you may be recorded by your mic" disclosure at shift start | RESEARCH-derived (RESEARCH-v0.5 §2.6). R-ASSIST-08: the ambient mic may pick up the real customer. Ethical and legal (one-party/two-party consent law) requires disclosure. The foreground service notification (Android requirement) + an in-app disclosure at shift start covers the learner's awareness. The customer's consent is the learner's responsibility (Praxis can't notify the customer). Flag for orchestrator: legal review of Canada consent law (PIPEDA) for ambient recording during coaching. | 0.65 | No disclosure (legal/ethical risk), explicit customer consent prompt (impractical — the customer isn't a Praxis user) |
| D-071 | Live Assist client architecture for v0.5 = tap-to-talk ONLY (no wake-word in v0.5) — React-Web (D-015) keeps the assist surface as a tap-to-talk web control; wake-word deferred to v0.6 with a React-Native or native Android app | RESEARCH-flagged decision (full autonomy). R-ASSIST-13: React-Web (v0.1, D-015) cannot run an Android background foreground service for on-device wake-word detection. Adding wake-word requires a React-Native upgrade or a separate native Android assist app — a client-architecture change too large for v0.5's scope. v0.5 ships assist as tap-to-talk (the existing fallback from D-058): learner taps a button to invoke an assist turn during a real shift. This preserves the "hands-free goal" as the v0.6 target while delivering the coaching/guardrail/context-binding value in v0.5 on the existing web client. D-058's wake-word is deferred, not abandoned. | 0.70 | Force React-Native in v0.5 (scope creep — client rewrite + assist feature together), defer all of v0.5 assist to v0.6 (no value delivered), ship wake-word on web (technically infeasible) |
| D-072 | Live Assist C-8 latency budget for v0.5 pilot = target <600ms (C-8) retained; accept ≤650ms as pilot tolerance with hardening in v0.6 — Piper TTS (D-065) + ≤150-token prompt (D-066) are the mitigations; if measurement shows >650ms, document as R-ASSIST-02 carried to v0.6 | RESEARCH-flagged decision (full autonomy). R-ASSIST-02: research estimates ~655-770ms all-cloud, ~655ms with Piper + lean prompt. C-8 is a binding constraint but v0.5 is a pilot — a 50ms tolerance (≤650ms) is acceptable if trending down, with <600ms as the v0.6 hardening target. The alternative (relax C-8 formally) weakens the constraint for all future milestones; the alternative (block v0.5 ship until <600ms) delays the safety-critical guardrail work. Accept pilot tolerance, measure in Phase 1, harden in v0.6. | 0.65 | Relax C-8 to 700ms (weakens constraint permanently), block v0.5 until <600ms (delays guardrail work), ignore the gap (unsafe) |
| D-073 | Live Assist PIPEDA consent-law review = defer to v0.5 Phase 1 implementation; document as R-ASSIST-08 in the grill — the ambient-mic legal question is a grill-axis candidate, not a Phase 0 blocker | RESEARCH-flagged decision (full autonomy). R-ASSIST-08: Canada PIPEDA + provincial consent law for ambient recording during coaching needs legal review. This is not a Phase 0 research blocker — the disclosure (D-070) is the engineering mitigation. Legal review runs in parallel with Phase 1 implementation. The grill (next stage) should include an axis on consent/privacy. If the grill returns a MUST for legal review before ship, schedule it before Phase 1 SHIP. | 0.60 | Block Phase 0 on legal review (over-cautious — no implementation yet), ignore the legal risk (unsafe), no disclosure (D-070 already addresses) |
Confidence updates from research
| ID | Before | After | Reason |
|---|---|---|---|
| D-003 | 0.75 | 0.95 | Both Ollama model IDs verified in catalog as real, current, cloud-hosted tags |
| D-007 | 0.80 | 0.90 | SQLite confirmed appropriate for v0.1 single-learner scale; no evidence favors alternatives |
| D-058 | 0.65 | 0.70 (REFINED) | Porcupine verified (on-device, offline, low-power, Android SDK, custom WW). MAU pricing / no recurring free tier contradicts the free-tier assumption — refined by D-064 (built-in WW for pilot, custom post-pilot, Vosk fallback). |
| D-059 | 0.70 | 0.82 | PraxisStore.get_progress() confirmed returns current_week; auto-detection impossible (C-4); learner declaration is the right model. |
| D-060 | 0.70 | 0.85 | 3-layer pattern confirmed as industry-standard; existing CustomerServiceGuardrail proves the regex output-filter approach. Refined by D-068 (regex + retry + canned fallback). |
| D-061 | 0.75 | 0.70 (AT RISK) | Estimated assist latency ~655-770ms (all-cloud) / ~655ms (Piper + lean prompt) — C-8 <600ms is at risk. Mitigations identified (D-065 Piper, D-066 lean prompt) but may not fully close the gap. Flag for orchestrator. |
| D-062 | 0.70 | 0.85 | Shift-bounded model confirmed as matching real CS work; no schema change to cohort_aggregates (new metric strings); on-session-end hook extended cleanly. |
| D-063 | 0.85 | 0.90 | SessionRecorder.end(schedule_mastery=False) for assist shifts confirmed — the mastery flow is practice-only by the existing flag. |
Target Users (v0.3: Canada pilot — Customer Service path)
| Persona | Description | Pain |
|---|---|---|
| Aspiring Adebayo → "Aspiring Alex" | 19–28, Canada. Recent secondary school grad. Smartphone, limited data. Wants a service job. | Can't afford vocational school. Needs to actually do the job. |
| Upskilling Ursula → "Upskilling Uma" | 25–40, Canada. Retail, hospitality, healthcare. Wants promotion/new role. | No time for courses. Learns on the job. |
| Frontline Felix | Customer service / sales / field tech agent, hired recently. | Manager has no time to coach. Wants quick on-shift practice. |
Success Metrics (Year-1 targets, post-v0.1)
| Metric | Target | Why |
|---|---|---|
| Active weekly learners | 100k | Engagement, not downloads |
| Sessions per learner / week | ≥5 | Habit formation |
| Mastery rate per path | ≥40% completion | Real learning |
| Median session length | 6–10 min | On-the-go use |
| Cost / active learner / month | ≤$3 | Sustainable |
| Reported job/promotion outcome | ≥25% | North star |
| NPS (learner) | ≥50 | Word-of-mouth growth |
Open Questions (for research/clarify phases)
- Will learners talk to their phone in public? (earbuds + "no one will know" framing)
- How to certify mastery credibly? (employer/agency recognition)
- Domain safety minimum HITL for health/electrical scenarios
- Voice cloning / impersonation disclosure
- Monetization model (deferred to Phase 1)
- Skills that should remain out of scope
References
- PRD v0.1 (this document's source)
- ARCHITECTURE.md — system architecture
- ROADMAP.md — phase breakdown
- REQUIREMENTS.md — formal requirements with REQ-IDs