Files
praxis/.ciagent/GRILL.md
T
Praxis CI 48cbd4a2b3 docs(P00): complete pre-execution phase
---ci---
phase: 0
milestone: v0.1
status: complete
requirements:
  covered: [REQ-VOICE-01, REQ-VOICE-02, REQ-VOICE-03, REQ-VOICE-04, REQ-SCEN-01, REQ-STATE-01, REQ-LLM-01, REQ-LLM-02, REQ-DEBRIEF-01, REQ-ORCH-01, REQ-ORCH-02, REQ-SCEN-FMT-01, REQ-NFR-LAT-01, REQ-NFR-SAFE-01, REQ-NFR-COST-01]
  partial: []
---/ci---
2026-08-01 12:49:34 +00:00

29 KiB

Praxis — v0.1 Foundation Grill (Red-Team Review)

Grill date: 2026-08-01 Griller: CIAgent (adversarial executive review) Mode: mechanical (autonomy full, no user interaction) Branch: phase/00-pre-execution Artifacts reviewed: PROJECT.md (D-001..D-020), ROADMAP.md, REQUIREMENTS.md, ARCHITECTURE.md, PERSONAS.md, RESEARCH.md (R1-R10), PLAN.md (D-P1-01..06), config.json, CHECKPOINT.json, git log (5 commits) Codebase state: planning artifacts only — no src/, server/, or client/ exists yet (expected at Phase 1 EXECUTE)


Verdict

Verdict PROCEED
Confidence 0.72
Binding decisions 8 (G-001..G-008)
Escalations 0 (all axes resolved with confidence ≥ 0.60)
Challenges posed 28 forcing questions across 10 axes; 9 produced material findings

One-line summary: v0.1 is a genuinely well-prepared foundation milestone with research-grounded, swappable architecture and front-loaded risk spikes. It is not, however, what its "pilot" framing implies: it is a tech-validation harness with no real learners, no timeline, no budget, no named sponsor, and all three thesis-defining constraints (2G, $100 phone, $3/learner) explicitly relaxed. The binding decisions below correct the framing and require two concrete refinements before EXECUTE (no-go action definition, recruitment-plan deferral). None block execution.


Per-Axis Findings

Axis 1 — The Business Case Itself — confidence 0.72

Forcing question Evidence Finding
What problem does this solve, and is it still top priority? PROJECT.md L9-15 (voice-first apprenticeship for resource-constrained environments); D-001 overrides PRD's Kenya → Canada The PRD thesis is "apprenticeship for resource-constrained environments on $100 Android over 2G." v0.1 relaxes both defining constraints (C-2 relaxed per REQUIREMENTS L116, C-3 relaxed per D-012). v0.1 validates the easy version of the problem on Canadian cloud infrastructure. The hard version (the actual moat per RESEARCH L279) remains unproven.
Is the Canada pilot a business case or tech validation? D-012 (PROJECT L82): "Canada pilot is a foundation/tech-validation milestone, not a unit-economics milestone" It is tech validation. D-012 admits it. This is honest but the surrounding "pilot" language (ROADMAP L33, PLAN L15) oversells it. No market entry is occurring.
What happens if R4 latency fails 600ms? ARCHITECTURE L76-78 (Piper mitigation ~550ms); PLAN SLICE-01 "go/no-go gate" A mitigation path exists (Piper, then self-hosted gemma4:e4b). But the go/no-go gate defines no explicit no-go actions — see G-003. If both mitigations fail, the project has no documented kill/scope-reduce trigger.
ROI against counterfactual? No ROI document exists; success metrics (PROJECT L107-117) are "Year-1 targets, post-v0.1" No counterfactual. Acceptable for a foundation milestone; would be a blocker for a funded market-entry pilot.

Axis verdict: Sound for a foundation milestone. The business case is tech-validation, honestly admitted in D-012. The risk is that v0.1's success could be misread as thesis validation when it validates only the voice loop. → G-001.


Axis 2 — Scope and Requirements — confidence 0.78

Forcing question Evidence Finding
Scope expanding, contracting, or stable? D-002 (v0.1 frozen), D-006..D-012 (7 ambiguities resolved), REQUIREMENTS L124-135 (explicit out-of-scope) Stable and frozen. Out-of-scope is comprehensive (14 items). This is well-handled — rare for a project in flux.
Is v0.1 SO thin it doesn't validate the thesis? PLAN L15 (one scenario, one branch, one voice, debrief, no mastery) v0.1 validates the daily loop (speak → AI responds → debrief). It does NOT validate apprenticeship (no mastery gates, no progression, no multi-scenario). Acceptable: the daily loop is the load-bearing wall; mastery is a later floor.
Does "one branch point" actually prove branching works? D-010 (PROJECT L80): one branch (escalate vs accept); PLAN TASK-03-06: branch classifier runs at session end via LLM-as-judge, offline from voice loop; D-P1-05 confirms "offline at session end" No. The "branch" is a post-hoc outcome label, not a runtime conversation fork. The conversation is linear; the branch is classified after the fact. Pipecat Flows is wired (TASK-03-03) but the branch does not change the conversation in-flight. The claim "exercises branching" (D-010 rationale) is overstated.
Hidden requirements disclosed late? None found — guardrails (D-019), data residency (R10), PIPEDA all surfaced in research Clean. No hidden regulatory/security requirements lurking.

Axis verdict: Scope is honest and frozen. The one overstatement is the branching claim. → G-002.


Axis 3 — Architecture and Technical Feasibility — confidence 0.75

Forcing question Evidence Finding
Has the architecture been validated by builders, not just sellers? RESEARCH L314 ("Measure, don't assume"); R1-R4 all "measure in Phase 1"; PLAN SLICE-01 is the measurement Architecture is research-grounded (web-verified, not vendor-pitched) but not yet builder-validated. SLICE-01 is the validation. Correct sequencing.
Integration surface — where does cost double? Three cloud hops (Deepgram + Ollama Cloud + Cartesia) + WebRTC + Pipecat + React SDK + SQLite Six integration points. Each is a place where latency or cost can surprise. The plan puts all swappable services behind interfaces from SLICE-02 (TASK-02-01..02-03) — correct risk management.
Is the ~670ms budget real? ARCHITECTURE L75 (all-cloud ~670ms, over 600ms); L76 (Piper ~550ms, 50ms margin); all numbers vendor-claimed, unmeasured The all-cloud path fails the target by 70ms on paper. The Piper mitigation has 50ms margin — and that's vendor-claimed, not measured. This is genuinely tight. SLICE-01 measures it. The risk is real but correctly front-loaded.
Is Piper fallback a real mitigation or hand-wave? D-014 (TTS behind interface, Piper pre-staged); R8 (Piper maintainer gap — OHF seeking maintainers); RESEARCH L135 Real but thin. Piper is a first-class Pipecat TTS service and is fast on CPU. But: (a) 50ms margin is slim, (b) R8 flags a maintainer sustainability risk, (c) Piper prosody is "good but not Cartesia-tier" — quality regression. It's a legitimate mitigation, not a hand-wave, but it trades quality for latency and has a dependency-health caveat.
Is Pipecat a safe foundation? D-017 (13.8k★, 11k+ commits, active); R6 (Ollama direct-API integration depth unverified) Yes for v0.1. Active, well-adopted, native integrations for all three services. R6 (unverified Ollama direct-API integration) is a real risk mitigated by SLICE-02 TASK-02-03 (thin adapter if Pipecat's Ollama service rejects custom host+bearer). Long-term: if Pipecat stagnates, Praxis can fork — but that's a future-milestone concern.
Is Ollama Cloud direct API a SPOF? D-020 (single vendor, US-hosted); R10 (PIPEDA data residency); R5 (tier throttling) Yes. Single vendor, single region (US), tier-based throttling. Mitigations: swappable LLM interface (D-020), self-host gemma4:e4b post-pilot path. For v0.1 single-learner, acceptable. R10 (PIPEDA) is low-medium and unresolved — flagged for monitoring, not blocking.

Axis verdict: Architecture is the strongest part of this project. Research-grounded, swappable, risk-front-loaded. The 670ms budget is the tightest constraint and has no margin, but SLICE-01 addresses it correctly. The one gap: the go/no-go gate has no defined no-go actions. → G-003.


Axis 4 — People, Skills, and Organization — confidence 0.70

Forcing question Evidence Finding
Key-person dependency? PERSONAS.md: 4 active personas; backend-engineer owns majority surface (Pipecat + all service integrations per PERSONAS L145) backend-engineer is the critical persona. It owns Pipecat server, Ollama/Deepgram/Cartesia/Piper adapters, guardrails, scenario runtime. If backend-engineer capacity is constrained, the critical path stalls. This is a concentration risk.
Resources allocated at claimed percentages? config.json: max_concurrent_agents 5; PLAN D-P1-04 (SLICE-03
Product owner with authority? D-001 ("user-directed" Canada override); no named PO The human "user" makes high-level decisions; the CI orchestrator handles execution prioritization. No named PO for day-to-day. Acceptable for an autonomous CI project but means prioritization is algorithmic, not market-informed.
Building capability they don't have? R6 (Pipecat + Ollama direct-API integration unverified); personas have no prior Pipecat track record Yes — first Pipecat integration. Mitigated by SLICE-02 verification task. Acceptable for a foundation milestone (learning-as-you-go is fine for prototypes/tech-validation; the plan treats it as such with early spikes).

Axis verdict: Thin but appropriate for an autonomous agent project. backend-engineer concentration is the structural risk. No binding decision — noted as a monitoring item.


Axis 5 — Timeline and Estimates — confidence 0.65

Forcing question Evidence Finding
Was the deadline set before or after scope? No deadline exists anywhere. ROADMAP.md: phases with no dates. PLAN.md: 5 slices, 3 waves, no duration estimates. There is no timeline. This is itself a grill finding.
Is missing timeline a blocker? CHECKPOINT.json (stage: plan); autonomy: full (no external deadline) For Phase 0 pre-execution in an autonomous CI project with no external deadline, the absence of a calendar timeline is defensible — you plan first, estimate later. But Phase 1 EXECUTE has no per-slice effort estimate either, which means no burn-rate tracking is possible.
Critical path + 3-month push risk? PLAN §3: SLICE-01 → SLICE-02 → (SLICE-03 ‖ SLICE-04) → SLICE-05 The single thing that would push by 3+ months: R4 latency failing even with Piper, forcing a self-hosted-LLM/edge architecture rethink. SLICE-01 is the de facto time-box on this risk.
Definition of done? PLAN §4: 10 explicit Phase 1 exit criteria Well-handled. 10 concrete, testable exit criteria. This compensates partially for the missing timeline — "done" is unambiguous even if "when" is not.
Estimates evidence-based? None exist No estimates at all. The wave structure is a sequencing estimate but not a duration estimate.

Axis verdict: Missing timeline is a finding but not a blocker for pre-execution. The 10 exit criteria provide a strong definition of done. → G-004 (add per-slice estimates at EXECUTE).


Axis 6 — Budget and Financial Realism — confidence 0.70

Forcing question Evidence Finding
Budget spent vs. remaining? No budget exists. No dollar amount, no token budget, no compute allocation defined anywhere. There is no budget to track. For an autonomous CI pilot, the "budget" is tokens/compute — and no token budget is defined.
Predictable cost drivers? RESEARCH L60 (Ollama tier pricing, not unit-economics-friendly at scale); Deepgram $0.0043/min; Cartesia per-char; WebRTC TURN/STUN if behind NAT Cost drivers are identified in research but no aggregate estimate exists. The Ollama tier model (Pro $20/Max $100) means pilot cost is plan-tier-based, not per-session — so logged per-session cost (TASK-04-04) will not map to at-scale unit economics.
Is v0.1 measuring things that inform $3/learner? SLICE-04 TASK-04-04 (per-session cost logging: tokens, minutes, chars, derived cents) Yes — the measurement infrastructure is correct. It logs the right inputs. But the outputs won't be representative: Canada + cloud + Ollama-tier pricing is the most expensive configuration, not the $3/learner target configuration (which requires self-hosted gemma4:e4b + Piper).
Burn rate / runway? No budget → no burn rate → no runway calculation Ungoverned. Acceptable for a pilot; would be a blocker for a funded delivery.
Budget contingent on something? D-004 (monetization deferred to Phase 1); D-012 (no enforced ceiling) No contingencies — because there's no budget to be contingent.

Axis verdict: Budget is hand-waved but honestly so (D-012 admits it's not a unit-economics milestone). The cost-logging infrastructure is the right v0.1 contribution. The gap: v0.1 logged costs will mislead if read as representative of at-scale economics. → G-005.


Axis 7 — Risks, Assumptions, and Dependencies — confidence 0.72

Forcing question Evidence Finding
Is R4 actually the biggest risk? RESEARCH L358 (R4: all-cloud ~670ms); PLAN SLICE-01 go/no-go R4 is the biggest technical risk and is well-handled. But it has a mitigation path (Piper, self-host). The risks below are less mitigated.
Accent robustness on real Canadian speech? D-013 (Deepgram "accent-robust" — vendor claim); R9 (French-Canadian code-switching, logged as low-risk) Unmeasurable in v0.1 — there are no real learners (D-007: hardcoded profile). Deepgram's accent robustness is vendor-claimed, not tested on real Canadian speech. This is arguably a bigger risk than R4 because it has no quick fix (retrain or switch ASR) and can't be validated until real learners exist.
Is the branch point too trivial? D-010 (one binary branch); TASK-03-06 (post-hoc LLM-as-judge) The branch is post-hoc, not runtime (see Axis 2). It proves the data model (branch field exists) but not the branching runtime (conversation forks in-flight).
LLM hallucinating outside Customer Service role? D-019 (guardrail ruleset); TASK-03-04 (unit test: "sue them" blocked) Guardrails are system-prompt + output filter. TASK-03-04 tests one case ("sue them"). No adversarial/jailbreak test of the guardrail. For Customer Service (low-risk domain), this is acceptable — but the guardrail layer's pluggability for high-risk domains (health/electrical) is untested under adversarial pressure.
Top 3 assumptions? (1) Pipecat integrates with Ollama direct API (R6); (2) Deepgram Nova-3 handles Canadian English (R1/R9); (3) Cartesia/Piper hits latency targets (R2/R4) All three are "measure in Phase 1" — correctly front-loaded. The fourth unstated assumption: that real learners will use this. No evidence.
Single killing risk? No recruitment plan; PERSONAS.md is personas, not recruitment No real learners. The entire "pilot" depends on ~50 real Canadian learners (implied by success metrics context) and there is no recruitment plan, no recruitment channel, no recruitment budget. v0.1 will produce a dev-harness demo, not a pilot. This is the biggest unflagged risk.
Pre-mortem (12 months, failed — why?) Inferred Most likely causes: (a) R4 can't hit 600ms even with Piper → architecture rethink; (b) voice loop works but debrief is generic → doesn't validate apprenticeship; (c) no real learners ever use it — dev demo that never reaches a population. (c) is the most likely.

Axis verdict: R4 is well-handled. The bigger risks are (1) no real-learner recruitment plan, (2) accent robustness unmeasurable without learners, (3) guardrail not adversarially tested, (4) post-hoc branching. → G-006.


Axis 8 — Governance, Decision-Making, and Communication — confidence 0.68

Forcing question Evidence Finding
Decision-maker when executives disagree? D-001 ("user-directed"); no governance body, no named sponsor The human "user" is the sole decision-maker. No sponsor, no committee. For an autonomous CI project, the orchestrator + user play this role. No disagreement-resolution mechanism exists — but with one decision-maker, none is needed yet.
Governance cadence / escalation pattern? config.json escalation_hooks (deploy, delete_data, merge_to_main); escalation_timeout 300s Escalation hooks exist for operational actions (deploy/delete/merge) but not for project-level risks (R4 failure, scope drift, recruitment failure). No cadence — the pipeline stages are the cadence.
Omissions from status reports? .ciagent artifacts are the status report Thorough on architecture/requirements/risks. Omit: timeline, budget, sponsor, recruitment plan, real-learner validation, no-go actions. These omissions are the grill findings.
Stop-the-project trigger? PLAN SLICE-01 "go/no-go gate" — but no-go actions undefined No explicit stop trigger. The SLICE-01 gate is the closest but its no-go branch is a blank. No pre-agreed kill criteria.

Axis verdict: Governance is minimal — appropriate for an autonomous CI project but with two gaps: no-go actions undefined, no project-level escalation for non-operational risks. → G-007 (ties to G-003).


Axis 9 — Change, Adoption, and Operational Readiness — confidence 0.80

Forcing question Evidence Finding
Who uses v0.1, how does their work change? D-007 (single hardcoded learner "Alex", no auth); PERSONAS.md (Aspiring Alex persona) No real users. v0.1's "learner" is a hardcoded SQLite row (learner-1, "Alex"). No real human will use v0.1. This is a dev harness, not a pilot.
Plan to get 50 real learners? None. No recruitment plan, no channel, no budget, no timeline for recruitment. PERSONAS.md L103 describes "Aspiring Alex" as a persona, not a recruitment target. Missing entirely. This is the most serious finding. The "pilot" framing (ROADMAP, PLAN) implies learners; the reality (D-007) is a hardcoded profile.
Ops/support involved now or handed finished product? No ops team; single pilot host (ARCHITECTURE L88-95) N/A for a dev harness. No production operations to hand off. Acceptable.
Rollback plan? Greenfield — no production system to roll back to N/A. Acceptable.
Success criteria validated with judges? PLAN §1.2 (10 tech exit criteria); no adoption/success criteria validated with learners Exit criteria are all technical (latency, DB rows, guardrail unit tests). No adoption criteria. No one has validated that "a learner completes a session" = success with actual learners.

Axis verdict: v0.1 has no real learners and no plan to get them. It is a tech-validation harness, not a pilot. This is the most serious finding — not because it blocks execution, but because the "pilot" framing is misleading. → G-008.


Meta — Closing Review — confidence 0.75

Forcing question Finding
What would the auditor flag? (1) No timeline; (2) no budget; (3) no named sponsor; (4) no recruitment plan; (5) "pilot" framing overstated; (6) branching is post-hoc not runtime; (7) go/no-go no-go actions undefined; (8) guardrail not adversarially tested; (9) thesis-critical constraints (C-2, C-3) all deferred.
What is the project NOT doing that it should? Recruiting real learners. Adversarially testing guardrails. Estimating timeline/budget. Defining no-go actions. Testing debrief quality (not just existence).
Simplest 80%-of-value version? v0.1 is already the simplest version. One scenario, one voice, no mastery. Correctly scoped. The over-scoping risk is low; the under-scoping risk (doesn't validate thesis) is real but acknowledged by design (D-002).
What must be true for success in 90 days? (a) R4 latency is measurable and has a viable path to <600ms — likely (SLICE-01); (b) voice loop works end-to-end — likely (SLICE-02); (c) debrief generates meaningful, non-generic coaching — unverified (no quality test in plan); (d) real learners use it — false today (no recruitment plan). (c) and (d) are the gaps.

Binding Decisions

ID Decision Rationale Confidence Alternatives
G-001 v0.1 is explicitly a tech-validation milestone, not market validation. The thesis-critical constraints (C-2: $100 Android/2G, C-3: $3/learner) are deferred and unmeasured. v0.1 success must not be reported as product-market-fit or thesis validation. D-012 admits "tech-validation, not unit-economics"; C-2/C-3 both relaxed per REQUIREMENTS L116-117. v0.1 validates the voice loop on the least hard configuration (Canada, cloud, high bandwidth). The moat (low-bandwidth/mobile/B2C-apprentice per RESEARCH L279) is unproven. 0.78 Claim thesis validation at v0.1 (false); enforce C-2/C-3 in v0.1 (premature, wrong milestone)
G-002 v0.1's branch point is a post-hoc outcome classification (LLM-as-judge at session end, offline), not a runtime conversation fork. The claim "exercises branching" (D-010 rationale) is overstated. Phase 2+ must validate true in-flight branching before claiming the scenario engine works. PLAN TASK-03-06 + D-P1-05 confirm classifier runs "offline at session end"; conversation is linear; Pipecat Flows is wired but the branch does not change in-flight behavior. 0.80 Redefine v0.1 branching as runtime (adds latency + complexity); drop the branch entirely (loses data-model validation)
G-003 The SLICE-01 go/no-go gate must define explicit no-go actions before EXECUTE: (a) if e2e >600ms with Cartesia but ≤600ms with Piper → swap TTS to Piper (SLICE-02 pre-stage); (b) if e2e >600ms even with Piper → evaluate self-hosted gemma4:e4b for LLM hop; (c) if e2e >600ms with both mitigations → escalate: reduce latency target for v0.1 or rethink architecture. "Measure and decide" without defined decisions is not a gate. PLAN L44/L227 call SLICE-01 a "go/no-go gate" but define no no-go branch. ARCHITECTURE L78 says "must be spiked" but not what failure triggers. A gate with no defined failure action is a measurement, not a gate. 0.75 Leave no-go undefined (current state — not a real gate); define a hard kill (too aggressive for a foundation milestone)
G-004 No calendar timeline is acceptable for v0.1 Phase 0 (pre-execution, autonomous project, no external deadline). Phase 1 EXECUTE should add per-slice rough effort estimates (even token-budget-order) to enable burn-rate tracking and parallelism planning. The 10 Phase 1 exit criteria (PLAN §4) compensate for the missing timeline by providing an unambiguous definition of done. No timeline in any document (ROADMAP, PLAN, CHECKPOINT). Defensible for pre-execution; not defensible for EXECUTE where parallelism (D-P1-04) and burn-rate need sizing. Exit criteria are strong (10 testable items). 0.65 Add full Gantt timeline now (premature for autonomous project); proceed with no estimates at EXECUTE (no burn-rate visibility)
G-005 v0.1 cost logging (SLICE-04 TASK-04-04) is the correct measurement infrastructure, but v0.1 logged costs will NOT be representative of at-scale per-learner cost. Ollama tier-based pricing (Pro/Max plan, not per-token) + Canada cloud + low volume = the most expensive configuration. The $3/learner target requires self-hosted gemma4:e4b + Piper (post-pilot path). Cost representativeness must be re-measured in a later milestone with self-hosted models before making unit-economics claims. RESEARCH L60 ("usage-tier pricing is not unit-economics-friendly at scale"); D-012 (no enforced ceiling); D-020 (self-host e4b is post-pilot path). The logged cost informs the measurement method, not the number. 0.72 Treat v0.1 logged cost as representative (false); enforce $3 ceiling in v0.1 (premature, D-012 rejects)
G-006 The single biggest unflagged v0.1 risk is the absence of a real-learner recruitment plan. v0.1 as scoped will produce a dev-harness demo (hardcoded learner-1 "Alex"), not a pilot with learners. This does not block tech validation (which can proceed without learners) but blocks any "pilot" claim. Accent robustness (R9) and adoption cannot be validated without real learners. Recruitment is deferred to a later milestone. D-007 (hardcoded profile, no auth); PERSONAS.md (persona roster, not recruitment plan); no recruitment plan/budget/channel in any document. The "pilot" language in ROADMAP/PLAN implies learners; the reality is a dev harness. 0.80 Block v0.1 until recruitment plan exists (too conservative for tech validation); claim pilot status at v0.1 (false)
G-007 The SLICE-01 go/no-go gate is the de facto stop-the-project trigger, but its no-go branch actions are currently undefined (ties to G-003). Additionally, no project-level escalation path exists for non-operational risks (R4 failure, scope drift, recruitment failure) — only operational hooks (deploy/delete/merge per config.json). Define no-go actions per G-003 before EXECUTE. config.json escalation_hooks cover operational actions only; PLAN L227 gate has no no-go definition; no stop trigger in any document. 0.70 Add a governance committee (overhead for autonomous project); proceed with no stop trigger (high-risk by definition)
G-008 v0.1 must be explicitly understood as a tech-validation harness, not a learner pilot. The "pilot" framing in ROADMAP L33 and PLAN L15 should be read as "tech pilot," not "learner pilot." Real-learner recruitment, adoption validation, and accent robustness on real speech are deferred to a later milestone. This is a framing correction, not a scope change — v0.1's technical scope is correct. D-007 (pilot harness, not production multi-user); D-012 (tech-validation milestone); G-006 (no recruitment plan). The technical scope (one scenario, one voice, debrief, SQLite) is right; the labeling oversells it. 0.82 Relabel as "v0.1 tech-validation" formally (would modify PROJECT/ROADMAP — grill surfaces, doesn't rewrite); proceed with "pilot" framing as-is (misleading)

Escalations

None. All nine axes plus meta resolved with confidence ≥ 0.60. The two findings closest to escalation threshold:

  1. No-go action definition (G-003/G-007, confidence 0.70-0.75): resolvable with evidence — the go/no-go gate exists, it just needs its no-go branch specified. Not an escalation; a binding pre-EXECUTE refinement.
  2. Real-learner recruitment (G-006/G-008, confidence 0.80): resolvable with evidence — D-007 and D-012 already admit v0.1 is a tech-validation harness. The binding decision makes the implication explicit and defers recruitment. Not an escalation; a framing correction.

Summary of Most Serious Findings

  1. "Pilot" is a misnomer (G-006, G-008). v0.1 has no real learners, no recruitment plan, no recruitment budget. It is a tech-validation harness with a hardcoded SQLite row ("Alex"). The technical scope is correct; the framing oversells it. Accent robustness and adoption are unvalidatable without learners.

  2. All thesis-defining constraints are deferred (G-001). The Praxis moat is "$100 Android on 2G at $3/learner" (RESEARCH L279). v0.1 relaxes C-2 (2G/device) and C-3 ($3/learner). It validates the voice loop on the easiest, most expensive configuration (Canada, cloud, high bandwidth, Ollama tier pricing). v0.1 success must not be reported as thesis validation.

  3. "Branching scenario" is post-hoc, not runtime (G-002). The branch is an LLM-as-judge classification at session end, offline from the voice loop. The conversation is linear. The data model (branch field) is validated; the branching runtime is not.

  4. Go/no-go gate has no no-go actions (G-003, G-007). SLICE-01 is called a "go/no-go gate" but defines no failure actions. A gate with no defined no-go branch is a measurement, not a gate. Must be specified before EXECUTE.

  5. No timeline, no budget, no sponsor (G-004, G-005). Defensible for Phase 0 pre-execution in an autonomous project, but EXECUTE needs per-slice estimates for burn-rate tracking. v0.1 cost logging is methodologically correct but its numbers won't represent at-scale economics (Ollama tier pricing ≠ per-token unit economics).

What's done well (to be clear-eyed): Research grounding (D-013..D-020 are web-verified, not vendor-pitched), swappable interfaces (TTS/LLM/guardrail all behind abstractions from SLICE-02), risk front-loading (SLICE-01 spike before building), explicit out-of-scope (14 items), 10 testable exit criteria, vertical-slice discipline (5 slices, each demoable). This is a well-prepared foundation. The findings above are framing corrections and pre-EXECUTE refinements, not structural rework.


End of grill report. Verdict: PROCEED at confidence 0.72. 8 binding decisions (G-001..G-008), 0 escalations. Escalations visible via ciagent audit. This grill surfaces findings; it does not rewrite PROJECT.md, ROADMAP.md, or REQUIREMENTS.md. Binding decisions that warrant spec changes must be promoted explicitly by the user (e.g., via ciagent-clarify or a follow-up CLARIFY stage).