v0.3 milestone merged to main. Mastery scoring + competency rubrics + verifiable credentials (formative-tier) shipped. 13/13 REQ-IDs covered. Next milestone: v0.4 (operator tier — cohort dashboard + auth + Postgres). ---ci--- project: praxis phase: 2 milestone: v0.3 status: complete milestone_complete: true milestone_merged_to_main: true ---/ci---
33 KiB
Praxis v0.3 CIAgent Plan — GRILL Verdict (Red-Team Review)
Reviewer: adversarial technology executive (red-team) Subject: v0.3 execution plan (Mastery Scoring + Competency Rubrics) — 2 phases, 15 slices, 70 tasks Stance: plan is unfeasible, over-scoped, and too costly until evidence forces otherwise Date: 2026-08-03 Binding status: This GRILL verdict must be cleared (MUSTs resolved, FIXs tracked) before EXECUTE is authorized. Artifacts reviewed: PLAN.md, PROJECT.md, REQUIREMENTS.md, RESEARCH.md (v0.3 section), ARCHITECTURE.md (v0.3 section), ROADMAP.md
Verdict Legend
- MUST — blocks execution until fixed. The plan cannot enter EXECUTE with this issue open.
- FIX — fix during execution, non-blocking. Tracked as a P1 condition in VERIFY.
- ACCEPT — proceed as-is. The evidence clears the challenge.
Axis 1 — Feasibility
Forcing question: Can this actually be built in 2 execution phases (70 tasks)? Is the scope realistic for one milestone, or is it 2 milestones pretending to be one?
Challenge: The v0.3 scope spans seven independent subsystems (rubric/mastery engine, IRT, scenario library + ≥6 authored scenarios, 6-week path engine, W3C VC 2.0 issuer with Ed25519 + Status List, operator auth + argon2id + slowapi, Postgres-in-LXC + asyncpg, cohort dashboard + k-anonymity aggregation + React UI). This is not a milestone — it is a program. The PLAN.md phase-split rationale (lines 14-23) openly admits the scope "is too large for one execution phase" and splits into P1/P2, but both phases ship under the same v0.3 milestone tag (v0.1.6). The 70-task count is artificially compressed: SLICE-12 (VC issuer) is 6 tasks for a W3C VC 2.0 + Ed25519 + JCS + Bitstring Status List + public verification endpoint + key rotation — that is at minimum a 10-12 task slice on its own, and SLICE-13 (cohort aggregation with k-anonymity + nightly reconciliation + on-session-end hook) is similarly under-tasked at 4 tasks.
Evidence:
- PLAN.md:14-23 — "The v0.3 scope … is too large for one execution phase."
- PLAN.md:32, 334 — P1 = 38 tasks, P2 = 32 tasks, total 70 (excludes P3 review).
- REQUIREMENTS.md:20-66 — 11 functional REQ-IDs + 9 NFRs = 20 active requirements, the largest single-milestone REQ surface in the project's history (v0.1 was ~14, v0.2 was 16 deploy + 4 NFR).
- RESEARCH.md:769 (R-VC-01) — "No batteries-included Python VC lib → ~200 LOC custom code" — 200 LOC of custom crypto code is not a 6-task slice; it is a liability that demands more tests than the plan allocates (only 2 test tasks: TASK-12-05, TASK-12-06).
- SLICE-13 (PLAN.md:516-546) — 4 tasks for: on-session-end hook, k-anonymity suppression SQL, nightly reconciliation cron, and tests. The nightly reconciliation job alone (recompute all 7-day windows from raw events, correct drift, idempotent upsert) is a 2-3 task effort.
Binding verdict: FIX — The plan is feasible as a 2-phase program, but only if it is honestly re-labeled. The milestone should ship as v0.3 (P1 mastery core, v0.1.4) and v0.3.1 (P2 operator tier, v0.1.5), with the v0.3 milestone release (v0.1.6) being the merge of two separately-shipped, separately-verified patches. Do not pretend P1+P2 is one milestone release. Additionally, re-task SLICE-12 and SLICE-13: add 2 tasks each (one for VC Status List edge cases + key rotation drill, one for reconciliation idempotency + race-condition test). This is non-blocking — the wave structure survives — but the task counts must be honest before EXECUTE.
Axis 2 — Scope
Forcing question: Is REQ-DASH-01 (cohort dashboard + multi-tenant + auth) really v0.3, or was it correctly deferred in v0.1/v0.2 for a reason? Does D-031 (override D-007) open a Pandora's box?
Challenge: REQ-DASH-01 was explicitly deferred in v0.1 (REQUIREMENTS.md:152, "later/deferred") and v0.2. ROADMAP.md:94 places the "Employer / program dashboard" at v0.8. The v0.3 plan pulls it forward three milestones with the justification that mastery scoring "needs" the operator view. But mastery scoring (REQ-MAST-01/02) and VC issuance (REQ-MAST-03) work without a cohort dashboard — the dashboard is an operator feature, not a learner feature. D-031 overrides D-007 (single-learner/no-auth) and introduces a hybrid SQLite+Postgres topology, operator auth, argon2id, slowapi, asyncpg, a second Docker service, k-anonymity aggregation, and a React operator UI — none of which is required for the learner-facing mastery gate to function. This is scope creep dressed as a dependency.
Evidence:
- ROADMAP.md:94 — "v0.8 | Employer / program dashboard" (original placement).
- PROJECT.md:129 (D-031) — "overrides D-007 for the cohort-dashboard surface" — confidence 0.75, the lowest-confidence decision that expands scope.
- REQUIREMENTS.md:44 (REQ-DASH-01) — "Forces multi-tenant + operator auth (D-031)" — the word "forces" is doing a lot of work. The mastery gate (REQ-MAST-02) does not depend on the dashboard.
- PLAN.md:18 — P1 "works standalone (learner can practice, score, progress) without the operator tier." — This is an admission that the operator tier is separable.
- PROJECT.md:147 (D-049) — failure-injection stays off, further confirming the learner-facing mastery layer is the real v0.3 deliverable.
Binding verdict: MUST — Split the milestone. Ship v0.3 = P1 only (mastery core + IRT + scenarios + paths + VC issuance, since VC issuance is triggered by the mastery gate and is learner-facing per D-048). Defer REQ-DASH-01 + REQ-AUTH-01 + REQ-MT-01/02 + REQ-NFR-DASH-01/02 + REQ-NFR-AUTH-01 + REQ-NFR-MT-01 to v0.4 (operator tier), restoring the original ROADMAP intent. D-031 does open a Pandora's box: every hybrid-DB system eventually faces the "which store is the source of truth?" question, and shipping it under a learner-milestone tag hides that risk. If the team insists on keeping the dashboard in v0.3, rebrand the milestone as "v0.3: Mastery + Operator Tier" and accept that this is a 2-milestone program — but the cleaner answer is to defer the dashboard.
Axis 3 — Cost
Forcing question: What is the maintenance cost of Postgres-in-LXC, asyncpg, argon2, pynacl, slowapi, and ~200 LOC custom VC code? Is R-VC-01 (custom VC code) a liability vs using a library?
Challenge: The v0.3 dependency surface grows by at least 5 new pip packages (asyncpg, argon2-cffi, slowapi, pynacl, canonicaljson, base58 — actually 6) plus a Postgres service. Each is a CVE vector, a version-pin maintenance burden, and a CI complexity adder. The ~200 LOC custom VC code (R-VC-01) is the most concerning: cryptographic code written by an AI agent is a liability regardless of test coverage. The W3C VC 2.0 + eddsa-jcs-2022 cryptosuite has subtle canonicalization edge cases (e.g., JSON number representation, key ordering, URI normalization) that unit-test round-trips do not catch — only interop tests against an independent verifier do, and the plan has zero interop tests.
Evidence:
- RESEARCH.md:769 (R-VC-01) — "~200 LOC custom code" — confidence 0.75. The mitigation is "unit-test signature/verify round-trip," which only proves the code is self-consistent, not that it is W3C-compliant.
- PLAN.md:504-513 (TASK-12-05, TASK-12-06) — VC tests are sign/verify round-trip, tamper detection, JCS determinism, status list, revocation, key rotation. No interop test against an external verifier (e.g., Verifiable Credential JS verifier, Digital Credentials Verifier).
- PROJECT.md:140 (D-042) — issuer key encrypted at rest with a root key from secrets. Key management is hand-rolled (init_issuer_key, encrypt, store, rotate). This is a security-engineer task, not a backend task, and the plan assigns it to security-engineer (good), but the rotation drill (D-042 "new key + old marked superseded") is not tested end-to-end except in TASK-12-06 which only checks "old VC still verifies against archived public key" — it does not test the operational rotation procedure (generate new key, archive old, re-sign new VCs, update verificationMethod URL).
- RESEARCH.md:772 (R-MT-01) — Postgres-in-LXC resource contention, confidence 0.65 — the lowest-confidence technical risk. Memory bump to 6GB is a guess, not a measurement.
Binding verdict: MUST — Two conditions before EXECUTE:
- Add a VC interop test (TASK-12-07): verify a Praxis-issued VC against at least one external W3C VC verifier (e.g., the
digitalbazaar/vc-verifieror a JS@digitalcredentials/vcverifier). Round-trip self-verification is insufficient for cryptographic claims. Without this, R-VC-01 is an unmitigated liability. - Add a key-rotation operational test (TASK-12-08): end-to-end drill — issue N VCs with key A, rotate to key B, issue M VCs with key B, verify all N+M VCs still verify (N against archived key A, M against active key B), revoke one of each, verify revocation. This is the one crypto procedure that, if broken, silently invalidates every credential ever issued.
The Postgres/argon2/slowapi maintenance cost is ACCEPT — these are well-maintained, widely-used libraries. The liability is concentrated in the custom VC code.
Axis 4 — Technical Risk
Forcing question: R-MAST-01 (N=3 thin for credential), R-AUTH-01 (Secure cookie + no TLS), R-MAST-02 (LLM hallucinated quotes), R-IRT-01 (cold start) — which are MUST-FIX before execution vs ACCEPT?
Challenge: The plan treats all four as "Open Questions Deferred to EXECUTE" (PLAN.md:687-693). That is insufficient. R-MAST-01 is a credibility risk: if the VC is labeled as a mastery credential and employers treat it as high-stakes, N=3 with G≈0.5-0.6 is defensible only if the credential is explicitly labeled formative. R-AUTH-01 is a security risk: relaxing the Secure cookie flag for a no-TLS pilot means session cookies travel in cleartext — if the operator bridge IP is on a shared network (vmbr0 DHCP), any host on the bridge can sniff the operator session. R-MAST-02 is the highest-confidence mitigation (fuzzy-match quotes), but the plan's fallback ("empty evidence + log warning") means a session could silently score as a zero with no learner-visible signal. R-IRT-01 is benign (cold-start fallback to fixed difficulty).
Evidence:
- RESEARCH.md:766 (R-MAST-01) — confidence 0.62, below the 0.70 decision threshold. Mitigation: "Label v0.3 VC as formative." This label is not in the PLAN.md VC payload (TASK-12-02) or the REQ-MAST-03 requirement text.
- RESEARCH.md:771 (R-AUTH-01) — "Secure cookie flag fails without TLS." PLAN.md:450 resolves this with
PRAXIS_COOKIE_SECURE=falseenv default. This ships a known-insecure default. - PLAN.md:154 (TASK-03-01) — "on final failure, fall back to empty evidence + log warning." Empty evidence → rule scorer has no signals → every criterion scores level 1 (fail) → scenario fails → learner sees a failed session with no explanation. This is a UX and fairness bug.
- RESEARCH.md:774 (R-IRT-01) — mitigation confidence 0.75, "fall back to scenario.difficulty until ≥5 observations." ACCEPT.
Binding verdict: MUST — Three conditions:
- R-MAST-01: Add
credentialTier: "formative"(or equivalent) to the VC payload (TASK-12-02) and to the verification endpoint response (TASK-12-04). Update REQ-MAST-03 to require this label. Without it, the credential is misleading. - R-AUTH-01: Do not ship
PRAXIS_COOKIE_SECURE=falseas a default. Either (a) require TLS for the operator surface (add a Traefik sidecar or Caddy in front of/api/operator/*), or (b) bind the operator surface to127.0.0.1only (loopback) so cookies never traverse the bridge. A cleartext cookie on a shared bridge is a MUST-FIX. - R-MAST-02: Change the fallback in TASK-03-01 from "empty evidence + log warning" to "empty evidence → mark scenario as
scoring_inconclusive→ do not count toward gate, do not penalize learner, surface 'technical issue, please retry' in the debrief." A silent fail-to-zero is unacceptable.
R-IRT-01: ACCEPT — cold-start fallback is sound.
Axis 5 — Requirements Coverage
Forcing question: Does the plan actually cover all 20 REQ-IDs, or are some hand-waved? Check the coverage matrix in PLAN.md against REQUIREMENTS.md.
Challenge: The PLAN.md coverage matrix (lines 657-683) claims "20 REQ-IDs covered, 0 partial, 0 deferred." Let me audit the suspicious ones.
Evidence (audit):
| REQ-ID | Claimed coverage | Actual coverage | Verdict |
|---|---|---|---|
| REQ-MAST-03 | P2 SLICE-12 "VC issuer" | SLICE-12 implements issuance + verification + revocation. But REQ-MAST-03 says "Issued when a mastery gate opens" — the trigger is in P1 SLICE-07 (TASK-07-01, "path_engine.check_gate + advance_week") and the issuance is in P2 SLICE-12. The P1→P2 handoff for VC issuance is not in any task — who calls issuer.issue_credential() when the week-final gate opens? TASK-07-01 says "(6) record mastery_gate_event" but does NOT call the VC issuer (VC issuer is P2). D-048 says "issue VC if week-final gate" but the plan splits the gate-open (P1) from the issuance (P2). Gap: no task wires the P1 gate-open event to the P2 VC issuer. |
FIX — add a task (either in SLICE-07 or SLICE-12) that defines the P1→P2 VC-issuance contract: a mastery_gate_events row with gate_opened_at is the trigger; P2's VC issuer polls/receives this event and issues. |
| REQ-NFR-MAST-02 | P1+P2 SLICE-07, 09 "gate auditability (SQLite + Postgres)" | SLICE-07 records the event in SQLite; SLICE-09 defines the Postgres mastery_gate_events table; but no task mirrors the SQLite event to Postgres. The "mirror" is implied but not tasked. |
FIX — TASK-13-01 (cohort aggregation hook) should explicitly mirror mastery_gate_events from SQLite to Postgres, or add a dedicated mirroring task. |
| REQ-SCEN-04 | P1 SLICE-02, 06 "expert-authored format + AI variation hooks" | SLICE-02 adds generated_from and intent_hash fields (the hook). SLICE-06 authors expert scenarios. But no task implements the AI-variation review pipeline (_pending/ dir → expert review → library promotion). RESEARCH.md:758 describes it; PLAN.md does not task it. |
ACCEPT — REQ-SCEN-04 says "AI-generated variations" with "expert review" — the hook is the schema field; the pipeline can be deferred. The plan is honest that AI variations are "in P2 or later" (SLICE-06 goal line 246). |
| REQ-NFR-DASH-02 | P2 SLICE-13 "freshness ≤24h" | SLICE-13 has a nightly reconciliation job (TASK-13-03) at 02:00. If the on-session-end hook (TASK-13-01) fails or lags, freshness depends on the nightly job. ≤24h is satisfied if the nightly job runs. But there is no task for monitoring or alerting on job failure. | FIX — add a health check for the nightly job (log last-run timestamp, surface in operator dashboard or /health). Non-blocking. |
| REQ-NFR-VC-02 | P2 SLICE-12 "revocation latency — within 1 sync of status list" | "1 sync" is undefined. Is it 1 sync of the status list blob? Is the status list in-memory or fetched on every verify? TASK-12-04 (verification endpoint) does not specify caching of the status list. | FIX — clarify in TASK-12-04: status list is fetched from Postgres on every verification (no cache), so revocation latency = next verify call. Non-blocking. |
Binding verdict: FIX — The coverage matrix is mostly honest (18/20 fully covered), but the P1→P2 VC-issuance wiring gap (REQ-MAST-03) is a real hole — without a task that defines the trigger contract, the VC issuer will be built but never called. Add the wiring task. The other three FIXs are minor clarifications.
Axis 6 — Architecture
Forcing question: Is hybrid SQLite + Postgres (D-031) a maintainable pattern or a future migration nightmare? Is "no cross-DB joins" realistic for the cohort dashboard queries?
Challenge: Hybrid polyglot persistence is a known anti-pattern when the two stores hold related data and there is no canonical source of truth. Here, mastery_gate_events exists in both SQLite (P1, the learner's local record) and Postgres (P2, the operator audit log). Which is canonical? If they diverge (e.g., SQLite write succeeds, Postgres mirror fails due to pool exhaustion), the cohort dashboard shows stale data while the learner sees correct data — and there is no reconciliation except the nightly job (which recomputes from Postgres mastery_gate_events, not from SQLite). This means the nightly job recomputes from a possibly-incomplete Postgres copy. The "no cross-DB joins" rule is realistic only if the cohort dashboard never needs to join learner-local data (e.g., θ distribution by path) with operator data — but the dashboard's "progression" and "failure patterns" views implicitly need both the learner's session outcomes (SQLite) and the operator's aggregate view (Postgres). The plan resolves this by aggregating at session-end (writing the aggregate to Postgres), so the dashboard reads only Postgres — but this means the aggregate is a derived copy, and the "no cross-DB joins" rule is maintained by duplicating data, not by query-time joins. This is workable but fragile.
Evidence:
- ARCHITECTURE.md:356-379 — "The two stores never share a session and never join via cross-DB FKs (
learner_refis an opaque string in Postgres)." — the design is clean if the mirror is reliable. - RESEARCH.md:746 — "No cross-DB joins via
learner_ref—learner_refis an opaque string, not a FK." — correct, butlearner_refis still a logical join key. If the SQLite learner is deleted and re-created, the Postgreslearner_refdangles. - PLAN.md:528-529 (TASK-13-01) — "after P1's mastery hooks fire, call
aggregator.upsert_aggregate(...)" — this is a synchronous call after the SQLite write, in the session-end path. If Postgres is down, does the session-end fail? The plan does not specify failure semantics. - RESEARCH.md:748 — migration strategy: "SQLite volume untouched → learner path never regresses." Good, but the operator path regresses if Postgres is down.
Binding verdict: FIX — Three conditions:
- Define failure semantics for the Postgres mirror (in TASK-13-01): if
upsert_aggregatefails (Postgres down, pool exhausted), the learner session-end must still succeed (SQLite write is canonical for the learner). The aggregate failure is logged and reconciled by the nightly job. This makes SQLite the learner-canonical store and Postgres the operator-derived store — state this explicitly in ARCHITECTURE.md. - Make
learner_refa stable, opaque, non-reusable identifier (e.g., a UUID generated once and stored in SQLite, never reused). Add this to TASK-02-01 or a new task. Without it, the "no FK" rule is a leaky abstraction. - Add a Postgres-readiness guard to the operator API: if Postgres is down,
/api/operator/cohort/*returns 503 (not 500 with a stack trace). Add to TASK-14-02.
The hybrid pattern is ACCEPT with these conditions — it is the correct pilot choice (don't migrate learner state to Postgres prematurely), but the failure semantics must be explicit.
Axis 7 — Testing
Forcing question: 70 tasks, but how many have tests? Is the test strategy (mocked LLM for evidence extraction, testcontainers for Postgres) viable, or are there untestable critical paths?
Challenge: Let me count test tasks across the plan.
Evidence (test task audit):
| Slice | Tasks | Test tasks | Test ratio |
|---|---|---|---|
| SLICE-01 | 4 | 1 (TASK-01-04) | 25% |
| SLICE-02 | 4 | 1 (TASK-02-04) | 25% |
| SLICE-03 | 5 | 2 (TASK-03-04, 03-05) | 40% |
| SLICE-04 | 4 | 2 (TASK-04-03, 04-04) | 50% |
| SLICE-05 | 4 | 1 (TASK-05-04) | 25% |
| SLICE-06 | 3 | 1 (TASK-06-03) | 33% |
| SLICE-07 | 4 | 2 (TASK-07-03, 07-04) | 50% |
| SLICE-08 | 3 | 3 (all test/verification) | 100% |
| SLICE-09 | 4 | 1 (TASK-09-04) | 25% |
| SLICE-10 | 3 | 0 | 0% — infra, acceptable |
| SLICE-11 | 5 | 2 (TASK-11-04, 11-05) | 40% |
| SLICE-12 | 6 | 2 (TASK-12-05, 12-06) | 33% — too low for crypto code |
| SLICE-13 | 4 | 1 (TASK-13-04) | 25% — too low for k-anonymity |
| SLICE-14 | 4 | 1 (TASK-14-04) | 25% |
| SLICE-15 | 4 | 1 (TASK-15-04) | 25% |
| SLICE-16 | 3 | 3 (all integration/verification) | 100% |
| Total | 70 | 24 | 34% |
Critical untestable paths:
- LLM evidence extraction (TASK-03-01) — the plan mocks the LLM (good for unit tests), but there is no test that runs against the real LLM with a real transcript. A mocked LLM proves the scoring logic, not that the extraction prompt works. This is a fundamentally untestable in CI path — the only test is manual/staging.
- k-anonymity suppression (TASK-13-02) — the test (TASK-13-04) checks "cell with 9 learners → suppressed, 10 → shown." But it does not test the differencing attack (comparing two adjacent windows to re-identify a learner who appears in one but not the other). RESEARCH.md:754 says "limit to pre-defined 2-D views to block differencing attacks" — but there is no test that the API enforces only pre-defined views (i.e., that an operator cannot request an arbitrary
path × week × outcome3-D view). - Nightly reconciliation (TASK-13-03) — no test for "reconciliation corrects drift." TASK-13-04 tests "reconciliation correctness" but not drift correction (insert a bad aggregate, run reconcile, verify it's fixed).
- VC verification endpoint (TASK-12-04) — tested via TASK-12-06, but only with Praxis-issued VCs. No interop test (see Axis 3).
Binding verdict: FIX — Four conditions:
- Add a real-LLM smoke test (in SLICE-08 or SLICE-16): run one session transcript through the actual deepseek-v4-flash:cloud evidence extractor and verify the output is valid JSON with fuzzy-matching quotes. This runs only in staging (requires OLLAMA_API_KEY), gated by an env flag. The mocked-LLM tests stay in CI.
- Add a k-anonymity differencing-attack test (TASK-13-04 extension): verify that the cohort API rejects arbitrary 3-D view requests, and that two adjacent 7-day windows cannot re-identify a single learner appearing in only one.
- Add a reconciliation drift-correction test (TASK-13-04 extension): insert a deliberately-wrong aggregate, run
reconcile_cohort(), verify it's corrected. - Add the VC interop test (per Axis 3, MUST condition).
The mocked-LLM + testcontainers strategy is ACCEPT for CI. The gaps are in integration and security testing, not unit testing.
Axis 8 — Phase Split
Forcing question: Is P1/P2 the right split? Should VC issuance (P2 SLICE-12) be in P1 with mastery gates (P1 SLICE-07) since they trigger on the same event? Is the P1→P2 dependency clean?
Challenge: The plan splits VC issuance (P2) from mastery-gate-open (P1) even though D-048 says "issue VC if week-final gate." This means P1 ships (v0.1.4) with mastery gates that open but no credential is issued — the learner reaches mastery and gets... nothing portable. The VC issuer arrives in P2 (v0.1.5). This is a user-visible gap: a learner who completes the path in v0.1.4 has no credential. The plan's phase-split rationale (lines 14-23) says P1 "works standalone (learner can practice, score, progress) without the operator tier" — but VC issuance is not the operator tier; it is a learner-facing consequence of mastery (D-048). VC issuance should be in P1.
Conversely, the operator auth + Postgres + cohort dashboard is correctly P2 — those are operator-tier.
Evidence:
- PROJECT.md:146 (D-048) — "issue VC if week-final gate" — VC issuance is a mastery-gate consequence, not an operator feature.
- PLAN.md:18 — "P1 works standalone" — but "standalone" here silently drops the VC, which is a REQ-MAST-03 requirement.
- PLAN.md:663 (coverage matrix) — REQ-MAST-03 is listed as P2 SLICE-12. But REQ-MAST-03 is a mastery requirement, not an operator requirement.
- PLAN.md:282 (TASK-07-01) — P1 session-end hook does steps 1-6 but step 6 is "record mastery_gate_event" — no VC issuance call. The VC issuance is orphaned in P2 with no trigger from P1.
Binding verdict: MUST — Move SLICE-12 (VC issuer) to P1, after SLICE-07 (mastery gates), as a new Wave-4 slice in P1 (parallel with SLICE-08). This requires:
- VC issuer needs
issuer_keysstorage — use SQLite for P1 (the issuer_keys table moves to SQLite for v0.3; Postgres takes over in v0.4 when the operator tier arrives). Or, if Postgres is required for VC, then Postgres must also move to P1 — which inflates P1 further and reinforces the Axis 2 verdict (split the milestone). - The cleaner resolution: defer VC issuance to v0.3.1 (P2) and accept that v0.1.4 (P1) ships mastery gates without credentials — but label this explicitly in the P1 ship notes ("VC issuance in v0.1.5"). Do not claim REQ-MAST-03 is covered in P1.
Either resolution is acceptable. The current plan — which implies VC issuance is triggered by P1's gate-open but tasks it in P2 with no wiring — is not acceptable. Pick one: (a) VC in P1 with SQLite-backed issuer keys, or (b) VC explicitly deferred to P2 with P1 shipping "mastery gates, no credential yet."
The P1→P2 dependency is otherwise clean (P2 reads P1's mastery_gate_events and session outcomes). ACCEPT on the dependency structure.
Axis 9 — Decisions
Forcing question: Are D-031..D-049 (12 clarify + 7 specify decisions) well-grounded, or are any below the 0.60 confidence threshold? Is D-049 (failure-injection stays off) a mistake given mastery scoring scores recovery from failure branches?
Challenge: Let me audit confidences against the 0.70 threshold (the project's apparent decision-acceptance floor).
Evidence (confidence audit of D-031..D-049):
| ID | Confidence | Below 0.70? | Verdict |
|---|---|---|---|
| D-031 | 0.75 | No | ACCEPT — but see Axis 2 (scope creep). |
| D-032 | 0.70 | At threshold | ACCEPT — N=3 is formative-only per R-MAST-01. |
| D-033 | 0.70 | At threshold | ACCEPT — W3C VC 2.0 is a stable standard. |
| D-034 | 0.70 | At threshold | ACCEPT — k=10 is the conventional minimum. |
| D-035 | 0.70 | At threshold | ACCEPT — 1PL/Rasch is the simplest IRT. |
| D-036 | 0.80 | No | ACCEPT. |
| D-037 | 0.75 | No | ACCEPT. |
| D-038 | 0.80 | No | ACCEPT — deterministic scoring is the right call. |
| D-039 | 0.80 | No | ACCEPT. |
| D-040 | 0.80 | No | ACCEPT. |
| D-041 | 0.75 | No | ACCEPT — but R-AUTH-01 (Secure cookie) is a MUST-FIX (Axis 4). |
| D-042 | 0.70 | At threshold | ACCEPT — but key rotation drill is a MUST (Axis 3). |
| D-043 | 0.80 | No | ACCEPT. |
| D-044 | 0.75 | No | ACCEPT. |
| D-045 | 0.70 | At threshold | ACCEPT — but failure semantics are a FIX (Axis 6). |
| D-046 | 0.80 | No | ACCEPT — θ in SQLite is correct. |
| D-047 | 0.70 | At threshold | ACCEPT — 6 scenarios is tight but defensible for formative. |
| D-048 | 0.75 | No | ACCEPT — but the P1/P2 split breaks the trigger wiring (Axis 8). |
| D-049 | 0.80 | No | See below. |
D-049 (failure-injection stays off): The challenge is whether this is a mistake. The rubric (SLICE-01) has a "de-escalation" criterion (weight 0.20), and RESEARCH.md:718 says "de-escalation up-weights to ~0.40 if the escalate branch triggers." The escalate branch is a naturally-occurring failure branch in cs_refund_ca_v01 (D-010), not an AI-provoked failure. So mastery scoring does score recovery from a failure branch — the naturally-occurring one. D-049 keeps AI-provoked failure injection off, which is correct: the rubric's de-escalation criterion is exercised by the existing branch, and adding AI-provoked failures would couple mastery scoring to a new feature (scope creep). D-049 is well-grounded.
However, there is a subtle gap: the rubric weights are static in the YAML (empathy 0.35, resolution 0.30, de-escalation 0.20, professionalism 0.15 per TASK-01-01). RESEARCH says de-escalation "up-weights to ~0.40 if the escalate branch triggers" — but TASK-01-01 does not mention dynamic re-weighting based on branch outcome. Either the weights are static (and the "up-weight" is a future feature) or they are dynamic (and the plan is missing a task). This is a FIX — clarify in TASK-01-01 whether weights are static or branch-dependent. If static, update RESEARCH.md to note the up-weight is deferred.
Binding verdict: ACCEPT on all D-031..D-049 confidences (none below 0.60; the floor is 0.70, which is the project's threshold). FIX on the de-escalation weight ambiguity (static vs dynamic) in TASK-01-01. D-049 is ACCEPT — failure-injection stays off is the correct call; the naturally-occurring escalate branch exercises the de-escalation criterion.
Summary Table
| # | Axis | Forcing question (short) | Verdict |
|---|---|---|---|
| 1 | Feasibility | 2 phases / 70 tasks realistic? | FIX — re-label as 2-milestone program; re-task SLICE-12/13 (+2 tasks each) |
| 2 | Scope | REQ-DASH-01 really v0.3? D-031 Pandora's box? | MUST — split milestone; defer dashboard to v0.4 (or rebrand honestly) |
| 3 | Cost | Custom VC code liability? Maintenance burden? | MUST — add VC interop test + key-rotation operational test before EXECUTE |
| 4 | Technical risk | R-MAST-01/R-AUTH-01/R-MAST-02/R-IRT-01 | MUST — label VC formative; fix Secure cookie; fix silent-fail-to-zero fallback |
| 5 | Requirements coverage | 20 REQ-IDs fully covered? | FIX — wire P1→P2 VC-issuance trigger; mirror SQLite→Postgres gate events; minor NFR clarifications |
| 6 | Architecture | Hybrid SQLite+Postgres maintainable? | FIX — define Postgres-failure semantics; stabilize learner_ref; add 503 guard |
| 7 | Testing | Test strategy viable? Untestable paths? | FIX — add real-LLM smoke test, differencing-attack test, drift-correction test, VC interop test |
| 8 | Phase split | VC issuance in P2 but triggers on P1 event? | MUST — move VC to P1 (SQLite-backed) OR explicitly defer to P2 with honest labeling |
| 9 | Decisions | D-031..D-049 below 0.60? D-049 a mistake? | ACCEPT — all confidences ≥0.70; D-049 correct; FIX de-escalation weight ambiguity |
Final Recommendation: GO-WITH-CONDITIONS
The v0.3 plan is not approved for EXECUTE as-is. It is a well-researched, well-structured plan that suffers from two structural flaws: (1) it is two milestones pretending to be one, and (2) it splits a learner-facing consequence (VC issuance) from its trigger (mastery gate) across a phase boundary without wiring.
MUST conditions (blocking — must be resolved in PLAN before EXECUTE):
-
Axis 2 — Split the milestone. Either (a) defer REQ-DASH-01 + operator tier to v0.4, shipping v0.3 = P1 + VC issuance only; or (b) rebrand v0.3 as a 2-milestone program (v0.3 + v0.3.1) with separate ship/verify cycles. Do not ship P1+P2 under one milestone tag.
-
Axis 3 — Add VC interop test + key-rotation operational test. Custom crypto code without interop verification is an unmitigated liability. Add TASK-12-07 (interop) and TASK-12-08 (rotation drill).
-
Axis 4 — Fix three technical risks. (a) Label VC as
formativein payload + verification response + REQ-MAST-03 text. (b) Do not shipPRAXIS_COOKIE_SECURE=falseas default — use TLS or loopback-binding for the operator surface. (c) Change evidence-extraction fallback from silent-fail-to-zero toscoring_inconclusivewith learner-visible retry signal. -
Axis 8 — Resolve the VC-issuance phase split. Either move SLICE-12 to P1 (with SQLite-backed issuer keys) or explicitly defer REQ-MAST-03 to P2 and label P1 as "mastery gates, no credential yet." The current plan's implicit wiring is a gap.
FIX conditions (non-blocking — tracked in VERIFY-P1/P2):
- Axis 1 — Re-task SLICE-12 and SLICE-13. Add 2 tasks each to honestly reflect the effort (VC edge cases + rotation drill; reconciliation idempotency + race test).
- Axis 5 — Wire the P1→P2 VC-issuance trigger (if VC stays in P2) and mirror SQLite→Postgres gate events explicitly in TASK-13-01.
- Axis 6 — Define Postgres-failure semantics (SQLite is learner-canonical, Postgres is operator-derived); stabilize
learner_refas a non-reusable UUID; add 503 guard on operator API. - Axis 7 — Add four tests: real-LLM smoke (staging-gated), k-anonymity differencing-attack, reconciliation drift-correction, VC interop (already a MUST).
- Axis 9 — Clarify de-escalation weight (static vs branch-dependent) in TASK-01-01.
ACCEPT items (proceed as-is):
- IRT 1PL/Rasch cold-start fallback (R-IRT-01).
- All decision confidences (D-031..D-049 ≥ 0.70, none below 0.60).
- D-049 (failure-injection stays off) — correct call.
- Mocked-LLM + testcontainers CI strategy.
- Hybrid SQLite+Postgres topology (with failure-semantics FIX).
- P1→P2 dependency structure (clean except for VC-issuance wiring).
Bottom line:
The plan is not unfeasible — the research is thorough, the architecture is sound, and the slice decomposition is reasonable. But it is over-scoped (two milestones in one tag) and under-tested in its highest-risk areas (custom crypto, k-anonymity, real-LLM extraction). Resolve the 4 MUST conditions, track the 5 FIX conditions, and this becomes a GO.