diff --git a/.ciagent/ARCHITECTURE.md b/.ciagent/ARCHITECTURE.md index 2fb56f6..86368c3 100644 --- a/.ciagent/ARCHITECTURE.md +++ b/.ciagent/ARCHITECTURE.md @@ -797,3 +797,85 @@ return `Skipped` when the resources are absent (`NoSuchBucket`/ `ResourceNotFoundException`). `RegressionReport.passed` is `all(r.status in ("Verified", "Skipped"))`. The gate passes at 18 Verified + 4 Skipped (0 Decayed/Broken). + +## v1.17 Addendum — Strategic Direction, Leadership Metrics & Unified Story (2026-08-04) + +The v1.17 milestone adds a telemetry/observability layer, a Decision +Ledger, a metrics export pipeline, a unified narrative deck, and a +durable strategic-direction artifact. This addendum documents the +architecture; the full research findings are in RESEARCH.md §v1.17. + +### New components + +| Component | Path | Purpose | +|-----------|------|---------| +| Event envelope | `core/metrics/event_envelope.py` | CloudEvents 1.0 envelope + `platform.*` semantic conventions (P1, REQ-187) | +| Per-run manifest writer | `core/metrics/run_manifest.py` | Emits `nova.run.started/completed/failed` events + writes `metrics/runs/.json` (P1, REQ-187) | +| Decision Ledger (SQLite) | `core/metrics/decision_ledger.py` | Extends `outbox_writer.py` → SQLite append-only hash-chain table; `ai.decision.made` + `attestation.recorded` events + outcome backfill (P1, REQ-188, D-121) | +| Infracost post-processor | `core/metrics/infracost_adapter.py` | Runs Infracost on plan JSON; emits `nova.cost.estimated{delta_usd}` (P1, REQ-187, D-120) | +| Metrics collector | `core/metrics/collector.py` | Reads all grounded signals (files + events) → SQLite cold store at `metrics/nova_metrics.db` (P2, REQ-189) | +| PowerBI export | `core/metrics/powerbi_export.py` | Emits CSV/JSON views to `metrics/powerbi/` (fact + dim + 8 deferred placeholder views) (P3, REQ-190) | +| Metrics schemas | `schemas/metrics_*.schema.json` | Schemas for all event types + fact/dim tables (P1–P2, REQ-187/189) | +| Metrics catalog | `docs/METRICS.md` + `docs/metrics/.md` | Canonical catalog + per-KPI definition-of-success docs (P4, REQ-195, D-127) | +| Unified narrative deck | `docs/presentations/nova-no-humans-platform.md` | Merged deck: Problem→Vision→How→Proof→Roadmap; x3 arc at deck+slide level (P5, REQ-196/197, D-130) | +| Strategic direction | `.ciagent/NORTH_STAR.md` | PO-authored durable vision/objectives/anti-goals/targets; read by CIAgent in every future `/ci-run` (P0, REQ-185/186) | + +### Modified components + +| Component | Change | Phase | +|-----------|--------|-------| +| `core/outbox_writer.py` | Extended to emit to SQLite append-only hash-chain table (Decision Ledger); `ai.decision.made` + `attestation.recorded` events added (P1, D-121) | P1 | +| `scripts/run_platform.sh` | Per-run manifest writer invoked; `$WORK/*.json` persisted to `metrics/runs/`; Infracost post-processor invoked after plan (P1) | P1 | +| `core/hitl_gates.py` | Emits `attestation.recorded` event to Decision Ledger on qa/prod/dr gate (P1, D-132) | P1 | +| `core/confidence_signal.py` | Emits `nova.confidence.computed` + `nova.ai.decision.made` events (P1, D-122) | P1 | +| `adapters/terraform/policy/checkov_adapter.py` | Emits `nova.policy.evaluated` event (P1) | P1 | +| `core/regression_verify.py` | Emits `nova.capability.verified` event; CAP-023 (metrics collector) + CAP-024 (deck structure) added (P1, P6) | P1, P6 | +| `pyproject.toml` | `addopts` gains `--junitxml=metrics/test-results.xml` + `--json-report` (P1, D-120) | P1 | +| `docs/presentations/` | Two old decks retired (deleted); unified deck added (P5, D-130) | P5 | + +### Telemetry/observability layer architecture (D-120) + +``` +┌─────────────────────────────────────────────────────────────────────┐ +│ Nova platform components (existing) │ +│ run_platform.sh · confidence_signal · checkov_adapter · │ +│ hitl_gates · regression_verify · outbox_writer · contract_ingestor │ +└──────────────────────┬──────────────────────────────────────────────┘ + │ CloudEvents 1.0 envelope (new emitters, P1) + ▼ +┌─────────────────────────────────────────────────────────────────────┐ +│ metrics/events.jsonl (append-only CloudEvents log) │ +│ metrics/runs/.json (per-run manifests) │ +│ metrics/decision_ledger.db (SQLite hash-chain, D-121) │ +│ metrics/test-results.xml (junit, P1) │ +└──────────────────────┬──────────────────────────────────────────────┘ + │ collector reads (P2) + ▼ +┌─────────────────────────────────────────────────────────────────────┐ +│ metrics/nova_metrics.db (SQLite cold store, D-126) │ +│ fact_run · fact_capability · fact_policy_check · fact_confidence │ +│ fact_test · fact_decision · fact_cost_estimate │ +│ dim_capability · dim_milestone │ +│ + 8 empty placeholder views (deferred metrics) │ +└──────────────────────┬──────────────────────────────────────────────┘ + │ powerbi_export (P3) + ▼ +┌─────────────────────────────────────────────────────────────────────┐ +│ metrics/powerbi/ (CSV/JSON views, folder connector, D-129) │ +│ → PowerBI dashboards (external) │ +└─────────────────────────────────────────────────────────────────────┘ +``` + +**Hot path: deferred (D-126).** No live ops dashboard; SQLite is +cold-only (batch/historical). The hot path activates when live AWS is +re-provisioned (D-096 lift). + +### NORTH_STAR integration point (REQ-186) + +`.ciagent/NORTH_STAR.md` is read by CIAgent in context-loading for all +future milestones. The integration mechanism (to be finalized in P4): +a reference from `PROJECT.md` + `ARCHITECTURE.md` (this section) + a +config entry in `config.json` (`strategic_direction_file: +".ciagent/NORTH_STAR.md"`) that the run workflow reads at SPECIFY. This +ensures the strategic direction survives across milestones without +being overwritten by status updates. diff --git a/.ciagent/CHECKPOINT.json b/.ciagent/CHECKPOINT.json index 3c60f0b..e8b82fd 100644 --- a/.ciagent/CHECKPOINT.json +++ b/.ciagent/CHECKPOINT.json @@ -1,13 +1,12 @@ { - "phase": 21, - "stage": "complete", - "milestone": "v1.16", - "phase_role": "final", + "phase": 0, + "stage": "grill", + "milestone": "v1.17", + "phase_role": "pre_execution", "attempts": 0, - "updated_at": "2026-07-30T16:05:00Z", - "milestone_complete": true, - "tag": "v1.15.26", - "release_id": 370, - "requirements": ["REQ-165", "REQ-166", "REQ-167", "REQ-168", "REQ-169", "REQ-170", "REQ-171", "REQ-172", "REQ-173", "REQ-174", "REQ-175", "REQ-176", "REQ-177", "REQ-178", "REQ-179", "REQ-180", "REQ-181", "REQ-182", "REQ-183", "REQ-184"], - "regression": {"Verified": 18, "Decayed": 0, "Broken": 0, "Skipped": 4} + "updated_at": "2026-08-04T21:15:00Z", + "milestone_complete": false, + "tag": null, + "requirements": ["REQ-185"], + "notes": "GRILL complete (interactive). 12 binding decisions applied: NORTH_STAR targets reclassified (E-003: 3 targets to Post-Pilot section; E-004: AI-Agent Intent Share to Future Horizons). Deck plan updated: slide 1 stake line (G-Q8), slide 4 benefit rewrite (G-Q9), slide 7 D-122 honesty sentence (G-Q4), Act 3->4 transition rewrite (G-Q13), slide 9 benefit reframe (G-Q14), slide 12 split into 12+13 (G-Q10), ROI formula inline + N=0 caveat (G-Q5/G-Q15), slide 14 preempt (G-Q11), slide 16 ask reframed as business decision (G-Q16). Deck now 16 main + 2 appendix = 18 slides." } \ No newline at end of file diff --git a/.ciagent/GRILL.md b/.ciagent/GRILL.md index 90ba90a..9e754c2 100644 --- a/.ciagent/GRILL.md +++ b/.ciagent/GRILL.md @@ -636,3 +636,262 @@ re-provision the bucket. YES, once G-111's criterion restatement + gate update are incorporated (into P9's must-haves). G-112/G-113 are phase-entry clarifications for P9/P12/P13. E-002 is deferred to P21. Confidence 0.85. + +--- + +# GRILL — v1.17 "Strategic Direction, Leadership Metrics & Unified Story" (2026-08-04) + +> **Griller:** CIAgent (red-team mode). **Milestone:** v1.17. **Axes:** 3 +> (NORTH_STAR alignment, Deck story & arc, Deck per-slide rigor) per PO +> direction. **Stance:** adversarial — presumed over-scoped / infeasible / +> storytelling-weak until evidence forced otherwise. + +## Evidence base + +- `NORTH_STAR.md` (183 lines, draft), `PLAN.md` (1,114 lines, deck rebuild + plan incl. slide-by-slide), `REQUIREMENTS.md` v1.17 (REQ-185..213), + `RESEARCH.md` v1.17 (signal inventory, scorecard, deferred-decision + ledger, deck research). +- Codebase cross-checks: `REGRESSION_REPORT.json` = **18 Verified + 4 + Skipped** (NOT "22/22 Verified" — the new deck plan correctly says + 18V+4S; the *existing* decks still claim 22/22). `PROJECT.md:495` = + **0 consumer adoption**. `docs/NO_HUMANS_THESIS.md`, `docs/METRICS.md`, + `metrics/` do not yet exist (P4/P5 deliverables — expected). +- Decisions locked (D-120..D-132) — not re-litigated. + +## The central contradiction + +**NORTH_STAR.md:111** states: *"Targets are committed, not aspirational."* +**PO's G-Q6 answer:** *"the goal is simply to target a high touchless +resolution rate, not to say we have reached those targets given there are +0 consumers."* + +These two statements are in direct conflict. "Committed, not aspirational" ++ "simply to target" = the document is lying about its own epistemic +status. This is the v1.10 decay root cause (PRE_MORTEM FM-3: decks +outrunning verified reality) repeating itself in the document meant to +prevent it. + +## Axis 1 — NORTH_STAR alignment + +### G-Q1 — Target with no backing REQ / placeholder +**Finding:** AI-Agent Intent Share (≥40%) is a committed 12–18mo target +(NORTH_STAR:128) with "placeholder view" claimed, but it is NOT among the +8 placeholder views in PLAN P3 (lines 309–315), and no REQ-185..213 builds +an emitter or placeholder for it. RESEARCH §3 marks it "future" with no +controlling decision ID (unlike every other deferred metric). NORTH_STAR:128 +falsely claims a placeholder view exists → violates the "no fabrication" +hard constraint. +**Verdict: BIND.** Add a 9th placeholder view OR move the target to a +"Future Horizons" section; correct NORTH_STAR:128. **Confidence: 0.90.** + +### G-Q2 — Anti-goal pursuit +**Finding:** No REQ builds an anti-goal. Deck title "No-Humans Infrastructure +Platform" is one weak slide away from violating anti-goal #3 (not removing +humans from accountability) — mitigation is entirely in slide 3's execution. +**Verdict: PASS (conditional on slide 3 landing the attestation model).** +**Confidence: 0.75.** + +### G-Q3 — Attestation clarification consistency +**Finding:** The attestation clarification is the most consistently +propagated concept in the plan — NORTH_STAR (3 places), REQUIREMENTS +(3 REQs), deck (3 slides). Well done. +**Verdict: PASS.** **Confidence: 0.92.** + +### G-Q4 — "AI decision" framing (D-122 honesty) +**Finding:** D-122 (confidence_signal + HITL gate, NOT an LLM) is cited on +slide 7 and required in NO_HUMANS_THESIS.md (REQ-213). BUT slide 7's +*Delivers* says "every AI decision captured" without ever telling the +audience what the "AI" is. The honesty is buried in a linked doc + a +decision ID the audience has never heard. +**Verdict: BIND.** Add one sentence to slide 7 *Delivers*: "Nova's 'AI +decision' is the confidence-gated policy engine, not an LLM planner +(D-122)." **Confidence: 0.85.** + +### G-Q5 — Secretly ungrounded metrics +**Finding:** The 8 deferred placeholder views cover their list. BUT (a) +AI-Agent Intent Share's placeholder is falsely claimed (G-Q1), and (b) +derived metrics (FTE Hours Saved, Platform ROI) are computed on zero +production runs yet shown on slide 12 without the zero-denominator caveat. +A "derived" metric from zero runs is technically not fabricated but is +misleading. +**Verdict: BIND.** (1) Resolve G-Q1; (2) slide 12 must annotate derived +metrics with "(computed on N internal runs; production-denominator activates +post-pilot)." **Confidence: 0.82.** + +### G-Q6 — 12–18mo target feasibility (0 consumers) +**Finding:** PO's answer ("simply to target") conflicts with NORTH_STAR:111 +("committed, not aspirational"). 3 "grounded (after P1)" targets (Touchless +Resolution, Human Escalation, AI Decision Accuracy) have scope "across +production estates" — but PROJECT.md:495 = 0 consumer adoption. The metric +IS computable on internal dev runs, but the target scope doesn't exist. +Marking "grounded" while the scope is absent is the overclaim the "no +fabrication" constraint exists to prevent. +**Verdict: BIND.** (1) Rewrite NORTH_STAR:111 → "Targets are committed +destinations; the grounding column records whether each is measurable this +milestone." (2) Reclassify the 3 targets to `partial — measurement pipeline +grounded on internal runs; production-estate scope activates post-pilot` +(the Cloud Spend Reduction precedent at NORTH_STAR:123). (3) Deck slide 5 +regroup as "Measurable today (internal runs)" vs "Activates post-pilot +(production estates)." Requires NORTH_STAR-CHANGE commit trailer (REQ-204). +**Confidence: 0.80.** + +## Axis 2 — Deck plan: story & arc + +### G-Q7 — Arc order (Problem→Vision→How→Proof→Roadmap vs Proof-first) +**Finding:** Current arc puts Proof at Act 4 (slides 10–13) — 40% of the +deck before a number. For a leadership audience that has seen 10+ milestone +decks, this risks losing the room by slide 4. BUT the "no-humans" thesis +is contentious; jumping to proof without the attestation model invites the +"removing humans from accountability" objection. The Vision act makes the +Proof credible. +**Verdict: PASS (marginal).** Defensible IF the Problem act is tight and +slide 3 front-loads the attestation clarification. **Confidence: 0.62.** + +### G-Q8 — x3 structure at deck level +**Finding:** Slide 1's 5-act preview is orienting (a table of contents), +not too much meta-structure. BUT it's also not a hook — it gives structure, +not stakes. A C-suite audience decides in the first 30 seconds. +**Verdict: BIND (minor).** Add one stake-establishing line to slide 1 +*Delivers* with a real number (18 verified, 0 consumers, honest deferral +list). **Confidence: 0.70.** + +### G-Q9 — Per-slide benefit callouts (substantive vs filler) +**Finding:** 4 of 17 closes are filler (slides 1, 4, 12, 15); 2 borderline +(8, A1). Worst offender: slide 12 (ROI) restates the *objective* ("ROI is +quantifiable") rather than giving the *number* or the *honest caveat*. +**Verdict: BIND.** Rewrite 4 filler closes. Slide 12's close must be: +"Benefit: you now know the ROI formula — (labor + cloud + avoided downtime) +÷ platform cost — and that it computes on internal runs today, with +production-denominator activating post-pilot." **Confidence: 0.78.** + +### G-Q10 — Deck length (17 slides) +**Finding:** 17 is at the upper bound but justifiable for 5 acts. The risk +is density, not length: slide 12 crams 6 metrics (Touchless, Human +Escalation, MTTR, Cost, FTE, ROI) into one slide — a wall of bullets. +**Verdict: BIND (minor).** Split slide 12 into "Zero-Touch Efficiency" +(Touchless, Human Escalation, MTTR) + "Cost & ROI" (Cost, FTE, ROI). Deck +→ 18 slides, each earning its place. **Confidence: 0.68.** + +### G-Q11 — "What's Deferred" slide (13) +**Finding:** The honesty strengthens the grounded claims BUT surfaces the +gap: Nova claims "no-humans in operations" while deferring the metrics +that would prove operations are healthy without humans (Live Infra Health, +SLA, Drift Auto-Reversal). A skeptical viewer notes the contradiction. +**Verdict: BIND.** Add preempt to slide 13: "These deferrals are about +*measurement infrastructure*, not about whether the platform runs without +humans — the platform runs autonomously today on internal runs; what's +deferred is the production-estate dashboard that would prove it at scale." +**Confidence: 0.75.** + +## Axis 3 — Deck plan: per-slide rigor + +### G-Q12 — Slide opening lines +**Finding:** The "This slide shows X" formula is orienting, not patronizing, +because each includes a stake-bearing clause. Consistent without being empty. +**Verdict: PASS.** **Confidence: 0.80.** + +### G-Q13 — Transitions (written vs hand-waved) +**Finding:** ~10 of 13 transitions are written (specific reference to prior +close). 3 are hand-waved (slides 8→9, 11→12, 13→14). Worst: the Act 3→4 +boundary (slide 8→9, How→Proof) — the most important transition in the deck +— is the weakest. +**Verdict: BIND.** Rewrite the 3 hand-waved transitions. The 8→9 Act +boundary must carry weight: "Having seen the gate model — autonomy in +operations, human in accountability — here is how Nova instruments itself +to prove that model at scale." **Confidence: 0.85.** + +### G-Q14 — Weakest slide (audience-loss point) +**Finding:** Slide 9 (Telemetry Architecture) is the audience-loss slide. +It's the 4th consecutive architecture slide (6,7,8,9), the most abstract +(CloudEvents, SQLite, PowerBI), its Benefit is about data plumbing not +business value, and it sits between the attestation matrix (slide 8, +emotionally resonant) and the Proof act (slide 10, the numbers) — between +the two things the audience came for. +**Verdict: BIND.** Compress slide 9 into slide 10 OR reframe its Benefit +from data plumbing to trust: "Benefit: you now know the proof you're about +to see isn't fabricated — every number traces to a file you can audit." +**Confidence: 0.78.** + +### G-Q15 — Proof act citation specificity +**Finding:** 5 of 6 Proof citations are specific (file paths + real numbers). +Gap: slide 12's derived metrics (FTE, ROI) cite "derived" without showing +the formula or the input count. +**Verdict: BIND (minor).** Show the ROI formula inline on slide 12 + the +N=0 production-runs caveat. **Confidence: 0.80.** + +### G-Q16 — Closing slide (15) — does the ask land? +**Finding:** THE ask is present but framed as insider language ("fund the +hot-path activation (post-D-096) + the tamper-evident ledger build-out +(D-083 lift)"). A leadership audience doesn't know what "hot-path +activation" means. The ask is a technical request, not a business decision +a leader can make in the room. +**Verdict: BIND.** Reframe slide 15's ask as a business decision: "The +ask: (1) approve a pilot estate to activate production-estate metrics +(unblocks D-096), and (2) approve the tamper-evident ledger build-out +(lifts D-083) — turning grounded claims into complete proof." Make it a +yes/no a leader can give. **Confidence: 0.82.** + +## Binding decisions (must resolve before SHIP) + +| G-ID | Axis | Verdict | What must change | Conf | +|---|---|---|---|---| +| G-Q1 | 1 | BIND | Add 9th placeholder view for AI-Agent Intent Share OR move to "Future Horizons"; correct NORTH_STAR:128 | 0.90 | +| G-Q4 | 1 | BIND | Add D-122 honesty sentence to slide 7 *Delivers* | 0.85 | +| G-Q5 | 1 | BIND | Annotate derived metrics on slide 12 with zero-run caveat | 0.82 | +| G-Q6 | 1 | BIND | Rewrite NORTH_STAR:111; reclassify 3 targets to `partial`; regroup deck slide 5. NORTH_STAR-CHANGE trailer required | 0.80 | +| G-Q8 | 2 | BIND (minor) | Add stake line with real number to slide 1 *Delivers* | 0.70 | +| G-Q9 | 2 | BIND | Rewrite 4 filler closes (slides 1, 4, 12, 15); slide 12 must give ROI formula + caveat | 0.78 | +| G-Q10 | 2 | BIND (minor) | Split slide 12 into two (Efficiency + Cost/ROI); deck → 18 slides | 0.68 | +| G-Q11 | 2 | BIND | Add preempt to slide 13 (deferrals are measurement infra, not whether platform runs without humans) | 0.75 | +| G-Q13 | 3 | BIND | Rewrite 3 hand-waved transitions (esp. Act 3→4 boundary 8→9) | 0.85 | +| G-Q14 | 3 | BIND | Compress slide 9 into slide 10 OR reframe its Benefit to trust | 0.78 | +| G-Q15 | 3 | BIND (minor) | Show ROI formula inline + N=0 caveat on slide 12 | 0.80 | +| G-Q16 | 3 | BIND | Reframe slide 15 ask as a business decision (pilot estate + ledger build-out) | 0.82 | + +**PASS (no change):** G-Q2 (anti-goals, conditional on slide 3), G-Q3 +(attestation consistency — excellent), G-Q7 (arc order — marginal), +G-Q12 (slide openings — formulaic but substantive). + +## Escalations (only the PO can decide) + +| E-ID | Question | Confidence | +|---|---|---| +| E-003 | Should the 3 "grounded (after P1)" targets with "production estates" scope be reclassified to `partial` (Cloud Spend precedent), or should "grounded" be redefined to mean "measurement pipeline grounded"? Changes a committed NORTH_STAR target's grounding label; requires NORTH_STAR-CHANGE trailer (REQ-204). | <0.60 | +| E-004 | Should AI-Agent Intent Share (≥40%) remain a "12–18mo Target" with no backing REQ/placeholder, or move to a "Future Horizons" section? Strategic-scope question (is agentic consumption a 12–18mo commitment or a longer horizon?). | <0.60 | + +## Overall verdict + +**🟡 REDUCE SCOPE / BINDING FIXES REQUIRED — not ready to ship as-is.** + +The plan is architecturally sound (metrics pipeline, Decision Ledger, +PowerBI export, x3 deck structure are well-designed and grounded). The +attestation clarification (G-Q3) is the best-propagated concept in the +plan. The regression-capability gate (CAP-023/024) is a credible safeguard. + +But the plan has one structural contradiction (NORTH_STAR:111 vs PO intent +vs grounding column) that infects 4 other findings (G-Q1, G-Q5, G-Q6, +G-Q9/slide 12). This is the v1.10 decay pattern (PRE_MORTEM FM-3) +repeating in the document meant to prevent it. The "no fabrication" hard +constraint is self-violated in two places (AI-Agent Intent Share placeholder +claim, derived-metrics-without-caveat) before a single slide is rendered. + +The deck plan is story-competent but not story-excellent. 4 benefit +callouts are filler, 3 transitions are hand-waved (incl. the critical +Act 3→4 boundary), slide 9 is the audience-loss slide, and the closing +ask is insider language. + +**12 binding decisions, 2 escalations.** None require re-architecting the +plan; all are edits to NORTH_STAR (2 rows + 1 line, with commit trailer), +the deck slide plan (4 slide rewrites, 1 split, 3 transition rewrites), +and one placeholder-view addition. Estimate: 1–2 phases of rework, not a +milestone restart. The plan does NOT need a revision loop — it needs +these 12 fixes applied in P0 (NORTH_STAR) and P5 (deck) before the +respective phases ship. Critical path unchanged. + +**Can the milestone proceed?** + +YES, once the 12 BIND decisions are incorporated (G-Q1/Q4/Q5/Q6 into P0 +NORTH_STAR + P5 deck plan; G-Q8/Q9/Q10/Q11/Q13/Q14/Q15/Q16 into P5 deck +plan). E-003/E-004 require PO decisions on NORTH_STAR target framing. +Confidence 0.80. diff --git a/.ciagent/NORTH_STAR.md b/.ciagent/NORTH_STAR.md new file mode 100644 index 0000000..de0f4f6 --- /dev/null +++ b/.ciagent/NORTH_STAR.md @@ -0,0 +1,211 @@ +# NORTH_STAR — Nova + +> **Status:** Draft (pending interactive GRILL → final) +> **Milestone:** v1.17 — Strategic Direction, Leadership Metrics & Unified Story +> **Owner:** Product Owner +> **Purpose:** Durable strategic intent. Read by CIAgent in every future +> `/ci-run` so the platform's direction survives across milestones. This +> is NOT a status document (that's PROJECT.md) and NOT an engineering +> architecture (that's the telemetry reference in RESEARCH.md/ +> ARCHITECTURE.md). It is the PO's committed direction: what we're +> building toward, what we refuse to build, and how we'll know we won. + +--- + +## Vision + +> **Infrastructure operations become invisible. Every environment +> provisioned, every incident healed, every risk remediated — by an +> autonomous system whose trustworthiness is provable, not promised. +> Human attestation remains required at stage gates — QA signs off for +> production, SRE greenlights based on operational readiness — but the +> operator is never in the loop of normal operations.** + +Nova is the autonomous infrastructure layer that lets product teams ship +without engaging an operator, and lets executives trust the AI not because +it never fails but because every decision is captured, scored, and +accountable. + +--- + +## Strategic Objectives (4) + +**1. Demonstrate production-grade zero-touch operations.** +Nova must run real customer estates with no human in the loop of normal +operations — autonomy as the default, not the demo. Stage-gate +attestation (QA for production, SRE for operational readiness) remains +human by design; operational escalations (AI confidence too low to +proceed) are the failure mode we drive toward zero. Everything else +collapses if autonomy isn't real. + +**2. Establish provable trust in AI decisions.** +Build the audit substrate — Decision Ledger, confidence scoring, circuit +breakers, blast-radius controls — that turns "autonomous" from a +marketing claim into a defensible one. Trust is the moat. Features can be +copied; an immutable, queryable decision history cannot. + +**3. Deliver compounding, quantifiable ROI for customers.** +Each quarter on Nova must reduce cloud spend, free engineering hours, and +avoid downtime measurably. If the CFO can't point to a number that +improves quarter-over-quarter, Nova fails its commercial test, regardless +of how clever the AI is. + +**4. Become the default substrate for agentic infrastructure consumption.** +AI agents are already becoming the largest consumers of cloud +infrastructure. Nova must be the platform through which those agents +declare, deploy, and verify infrastructure — not a vendor scrambling into +that market two quarters late. + +--- + +## Anti-Goals (5 — what Nova is fundamentally NOT) + +1. **Not a Terraform, Kubernetes, or hyperscaler competitor.** We + orchestrate them. Replacing them is the most expensive possible + distraction from the value we create. +2. **Not a general-purpose AI agent platform.** We are purpose-built for + infrastructure operations. Breadth here produces shallow tools; depth + here wins the category. +3. **Not a system that removes humans from accountability.** Only from + operations. Every AI decision lands in an immutable ledger. Every + stage-gate promotion (qa/prod/dr) requires a human attestation recorded + with approver identity, separation-of-duties check, and the 8-concern + evidence matrix. The absence of an operator is never the absence of a + record. +4. **Not for legacy, untagged, or freeform infrastructure.** Nova requires + Terraform-managed, policy-aligned, fully-tagged inputs. We optimize for + the disciplined 95%, not the chaotic 5%. +5. **Not sold to operators.** Nova is sold to leadership on outcomes — + cost, velocity, risk. Selling to operators inverts the incentive and + breaks the autonomy thesis. + +--- + +## Non-Goals (v1.17 milestone scope — deferred work, not permanent boundaries) + +> Anti-Goals are what Nova *fundamentally is not*. Non-Goals are what we +> *will not do this milestone* — deferred work, not permanent boundaries. +> Each Non-Goal cites the controlling decision ID. + +1. **Live AWS re-provisioning** (deferred — D-096). Metrics that require + live infrastructure ship as placeholder PowerBI views with documented + schemas. +2. **Onboarding auto-grant** (deferred — D-113/D-114/D-119). Only the + request-path metric is grounded; the requested→granted funnel is a + placeholder. +3. **ML anomaly-forecasting / predictive remediation** (no emitter today). + The Predictive-vs-Reactive metric ships as a placeholder. +4. **Drift detection scheduled job** (deferred — D-096 + no scheduler). + Drift metrics ship as placeholders. +5. **Live cost CUR reconciliation** (deferred — D-096). Pre-apply Infracost + estimates are grounded; actual-spend reconciliation is a placeholder. +6. **S3 Object Lock / JWS tamper-evident ledger** (deferred — D-083). The + Decision Ledger uses a local SQLite hash-chain this milestone; the + Object-Lock/JWS build-out is a future milestone. +7. **Multi-cloud support** (Azure/GCP/K8s). Nova is AWS-only this milestone. + +--- + +## 12–18 Month Targets + +Targets are committed, not aspirational. Each is a number a board member +can repeat back to us. The grounding column records whether the metric is +measurable this milestone, and if not, what blocks it. + +> **Honesty note (GRILL G-Q6 binding):** Nova has 0 consumer adoption +> today (`PROJECT.md:495`). Three targets (Touchless Resolution, Human +> Escalation, AI Decision Accuracy) are scoped "across production +> estates" — the measurement *pipeline* is grounded this milestone, but +> the *denominator* is zero until a pilot estate activates. These +> targets are reclassified as **Post-Pilot** (the pipeline works; the +> numbers fill when consumers exist). This is the same honesty model as +> Cloud Spend Reduction (partial: pipeline grounded, actuals deferred). + +### Current-milestone targets (grounded or derived this milestone) + +| Domain | Target | Grounding (v1.17) | Note | +|---|---|---|---| +| **MTTR (p95)** | < 60 seconds | grounded (platform-run MTTR) | apply.failed → successful retry; infra-incident MTTR deferred (no incident detection) | +| **Cloud Spend Reduction** | ≥ 25% on pilot estates vs. 12-month pre-Nova baseline | partial | pre-apply estimate grounded (Infracost); actual-spend deferred (D-096 CUR) | +| **L1 / L2 Ops Hours Avoided** | ≥ 70% of pre-Nova FTE allocation | derived | formula over run count × manual baseline (computed on N internal runs; production-denominator activates post-pilot) | +| **Platform ROI** | ≥ 250% measured annually | derived | formula (labor savings + cloud savings + avoided downtime) ÷ platform op cost (computed on N internal runs; production-denominator activates post-pilot) | +| **Decision Ledger Coverage** | 100% of AI actions with backfilled outcome | grounded (this milestone builds it) | outbox_writer.py → SQLite hash-chain | +| **Attestation Coverage** | 100% of prod/dr promotions attested by a human | grounded | hitl_gates.py + outbox approver_* attributes; separation-of-duties on prod | + +### Post-Pilot targets (pipeline grounded this milestone; denominator activates when a pilot estate runs) + +| Domain | Target | Grounding (v1.17) | Note | +|---|---|---|---| +| **Touchless Resolution Rate** | ≥ 99% across production estates | partial (pipeline grounded; denominator = 0 today) | runs completing without *operational* HITL block ÷ total runs (attestation gates excluded); activates post-pilot | +| **Human Escalation Frequency** | < 0.1% of platform actions | partial (pipeline grounded; denominator = 0 today) | *operational* HITL blocks only (confidence-driven); attestation sign-offs excluded; activates post-pilot | +| **AI Decision Accuracy** | ≥ 99.5% (no rollback, no follow-up incident within 5 min of action) | partial (pipeline grounded; denominator = 0 today) | decisions not followed by apply.failed/incident within 5min; activates post-pilot | + +### Deferred targets (measurement requires future systems) + +| Domain | Target | Grounding (v1.17) | Note | +|---|---|---|---| +| **Predictive vs. Reactive Ratio** | ≥ 3 : 1 (prevention dominates reaction) | deferred | requires ML forecasting service (future emitter) | +| **Drift Auto-Reversal Rate** | ≥ 95% within one detection cycle | deferred | requires drift detection (D-096 + scheduler) | + +> Committed targets whose measurement is deferred remain committed — the +> target is the destination; the metric is the odometer, and some +> odometers aren't built yet. Each deferred metric ships as a placeholder +> PowerBI view + a definition-of-success doc recording the dependency. +> Post-Pilot targets are committed targets whose measurement pipeline is +> grounded this milestone; the numbers activate when a pilot estate runs. + +### Future Horizons (strategic direction, not committed targets) + +| Domain | Aspiration | Note | +|---|---|---| +| **AI-Agent Intent Share** | ≥ 40% of total intent volume originated by non-human consumers | Strategic Objective #4 direction. No backing requirement, no placeholder view, no emitter today. Moves to a committed target when agentic consumption is real. | + +--- + +## Success Criteria (v1.17 — what constitutes success for THIS milestone) + +> Distinct from the 12–18mo targets: those are the destination. These are +> the milestone's exit criteria. + +v1.17 is a success if: + +1. **Decision Ledger emits `ai.decision.made` for 100% of platform runs** + with outcome backfill, AND **`attestation.recorded` events for 100% + of qa/prod/dr promotions** (event completeness — all 3 gates captured; + grounded in `outbox_writer.py` → SQLite hash-chain; honors D-083). + The **Attestation Coverage metric** (target 100%) measures prod/dr + promotions specifically — see REQ-194. +2. **`docs/METRICS.md` catalogs every executive KPI** with a `grounded` / + `derived` / `deferred` status, a source file or decision ID, and a + per-KPI definition-of-success doc in `docs/metrics/`. +3. **The PowerBI export produces all fact/dimension views** + 8 empty + placeholder views for deferred metrics (with documented schemas ready + to fill when their blocking decisions lift). +4. **The unified narrative deck ships** with the x3 arc + (Problem→Vision→How→Proof→Roadmap) at deck + slide level, per-slide + benefit callouts, and fluid transitions; both old decks retired. +5. **`NORTH_STAR.md` is wired into CIAgent context-loading** so every + future `/ci-run` reads it. +6. **CAP-023 (metrics collector) + CAP-024 (deck structure) pass** in the + regression gate. + +--- + +## What "won" looks like + +By month 18, Nova is the layer enterprise leadership points to when they +say *"we don't have an infrastructure ops team anymore, and the audit +trail is stronger than it ever was"* — and it is the default substrate +their AI engineering teams reach for first when an agent needs to deploy. + +--- + +## Relationship to v1.17 engineering + +- **Pillar A (this file):** strategic direction — durable, PO-authored. +- **Pillar B (engineering):** the telemetry reference architecture + (adapted from the PO's technical-direction input) lives in + RESEARCH.md/ARCHITECTURE.md. It is the *how*; this file is the *why*. +- **Pillar C (story):** the unified narrative deck proves Pillars A+B to + leadership. The deck's Proof section cites grounded metrics; its + Roadmap section cites deferred targets honestly. \ No newline at end of file diff --git a/.ciagent/PERSONAS.md b/.ciagent/PERSONAS.md index 9d96375..3793226 100644 --- a/.ciagent/PERSONAS.md +++ b/.ciagent/PERSONAS.md @@ -1,23 +1,24 @@ --- project: acdl -milestone: v1.16 -generated_at: 2026-07-30 +milestone: v1.17 +generated_at: 2026-08-04 generator: lead-developer verification_toolchain: typecheck: "terraform validate && python3 -m py_compile core/**/*.py && python3 -m jsonschema schemas/*.schema.json" - test: "bash scripts/run_regression.sh # 22-capability gate (D-091/D-118)" + test: "bash scripts/run_regression.sh # 22-capability gate (D-091/D-118) + CAP-023/024 (v1.17)" build: "bash scripts/run_ci.sh # full local CI reproduction (lint+test+check-only)" note: | - Nova (formerly ACDL) has no package.json. The execute/verify/ship - workflows substitute `terraform validate` + `python -m py_compile` + - JSON Schema validation for npm run typecheck, the regression gate - (D-091, 22 capabilities) for npm test, and `bash scripts/run_ci.sh` - for npm run build. v1.11 testing is pipeline-driven (D-102); - v1.16 is NFR-only (no live apply by default; NOVA_LIFECYCLE_MODE= - plan). Roster carries forward from v1.11/v1.14/v1.15 unchanged. - frontend-engineer stays inactive (no frontend; decks are markdown = - lead-developer territory). No custom personas needed (no new - domains — onboarding is backend-engineer + data-engineer territory). + v1.17 adds a telemetry/observability layer (metrics emitters, SQLite + cold store, PowerBI export, Decision Ledger) + a unified narrative + deck + a durable NORTH_STAR.md. Three active personas: lead-developer + (coordination + deck narrative co-author), backend-engineer (event + emitters, outbox_writer extension, Infracost adapter), data-engineer + (SQLite store, schemas, PowerBI views, metrics collector). frontend- + engineer stays deactivated (no Nova web UI — dashboards are PowerBI, + not a Nova-built frontend; decks are markdown = lead-developer + territory). No new custom personas needed — the metrics domain maps + cleanly to data-engineer (schema/store/export) + backend-engineer + (emitters/instrumentation). --- # ACDL — Persona Roster (project-level, v1.11 RESTART) @@ -251,3 +252,134 @@ The regression gate (22 capabilities) must stay **22/22 Verified** throughout v1.16 — simplification must not regress any capability (D-118). P9 (end of Wave 2) and P21 (milestone complete) run the gate; P14 (end of Wave 3) is an offline mid-milestone checkpoint. + +--- + +# v1.17 Persona Roster — Strategic Direction, Leadership Metrics & Unified Story + +> v1.17 adds a telemetry/observability layer (P1–P3), a metrics catalog +> + NORTH_STAR integration (P4), a unified narrative deck (P5), a +> regression capability (P6), and a final review/ship (P7). Three +> active personas; frontend-engineer stays deactivated (no Nova web UI +> — dashboards are PowerBI, not a Nova-built frontend). + +## Active personas + +### lead-developer +- **Domain:** coordination + deck narrative +- **Active:** true +- **Phase-specific:** false +- **Reason:** Owns CIAgent metadata, the NORTH_STAR.md authoring + process (P0), the milestone decomposition, the unified narrative deck + co-authoring (P5 — the deck is markdown, which is lead-developer + territory per the established convention), and the final review/ship + (P7). Arbitrates persona conflicts (e.g., backend vs data on the + emitter/store boundary). +- **Territory:** `.ciagent/NORTH_STAR.md`, `.ciagent/PROJECT.md`, + `.ciagent/REQUIREMENTS.md`, `.ciagent/PLAN.md`, `.ciagent/RESEARCH.md`, + `.ciagent/ARCHITECTURE.md`, `docs/presentations/nova-no-humans-platform.md` + (NEW — unified deck source of truth), `docs/presentations/nova-no-humans-platform-marp.md`, + `docs/presentations/nova-no-humans-platform-talking-points.md`, + `docs/METRICS.md`, `docs/metrics/*.md` (per-KPI definition docs). + +### backend-engineer +- **Domain:** backend (event emitters + instrumentation) +- **Active:** true +- **Phase-specific:** false +- **Reason:** Owns the event emitters (P1): the CloudEvents envelope, + the per-run manifest writer, the `outbox_writer.py` extension to the + SQLite Decision Ledger, the Infracost post-processor, the + `hitl_gates.py` attestation event emission, the `confidence_signal.py` + decision event emission, the `checkov_adapter.py` policy event + emission, and the pytest `--junitxml` addopts change. Also owns the + `regression_verify.py` CAP-023/024 additions (P6). The emitter work + is the bridge between existing Nova components and the new metrics + layer — it touches the code paths that already exist. +- **Territory:** `core/metrics/event_envelope.py` (NEW), + `core/metrics/run_manifest.py` (NEW), + `core/metrics/infracost_adapter.py` (NEW), + `core/metrics/decision_ledger.py` (NEW — extends outbox_writer), + `core/outbox_writer.py` (extend to SQLite), + `core/hitl_gates.py` (emit attestation.recorded), + `core/confidence_signal.py` (emit ai.decision.made), + `adapters/terraform/policy/checkov_adapter.py` (emit policy.evaluated), + `scripts/run_platform.sh` (invoke manifest writer + Infracost), + `core/regression_verify.py` (CAP-023/024), + `pyproject.toml` (addopts --junitxml), + `tests/test_metrics_emitters.py` (NEW), + `tests/test_decision_ledger.py` (NEW). + +### data-engineer +- **Domain:** data (schema, SQLite store, PowerBI export) +- **Active:** true +- **Phase-specific:** false +- **Reason:** Reactivated with a new territory for v1.17: the metrics + collector (P2) and the PowerBI export (P3). Owns the schema design + (metrics_*.schema.json), the SQLite cold store (nova_metrics.db), the + fact/dimension table design, the 8 deferred placeholder views, and + the CSV/JSON export. The data-engineer's schema-first constraint + applies: all event types and fact/dim tables have JSON Schema + definitions before any code is written. The collector reads files + + events → SQLite; the export reads SQLite → CSV/JSON. This is the + heaviest data-territory work since v1.11's terraform modules. +- **Territory:** `core/metrics/collector.py` (NEW), + `core/metrics/powerbi_export.py` (NEW), + `schemas/metrics_*.schema.json` (NEW — event + fact/dim schemas), + `metrics/nova_metrics.db` (NEW — SQLite cold store), + `metrics/powerbi/` (NEW — CSV/JSON export dir), + `docs/METRICS_VIEWS.md` (NEW — schema doc for PowerBI views), + `tests/test_metrics_collector.py` (NEW), + `tests/test_powerbi_export.py` (NEW). + +## Deactivated personas + +### frontend-engineer +- **Domain:** frontend +- **Active:** false +- **Phase-specific:** false +- **Reason:** v1.17 has no Nova web UI. The leadership dashboards are + PowerBI (an external tool that ingests CSV/JSON files), not a + Nova-built frontend. The decks are markdown (lead-developer + territory). frontend-engineer stays deactivated, consistent with + v1.11–v1.16. Reactivates if a future milestone builds a Nova web UI. + +### lambda-engineer, platform-engineer, security-engineer +- **Active:** false (carried forward from v1.11) +- **Reason:** v1.17 does not touch the Lambda (beyond emitting events + from the existing hitl_gates/attestation_matrix), does not do IR- + shaped module authoring, and does not touch security adapters beyond + emitting policy.evaluated events. The existing components are + instrumented, not rewritten. + +## v1.17 phase assignment + +| Phase | Primary persona | Supporting | Territory | +|-------|----------------|------------|-----------| +| P0 pre-execution | lead-developer | — | `.ciagent/NORTH_STAR.md`, `PROJECT.md`, `REQUIREMENTS.md`, `RESEARCH.md`, `ARCHITECTURE.md`, `PERSONAS.md`, `PLAN.md` | +| P1 event-emitters | backend-engineer | data-engineer (schemas) | `core/metrics/event_envelope.py`, `run_manifest.py`, `decision_ledger.py`, `infracost_adapter.py`, `outbox_writer.py`, `hitl_gates.py`, `confidence_signal.py`, `checkov_adapter.py`, `run_platform.sh`, `pyproject.toml` | +| P2 metrics-collector | data-engineer | backend-engineer (event formats) | `core/metrics/collector.py`, `schemas/metrics_*.schema.json`, `metrics/nova_metrics.db` | +| P3 powerbi-export | data-engineer | — | `core/metrics/powerbi_export.py`, `metrics/powerbi/`, `docs/METRICS_VIEWS.md` | +| P4 metrics-catalog + north-star-integration | lead-developer | data-engineer (metric definitions) | `docs/METRICS.md`, `docs/metrics/*.md`, `PROJECT.md`, `ARCHITECTURE.md`, `config.json` | +| P5 deck-rebuild | lead-developer | — | `docs/presentations/nova-no-humans-platform*.md`, retire old decks | +| P6 regression-capability | backend-engineer | data-engineer (CAP-023 schema) | `core/regression_verify.py` (CAP-023, CAP-024) | +| P7 final-review-ship | lead-developer | all active (review) | `.ciagent/**`, review + audit + ship | + +## v1.17 domain priority + +`backend → data → lead` (the emitter work in P1 is the foundation; +data-engineer's collector + export in P2–P3 depends on P1's event +formats; lead-developer's catalog + deck in P4–P5 depends on the +metrics being grounded). + +## v1.17 verification toolchain + +``` +typecheck: terraform validate && python3 -m py_compile core/**/*.py adapters/**/*.py +test: bash scripts/run_regression.sh # 22-capability gate + CAP-023/024 (v1.17) +build: bash scripts/run_ci.sh # full local CI reproduction +``` + +The regression gate (22 capabilities + CAP-023 metrics collector + +CAP-024 deck structure) must pass at P6 and P7. CAP-009 (offline pytest +suite) must remain Verified after the `--junitxml` addopts change +(assumption A5). diff --git a/.ciagent/PLAN.md b/.ciagent/PLAN.md index f602e44..efe65e7 100644 --- a/.ciagent/PLAN.md +++ b/.ciagent/PLAN.md @@ -1,420 +1,1182 @@ --- phase: P0 name: pre-execution -milestone: v1.16 -requirements: [REQ-165, REQ-166, REQ-167, REQ-168, REQ-169, REQ-170, REQ-171, REQ-172, REQ-173, REQ-174, REQ-175, REQ-176, REQ-177, REQ-178, REQ-179, REQ-180, REQ-181, REQ-182, REQ-183, REQ-184] +milestone: v1.17 +requirements: [REQ-185, REQ-186, REQ-187, REQ-188, REQ-189, REQ-190, REQ-191, REQ-192, REQ-193, REQ-194, REQ-195, REQ-196, REQ-197, REQ-198, REQ-199, REQ-200, REQ-201, REQ-202, REQ-203, REQ-204, REQ-205, REQ-206, REQ-207, REQ-208, REQ-209, REQ-210, REQ-211, REQ-212, REQ-213] wave: 0 depends_on: [] --- -# v1.16 — Nova Simplification Plan (20 execution phases + 1 final) +# v1.17 — Strategic Direction, Leadership Metrics & Unified Story (Plan) -**Milestone:** v1.16 (Nova Simplification — NFR) -**Type:** NFR (all phases fix/chore/docs/refactor/test). The final -phase's patch IS the deliverable — no separate milestone tag. Tags run -on the v1.15.x line: `v1.15.5` (P0) → `v1.15.6..v1.15.25` (P1–P20) → -`v1.15.26` (P21 final = milestone release). +**Milestone:** v1.17 — Strategic Direction, Leadership Metrics & Unified Story +**Type:** Feature (P1–P3 feat; P4 docs; P5 docs+test; P6 test; P7 review+audit+ship; +P8 final). Progressive patches; the final phase's patch IS the milestone +release. Tags run on the v1.16.x line: `v1.16.0` (P0) → `v1.16.1..v1.16.7` +(P1–P7) → `v1.16.8` (P8 final = milestone release). +**Branch:** `milestone/v1.17-strategic-metrics-deck` (branched off the v1.16 +complete merge). Execution phases branch `phase/NN-*` → merge to milestone +branch → tag patch on the v1.16.x line. +**Tags:** `metrics`, `telemetry`, `decision-ledger`, `powerbi`, `deck`, +`north-star`, `no-humans-thesis`, `regression-capability` +**Decisions (locked, D-120..D-132 — do NOT re-open):** +D-120 Nova-native + Infracost, drift deferred · D-121 Decision Ledger = +outbox_writer → SQLite hash-chain · D-122 AI decision = confidence_signal + +HITL gate · D-123 8 deferred metrics ship as empty placeholder views · +D-125 Hybrid events/files · D-126 Cold-only SQLite · D-127 Per-KPI +definition-of-success docs · D-128 metrics/ at repo root · D-129 PowerBI = +CSV/JSON folder connector · D-130 Deck arc Problem→Vision→How→Proof→Roadmap, +both old decks retired · D-131 MTTR = platform-run only · D-132 Attestation +instrumentation = emit attestation.recorded events. -**Objective:** A 20-phase NFR sweep (no new features) themed around five -user-directed axes: Simplify without regressions, Security, -Maintainability, User/Developer Experience, No Humans Onboarding Flow. -Clears the fresh debt the v1.15 rebrand left, delivers genuine -simplification, and implements the first self-service onboarding -request path (request-path only; real AWS provisioning deferred, D-113). +**Objective (three pillars):** +- **(A) Strategic Direction** — encode the PO's strategic direction in a + durable `NORTH_STAR.md` read by CIAgent in every future `/ci-run`. +- **(B) Leadership Metrics + PowerBI** — instrument Nova to collect, + aggregate, and surface leadership-grade metrics that prove the "no-humans" + autonomous-infrastructure value proposition — grounded in signals Nova + actually emits, derived via documented formulas, or explicitly deferred + with a decision ID — flowing into PowerBI-ready views. +- **(C) Unified Narrative Deck** — merge the two existing decks into one + unified narrative deck with the "tell them x3" arc at deck + slide level, + per-slide benefit callouts, and fluid transitions. -## Wave ordering +**Hard constraint:** DO NOT make anything up. Every metric carries a +`grounded` / `derived` / `deferred` status with a source file or decision +ID. Deferred metrics ship as empty PowerBI placeholder views with +documented schemas. -- **Wave 1 (P1–P4): correctness + brand regression fixes.** P1 first — - the state-bucket drift (`adapter.py:117` emits `acdl-tfstate-*` while - the live bucket is `nova-tfstate-*`) and the Kyverno policy - contradiction (enforces `acdl:*` labels that `nova_tagging.py` hard- - fails) are the highest-severity findings, both correctness regressions - left by the rebrand. P2–P4 independent brand/dead-code/except work. -- **Wave 2 (P5–P9): simplify without regressions.** P5 before P6/P9 - (regression-verify dedup is independent; P6/P9 both touch - `run_platform.sh`). P8 changes the workflow byte-identity test → - generator (D-115). P9 must run the regression gate (D-118) at the end - of Wave 2 — 22/22 capabilities must stay Verified. -- **Wave 3 (P10–P14): security + maintainability.** P10 before P11 - (identity enforcement before payload validation). P12/P13 independent - file splits. P14 mid-milestone checkpoint (offline) at end of Wave 3. -- **Wave 4 (P15–P17): developer experience.** Independent; P17 last - (reflects the consolidated path after P15/P16 land). -- **Wave 5 (P18–P20): no-humans onboarding (request-path only).** P18 - (schema + Lambda action) before P19 (env-file autogen consumes the - schema) before P20 (cross-account role, offline-proven per D-114). -- **Final (P21): review + audit + milestone ship.** +--- + +## Wave Overview + +| Wave | Phases | Theme | Dependency rationale | +|------|--------|-------|----------------------| +| **Wave 1** | P1 | Event emitters — the foundation | Everything depends on events being emitted. P1 establishes the CloudEvents envelope, per-run manifests, Decision Ledger, Infracost adapter, and attestation/confidence/policy event emission. P2/P3/P4/P5 all consume P1's event formats. | +| **Wave 2** | P2, P3 | Collector + PowerBI export | P2 (collector) reads P1's events/files → SQLite cold store. P3 (export) reads P2's SQLite → CSV/JSON views. P2 + P3 can partially parallelize: P3's view schemas can be authored against P2's schemas before P2's collector code is complete, but P3's export code needs P2's SQLite to exist. | +| **Wave 3** | P4, P5 | Metrics catalog + deck rebuild | P4 (catalog + NORTH_STAR integration) depends on metrics being grounded (P2/P3 done). P5 (deck) depends on P4's `METRICS.md` for the Proof section's grounded citations. P4 + P5 can partially parallelize: P5's Problem/Vision/How acts don't need P4; P5's Proof act needs P4's catalog. | +| **Wave 4** | P6, P7 | Regression capability + final review/ship | P6 (CAP-023/024) depends on P2/P3 (collector) + P5 (deck) existing. P7 (review/audit/ship) depends on all prior phases. | +| **Final** | P8 | Milestone ship | Merge to main, tag `v1.16.8`, Gitea release, delete milestone branches. | + +**Dependency chain (critical path):** +P1 → P2 → P3 → P4 → P5(Proof) → P6 → P7 → P8 + +**Parallelization opportunities:** +- P2 schemas + P3 view schemas can be authored concurrently (Wave 2 entry). +- P4 `docs/metrics/*.md` per-KPI docs + P5 Problem/Vision/How acts can be + authored concurrently (Wave 3 entry); P5 Proof act waits for P4 `METRICS.md`. +- P6 CAP-023 (collector) test can be drafted while P5 finishes (the test + needs P2's collector to exist, which it does by Wave 4). + +--- + +## Per-Phase Vertical-Slice Plans + +### Phase P1 — event-emitters (Wave 1, feat) + +**Goal:** Instrument every Nova decision point to emit structured CloudEvents +1.0 events + persist ephemeral `$WORK/*.json` as durable artifacts + extend +`outbox_writer.py` into the SQLite Decision Ledger. After P1, the metrics +layer has all the raw signals it needs — no downstream phase invents new +signals. + +**Requirements covered:** REQ-187, REQ-188, REQ-205 (emitter half), +REQ-206 (emitter half). + +**Primary persona:** backend-engineer. **Supporting:** data-engineer +(event schemas). + +**Tasks (vertical slices):** + +1. **CloudEvents envelope + schemas** — `core/metrics/event_envelope.py` + defines the CloudEvents 1.0 envelope + `platform.*` semantic conventions + (specversion, id, source, type, time, subject, datacontenttype, platform + block, data). `schemas/metrics_event.schema.json` validates the envelope. + `schemas/metrics_run_manifest.schema.json` validates per-run manifests. + - *Acceptance:* `python -m jsonschema` validates a sample event against + the schema; `tests/test_metrics_emitters.py::test_envelope` passes. + +2. **Per-run manifest writer** — `core/metrics/run_manifest.py` emits + `nova.run.started`, `nova.run.completed`, `nova.run.failed` events with + (run_id, contractId, env, stages×durations, exit, confidence, HITL block + count). Writes `metrics/runs/.json`. `scripts/run_platform.sh` + invokes the writer at run start + run end. + - *Acceptance:* a `--check-only` run produces `metrics/runs/.json` + with a valid manifest; `test_run_manifest` passes. + +3. **Persist ephemeral `$WORK/*.json`** — `run_platform.sh` copies + `$WORK/pcr.json`, `signal.json`, `event.json`, `outbox_item.json`, + `stack.json` to `metrics/runs//` as durable artifacts (the + ephemeral `$WORK` copies remain for the running pipeline; the persisted + copies are the metrics source of truth). + - *Acceptance:* after a run, `metrics/runs//pcr.json` exists and + matches `$WORK/pcr.json`; a test asserts the copy. + +4. **pytest addopts** — `pyproject.toml` `addopts` gains + `--junitxml=metrics/test-results.xml --json-report --cov=core --cov=adapters + --cov-report=json:metrics/coverage.json`. CAP-009 (offline pytest suite + passes) must remain Verified (assumption A5 — additive flags). + - *Acceptance:* `bash scripts/run_ci.sh` exits 0; `metrics/test-results.xml` + + `metrics/coverage.json` exist; regression gate 22/22 (run at P6, but + P1 must not break any cap locally). + +5. **Infracost post-processor** — `core/metrics/infracost_adapter.py` runs + Infracost on `terraform show -json plan.tfplan` (offline, reads plan JSON, + no live AWS). Emits `nova.cost.estimated{delta_usd}`. Degrades gracefully + (omits the event, logs a warning) when Infracost CLI is absent (A6). + `run_platform.sh` invokes it after the plan stage. + - *Acceptance:* when Infracost is available, `metrics/runs//` + contains a `cost_estimate.json`; when absent, the run still exits 0; + `test_infracost_adapter` passes (mock the CLI). + +6. **Decision Ledger (SQLite hash-chain)** — `core/metrics/decision_ledger.py` + extends `outbox_writer.py` to emit to a SQLite append-only table + (`metrics/decision_ledger.db`) with a hash chain (`prev_hash` + own + `hash`, SHA-256). Emits `ai.decision.made` events (decision_id=run_id, + chosen_action=band outcome, confidence=score, alternatives=perInput + breakdown, human_override=HITL block) with outcome backfill from + `apply.completed`. Honors D-083 (no S3 Object Lock/JWS — local SQLite + hash-chain only). + - *Acceptance:* `metrics/decision_ledger.db` exists after a run; the + hash chain verifies (`verify-chain` returns 0 broken); `test_decision_ledger` + passes. + +7. **Attestation event emission** — `core/hitl_gates.py` emits + `attestation.recorded` events to the Decision Ledger on qa/prod/dr gates + (approver, env, concerns, result). D-132. (Dev skips — autonomous.) + - *Acceptance:* a mocked qa gate produces an `attestation.recorded` row + in the Decision Ledger; `test_attestation_event` passes. + +8. **Confidence decision event emission** — `core/confidence_signal.py` + emits `nova.confidence.computed` + `nova.ai.decision.made` events (D-122: + the "AI decision" is the confidence-gated policy engine, not an LLM). + - *Acceptance:* a confidence computation produces both events in + `metrics/events.jsonl`; `test_confidence_event` passes. + +9. **Policy event emission** — `adapters/terraform/policy/checkov_adapter.py` + emits `nova.policy.evaluated` events (rule count, pass/fail/skipped, + severity breakdown). + - *Acceptance:* a Checkov run produces a `nova.policy.evaluated` event; + `test_policy_event` passes. + +10. **Lifecycle success-rate emitter** — each lifecycle run writes + `metrics/lifecycle/-.json` (module, env, phase + apply/modify/destroy, result, duration_ms). REQ-205 emitter half. + - *Acceptance:* a mocked lifecycle run produces the JSON; the emitter + test passes. + +11. **Capability event emission** — `core/regression_verify.py` emits + `nova.capability.verified` events (capability ID, status, tier, duration). + - *Acceptance:* a regression run produces `nova.capability.verified` + events; `test_capability_event` passes. + +**Must-haves (phase ships only if ALL true):** +- `core/metrics/event_envelope.py`, `run_manifest.py`, + `infracost_adapter.py`, `decision_ledger.py` exist and are tested. +- `metrics/events.jsonl` is appended to on every run (CloudEvents 1.0 + envelope, valid against `schemas/metrics_event.schema.json`). +- `metrics/runs/.json` manifest exists after every run. +- `metrics/decision_ledger.db` exists with a verified hash chain. +- `outbox_writer.py` extended to write to the SQLite Decision Ledger. +- `hitl_gates.py` emits `attestation.recorded` (D-132). +- `confidence_signal.py` emits `nova.confidence.computed` + + `nova.ai.decision.made` (D-122). +- `checkov_adapter.py` emits `nova.policy.evaluated`. +- `pyproject.toml` addopts include `--junitxml` + `--json-report` + `--cov`. +- `bash scripts/run_ci.sh` exits 0. +- No existing capability regresses (22/22 locally). + +**Risks + mitigations:** +- *Risk:* `--junitxml`/`--cov` addopts break the existing test suite. + *Mitigation:* A5 (additive flags); verify CAP-009 stays Verified locally + before merging. +- *Risk:* Infracost CLI not available in CI. *Mitigation:* A6 — degraded + mode (omit event, log warning, don't fail the run). +- *Risk:* SQLite hash-chain corruption on concurrent writes. *Mitigation:* + single-writer model (the run manifest writer is the only writer per run); + WAL mode + `BEGIN IMMEDIATE`. +- *Risk:* Event schema drift between emitters and collector. *Mitigation:* + schemas authored first (task 1); all emitters validate against the schema + before writing. + +--- + +### Phase P2 — metrics-collector (Wave 2, feat) + +**Goal:** Read all grounded signals (files + events) into a normalized +SQLite cold store at `metrics/nova_metrics.db` with idempotent re-runs. +After P2, the metrics layer has a queryable store — P3 exports it, P4 +catalogs it. + +**Requirements covered:** REQ-189, REQ-200, REQ-201, REQ-205 (collector +half), REQ-206 (collector half), REQ-207. + +**Primary persona:** data-engineer. **Supporting:** backend-engineer +(event formats). + +**Tasks (vertical slices):** + +1. **Fact/dimension schemas** — `schemas/metrics_fact_run.schema.json`, + `schemas/metrics_fact_capability.schema.json`, + `schemas/metrics_fact_policy_check.schema.json`, + `schemas/metrics_fact_confidence.schema.json`, + `schemas/metrics_fact_test.schema.json`, + `schemas/metrics_fact_decision.schema.json`, + `schemas/metrics_fact_cost_estimate.schema.json`, + `schemas/metrics_fact_lifecycle.schema.json`, + `schemas/metrics_dim_capability.schema.json`, + `schemas/metrics_dim_milestone.schema.json`. Schema-first (data-engineer + constraint): all schemas exist before any collector code. + - *Acceptance:* all schemas validate sample rows; `python -m jsonschema` + passes for each. + +2. **Collector core** — `core/metrics/collector.py` reads: + - `REGRESSION_REPORT.json` → `fact_capability` + `dim_capability`. + - `metrics/runs/*.json` → `fact_run`. + - `metrics/test-results.xml` (junit) → `fact_test`. + - `metrics/coverage.json` → `fact_test.coverage` column. + - `metrics/runs//pcr.json` → `fact_policy_check`. + - `metrics/runs//signal.json` → `fact_confidence`. + - `metrics/decision_ledger.db` → `fact_decision`. + - `metrics/runs//cost_estimate.json` → `fact_cost_estimate`. + - `metrics/lifecycle/*.json` → `fact_lifecycle`. + - `CHECKPOINT.json` → `dim_milestone`. + Writes to `metrics/nova_metrics.db` (SQLite cold store, D-126). + - *Acceptance:* after a run + collector invocation, + `metrics/nova_metrics.db` has all fact/dim tables populated; + `test_metrics_collector` passes. + +3. **Idempotent re-runs** — the collector is idempotent: re-running it + produces identical row counts + a verified chain. REQ-200. + - *Acceptance:* `test_metrics_collector_idempotent` passes (two runs → + identical row counts + chain verified). + +4. **Decision Ledger CLI** — `core/metrics/decision_ledger_cli.py` supports + `query`, `verify-chain`, `stats`, `export`, `replay`. `verify-chain` + detects broken hashes; `replay` prints ordered events. REQ-207. + - *Acceptance:* `decision_ledger_cli.py verify-chain` exits 0 on a clean + chain, exits 1 on a tampered chain; `test_decision_ledger_cli` passes. + +5. **Metrics README** — `metrics/README.md` documents regenerable vs + append-only artifacts + the restore procedure (the cold store is + regenerable from the raw signals; the Decision Ledger is append-only). + REQ-201. + - *Acceptance:* `metrics/README.md` exists with the two categories + a + restore procedure section. + +**Must-haves:** +- `core/metrics/collector.py` exists and is tested. +- `metrics/nova_metrics.db` is produced with all fact/dim tables. +- Idempotent re-runs (REQ-200) verified by test. +- `core/metrics/decision_ledger_cli.py` exists with all 5 subcommands. +- `metrics/README.md` documents regenerable vs append-only + restore. +- `bash scripts/run_ci.sh` exits 0. + +**Risks + mitigations:** +- *Risk:* Schema drift between P1's event formats and P2's fact schemas. + *Mitigation:* data-engineer authors both; backend-engineer reviews the + event-format alignment. +- *Risk:* Junit XML parsing edge cases (test names with special chars). + *Mitigation:* use `xml.etree.ElementTree` with XPath; test with a fixture + containing edge-case names. + +--- + +### Phase P3 — powerbi-export (Wave 2, feat) + +**Goal:** Emit CSV/JSON views from the SQLite cold store to +`metrics/powerbi/` — fact + dimension views + 8 empty placeholder views +for deferred metrics. After P3, a PowerBI folder-connector dashboard can +be built. + +**Requirements covered:** REQ-190, REQ-199, REQ-208, REQ-209 (P3 half), +REQ-205 (view half). + +**Primary persona:** data-engineer. + +**Tasks (vertical slices):** + +1. **PowerBI export core** — `core/metrics/powerbi_export.py` reads + `metrics/nova_metrics.db` and emits CSV/JSON views to `metrics/powerbi/`: + `fact_run.csv`, `fact_capability.csv`, `fact_policy_check.csv`, + `fact_confidence.csv`, `fact_test.csv`, `fact_decision.csv`, + `fact_cost_estimate.csv`, `fact_lifecycle.csv`, `dim_capability.csv`, + `dim_milestone.csv`. D-129 (CSV/JSON folder connector). + - *Acceptance:* after `powerbi_export.py` runs, all 10 CSV files exist + in `metrics/powerbi/` with non-empty content (given a populated cold + store); `test_powerbi_export` passes. + +2. **8 deferred placeholder views** — empty CSV files with documented + schemas (headers only, no data rows) for the 8 deferred metrics: + (1) Live Infrastructure Health, (2) Live Outbox Write Rate, + (3) Tamper-Evident Ledger Checkpoints, (4) Onboarding Funnel + (requested→granted), (5) Drift Auto-Reversal Rate, (6) Live CUR + Reconciliation, (7) SLA / Unplanned Downtime, (8) Predictive vs Reactive + Ratio. D-123. Each has a header row documenting the columns + a comment + row citing the blocking decision ID. + - *Acceptance:* all 8 placeholder CSVs exist with header rows + a + decision-ID comment; `test_placeholder_views` passes. + +3. **METRICS_VIEWS.md data dictionary** — `docs/METRICS_VIEWS.md` has a + per-column data-dictionary table (column, type, source/formula, unit, + grounded/derived/deferred status) for every view. REQ-209 (P3 half). + - *Acceptance:* `docs/METRICS_VIEWS.md` exists with a complete + per-column table covering all 18 views (10 fact/dim + 8 placeholder). + +4. **NOVA_DASHBOARD_README.md** — `metrics/powerbi/NOVA_DASHBOARD_README.md` + documents the folder-connector import path + a starter visual model + + a reference screenshot placeholder. REQ-208. + - *Acceptance:* the README exists with import steps + visual model + description. + +5. **Schema validation in CI** — `run_ci.sh` validates + `metrics/powerbi/*.json` + a sample `metrics/events.jsonl` against their + schemas; exits 0. REQ-199. + - *Acceptance:* `bash scripts/run_ci.sh` validates the PowerBI JSON + exports + a sample events file; exits 0. + +**Must-haves:** +- `core/metrics/powerbi_export.py` exists and is tested. +- `metrics/powerbi/` contains all 10 fact/dim CSVs + 8 placeholder CSVs. +- `docs/METRICS_VIEWS.md` has the per-column data dictionary. +- `metrics/powerbi/NOVA_DASHBOARD_README.md` exists. +- `run_ci.sh` schema validation (REQ-199) passes. +- `bash scripts/run_ci.sh` exits 0. + +**Risks + mitigations:** +- *Risk:* Placeholder view schemas diverge from what the future emitter + will produce. *Mitigation:* the schema is documented in the header row + + METRICS_VIEWS.md; the future emitter must conform to the documented + schema. +- *Risk:* PowerBI folder connector quirks (CSV encoding, delimiters). + *Mitigation:* UTF-8 + comma-delimited; documented in the README. + +--- + +### Phase P4 — metrics-catalog + north-star-integration (Wave 3, docs) + +**Goal:** Catalog every executive KPI in `docs/METRICS.md` with +grounded/derived/deferred status + per-KPI definition-of-success docs. +Wire `NORTH_STAR.md` into CIAgent context-loading so every future +`/ci-run` reads it. Produce the trust-snapshot report, the deferred-metrics +roadmap, the confidence-gate halt rate metric, and the no-humans thesis +brief. After P4, the metrics layer is fully documented and the strategic +direction is durable. + +**Requirements covered:** REQ-186, REQ-191, REQ-192, REQ-193, REQ-194, +REQ-195, REQ-204, REQ-209 (P4 half), REQ-210, REQ-211, REQ-212, REQ-213 +(P4 half). + +**Primary persona:** lead-developer. **Supporting:** data-engineer +(metric definitions). + +**Tasks (vertical slices):** + +1. **METRICS.md catalog** — `docs/METRICS.md` catalogs every executive KPI + with: name, NORTH_STAR target, `grounded`/`derived`/`deferred` status, + source file or decision ID, and a link to the per-KPI definition doc. + REQ-195. Covers all metrics from the scorecard (RESEARCH.md §3): + Touchless Resolution Rate, Human Escalation Frequency, MTTR (platform-run), + AI Decision Accuracy, Decision Ledger Coverage, Attestation Coverage, + Capability Health, Confidence Distribution, Policy Pass Rate, Test + Count/Pass Rate, Provisioning Lead Time, Deployment Frequency, + Cost Estimates (Infracost), FTE Hours Saved, Platform ROI, + Confidence-Gate Halt Rate, + the 8 deferred metrics. + - *Acceptance:* `docs/METRICS.md` exists; every KPI has a status badge + + a source link; a grep confirms no KPI is missing a status. + +2. **Per-KPI definition-of-success docs** — `docs/metrics/.md` for + every KPI (D-127). Each doc defines: the metric, the formula, the + grounding status, the source file, the definition of success (what + number = "won"), and the deferred dependency (if applicable). + - *Acceptance:* `docs/metrics/` contains one `.md` per KPI; each doc + has all 5 sections. + +3. **Zero-touch efficiency metrics docs** — REQ-191: Autonomous Resolution + Rate, Human Escalation Frequency, AI Decision Accuracy, MTTD/MTTR + (platform-run, D-131). Documented in METRICS.md + per-KPI docs with + the attestation exclusion clarification (attestation gates are designed + controls, not escalations). + - *Acceptance:* the 4 metrics have per-KPI docs with the correct + formulas + attestation exclusion language. + +4. **Velocity metrics docs** — REQ-192: Provisioning Lead Time + (apply.completed.time − intent.received.time), Deployment Frequency + (count(apply.completed) per day). Self-Healing Velocity deferred. + - *Acceptance:* the 2 metrics have per-KPI docs; the deferral is + documented. + +5. **Financial & cost-ROI metrics docs** — REQ-193: FTE Hours Saved + (derived), Cost Savings via Infracost (grounded), Cost Efficiency Ratio + (derived), Platform ROI (derived formula). Live CUR deferred (D-096). + - *Acceptance:* the 4 metrics have per-KPI docs with formulas; the CUR + deferral cites D-096. + +6. **Reliability, security & compliance metrics docs** — REQ-194: + Zero-Trust Policy Compliance Rate (from pcr.json), Attestation Coverage + (prod/dr promotions attested by a human ÷ total prod/dr promotions; + grounded in `hitl_gates.py` + outbox `approver_*` attributes). Uptime, + Patch Remediation, SLA/downtime deferred (D-096). **Attestation Coverage + is canonically owned here (REQ-194), not in REQ-191.** + - *Acceptance:* the 2 grounded metrics have per-KPI docs; the 3 deferred + metrics have deferral docs citing D-096. + +7. **NORTH_STAR integration** — REQ-186: `NORTH_STAR.md` is referenced from + `PROJECT.md` (a "Strategic Direction" section pointing to it) + + `ARCHITECTURE.md` (the v1.17 addendum already references it). `config.json` + gains `strategic_direction_file: ".ciagent/NORTH_STAR.md"` so the run + workflow reads it at SPECIFY. + - *Acceptance:* `PROJECT.md` has a Strategic Direction section; + `config.json` has the `strategic_direction_file` key; a test confirms + the file is readable. + +8. **NORTH_STAR diff-check in CI** — REQ-204: `run_ci.sh` includes + `check_north_star_diff` that fails when Vision/Objectives/Anti-Goals/ + Targets sections change without a `NORTH_STAR-CHANGE:` commit trailer. + - *Acceptance:* a test commit changing a Target without the trailer + fails the check; a commit with the trailer passes. + +9. **Deferred-metrics activation roadmap** — `docs/METRICS_DEFERRED_ROADMAP.md` + lists 8 deferred metrics + onboarding-grant half with {blocking decision, + unblock requirement, candidate milestone} + a "Hot-Path Activation + (post-D-096)" section (Nova-native only, D-120) + "Re-evaluation + Triggers" section. REQ-210. + - *Acceptance:* the roadmap exists with all 8 + the onboarding-grant + half + the 2 sections. + +10. **Trust-snapshot report** — `core/metrics/trust_snapshot.py` emits + `metrics/TRUST_SNAPSHOT.md` with 5 trust metrics (Decision Ledger + Coverage, Attestation Coverage, Capability Health, AI Decision + Accuracy, Confidence-Gate Halt Rate) + chain-integrity verdict + + snapshot hash. Runs offline. REQ-211. + - *Acceptance:* `metrics/TRUST_SNAPSHOT.md` exists after running + `trust_snapshot.py`; the 5 metrics + verdict + hash are present; + `test_trust_snapshot` passes. + +11. **Confidence-Gate Halt Rate metric** — REQ-212: `docs/METRICS.md` + + trust snapshot include "Confidence-Gate Halt Rate" (signal.json + band=halt ÷ total runs). PowerBI view includes it (added to + `fact_confidence` projection in P3's export — coordinate with P3). + - *Acceptance:* METRICS.md has the metric; the trust snapshot includes + it; the PowerBI export includes a column for it. + +12. **No-humans thesis brief** — `docs/NO_HUMANS_THESIS.md` defines the + thesis, grounded proof metrics, deferred proof metrics, and explicit + anti-claims (incl. D-122 honesty: the "AI" is the confidence-gated + policy engine, not an LLM). REQ-213 (P4 half). The unified deck's + Vision act cites it (P5). + - *Acceptance:* `docs/NO_HUMANS_THESIS.md` exists with all 4 sections; + the anti-claims section explicitly addresses D-122. + +**Must-haves:** +- `docs/METRICS.md` catalogs every KPI with status + source. +- `docs/metrics/*.md` per-KPI docs exist for every KPI. +- `NORTH_STAR.md` referenced from PROJECT.md + ARCHITECTURE.md + config.json. +- `run_ci.sh` includes `check_north_star_diff` (REQ-204). +- `docs/METRICS_DEFERRED_ROADMAP.md` exists (REQ-210). +- `core/metrics/trust_snapshot.py` + `metrics/TRUST_SNAPSHOT.md` (REQ-211). +- Confidence-Gate Halt Rate in METRICS.md + trust snapshot + PowerBI (REQ-212). +- `docs/NO_HUMANS_THESIS.md` exists (REQ-213 P4 half). +- `bash scripts/run_ci.sh` exits 0. + +**Risks + mitigations:** +- *Risk:* KPI definitions drift from NORTH_STAR targets. *Mitigation:* + the catalog cross-references NORTH_STAR target rows; the diff-check + (REQ-204) catches NORTH_STAR changes. +- *Risk:* The no-humans thesis overclaims. *Mitigation:* D-122 honesty + constraint — the anti-claims section explicitly states the "AI" is the + confidence-gated policy engine; A3. + +--- + +### Phase P5 — deck-rebuild (Wave 3, docs+test) + +**Goal:** Merge the two existing decks into one unified narrative deck +"Nova — The No-Humans Infrastructure Platform" with the 5-act arc +(Problem → Vision → How → Proof → Roadmap), x3 structure at deck + slide +level, per-slide benefit callouts, fluid transitions, a metrics glossary +appendix slide, a "what's deferred" slide, and the no-humans thesis cited +in the Vision act. Retire both old decks. Re-run the 4-step deck process +(source `.md` → Marp → HTML → talking-points). + +**Requirements covered:** REQ-196, REQ-197, REQ-202, REQ-203, REQ-213 +(P5 half). + +**Primary persona:** lead-developer. + +**Tasks (vertical slices):** + +1. **Unified deck source markdown** — `docs/presentations/nova-no-humans-platform.md` + is the single source of truth (the full slide-by-slide plan is in the + "Deck Rebuild Plan" section below). The 5-act arc with x3 at deck level + (opening = arc preview, body = tell them, closing = recap + ask) + x3 + per slide (opens with what it covers, delivers, closes with benefit + callout). Fluid transitions written into each slide's opening line. + REQ-196, REQ-197. + - *Acceptance:* the source `.md` exists with all slides from the deck + plan below; each slide has the 3-part structure; transitions are + written. + +2. **Marp deck** — `docs/presentations/nova-no-humans-platform-marp.md` + (Marp-formatted with the S&P visual theme, `sp-theme.json` unchanged). + - *Acceptance:* the Marp deck renders to HTML with the correct slide + count + theme. + +3. **HTML render** — `docs/presentations/nova-no-humans-platform.html` + (re-rendered from the Marp deck). + - *Acceptance:* the HTML exists and opens with the correct title slide. + +4. **Talking points** — `docs/presentations/nova-no-humans-platform-talking-points.md` + (distilled from the Marp deck, one section per slide with speaker notes). + - *Acceptance:* the talking-points file exists with one section per + slide. + +5. **Metrics glossary appendix slide** — REQ-202: the deck has a + "Metrics Glossary" appendix slide with one-line KPI definitions + + grounding badges (grounded/derived/deferred). + - *Acceptance:* the glossary slide exists with all KPIs + badges. + +6. **"What's Deferred — and Why" slide** — REQ-203: the deck has a slide + pairing each of 8 deferred metrics with its blocking decision ID. + - *Acceptance:* the deferred slide exists with all 8 + decision IDs. + +7. **No-humans thesis cited in Vision act** — REQ-213 (P5 half): the + Vision act cites `docs/NO_HUMANS_THESIS.md` (the thesis, grounded proof, + deferred proof, anti-claims). + - *Acceptance:* the Vision act slides reference the thesis brief. + +8. **Retire both old decks** — delete `how-the-platform-works.md` + + `-marp.md` + `.html` + `-talking-points.md` + `the-developer-experience.md` + + `-marp.md` + `.html` + `-talking-points.md`. D-130. + - *Acceptance:* a grep confirms the old deck files are deleted; no + references to them remain in the repo. + +**Must-haves:** +- `docs/presentations/nova-no-humans-platform.md` (+ marp + html + + talking-points) exists with the full slide plan. +- x3 structure at deck + slide level (REQ-197). +- Per-slide benefit callouts (REQ-197). +- Fluid transitions written into each slide (REQ-197). +- Metrics glossary appendix slide (REQ-202). +- "What's Deferred" slide (REQ-203). +- No-humans thesis cited in Vision act (REQ-213 P5 half). +- Both old decks deleted (D-130). +- `bash scripts/run_ci.sh` exits 0. + +**Risks + mitigations:** +- *Risk:* The deck claims a metric that isn't grounded yet. *Mitigation:* + P5 Proof act depends on P4's METRICS.md; every cited metric has a + grounded source file verified by the catalog. +- *Risk:* The old decks are referenced by other docs. *Mitigation:* grep + for references before deletion; update or remove them. + +--- + +### Phase P6 — regression-capability (Wave 4, test) + +**Goal:** Add CAP-023 (metrics collector runs, emits expected schema) + +CAP-024 (deck structure: slide count, x3 present, per-slide benefit +present) to `core/regression_verify.py`. After P6, the regression gate +protects the metrics layer + the deck structure. + +**Requirements covered:** REQ-198. + +**Primary persona:** backend-engineer. **Supporting:** data-engineer +(CAP-023 schema). + +**Tasks (vertical slices):** + +1. **CAP-023 — metrics collector** — `core/regression_verify.py` gains a + `CAP-023` check: runs `core/metrics/collector.py` against a fixture + metrics dir, asserts the SQLite cold store has all fact/dim tables with + the expected schema, asserts idempotent re-run. Tags Verified/Decayed/Broken. + - *Acceptance:* `CAP-023` returns Verified when the collector produces + the correct schema; `test_regression_cap023` passes. + +2. **CAP-024 — deck structure** — `core/regression_verify.py` gains a + `CAP-024` check: parses `docs/presentations/nova-no-humans-platform.md`, + asserts (a) slide count is in the expected range (12–20), (b) the x3 + structure is present (opening arc preview + closing recap), (c) each + slide has a benefit callout. Tags Verified/Decayed/Broken. + - *Acceptance:* `CAP-024` returns Verified when the deck meets all 3 + criteria; `test_regression_cap024` passes. + +3. **Regression gate run** — `bash scripts/run_regression.sh` runs the + full gate (22 prior capabilities + CAP-023 + CAP-024 = 24 total). All + must pass (Verified or Skipped per D-118). + - *Acceptance:* the regression report shows 24 capabilities, all + Verified or Skipped, 0 Decayed/Broken. + +**Must-haves:** +- `CAP-023` + `CAP-024` in `core/regression_verify.py`. +- `bash scripts/run_regression.sh` passes (24 capabilities, 0 Broken). +- `bash scripts/run_ci.sh` exits 0. + +**Risks + mitigations:** +- *Risk:* CAP-024's slide-count range is too tight and breaks on minor + deck edits. *Mitigation:* the range is 12–20 (generous); the check + focuses on structure (x3 + benefit callouts), not exact count. + +--- + +### Phase P7 — final-review-ship (Wave 4, review+audit+ship) + +**Goal:** Multi-persona review (incl. deck story quality), audit, and +milestone ship. After P7, v1.17 is complete and ready for the final merge. + +**Requirements covered:** all (review gate). + +**Primary persona:** lead-developer. **Supporting:** all active personas +(review participation). + +**Tasks (vertical slices):** + +1. **Multi-persona review** — each active persona reviews their territory: + - backend-engineer: event emitters, Decision Ledger, Infracost adapter, + regression CAP-023/024 code. + - data-engineer: collector, PowerBI export, schemas, data dictionary. + - lead-developer: NORTH_STAR integration, METRICS.md catalog, deck + narrative, no-humans thesis. + - Deck story quality review: the lead-developer reviews the deck for + narrative coherence, fluidity, and benefit-callout quality. + - *Acceptance:* review findings recorded; P0/P1 findings fixed before + ship; P2 findings logged for future milestones. + +2. **Audit** — verify: + - All 29 requirements (REQ-185..213) have a status of `complete` in + the traceability table. + - No stale claims in the deck (every metric citation has a grounded + source). + - `NORTH_STAR.md` is readable + referenced. + - The regression gate passes (24 capabilities). + - `bash scripts/run_ci.sh` exits 0. + - *Acceptance:* audit PASS recorded in `---ci---` block. + +3. **Milestone completion** — update `PROJECT.md`, `ROADMAP.md`, + `REQUIREMENTS.md` traceability to mark v1.17 complete. Tag `v1.16.7` + (P7 patch on the v1.16.x line). + - *Acceptance:* `PROJECT.md` reflects v1.17 complete; tag `v1.16.7` + exists. + +**Must-haves:** +- All 29 requirements marked complete. +- Multi-persona review complete (incl. deck story quality). +- Audit PASS. +- Regression gate 24/24 (Verified or Skipped). +- `bash scripts/run_ci.sh` exits 0. +- Tag `v1.16.7` exists. + +**Risks + mitigations:** +- *Risk:* Review surfaces a P0 finding late. *Mitigation:* the review is + scoped to each persona's territory; findings are fixed before the audit + step. + +--- + +### Phase P8 — milestone-ship (Final) + +**Goal:** Merge the milestone branch to main, tag `v1.16.8` (the milestone +release), publish the Gitea release, and delete the milestone branches. + +**Requirements covered:** all (ship gate). + +**Primary persona:** lead-developer. + +**Tasks (vertical slices):** + +1. **Merge to main** — merge `milestone/v1.17-strategic-metrics-deck` → + `main`. + - *Acceptance:* `main` contains all v1.17 commits; `git log main` shows + the milestone merge. + +2. **Tag + release** — tag `v1.16.8` on main; publish the Gitea release + (`Nova v1.16.8 — Strategic Direction, Leadership Metrics & Unified Story`) + with the release notes summarizing the three pillars. + - *Acceptance:* tag `v1.16.8` exists; Gitea release published (release + ID recorded). + +3. **Delete milestone branches** — delete `milestone/v1.17-strategic-metrics-deck` + + all `phase/NN-*` branches. + - *Acceptance:* `git branch -r` shows no v1.17 milestone/phase branches. + +**Must-haves:** +- `main` has the v1.17 merge. +- Tag `v1.16.8` exists. +- Gitea release published. +- Milestone + phase branches deleted. + +**Risks + mitigations:** +- *Risk:* Merge conflicts on main. *Mitigation:* the milestone branch is + off the v1.16 complete merge; rebase before merge if needed. + +--- + +## Deck Rebuild Plan + +> The unified deck: **"Nova — The No-Humans Infrastructure Platform."** +> 5-act arc: Problem → Vision → How → Proof → Roadmap. x3 at deck level +> (opening = arc preview, body = tell them, closing = recap + ask) + x3 +> per slide (opens with what it covers, delivers, closes with benefit +> callout). Fluid transitions written into each slide's opening line. +> Act indicator in the Marp footer (`Act N/5: `). + +### Deck-level x3 structure + +| Level | "What I'm going to tell you" | "Tell them" | "What I told you" | +|-------|------------------------------|-------------|-------------------| +| **Deck** | Slide 1 (arc preview: Problem→Vision→How→Proof→Roadmap; with 18V+0-consumer stake line) | Slides 2–15 (the 5 acts, 14 slides) | Slide 16 (recap of 5 acts + the business-decision ask) | +| **Per slide** | Opening line: "This slide shows X" | Body: bullets/diagram/table | Closing line: "Benefit: you now know Y" | + +### Act 1 — Problem (2 slides) + +> **Transition into Act 1:** (none — this is the opening; the arc preview +> slide sets up all 5 acts). + +**Slide 1 — Arc Preview (the "what I'm going to tell you" deck-level opening)** +- *Opens:* "This deck proves Nova is the no-humans infrastructure platform — + and shows you the metrics that make the claim defensible." +- *Stake line (G-Q8 binding):* "Today: 18 capabilities verified, 0 consumer + estates in production. This deck shows what's proven, what's pipeline-ready, + and what's honestly deferred." +- *Delivers:* The 5-act arc as a visual roadmap: Problem → Vision → How → + Proof → Roadmap. One-line summary per act. +- *Closes:* "Benefit: you leave this deck knowing which claims are proven + today, which are pipeline-ready, and which are deferred with a documented + unblock path — no marketing, just grounded evidence." +- *Grounded metrics cited:* 18 Verified + 4 Skipped (source: + `REGRESSION_REPORT.json`); 0 consumers (source: `PROJECT.md:495`). +- *Deferred metrics:* none. + +**Slide 2 — The No-Humans Imperative** +- *Opens:* "This slide shows why the operator is the bottleneck — and why + removing them from operations (not accountability) is the imperative." +- *Delivers:* The cost of humans-in-the-loop: L1/L2 ops hours, escalation + latency, the trust gap (autonomous claims without proof). Cites the + no-humans thesis (`docs/NO_HUMANS_THESIS.md`). +- *Closes:* "Benefit: you now know the problem framing — autonomy in + operations, human at stage gates, is the path forward." +- *Grounded metrics cited:* none (problem framing). +- *Deferred metrics:* none. +- *Transition into Act 2:* "Having defined the problem, here is Nova's + strategic direction toward solving it." + +### Act 2 — Vision/Direction (3 slides) + +> **Transition into Act 2:** "Having defined the problem, here is Nova's +> strategic direction toward solving it." + +**Slide 3 — Nova's Vision** +- *Opens:* "This slide states Nova's vision — infrastructure operations + become invisible, with provable trust." +- *Delivers:* The NORTH_STAR vision statement verbatim. The attestation + model: human attestation required at stage gates (QA for production, SRE + for operational readiness); autonomy in operations, not in + accountability. Cites `docs/NO_HUMANS_THESIS.md` (the thesis, grounded + proof, deferred proof, anti-claims incl. D-122 honesty). +- *Closes:* "Benefit: you now know the destination — invisible operations + with provable trust, not promised trust." +- *Grounded metrics cited:* none (vision). +- *Deferred metrics:* none. +- *Transition:* "The vision is ambitious — here are the 4 strategic + objectives that make it concrete." + +**Slide 4 — Strategic Objectives + Anti-Goals** +- *Opens:* "This slide pairs what Nova is building toward (4 objectives) + with what Nova refuses to build (5 anti-goals)." +- *Delivers:* The 4 strategic objectives (zero-touch ops, provable trust, + compounding ROI, default substrate for agentic consumption) + the 5 + anti-goals (not a hyperscaler competitor, not a general AI platform, not + removing humans from accountability, not for legacy infra, not sold to + operators). From `NORTH_STAR.md`. +- *Closes:* "Benefit: you now know the scope boundaries — Nova is + purpose-built for infrastructure operations, sold to leadership on + outcomes, and explicitly not a general-purpose AI platform or a + hyperscaler competitor." +- *Grounded metrics cited:* none (direction). +- *Deferred metrics:* none. +- *Transition:* "The objectives are committed to measurable targets — + here is the 12–18 month scorecard, with honest grounding status." + +**Slide 5 — 12–18 Month Targets (the scorecard)** +- *Opens:* "This slide shows the committed targets — numbers a board + member can repeat back — with their grounding status." +- *Delivers:* The NORTH_STAR targets table with the grounding column: + Touchless Resolution Rate ≥99% (grounded), Human Escalation <0.1% + (grounded), MTTR <60s (grounded, platform-run), AI Decision Accuracy + ≥99.5% (grounded), Decision Ledger Coverage 100% (grounded), + Attestation Coverage 100% (grounded), Cloud Spend Reduction ≥25% + (partial — Infracost grounded, CUR deferred), Platform ROI ≥250% + (derived), + deferred targets (Predictive vs Reactive, Drift Auto-Reversal, + AI-Agent Intent Share) marked **Planned**. +- *Closes:* "Benefit: you now know the destination numbers — and which + ones are measurable today vs deferred honestly." +- *Grounded metrics cited:* Touchless Resolution Rate, Human Escalation + Frequency, MTTR, AI Decision Accuracy, Decision Ledger Coverage, + Attestation Coverage — all `grounded` with source files. +- *Deferred metrics marked Planned:* Predictive vs Reactive, Drift + Auto-Reversal, AI-Agent Intent Share. +- *Transition into Act 3:* "The targets are committed — here is how Nova + works to achieve them." + +### Act 3 — How it works (4 slides) + +> **Transition into Act 3:** "The targets are committed — here is how +> Nova works to achieve them." + +**Slide 6 — The Platform Pipeline** +- *Opens:* "This slide shows the contract-to-evidence pipeline — how + intent becomes verified infrastructure without an operator." +- *Delivers:* The pipeline flow: contract → resolver → adapter → terraform + plan → Checkov (policy) → confidence signal → HITL gate (dev autonomous; + qa/prod/dr attested) → apply → evidence. Mermaid diagram. Grounded in + `scripts/run_platform.sh` + `core/contract_resolver.py` + + `adapters/terraform/adapter.py` + `core/confidence_signal.py`. +- *Closes:* "Benefit: you now know the path from intent to evidence — + and where the human appears (stage gates only)." +- *Grounded metrics cited:* none (architecture). +- *Deferred metrics:* none. +- *Transition:* "The pipeline produces decisions — here is how every + decision is captured and made accountable." + +**Slide 7 — The Decision Ledger** +- *Opens:* "This slide shows the Decision Ledger — every AI decision + captured with confidence, alternatives, and outcome." +- *Delivers:* The Decision Ledger architecture: `outbox_writer.py` + extended → SQLite append-only hash-chain table. `ai.decision.made` + events (decision_id=run_id, chosen_action=band, confidence=score, + alternatives=perInput, human_override=HITL block) with outcome backfill + from `apply.completed`. `attestation.recorded` events for qa/prod/dr. + D-121, D-122, D-132. Honors D-083 (no S3 Object Lock/JWS — local + hash-chain this milestone). +- **D-122 honesty sentence (G-Q4 binding):** "Nova's 'AI' is the + confidence-gated policy engine (confidence_signal + HITL gate), not + an LLM planner. The Decision Ledger captures this real decision path — + not a fabricated 'AI agent' that doesn't exist yet." +- *Closes:* "Benefit: you now know why 'autonomous' is defensible — every + decision is immutable, queryable, and accountable. And you know exactly + what 'AI' means here: a confidence-gated policy engine, not a black-box + LLM." +- *Grounded metrics cited:* Decision Ledger Coverage 100% (source: + `core/metrics/decision_ledger.py` + `metrics/decision_ledger.db`). +- *Deferred metrics marked Planned:* Tamper-Evident Ledger Checkpoints + (D-083). +- *Transition:* "Decisions are captured — here is how stage-gate + attestation keeps humans in accountability." + +**Slide 8 — The 8-Concern Attestation Matrix** +- *Opens:* "This slide shows the 8-concern attestation matrix — the + designed controls that keep humans at stage gates." +- *Delivers:* The 8 concerns (functional, performance, security posture, + contract NFRs, operational readiness, incident response, capacity/cost, + resilience). Offline-testable concerns run for real; operator-supplied + concerns accept signed evidence artifacts. Separation-of-duties on prod. + Grounded in `core/attestation_matrix.py` + `core/hitl_gates.py`. +- *Closes:* "Benefit: you now know the gate model — autonomy in + operations, human in accountability, by design." +- *Grounded metrics cited:* Attestation Coverage 100% (source: + `core/hitl_gates.py` + outbox `approver_*` attributes). +- *Deferred metrics:* none. +- *Transition into Act 4 (G-Q13 binding — rewritten):* "You've now seen + how Nova works — the pipeline, the Decision Ledger, the attestation + gates. But 'how it works' is not 'proof it works.' The next four slides + show the measured evidence: capability health, trust metrics, efficiency, + and cost — every number grounded in a real file, not a marketing claim." + +**Slide 9 — Telemetry Architecture (G-Q14 binding — benefit reframed from data plumbing to trust)** +- *Opens:* "This slide shows how Nova instruments itself — the + CloudEvents envelope, the cold store, and the PowerBI export." +- *Delivers:* The telemetry architecture diagram (from ARCHITECTURE.md + v1.17 addendum): platform components → CloudEvents 1.0 envelope → + `metrics/events.jsonl` + `metrics/runs/` + `metrics/decision_ledger.db` + → collector → `metrics/nova_metrics.db` (SQLite cold store) → + `metrics/powerbi/` (CSV/JSON views) → PowerBI. D-120 (Nova-native), + D-125 (hybrid events/files), D-126 (cold-only). +- *Closes:* "Benefit: you now know that every metric in this deck is + traceable to a real emitted event — the architecture IS the trust + substrate. When a CFO asks 'where does this number come from?', the + answer is a file path, not a Slack thread." +- *Grounded metrics cited:* none (architecture). +- *Deferred metrics marked Planned:* Hot-path (live ops dashboard) — D-126. +- *Transition into Act 4:* "The architecture is sound — here is the + measured proof." + +### Act 4 — Proof (4 slides) + +> **Transition into Act 4:** "The architecture is sound — here is the +> measured proof." + +**Slide 10 — Capability Health + Confidence Distribution** +- *Opens:* "This slide shows the grounded proof: capability health and + confidence distribution from real runs." +- *Delivers:* Capability health: 18 Verified + 4 Skipped (post-D-096 + teardown) from `.ciagent/REGRESSION_REPORT.json`. Confidence + distribution: from `metrics/nova_metrics.db` `fact_confidence` — score + histogram, band breakdown (pass/halt). The honesty model: Skipped is + honest (resources torn down per D-096), not a failure. +- *Closes:* "Benefit: you now know the platform is verified — 18 + capabilities pass, 4 are honestly skipped, 0 broken." +- *Grounded metrics cited:* Capability Health (source: + `REGRESSION_REPORT.json`), Confidence Distribution (source: + `metrics/nova_metrics.db` `fact_confidence`). +- *Deferred metrics:* none. +- *Transition:* "Capability health is necessary — here is the trust + substrate that makes autonomy defensible." + +**Slide 11 — Decision Ledger + Attestation Coverage** +- *Opens:* "This slide shows the trust metrics — Decision Ledger coverage + and attestation coverage, both 100%." +- *Delivers:* Decision Ledger Coverage: 100% of platform runs emit + `ai.decision.made` with outcome backfill (source: + `metrics/decision_ledger.db`). Attestation Coverage: 100% of prod/dr + promotions attested by a human (source: `hitl_gates.py` + outbox + `approver_*` attributes). AI Decision Accuracy: decisions not followed + by apply.failed/incident within 5min. The trust-snapshot report + (`metrics/TRUST_SNAPSHOT.md`) with chain-integrity verdict. +- *Closes:* "Benefit: you now know the trust is provable — not a marketing + claim, a queryable record." +- *Grounded metrics cited:* Decision Ledger Coverage, Attestation + Coverage, AI Decision Accuracy (source: `metrics/decision_ledger.db` + + `metrics/TRUST_SNAPSHOT.md`). +- *Deferred metrics:* Tamper-Evident Ledger Checkpoints (D-083) — Planned. +- *Transition:* "Trust is provable — here is the operational efficiency + that makes the ROI real." + +**Slide 12 — Zero-Touch Efficiency (G-Q10 binding — split from old slide 12)** +- *Opens:* "This slide shows the zero-touch efficiency metrics — + touchless resolution, human escalation, and MTTR." +- *Delivers:* Touchless Resolution Rate (runs without operational HITL + block ÷ total; attestation gates excluded). Human Escalation Frequency + (operational HITL blocks only). MTTR (platform-run: apply.failed → + successful retry, D-131). **Post-Pilot caveat (G-Q5 binding):** these + three metrics are computed on N internal runs today; the + production-denominator activates when a pilot estate runs (see + NORTH_STAR Post-Pilot Targets section). +- *Closes:* "Benefit: you now know the zero-touch efficiency is + measurable — the pipeline works today on internal runs, and the + denominator expands to production estates when a pilot activates." +- *Grounded metrics cited:* Touchless Resolution Rate, Human Escalation + Frequency, MTTR (source: `metrics/nova_metrics.db` `fact_run`). +- *Derived metrics:* none on this slide. +- *Deferred metrics marked Planned:* Self-Healing Velocity (no + auto-remediator). +- *Transition:* "Efficiency is half the ROI story — here is the cost + side." + +**Slide 13 — Cost & ROI (G-Q10 binding — split from old slide 12; G-Q15 binding — formula inline + N=0 caveat)** +- *Opens:* "This slide shows the cost estimates and the ROI formula — + with honest caveats about the current denominator." +- *Delivers:* Cost Estimates via Infracost (pre-apply, grounded). + **ROI formula shown inline (G-Q15 binding):** `Platform ROI = (FTE + hours saved × blended rate + cloud savings + avoided downtime) ÷ + platform op cost`. **N=0 caveat (G-Q5/G-Q15 binding):** "These + derived metrics are computed on N internal runs today; the + production-denominator activates post-pilot. The formula is grounded; + the production numbers are not yet." FTE Hours Saved (derived). Platform + ROI (derived formula). The grounded/derived/deferred honesty model. +- *Closes:* "Benefit: you now know the ROI formula — and you know it's + computed on internal runs today, not fabricated production numbers. + The formula is ready; the production denominator activates with a + pilot." +- *Grounded metrics cited:* Cost Estimates (source: + `metrics/nova_metrics.db` `fact_cost_estimate`). +- *Derived metrics:* FTE Hours Saved, Platform ROI (formula shown inline). +- *Deferred metrics marked Planned:* Live CUR Reconciliation (D-096), + Drift Auto-Reversal (D-096). +- *Transition:* "The proof is grounded — here is what is honestly + deferred." + +**Slide 14 — What's Deferred — and Why (G-Q11 binding — preempt: deferrals are measurement infra, not whether the platform runs without humans)** +- *Opens:* "This slide pairs each deferred metric with its blocking + decision — honesty about what isn't measured yet." +- **Preempt (G-Q11 binding):** "To be clear: these deferrals are + *measurement infrastructure*, not whether the platform runs without + humans. The platform IS autonomous in operations. What's deferred is + the *evidence pipeline* for certain metrics (live infra health, drift + detection, predictive remediation) — not the autonomy itself." +- *Delivers:* The 8 deferred metrics + onboarding-grant half, each paired + with its blocking decision ID: (1) Live Infrastructure Health — D-096, + (2) Live Outbox Write Rate — D-096, (3) Tamper-Evident Ledger + Checkpoints — D-083, (4) Onboarding Funnel (granted) — D-113/D-114/D-119, + (5) Drift Auto-Reversal — D-096 + no scheduler, (6) Live CUR + Reconciliation — D-096, (7) SLA / Unplanned Downtime — D-096, + (8) Predictive vs Reactive — future emitter. From + `docs/METRICS_DEFERRED_ROADMAP.md`. +- *Closes:* "Benefit: you now know the boundaries — what Nova measures + today, and exactly what blocks the rest. The autonomy is real; the + measurement gaps are documented." +- *Grounded metrics cited:* none (deferral honesty). +- *Deferred metrics:* all 8 + onboarding-grant half, each with decision ID. +- *Transition into Act 5:* "The proof is honest — here is the roadmap + from here to the 12–18 month targets." + +### Act 5 — Roadmap/Ask (2 slides) + +> **Transition into Act 5:** "The proof is honest — here is the roadmap +> from here to the 12–18 month targets." + +**Slide 15 — Roadmap to the North Star** +- *Opens:* "This slide shows the path from v1.17's grounded metrics to + the 12–18 month targets — the unblock path for each deferred metric." +- *Delivers:* The deferred-metrics activation roadmap (from + `docs/METRICS_DEFERRED_ROADMAP.md`): each deferred metric → blocking + decision → unblock requirement → candidate milestone. The hot-path + activation section (post-D-096, Nova-native only, D-120). Re-evaluation + triggers. +- *Closes:* "Benefit: you now know the path — every deferred metric has + an unblock requirement and a candidate milestone." +- *Grounded metrics cited:* none (roadmap). +- *Deferred metrics:* all 8 referenced with unblock paths. +- *Transition:* "The roadmap is clear — here is the recap and the ask." + +**Slide 16 — Recap + Ask (the "what I told you" deck-level closing; G-Q16 binding — ask reframed as a business decision)** +- *Opens:* "This slide recaps the 5 acts and states the ask." +- *Delivers:* Recap: Problem (operator bottleneck) → Vision (invisible + ops, provable trust) → How (pipeline + Decision Ledger + attestation) → + Proof (18V+4S, 100% ledger coverage, 100% attestation, grounded ROI + formula) → Roadmap (deferred metrics have unblock paths). **The ask + (G-Q16 binding — reframed as a business decision, not insider + language):** "The ask is a business decision: approve a pilot estate + to activate the production-denominator metrics (Touchless Resolution, + Human Escalation, AI Decision Accuracy), and approve the tamper- + evident ledger build-out (D-083 lift) to move from local hash-chain + to S3 Object Lock + JWS. These two decisions move Nova from + 'pipeline-ready' to 'production-proven.'" +- *Closes:* "Benefit: you leave with a clear business decision to make + — approve a pilot + the ledger build-out — and the confidence that + every claim in this deck is grounded, derived, or honestly deferred." +- *Grounded metrics cited:* Capability Health, Decision Ledger Coverage, + Attestation Coverage (recap). +- *Deferred metrics:* referenced as the ask. + +### Appendix slides (2 slides) + +**Slide A1 — Metrics Glossary** +- *Opens:* "This appendix defines every KPI in one line with its grounding + badge." +- *Delivers:* One-line definitions for all KPIs with grounded/derived/ + deferred badges. REQ-202. +- *Closes:* "Benefit: you now have a reference for every metric mentioned + in the deck." +- *Grounded metrics cited:* all (glossary). +- *Deferred metrics:* all (badged). + +**Slide A2 — Operating Model & Cost** +- *Opens:* "This appendix shows the real cost figures + the zero-cost + steady state." +- *Delivers:* `COST.md` figures ($0.001883 / 8 days, ~$0.007/mo, + S3-dominated, zero BAU compute) + the zero-cost-steady-state / D-096 + teardown claim. References the pre-mortem (`PRE_MORTEM.md`: v1.10 decay + root cause + four forward failure modes + structural mitigations). +- *Closes:* "Benefit: you now know the operating cost is negligible — and + the structural mitigation that prevents decay." +- *Grounded metrics cited:* Cost figures (source: `COST.md`). +- *Deferred metrics:* none. + +### Fluidity strategy + +1. **Every slide's opening line references the previous slide's close.** + Each slide above has an explicit transition sentence. No disjointed + jumps. The Act 3→4 boundary (slide 9→10) was rewritten per G-Q13 + binding: "But 'how it works' is not 'proof it works.'" +2. **Act indicator in the Marp footer.** `Act N/5: ` keeps the + audience oriented. Configured in the Marp theme. +3. **The arc is visible.** Slide 1 (arc preview + stake line) + slide 16 + (recap + business-decision ask) bookend the deck. The audience always + knows where they are in the 5-act structure. +4. **Per-slide benefit callout is the last line.** Every slide closes with + "Benefit: ..." — the audience leaves each slide with a takeaway, not a + cliffhanger. Benefit callouts rewritten per G-Q9 binding (slides 1, 4, + 13, 16 now give specific value, not generic restatements). +5. **The Proof act is the centerpiece.** It is 5 slides (the longest act, + expanded from 4 per G-Q10 binding: slide 12 split into Zero-Touch + Efficiency + Cost & ROI) because the PO's direction is "prove it, don't + promise it." The grounded/derived/deferred honesty model is the + narrative spine of the Proof act. +6. **Deferred metrics are shown, not hidden.** Slide 14 ("What's Deferred + — and Why") pairs each deferred metric with its blocking decision, + with a preempt (G-Q11 binding) clarifying that deferrals are + measurement infrastructure, not whether the platform runs without + humans. +7. **The D-122 honesty sentence on slide 7.** The deck explicitly states + that Nova's "AI" is the confidence-gated policy engine, not an LLM + planner — per G-Q4 binding. This prevents the "no fabrication" + constraint from being violated by implication. +8. **Derived metrics carry the N=0 caveat.** Slides 12 and 13 annotate + derived metrics (FTE, ROI) with "computed on N internal runs; + production-denominator activates post-pilot" — per G-Q5/G-Q15 binding. + The ROI formula is shown inline (G-Q15). + +### Deck file inventory (after P5) + +| File | Status | +|------|--------| +| `docs/presentations/nova-no-humans-platform.md` | NEW (source of truth, 16 main + 2 appendix slides per G-Q10 split) | +| `docs/presentations/nova-no-humans-platform-marp.md` | NEW (Marp) | +| `docs/presentations/nova-no-humans-platform.html` | NEW (rendered) | +| `docs/presentations/nova-no-humans-platform-talking-points.md` | NEW (talking points) | +| `docs/presentations/how-the-platform-works.md` | DELETED (retired, D-130) | +| `docs/presentations/how-the-platform-works-marp.md` | DELETED | +| `docs/presentations/how-the-platform-works.html` | DELETED | +| `docs/presentations/how-the-platform-works-talking-points.md` | DELETED | +| `docs/presentations/the-developer-experience.md` | DELETED (retired, D-130) | +| `docs/presentations/the-developer-experience-marp.md` | DELETED | +| `docs/presentations/the-developer-experience.html` | DELETED | +| `docs/presentations/the-developer-experience-talking-points.md` | DELETED | + +--- + +## Wave Dependency Graph + +``` +Wave 1 Wave 2 Wave 3 Wave 4 Final + ┌──────────────────┐ ┌──────────────────┐ ┌──────────────┐ +P1 (event emitters)──┤P2 (collector) │ │P4 (catalog + │ │P6 (regression│ P8 + │ P3 (powerbi │──▶│ NORTH_STAR │──▶│ capability) │──▶(ship) + │ export) │ │ integration) │ │P7 (review + │ + └──────────────────┘ │P5 (deck rebuild) │ │ audit + ship)│ + └──────────────────┘ └──────────────┘ + +Critical path: +P1 ──▶ P2 ──▶ P3 ──▶ P4 ──▶ P5(Proof) ──▶ P6 ──▶ P7 ──▶ P8 + +Parallelization: + Wave 2: P2 schemas + P3 view schemas can be authored concurrently. + Wave 3: P4 docs/metrics/*.md + P5 Problem/Vision/How acts can be authored + concurrently; P5 Proof act waits for P4 METRICS.md. + Wave 4: P6 CAP-023 test can be drafted while P5 finishes. +``` + +**Dependency details:** + +| Phase | Depends on | Blocks | +|-------|------------|--------| +| P1 | (none — foundation) | P2, P3, P4, P5, P6 | +| P2 | P1 (event formats) | P3 (SQLite store), P4 (catalog sources), P6 (CAP-023) | +| P3 | P2 (SQLite store) | P4 (PowerBI view references), P6 (CAP-023 schema) | +| P4 | P2 + P3 (grounded metrics) | P5 (Proof act citations), P6 (CAP-024 deck structure) | +| P5 | P4 (METRICS.md for Proof act) | P6 (CAP-024 deck structure) | +| P6 | P2 + P3 (CAP-023) + P5 (CAP-024) | P7 (regression gate must pass) | +| P7 | P1–P6 (all prior phases) | P8 (audit must pass) | +| P8 | P7 (milestone complete) | (none — terminal) | + +--- ## Execution approach - **Per-phase ship:** each execution phase merges `phase/NN-*` → - `milestone/v1.16-nova-simplification` and tags a patch on the v1.15.x - line (`v1.15.6` = P1 ... `v1.15.26` = P21). -- **Verification:** 4-layer verify (structural/behavioral/security/ - quality) per phase; the regression gate (D-091, 22 capabilities) runs - at P9 (end of Wave 2) and P21 (milestone complete) per D-118. -- **No live AWS:** `NOVA_LIFECYCLE_MODE=plan` default; terraform changes - validated via `terraform validate` + `--check-only`. P20 cross-account - Terraform is offline-proven only (D-114). + `milestone/v1.17-strategic-metrics-deck` and tags a patch on the v1.16.x + line (`v1.16.1` = P1 ... `v1.16.7` = P7, `v1.16.8` = P8 final). +- **Verification:** 4-layer verify (structural/behavioral/security/quality) + per phase; the regression gate (D-091, 22 prior + CAP-023 + CAP-024 = 24 + capabilities) runs at P6 and P7. +- **No live AWS:** `NOVA_LIFECYCLE_MODE=plan` default; all metrics that + require live AWS ship as placeholder views (D-096). Infracost runs + offline (reads plan JSON, A6). - **Test discipline:** each phase that changes runtime code adds/updates tests; `bash scripts/run_ci.sh` exits 0 at every phase boundary. - -## Wave 1 — Correctness + Brand Regression Fixes (P1–P4) - -### Phase P1 — state-bucket-and-kyverno-rebrand-fix (REQ-165) -- **Lead:** backend-engineer; **Contributor:** data-engineer (kyverno) -- **Must-haves:** - - `adapters/terraform/adapter.py:117` `state_bucket = - f"acdl-tfstate-{account_id}-us-east-1"` → `f"nova-tfstate-{account_id}-us-east-1"`. - - `adapters/kyverno/policies/require-resource-labels.yml`: annotation - title `Require ACDL Resource Labels` → `Require Nova Resource Labels`; - rule names `require-acdl-owner-label`/`require-acdl-environment-label` - → `require-nova-owner-label`/`require-nova-environment-label`; - messages + patterns `acdl:owner`/`acdl:environment` → `nova:owner`/ - `nova:environment`. - - Update any test fixtures referencing the old bucket name / label keys. -- **Verify:** `terraform validate` (adapter-emitted); pytest passes; - `run_ci.sh` exits 0; regression gate 22/22 (run at P9, but P1 must not - break any cap locally). - -### Phase P2 — user-facing-acdl-to-nova-sweep (REQ-166) -- **Lead:** lead-developer; **Contributor:** backend-engineer -- **Must-haves:** - - `core/environment_check.py:59,61` onboarding message header/body - "ACDL" → "Nova". - - `core/lambda/contract_ingestor.py:145` alert title `[ACDL-ALERT]` → - `[NOVA-ALERT]`; `:191` issue body "ACDL platform Lambda" → "Nova - platform Lambda". - - `scripts/post_stage_comment.sh:39` PR comment header "ACDL Stage" → - "Nova Stage"; `:46` footer "ACDL deploy pipeline" → "Nova deploy - pipeline". - - `scripts/run_ci.sh:39` CI banner "ACDL CI Pipeline" → "Nova CI - Pipeline". - - Module docstrings: `core/contract_resolver.py:1,474`, - `core/confidence_signal.py:1`, `adapters/terraform/adapter.py:1`, - `adapters/kyverno/kyverno_adapter.py:1`, `adapters/wiz/wiz_adapter.py:1`, - `adapters/README.md:1`, `adapters/kyverno/README.md:4,18` → Nova. - - Update tests that assert these strings. -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P3 — dead-code-and-stale-prefix-cleanup (REQ-167) -- **Lead:** lead-developer -- **Must-haves:** - - `scripts/run_platform.sh:153` remove the dead - `export ACDL_ENVIRONMENT_OVERRIDE=...` line (comment says "removed - in P5" but the line is present). - - Stale dual-read comments: drop the "ACDL_* fallback until P5" / - "dual-read NOVA_* first, ACDL_* fallback per G-106" comments in - `core/local_emulators.py:15-16,503,505`, - `core/regression_verify.py:318-319,333`, and the lifecycle scripts - (the G-106 fallback is retired per `core/env.py:4-5`). - - `acdl_*` temp-dir prefixes → `nova_*`: `core/local_emulators.py:71,252` - (`acdl_outbox_`/`acdl_tfstate_`), `core/regression_verify.py:183,234` - (`acdl_regr_`/`acdl_outbox_`), `scripts/run_pattern_plan.sh:29`, - `scripts/run_primitive_plan.sh:29`, `scripts/run_lifecycle_test.sh:41`, - `scripts/run_lifecycle_destroy.sh:36`. - - `core/regression_verify.py:214` interpolation fixture `acdl-` → `nova-` - (or make it a clearly-generic token). -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P4 — migrate-ssm-except-narrowing (REQ-168) -- **Lead:** backend-engineer -- **Must-haves:** - - `scripts/migrate_ssm_paths.py:113` `except Exception: pass` → - narrow to `ParameterNotFound` + structured log on the non- - ParameterNotFound path. - - Narrow `core/output_publisher.py:112,182` `except Exception` → - specific `(ClientError, OSError)` + structured stderr log. - - Test that a non-ParameterNotFound error is raised (not swallowed). -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -## Wave 2 — Simplify Without Regressions (P5–P9) - -### Phase P5 — regression-verify-dedup (REQ-169) -- **Lead:** backend-engineer -- **Must-haves:** - - Extract `_check_live_terraform_plan(contract_path, label)` from the - two ~95% identical methods `_check_live_terraform_plan_microservice` - + `_check_live_terraform_plan_static_assets` (~35 lines saved). - - Extract `_check_resolver(contract_path)` from - `_check_resolver_static_assets` + `_check_resolver_microservice`. - - Extract `_assert_contracts_resolve(module_dir)` from the duplicated - lifecycle-contract-resolve block in - `_check_lifecycle_module_terraform` + `_check_lifecycle_l2_module`. - - Behavior preserved (the regression gate output is unchanged). -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P6 — run-platform-deadcode-and-hitl-fn (REQ-170) -- **Lead:** lead-developer -- **Must-haves:** - - Extract the duplicated HITL attestation block (`:336-350` + `:452-466`) - into a shell function `run_hitl_gate()` invoked at both sites (~14 - lines saved). - - `scripts/run_platform.sh:145` hardcoded `CONTRACT_ID` UUID → - `NOVA_CONTRACT_ID` env with the existing UUID as default. - - `scripts/run_platform.sh:146` `WORK="/tmp/acdl_platform_run_v18"` → - `WORK="${NOVA_WORK_DIR:-/tmp/nova_platform_run}"` (drop the stale - `v18` stamp + `acdl_` prefix). - - Drop the stale brand comment `run_platform.sh:2` "the ACDL platform - pipeline" → "the Nova platform pipeline". -- **Verify:** pytest passes; `run_ci.sh` exits 0; `run_platform.sh - --check-only` exits 0. - -### Phase P7 — contract-resolver-envloader-and-kind (REQ-171) -- **Lead:** backend-engineer -- **Must-haves:** - - `core/contract_resolver.py:50-68` `_load_env` → import - `core/environment_check.py:load()` (dedup; both load + placeholder - warning). - - Add a `kind` field (`"l1"` / `"l2"`) to each `modules/registry.json` - entry; the resolver reads `kind` directly instead of the fragile - `is_l2 = "l2" in interface_path or "composition" in interface_path` - heuristic (`contract_resolver.py:540`). - - Collapse the redundant `kind` computation (`:584-589`) → - `kind = "l2" if (multi_module or any_l2) else "l1"` (after the - registry `kind` field is authoritative, simplify further). -- **Verify:** pytest passes; `run_ci.sh` exits 0; resolver behavior - unchanged (all contracts still resolve to the same stacks). - -### Phase P8 — workflow-generator-dedup (REQ-172) -- **Lead:** lead-developer; **Contributor:** backend-engineer (test) -- **Must-haves:** - - Author `scripts/sync_workflows.py` — reads one source workflow per - pair (e.g. `workflows-src/ci.yml`, `workflows-src/deploy.yml`, - `workflows-src/modules-lifecycle.yml`) and writes byte-identical - copies to both `.gitea/workflows/` and `.github/workflows/`. - Establish the `workflows-src/` dir as the single source. - - Replace the byte-identity assertions in - `tests/test_pipeline_contract.py` with a "generated outputs match - committed files" test (run `sync_workflows.py --check` → exit 0 if - the committed files match the generated output, non-zero + diff if - drift). - - Migrate the 3 existing pairs to the `workflows-src/` source; remove - the hand-maintained duplicates (the generator owns them). -- **Verify:** `python3 scripts/sync_workflows.py --check` exits 0; - pytest passes; `run_ci.sh` exits 0; the 4 GitHub-only workflows are - untouched (they have no pair). - -### Phase P9 — run-platform-split (REQ-173) -- **Lead:** lead-developer -- **Must-haves:** - - Extract the decommission block (`scripts/run_platform.sh:180-237`) - into `scripts/run_decommission.sh` (sourced or invoked). - - Extract the uptime block (`:520-606`) into `scripts/run_uptime.sh`. - - `run_platform.sh` invokes the helpers; behavior unchanged. - - **G-112 binding:** the helpers are **`source`d** (shared shell env), - not invoked as subshells — the extracted blocks reference - `run_platform.sh`-local vars (`NOVA_CONTRACT_ID`/`NOVA_WORK_DIR` from - P6); a subshell would not inherit them. - - **G-111 binding:** update `core/regression_verify.py` CAP-015/016 - checks — when the live resource is absent - (`ResourceNotFoundException`/`404`), mark `Skipped (post-teardown, - D-096)` not `Decayed`, so a clean local run reports 20/20 Verified + - 2 Skipped (not a strict-`all` failure on the known teardown state). - - **Run the regression gate (D-118, end of Wave 2):** **20/22 Verified** - is the passing bar (CAP-015/016 Skipped — post-v1.11-teardown steady - state, D-096; re-provisioning is a future feature, not an NFR). Any - non-Verified/non-Skipped capability halts Wave 3. -- **Verify:** pytest passes; `run_ci.sh` exits 0; `run_platform.sh - --check-only` exits 0; **regression gate 20/22 Verified + 2 Skipped**. - -## Wave 3 — Security + Maintainability (P10–P14) - -### Phase P10 — contract-ingestor-defense-in-depth (REQ-174) -- **Lead:** backend-engineer; **Contributor:** lead-developer (review) -- **Must-haves:** - - `core/lambda/contract_ingestor.py:251-252` `if not caller_arn: pass` - → fail closed: return a 401/403 with a clear message when IAM identity - is absent (defense-in-depth; ABAC layer still the primary control). - - `core/lambda/contract_ingestor.py:269` hardcoded - `valid_envs = {"dev","qa","prod","dr"}` → derive from the - `core/environments/` directory (list `*.json` filenames). - - Document the ABAC reliance explicitly in the function docstring + - ARCHITECTURE.md. - - Test: a request without IAM identity is rejected; a request with an - unknown environment is rejected. -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P11 — contract-ingestor-payload-validation (REQ-175) -- **Lead:** backend-engineer -- **Must-haves:** - - `submit_contract`: size-cap the `contract` blob (e.g. 256 KB) before - the DynamoDB write; reject oversized payloads with 413. - - Schema-validate the contract blob against `schemas/contract.schema.json` - before the write; reject invalid with 400. - - Consistent caps: `error` and `stackTrace` use the same cap (align the - 10k vs 2k inconsistency). - - Tests for size-limit + schema-rejection paths. -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P12 — split-contract-resolver (REQ-176) -- **Lead:** backend-engineer -- **Must-haves:** - - Split `core/contract_resolver.py` (638 lines) into: - `core/contract_resolve.py` (the resolve + interpolation core), - `core/decommission_transform.py` (the decommission zero-counts - transform), `core/contract_resolver_cli.py` (the `__main__` CLI). - - `core/contract_resolver.py` becomes a thin re-export shim for - backwards compat (existing imports keep working). - - **G-113 binding:** import direction is one-way — split modules - import only each other + stdlib; the re-export shim imports the - split modules; nothing imports the shim except external callers - (prevents the latent cycle shim → split → split → shim). - - Behavior unchanged; all tests pass without modification. -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P13 — split-regression-verify (REQ-177) -- **Lead:** backend-engineer -- **Must-haves:** - - Split `core/regression_verify.py` (670 lines) into: - `core/regression_capabilities.py` (the CAP-001..022 checks), - `core/regression_live_plan.py` (the shared live-plan helpers from - P5), `core/regression_verify_cli.py` (the `__main__` CLI + - `run_regression` orchestration). - - `core/regression_verify.py` becomes a thin re-export shim. - - Behavior unchanged; the regression gate output is identical. -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P14 — schema-driven-outputs-and-cache (REQ-178) -- **Lead:** backend-engineer; **Contributor:** data-engineer (interface.json) -- **Must-haves:** - - `core/output_publisher.py:38-55` `SAFE_OUTPUT_NAMES` hardcoded set → - derived from `modules/l1/*/interface.json` `outputs[].sensitive` - annotations (non-sensitive outputs are safe to publish). - - `core/contract_resolver.py:498,617` (now in the split module) — - cache loaded JSON schemas in a module-level dict (avoid re-reading - from disk each resolve call). - - **Mid-milestone checkpoint (offline):** regression gate spot-check - (not the full P9/P21 gate); confirm Wave 3 introduced no regressions. -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -## Wave 4 — Developer Experience (P15–P17) - -### Phase P15 — run-platform-help-and-flags-doc (REQ-179) -- **Lead:** lead-developer -- **Must-haves:** - - `scripts/run_platform.sh` add a real `--help` / `-h` flag that - prints all flags + a one-line description each (`--check-only`, - `--plan-only`, `--apply`, `--destroy`, `--quiet`, `--deploy-uptime`, - `--decommission`, `--local`, `--environment`). The current `:82` - reject-unknown-flags path must allow `--help` to print + exit 0. - - Document `--deploy-uptime` in the header comment block (currently - used at `:532` but absent from the header). - - Surface `--local` (D-092 local emulating tier) in the README "How to - run" section. -- **Verify:** `run_platform.sh --help` exits 0 and lists all flags; - pytest passes; `run_ci.sh` exits 0. - -### Phase P16 — workflows-readme-catalog (REQ-180) -- **Lead:** lead-developer -- **Must-haves:** - - Author `.github/workflows/README.md` cataloging all 7 workflows: - `ci.yml`, `deploy.yml`, `platform-test.yml`, `primitives-plan.yml`, - `patterns-plan.yml`, `release.yml`, `modules-lifecycle.yml`. For - each: trigger (`on:`), inputs (reusable-workflow `workflow_call` - inputs), required secrets, and one-line purpose. - - Note which 3 are byte-identical Gitea mirrors (post-P8, generated by - `sync_workflows.py`) and which 4 are GitHub-only (Gitea act_runner - feature gaps). - - Add a `tests/test_docs_coverage.py` assertion that the README exists - + lists all 7 workflow filenames. -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P17 — getting-started-consolidation (REQ-181) -- **Lead:** lead-developer -- **Must-haves:** - - Consolidate the README "How to run" into a single getting-started - section: **offline happy path first** (`bash scripts/run_ci.sh` + - `bash scripts/run_platform.sh --check-only` / `--local` — no AWS - needed), then the **AWS path** (bootstrap + `--apply`). - - Remove the fragmented 3-step bootstrap as the lead; demote it to - the AWS-path subsection. - - Cross-link `docs/CONSUMER_GUIDE.md` for the consumer contract model. -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -## Wave 5 — No Humans Onboarding Flow (P18–P20) - -### Phase P18 — onboarding-schema-and-lambda-action (REQ-182) -- **Lead:** backend-engineer; **Contributor:** lead-developer (schema) -- **Must-haves:** - - Author `schemas/onboarding.schema.json` (JSON Schema draft 2020-12): - required fields `consumerRepo` (string, format), `requestedEnvironment` - (string, enum from environments dir), `ownerId` (string), `billingTag` - (string); optional `notes`. - - `core/lambda/contract_ingestor.py` add an `onboard_consumer` action - (D-119): validates the payload against the onboarding schema, writes - a `pending` row to `nova-contracts` (PK `consumerRepo`, SK - `onboarding##`, status `pending`). - No AWS resources created (D-113). - - Tests: valid onboarding request writes a pending row; invalid request - rejected with 400; offline-testable via moto/local Lambda stub. -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P19 — onboarding-envfile-autogen (REQ-183) -- **Lead:** backend-engineer; **Contributor:** lead-developer (docs) -- **Must-haves:** - - Author `core/onboarding.py` with `generate_env_file(request, - template_env="dev")` — produces a `.json` from a consumer - onboarding request (fills `account_id` placeholder, `ownerId`, - `billingTag` into the env template). Emits the file + a git patch / - PR-branch instruction. - - Rebrand `core/environment_check.py:57-77` onboarding message to - Nova; replace the "1. Contact the platform team" handoff with the - self-service request path: "Run `nova onboard` (or POST to the - Lambda `onboard_consumer` action) to request an environment; the - platform generates a binding + opens a PR." - - Update `core/environments/README.md:34-37` — self-service request - path is now implemented (real provisioning still a future feature). - - Tests: `generate_env_file` produces a valid env JSON; the rebranded - message no longer says "contact the platform team". -- **Verify:** pytest passes; `run_ci.sh` exits 0. - -### Phase P20 — cross-account-role-automation-offline (REQ-184) -- **Lead:** data-engineer; **Contributor:** backend-engineer (ABAC) -- **Must-haves:** - - Author `terraform/onboarding/` (new dir): `main.tf` defining the - consumer deploy-role + `nova:owner` ABAC tag grant (cross-account - IAM role + trust policy + tag-based permission boundary). Variables - for `consumer_repo`, `owner_id`, `account_id`. - - `terraform validate` passes; `terraform plan` (offline / no live - apply per D-114) produces the expected role + policy. - - Document the onboarding Terraform in `docs/ONBOARDING.md` — the - request path (P18) → env-file autogen (P19) → role grant (P20, this - phase, offline-proven; live apply deferred). - - Tests: `terraform validate` for the onboarding module; a - `test_onboarding_terraform.py` asserting the module validates. -- **Verify:** `terraform validate` (onboarding module) passes; pytest - passes; `run_ci.sh` exits 0. - -## Final Phase — P21 — final-review-ship - -- **Lead:** lead-developer; **Contributors:** all active (review) -- **Must-haves:** - - Multi-persona code review across all v1.16 phases (ci-code-reviewer). - Auto-apply P0 fixes; flag P1+ for post-hoc review. If P1+ found, fix - in this phase (not loop back to EXECUTE). - - Audit (ciagent-audit): reconstruction test (git log matches - `.ciagent/` files), file discipline, branch hygiene, commit - discipline. Fix critical issues in this phase. - - **Run the regression gate (D-118, milestone complete):** **20/22 - Verified** (CAP-015/016 Skipped — post-teardown steady state, D-096). - - Update `.ciagent/REQUIREMENTS.md` — mark REQ-165..184 complete. - - Update `.ciagent/ROADMAP.md` — mark v1.16 complete. - - Update `.ciagent/PROJECT.md` — v1.16 complete summary. - - Ship: merge `phase/21-final-review-ship` → - `milestone/v1.16-nova-simplification`; merge milestone → `main`; - tag `v1.15.26` (= milestone release); create Gitea release with full - milestone summary. - - Clear CHECKPOINT.json (milestone complete). - -## Success Criteria (milestone gate) - -- All 20 requirements (REQ-165..184) satisfied; 0 partial. -- Regression gate **20/22 Verified + 2 Skipped** at P9 + P21 (D-118, - G-111; CAP-015/016 are the post-v1.11-teardown steady state, D-096). -- `bash scripts/run_ci.sh` exits 0 at every phase boundary. -- Review: 0 new P0; P1+ flagged or auto-fixed. -- Audit: clean; reconstruction test passes. -- Tag `v1.15.26` created; milestone merged to main. -- Onboarding request path implemented (P18–P20); real AWS provisioning - explicitly deferred (D-113, D-114). \ No newline at end of file +- **No fabrication:** every metric carries a grounded/derived/deferred + status with a source file or decision ID. No fabricated numbers in any + deck slide or METRICS.md entry. +- **Decision discipline:** D-120..D-132 are locked. This plan does not + re-open any locked decision. If a decision needs revisiting, it goes + through the GRILL, not the plan. \ No newline at end of file diff --git a/.ciagent/PROJECT.md b/.ciagent/PROJECT.md index 1c8a60a..2ffa95a 100644 --- a/.ciagent/PROJECT.md +++ b/.ciagent/PROJECT.md @@ -1072,4 +1072,65 @@ conversation; D-117..D-119 resolved at CLARIFY. | D-116 | Drift fixes = P1 of v1.16 (not a hotfix to main). | User chose "P1 of v1.16." The state-bucket drift (`adapter.py:117`) and Kyverno label contradiction are correctness regressions but latent in plan-only mode (no live apply in the default path), so they are not an active outage. Fixing them as P1 keeps the milestone self-contained. | P1 fixes both; no hotfix to main. | | D-117 | v1.14 NFR categories are NOT re-proposed. | v1.14 already swept over-broad excepts (REQ-141), hardcoded account-ID (REQ-142), IAM `Resource:"*"` scoping (REQ-143), contractId/env validation (REQ-144), `.gitignore` catch-all (REQ-146), `--kube-version` removal (REQ-147), orphan cleanup (REQ-148), `set -euo pipefail` parity (REQ-150). v1.16 finds NEW residual signals (the v1.15 rebrand left a fresh debt layer) and does not duplicate completed work. | Wave 1–5 target only fresh debt. | | D-118 | Regression gate (D-091) gates Wave 2 completion and P21. | "Simplify without regressions" is only credible if the regression gate runs after the simplification wave. The gate runs after P9 (Wave 2 done) and at P21 (milestone complete); any non-Verified capability halts W3. Mid-milestone checkpoint after P14 (offline). | P9 + P21 run the gate; P14 checkpoint. | -| D-119 | `onboard_consumer` action stores a CMDB row pending grant (not auto-provisions). | The request-path-only scope (D-113) means the Lambda accepts an onboarding request and writes a `pending` row to `nova-contracts` (or a new `nova-onboarding` partition key); the platform automation that grants the ABAC role is the P20 Terraform (offline-proven). No AWS resources are created by the Lambda action itself. | P18 writes the pending row; P20 proves the grant Terraform offline. | \ No newline at end of file +| D-119 | `onboard_consumer` action stores a CMDB row pending grant (not auto-provisions). | The request-path-only scope (D-113) means the Lambda accepts an onboarding request and writes a `pending` row to `nova-contracts` (or a new `nova-onboarding` partition key); the platform automation that grants the ABAC role is the P20 Terraform (offline-proven). No AWS resources are created by the Lambda action itself. | P18 writes the pending row; P20 proves the grant Terraform offline. | + +## Objective for Milestone v1.17 (active — Strategic Direction, Leadership Metrics & Unified Story) + +**Milestone type:** Feature (P1–P3 feat; P4 docs; P5 docs+test; P6 test; +P7 review+audit+ship). Tags on the v1.16.x line: `v1.16.0` (P0) → +`v1.16.1..v1.16.7` (P1–P7) → `v1.16.8` (P8 final = milestone release). + +**Three pillars:** + +- **Pillar A — Strategic Direction.** A durable, PO-authored + `.ciagent/NORTH_STAR.md` encodes the platform's vision, 4 strategic + objectives, 5 anti-goals, v1.17 non-goals, 12–18mo targets (with a + grounding column), and success criteria. CIAgent reads it in every + future `/ci-run` so the direction survives across milestones. The + attestation clarification is reflected: human attestation required at + stage gates (QA for production, SRE for operational readiness); autonomy + in operations, not in accountability. + +- **Pillar B — Leadership Metrics + PowerBI.** Instrument Nova to + collect, aggregate, and surface leadership-grade metrics that prove the + "no-humans" autonomous-infrastructure value proposition. Nova-native + minimal tech (CloudEvents 1.0 envelope, JSONL event log, SQLite cold + store, hash-chained Decision Ledger via `outbox_writer.py` extension) + + Infracost for pre-apply cost estimates. Hybrid model: existing + file-based signals (REGRESSION_REPORT.json, pcr.json, signal.json, + junit XML) are sources the collector reads and projects into events; + new emitters emit CloudEvents directly. PowerBI export = CSV/JSON + views (fact + dimension tables + 8 empty placeholder views for + deferred metrics). **Hard constraint: DO NOT make anything up.** Every + metric is `grounded` (cites source file + schema), `derived` + (documented formula), or `deferred` (cites decision ID — D-096/D-083/ + D-113/D-114/D-119). The 8 deferred metrics: drift detection, GreenOps/ + carbon, predictive/reactive, live CUR reconciliation, multi-cloud, + red-team MTTR, self-healing velocity, SLA/downtime. + +- **Pillar C — Unified Narrative Deck.** Merge the two existing decks + (`how-the-platform-works` + `the-developer-experience`) into one unified + narrative deck "Nova — The No-Humans Infrastructure Platform" with a + single arc: Problem → Vision/Direction (NORTH_STAR) → How it works → + Proof (metrics) → Roadmap/Ask. The "tell them x3" structure applies at + deck level AND per slide (each slide opens with what it covers, + delivers, closes with an explicit "benefit of this stage" callout). + Fluid transitions between slides. Both old decks retired. + +**Key decisions resolved in the planning conversation (D-120+):** + +| ID | Decision | Rationale | Outcome | +|----|----------|-----------|---------| +| D-120 | Tech stack = Nova-native + Infracost, drift deferred. | The PO's technical-direction document specifies Kafka/Prometheus/ClickHouse/QLDB/OTel — none exist in Nova today. Adopt the PRINCIPLES (events as source of truth, CloudEvents envelope, decision ledger, definition-of-success docs, dashboards-as-projections) but implement with Nova-native minimal tech (JSONL + SQLite + hash-chained ledger). No Kafka/Prometheus/ClickHouse/QLDB. Infracost adopted (runs offline on plan JSON). Drift detection deferred (D-096 + no scheduler). | P1–P3 use Nova-native tech; Infracost in P1; drift deferred. | +| D-121 | Decision Ledger = extend outbox_writer.py → SQLite append-only hash chain. | The direction's #1 priority is the Decision Ledger. Nova already has a hash-chained outbox (outbox_writer.py). Extend it to a SQLite append-only table with hash chain; add ai.decision.made + attestation.recorded events. Honors D-083 (no S3 Object Lock/JWS). | P1 extends outbox_writer; ledger is SQLite hash-chain. | +| D-122 | AI Planner framing = map Nova's real decision points. | The direction assumes an "AI Planner/Reasoner" (planner-v3.2). Nova's actual decision path is confidence_signal + HITL gate. Model ai.decision.made from confidence_signal (decision_id=run_id, chosen_action=band, confidence=score, alternatives=perInput, human_override=HITL block). LLM planner marked future/aspirational. | P1 emits honest decision events; no fabricated LLM. | +| D-123 | Deferred metrics = all 8 (drift, GreenOps, predictive/reactive, live CUR, multi-cloud, red-team MTTR, self-healing, SLA/downtime). | These require live AWS (D-096) or new external systems. Ship as empty PowerBI placeholder views with documented schemas. | P3 ships 8 placeholder views; METRICS.md marks them deferred. | +| D-124 | NORTH_STAR = strategy; tech direction = engineering input. | The PO's technical-direction document is engineering architecture, not strategy. NORTH_STAR.md captures strategic vision/objectives/anti-goals (PO-authored). The tech direction becomes the telemetry reference architecture section in RESEARCH.md/ARCHITECTURE.md, cited by NORTH_STAR's engineering objectives. | P0 writes NORTH_STAR; RESEARCH writes the telemetry reference. | +| D-125 | Events vs files = hybrid. | Existing file-based signals (REGRESSION_REPORT.json, pcr.json, signal.json, junit) stay as files; the collector reads them and emits normalized CloudEvents into JSONL + SQLite. New emitters emit CloudEvents directly. | P2 collector reads files + events. | +| D-126 | Hot/cold split = cold-only SQLite (hot path deferred). | Nova has no live ops dashboard (no live AWS, D-096). The SQLite store is cold-only (batch/historical). The hot path is documented as deferred. | P2 SQLite is cold-only. | +| D-127 | Definition-of-success = per-KPI docs. | The direction's §11 requires a definition-of-success doc for every executive KPI. Adopt this standard; docs live in `docs/metrics/`. | P4 writes per-KPI docs. | +| D-128 | Storage location = metrics/ at repo root. | metrics/runs/ (per-run manifests), metrics/nova_metrics.db (SQLite), metrics/events.jsonl (event log), metrics/powerbi/ (export). | P1–P3 use metrics/ at repo root. | +| D-129 | PowerBI delivery = CSV/JSON files, folder connector. | Nova is offline-first; no live connector to a running service. PowerBI ingests via the folder connector. | P3 emits CSV/JSON to metrics/powerbi/. | +| D-130 | Deck arc = Problem → Vision → How → Proof → Roadmap. | The unified narrative deck's 5-act structure. x3 arc at deck + slide level. Per-slide benefit callouts. Fluid transitions. Both old decks retired. | P5 builds the unified deck; old decks deleted. | +| D-131 | MTTR scope = platform-run MTTR. | The <60s MTTR target refers to platform-run failures (apply.failed → successful retry), not infra-incident MTTR (no incident detection system). Infra-incident MTTR deferred. | P4 grounds platform-run MTTR. | +| D-132 | Attestation instrumentation = emit attestation.recorded events. | The attestation system (hitl_gates.py + attestation_matrix.py + separation_of_duties.py) already exists. Instrument it: emit attestation.recorded events into the Decision Ledger + PowerBI. Attestation Coverage = 100% target grounded from outbox approver_* attributes. | P1 emits attestation events; P4 grounds Attestation Coverage. | \ No newline at end of file diff --git a/.ciagent/REQUIREMENTS.md b/.ciagent/REQUIREMENTS.md index 60ba2c1..2274a2a 100644 --- a/.ciagent/REQUIREMENTS.md +++ b/.ciagent/REQUIREMENTS.md @@ -956,3 +956,227 @@ simplification and the first self-service onboarding request path. scoping (REQ-143), contractId/env validation (REQ-144), `.gitignore` catch-all (REQ-146), `--kube-version` removal (REQ-147), orphan cleanup (REQ-148), `set -euo pipefail` parity (REQ-150). + +## v1.17 — Strategic Direction, Leadership Metrics & Unified Story + +**Milestone type:** Feature (P1–P3 feat; P4 docs; P5 docs+test; P6 test; +P7 review+audit+ship). Progressive patches; the final phase's patch IS +the milestone release. Tags run on the v1.16.x line: `v1.16.0` (P0) → +`v1.16.1..v1.16.7` (P1–P7) → `v1.16.8` (P8 final = milestone release). + +**Objective:** Three pillars. (A) Encode the PO's strategic direction in +a durable `NORTH_STAR.md` read by CIAgent in every future `/ci-run`. +(B) Instrument Nova to collect, aggregate, and surface leadership-grade +metrics that prove the "no-humans" autonomous-infrastructure value +proposition — grounded in signals Nova actually emits, derived via +documented formulas, or explicitly deferred with a decision ID — flowing +into PowerBI-ready views. (C) Merge the two existing decks into one +unified narrative deck with the "tell them x3" arc at deck + slide level, +per-slide benefit callouts, and fluid transitions. + +**Hard constraint:** DO NOT make anything up. Every metric carries a +`grounded` / `derived` / `deferred` status with a source file or +decision ID. Deferred metrics ship as empty PowerBI placeholder views +with documented schemas. + +### Requirements + +**Pillar A — Strategic Direction** + +- **REQ-185** — `.ciagent/NORTH_STAR.md` is PO-authored with Vision, + Strategic Objectives (4), Anti-Goals (5), Non-Goals (v1.17 scope), + 12–18mo Targets (with grounding column), and Success Criteria. The + attestation clarification is reflected: human attestation required at + stage gates (QA for production, SRE for operational readiness); + autonomy in operations, not in accountability. (Phase P0) +- **REQ-186** — CIAgent reads `NORTH_STAR.md` in context-loading for all + future milestones; the file is referenced from PROJECT.md and + ARCHITECTURE.md so the strategic direction survives across milestones. + (Phase P4) + +**Pillar B — Leadership Metrics + PowerBI** + +- **REQ-187** — Event emitters: a CloudEvents 1.0 envelope is adopted; + a per-run manifest writer emits structured events (run_id, contractId, + env, stages×durations, exit, confidence, HITL block count) to + `metrics/runs/`; existing ephemeral `$WORK/*.json` (pcr, signal, + event, outbox, stack) are persisted as durable artifacts; pytest + `addopts` gains `--junitxml`+`--json-report`; Infracost runs as a + plan post-processor emitting `cost.estimated{delta_usd}` (offline). + (Phase P1) +- **REQ-188** — Decision Ledger: `outbox_writer.py` is extended to emit + to a SQLite append-only table with hash chain; `ai.decision.made` + events are modeled from Nova's real decision points (decision_id=run_id, + chosen_action=band outcome, confidence=score, alternatives=perInput + breakdown, human_override=HITL block) with outcome backfill from + apply.completed; `attestation.recorded` events capture qa/prod/dr + sign-offs (approver, env, concerns, result). Honors D-083 (no S3 Object + Lock/JWS). (Phase P1) +- **REQ-189** — Metrics collector: `core/metrics/collector.py` + + `schemas/metrics_*.schema.json` read all grounded signals + (REGRESSION_REPORT.json, per-run manifests, junit XML, pcr.json, + signal.json, COST.md, decision ledger) → normalized SQLite cold store + at `metrics/nova_metrics.db`; idempotent re-runs. (Phase P2) +- **REQ-190** — PowerBI export: `core/metrics/powerbi_export.py` emits + CSV/JSON views to `metrics/powerbi/` (fact_run, fact_capability, + fact_policy_check, fact_confidence, fact_test, fact_decision, + fact_cost_estimate, dim_capability, dim_milestone + 8 empty + placeholder views for deferred metrics with documented schemas) + + `docs/METRICS_VIEWS.md` schema doc. (Phase P3) +- **REQ-191** — Zero-touch efficiency metrics: Autonomous Resolution + Rate (runs without operational HITL block ÷ total; attestation gates + excluded), Human Escalation Frequency (operational HITL blocks only), + AI Decision Accuracy (decisions not followed by apply.failed/incident + within 5min), MTTD/MTTR (platform-run: apply.failed → successful + retry). (Attestation Coverage is owned by REQ-194, not here.) + (Phase P4) +- **REQ-192** — Velocity metrics: Provisioning Lead Time + (apply.completed.time − intent.received.time), Deployment Frequency + (count(apply.completed) per day). Self-Healing Velocity deferred (no + auto-remediator). (Phase P4) +- **REQ-193** — Financial & cost-ROI metrics: FTE Hours Saved (derived: + run count × manual baseline), Cost Savings via Infracost estimates + (grounded), Cost Efficiency Ratio (derived), Platform ROI (derived + formula). Live CUR reconciliation deferred (D-096). (Phase P4) +- **REQ-194** — Reliability, security & compliance metrics: Zero-Trust + Policy Compliance Rate (from pcr.json), Attestation Coverage (prod/dr + promotions attested by a human ÷ total prod/dr promotions; grounded in + hitl_gates.py + outbox approver_* attributes; canonical owner of this + metric). Uptime, Patch Remediation, SLA/downtime deferred (D-096). + (Phase P4) +- **REQ-195** — Metrics catalog doc: `docs/METRICS.md` catalogs every + executive KPI with `grounded`/`derived`/`deferred` status, source + file or decision ID, and a per-KPI definition-of-success doc in + `docs/metrics/.md`. (Phase P4) + +**Pillar C — Unified Narrative Deck** + +- **REQ-196** — The two existing decks (`how-the-platform-works` + + `the-developer-experience`) are merged into one unified narrative deck + "Nova — The No-Humans Infrastructure Platform" with a single arc: + Problem → Vision/Direction (NORTH_STAR) → How it works → Proof + (metrics) → Roadmap/Ask. The x3 structure ("tell them what you're + going to tell them → tell them → tell them what you told them") applies + at deck level (opening = arc; body = tell them; closing = recap + ask). + Both old decks are retired (all derived artifacts deleted). (Phase P5) +- **REQ-197** — Each slide has the x3 structure (opens with what it + covers, delivers, closes with an explicit "benefit of this stage" + callout) + fluid transitions between slides (no disjointed jumps). + The 4-step deck process (source `.md` → Marp → HTML → talking-points) + is re-run for the unified deck. (Phase P5) + +**Cross-cutting** + +- **REQ-198** — Regression capability: CAP-023 (metrics collector runs, + emits expected schema) + CAP-024 (deck structure: slide count, x3 + present, per-slide benefit present) added to `core/regression_verify.py`. + (Phase P6) + +**Ideation enhancements (REQ-199..213 — additive, within D-120..D-132)** + +- **REQ-199** — Metrics schema validation in CI: `run_ci.sh` validates + `metrics/powerbi/*.json` + a sample `metrics/events.jsonl` against + their schemas; exits 0. (Phase P3) +- **REQ-200** — Idempotent collector re-run test: `test_metrics_collector_idempotent` + passes (two runs → identical row counts + chain verified). (Phase P2) +- **REQ-201** — Metrics store backup/restore doc: `metrics/README.md` + documents regenerable vs append-only artifacts + restore procedure. + (Phase P2) +- **REQ-202** — Metrics glossary appendix slide: the unified deck has a + "Metrics Glossary" appendix slide with one-line KPI definitions + + grounding badges. (Phase P5) +- **REQ-203** — "What's Deferred — and Why" slide: the unified deck has + a slide pairing each of 8 deferred metrics with its blocking decision + ID. (Phase P5) +- **REQ-204** — NORTH_STAR diff-check in CI: `run_ci.sh` includes + `check_north_star_diff` that fails when Vision/Objectives/Anti-Goals/ + Targets sections change without a `NORTH_STAR-CHANGE:` commit trailer. + (Phase P4) +- **REQ-205** — Per-module lifecycle success-rate report: each lifecycle + run writes `metrics/lifecycle/-.json`; collector projects + into `fact_lifecycle`; PowerBI "Module Lifecycle Health" view. (Phase + P1 emitter + P2 collector + P3 view) +- **REQ-206** — Code coverage trend emission: `pyproject.toml` addopts + gains `--cov=core --cov=adapters --cov-report=json:metrics/coverage.json`; + collector ingests; `fact_test` carries a coverage column. (Phase P1 + + P2) +- **REQ-207** — Decision Ledger CLI: `core/metrics/decision_ledger_cli.py` + supports `query`, `verify-chain`, `stats`, `export`, `replay`; + `verify-chain` detects broken hashes; `replay` prints ordered events; + tests pass offline. (Phase P2) +- **REQ-208** — PowerBI starter dashboard README: `metrics/powerbi/NOVA_DASHBOARD_README.md` + documents folder-connector import + starter visual model + reference + screenshot. (Phase P3) +- **REQ-209** — PowerBI column-level data dictionary: `docs/METRICS_VIEWS.md` + has a per-column data-dictionary table (column, type, source/formula, + unit, grounded/derived/deferred status). (Phase P3/P4) +- **REQ-210** — Deferred-metrics activation roadmap: `docs/METRICS_DEFERRED_ROADMAP.md` + lists 8 deferred metrics + onboarding-grant half with {blocking + decision, unblock requirement, candidate milestone} + a "Hot-Path + Activation (post-D-096)" section (Nova-native only, D-120) + + "Re-evaluation Triggers" section. (Phase P4) +- **REQ-211** — Trust-snapshot report: `core/metrics/trust_snapshot.py` + emits `metrics/TRUST_SNAPSHOT.md` with 5 trust metrics (Decision Ledger + Coverage, Attestation Coverage, Capability Health, AI Decision + Accuracy, Confidence-Gate Halt Rate) + chain-integrity verdict + + snapshot hash; runs offline. (Phase P4) +- **REQ-212** — Confidence-Gate Halt Rate metric: `docs/METRICS.md` + + trust snapshot include "Confidence-Gate Halt Rate" (signal.json + band=halt ÷ total runs); PowerBI view includes it. (Phase P4) +- **REQ-213** — "No-humans" thesis defensibility brief: `docs/NO_HUMANS_THESIS.md` + defines the thesis, grounded proof metrics, deferred proof metrics, + and explicit anti-claims (incl. D-122 honesty); the unified deck's + Vision act cites it. (Phase P4/P5) + +### v1.17 Traceability + +| Requirement | Phase | Status | +|-------------|-------|--------| +| REQ-185 | P0 | in_progress | +| REQ-186 | P4 | pending | +| REQ-187 | P1 | pending | +| REQ-188 | P1 | pending | +| REQ-189 | P2 | pending | +| REQ-190 | P3 | pending | +| REQ-191 | P4 | pending | +| REQ-192 | P4 | pending | +| REQ-193 | P4 | pending | +| REQ-194 | P4 | pending | +| REQ-195 | P4 | pending | +| REQ-196 | P5 | pending | +| REQ-197 | P5 | pending | +| REQ-198 | P6 | pending | +| REQ-199 | P3 | pending | +| REQ-200 | P2 | pending | +| REQ-201 | P2 | pending | +| REQ-202 | P5 | pending | +| REQ-203 | P5 | pending | +| REQ-204 | P4 | pending | +| REQ-205 | P1+P2+P3 | pending | +| REQ-206 | P1+P2 | pending | +| REQ-207 | P2 | pending | +| REQ-208 | P3 | pending | +| REQ-209 | P3/P4 | pending | +| REQ-210 | P4 | pending | +| REQ-211 | P4 | pending | +| REQ-212 | P4 | pending | +| REQ-213 | P4/P5 | pending | + +### Out of Scope (v1.17) + +- Live AWS re-provisioning (D-096) — metrics requiring live + infrastructure ship as placeholder views. +- Onboarding auto-grant (D-113/D-114/D-119) — only the request-path + metric is grounded. +- ML anomaly-forecasting / predictive remediation — no emitter today; + Predictive-vs-Reactive metric ships as a placeholder. +- Drift detection scheduled job (D-096 + no scheduler) — drift metrics + ship as placeholders. +- Live cost CUR reconciliation (D-096) — Infracost pre-apply estimates + are grounded; actuals are not. +- S3 Object Lock / JWS tamper-evident ledger (D-083) — Decision Ledger + uses a local SQLite hash-chain this milestone. +- Multi-cloud support (Azure/GCP/K8s) — Nova is AWS-only this milestone. +- A third deck — the two existing decks merge into one; no new + standalone metrics deck. +- A Nova web UI — dashboards are PowerBI, not a Nova-built frontend. diff --git a/.ciagent/RESEARCH.md b/.ciagent/RESEARCH.md index 8843f26..755aed5 100644 --- a/.ciagent/RESEARCH.md +++ b/.ciagent/RESEARCH.md @@ -1188,3 +1188,308 @@ stays a future feature (D-113). - A4 (0.85): The regression gate (D-091, D-118) at P9 and P21 confirms "simplify without regressions" — 22/22 capabilities must stay Verified. The gate is the credible control for the simplification wave. + +--- + +# v1.17 Research — Strategic Direction, Leadership Metrics & Unified Story + +> Phase: research (P0). Milestone: v1.17. Status: research. +> Researcher: ci-researcher + explore agent (signal inventory). +> Autonomy: full. Decisions D-120..D-132 locked in the planning +> conversation (PROJECT.md). NORTH_STAR.md drafted (pending GRILL). + +## 1. Telemetry Signal Inventory (grounding audit) + +**Methodology:** every claim below is grounded in a concrete file path + +line number in `/root/acdl`. No speculation. The explore agent performed +a full sweep of the repo. The finding: **Nova has no metrics/telemetry/ +dashboard aggregation layer today.** What exists is a set of discrete, +structured, file-based signal artifacts (JSON reports, JSONL logs, +hash-chained outbox events, PR comments, Checkov JSON) plus unstructured +stdout logs. A metrics milestone must aggregate these existing signals +— it must not invent new ones without first adding emitters. + +### (a) Signals that EXIST TODAY and are STRUCTURED (groundable) + +| Signal | File / Emitter | Schema | Persistent? | +|--------|---------------|--------|-------------| +| Regression report (22 caps, status, duration_ms, gate) | `.ciagent/REGRESSION_REPORT.json` ← `core/regression_verify.py:643-667` | `regression_verify.py:82-91` | **Yes** (committed file) | +| Regression report (markdown mirror) | `.ciagent/REGRESSION_REPORT.md` | same | Yes | +| Checkpoint (milestone/phase/tag/regression summary) | `.ciagent/CHECKPOINT.json` (CIAgent-managed) | ad-hoc | Yes | +| PolicyCheckResult list (per-rule pass/fail/severity/resourceRef) | `$WORK/pcr.json` ← `run_platform.sh:395` + `checkov_adapter.py:50-71` | `schemas/policy_check_result.schema.json` | **No** (ephemeral `/tmp/`) | +| Confidence signal (score, band, perInput, reasonCodes) | `$WORK/signal.json` ← `run_platform.sh:412-426` + `confidence_signal.py:60-65` | `confidence_signal.py:60-65` | No (ephemeral) | +| Outbox event (hash-chained, CONFIDENCE_COMPUTED) | `$WORK/event.json` + `$WORK/outbox_item.json` ← `run_platform.sh:444-459` + `outbox_writer.py:44-56` | `audit_ledger_design.md:44-45,81-97` | No (ephemeral; live DynamoDB torn down D-096) | +| Resolved Target Stack | `$WORK/stack.json` ← `contract_resolver.py:581-603` | `schemas/stack.schema.json` | No (ephemeral) | +| Lambda return bodies (submit/report_error/validate_cr/onboard) | `core/lambda/contract_ingestor.py:171,265,284,392,446` | ad-hoc JSON | No (Lambda not live; local stub only) | +| DynamoDB CMDB rows (submitted/pending contracts) | `nova-contracts` table ← `contract_ingestor.py:160-170,433-445` | ad-hoc | **No** (table torn down D-096) | +| SSM parameters (deploy outputs) | `/nova///` ← `output_publisher.py:123-156` | ad-hoc | No (live AWS, torn down) | +| PR stage comment (mode, runId) | GitHub PR API ← `post_stage_comment.sh:34-48` + `deploy.yml:141` | markdown table | Yes (GitHub) | +| PR deploy-outputs comment | GitHub PR API ← `output_publisher.py:159-189` | markdown table | Yes (GitHub) | +| GitHub issue (deploy failure alert) | GitHub API ← `contract_ingestor.py:179-290` + `deploy.yml:143-152` | issue body | Yes (GitHub) | +| Local E2E result (stack_name, tier, outbox_events, chain_verified, lambda_status) | stdout JSON ← `core/local_emulators.py:498-508,519` | ad-hoc | No (stdout) | +| HITL gate result | `core/hitl_gates.py:87,90` + `run_platform.sh:179-185` | stdout `HITL PASS/BLOCK` | No (stdout) | +| Attestation matrix result | `core/attestation_matrix.py:184,187` | stdout `ATTESTATION PASS/BLOCK` | No (stdout) | +| Cost figures | `.ciagent/COST.md` (manual Cost Explorer query) | markdown table | Yes (manual, not automated) | + +### (b) Signals that EXIST but are UNSTRUCTURED (log-only) + +| Signal | Source | Format | +|--------|--------|--------| +| CI pipeline result | `scripts/run_ci.sh:70-71` | stdout banner `=== CI PIPELINE OK ===` | +| Platform stage banners + summaries | `scripts/run_platform.sh:222,241,258,263,315,383,411,442,463,490,496` | stdout `=== Step N: ... ===` + summary lines | +| Terraform init/validate/plan/apply/destroy logs | `$WORK/tf-*.log` ← `run_platform.sh:320,324,328,352,375` | raw terraform stdout (via `tee`) | +| Lifecycle test results | `scripts/run_lifecycle_test.sh` etc. | exit code only (no report file) | +| Decommission step counts | `scripts/run_decommission.sh:40,54` | stdout `decommission step N: M resources...` | +| Uptime endpoint count | `scripts/run_uptime.sh:72,87` | stdout `uptime: N endpoint(s) to monitor` | +| Onboarding prompt | `core/environment_check.py:57-81` | stdout text block | +| Pytest results | `pyproject.toml:25` (`-v --tb=short`) | stdout only (no junit/json) | +| sync_workflows result | `scripts/sync_workflows.py:56,53` | stdout `OK: 3 workflow pairs match` / `DRIFT: ...` | + +### (c) Proposed executive metrics with NO grounding today (DEFERRED) + +| Proposed metric | Why no grounding | Controlling decision | +|------------------|------------------|---------------------| +| Live infrastructure health (ECS running count, ALB 5xx, RPS) | Live AWS torn down; CAP-013..016 Skipped | **D-096** | +| Live outbox write rate / ledger append latency | DynamoDB outbox table absent | **D-096** | +| Tamper-evident ledger checkpoint count / JWS signature rate | S3 Object Lock + JWS + async worker deferred | **D-083** | +| Onboarding funnel: requested → granted conversion | Only "requested" (pending row) is emitted; no grant event | **D-113, D-114, D-119** | +| Time-to-provision (onboarding SLA) | Real AWS provisioning deferred | **D-113** | +| Cross-account role grant count | Offline-proven only, no live apply | **D-114** | +| Drift detection (scheduled terraform plan -detailed-exitcode) | Needs live AWS workspaces + a scheduler Nova doesn't have | **D-096** + no scheduler | +| GreenOps / carbon (WattTime/Electricity Maps API) | No grounding; new external API | future emitter | +| Predictive vs Reactive ratio | Requires an ML anomaly-forecasting service | future emitter | +| Multi-cloud normalization (Azure/GCP/K8s, FOCUS spec) | Nova is AWS-only | future | +| Red Team MTTR | No red-team program exists | future | +| Self-healing velocity | Nova has no auto-remediator | future emitter | +| SLA / unplanned downtime | Needs live service uptime monitoring against SLOs | **D-096** | +| Per-module lifecycle success rate over time | No structured report file written; only exit code | gap (no decision) | +| Test pass rate / test count time-series | No junit/json reporter configured | gap (add `--junitxml` to addopts) | +| Code coverage trend | `pytest-cov` installed but not in `addopts` | gap | +| Deploy frequency / lead time / MTTR (DORA) | No deploy-event emitter; pipeline runs not counted | gap | +| Policy pass rate time-series | `pcr.json` emitted but ephemeral; not persisted | gap (D-096 blocks live persistence) | +| Confidence score distribution over time | `signal.json` emitted but ephemeral | gap | +| Consumer adoption count / active consumers | `PROJECT.md:487` explicitly states "0 consumer adoption today" | honest scope | +| Cost time-series (automated) | `COST.md` is a one-shot manual query; no automated emitter | gap | + +**Bottom line:** the single richest existing structured signal is +`.ciagent/REGRESSION_REPORT.json` (22 capabilities × {status, tier, +duration_ms, detail} + summary counts + boolean gate). The next richest +is the per-run `$WORK/*.json` family (pcr.json, signal.json, event.json, +stack.json) — but these are **ephemeral** and **not persisted in CI**. +The lowest-friction grounding for a "no-humans" dashboard is therefore: +(1) regression report → capability health, (2) PR comments + GitHub +issues → deploy/failure activity, (3) add `--junitxml` to pytest → test +trend, (4) persist `$WORK/*.json` → policy/confidence/outbox time-series, +(5) extend outbox_writer → Decision Ledger, (6) add Infracost → +pre-apply cost estimates. + +## 2. Telemetry Reference Architecture (Nova-native adaptation) + +The PO provided a full distributed-system telemetry reference +architecture (CloudEvents 1.0 envelope, OpenTelemetry SDK, Kafka/NATS +event bus, Prometheus hot path, ClickHouse warehouse, QLDB decision +ledger, Infracost, drift detection, ML anomaly forecasting). Per +D-120, we adopt the **principles** but implement with **Nova-native +minimal tech**. The mapping: + +| Direction's principle | Nova-native implementation (v1.17) | +|---|---| +| Events are the source of truth; dashboards are projections | Hybrid (D-125): existing file signals stay as files; collector reads them and emits normalized CloudEvents into `metrics/events.jsonl` + SQLite. New emitters emit CloudEvents directly. | +| Every AI action is logged with confidence + alternatives | Decision Ledger (D-121): `outbox_writer.py` extended → SQLite append-only hash-chain table. `ai.decision.made` modeled from confidence_signal (D-122): decision_id=run_id, chosen_action=band, confidence=score, alternatives=perInput, human_override=HITL block. | +| Hot/cold storage split | Cold-only SQLite (D-126): `metrics/nova_metrics.db`. Hot path deferred (no live ops, D-096). | +| Read-only external integrators | Infracost (pre-apply, offline, reads plan JSON). Cloud billing CUR deferred (D-096). Carbon APIs deferred (future). | +| CloudEvents 1.0 envelope | Adopted. `core/metrics/event_envelope.py` defines the envelope + `platform.*` semantic conventions. | +| Decision Ledger = append-only with hash chain + outcome backfill | SQLite append-only table with hash chain (D-121). Outcome backfilled from apply.completed via decision_id → request_id correlation. Honors D-083 (no S3 Object Lock/JWS). | +| Cost governance: mandatory tags + Infracost pre-apply | Nova already enforces `nova:*` tags (nova_tagging.py, hard mode). Infracost added as plan post-processor (D-120). Post-apply CUR deferred (D-096). | +| Definition-of-success docs for every KPI | Per-KPI docs in `docs/metrics/` (D-127). | +| Replay-ability | SQLite store + JSONL event log are replayable by design. | + +### CloudEvents envelope (Nova-native) + +```json +{ + "specversion": "1.0", + "id": "", + "source": "nova.platform", + "type": "nova.run.completed", + "time": "", + "subject": "/", + "datacontenttype": "application/json", + "platform": { + "tenant_id": "acdl", + "run_id": "run-", + "contract_id": "", + "environment": "dev|qa|prod|dr", + "actor": {"type": "confidence-gate", "id": "confidence_signal"}, + "trace_id": "" + }, + "data": { + "duration_ms": 4800, + "stages": ["resolve", "adapt", "validate", "plan", "apply"], + "exit_code": 0, + "confidence": {"score": 0.94, "band": "pass", "perInput": {...}}, + "policy": {"passed": 12, "failed": 0, "skipped": 0}, + "hitl": {"gate": "dev", "result": "autonomous", "block": false}, + "cost_estimate_usd": -12.40, + "decision_id": "run-", + "outcome": "succeeded" + } +} +``` + +### Core event types (Nova-native minimum viable set) + +| Event type | Emitted by | Purpose | Grounding | +|---|---|---|---| +| `nova.run.started` | run_platform.sh | Measures demand; provisioning lead time start | new emitter (P1) | +| `nova.run.completed` | run_platform.sh | Run count, stage durations, exit, MTTR | new emitter (P1) | +| `nova.run.failed` | run_platform.sh | Failure count, MTTR numerator | new emitter (P1) | +| `nova.policy.evaluated` | checkov_adapter.py | Policy pass rate, compliance KPIs | grounded (pcr.json → P1 persists) | +| `nova.confidence.computed` | confidence_signal.py | Confidence distribution, decision accuracy | grounded (signal.json → P1 persists) | +| `nova.ai.decision.made` | outbox_writer.py (extended) | Decision Ledger entry | grounded (D-121, D-122) | +| `nova.attestation.recorded` | hitl_gates.py | Attestation Coverage, human-in-the-loop audit | grounded (D-132) | +| `nova.cost.estimated` | Infracost post-processor | Pre-apply cost estimate | new emitter (P1, Infracost) | +| `nova.capability.verified` | regression_verify.py | Capability health, regression gate | grounded (REGRESSION_REPORT.json) | +| `nova.test.completed` | pytest (junit XML) | Test count, pass rate | new (P1 adds --junitxml) | + +## 3. Metric-to-Signal Scorecard (the "no fabrication" contract) + +| Executive metric (NORTH_STAR target) | Status | Source / formula | Decision | +|---|---|---|---| +| Touchless Resolution Rate ≥99% | grounded (after P1) | runs without operational HITL block ÷ total runs (attestation gates excluded) | D-122, D-132 | +| Human Escalation Frequency <0.1% | grounded (after P1) | operational HITL blocks ÷ total runs (attestation sign-offs excluded) | D-122, D-132 | +| MTTR (p95) <60s | grounded (platform-run) | apply.failed.time → successful retry.time | D-131 | +| Predictive vs Reactive ≥3:1 | **deferred** | requires ML forecasting (future emitter) | future | +| AI Decision Accuracy ≥99.5% | grounded (after decision ledger) | decisions not followed by apply.failed/incident within 5min | D-121, D-122 | +| Drift Auto-Reversal ≥95% | **deferred** | requires drift detection (D-096 + scheduler) | D-096 | +| Cloud Spend Reduction ≥25% | partial | pre-apply estimate grounded (Infracost); actuals deferred (D-096 CUR) | D-120 | +| L1/L2 Ops Hours Avoided ≥70% | derived | formula: run count × manual baseline minutes × blended rate | D-127 | +| Platform ROI ≥250% | derived | formula: (labor savings + cloud savings + avoided downtime) ÷ platform op cost | D-127 | +| Decision Ledger Coverage 100% | grounded (this milestone) | outbox_writer.py → SQLite hash-chain | D-121 | +| Attestation Coverage 100% | grounded | hitl_gates.py + outbox approver_* attributes; prod/dr | D-132 | +| AI-Agent Intent Share ≥40% | future | no AI-agent consumers today; placeholder view | future | +| Capability health (18V+4S) | grounded | REGRESSION_REPORT.json | existing | +| Confidence score distribution | grounded (after P1) | signal.json → decision ledger | D-121 | +| Policy pass rate | grounded (after P1) | pcr.json → persisted | D-120 | +| Test count / pass rate | grounded (after P1) | pytest --junitxml | D-120 | +| Provisioning Lead Time | grounded (after P1) | run.started → run.completed | D-120 | +| Deployment Frequency | grounded (after P1) | count(run.completed) per day | D-120 | +| Deploy-failure alert count | grounded | GitHub issues via Lambda report_error (D-055) | existing | +| Cost figures (actuals) | manual one-shot | COST.md (Cost Explorer query) | existing | +| FTE Hours Saved / TRV | derived | formula over run count + COST.md | D-127 | +| Self-healing velocity | **deferred** | no auto-remediator | future | +| SLA / unplanned downtime | **deferred** | needs live service uptime (D-096) | D-096 | +| GreenOps / carbon | **deferred** | WattTime/Electricity Maps API (future) | future | +| Red Team MTTR | **deferred** | no red-team program | future | +| Multi-cloud normalization | **deferred** | Nova is AWS-only | future | +| Live CUR reconciliation | **deferred** | needs live AWS billing (D-096) | D-096 | + +## 4. Deferred-Decision Ledger (constraints on this milestone) + +| Decision | Scope | Grounding impact | +|----------|-------|------------------| +| D-096 | Live AWS torn down post-v1.11 | BLOCKS all live-AWS metrics (CAP-013..016 Skipped; live outbox; live state bucket; live CUR) | +| D-083 | S3 Object Lock + JWS + async worker deferred | BLOCKS tamper-evident ledger; v1.17 uses SQLite hash-chain instead | +| D-113/D-114/D-119 | Onboarding = request-path only; no auto-grant | BLOCKS onboarding funnel "granted" half | +| D-091/D-118 | Regression gate (D-091) gates milestone completion | ENABLES the strongest metric signal (REGRESSION_REPORT.json) | +| D-092 | Local emulating adapters | ENABLES offline E2E metrics (CAP-011/012) | +| D-055 | report_error Lambda action creates GitHub issues | ENABLES deploy-failure alert metric | +| D-050 | Publish deploy outputs to SSM + GitHub PR comment | ENABLES outputs-published metric | +| D-054/D-043/D-109 | Nova tagging standard (hard mode) | ENABLES tagging-compliance metric | +| D-084 | 8-concern attestation matrix | ENABLES attestation metrics (operator-supplied evidence artifacts) | +| D-089 | Signature verification skipped when signing key unset (dev/CI) | Signature metrics are no-ops in dev | + +## 5. Deck-Storytelling Research (x3 arc + per-slide benefit) + +### The "tell them x3" structure + +The PO's direction: "Tell them what you're going to tell them, then tell +them, then tell them what you told them." Applied at two levels: + +**Deck level (the 5-act arc):** +1. **Opening slide** = "what I'm going to tell you" — the full arc + preview: Problem → Vision → How → Proof → Roadmap. +2. **Body** (acts 1–5) = "tell them" — each act delivers its content. +3. **Closing slide** = "what I told you" — recap of the 5 acts + the ask. + +**Per slide:** +1. **Slide opens** with what it'll cover (1 line: "This slide shows X"). +2. **Slide delivers** the content (bullets, diagram, or table). +3. **Slide closes** with an explicit **"benefit of this stage" callout** + (1 line: "Benefit: you now know Y" or "Why this matters: Z"). + +### Fluidity conventions + +- **Transitions are written, not hand-waved.** Each slide's opening line + references the previous slide's close ("Having seen X, now consider Y"). +- **No disjointed jumps.** If a topic shift is needed, a bridge slide or + a transition sentence carries the audience across. +- **The arc is visible.** A small "act indicator" in the Marp footer + (e.g., `Act 3/5: How it works`) keeps the audience oriented. + +### Existing deck inventory (to be retired) + +Two decks exist today in `docs/presentations/`: +- `how-the-platform-works.md` (32,916 bytes) → marp → html → talking-points +- `the-developer-experience.md` (27,509 bytes) → marp → html → talking-points + +Both follow a 4-step process (source `.md` → Marp → HTML → talking-points) +documented in `docs/presentations/README.md`. Per D-130, both are merged +into one unified narrative deck and retired. + +### Grounded metrics already cited in existing decks + +- "22/22 auto-verifiable capabilities Verified" — **stale** vs current + REGRESSION_REPORT.json (18V+4S post-D-096). The unified deck must + derive this from the report, not copy the stale claim. +- Confidence thresholds: dev ≥0.50, qa ≥0.75, prod ≥0.90, dr ≥0.95 — + grounded in `core/confidence_signal.py:57` (THRESHOLDS). +- RPO = 0 (evidence write synchronous) — grounded in + `core/audit_ledger_design.md:27,103`. +- Cost figures — `how-the-platform-works.md:461`; cites COST.md. +- Confidence signal 6 inputs + weights — grounded in + `core/confidence_signal.py:40-47`. +- "~80-line stateless adapter" vs "918-line monolith" — grounded in + ROADMAP/RESEARCH prose. + +### Planned deck structure (for PLAN to detail) + +The unified deck "Nova — The No-Humans Infrastructure Platform": + +| Act | Slides | Content | Proof source | +|---|---|---|---| +| 1. Problem | 2–3 | The no-humans imperative; why operators are the bottleneck; the trust gap | NORTH_STAR vision | +| 2. Vision/Direction | 2–3 | Nova's vision; 4 strategic objectives; anti-goals; the attestation model (autonomy in operations, human at stage gates) | NORTH_STAR | +| 3. How it works | 3–4 | Contract → resolver → adapter → confidence → HITL gate; the Decision Ledger; the 8-concern attestation matrix | code grounding | +| 4. Proof (metrics) | 3–4 | Capability health (18V+4S); confidence distribution; policy pass rate; Decision Ledger coverage; Attestation Coverage; cost estimates; the grounded/derived/deferred honesty model | metrics export | +| 5. Roadmap/Ask | 2 | 12–18mo targets (committed); deferred metrics (honest); the ask | NORTH_STAR targets | + +Total: ~12–16 slides. Opening = arc preview; closing = recap + ask. + +## 6. Assumptions logged (v1.17) + +- A1 (0.9): No live AWS access during execution (consistent with + v1.11–v1.16). All metrics that require live AWS ship as placeholder + views. The Infracost integration runs offline (reads plan JSON). +- A2 (0.85): The Decision Ledger SQLite hash-chain is sufficient for + v1.17's audit needs. The full tamper-evident ledger (S3 Object Lock + + JWS, D-083) is a future milestone. The hash-chain provides + append-only + integrity verification locally. +- A3 (0.8): The "AI decision" framing (D-122) is honest: Nova's "AI" is + the confidence-gated policy engine (confidence_signal + HITL gate), + not an LLM planner. The deck and METRICS.md must frame this accurately + — overclaiming "AI" would violate the "no fabrication" constraint. +- A4 (0.85): The unified deck's "Proof" section cites only grounded + metrics with real numbers. Deferred metrics are shown as "Planned" + with the `Planned` badge. No + fabricated numbers in any slide. +- A5 (0.8): `--junitxml` + `--json-report` added to pytest addopts + does not break the existing test suite (the flags are additive; pytest + continues to run normally). CAP-009 (offline pytest suite passes) + must remain Verified after the change. +- A6 (0.75): Infracost is available as a CLI tool that can be installed + in the CI environment and run locally. It reads `terraform plan + -out=plan.tfplan` + `terraform show -json plan.tfplan` to produce a + cost estimate. No live AWS access required. If Infracost is not + available, the `cost.estimated` event is omitted (degraded mode, not + a failure). diff --git a/.ciagent/config.json b/.ciagent/config.json index 34b8826..b433935 100644 --- a/.ciagent/config.json +++ b/.ciagent/config.json @@ -8,7 +8,7 @@ ], "active_project": "acdl", "active_projects": ["acdl"], - "active_milestone": "v1.16", + "active_milestone": "v1.17", "autonomy": { "level": "full", "escalation_hooks": ["deploy", "delete_data", "merge_to_main"],