Merge phase/00-pre-execution — v1.16.0 (v1.17 P0 pre-execution complete: NORTH_STAR + metrics plan + deck plan + grill)

This commit is contained in:
Jon Chery
2026-08-04 19:44:40 +00:00
10 changed files with 2463 additions and 428 deletions
+82
View File
@@ -797,3 +797,85 @@ return `Skipped` when the resources are absent (`NoSuchBucket`/
`ResourceNotFoundException`). `RegressionReport.passed` is
`all(r.status in ("Verified", "Skipped"))`. The gate passes at 18
Verified + 4 Skipped (0 Decayed/Broken).
## v1.17 Addendum — Strategic Direction, Leadership Metrics & Unified Story (2026-08-04)
The v1.17 milestone adds a telemetry/observability layer, a Decision
Ledger, a metrics export pipeline, a unified narrative deck, and a
durable strategic-direction artifact. This addendum documents the
architecture; the full research findings are in RESEARCH.md §v1.17.
### New components
| Component | Path | Purpose |
|-----------|------|---------|
| Event envelope | `core/metrics/event_envelope.py` | CloudEvents 1.0 envelope + `platform.*` semantic conventions (P1, REQ-187) |
| Per-run manifest writer | `core/metrics/run_manifest.py` | Emits `nova.run.started/completed/failed` events + writes `metrics/runs/<run_id>.json` (P1, REQ-187) |
| Decision Ledger (SQLite) | `core/metrics/decision_ledger.py` | Extends `outbox_writer.py` → SQLite append-only hash-chain table; `ai.decision.made` + `attestation.recorded` events + outcome backfill (P1, REQ-188, D-121) |
| Infracost post-processor | `core/metrics/infracost_adapter.py` | Runs Infracost on plan JSON; emits `nova.cost.estimated{delta_usd}` (P1, REQ-187, D-120) |
| Metrics collector | `core/metrics/collector.py` | Reads all grounded signals (files + events) → SQLite cold store at `metrics/nova_metrics.db` (P2, REQ-189) |
| PowerBI export | `core/metrics/powerbi_export.py` | Emits CSV/JSON views to `metrics/powerbi/` (fact + dim + 8 deferred placeholder views) (P3, REQ-190) |
| Metrics schemas | `schemas/metrics_*.schema.json` | Schemas for all event types + fact/dim tables (P1P2, REQ-187/189) |
| Metrics catalog | `docs/METRICS.md` + `docs/metrics/<kpi>.md` | Canonical catalog + per-KPI definition-of-success docs (P4, REQ-195, D-127) |
| Unified narrative deck | `docs/presentations/nova-no-humans-platform.md` | Merged deck: Problem→Vision→How→Proof→Roadmap; x3 arc at deck+slide level (P5, REQ-196/197, D-130) |
| Strategic direction | `.ciagent/NORTH_STAR.md` | PO-authored durable vision/objectives/anti-goals/targets; read by CIAgent in every future `/ci-run` (P0, REQ-185/186) |
### Modified components
| Component | Change | Phase |
|-----------|--------|-------|
| `core/outbox_writer.py` | Extended to emit to SQLite append-only hash-chain table (Decision Ledger); `ai.decision.made` + `attestation.recorded` events added (P1, D-121) | P1 |
| `scripts/run_platform.sh` | Per-run manifest writer invoked; `$WORK/*.json` persisted to `metrics/runs/`; Infracost post-processor invoked after plan (P1) | P1 |
| `core/hitl_gates.py` | Emits `attestation.recorded` event to Decision Ledger on qa/prod/dr gate (P1, D-132) | P1 |
| `core/confidence_signal.py` | Emits `nova.confidence.computed` + `nova.ai.decision.made` events (P1, D-122) | P1 |
| `adapters/terraform/policy/checkov_adapter.py` | Emits `nova.policy.evaluated` event (P1) | P1 |
| `core/regression_verify.py` | Emits `nova.capability.verified` event; CAP-023 (metrics collector) + CAP-024 (deck structure) added (P1, P6) | P1, P6 |
| `pyproject.toml` | `addopts` gains `--junitxml=metrics/test-results.xml` + `--json-report` (P1, D-120) | P1 |
| `docs/presentations/` | Two old decks retired (deleted); unified deck added (P5, D-130) | P5 |
### Telemetry/observability layer architecture (D-120)
```
┌─────────────────────────────────────────────────────────────────────┐
│ Nova platform components (existing) │
│ run_platform.sh · confidence_signal · checkov_adapter · │
│ hitl_gates · regression_verify · outbox_writer · contract_ingestor │
└──────────────────────┬──────────────────────────────────────────────┘
│ CloudEvents 1.0 envelope (new emitters, P1)
┌─────────────────────────────────────────────────────────────────────┐
│ metrics/events.jsonl (append-only CloudEvents log) │
│ metrics/runs/<run_id>.json (per-run manifests) │
│ metrics/decision_ledger.db (SQLite hash-chain, D-121) │
│ metrics/test-results.xml (junit, P1) │
└──────────────────────┬──────────────────────────────────────────────┘
│ collector reads (P2)
┌─────────────────────────────────────────────────────────────────────┐
│ metrics/nova_metrics.db (SQLite cold store, D-126) │
│ fact_run · fact_capability · fact_policy_check · fact_confidence │
│ fact_test · fact_decision · fact_cost_estimate │
│ dim_capability · dim_milestone │
│ + 8 empty placeholder views (deferred metrics) │
└──────────────────────┬──────────────────────────────────────────────┘
│ powerbi_export (P3)
┌─────────────────────────────────────────────────────────────────────┐
│ metrics/powerbi/ (CSV/JSON views, folder connector, D-129) │
│ → PowerBI dashboards (external) │
└─────────────────────────────────────────────────────────────────────┘
```
**Hot path: deferred (D-126).** No live ops dashboard; SQLite is
cold-only (batch/historical). The hot path activates when live AWS is
re-provisioned (D-096 lift).
### NORTH_STAR integration point (REQ-186)
`.ciagent/NORTH_STAR.md` is read by CIAgent in context-loading for all
future milestones. The integration mechanism (to be finalized in P4):
a reference from `PROJECT.md` + `ARCHITECTURE.md` (this section) + a
config entry in `config.json` (`strategic_direction_file:
".ciagent/NORTH_STAR.md"`) that the run workflow reads at SPECIFY. This
ensures the strategic direction survives across milestones without
being overwritten by status updates.
+9 -10
View File
@@ -1,13 +1,12 @@
{
"phase": 21,
"stage": "complete",
"milestone": "v1.16",
"phase_role": "final",
"phase": 0,
"stage": "grill",
"milestone": "v1.17",
"phase_role": "pre_execution",
"attempts": 0,
"updated_at": "2026-07-30T16:05:00Z",
"milestone_complete": true,
"tag": "v1.15.26",
"release_id": 370,
"requirements": ["REQ-165", "REQ-166", "REQ-167", "REQ-168", "REQ-169", "REQ-170", "REQ-171", "REQ-172", "REQ-173", "REQ-174", "REQ-175", "REQ-176", "REQ-177", "REQ-178", "REQ-179", "REQ-180", "REQ-181", "REQ-182", "REQ-183", "REQ-184"],
"regression": {"Verified": 18, "Decayed": 0, "Broken": 0, "Skipped": 4}
"updated_at": "2026-08-04T21:15:00Z",
"milestone_complete": false,
"tag": null,
"requirements": ["REQ-185"],
"notes": "GRILL complete (interactive). 12 binding decisions applied: NORTH_STAR targets reclassified (E-003: 3 targets to Post-Pilot section; E-004: AI-Agent Intent Share to Future Horizons). Deck plan updated: slide 1 stake line (G-Q8), slide 4 benefit rewrite (G-Q9), slide 7 D-122 honesty sentence (G-Q4), Act 3->4 transition rewrite (G-Q13), slide 9 benefit reframe (G-Q14), slide 12 split into 12+13 (G-Q10), ROI formula inline + N=0 caveat (G-Q5/G-Q15), slide 14 preempt (G-Q11), slide 16 ask reframed as business decision (G-Q16). Deck now 16 main + 2 appendix = 18 slides."
}
+259
View File
@@ -636,3 +636,262 @@ re-provision the bucket.
YES, once G-111's criterion restatement + gate update are incorporated
(into P9's must-haves). G-112/G-113 are phase-entry clarifications for
P9/P12/P13. E-002 is deferred to P21. Confidence 0.85.
---
# GRILL — v1.17 "Strategic Direction, Leadership Metrics & Unified Story" (2026-08-04)
> **Griller:** CIAgent (red-team mode). **Milestone:** v1.17. **Axes:** 3
> (NORTH_STAR alignment, Deck story & arc, Deck per-slide rigor) per PO
> direction. **Stance:** adversarial — presumed over-scoped / infeasible /
> storytelling-weak until evidence forced otherwise.
## Evidence base
- `NORTH_STAR.md` (183 lines, draft), `PLAN.md` (1,114 lines, deck rebuild
plan incl. slide-by-slide), `REQUIREMENTS.md` v1.17 (REQ-185..213),
`RESEARCH.md` v1.17 (signal inventory, scorecard, deferred-decision
ledger, deck research).
- Codebase cross-checks: `REGRESSION_REPORT.json` = **18 Verified + 4
Skipped** (NOT "22/22 Verified" — the new deck plan correctly says
18V+4S; the *existing* decks still claim 22/22). `PROJECT.md:495` =
**0 consumer adoption**. `docs/NO_HUMANS_THESIS.md`, `docs/METRICS.md`,
`metrics/` do not yet exist (P4/P5 deliverables — expected).
- Decisions locked (D-120..D-132) — not re-litigated.
## The central contradiction
**NORTH_STAR.md:111** states: *"Targets are committed, not aspirational."*
**PO's G-Q6 answer:** *"the goal is simply to target a high touchless
resolution rate, not to say we have reached those targets given there are
0 consumers."*
These two statements are in direct conflict. "Committed, not aspirational"
+ "simply to target" = the document is lying about its own epistemic
status. This is the v1.10 decay root cause (PRE_MORTEM FM-3: decks
outrunning verified reality) repeating itself in the document meant to
prevent it.
## Axis 1 — NORTH_STAR alignment
### G-Q1 — Target with no backing REQ / placeholder
**Finding:** AI-Agent Intent Share (≥40%) is a committed 1218mo target
(NORTH_STAR:128) with "placeholder view" claimed, but it is NOT among the
8 placeholder views in PLAN P3 (lines 309315), and no REQ-185..213 builds
an emitter or placeholder for it. RESEARCH §3 marks it "future" with no
controlling decision ID (unlike every other deferred metric). NORTH_STAR:128
falsely claims a placeholder view exists → violates the "no fabrication"
hard constraint.
**Verdict: BIND.** Add a 9th placeholder view OR move the target to a
"Future Horizons" section; correct NORTH_STAR:128. **Confidence: 0.90.**
### G-Q2 — Anti-goal pursuit
**Finding:** No REQ builds an anti-goal. Deck title "No-Humans Infrastructure
Platform" is one weak slide away from violating anti-goal #3 (not removing
humans from accountability) — mitigation is entirely in slide 3's execution.
**Verdict: PASS (conditional on slide 3 landing the attestation model).**
**Confidence: 0.75.**
### G-Q3 — Attestation clarification consistency
**Finding:** The attestation clarification is the most consistently
propagated concept in the plan — NORTH_STAR (3 places), REQUIREMENTS
(3 REQs), deck (3 slides). Well done.
**Verdict: PASS.** **Confidence: 0.92.**
### G-Q4 — "AI decision" framing (D-122 honesty)
**Finding:** D-122 (confidence_signal + HITL gate, NOT an LLM) is cited on
slide 7 and required in NO_HUMANS_THESIS.md (REQ-213). BUT slide 7's
*Delivers* says "every AI decision captured" without ever telling the
audience what the "AI" is. The honesty is buried in a linked doc + a
decision ID the audience has never heard.
**Verdict: BIND.** Add one sentence to slide 7 *Delivers*: "Nova's 'AI
decision' is the confidence-gated policy engine, not an LLM planner
(D-122)." **Confidence: 0.85.**
### G-Q5 — Secretly ungrounded metrics
**Finding:** The 8 deferred placeholder views cover their list. BUT (a)
AI-Agent Intent Share's placeholder is falsely claimed (G-Q1), and (b)
derived metrics (FTE Hours Saved, Platform ROI) are computed on zero
production runs yet shown on slide 12 without the zero-denominator caveat.
A "derived" metric from zero runs is technically not fabricated but is
misleading.
**Verdict: BIND.** (1) Resolve G-Q1; (2) slide 12 must annotate derived
metrics with "(computed on N internal runs; production-denominator activates
post-pilot)." **Confidence: 0.82.**
### G-Q6 — 1218mo target feasibility (0 consumers)
**Finding:** PO's answer ("simply to target") conflicts with NORTH_STAR:111
("committed, not aspirational"). 3 "grounded (after P1)" targets (Touchless
Resolution, Human Escalation, AI Decision Accuracy) have scope "across
production estates" — but PROJECT.md:495 = 0 consumer adoption. The metric
IS computable on internal dev runs, but the target scope doesn't exist.
Marking "grounded" while the scope is absent is the overclaim the "no
fabrication" constraint exists to prevent.
**Verdict: BIND.** (1) Rewrite NORTH_STAR:111 → "Targets are committed
destinations; the grounding column records whether each is measurable this
milestone." (2) Reclassify the 3 targets to `partial — measurement pipeline
grounded on internal runs; production-estate scope activates post-pilot`
(the Cloud Spend Reduction precedent at NORTH_STAR:123). (3) Deck slide 5
regroup as "Measurable today (internal runs)" vs "Activates post-pilot
(production estates)." Requires NORTH_STAR-CHANGE commit trailer (REQ-204).
**Confidence: 0.80.**
## Axis 2 — Deck plan: story & arc
### G-Q7 — Arc order (Problem→Vision→How→Proof→Roadmap vs Proof-first)
**Finding:** Current arc puts Proof at Act 4 (slides 1013) — 40% of the
deck before a number. For a leadership audience that has seen 10+ milestone
decks, this risks losing the room by slide 4. BUT the "no-humans" thesis
is contentious; jumping to proof without the attestation model invites the
"removing humans from accountability" objection. The Vision act makes the
Proof credible.
**Verdict: PASS (marginal).** Defensible IF the Problem act is tight and
slide 3 front-loads the attestation clarification. **Confidence: 0.62.**
### G-Q8 — x3 structure at deck level
**Finding:** Slide 1's 5-act preview is orienting (a table of contents),
not too much meta-structure. BUT it's also not a hook — it gives structure,
not stakes. A C-suite audience decides in the first 30 seconds.
**Verdict: BIND (minor).** Add one stake-establishing line to slide 1
*Delivers* with a real number (18 verified, 0 consumers, honest deferral
list). **Confidence: 0.70.**
### G-Q9 — Per-slide benefit callouts (substantive vs filler)
**Finding:** 4 of 17 closes are filler (slides 1, 4, 12, 15); 2 borderline
(8, A1). Worst offender: slide 12 (ROI) restates the *objective* ("ROI is
quantifiable") rather than giving the *number* or the *honest caveat*.
**Verdict: BIND.** Rewrite 4 filler closes. Slide 12's close must be:
"Benefit: you now know the ROI formula — (labor + cloud + avoided downtime)
÷ platform cost — and that it computes on internal runs today, with
production-denominator activating post-pilot." **Confidence: 0.78.**
### G-Q10 — Deck length (17 slides)
**Finding:** 17 is at the upper bound but justifiable for 5 acts. The risk
is density, not length: slide 12 crams 6 metrics (Touchless, Human
Escalation, MTTR, Cost, FTE, ROI) into one slide — a wall of bullets.
**Verdict: BIND (minor).** Split slide 12 into "Zero-Touch Efficiency"
(Touchless, Human Escalation, MTTR) + "Cost & ROI" (Cost, FTE, ROI). Deck
→ 18 slides, each earning its place. **Confidence: 0.68.**
### G-Q11 — "What's Deferred" slide (13)
**Finding:** The honesty strengthens the grounded claims BUT surfaces the
gap: Nova claims "no-humans in operations" while deferring the metrics
that would prove operations are healthy without humans (Live Infra Health,
SLA, Drift Auto-Reversal). A skeptical viewer notes the contradiction.
**Verdict: BIND.** Add preempt to slide 13: "These deferrals are about
*measurement infrastructure*, not about whether the platform runs without
humans — the platform runs autonomously today on internal runs; what's
deferred is the production-estate dashboard that would prove it at scale."
**Confidence: 0.75.**
## Axis 3 — Deck plan: per-slide rigor
### G-Q12 — Slide opening lines
**Finding:** The "This slide shows X" formula is orienting, not patronizing,
because each includes a stake-bearing clause. Consistent without being empty.
**Verdict: PASS.** **Confidence: 0.80.**
### G-Q13 — Transitions (written vs hand-waved)
**Finding:** ~10 of 13 transitions are written (specific reference to prior
close). 3 are hand-waved (slides 8→9, 11→12, 13→14). Worst: the Act 3→4
boundary (slide 8→9, How→Proof) — the most important transition in the deck
— is the weakest.
**Verdict: BIND.** Rewrite the 3 hand-waved transitions. The 8→9 Act
boundary must carry weight: "Having seen the gate model — autonomy in
operations, human in accountability — here is how Nova instruments itself
to prove that model at scale." **Confidence: 0.85.**
### G-Q14 — Weakest slide (audience-loss point)
**Finding:** Slide 9 (Telemetry Architecture) is the audience-loss slide.
It's the 4th consecutive architecture slide (6,7,8,9), the most abstract
(CloudEvents, SQLite, PowerBI), its Benefit is about data plumbing not
business value, and it sits between the attestation matrix (slide 8,
emotionally resonant) and the Proof act (slide 10, the numbers) — between
the two things the audience came for.
**Verdict: BIND.** Compress slide 9 into slide 10 OR reframe its Benefit
from data plumbing to trust: "Benefit: you now know the proof you're about
to see isn't fabricated — every number traces to a file you can audit."
**Confidence: 0.78.**
### G-Q15 — Proof act citation specificity
**Finding:** 5 of 6 Proof citations are specific (file paths + real numbers).
Gap: slide 12's derived metrics (FTE, ROI) cite "derived" without showing
the formula or the input count.
**Verdict: BIND (minor).** Show the ROI formula inline on slide 12 + the
N=0 production-runs caveat. **Confidence: 0.80.**
### G-Q16 — Closing slide (15) — does the ask land?
**Finding:** THE ask is present but framed as insider language ("fund the
hot-path activation (post-D-096) + the tamper-evident ledger build-out
(D-083 lift)"). A leadership audience doesn't know what "hot-path
activation" means. The ask is a technical request, not a business decision
a leader can make in the room.
**Verdict: BIND.** Reframe slide 15's ask as a business decision: "The
ask: (1) approve a pilot estate to activate production-estate metrics
(unblocks D-096), and (2) approve the tamper-evident ledger build-out
(lifts D-083) — turning grounded claims into complete proof." Make it a
yes/no a leader can give. **Confidence: 0.82.**
## Binding decisions (must resolve before SHIP)
| G-ID | Axis | Verdict | What must change | Conf |
|---|---|---|---|---|
| G-Q1 | 1 | BIND | Add 9th placeholder view for AI-Agent Intent Share OR move to "Future Horizons"; correct NORTH_STAR:128 | 0.90 |
| G-Q4 | 1 | BIND | Add D-122 honesty sentence to slide 7 *Delivers* | 0.85 |
| G-Q5 | 1 | BIND | Annotate derived metrics on slide 12 with zero-run caveat | 0.82 |
| G-Q6 | 1 | BIND | Rewrite NORTH_STAR:111; reclassify 3 targets to `partial`; regroup deck slide 5. NORTH_STAR-CHANGE trailer required | 0.80 |
| G-Q8 | 2 | BIND (minor) | Add stake line with real number to slide 1 *Delivers* | 0.70 |
| G-Q9 | 2 | BIND | Rewrite 4 filler closes (slides 1, 4, 12, 15); slide 12 must give ROI formula + caveat | 0.78 |
| G-Q10 | 2 | BIND (minor) | Split slide 12 into two (Efficiency + Cost/ROI); deck → 18 slides | 0.68 |
| G-Q11 | 2 | BIND | Add preempt to slide 13 (deferrals are measurement infra, not whether platform runs without humans) | 0.75 |
| G-Q13 | 3 | BIND | Rewrite 3 hand-waved transitions (esp. Act 3→4 boundary 8→9) | 0.85 |
| G-Q14 | 3 | BIND | Compress slide 9 into slide 10 OR reframe its Benefit to trust | 0.78 |
| G-Q15 | 3 | BIND (minor) | Show ROI formula inline + N=0 caveat on slide 12 | 0.80 |
| G-Q16 | 3 | BIND | Reframe slide 15 ask as a business decision (pilot estate + ledger build-out) | 0.82 |
**PASS (no change):** G-Q2 (anti-goals, conditional on slide 3), G-Q3
(attestation consistency — excellent), G-Q7 (arc order — marginal),
G-Q12 (slide openings — formulaic but substantive).
## Escalations (only the PO can decide)
| E-ID | Question | Confidence |
|---|---|---|
| E-003 | Should the 3 "grounded (after P1)" targets with "production estates" scope be reclassified to `partial` (Cloud Spend precedent), or should "grounded" be redefined to mean "measurement pipeline grounded"? Changes a committed NORTH_STAR target's grounding label; requires NORTH_STAR-CHANGE trailer (REQ-204). | <0.60 |
| E-004 | Should AI-Agent Intent Share (≥40%) remain a "1218mo Target" with no backing REQ/placeholder, or move to a "Future Horizons" section? Strategic-scope question (is agentic consumption a 1218mo commitment or a longer horizon?). | <0.60 |
## Overall verdict
**🟡 REDUCE SCOPE / BINDING FIXES REQUIRED — not ready to ship as-is.**
The plan is architecturally sound (metrics pipeline, Decision Ledger,
PowerBI export, x3 deck structure are well-designed and grounded). The
attestation clarification (G-Q3) is the best-propagated concept in the
plan. The regression-capability gate (CAP-023/024) is a credible safeguard.
But the plan has one structural contradiction (NORTH_STAR:111 vs PO intent
vs grounding column) that infects 4 other findings (G-Q1, G-Q5, G-Q6,
G-Q9/slide 12). This is the v1.10 decay pattern (PRE_MORTEM FM-3)
repeating in the document meant to prevent it. The "no fabrication" hard
constraint is self-violated in two places (AI-Agent Intent Share placeholder
claim, derived-metrics-without-caveat) before a single slide is rendered.
The deck plan is story-competent but not story-excellent. 4 benefit
callouts are filler, 3 transitions are hand-waved (incl. the critical
Act 3→4 boundary), slide 9 is the audience-loss slide, and the closing
ask is insider language.
**12 binding decisions, 2 escalations.** None require re-architecting the
plan; all are edits to NORTH_STAR (2 rows + 1 line, with commit trailer),
the deck slide plan (4 slide rewrites, 1 split, 3 transition rewrites),
and one placeholder-view addition. Estimate: 12 phases of rework, not a
milestone restart. The plan does NOT need a revision loop — it needs
these 12 fixes applied in P0 (NORTH_STAR) and P5 (deck) before the
respective phases ship. Critical path unchanged.
**Can the milestone proceed?**
YES, once the 12 BIND decisions are incorporated (G-Q1/Q4/Q5/Q6 into P0
NORTH_STAR + P5 deck plan; G-Q8/Q9/Q10/Q11/Q13/Q14/Q15/Q16 into P5 deck
plan). E-003/E-004 require PO decisions on NORTH_STAR target framing.
Confidence 0.80.
+211
View File
@@ -0,0 +1,211 @@
# NORTH_STAR — Nova
> **Status:** Draft (pending interactive GRILL → final)
> **Milestone:** v1.17 — Strategic Direction, Leadership Metrics & Unified Story
> **Owner:** Product Owner
> **Purpose:** Durable strategic intent. Read by CIAgent in every future
> `/ci-run` so the platform's direction survives across milestones. This
> is NOT a status document (that's PROJECT.md) and NOT an engineering
> architecture (that's the telemetry reference in RESEARCH.md/
> ARCHITECTURE.md). It is the PO's committed direction: what we're
> building toward, what we refuse to build, and how we'll know we won.
---
## Vision
> **Infrastructure operations become invisible. Every environment
> provisioned, every incident healed, every risk remediated — by an
> autonomous system whose trustworthiness is provable, not promised.
> Human attestation remains required at stage gates — QA signs off for
> production, SRE greenlights based on operational readiness — but the
> operator is never in the loop of normal operations.**
Nova is the autonomous infrastructure layer that lets product teams ship
without engaging an operator, and lets executives trust the AI not because
it never fails but because every decision is captured, scored, and
accountable.
---
## Strategic Objectives (4)
**1. Demonstrate production-grade zero-touch operations.**
Nova must run real customer estates with no human in the loop of normal
operations — autonomy as the default, not the demo. Stage-gate
attestation (QA for production, SRE for operational readiness) remains
human by design; operational escalations (AI confidence too low to
proceed) are the failure mode we drive toward zero. Everything else
collapses if autonomy isn't real.
**2. Establish provable trust in AI decisions.**
Build the audit substrate — Decision Ledger, confidence scoring, circuit
breakers, blast-radius controls — that turns "autonomous" from a
marketing claim into a defensible one. Trust is the moat. Features can be
copied; an immutable, queryable decision history cannot.
**3. Deliver compounding, quantifiable ROI for customers.**
Each quarter on Nova must reduce cloud spend, free engineering hours, and
avoid downtime measurably. If the CFO can't point to a number that
improves quarter-over-quarter, Nova fails its commercial test, regardless
of how clever the AI is.
**4. Become the default substrate for agentic infrastructure consumption.**
AI agents are already becoming the largest consumers of cloud
infrastructure. Nova must be the platform through which those agents
declare, deploy, and verify infrastructure — not a vendor scrambling into
that market two quarters late.
---
## Anti-Goals (5 — what Nova is fundamentally NOT)
1. **Not a Terraform, Kubernetes, or hyperscaler competitor.** We
orchestrate them. Replacing them is the most expensive possible
distraction from the value we create.
2. **Not a general-purpose AI agent platform.** We are purpose-built for
infrastructure operations. Breadth here produces shallow tools; depth
here wins the category.
3. **Not a system that removes humans from accountability.** Only from
operations. Every AI decision lands in an immutable ledger. Every
stage-gate promotion (qa/prod/dr) requires a human attestation recorded
with approver identity, separation-of-duties check, and the 8-concern
evidence matrix. The absence of an operator is never the absence of a
record.
4. **Not for legacy, untagged, or freeform infrastructure.** Nova requires
Terraform-managed, policy-aligned, fully-tagged inputs. We optimize for
the disciplined 95%, not the chaotic 5%.
5. **Not sold to operators.** Nova is sold to leadership on outcomes —
cost, velocity, risk. Selling to operators inverts the incentive and
breaks the autonomy thesis.
---
## Non-Goals (v1.17 milestone scope — deferred work, not permanent boundaries)
> Anti-Goals are what Nova *fundamentally is not*. Non-Goals are what we
> *will not do this milestone* — deferred work, not permanent boundaries.
> Each Non-Goal cites the controlling decision ID.
1. **Live AWS re-provisioning** (deferred — D-096). Metrics that require
live infrastructure ship as placeholder PowerBI views with documented
schemas.
2. **Onboarding auto-grant** (deferred — D-113/D-114/D-119). Only the
request-path metric is grounded; the requested→granted funnel is a
placeholder.
3. **ML anomaly-forecasting / predictive remediation** (no emitter today).
The Predictive-vs-Reactive metric ships as a placeholder.
4. **Drift detection scheduled job** (deferred — D-096 + no scheduler).
Drift metrics ship as placeholders.
5. **Live cost CUR reconciliation** (deferred — D-096). Pre-apply Infracost
estimates are grounded; actual-spend reconciliation is a placeholder.
6. **S3 Object Lock / JWS tamper-evident ledger** (deferred — D-083). The
Decision Ledger uses a local SQLite hash-chain this milestone; the
Object-Lock/JWS build-out is a future milestone.
7. **Multi-cloud support** (Azure/GCP/K8s). Nova is AWS-only this milestone.
---
## 1218 Month Targets
Targets are committed, not aspirational. Each is a number a board member
can repeat back to us. The grounding column records whether the metric is
measurable this milestone, and if not, what blocks it.
> **Honesty note (GRILL G-Q6 binding):** Nova has 0 consumer adoption
> today (`PROJECT.md:495`). Three targets (Touchless Resolution, Human
> Escalation, AI Decision Accuracy) are scoped "across production
> estates" — the measurement *pipeline* is grounded this milestone, but
> the *denominator* is zero until a pilot estate activates. These
> targets are reclassified as **Post-Pilot** (the pipeline works; the
> numbers fill when consumers exist). This is the same honesty model as
> Cloud Spend Reduction (partial: pipeline grounded, actuals deferred).
### Current-milestone targets (grounded or derived this milestone)
| Domain | Target | Grounding (v1.17) | Note |
|---|---|---|---|
| **MTTR (p95)** | < 60 seconds | grounded (platform-run MTTR) | apply.failed → successful retry; infra-incident MTTR deferred (no incident detection) |
| **Cloud Spend Reduction** | ≥ 25% on pilot estates vs. 12-month pre-Nova baseline | partial | pre-apply estimate grounded (Infracost); actual-spend deferred (D-096 CUR) |
| **L1 / L2 Ops Hours Avoided** | ≥ 70% of pre-Nova FTE allocation | derived | formula over run count × manual baseline (computed on N internal runs; production-denominator activates post-pilot) |
| **Platform ROI** | ≥ 250% measured annually | derived | formula (labor savings + cloud savings + avoided downtime) ÷ platform op cost (computed on N internal runs; production-denominator activates post-pilot) |
| **Decision Ledger Coverage** | 100% of AI actions with backfilled outcome | grounded (this milestone builds it) | outbox_writer.py → SQLite hash-chain |
| **Attestation Coverage** | 100% of prod/dr promotions attested by a human | grounded | hitl_gates.py + outbox approver_* attributes; separation-of-duties on prod |
### Post-Pilot targets (pipeline grounded this milestone; denominator activates when a pilot estate runs)
| Domain | Target | Grounding (v1.17) | Note |
|---|---|---|---|
| **Touchless Resolution Rate** | ≥ 99% across production estates | partial (pipeline grounded; denominator = 0 today) | runs completing without *operational* HITL block ÷ total runs (attestation gates excluded); activates post-pilot |
| **Human Escalation Frequency** | < 0.1% of platform actions | partial (pipeline grounded; denominator = 0 today) | *operational* HITL blocks only (confidence-driven); attestation sign-offs excluded; activates post-pilot |
| **AI Decision Accuracy** | ≥ 99.5% (no rollback, no follow-up incident within 5 min of action) | partial (pipeline grounded; denominator = 0 today) | decisions not followed by apply.failed/incident within 5min; activates post-pilot |
### Deferred targets (measurement requires future systems)
| Domain | Target | Grounding (v1.17) | Note |
|---|---|---|---|
| **Predictive vs. Reactive Ratio** | ≥ 3 : 1 (prevention dominates reaction) | deferred | requires ML forecasting service (future emitter) |
| **Drift Auto-Reversal Rate** | ≥ 95% within one detection cycle | deferred | requires drift detection (D-096 + scheduler) |
> Committed targets whose measurement is deferred remain committed — the
> target is the destination; the metric is the odometer, and some
> odometers aren't built yet. Each deferred metric ships as a placeholder
> PowerBI view + a definition-of-success doc recording the dependency.
> Post-Pilot targets are committed targets whose measurement pipeline is
> grounded this milestone; the numbers activate when a pilot estate runs.
### Future Horizons (strategic direction, not committed targets)
| Domain | Aspiration | Note |
|---|---|---|
| **AI-Agent Intent Share** | ≥ 40% of total intent volume originated by non-human consumers | Strategic Objective #4 direction. No backing requirement, no placeholder view, no emitter today. Moves to a committed target when agentic consumption is real. |
---
## Success Criteria (v1.17 — what constitutes success for THIS milestone)
> Distinct from the 1218mo targets: those are the destination. These are
> the milestone's exit criteria.
v1.17 is a success if:
1. **Decision Ledger emits `ai.decision.made` for 100% of platform runs**
with outcome backfill, AND **`attestation.recorded` events for 100%
of qa/prod/dr promotions** (event completeness — all 3 gates captured;
grounded in `outbox_writer.py` → SQLite hash-chain; honors D-083).
The **Attestation Coverage metric** (target 100%) measures prod/dr
promotions specifically — see REQ-194.
2. **`docs/METRICS.md` catalogs every executive KPI** with a `grounded` /
`derived` / `deferred` status, a source file or decision ID, and a
per-KPI definition-of-success doc in `docs/metrics/`.
3. **The PowerBI export produces all fact/dimension views** + 8 empty
placeholder views for deferred metrics (with documented schemas ready
to fill when their blocking decisions lift).
4. **The unified narrative deck ships** with the x3 arc
(Problem→Vision→How→Proof→Roadmap) at deck + slide level, per-slide
benefit callouts, and fluid transitions; both old decks retired.
5. **`NORTH_STAR.md` is wired into CIAgent context-loading** so every
future `/ci-run` reads it.
6. **CAP-023 (metrics collector) + CAP-024 (deck structure) pass** in the
regression gate.
---
## What "won" looks like
By month 18, Nova is the layer enterprise leadership points to when they
say *"we don't have an infrastructure ops team anymore, and the audit
trail is stronger than it ever was"* — and it is the default substrate
their AI engineering teams reach for first when an agent needs to deploy.
---
## Relationship to v1.17 engineering
- **Pillar A (this file):** strategic direction — durable, PO-authored.
- **Pillar B (engineering):** the telemetry reference architecture
(adapted from the PO's technical-direction input) lives in
RESEARCH.md/ARCHITECTURE.md. It is the *how*; this file is the *why*.
- **Pillar C (story):** the unified narrative deck proves Pillars A+B to
leadership. The deck's Proof section cites grounded metrics; its
Roadmap section cites deferred targets honestly.
+145 -13
View File
@@ -1,23 +1,24 @@
---
project: acdl
milestone: v1.16
generated_at: 2026-07-30
milestone: v1.17
generated_at: 2026-08-04
generator: lead-developer
verification_toolchain:
typecheck: "terraform validate && python3 -m py_compile core/**/*.py && python3 -m jsonschema schemas/*.schema.json"
test: "bash scripts/run_regression.sh # 22-capability gate (D-091/D-118)"
test: "bash scripts/run_regression.sh # 22-capability gate (D-091/D-118) + CAP-023/024 (v1.17)"
build: "bash scripts/run_ci.sh # full local CI reproduction (lint+test+check-only)"
note: |
Nova (formerly ACDL) has no package.json. The execute/verify/ship
workflows substitute `terraform validate` + `python -m py_compile` +
JSON Schema validation for npm run typecheck, the regression gate
(D-091, 22 capabilities) for npm test, and `bash scripts/run_ci.sh`
for npm run build. v1.11 testing is pipeline-driven (D-102);
v1.16 is NFR-only (no live apply by default; NOVA_LIFECYCLE_MODE=
plan). Roster carries forward from v1.11/v1.14/v1.15 unchanged.
frontend-engineer stays inactive (no frontend; decks are markdown =
lead-developer territory). No custom personas needed (no new
domains — onboarding is backend-engineer + data-engineer territory).
v1.17 adds a telemetry/observability layer (metrics emitters, SQLite
cold store, PowerBI export, Decision Ledger) + a unified narrative
deck + a durable NORTH_STAR.md. Three active personas: lead-developer
(coordination + deck narrative co-author), backend-engineer (event
emitters, outbox_writer extension, Infracost adapter), data-engineer
(SQLite store, schemas, PowerBI views, metrics collector). frontend-
engineer stays deactivated (no Nova web UI — dashboards are PowerBI,
not a Nova-built frontend; decks are markdown = lead-developer
territory). No new custom personas needed — the metrics domain maps
cleanly to data-engineer (schema/store/export) + backend-engineer
(emitters/instrumentation).
---
# ACDL — Persona Roster (project-level, v1.11 RESTART)
@@ -251,3 +252,134 @@ The regression gate (22 capabilities) must stay **22/22 Verified**
throughout v1.16 — simplification must not regress any capability
(D-118). P9 (end of Wave 2) and P21 (milestone complete) run the gate;
P14 (end of Wave 3) is an offline mid-milestone checkpoint.
---
# v1.17 Persona Roster — Strategic Direction, Leadership Metrics & Unified Story
> v1.17 adds a telemetry/observability layer (P1P3), a metrics catalog
> + NORTH_STAR integration (P4), a unified narrative deck (P5), a
> regression capability (P6), and a final review/ship (P7). Three
> active personas; frontend-engineer stays deactivated (no Nova web UI
> — dashboards are PowerBI, not a Nova-built frontend).
## Active personas
### lead-developer
- **Domain:** coordination + deck narrative
- **Active:** true
- **Phase-specific:** false
- **Reason:** Owns CIAgent metadata, the NORTH_STAR.md authoring
process (P0), the milestone decomposition, the unified narrative deck
co-authoring (P5 — the deck is markdown, which is lead-developer
territory per the established convention), and the final review/ship
(P7). Arbitrates persona conflicts (e.g., backend vs data on the
emitter/store boundary).
- **Territory:** `.ciagent/NORTH_STAR.md`, `.ciagent/PROJECT.md`,
`.ciagent/REQUIREMENTS.md`, `.ciagent/PLAN.md`, `.ciagent/RESEARCH.md`,
`.ciagent/ARCHITECTURE.md`, `docs/presentations/nova-no-humans-platform.md`
(NEW — unified deck source of truth), `docs/presentations/nova-no-humans-platform-marp.md`,
`docs/presentations/nova-no-humans-platform-talking-points.md`,
`docs/METRICS.md`, `docs/metrics/*.md` (per-KPI definition docs).
### backend-engineer
- **Domain:** backend (event emitters + instrumentation)
- **Active:** true
- **Phase-specific:** false
- **Reason:** Owns the event emitters (P1): the CloudEvents envelope,
the per-run manifest writer, the `outbox_writer.py` extension to the
SQLite Decision Ledger, the Infracost post-processor, the
`hitl_gates.py` attestation event emission, the `confidence_signal.py`
decision event emission, the `checkov_adapter.py` policy event
emission, and the pytest `--junitxml` addopts change. Also owns the
`regression_verify.py` CAP-023/024 additions (P6). The emitter work
is the bridge between existing Nova components and the new metrics
layer — it touches the code paths that already exist.
- **Territory:** `core/metrics/event_envelope.py` (NEW),
`core/metrics/run_manifest.py` (NEW),
`core/metrics/infracost_adapter.py` (NEW),
`core/metrics/decision_ledger.py` (NEW — extends outbox_writer),
`core/outbox_writer.py` (extend to SQLite),
`core/hitl_gates.py` (emit attestation.recorded),
`core/confidence_signal.py` (emit ai.decision.made),
`adapters/terraform/policy/checkov_adapter.py` (emit policy.evaluated),
`scripts/run_platform.sh` (invoke manifest writer + Infracost),
`core/regression_verify.py` (CAP-023/024),
`pyproject.toml` (addopts --junitxml),
`tests/test_metrics_emitters.py` (NEW),
`tests/test_decision_ledger.py` (NEW).
### data-engineer
- **Domain:** data (schema, SQLite store, PowerBI export)
- **Active:** true
- **Phase-specific:** false
- **Reason:** Reactivated with a new territory for v1.17: the metrics
collector (P2) and the PowerBI export (P3). Owns the schema design
(metrics_*.schema.json), the SQLite cold store (nova_metrics.db), the
fact/dimension table design, the 8 deferred placeholder views, and
the CSV/JSON export. The data-engineer's schema-first constraint
applies: all event types and fact/dim tables have JSON Schema
definitions before any code is written. The collector reads files +
events → SQLite; the export reads SQLite → CSV/JSON. This is the
heaviest data-territory work since v1.11's terraform modules.
- **Territory:** `core/metrics/collector.py` (NEW),
`core/metrics/powerbi_export.py` (NEW),
`schemas/metrics_*.schema.json` (NEW — event + fact/dim schemas),
`metrics/nova_metrics.db` (NEW — SQLite cold store),
`metrics/powerbi/` (NEW — CSV/JSON export dir),
`docs/METRICS_VIEWS.md` (NEW — schema doc for PowerBI views),
`tests/test_metrics_collector.py` (NEW),
`tests/test_powerbi_export.py` (NEW).
## Deactivated personas
### frontend-engineer
- **Domain:** frontend
- **Active:** false
- **Phase-specific:** false
- **Reason:** v1.17 has no Nova web UI. The leadership dashboards are
PowerBI (an external tool that ingests CSV/JSON files), not a
Nova-built frontend. The decks are markdown (lead-developer
territory). frontend-engineer stays deactivated, consistent with
v1.11v1.16. Reactivates if a future milestone builds a Nova web UI.
### lambda-engineer, platform-engineer, security-engineer
- **Active:** false (carried forward from v1.11)
- **Reason:** v1.17 does not touch the Lambda (beyond emitting events
from the existing hitl_gates/attestation_matrix), does not do IR-
shaped module authoring, and does not touch security adapters beyond
emitting policy.evaluated events. The existing components are
instrumented, not rewritten.
## v1.17 phase assignment
| Phase | Primary persona | Supporting | Territory |
|-------|----------------|------------|-----------|
| P0 pre-execution | lead-developer | — | `.ciagent/NORTH_STAR.md`, `PROJECT.md`, `REQUIREMENTS.md`, `RESEARCH.md`, `ARCHITECTURE.md`, `PERSONAS.md`, `PLAN.md` |
| P1 event-emitters | backend-engineer | data-engineer (schemas) | `core/metrics/event_envelope.py`, `run_manifest.py`, `decision_ledger.py`, `infracost_adapter.py`, `outbox_writer.py`, `hitl_gates.py`, `confidence_signal.py`, `checkov_adapter.py`, `run_platform.sh`, `pyproject.toml` |
| P2 metrics-collector | data-engineer | backend-engineer (event formats) | `core/metrics/collector.py`, `schemas/metrics_*.schema.json`, `metrics/nova_metrics.db` |
| P3 powerbi-export | data-engineer | — | `core/metrics/powerbi_export.py`, `metrics/powerbi/`, `docs/METRICS_VIEWS.md` |
| P4 metrics-catalog + north-star-integration | lead-developer | data-engineer (metric definitions) | `docs/METRICS.md`, `docs/metrics/*.md`, `PROJECT.md`, `ARCHITECTURE.md`, `config.json` |
| P5 deck-rebuild | lead-developer | — | `docs/presentations/nova-no-humans-platform*.md`, retire old decks |
| P6 regression-capability | backend-engineer | data-engineer (CAP-023 schema) | `core/regression_verify.py` (CAP-023, CAP-024) |
| P7 final-review-ship | lead-developer | all active (review) | `.ciagent/**`, review + audit + ship |
## v1.17 domain priority
`backend → data → lead` (the emitter work in P1 is the foundation;
data-engineer's collector + export in P2P3 depends on P1's event
formats; lead-developer's catalog + deck in P4P5 depends on the
metrics being grounded).
## v1.17 verification toolchain
```
typecheck: terraform validate && python3 -m py_compile core/**/*.py adapters/**/*.py
test: bash scripts/run_regression.sh # 22-capability gate + CAP-023/024 (v1.17)
build: bash scripts/run_ci.sh # full local CI reproduction
```
The regression gate (22 capabilities + CAP-023 metrics collector +
CAP-024 deck structure) must pass at P6 and P7. CAP-009 (offline pytest
suite) must remain Verified after the `--junitxml` addopts change
(assumption A5).
+1165 -403
View File
File diff suppressed because it is too large Load Diff
+62 -1
View File
@@ -1072,4 +1072,65 @@ conversation; D-117..D-119 resolved at CLARIFY.
| D-116 | Drift fixes = P1 of v1.16 (not a hotfix to main). | User chose "P1 of v1.16." The state-bucket drift (`adapter.py:117`) and Kyverno label contradiction are correctness regressions but latent in plan-only mode (no live apply in the default path), so they are not an active outage. Fixing them as P1 keeps the milestone self-contained. | P1 fixes both; no hotfix to main. |
| D-117 | v1.14 NFR categories are NOT re-proposed. | v1.14 already swept over-broad excepts (REQ-141), hardcoded account-ID (REQ-142), IAM `Resource:"*"` scoping (REQ-143), contractId/env validation (REQ-144), `.gitignore` catch-all (REQ-146), `--kube-version` removal (REQ-147), orphan cleanup (REQ-148), `set -euo pipefail` parity (REQ-150). v1.16 finds NEW residual signals (the v1.15 rebrand left a fresh debt layer) and does not duplicate completed work. | Wave 15 target only fresh debt. |
| D-118 | Regression gate (D-091) gates Wave 2 completion and P21. | "Simplify without regressions" is only credible if the regression gate runs after the simplification wave. The gate runs after P9 (Wave 2 done) and at P21 (milestone complete); any non-Verified capability halts W3. Mid-milestone checkpoint after P14 (offline). | P9 + P21 run the gate; P14 checkpoint. |
| D-119 | `onboard_consumer` action stores a CMDB row pending grant (not auto-provisions). | The request-path-only scope (D-113) means the Lambda accepts an onboarding request and writes a `pending` row to `nova-contracts` (or a new `nova-onboarding` partition key); the platform automation that grants the ABAC role is the P20 Terraform (offline-proven). No AWS resources are created by the Lambda action itself. | P18 writes the pending row; P20 proves the grant Terraform offline. |
| D-119 | `onboard_consumer` action stores a CMDB row pending grant (not auto-provisions). | The request-path-only scope (D-113) means the Lambda accepts an onboarding request and writes a `pending` row to `nova-contracts` (or a new `nova-onboarding` partition key); the platform automation that grants the ABAC role is the P20 Terraform (offline-proven). No AWS resources are created by the Lambda action itself. | P18 writes the pending row; P20 proves the grant Terraform offline. |
## Objective for Milestone v1.17 (active — Strategic Direction, Leadership Metrics & Unified Story)
**Milestone type:** Feature (P1P3 feat; P4 docs; P5 docs+test; P6 test;
P7 review+audit+ship). Tags on the v1.16.x line: `v1.16.0` (P0) →
`v1.16.1..v1.16.7` (P1P7) → `v1.16.8` (P8 final = milestone release).
**Three pillars:**
- **Pillar A — Strategic Direction.** A durable, PO-authored
`.ciagent/NORTH_STAR.md` encodes the platform's vision, 4 strategic
objectives, 5 anti-goals, v1.17 non-goals, 1218mo targets (with a
grounding column), and success criteria. CIAgent reads it in every
future `/ci-run` so the direction survives across milestones. The
attestation clarification is reflected: human attestation required at
stage gates (QA for production, SRE for operational readiness); autonomy
in operations, not in accountability.
- **Pillar B — Leadership Metrics + PowerBI.** Instrument Nova to
collect, aggregate, and surface leadership-grade metrics that prove the
"no-humans" autonomous-infrastructure value proposition. Nova-native
minimal tech (CloudEvents 1.0 envelope, JSONL event log, SQLite cold
store, hash-chained Decision Ledger via `outbox_writer.py` extension)
+ Infracost for pre-apply cost estimates. Hybrid model: existing
file-based signals (REGRESSION_REPORT.json, pcr.json, signal.json,
junit XML) are sources the collector reads and projects into events;
new emitters emit CloudEvents directly. PowerBI export = CSV/JSON
views (fact + dimension tables + 8 empty placeholder views for
deferred metrics). **Hard constraint: DO NOT make anything up.** Every
metric is `grounded` (cites source file + schema), `derived`
(documented formula), or `deferred` (cites decision ID — D-096/D-083/
D-113/D-114/D-119). The 8 deferred metrics: drift detection, GreenOps/
carbon, predictive/reactive, live CUR reconciliation, multi-cloud,
red-team MTTR, self-healing velocity, SLA/downtime.
- **Pillar C — Unified Narrative Deck.** Merge the two existing decks
(`how-the-platform-works` + `the-developer-experience`) into one unified
narrative deck "Nova — The No-Humans Infrastructure Platform" with a
single arc: Problem → Vision/Direction (NORTH_STAR) → How it works →
Proof (metrics) → Roadmap/Ask. The "tell them x3" structure applies at
deck level AND per slide (each slide opens with what it covers,
delivers, closes with an explicit "benefit of this stage" callout).
Fluid transitions between slides. Both old decks retired.
**Key decisions resolved in the planning conversation (D-120+):**
| ID | Decision | Rationale | Outcome |
|----|----------|-----------|---------|
| D-120 | Tech stack = Nova-native + Infracost, drift deferred. | The PO's technical-direction document specifies Kafka/Prometheus/ClickHouse/QLDB/OTel — none exist in Nova today. Adopt the PRINCIPLES (events as source of truth, CloudEvents envelope, decision ledger, definition-of-success docs, dashboards-as-projections) but implement with Nova-native minimal tech (JSONL + SQLite + hash-chained ledger). No Kafka/Prometheus/ClickHouse/QLDB. Infracost adopted (runs offline on plan JSON). Drift detection deferred (D-096 + no scheduler). | P1P3 use Nova-native tech; Infracost in P1; drift deferred. |
| D-121 | Decision Ledger = extend outbox_writer.py → SQLite append-only hash chain. | The direction's #1 priority is the Decision Ledger. Nova already has a hash-chained outbox (outbox_writer.py). Extend it to a SQLite append-only table with hash chain; add ai.decision.made + attestation.recorded events. Honors D-083 (no S3 Object Lock/JWS). | P1 extends outbox_writer; ledger is SQLite hash-chain. |
| D-122 | AI Planner framing = map Nova's real decision points. | The direction assumes an "AI Planner/Reasoner" (planner-v3.2). Nova's actual decision path is confidence_signal + HITL gate. Model ai.decision.made from confidence_signal (decision_id=run_id, chosen_action=band, confidence=score, alternatives=perInput, human_override=HITL block). LLM planner marked future/aspirational. | P1 emits honest decision events; no fabricated LLM. |
| D-123 | Deferred metrics = all 8 (drift, GreenOps, predictive/reactive, live CUR, multi-cloud, red-team MTTR, self-healing, SLA/downtime). | These require live AWS (D-096) or new external systems. Ship as empty PowerBI placeholder views with documented schemas. | P3 ships 8 placeholder views; METRICS.md marks them deferred. |
| D-124 | NORTH_STAR = strategy; tech direction = engineering input. | The PO's technical-direction document is engineering architecture, not strategy. NORTH_STAR.md captures strategic vision/objectives/anti-goals (PO-authored). The tech direction becomes the telemetry reference architecture section in RESEARCH.md/ARCHITECTURE.md, cited by NORTH_STAR's engineering objectives. | P0 writes NORTH_STAR; RESEARCH writes the telemetry reference. |
| D-125 | Events vs files = hybrid. | Existing file-based signals (REGRESSION_REPORT.json, pcr.json, signal.json, junit) stay as files; the collector reads them and emits normalized CloudEvents into JSONL + SQLite. New emitters emit CloudEvents directly. | P2 collector reads files + events. |
| D-126 | Hot/cold split = cold-only SQLite (hot path deferred). | Nova has no live ops dashboard (no live AWS, D-096). The SQLite store is cold-only (batch/historical). The hot path is documented as deferred. | P2 SQLite is cold-only. |
| D-127 | Definition-of-success = per-KPI docs. | The direction's §11 requires a definition-of-success doc for every executive KPI. Adopt this standard; docs live in `docs/metrics/`. | P4 writes per-KPI docs. |
| D-128 | Storage location = metrics/ at repo root. | metrics/runs/ (per-run manifests), metrics/nova_metrics.db (SQLite), metrics/events.jsonl (event log), metrics/powerbi/ (export). | P1P3 use metrics/ at repo root. |
| D-129 | PowerBI delivery = CSV/JSON files, folder connector. | Nova is offline-first; no live connector to a running service. PowerBI ingests via the folder connector. | P3 emits CSV/JSON to metrics/powerbi/. |
| D-130 | Deck arc = Problem → Vision → How → Proof → Roadmap. | The unified narrative deck's 5-act structure. x3 arc at deck + slide level. Per-slide benefit callouts. Fluid transitions. Both old decks retired. | P5 builds the unified deck; old decks deleted. |
| D-131 | MTTR scope = platform-run MTTR. | The <60s MTTR target refers to platform-run failures (apply.failed → successful retry), not infra-incident MTTR (no incident detection system). Infra-incident MTTR deferred. | P4 grounds platform-run MTTR. |
| D-132 | Attestation instrumentation = emit attestation.recorded events. | The attestation system (hitl_gates.py + attestation_matrix.py + separation_of_duties.py) already exists. Instrument it: emit attestation.recorded events into the Decision Ledger + PowerBI. Attestation Coverage = 100% target grounded from outbox approver_* attributes. | P1 emits attestation events; P4 grounds Attestation Coverage. |
+224
View File
@@ -956,3 +956,227 @@ simplification and the first self-service onboarding request path.
scoping (REQ-143), contractId/env validation (REQ-144), `.gitignore`
catch-all (REQ-146), `--kube-version` removal (REQ-147), orphan
cleanup (REQ-148), `set -euo pipefail` parity (REQ-150).
## v1.17 — Strategic Direction, Leadership Metrics & Unified Story
**Milestone type:** Feature (P1P3 feat; P4 docs; P5 docs+test; P6 test;
P7 review+audit+ship). Progressive patches; the final phase's patch IS
the milestone release. Tags run on the v1.16.x line: `v1.16.0` (P0) →
`v1.16.1..v1.16.7` (P1P7) → `v1.16.8` (P8 final = milestone release).
**Objective:** Three pillars. (A) Encode the PO's strategic direction in
a durable `NORTH_STAR.md` read by CIAgent in every future `/ci-run`.
(B) Instrument Nova to collect, aggregate, and surface leadership-grade
metrics that prove the "no-humans" autonomous-infrastructure value
proposition — grounded in signals Nova actually emits, derived via
documented formulas, or explicitly deferred with a decision ID — flowing
into PowerBI-ready views. (C) Merge the two existing decks into one
unified narrative deck with the "tell them x3" arc at deck + slide level,
per-slide benefit callouts, and fluid transitions.
**Hard constraint:** DO NOT make anything up. Every metric carries a
`grounded` / `derived` / `deferred` status with a source file or
decision ID. Deferred metrics ship as empty PowerBI placeholder views
with documented schemas.
### Requirements
**Pillar A — Strategic Direction**
- **REQ-185**`.ciagent/NORTH_STAR.md` is PO-authored with Vision,
Strategic Objectives (4), Anti-Goals (5), Non-Goals (v1.17 scope),
1218mo Targets (with grounding column), and Success Criteria. The
attestation clarification is reflected: human attestation required at
stage gates (QA for production, SRE for operational readiness);
autonomy in operations, not in accountability. (Phase P0)
- **REQ-186** — CIAgent reads `NORTH_STAR.md` in context-loading for all
future milestones; the file is referenced from PROJECT.md and
ARCHITECTURE.md so the strategic direction survives across milestones.
(Phase P4)
**Pillar B — Leadership Metrics + PowerBI**
- **REQ-187** — Event emitters: a CloudEvents 1.0 envelope is adopted;
a per-run manifest writer emits structured events (run_id, contractId,
env, stages×durations, exit, confidence, HITL block count) to
`metrics/runs/`; existing ephemeral `$WORK/*.json` (pcr, signal,
event, outbox, stack) are persisted as durable artifacts; pytest
`addopts` gains `--junitxml`+`--json-report`; Infracost runs as a
plan post-processor emitting `cost.estimated{delta_usd}` (offline).
(Phase P1)
- **REQ-188** — Decision Ledger: `outbox_writer.py` is extended to emit
to a SQLite append-only table with hash chain; `ai.decision.made`
events are modeled from Nova's real decision points (decision_id=run_id,
chosen_action=band outcome, confidence=score, alternatives=perInput
breakdown, human_override=HITL block) with outcome backfill from
apply.completed; `attestation.recorded` events capture qa/prod/dr
sign-offs (approver, env, concerns, result). Honors D-083 (no S3 Object
Lock/JWS). (Phase P1)
- **REQ-189** — Metrics collector: `core/metrics/collector.py` +
`schemas/metrics_*.schema.json` read all grounded signals
(REGRESSION_REPORT.json, per-run manifests, junit XML, pcr.json,
signal.json, COST.md, decision ledger) → normalized SQLite cold store
at `metrics/nova_metrics.db`; idempotent re-runs. (Phase P2)
- **REQ-190** — PowerBI export: `core/metrics/powerbi_export.py` emits
CSV/JSON views to `metrics/powerbi/` (fact_run, fact_capability,
fact_policy_check, fact_confidence, fact_test, fact_decision,
fact_cost_estimate, dim_capability, dim_milestone + 8 empty
placeholder views for deferred metrics with documented schemas) +
`docs/METRICS_VIEWS.md` schema doc. (Phase P3)
- **REQ-191** — Zero-touch efficiency metrics: Autonomous Resolution
Rate (runs without operational HITL block ÷ total; attestation gates
excluded), Human Escalation Frequency (operational HITL blocks only),
AI Decision Accuracy (decisions not followed by apply.failed/incident
within 5min), MTTD/MTTR (platform-run: apply.failed → successful
retry). (Attestation Coverage is owned by REQ-194, not here.)
(Phase P4)
- **REQ-192** — Velocity metrics: Provisioning Lead Time
(apply.completed.time intent.received.time), Deployment Frequency
(count(apply.completed) per day). Self-Healing Velocity deferred (no
auto-remediator). (Phase P4)
- **REQ-193** — Financial & cost-ROI metrics: FTE Hours Saved (derived:
run count × manual baseline), Cost Savings via Infracost estimates
(grounded), Cost Efficiency Ratio (derived), Platform ROI (derived
formula). Live CUR reconciliation deferred (D-096). (Phase P4)
- **REQ-194** — Reliability, security & compliance metrics: Zero-Trust
Policy Compliance Rate (from pcr.json), Attestation Coverage (prod/dr
promotions attested by a human ÷ total prod/dr promotions; grounded in
hitl_gates.py + outbox approver_* attributes; canonical owner of this
metric). Uptime, Patch Remediation, SLA/downtime deferred (D-096).
(Phase P4)
- **REQ-195** — Metrics catalog doc: `docs/METRICS.md` catalogs every
executive KPI with `grounded`/`derived`/`deferred` status, source
file or decision ID, and a per-KPI definition-of-success doc in
`docs/metrics/<kpi>.md`. (Phase P4)
**Pillar C — Unified Narrative Deck**
- **REQ-196** — The two existing decks (`how-the-platform-works` +
`the-developer-experience`) are merged into one unified narrative deck
"Nova — The No-Humans Infrastructure Platform" with a single arc:
Problem → Vision/Direction (NORTH_STAR) → How it works → Proof
(metrics) → Roadmap/Ask. The x3 structure ("tell them what you're
going to tell them → tell them → tell them what you told them") applies
at deck level (opening = arc; body = tell them; closing = recap + ask).
Both old decks are retired (all derived artifacts deleted). (Phase P5)
- **REQ-197** — Each slide has the x3 structure (opens with what it
covers, delivers, closes with an explicit "benefit of this stage"
callout) + fluid transitions between slides (no disjointed jumps).
The 4-step deck process (source `.md` → Marp → HTML → talking-points)
is re-run for the unified deck. (Phase P5)
**Cross-cutting**
- **REQ-198** — Regression capability: CAP-023 (metrics collector runs,
emits expected schema) + CAP-024 (deck structure: slide count, x3
present, per-slide benefit present) added to `core/regression_verify.py`.
(Phase P6)
**Ideation enhancements (REQ-199..213 — additive, within D-120..D-132)**
- **REQ-199** — Metrics schema validation in CI: `run_ci.sh` validates
`metrics/powerbi/*.json` + a sample `metrics/events.jsonl` against
their schemas; exits 0. (Phase P3)
- **REQ-200** — Idempotent collector re-run test: `test_metrics_collector_idempotent`
passes (two runs → identical row counts + chain verified). (Phase P2)
- **REQ-201** — Metrics store backup/restore doc: `metrics/README.md`
documents regenerable vs append-only artifacts + restore procedure.
(Phase P2)
- **REQ-202** — Metrics glossary appendix slide: the unified deck has a
"Metrics Glossary" appendix slide with one-line KPI definitions +
grounding badges. (Phase P5)
- **REQ-203** — "What's Deferred — and Why" slide: the unified deck has
a slide pairing each of 8 deferred metrics with its blocking decision
ID. (Phase P5)
- **REQ-204** — NORTH_STAR diff-check in CI: `run_ci.sh` includes
`check_north_star_diff` that fails when Vision/Objectives/Anti-Goals/
Targets sections change without a `NORTH_STAR-CHANGE:` commit trailer.
(Phase P4)
- **REQ-205** — Per-module lifecycle success-rate report: each lifecycle
run writes `metrics/lifecycle/<module>-<env>.json`; collector projects
into `fact_lifecycle`; PowerBI "Module Lifecycle Health" view. (Phase
P1 emitter + P2 collector + P3 view)
- **REQ-206** — Code coverage trend emission: `pyproject.toml` addopts
gains `--cov=core --cov=adapters --cov-report=json:metrics/coverage.json`;
collector ingests; `fact_test` carries a coverage column. (Phase P1 +
P2)
- **REQ-207** — Decision Ledger CLI: `core/metrics/decision_ledger_cli.py`
supports `query`, `verify-chain`, `stats`, `export`, `replay`;
`verify-chain` detects broken hashes; `replay` prints ordered events;
tests pass offline. (Phase P2)
- **REQ-208** — PowerBI starter dashboard README: `metrics/powerbi/NOVA_DASHBOARD_README.md`
documents folder-connector import + starter visual model + reference
screenshot. (Phase P3)
- **REQ-209** — PowerBI column-level data dictionary: `docs/METRICS_VIEWS.md`
has a per-column data-dictionary table (column, type, source/formula,
unit, grounded/derived/deferred status). (Phase P3/P4)
- **REQ-210** — Deferred-metrics activation roadmap: `docs/METRICS_DEFERRED_ROADMAP.md`
lists 8 deferred metrics + onboarding-grant half with {blocking
decision, unblock requirement, candidate milestone} + a "Hot-Path
Activation (post-D-096)" section (Nova-native only, D-120) +
"Re-evaluation Triggers" section. (Phase P4)
- **REQ-211** — Trust-snapshot report: `core/metrics/trust_snapshot.py`
emits `metrics/TRUST_SNAPSHOT.md` with 5 trust metrics (Decision Ledger
Coverage, Attestation Coverage, Capability Health, AI Decision
Accuracy, Confidence-Gate Halt Rate) + chain-integrity verdict +
snapshot hash; runs offline. (Phase P4)
- **REQ-212** — Confidence-Gate Halt Rate metric: `docs/METRICS.md` +
trust snapshot include "Confidence-Gate Halt Rate" (signal.json
band=halt ÷ total runs); PowerBI view includes it. (Phase P4)
- **REQ-213** — "No-humans" thesis defensibility brief: `docs/NO_HUMANS_THESIS.md`
defines the thesis, grounded proof metrics, deferred proof metrics,
and explicit anti-claims (incl. D-122 honesty); the unified deck's
Vision act cites it. (Phase P4/P5)
### v1.17 Traceability
| Requirement | Phase | Status |
|-------------|-------|--------|
| REQ-185 | P0 | in_progress |
| REQ-186 | P4 | pending |
| REQ-187 | P1 | pending |
| REQ-188 | P1 | pending |
| REQ-189 | P2 | pending |
| REQ-190 | P3 | pending |
| REQ-191 | P4 | pending |
| REQ-192 | P4 | pending |
| REQ-193 | P4 | pending |
| REQ-194 | P4 | pending |
| REQ-195 | P4 | pending |
| REQ-196 | P5 | pending |
| REQ-197 | P5 | pending |
| REQ-198 | P6 | pending |
| REQ-199 | P3 | pending |
| REQ-200 | P2 | pending |
| REQ-201 | P2 | pending |
| REQ-202 | P5 | pending |
| REQ-203 | P5 | pending |
| REQ-204 | P4 | pending |
| REQ-205 | P1+P2+P3 | pending |
| REQ-206 | P1+P2 | pending |
| REQ-207 | P2 | pending |
| REQ-208 | P3 | pending |
| REQ-209 | P3/P4 | pending |
| REQ-210 | P4 | pending |
| REQ-211 | P4 | pending |
| REQ-212 | P4 | pending |
| REQ-213 | P4/P5 | pending |
### Out of Scope (v1.17)
- Live AWS re-provisioning (D-096) — metrics requiring live
infrastructure ship as placeholder views.
- Onboarding auto-grant (D-113/D-114/D-119) — only the request-path
metric is grounded.
- ML anomaly-forecasting / predictive remediation — no emitter today;
Predictive-vs-Reactive metric ships as a placeholder.
- Drift detection scheduled job (D-096 + no scheduler) — drift metrics
ship as placeholders.
- Live cost CUR reconciliation (D-096) — Infracost pre-apply estimates
are grounded; actuals are not.
- S3 Object Lock / JWS tamper-evident ledger (D-083) — Decision Ledger
uses a local SQLite hash-chain this milestone.
- Multi-cloud support (Azure/GCP/K8s) — Nova is AWS-only this milestone.
- A third deck — the two existing decks merge into one; no new
standalone metrics deck.
- A Nova web UI — dashboards are PowerBI, not a Nova-built frontend.
+305
View File
@@ -1188,3 +1188,308 @@ stays a future feature (D-113).
- A4 (0.85): The regression gate (D-091, D-118) at P9 and P21 confirms
"simplify without regressions" — 22/22 capabilities must stay Verified.
The gate is the credible control for the simplification wave.
---
# v1.17 Research — Strategic Direction, Leadership Metrics & Unified Story
> Phase: research (P0). Milestone: v1.17. Status: research.
> Researcher: ci-researcher + explore agent (signal inventory).
> Autonomy: full. Decisions D-120..D-132 locked in the planning
> conversation (PROJECT.md). NORTH_STAR.md drafted (pending GRILL).
## 1. Telemetry Signal Inventory (grounding audit)
**Methodology:** every claim below is grounded in a concrete file path +
line number in `/root/acdl`. No speculation. The explore agent performed
a full sweep of the repo. The finding: **Nova has no metrics/telemetry/
dashboard aggregation layer today.** What exists is a set of discrete,
structured, file-based signal artifacts (JSON reports, JSONL logs,
hash-chained outbox events, PR comments, Checkov JSON) plus unstructured
stdout logs. A metrics milestone must aggregate these existing signals
— it must not invent new ones without first adding emitters.
### (a) Signals that EXIST TODAY and are STRUCTURED (groundable)
| Signal | File / Emitter | Schema | Persistent? |
|--------|---------------|--------|-------------|
| Regression report (22 caps, status, duration_ms, gate) | `.ciagent/REGRESSION_REPORT.json``core/regression_verify.py:643-667` | `regression_verify.py:82-91` | **Yes** (committed file) |
| Regression report (markdown mirror) | `.ciagent/REGRESSION_REPORT.md` | same | Yes |
| Checkpoint (milestone/phase/tag/regression summary) | `.ciagent/CHECKPOINT.json` (CIAgent-managed) | ad-hoc | Yes |
| PolicyCheckResult list (per-rule pass/fail/severity/resourceRef) | `$WORK/pcr.json``run_platform.sh:395` + `checkov_adapter.py:50-71` | `schemas/policy_check_result.schema.json` | **No** (ephemeral `/tmp/`) |
| Confidence signal (score, band, perInput, reasonCodes) | `$WORK/signal.json``run_platform.sh:412-426` + `confidence_signal.py:60-65` | `confidence_signal.py:60-65` | No (ephemeral) |
| Outbox event (hash-chained, CONFIDENCE_COMPUTED) | `$WORK/event.json` + `$WORK/outbox_item.json``run_platform.sh:444-459` + `outbox_writer.py:44-56` | `audit_ledger_design.md:44-45,81-97` | No (ephemeral; live DynamoDB torn down D-096) |
| Resolved Target Stack | `$WORK/stack.json``contract_resolver.py:581-603` | `schemas/stack.schema.json` | No (ephemeral) |
| Lambda return bodies (submit/report_error/validate_cr/onboard) | `core/lambda/contract_ingestor.py:171,265,284,392,446` | ad-hoc JSON | No (Lambda not live; local stub only) |
| DynamoDB CMDB rows (submitted/pending contracts) | `nova-contracts` table ← `contract_ingestor.py:160-170,433-445` | ad-hoc | **No** (table torn down D-096) |
| SSM parameters (deploy outputs) | `/nova/<env>/<contractId>/<name>``output_publisher.py:123-156` | ad-hoc | No (live AWS, torn down) |
| PR stage comment (mode, runId) | GitHub PR API ← `post_stage_comment.sh:34-48` + `deploy.yml:141` | markdown table | Yes (GitHub) |
| PR deploy-outputs comment | GitHub PR API ← `output_publisher.py:159-189` | markdown table | Yes (GitHub) |
| GitHub issue (deploy failure alert) | GitHub API ← `contract_ingestor.py:179-290` + `deploy.yml:143-152` | issue body | Yes (GitHub) |
| Local E2E result (stack_name, tier, outbox_events, chain_verified, lambda_status) | stdout JSON ← `core/local_emulators.py:498-508,519` | ad-hoc | No (stdout) |
| HITL gate result | `core/hitl_gates.py:87,90` + `run_platform.sh:179-185` | stdout `HITL PASS/BLOCK` | No (stdout) |
| Attestation matrix result | `core/attestation_matrix.py:184,187` | stdout `ATTESTATION PASS/BLOCK` | No (stdout) |
| Cost figures | `.ciagent/COST.md` (manual Cost Explorer query) | markdown table | Yes (manual, not automated) |
### (b) Signals that EXIST but are UNSTRUCTURED (log-only)
| Signal | Source | Format |
|--------|--------|--------|
| CI pipeline result | `scripts/run_ci.sh:70-71` | stdout banner `=== CI PIPELINE OK ===` |
| Platform stage banners + summaries | `scripts/run_platform.sh:222,241,258,263,315,383,411,442,463,490,496` | stdout `=== Step N: ... ===` + summary lines |
| Terraform init/validate/plan/apply/destroy logs | `$WORK/tf-*.log``run_platform.sh:320,324,328,352,375` | raw terraform stdout (via `tee`) |
| Lifecycle test results | `scripts/run_lifecycle_test.sh` etc. | exit code only (no report file) |
| Decommission step counts | `scripts/run_decommission.sh:40,54` | stdout `decommission step N: M resources...` |
| Uptime endpoint count | `scripts/run_uptime.sh:72,87` | stdout `uptime: N endpoint(s) to monitor` |
| Onboarding prompt | `core/environment_check.py:57-81` | stdout text block |
| Pytest results | `pyproject.toml:25` (`-v --tb=short`) | stdout only (no junit/json) |
| sync_workflows result | `scripts/sync_workflows.py:56,53` | stdout `OK: 3 workflow pairs match` / `DRIFT: ...` |
### (c) Proposed executive metrics with NO grounding today (DEFERRED)
| Proposed metric | Why no grounding | Controlling decision |
|------------------|------------------|---------------------|
| Live infrastructure health (ECS running count, ALB 5xx, RPS) | Live AWS torn down; CAP-013..016 Skipped | **D-096** |
| Live outbox write rate / ledger append latency | DynamoDB outbox table absent | **D-096** |
| Tamper-evident ledger checkpoint count / JWS signature rate | S3 Object Lock + JWS + async worker deferred | **D-083** |
| Onboarding funnel: requested → granted conversion | Only "requested" (pending row) is emitted; no grant event | **D-113, D-114, D-119** |
| Time-to-provision (onboarding SLA) | Real AWS provisioning deferred | **D-113** |
| Cross-account role grant count | Offline-proven only, no live apply | **D-114** |
| Drift detection (scheduled terraform plan -detailed-exitcode) | Needs live AWS workspaces + a scheduler Nova doesn't have | **D-096** + no scheduler |
| GreenOps / carbon (WattTime/Electricity Maps API) | No grounding; new external API | future emitter |
| Predictive vs Reactive ratio | Requires an ML anomaly-forecasting service | future emitter |
| Multi-cloud normalization (Azure/GCP/K8s, FOCUS spec) | Nova is AWS-only | future |
| Red Team MTTR | No red-team program exists | future |
| Self-healing velocity | Nova has no auto-remediator | future emitter |
| SLA / unplanned downtime | Needs live service uptime monitoring against SLOs | **D-096** |
| Per-module lifecycle success rate over time | No structured report file written; only exit code | gap (no decision) |
| Test pass rate / test count time-series | No junit/json reporter configured | gap (add `--junitxml` to addopts) |
| Code coverage trend | `pytest-cov` installed but not in `addopts` | gap |
| Deploy frequency / lead time / MTTR (DORA) | No deploy-event emitter; pipeline runs not counted | gap |
| Policy pass rate time-series | `pcr.json` emitted but ephemeral; not persisted | gap (D-096 blocks live persistence) |
| Confidence score distribution over time | `signal.json` emitted but ephemeral | gap |
| Consumer adoption count / active consumers | `PROJECT.md:487` explicitly states "0 consumer adoption today" | honest scope |
| Cost time-series (automated) | `COST.md` is a one-shot manual query; no automated emitter | gap |
**Bottom line:** the single richest existing structured signal is
`.ciagent/REGRESSION_REPORT.json` (22 capabilities × {status, tier,
duration_ms, detail} + summary counts + boolean gate). The next richest
is the per-run `$WORK/*.json` family (pcr.json, signal.json, event.json,
stack.json) — but these are **ephemeral** and **not persisted in CI**.
The lowest-friction grounding for a "no-humans" dashboard is therefore:
(1) regression report → capability health, (2) PR comments + GitHub
issues → deploy/failure activity, (3) add `--junitxml` to pytest → test
trend, (4) persist `$WORK/*.json` → policy/confidence/outbox time-series,
(5) extend outbox_writer → Decision Ledger, (6) add Infracost →
pre-apply cost estimates.
## 2. Telemetry Reference Architecture (Nova-native adaptation)
The PO provided a full distributed-system telemetry reference
architecture (CloudEvents 1.0 envelope, OpenTelemetry SDK, Kafka/NATS
event bus, Prometheus hot path, ClickHouse warehouse, QLDB decision
ledger, Infracost, drift detection, ML anomaly forecasting). Per
D-120, we adopt the **principles** but implement with **Nova-native
minimal tech**. The mapping:
| Direction's principle | Nova-native implementation (v1.17) |
|---|---|
| Events are the source of truth; dashboards are projections | Hybrid (D-125): existing file signals stay as files; collector reads them and emits normalized CloudEvents into `metrics/events.jsonl` + SQLite. New emitters emit CloudEvents directly. |
| Every AI action is logged with confidence + alternatives | Decision Ledger (D-121): `outbox_writer.py` extended → SQLite append-only hash-chain table. `ai.decision.made` modeled from confidence_signal (D-122): decision_id=run_id, chosen_action=band, confidence=score, alternatives=perInput, human_override=HITL block. |
| Hot/cold storage split | Cold-only SQLite (D-126): `metrics/nova_metrics.db`. Hot path deferred (no live ops, D-096). |
| Read-only external integrators | Infracost (pre-apply, offline, reads plan JSON). Cloud billing CUR deferred (D-096). Carbon APIs deferred (future). |
| CloudEvents 1.0 envelope | Adopted. `core/metrics/event_envelope.py` defines the envelope + `platform.*` semantic conventions. |
| Decision Ledger = append-only with hash chain + outcome backfill | SQLite append-only table with hash chain (D-121). Outcome backfilled from apply.completed via decision_id → request_id correlation. Honors D-083 (no S3 Object Lock/JWS). |
| Cost governance: mandatory tags + Infracost pre-apply | Nova already enforces `nova:*` tags (nova_tagging.py, hard mode). Infracost added as plan post-processor (D-120). Post-apply CUR deferred (D-096). |
| Definition-of-success docs for every KPI | Per-KPI docs in `docs/metrics/` (D-127). |
| Replay-ability | SQLite store + JSONL event log are replayable by design. |
### CloudEvents envelope (Nova-native)
```json
{
"specversion": "1.0",
"id": "<uuid>",
"source": "nova.platform",
"type": "nova.run.completed",
"time": "<ISO8601>",
"subject": "<contractId>/<env>",
"datacontenttype": "application/json",
"platform": {
"tenant_id": "acdl",
"run_id": "run-<epoch>",
"contract_id": "<uuid>",
"environment": "dev|qa|prod|dr",
"actor": {"type": "confidence-gate", "id": "confidence_signal"},
"trace_id": "<run_id>"
},
"data": {
"duration_ms": 4800,
"stages": ["resolve", "adapt", "validate", "plan", "apply"],
"exit_code": 0,
"confidence": {"score": 0.94, "band": "pass", "perInput": {...}},
"policy": {"passed": 12, "failed": 0, "skipped": 0},
"hitl": {"gate": "dev", "result": "autonomous", "block": false},
"cost_estimate_usd": -12.40,
"decision_id": "run-<epoch>",
"outcome": "succeeded"
}
}
```
### Core event types (Nova-native minimum viable set)
| Event type | Emitted by | Purpose | Grounding |
|---|---|---|---|
| `nova.run.started` | run_platform.sh | Measures demand; provisioning lead time start | new emitter (P1) |
| `nova.run.completed` | run_platform.sh | Run count, stage durations, exit, MTTR | new emitter (P1) |
| `nova.run.failed` | run_platform.sh | Failure count, MTTR numerator | new emitter (P1) |
| `nova.policy.evaluated` | checkov_adapter.py | Policy pass rate, compliance KPIs | grounded (pcr.json → P1 persists) |
| `nova.confidence.computed` | confidence_signal.py | Confidence distribution, decision accuracy | grounded (signal.json → P1 persists) |
| `nova.ai.decision.made` | outbox_writer.py (extended) | Decision Ledger entry | grounded (D-121, D-122) |
| `nova.attestation.recorded` | hitl_gates.py | Attestation Coverage, human-in-the-loop audit | grounded (D-132) |
| `nova.cost.estimated` | Infracost post-processor | Pre-apply cost estimate | new emitter (P1, Infracost) |
| `nova.capability.verified` | regression_verify.py | Capability health, regression gate | grounded (REGRESSION_REPORT.json) |
| `nova.test.completed` | pytest (junit XML) | Test count, pass rate | new (P1 adds --junitxml) |
## 3. Metric-to-Signal Scorecard (the "no fabrication" contract)
| Executive metric (NORTH_STAR target) | Status | Source / formula | Decision |
|---|---|---|---|
| Touchless Resolution Rate ≥99% | grounded (after P1) | runs without operational HITL block ÷ total runs (attestation gates excluded) | D-122, D-132 |
| Human Escalation Frequency <0.1% | grounded (after P1) | operational HITL blocks ÷ total runs (attestation sign-offs excluded) | D-122, D-132 |
| MTTR (p95) <60s | grounded (platform-run) | apply.failed.time → successful retry.time | D-131 |
| Predictive vs Reactive ≥3:1 | **deferred** | requires ML forecasting (future emitter) | future |
| AI Decision Accuracy ≥99.5% | grounded (after decision ledger) | decisions not followed by apply.failed/incident within 5min | D-121, D-122 |
| Drift Auto-Reversal ≥95% | **deferred** | requires drift detection (D-096 + scheduler) | D-096 |
| Cloud Spend Reduction ≥25% | partial | pre-apply estimate grounded (Infracost); actuals deferred (D-096 CUR) | D-120 |
| L1/L2 Ops Hours Avoided ≥70% | derived | formula: run count × manual baseline minutes × blended rate | D-127 |
| Platform ROI ≥250% | derived | formula: (labor savings + cloud savings + avoided downtime) ÷ platform op cost | D-127 |
| Decision Ledger Coverage 100% | grounded (this milestone) | outbox_writer.py → SQLite hash-chain | D-121 |
| Attestation Coverage 100% | grounded | hitl_gates.py + outbox approver_* attributes; prod/dr | D-132 |
| AI-Agent Intent Share ≥40% | future | no AI-agent consumers today; placeholder view | future |
| Capability health (18V+4S) | grounded | REGRESSION_REPORT.json | existing |
| Confidence score distribution | grounded (after P1) | signal.json → decision ledger | D-121 |
| Policy pass rate | grounded (after P1) | pcr.json → persisted | D-120 |
| Test count / pass rate | grounded (after P1) | pytest --junitxml | D-120 |
| Provisioning Lead Time | grounded (after P1) | run.started → run.completed | D-120 |
| Deployment Frequency | grounded (after P1) | count(run.completed) per day | D-120 |
| Deploy-failure alert count | grounded | GitHub issues via Lambda report_error (D-055) | existing |
| Cost figures (actuals) | manual one-shot | COST.md (Cost Explorer query) | existing |
| FTE Hours Saved / TRV | derived | formula over run count + COST.md | D-127 |
| Self-healing velocity | **deferred** | no auto-remediator | future |
| SLA / unplanned downtime | **deferred** | needs live service uptime (D-096) | D-096 |
| GreenOps / carbon | **deferred** | WattTime/Electricity Maps API (future) | future |
| Red Team MTTR | **deferred** | no red-team program | future |
| Multi-cloud normalization | **deferred** | Nova is AWS-only | future |
| Live CUR reconciliation | **deferred** | needs live AWS billing (D-096) | D-096 |
## 4. Deferred-Decision Ledger (constraints on this milestone)
| Decision | Scope | Grounding impact |
|----------|-------|------------------|
| D-096 | Live AWS torn down post-v1.11 | BLOCKS all live-AWS metrics (CAP-013..016 Skipped; live outbox; live state bucket; live CUR) |
| D-083 | S3 Object Lock + JWS + async worker deferred | BLOCKS tamper-evident ledger; v1.17 uses SQLite hash-chain instead |
| D-113/D-114/D-119 | Onboarding = request-path only; no auto-grant | BLOCKS onboarding funnel "granted" half |
| D-091/D-118 | Regression gate (D-091) gates milestone completion | ENABLES the strongest metric signal (REGRESSION_REPORT.json) |
| D-092 | Local emulating adapters | ENABLES offline E2E metrics (CAP-011/012) |
| D-055 | report_error Lambda action creates GitHub issues | ENABLES deploy-failure alert metric |
| D-050 | Publish deploy outputs to SSM + GitHub PR comment | ENABLES outputs-published metric |
| D-054/D-043/D-109 | Nova tagging standard (hard mode) | ENABLES tagging-compliance metric |
| D-084 | 8-concern attestation matrix | ENABLES attestation metrics (operator-supplied evidence artifacts) |
| D-089 | Signature verification skipped when signing key unset (dev/CI) | Signature metrics are no-ops in dev |
## 5. Deck-Storytelling Research (x3 arc + per-slide benefit)
### The "tell them x3" structure
The PO's direction: "Tell them what you're going to tell them, then tell
them, then tell them what you told them." Applied at two levels:
**Deck level (the 5-act arc):**
1. **Opening slide** = "what I'm going to tell you" — the full arc
preview: Problem → Vision → How → Proof → Roadmap.
2. **Body** (acts 15) = "tell them" — each act delivers its content.
3. **Closing slide** = "what I told you" — recap of the 5 acts + the ask.
**Per slide:**
1. **Slide opens** with what it'll cover (1 line: "This slide shows X").
2. **Slide delivers** the content (bullets, diagram, or table).
3. **Slide closes** with an explicit **"benefit of this stage" callout**
(1 line: "Benefit: you now know Y" or "Why this matters: Z").
### Fluidity conventions
- **Transitions are written, not hand-waved.** Each slide's opening line
references the previous slide's close ("Having seen X, now consider Y").
- **No disjointed jumps.** If a topic shift is needed, a bridge slide or
a transition sentence carries the audience across.
- **The arc is visible.** A small "act indicator" in the Marp footer
(e.g., `Act 3/5: How it works`) keeps the audience oriented.
### Existing deck inventory (to be retired)
Two decks exist today in `docs/presentations/`:
- `how-the-platform-works.md` (32,916 bytes) → marp → html → talking-points
- `the-developer-experience.md` (27,509 bytes) → marp → html → talking-points
Both follow a 4-step process (source `.md` → Marp → HTML → talking-points)
documented in `docs/presentations/README.md`. Per D-130, both are merged
into one unified narrative deck and retired.
### Grounded metrics already cited in existing decks
- "22/22 auto-verifiable capabilities Verified" — **stale** vs current
REGRESSION_REPORT.json (18V+4S post-D-096). The unified deck must
derive this from the report, not copy the stale claim.
- Confidence thresholds: dev ≥0.50, qa ≥0.75, prod ≥0.90, dr ≥0.95 —
grounded in `core/confidence_signal.py:57` (THRESHOLDS).
- RPO = 0 (evidence write synchronous) — grounded in
`core/audit_ledger_design.md:27,103`.
- Cost figures — `how-the-platform-works.md:461`; cites COST.md.
- Confidence signal 6 inputs + weights — grounded in
`core/confidence_signal.py:40-47`.
- "~80-line stateless adapter" vs "918-line monolith" — grounded in
ROADMAP/RESEARCH prose.
### Planned deck structure (for PLAN to detail)
The unified deck "Nova — The No-Humans Infrastructure Platform":
| Act | Slides | Content | Proof source |
|---|---|---|---|
| 1. Problem | 23 | The no-humans imperative; why operators are the bottleneck; the trust gap | NORTH_STAR vision |
| 2. Vision/Direction | 23 | Nova's vision; 4 strategic objectives; anti-goals; the attestation model (autonomy in operations, human at stage gates) | NORTH_STAR |
| 3. How it works | 34 | Contract → resolver → adapter → confidence → HITL gate; the Decision Ledger; the 8-concern attestation matrix | code grounding |
| 4. Proof (metrics) | 34 | Capability health (18V+4S); confidence distribution; policy pass rate; Decision Ledger coverage; Attestation Coverage; cost estimates; the grounded/derived/deferred honesty model | metrics export |
| 5. Roadmap/Ask | 2 | 1218mo targets (committed); deferred metrics (honest); the ask | NORTH_STAR targets |
Total: ~1216 slides. Opening = arc preview; closing = recap + ask.
## 6. Assumptions logged (v1.17)
- A1 (0.9): No live AWS access during execution (consistent with
v1.11v1.16). All metrics that require live AWS ship as placeholder
views. The Infracost integration runs offline (reads plan JSON).
- A2 (0.85): The Decision Ledger SQLite hash-chain is sufficient for
v1.17's audit needs. The full tamper-evident ledger (S3 Object Lock +
JWS, D-083) is a future milestone. The hash-chain provides
append-only + integrity verification locally.
- A3 (0.8): The "AI decision" framing (D-122) is honest: Nova's "AI" is
the confidence-gated policy engine (confidence_signal + HITL gate),
not an LLM planner. The deck and METRICS.md must frame this accurately
— overclaiming "AI" would violate the "no fabrication" constraint.
- A4 (0.85): The unified deck's "Proof" section cites only grounded
metrics with real numbers. Deferred metrics are shown as "Planned"
with the `<span class="badge planned">Planned</span>` badge. No
fabricated numbers in any slide.
- A5 (0.8): `--junitxml` + `--json-report` added to pytest addopts
does not break the existing test suite (the flags are additive; pytest
continues to run normally). CAP-009 (offline pytest suite passes)
must remain Verified after the change.
- A6 (0.75): Infracost is available as a CLI tool that can be installed
in the CI environment and run locally. It reads `terraform plan
-out=plan.tfplan` + `terraform show -json plan.tfplan` to produce a
cost estimate. No live AWS access required. If Infracost is not
available, the `cost.estimated` event is omitted (degraded mode, not
a failure).
+1 -1
View File
@@ -8,7 +8,7 @@
],
"active_project": "acdl",
"active_projects": ["acdl"],
"active_milestone": "v1.16",
"active_milestone": "v1.17",
"autonomy": {
"level": "full",
"escalation_hooks": ["deploy", "delete_data", "merge_to_main"],