Compare commits
9 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 1db5ca8286 | |||
| 93ae9e4a39 | |||
| d3179fff37 | |||
| c029b102a3 | |||
| 2806c6c3ed | |||
| e048acd4dd | |||
| 9421442afd | |||
| bb43d94563 | |||
| ee5c372e65 |
@@ -734,148 +734,3 @@ stay **16/16 Verified** throughout the rebrand. P2/P3/P4 update test
|
||||
fixtures that reference `ACDL`/`acdl` so the gate stays green. No
|
||||
capability is added, removed, or reclassified in v1.15 — the rebrand is
|
||||
nomenclature + identifiers, not behavior.
|
||||
|
||||
---
|
||||
|
||||
## v1.16 Addendum — Nova Simplification (NFR, 2026-07-30)
|
||||
|
||||
The v1.16 NFR milestone added 6 new code components + 1 new Terraform
|
||||
module + 1 new schema, all documented here for the architecture record.
|
||||
|
||||
### New components
|
||||
|
||||
| Component | Path | Purpose |
|
||||
|-----------|------|---------|
|
||||
| Onboarding request handler | `core/onboarding.py` | `generate_env_file(request, template_env)` — produces a `<env>.json` from a consumer onboarding request (P19, REQ-183). CLI entry point for self-service env-file generation. |
|
||||
| Decommission transform | `core/decommission_transform.py` | `decommission_transform(stack)` — zero counts + disable deletion protection (REQ-92). Extracted from contract_resolver (P12, REQ-176). |
|
||||
| Contract resolver CLI | `core/contract_resolver_cli.py` | `main()` CLI entry point — resolves a contract YAML to a Target Stack JSON. Extracted from contract_resolver (P12, REQ-176). |
|
||||
| Regression verify CLI | `core/regression_verify_cli.py` | `main()` CLI entry point — runs the regression gate + writes the report. Extracted from regression_verify (P13, REQ-177). |
|
||||
| Workflow sync generator | `scripts/sync_workflows.py` | `--check`/`--write` — generates the 3 byte-identical Gitea+GitHub workflow pairs from `workflows-src/` (P8, REQ-172). |
|
||||
| Onboarding Terraform | `terraform/onboarding/` | `aws_iam_role.consumer_deploy` + `aws_iam_role_policy.consumer_invoke` (ABAC `nova:owner` tag). Offline-proven only (P20, REQ-184, D-114). |
|
||||
|
||||
### Modified components
|
||||
|
||||
| Component | Change | Phase |
|
||||
|-----------|--------|-------|
|
||||
| `core/contract_resolver.py` | `_load_env` delegates to `environment_check.load()` (dedup); `is_l2` uses registry `kind` field; `_load_schema` caches schemas; `decommission_transform` + CLI re-export shim (P12). | P7, P12, P14 |
|
||||
| `core/regression_verify.py` | Dedup helpers (`_check_resolver`, `_check_live_terraform_plan`, `_assert_contracts_resolve`); CAP-013..016 `Skipped` on post-teardown (G-111); `passed` accepts Skipped; CLI re-export shim (P13). | P5, P9, P13 |
|
||||
| `core/lambda/contract_ingestor.py` | Fail closed on missing IAM identity (P10); env enum from `core/environments/` (P10); payload size cap + schema validation (P11); `onboard_consumer` action (P18); `[NOVA-ALERT]` rebrand (P2). | P2, P10, P11, P18 |
|
||||
| `core/output_publisher.py` | `SAFE_OUTPUT_NAMES` schema-driven from `interface.json`; narrowed excepts; `urllib.error` import (P4, P14). | P4, P14 |
|
||||
| `core/environment_check.py` | Onboarding message rebranded Nova + self-service request path (P2, P19). | P2, P19 |
|
||||
| `core/local_emulators.py` | `LocalLambdaStub` sets `NOVA_LAMBDA_LOCAL_BYPASS`; stale dual-read comments + `acdl_*` prefixes removed (P3, P10). | P3, P10 |
|
||||
| `scripts/run_platform.sh` | `--help` flag; `run_hitl_gate()` fn; `NOVA_CONTRACT_ID`/`NOVA_WORK_DIR` config; decommission + uptime blocks extracted to sourced helpers (P6, P9, P15). | P6, P9, P15 |
|
||||
| `adapters/terraform/adapter.py` | State bucket `nova-tfstate-*` (P1); module docstring Nova (P2). | P1, P2 |
|
||||
| `adapters/kyverno/policies/require-resource-labels.yml` | `nova:*` labels (not `acdl:*`) (P1). | P1 |
|
||||
| `modules/registry.json` | `kind` field (`l1`/`l2`) on all 14 entries (P7). | P7 |
|
||||
|
||||
### New schema
|
||||
|
||||
- `schemas/onboarding.schema.json` — the self-service onboarding request
|
||||
(consumerRepo, requestedEnvironment, ownerId, billingTag). P18, REQ-182.
|
||||
|
||||
### Onboarding request-path architecture (D-113)
|
||||
|
||||
The no-humans onboarding flow is a 3-step request path (real AWS
|
||||
provisioning deferred):
|
||||
|
||||
```
|
||||
Consumer → POST Lambda (onboard_consumer) → pending CMDB row (P18)
|
||||
→ core/onboarding.py → <env>.json binding file (P19)
|
||||
→ terraform/onboarding/ → cross-account role + ABAC tag (P20, offline)
|
||||
```
|
||||
|
||||
The Lambda Function URL (IAM auth) + `consumer_invoke_policy.json` (ABAC
|
||||
`nova:owner`) are the transport; the request is accepted + a binding
|
||||
generated + the role Terraform proven offline. No AWS resources are
|
||||
created by the request path (D-113/D-114).
|
||||
|
||||
### Regression gate (G-111 binding)
|
||||
|
||||
The regression gate (D-091) now treats `Skipped` as acceptable for the
|
||||
post-v1.11-teardown steady state (D-096): CAP-013..016 (live-AWS tier)
|
||||
return `Skipped` when the resources are absent (`NoSuchBucket`/
|
||||
`ResourceNotFoundException`). `RegressionReport.passed` is
|
||||
`all(r.status in ("Verified", "Skipped"))`. The gate passes at 18
|
||||
Verified + 4 Skipped (0 Decayed/Broken).
|
||||
|
||||
## v1.17 Addendum — Strategic Direction, Leadership Metrics & Unified Story (2026-08-04)
|
||||
|
||||
The v1.17 milestone adds a telemetry/observability layer, a Decision
|
||||
Ledger, a metrics export pipeline, a unified narrative deck, and a
|
||||
durable strategic-direction artifact. This addendum documents the
|
||||
architecture; the full research findings are in RESEARCH.md §v1.17.
|
||||
|
||||
### New components
|
||||
|
||||
| Component | Path | Purpose |
|
||||
|-----------|------|---------|
|
||||
| Event envelope | `core/metrics/event_envelope.py` | CloudEvents 1.0 envelope + `platform.*` semantic conventions (P1, REQ-187) |
|
||||
| Per-run manifest writer | `core/metrics/run_manifest.py` | Emits `nova.run.started/completed/failed` events + writes `metrics/runs/<run_id>.json` (P1, REQ-187) |
|
||||
| Decision Ledger (SQLite) | `core/metrics/decision_ledger.py` | Extends `outbox_writer.py` → SQLite append-only hash-chain table; `ai.decision.made` + `attestation.recorded` events + outcome backfill (P1, REQ-188, D-121) |
|
||||
| Infracost post-processor | `core/metrics/infracost_adapter.py` | Runs Infracost on plan JSON; emits `nova.cost.estimated{delta_usd}` (P1, REQ-187, D-120) |
|
||||
| Metrics collector | `core/metrics/collector.py` | Reads all grounded signals (files + events) → SQLite cold store at `metrics/nova_metrics.db` (P2, REQ-189) |
|
||||
| PowerBI export | `core/metrics/powerbi_export.py` | Emits CSV/JSON views to `metrics/powerbi/` (fact + dim + 8 deferred placeholder views) (P3, REQ-190) |
|
||||
| Metrics schemas | `schemas/metrics_*.schema.json` | Schemas for all event types + fact/dim tables (P1–P2, REQ-187/189) |
|
||||
| Metrics catalog | `docs/METRICS.md` + `docs/metrics/<kpi>.md` | Canonical catalog + per-KPI definition-of-success docs (P4, REQ-195, D-127) |
|
||||
| Unified narrative deck | `docs/presentations/nova-no-humans-platform.md` | Merged deck: Problem→Vision→How→Proof→Roadmap; x3 arc at deck+slide level (P5, REQ-196/197, D-130) |
|
||||
| Strategic direction | `.ciagent/NORTH_STAR.md` | PO-authored durable vision/objectives/anti-goals/targets; read by CIAgent in every future `/ci-run` (P0, REQ-185/186) |
|
||||
|
||||
### Modified components
|
||||
|
||||
| Component | Change | Phase |
|
||||
|-----------|--------|-------|
|
||||
| `core/outbox_writer.py` | Extended to emit to SQLite append-only hash-chain table (Decision Ledger); `ai.decision.made` + `attestation.recorded` events added (P1, D-121) | P1 |
|
||||
| `scripts/run_platform.sh` | Per-run manifest writer invoked; `$WORK/*.json` persisted to `metrics/runs/`; Infracost post-processor invoked after plan (P1) | P1 |
|
||||
| `core/hitl_gates.py` | Emits `attestation.recorded` event to Decision Ledger on qa/prod/dr gate (P1, D-132) | P1 |
|
||||
| `core/confidence_signal.py` | Emits `nova.confidence.computed` + `nova.ai.decision.made` events (P1, D-122) | P1 |
|
||||
| `adapters/terraform/policy/checkov_adapter.py` | Emits `nova.policy.evaluated` event (P1) | P1 |
|
||||
| `core/regression_verify.py` | Emits `nova.capability.verified` event; CAP-023 (metrics collector) + CAP-024 (deck structure) added (P1, P6) | P1, P6 |
|
||||
| `pyproject.toml` | `addopts` gains `--junitxml=metrics/test-results.xml` + `--json-report` (P1, D-120) | P1 |
|
||||
| `docs/presentations/` | Two old decks retired (deleted); unified deck added (P5, D-130) | P5 |
|
||||
|
||||
### Telemetry/observability layer architecture (D-120)
|
||||
|
||||
```
|
||||
┌─────────────────────────────────────────────────────────────────────┐
|
||||
│ Nova platform components (existing) │
|
||||
│ run_platform.sh · confidence_signal · checkov_adapter · │
|
||||
│ hitl_gates · regression_verify · outbox_writer · contract_ingestor │
|
||||
└──────────────────────┬──────────────────────────────────────────────┘
|
||||
│ CloudEvents 1.0 envelope (new emitters, P1)
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────┐
|
||||
│ metrics/events.jsonl (append-only CloudEvents log) │
|
||||
│ metrics/runs/<run_id>.json (per-run manifests) │
|
||||
│ metrics/decision_ledger.db (SQLite hash-chain, D-121) │
|
||||
│ metrics/test-results.xml (junit, P1) │
|
||||
└──────────────────────┬──────────────────────────────────────────────┘
|
||||
│ collector reads (P2)
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────┐
|
||||
│ metrics/nova_metrics.db (SQLite cold store, D-126) │
|
||||
│ fact_run · fact_capability · fact_policy_check · fact_confidence │
|
||||
│ fact_test · fact_decision · fact_cost_estimate │
|
||||
│ dim_capability · dim_milestone │
|
||||
│ + 8 empty placeholder views (deferred metrics) │
|
||||
└──────────────────────┬──────────────────────────────────────────────┘
|
||||
│ powerbi_export (P3)
|
||||
▼
|
||||
┌─────────────────────────────────────────────────────────────────────┐
|
||||
│ metrics/powerbi/ (CSV/JSON views, folder connector, D-129) │
|
||||
│ → PowerBI dashboards (external) │
|
||||
└─────────────────────────────────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
**Hot path: deferred (D-126).** No live ops dashboard; SQLite is
|
||||
cold-only (batch/historical). The hot path activates when live AWS is
|
||||
re-provisioned (D-096 lift).
|
||||
|
||||
### NORTH_STAR integration point (REQ-186)
|
||||
|
||||
`.ciagent/NORTH_STAR.md` is read by CIAgent in context-loading for all
|
||||
future milestones. The integration mechanism (to be finalized in P4):
|
||||
a reference from `PROJECT.md` + `ARCHITECTURE.md` (this section) + a
|
||||
config entry in `config.json` (`strategic_direction_file:
|
||||
".ciagent/NORTH_STAR.md"`) that the run workflow reads at SPECIFY. This
|
||||
ensures the strategic direction survives across milestones without
|
||||
being overwritten by status updates.
|
||||
|
||||
@@ -462,92 +462,3 @@ status: complete
|
||||
phase_role: final
|
||||
audit: pass
|
||||
---/ci---
|
||||
|
||||
---
|
||||
|
||||
## v1.16 Post-Milestone Audit (2026-07-30)
|
||||
|
||||
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
CIAgent ► AUDIT REPORT
|
||||
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
|
||||
|
||||
**Reconstruction: PASS** — 4 commits since v1.15.4 base (787a649), 3 with
|
||||
`---ci---` blocks (1 merge commit without blocks, per convention — the
|
||||
squash-merge summary IS the record). Reconstructed state: phase 21,
|
||||
milestone v1.16, complete, tag v1.15.26, release 370, REQ-165..184
|
||||
covered. Matches CHECKPOINT.json + REQUIREMENTS.md + ROADMAP.md.
|
||||
|
||||
**.ciagent/ Files: 15 checked.**
|
||||
- config.json: valid JSON; active_milestone v1.16, active_project acdl,
|
||||
projects[] length 1. **PASS.**
|
||||
- PROJECT.md: v1.16 Objective (complete) + Key Decisions D-113..D-119
|
||||
present. 44 section headers. **PASS.**
|
||||
- ROADMAP.md: v1.16 section with P0–P21, all complete; tags v1.15.5..26.
|
||||
**PASS.**
|
||||
- REQUIREMENTS.md: v1.16 traceability 20/20 REQ-165..184 complete.
|
||||
**PASS.**
|
||||
- ARCHITECTURE.md: **FIXED DURING AUDIT** — 0 v1.16 references → v1.16
|
||||
addendum added (6 new components, 10 modified components, new schema,
|
||||
onboarding request-path architecture, regression gate G-111). **PASS
|
||||
(after fix).**
|
||||
- CHECKPOINT.json: valid JSON; phase=21, stage=complete,
|
||||
milestone_complete=true, tag=v1.15.26, release_id=370. **PASS.**
|
||||
- PERSONAS.md: v1.16 addendum present (8 references). **PASS.**
|
||||
- GRILL.md: v1.16 grill present (G-111..G-113, E-002). **PASS.**
|
||||
- RESEARCH.md: v1.16 addendum present (R1..R6). **PASS.**
|
||||
- PLAN.md: v1.16 20-phase + final plan present. **PASS.**
|
||||
- REVIEW.md: **FIXED DURING AUDIT** — 0 v1.16 references → reconstructed
|
||||
with v1.16 P21 final review content (0 P0, 0 P1, 2 P2 post-hoc). **PASS
|
||||
(after fix).**
|
||||
- AUDIT.md: this file (v1.16 audit recorded). **PASS.**
|
||||
- CAPABILITY_INVENTORY.md: not modified in v1.16 (no capability changes).
|
||||
**PASS.**
|
||||
- COST.md: not modified in v1.16 (no cost changes — offline-only). **PASS.**
|
||||
- IAM_POLICY.md: not modified in v1.16 (no IAM policy changes —
|
||||
onboarding Terraform is offline-proven, not applied). **PASS.**
|
||||
|
||||
**Branches: 0 v1.16 phase branches, 0 v1.16 milestone branches** (all
|
||||
cleaned up post-merge). Prior-milestone branches (v1.14 P1-P20, v1.11
|
||||
P56-P59) remain locally — historical, harmless, documented in ROADMAP.
|
||||
No v1.16 orphans. **PASS.**
|
||||
|
||||
**Commits: 4 total in v1.16 range, 3 with `---ci---` blocks, 1 merge
|
||||
commit without (per convention), 0 unresolved escalations.** The
|
||||
squash-merge strategy collapsed 20 phase branches + the milestone into
|
||||
the merge commit `f83b974`; the phase-level `---ci---` blocks lived in
|
||||
the (now-deleted) phase-branch commits. The milestone-level `---ci---`
|
||||
block (commit `58fa7a6`) records the final state. **PASS.**
|
||||
|
||||
**Audit Checks (runAuditChecks):**
|
||||
1. HEAD on main (milestone complete) — **PASS**
|
||||
2. CHECKPOINT.json exists — **PASS**
|
||||
3. CHECKPOINT consistent with latest `---ci---` (phase 21, v1.16,
|
||||
complete, v1.15.26, release 370) — **PASS**
|
||||
4. Report template exists (`opencode/ci/references/report-template.md`)
|
||||
— **PASS**
|
||||
5. No pending escalations (grill E-002 auto-resolved at P21; 0
|
||||
unresolved) — **PASS**
|
||||
6. Milestone version in config (v1.16) consistent with checkpoint —
|
||||
**PASS**
|
||||
|
||||
**Issues fixed during audit:**
|
||||
- ARCHITECTURE.md missing v1.16 addendum (0 references → added: 6 new
|
||||
components, 10 modified, new schema, onboarding architecture, G-111
|
||||
gate).
|
||||
- REVIEW.md held v1.11 content → reconstructed with v1.16 P21 final
|
||||
review (0 P0, 0 P1, 2 P2 post-hoc accepted).
|
||||
|
||||
**Verdict: PASS** — Project state is fully reconstructable from git log.
|
||||
All 6 audit checks pass. 2 auto-fixed issues (ARCHITECTURE.md addendum +
|
||||
REVIEW.md reconstruction) were file-discipline gaps, not structural
|
||||
defects. 20/20 requirements complete; regression gate 18V+4S; milestone
|
||||
merged to main; tag v1.15.26; release 370.
|
||||
|
||||
---ci---
|
||||
project: acdl
|
||||
phase: 21
|
||||
milestone: v1.16
|
||||
status: complete
|
||||
phase_role: final
|
||||
audit: pass
|
||||
---/ci---
|
||||
|
||||
@@ -1,13 +1,9 @@
|
||||
{
|
||||
"phase": 0,
|
||||
"stage": "complete",
|
||||
"milestone": "v1.17",
|
||||
"phase_role": "pre_execution",
|
||||
"phase": 1,
|
||||
"stage": "execute",
|
||||
"milestone": "v1.16",
|
||||
"phase_role": "execution",
|
||||
"attempts": 0,
|
||||
"updated_at": "2026-08-04T21:30:00Z",
|
||||
"milestone_complete": false,
|
||||
"tag": "v1.16.0",
|
||||
"release_id": 441,
|
||||
"requirements": ["REQ-185"],
|
||||
"notes": "Phase 0 complete. NORTH_STAR.md authored. 29 requirements (REQ-185..213). Telemetry reference architecture + metric scorecard. Deck rebuild plan (18 slides). Interactive GRILL: 12 binding decisions applied. Tag v1.16.0 pushed. Gitea release 441 created. Ready for execution phases P1..P7 + final P8."
|
||||
"updated_at": "2026-07-30T15:30:00Z",
|
||||
"milestone_complete": false
|
||||
}
|
||||
@@ -636,262 +636,3 @@ re-provision the bucket.
|
||||
YES, once G-111's criterion restatement + gate update are incorporated
|
||||
(into P9's must-haves). G-112/G-113 are phase-entry clarifications for
|
||||
P9/P12/P13. E-002 is deferred to P21. Confidence 0.85.
|
||||
|
||||
---
|
||||
|
||||
# GRILL — v1.17 "Strategic Direction, Leadership Metrics & Unified Story" (2026-08-04)
|
||||
|
||||
> **Griller:** CIAgent (red-team mode). **Milestone:** v1.17. **Axes:** 3
|
||||
> (NORTH_STAR alignment, Deck story & arc, Deck per-slide rigor) per PO
|
||||
> direction. **Stance:** adversarial — presumed over-scoped / infeasible /
|
||||
> storytelling-weak until evidence forced otherwise.
|
||||
|
||||
## Evidence base
|
||||
|
||||
- `NORTH_STAR.md` (183 lines, draft), `PLAN.md` (1,114 lines, deck rebuild
|
||||
plan incl. slide-by-slide), `REQUIREMENTS.md` v1.17 (REQ-185..213),
|
||||
`RESEARCH.md` v1.17 (signal inventory, scorecard, deferred-decision
|
||||
ledger, deck research).
|
||||
- Codebase cross-checks: `REGRESSION_REPORT.json` = **18 Verified + 4
|
||||
Skipped** (NOT "22/22 Verified" — the new deck plan correctly says
|
||||
18V+4S; the *existing* decks still claim 22/22). `PROJECT.md:495` =
|
||||
**0 consumer adoption**. `docs/NO_HUMANS_THESIS.md`, `docs/METRICS.md`,
|
||||
`metrics/` do not yet exist (P4/P5 deliverables — expected).
|
||||
- Decisions locked (D-120..D-132) — not re-litigated.
|
||||
|
||||
## The central contradiction
|
||||
|
||||
**NORTH_STAR.md:111** states: *"Targets are committed, not aspirational."*
|
||||
**PO's G-Q6 answer:** *"the goal is simply to target a high touchless
|
||||
resolution rate, not to say we have reached those targets given there are
|
||||
0 consumers."*
|
||||
|
||||
These two statements are in direct conflict. "Committed, not aspirational"
|
||||
+ "simply to target" = the document is lying about its own epistemic
|
||||
status. This is the v1.10 decay root cause (PRE_MORTEM FM-3: decks
|
||||
outrunning verified reality) repeating itself in the document meant to
|
||||
prevent it.
|
||||
|
||||
## Axis 1 — NORTH_STAR alignment
|
||||
|
||||
### G-Q1 — Target with no backing REQ / placeholder
|
||||
**Finding:** AI-Agent Intent Share (≥40%) is a committed 12–18mo target
|
||||
(NORTH_STAR:128) with "placeholder view" claimed, but it is NOT among the
|
||||
8 placeholder views in PLAN P3 (lines 309–315), and no REQ-185..213 builds
|
||||
an emitter or placeholder for it. RESEARCH §3 marks it "future" with no
|
||||
controlling decision ID (unlike every other deferred metric). NORTH_STAR:128
|
||||
falsely claims a placeholder view exists → violates the "no fabrication"
|
||||
hard constraint.
|
||||
**Verdict: BIND.** Add a 9th placeholder view OR move the target to a
|
||||
"Future Horizons" section; correct NORTH_STAR:128. **Confidence: 0.90.**
|
||||
|
||||
### G-Q2 — Anti-goal pursuit
|
||||
**Finding:** No REQ builds an anti-goal. Deck title "No-Humans Infrastructure
|
||||
Platform" is one weak slide away from violating anti-goal #3 (not removing
|
||||
humans from accountability) — mitigation is entirely in slide 3's execution.
|
||||
**Verdict: PASS (conditional on slide 3 landing the attestation model).**
|
||||
**Confidence: 0.75.**
|
||||
|
||||
### G-Q3 — Attestation clarification consistency
|
||||
**Finding:** The attestation clarification is the most consistently
|
||||
propagated concept in the plan — NORTH_STAR (3 places), REQUIREMENTS
|
||||
(3 REQs), deck (3 slides). Well done.
|
||||
**Verdict: PASS.** **Confidence: 0.92.**
|
||||
|
||||
### G-Q4 — "AI decision" framing (D-122 honesty)
|
||||
**Finding:** D-122 (confidence_signal + HITL gate, NOT an LLM) is cited on
|
||||
slide 7 and required in NO_HUMANS_THESIS.md (REQ-213). BUT slide 7's
|
||||
*Delivers* says "every AI decision captured" without ever telling the
|
||||
audience what the "AI" is. The honesty is buried in a linked doc + a
|
||||
decision ID the audience has never heard.
|
||||
**Verdict: BIND.** Add one sentence to slide 7 *Delivers*: "Nova's 'AI
|
||||
decision' is the confidence-gated policy engine, not an LLM planner
|
||||
(D-122)." **Confidence: 0.85.**
|
||||
|
||||
### G-Q5 — Secretly ungrounded metrics
|
||||
**Finding:** The 8 deferred placeholder views cover their list. BUT (a)
|
||||
AI-Agent Intent Share's placeholder is falsely claimed (G-Q1), and (b)
|
||||
derived metrics (FTE Hours Saved, Platform ROI) are computed on zero
|
||||
production runs yet shown on slide 12 without the zero-denominator caveat.
|
||||
A "derived" metric from zero runs is technically not fabricated but is
|
||||
misleading.
|
||||
**Verdict: BIND.** (1) Resolve G-Q1; (2) slide 12 must annotate derived
|
||||
metrics with "(computed on N internal runs; production-denominator activates
|
||||
post-pilot)." **Confidence: 0.82.**
|
||||
|
||||
### G-Q6 — 12–18mo target feasibility (0 consumers)
|
||||
**Finding:** PO's answer ("simply to target") conflicts with NORTH_STAR:111
|
||||
("committed, not aspirational"). 3 "grounded (after P1)" targets (Touchless
|
||||
Resolution, Human Escalation, AI Decision Accuracy) have scope "across
|
||||
production estates" — but PROJECT.md:495 = 0 consumer adoption. The metric
|
||||
IS computable on internal dev runs, but the target scope doesn't exist.
|
||||
Marking "grounded" while the scope is absent is the overclaim the "no
|
||||
fabrication" constraint exists to prevent.
|
||||
**Verdict: BIND.** (1) Rewrite NORTH_STAR:111 → "Targets are committed
|
||||
destinations; the grounding column records whether each is measurable this
|
||||
milestone." (2) Reclassify the 3 targets to `partial — measurement pipeline
|
||||
grounded on internal runs; production-estate scope activates post-pilot`
|
||||
(the Cloud Spend Reduction precedent at NORTH_STAR:123). (3) Deck slide 5
|
||||
regroup as "Measurable today (internal runs)" vs "Activates post-pilot
|
||||
(production estates)." Requires NORTH_STAR-CHANGE commit trailer (REQ-204).
|
||||
**Confidence: 0.80.**
|
||||
|
||||
## Axis 2 — Deck plan: story & arc
|
||||
|
||||
### G-Q7 — Arc order (Problem→Vision→How→Proof→Roadmap vs Proof-first)
|
||||
**Finding:** Current arc puts Proof at Act 4 (slides 10–13) — 40% of the
|
||||
deck before a number. For a leadership audience that has seen 10+ milestone
|
||||
decks, this risks losing the room by slide 4. BUT the "no-humans" thesis
|
||||
is contentious; jumping to proof without the attestation model invites the
|
||||
"removing humans from accountability" objection. The Vision act makes the
|
||||
Proof credible.
|
||||
**Verdict: PASS (marginal).** Defensible IF the Problem act is tight and
|
||||
slide 3 front-loads the attestation clarification. **Confidence: 0.62.**
|
||||
|
||||
### G-Q8 — x3 structure at deck level
|
||||
**Finding:** Slide 1's 5-act preview is orienting (a table of contents),
|
||||
not too much meta-structure. BUT it's also not a hook — it gives structure,
|
||||
not stakes. A C-suite audience decides in the first 30 seconds.
|
||||
**Verdict: BIND (minor).** Add one stake-establishing line to slide 1
|
||||
*Delivers* with a real number (18 verified, 0 consumers, honest deferral
|
||||
list). **Confidence: 0.70.**
|
||||
|
||||
### G-Q9 — Per-slide benefit callouts (substantive vs filler)
|
||||
**Finding:** 4 of 17 closes are filler (slides 1, 4, 12, 15); 2 borderline
|
||||
(8, A1). Worst offender: slide 12 (ROI) restates the *objective* ("ROI is
|
||||
quantifiable") rather than giving the *number* or the *honest caveat*.
|
||||
**Verdict: BIND.** Rewrite 4 filler closes. Slide 12's close must be:
|
||||
"Benefit: you now know the ROI formula — (labor + cloud + avoided downtime)
|
||||
÷ platform cost — and that it computes on internal runs today, with
|
||||
production-denominator activating post-pilot." **Confidence: 0.78.**
|
||||
|
||||
### G-Q10 — Deck length (17 slides)
|
||||
**Finding:** 17 is at the upper bound but justifiable for 5 acts. The risk
|
||||
is density, not length: slide 12 crams 6 metrics (Touchless, Human
|
||||
Escalation, MTTR, Cost, FTE, ROI) into one slide — a wall of bullets.
|
||||
**Verdict: BIND (minor).** Split slide 12 into "Zero-Touch Efficiency"
|
||||
(Touchless, Human Escalation, MTTR) + "Cost & ROI" (Cost, FTE, ROI). Deck
|
||||
→ 18 slides, each earning its place. **Confidence: 0.68.**
|
||||
|
||||
### G-Q11 — "What's Deferred" slide (13)
|
||||
**Finding:** The honesty strengthens the grounded claims BUT surfaces the
|
||||
gap: Nova claims "no-humans in operations" while deferring the metrics
|
||||
that would prove operations are healthy without humans (Live Infra Health,
|
||||
SLA, Drift Auto-Reversal). A skeptical viewer notes the contradiction.
|
||||
**Verdict: BIND.** Add preempt to slide 13: "These deferrals are about
|
||||
*measurement infrastructure*, not about whether the platform runs without
|
||||
humans — the platform runs autonomously today on internal runs; what's
|
||||
deferred is the production-estate dashboard that would prove it at scale."
|
||||
**Confidence: 0.75.**
|
||||
|
||||
## Axis 3 — Deck plan: per-slide rigor
|
||||
|
||||
### G-Q12 — Slide opening lines
|
||||
**Finding:** The "This slide shows X" formula is orienting, not patronizing,
|
||||
because each includes a stake-bearing clause. Consistent without being empty.
|
||||
**Verdict: PASS.** **Confidence: 0.80.**
|
||||
|
||||
### G-Q13 — Transitions (written vs hand-waved)
|
||||
**Finding:** ~10 of 13 transitions are written (specific reference to prior
|
||||
close). 3 are hand-waved (slides 8→9, 11→12, 13→14). Worst: the Act 3→4
|
||||
boundary (slide 8→9, How→Proof) — the most important transition in the deck
|
||||
— is the weakest.
|
||||
**Verdict: BIND.** Rewrite the 3 hand-waved transitions. The 8→9 Act
|
||||
boundary must carry weight: "Having seen the gate model — autonomy in
|
||||
operations, human in accountability — here is how Nova instruments itself
|
||||
to prove that model at scale." **Confidence: 0.85.**
|
||||
|
||||
### G-Q14 — Weakest slide (audience-loss point)
|
||||
**Finding:** Slide 9 (Telemetry Architecture) is the audience-loss slide.
|
||||
It's the 4th consecutive architecture slide (6,7,8,9), the most abstract
|
||||
(CloudEvents, SQLite, PowerBI), its Benefit is about data plumbing not
|
||||
business value, and it sits between the attestation matrix (slide 8,
|
||||
emotionally resonant) and the Proof act (slide 10, the numbers) — between
|
||||
the two things the audience came for.
|
||||
**Verdict: BIND.** Compress slide 9 into slide 10 OR reframe its Benefit
|
||||
from data plumbing to trust: "Benefit: you now know the proof you're about
|
||||
to see isn't fabricated — every number traces to a file you can audit."
|
||||
**Confidence: 0.78.**
|
||||
|
||||
### G-Q15 — Proof act citation specificity
|
||||
**Finding:** 5 of 6 Proof citations are specific (file paths + real numbers).
|
||||
Gap: slide 12's derived metrics (FTE, ROI) cite "derived" without showing
|
||||
the formula or the input count.
|
||||
**Verdict: BIND (minor).** Show the ROI formula inline on slide 12 + the
|
||||
N=0 production-runs caveat. **Confidence: 0.80.**
|
||||
|
||||
### G-Q16 — Closing slide (15) — does the ask land?
|
||||
**Finding:** THE ask is present but framed as insider language ("fund the
|
||||
hot-path activation (post-D-096) + the tamper-evident ledger build-out
|
||||
(D-083 lift)"). A leadership audience doesn't know what "hot-path
|
||||
activation" means. The ask is a technical request, not a business decision
|
||||
a leader can make in the room.
|
||||
**Verdict: BIND.** Reframe slide 15's ask as a business decision: "The
|
||||
ask: (1) approve a pilot estate to activate production-estate metrics
|
||||
(unblocks D-096), and (2) approve the tamper-evident ledger build-out
|
||||
(lifts D-083) — turning grounded claims into complete proof." Make it a
|
||||
yes/no a leader can give. **Confidence: 0.82.**
|
||||
|
||||
## Binding decisions (must resolve before SHIP)
|
||||
|
||||
| G-ID | Axis | Verdict | What must change | Conf |
|
||||
|---|---|---|---|---|
|
||||
| G-Q1 | 1 | BIND | Add 9th placeholder view for AI-Agent Intent Share OR move to "Future Horizons"; correct NORTH_STAR:128 | 0.90 |
|
||||
| G-Q4 | 1 | BIND | Add D-122 honesty sentence to slide 7 *Delivers* | 0.85 |
|
||||
| G-Q5 | 1 | BIND | Annotate derived metrics on slide 12 with zero-run caveat | 0.82 |
|
||||
| G-Q6 | 1 | BIND | Rewrite NORTH_STAR:111; reclassify 3 targets to `partial`; regroup deck slide 5. NORTH_STAR-CHANGE trailer required | 0.80 |
|
||||
| G-Q8 | 2 | BIND (minor) | Add stake line with real number to slide 1 *Delivers* | 0.70 |
|
||||
| G-Q9 | 2 | BIND | Rewrite 4 filler closes (slides 1, 4, 12, 15); slide 12 must give ROI formula + caveat | 0.78 |
|
||||
| G-Q10 | 2 | BIND (minor) | Split slide 12 into two (Efficiency + Cost/ROI); deck → 18 slides | 0.68 |
|
||||
| G-Q11 | 2 | BIND | Add preempt to slide 13 (deferrals are measurement infra, not whether platform runs without humans) | 0.75 |
|
||||
| G-Q13 | 3 | BIND | Rewrite 3 hand-waved transitions (esp. Act 3→4 boundary 8→9) | 0.85 |
|
||||
| G-Q14 | 3 | BIND | Compress slide 9 into slide 10 OR reframe its Benefit to trust | 0.78 |
|
||||
| G-Q15 | 3 | BIND (minor) | Show ROI formula inline + N=0 caveat on slide 12 | 0.80 |
|
||||
| G-Q16 | 3 | BIND | Reframe slide 15 ask as a business decision (pilot estate + ledger build-out) | 0.82 |
|
||||
|
||||
**PASS (no change):** G-Q2 (anti-goals, conditional on slide 3), G-Q3
|
||||
(attestation consistency — excellent), G-Q7 (arc order — marginal),
|
||||
G-Q12 (slide openings — formulaic but substantive).
|
||||
|
||||
## Escalations (only the PO can decide)
|
||||
|
||||
| E-ID | Question | Confidence |
|
||||
|---|---|---|
|
||||
| E-003 | Should the 3 "grounded (after P1)" targets with "production estates" scope be reclassified to `partial` (Cloud Spend precedent), or should "grounded" be redefined to mean "measurement pipeline grounded"? Changes a committed NORTH_STAR target's grounding label; requires NORTH_STAR-CHANGE trailer (REQ-204). | <0.60 |
|
||||
| E-004 | Should AI-Agent Intent Share (≥40%) remain a "12–18mo Target" with no backing REQ/placeholder, or move to a "Future Horizons" section? Strategic-scope question (is agentic consumption a 12–18mo commitment or a longer horizon?). | <0.60 |
|
||||
|
||||
## Overall verdict
|
||||
|
||||
**🟡 REDUCE SCOPE / BINDING FIXES REQUIRED — not ready to ship as-is.**
|
||||
|
||||
The plan is architecturally sound (metrics pipeline, Decision Ledger,
|
||||
PowerBI export, x3 deck structure are well-designed and grounded). The
|
||||
attestation clarification (G-Q3) is the best-propagated concept in the
|
||||
plan. The regression-capability gate (CAP-023/024) is a credible safeguard.
|
||||
|
||||
But the plan has one structural contradiction (NORTH_STAR:111 vs PO intent
|
||||
vs grounding column) that infects 4 other findings (G-Q1, G-Q5, G-Q6,
|
||||
G-Q9/slide 12). This is the v1.10 decay pattern (PRE_MORTEM FM-3)
|
||||
repeating in the document meant to prevent it. The "no fabrication" hard
|
||||
constraint is self-violated in two places (AI-Agent Intent Share placeholder
|
||||
claim, derived-metrics-without-caveat) before a single slide is rendered.
|
||||
|
||||
The deck plan is story-competent but not story-excellent. 4 benefit
|
||||
callouts are filler, 3 transitions are hand-waved (incl. the critical
|
||||
Act 3→4 boundary), slide 9 is the audience-loss slide, and the closing
|
||||
ask is insider language.
|
||||
|
||||
**12 binding decisions, 2 escalations.** None require re-architecting the
|
||||
plan; all are edits to NORTH_STAR (2 rows + 1 line, with commit trailer),
|
||||
the deck slide plan (4 slide rewrites, 1 split, 3 transition rewrites),
|
||||
and one placeholder-view addition. Estimate: 1–2 phases of rework, not a
|
||||
milestone restart. The plan does NOT need a revision loop — it needs
|
||||
these 12 fixes applied in P0 (NORTH_STAR) and P5 (deck) before the
|
||||
respective phases ship. Critical path unchanged.
|
||||
|
||||
**Can the milestone proceed?**
|
||||
|
||||
YES, once the 12 BIND decisions are incorporated (G-Q1/Q4/Q5/Q6 into P0
|
||||
NORTH_STAR + P5 deck plan; G-Q8/Q9/Q10/Q11/Q13/Q14/Q15/Q16 into P5 deck
|
||||
plan). E-003/E-004 require PO decisions on NORTH_STAR target framing.
|
||||
Confidence 0.80.
|
||||
|
||||
@@ -1,211 +0,0 @@
|
||||
# NORTH_STAR — Nova
|
||||
|
||||
> **Status:** Draft (pending interactive GRILL → final)
|
||||
> **Milestone:** v1.17 — Strategic Direction, Leadership Metrics & Unified Story
|
||||
> **Owner:** Product Owner
|
||||
> **Purpose:** Durable strategic intent. Read by CIAgent in every future
|
||||
> `/ci-run` so the platform's direction survives across milestones. This
|
||||
> is NOT a status document (that's PROJECT.md) and NOT an engineering
|
||||
> architecture (that's the telemetry reference in RESEARCH.md/
|
||||
> ARCHITECTURE.md). It is the PO's committed direction: what we're
|
||||
> building toward, what we refuse to build, and how we'll know we won.
|
||||
|
||||
---
|
||||
|
||||
## Vision
|
||||
|
||||
> **Infrastructure operations become invisible. Every environment
|
||||
> provisioned, every incident healed, every risk remediated — by an
|
||||
> autonomous system whose trustworthiness is provable, not promised.
|
||||
> Human attestation remains required at stage gates — QA signs off for
|
||||
> production, SRE greenlights based on operational readiness — but the
|
||||
> operator is never in the loop of normal operations.**
|
||||
|
||||
Nova is the autonomous infrastructure layer that lets product teams ship
|
||||
without engaging an operator, and lets executives trust the AI not because
|
||||
it never fails but because every decision is captured, scored, and
|
||||
accountable.
|
||||
|
||||
---
|
||||
|
||||
## Strategic Objectives (4)
|
||||
|
||||
**1. Demonstrate production-grade zero-touch operations.**
|
||||
Nova must run real customer estates with no human in the loop of normal
|
||||
operations — autonomy as the default, not the demo. Stage-gate
|
||||
attestation (QA for production, SRE for operational readiness) remains
|
||||
human by design; operational escalations (AI confidence too low to
|
||||
proceed) are the failure mode we drive toward zero. Everything else
|
||||
collapses if autonomy isn't real.
|
||||
|
||||
**2. Establish provable trust in AI decisions.**
|
||||
Build the audit substrate — Decision Ledger, confidence scoring, circuit
|
||||
breakers, blast-radius controls — that turns "autonomous" from a
|
||||
marketing claim into a defensible one. Trust is the moat. Features can be
|
||||
copied; an immutable, queryable decision history cannot.
|
||||
|
||||
**3. Deliver compounding, quantifiable ROI for customers.**
|
||||
Each quarter on Nova must reduce cloud spend, free engineering hours, and
|
||||
avoid downtime measurably. If the CFO can't point to a number that
|
||||
improves quarter-over-quarter, Nova fails its commercial test, regardless
|
||||
of how clever the AI is.
|
||||
|
||||
**4. Become the default substrate for agentic infrastructure consumption.**
|
||||
AI agents are already becoming the largest consumers of cloud
|
||||
infrastructure. Nova must be the platform through which those agents
|
||||
declare, deploy, and verify infrastructure — not a vendor scrambling into
|
||||
that market two quarters late.
|
||||
|
||||
---
|
||||
|
||||
## Anti-Goals (5 — what Nova is fundamentally NOT)
|
||||
|
||||
1. **Not a Terraform, Kubernetes, or hyperscaler competitor.** We
|
||||
orchestrate them. Replacing them is the most expensive possible
|
||||
distraction from the value we create.
|
||||
2. **Not a general-purpose AI agent platform.** We are purpose-built for
|
||||
infrastructure operations. Breadth here produces shallow tools; depth
|
||||
here wins the category.
|
||||
3. **Not a system that removes humans from accountability.** Only from
|
||||
operations. Every AI decision lands in an immutable ledger. Every
|
||||
stage-gate promotion (qa/prod/dr) requires a human attestation recorded
|
||||
with approver identity, separation-of-duties check, and the 8-concern
|
||||
evidence matrix. The absence of an operator is never the absence of a
|
||||
record.
|
||||
4. **Not for legacy, untagged, or freeform infrastructure.** Nova requires
|
||||
Terraform-managed, policy-aligned, fully-tagged inputs. We optimize for
|
||||
the disciplined 95%, not the chaotic 5%.
|
||||
5. **Not sold to operators.** Nova is sold to leadership on outcomes —
|
||||
cost, velocity, risk. Selling to operators inverts the incentive and
|
||||
breaks the autonomy thesis.
|
||||
|
||||
---
|
||||
|
||||
## Non-Goals (v1.17 milestone scope — deferred work, not permanent boundaries)
|
||||
|
||||
> Anti-Goals are what Nova *fundamentally is not*. Non-Goals are what we
|
||||
> *will not do this milestone* — deferred work, not permanent boundaries.
|
||||
> Each Non-Goal cites the controlling decision ID.
|
||||
|
||||
1. **Live AWS re-provisioning** (deferred — D-096). Metrics that require
|
||||
live infrastructure ship as placeholder PowerBI views with documented
|
||||
schemas.
|
||||
2. **Onboarding auto-grant** (deferred — D-113/D-114/D-119). Only the
|
||||
request-path metric is grounded; the requested→granted funnel is a
|
||||
placeholder.
|
||||
3. **ML anomaly-forecasting / predictive remediation** (no emitter today).
|
||||
The Predictive-vs-Reactive metric ships as a placeholder.
|
||||
4. **Drift detection scheduled job** (deferred — D-096 + no scheduler).
|
||||
Drift metrics ship as placeholders.
|
||||
5. **Live cost CUR reconciliation** (deferred — D-096). Pre-apply Infracost
|
||||
estimates are grounded; actual-spend reconciliation is a placeholder.
|
||||
6. **S3 Object Lock / JWS tamper-evident ledger** (deferred — D-083). The
|
||||
Decision Ledger uses a local SQLite hash-chain this milestone; the
|
||||
Object-Lock/JWS build-out is a future milestone.
|
||||
7. **Multi-cloud support** (Azure/GCP/K8s). Nova is AWS-only this milestone.
|
||||
|
||||
---
|
||||
|
||||
## 12–18 Month Targets
|
||||
|
||||
Targets are committed, not aspirational. Each is a number a board member
|
||||
can repeat back to us. The grounding column records whether the metric is
|
||||
measurable this milestone, and if not, what blocks it.
|
||||
|
||||
> **Honesty note (GRILL G-Q6 binding):** Nova has 0 consumer adoption
|
||||
> today (`PROJECT.md:495`). Three targets (Touchless Resolution, Human
|
||||
> Escalation, AI Decision Accuracy) are scoped "across production
|
||||
> estates" — the measurement *pipeline* is grounded this milestone, but
|
||||
> the *denominator* is zero until a pilot estate activates. These
|
||||
> targets are reclassified as **Post-Pilot** (the pipeline works; the
|
||||
> numbers fill when consumers exist). This is the same honesty model as
|
||||
> Cloud Spend Reduction (partial: pipeline grounded, actuals deferred).
|
||||
|
||||
### Current-milestone targets (grounded or derived this milestone)
|
||||
|
||||
| Domain | Target | Grounding (v1.17) | Note |
|
||||
|---|---|---|---|
|
||||
| **MTTR (p95)** | < 60 seconds | grounded (platform-run MTTR) | apply.failed → successful retry; infra-incident MTTR deferred (no incident detection) |
|
||||
| **Cloud Spend Reduction** | ≥ 25% on pilot estates vs. 12-month pre-Nova baseline | partial | pre-apply estimate grounded (Infracost); actual-spend deferred (D-096 CUR) |
|
||||
| **L1 / L2 Ops Hours Avoided** | ≥ 70% of pre-Nova FTE allocation | derived | formula over run count × manual baseline (computed on N internal runs; production-denominator activates post-pilot) |
|
||||
| **Platform ROI** | ≥ 250% measured annually | derived | formula (labor savings + cloud savings + avoided downtime) ÷ platform op cost (computed on N internal runs; production-denominator activates post-pilot) |
|
||||
| **Decision Ledger Coverage** | 100% of AI actions with backfilled outcome | grounded (this milestone builds it) | outbox_writer.py → SQLite hash-chain |
|
||||
| **Attestation Coverage** | 100% of prod/dr promotions attested by a human | grounded | hitl_gates.py + outbox approver_* attributes; separation-of-duties on prod |
|
||||
|
||||
### Post-Pilot targets (pipeline grounded this milestone; denominator activates when a pilot estate runs)
|
||||
|
||||
| Domain | Target | Grounding (v1.17) | Note |
|
||||
|---|---|---|---|
|
||||
| **Touchless Resolution Rate** | ≥ 99% across production estates | partial (pipeline grounded; denominator = 0 today) | runs completing without *operational* HITL block ÷ total runs (attestation gates excluded); activates post-pilot |
|
||||
| **Human Escalation Frequency** | < 0.1% of platform actions | partial (pipeline grounded; denominator = 0 today) | *operational* HITL blocks only (confidence-driven); attestation sign-offs excluded; activates post-pilot |
|
||||
| **AI Decision Accuracy** | ≥ 99.5% (no rollback, no follow-up incident within 5 min of action) | partial (pipeline grounded; denominator = 0 today) | decisions not followed by apply.failed/incident within 5min; activates post-pilot |
|
||||
|
||||
### Deferred targets (measurement requires future systems)
|
||||
|
||||
| Domain | Target | Grounding (v1.17) | Note |
|
||||
|---|---|---|---|
|
||||
| **Predictive vs. Reactive Ratio** | ≥ 3 : 1 (prevention dominates reaction) | deferred | requires ML forecasting service (future emitter) |
|
||||
| **Drift Auto-Reversal Rate** | ≥ 95% within one detection cycle | deferred | requires drift detection (D-096 + scheduler) |
|
||||
|
||||
> Committed targets whose measurement is deferred remain committed — the
|
||||
> target is the destination; the metric is the odometer, and some
|
||||
> odometers aren't built yet. Each deferred metric ships as a placeholder
|
||||
> PowerBI view + a definition-of-success doc recording the dependency.
|
||||
> Post-Pilot targets are committed targets whose measurement pipeline is
|
||||
> grounded this milestone; the numbers activate when a pilot estate runs.
|
||||
|
||||
### Future Horizons (strategic direction, not committed targets)
|
||||
|
||||
| Domain | Aspiration | Note |
|
||||
|---|---|---|
|
||||
| **AI-Agent Intent Share** | ≥ 40% of total intent volume originated by non-human consumers | Strategic Objective #4 direction. No backing requirement, no placeholder view, no emitter today. Moves to a committed target when agentic consumption is real. |
|
||||
|
||||
---
|
||||
|
||||
## Success Criteria (v1.17 — what constitutes success for THIS milestone)
|
||||
|
||||
> Distinct from the 12–18mo targets: those are the destination. These are
|
||||
> the milestone's exit criteria.
|
||||
|
||||
v1.17 is a success if:
|
||||
|
||||
1. **Decision Ledger emits `ai.decision.made` for 100% of platform runs**
|
||||
with outcome backfill, AND **`attestation.recorded` events for 100%
|
||||
of qa/prod/dr promotions** (event completeness — all 3 gates captured;
|
||||
grounded in `outbox_writer.py` → SQLite hash-chain; honors D-083).
|
||||
The **Attestation Coverage metric** (target 100%) measures prod/dr
|
||||
promotions specifically — see REQ-194.
|
||||
2. **`docs/METRICS.md` catalogs every executive KPI** with a `grounded` /
|
||||
`derived` / `deferred` status, a source file or decision ID, and a
|
||||
per-KPI definition-of-success doc in `docs/metrics/`.
|
||||
3. **The PowerBI export produces all fact/dimension views** + 8 empty
|
||||
placeholder views for deferred metrics (with documented schemas ready
|
||||
to fill when their blocking decisions lift).
|
||||
4. **The unified narrative deck ships** with the x3 arc
|
||||
(Problem→Vision→How→Proof→Roadmap) at deck + slide level, per-slide
|
||||
benefit callouts, and fluid transitions; both old decks retired.
|
||||
5. **`NORTH_STAR.md` is wired into CIAgent context-loading** so every
|
||||
future `/ci-run` reads it.
|
||||
6. **CAP-023 (metrics collector) + CAP-024 (deck structure) pass** in the
|
||||
regression gate.
|
||||
|
||||
---
|
||||
|
||||
## What "won" looks like
|
||||
|
||||
By month 18, Nova is the layer enterprise leadership points to when they
|
||||
say *"we don't have an infrastructure ops team anymore, and the audit
|
||||
trail is stronger than it ever was"* — and it is the default substrate
|
||||
their AI engineering teams reach for first when an agent needs to deploy.
|
||||
|
||||
---
|
||||
|
||||
## Relationship to v1.17 engineering
|
||||
|
||||
- **Pillar A (this file):** strategic direction — durable, PO-authored.
|
||||
- **Pillar B (engineering):** the telemetry reference architecture
|
||||
(adapted from the PO's technical-direction input) lives in
|
||||
RESEARCH.md/ARCHITECTURE.md. It is the *how*; this file is the *why*.
|
||||
- **Pillar C (story):** the unified narrative deck proves Pillars A+B to
|
||||
leadership. The deck's Proof section cites grounded metrics; its
|
||||
Roadmap section cites deferred targets honestly.
|
||||
+13
-145
@@ -1,24 +1,23 @@
|
||||
---
|
||||
project: acdl
|
||||
milestone: v1.17
|
||||
generated_at: 2026-08-04
|
||||
milestone: v1.16
|
||||
generated_at: 2026-07-30
|
||||
generator: lead-developer
|
||||
verification_toolchain:
|
||||
typecheck: "terraform validate && python3 -m py_compile core/**/*.py && python3 -m jsonschema schemas/*.schema.json"
|
||||
test: "bash scripts/run_regression.sh # 22-capability gate (D-091/D-118) + CAP-023/024 (v1.17)"
|
||||
test: "bash scripts/run_regression.sh # 22-capability gate (D-091/D-118)"
|
||||
build: "bash scripts/run_ci.sh # full local CI reproduction (lint+test+check-only)"
|
||||
note: |
|
||||
v1.17 adds a telemetry/observability layer (metrics emitters, SQLite
|
||||
cold store, PowerBI export, Decision Ledger) + a unified narrative
|
||||
deck + a durable NORTH_STAR.md. Three active personas: lead-developer
|
||||
(coordination + deck narrative co-author), backend-engineer (event
|
||||
emitters, outbox_writer extension, Infracost adapter), data-engineer
|
||||
(SQLite store, schemas, PowerBI views, metrics collector). frontend-
|
||||
engineer stays deactivated (no Nova web UI — dashboards are PowerBI,
|
||||
not a Nova-built frontend; decks are markdown = lead-developer
|
||||
territory). No new custom personas needed — the metrics domain maps
|
||||
cleanly to data-engineer (schema/store/export) + backend-engineer
|
||||
(emitters/instrumentation).
|
||||
Nova (formerly ACDL) has no package.json. The execute/verify/ship
|
||||
workflows substitute `terraform validate` + `python -m py_compile` +
|
||||
JSON Schema validation for npm run typecheck, the regression gate
|
||||
(D-091, 22 capabilities) for npm test, and `bash scripts/run_ci.sh`
|
||||
for npm run build. v1.11 testing is pipeline-driven (D-102);
|
||||
v1.16 is NFR-only (no live apply by default; NOVA_LIFECYCLE_MODE=
|
||||
plan). Roster carries forward from v1.11/v1.14/v1.15 unchanged.
|
||||
frontend-engineer stays inactive (no frontend; decks are markdown =
|
||||
lead-developer territory). No custom personas needed (no new
|
||||
domains — onboarding is backend-engineer + data-engineer territory).
|
||||
---
|
||||
|
||||
# ACDL — Persona Roster (project-level, v1.11 RESTART)
|
||||
@@ -252,134 +251,3 @@ The regression gate (22 capabilities) must stay **22/22 Verified**
|
||||
throughout v1.16 — simplification must not regress any capability
|
||||
(D-118). P9 (end of Wave 2) and P21 (milestone complete) run the gate;
|
||||
P14 (end of Wave 3) is an offline mid-milestone checkpoint.
|
||||
|
||||
---
|
||||
|
||||
# v1.17 Persona Roster — Strategic Direction, Leadership Metrics & Unified Story
|
||||
|
||||
> v1.17 adds a telemetry/observability layer (P1–P3), a metrics catalog
|
||||
> + NORTH_STAR integration (P4), a unified narrative deck (P5), a
|
||||
> regression capability (P6), and a final review/ship (P7). Three
|
||||
> active personas; frontend-engineer stays deactivated (no Nova web UI
|
||||
> — dashboards are PowerBI, not a Nova-built frontend).
|
||||
|
||||
## Active personas
|
||||
|
||||
### lead-developer
|
||||
- **Domain:** coordination + deck narrative
|
||||
- **Active:** true
|
||||
- **Phase-specific:** false
|
||||
- **Reason:** Owns CIAgent metadata, the NORTH_STAR.md authoring
|
||||
process (P0), the milestone decomposition, the unified narrative deck
|
||||
co-authoring (P5 — the deck is markdown, which is lead-developer
|
||||
territory per the established convention), and the final review/ship
|
||||
(P7). Arbitrates persona conflicts (e.g., backend vs data on the
|
||||
emitter/store boundary).
|
||||
- **Territory:** `.ciagent/NORTH_STAR.md`, `.ciagent/PROJECT.md`,
|
||||
`.ciagent/REQUIREMENTS.md`, `.ciagent/PLAN.md`, `.ciagent/RESEARCH.md`,
|
||||
`.ciagent/ARCHITECTURE.md`, `docs/presentations/nova-no-humans-platform.md`
|
||||
(NEW — unified deck source of truth), `docs/presentations/nova-no-humans-platform-marp.md`,
|
||||
`docs/presentations/nova-no-humans-platform-talking-points.md`,
|
||||
`docs/METRICS.md`, `docs/metrics/*.md` (per-KPI definition docs).
|
||||
|
||||
### backend-engineer
|
||||
- **Domain:** backend (event emitters + instrumentation)
|
||||
- **Active:** true
|
||||
- **Phase-specific:** false
|
||||
- **Reason:** Owns the event emitters (P1): the CloudEvents envelope,
|
||||
the per-run manifest writer, the `outbox_writer.py` extension to the
|
||||
SQLite Decision Ledger, the Infracost post-processor, the
|
||||
`hitl_gates.py` attestation event emission, the `confidence_signal.py`
|
||||
decision event emission, the `checkov_adapter.py` policy event
|
||||
emission, and the pytest `--junitxml` addopts change. Also owns the
|
||||
`regression_verify.py` CAP-023/024 additions (P6). The emitter work
|
||||
is the bridge between existing Nova components and the new metrics
|
||||
layer — it touches the code paths that already exist.
|
||||
- **Territory:** `core/metrics/event_envelope.py` (NEW),
|
||||
`core/metrics/run_manifest.py` (NEW),
|
||||
`core/metrics/infracost_adapter.py` (NEW),
|
||||
`core/metrics/decision_ledger.py` (NEW — extends outbox_writer),
|
||||
`core/outbox_writer.py` (extend to SQLite),
|
||||
`core/hitl_gates.py` (emit attestation.recorded),
|
||||
`core/confidence_signal.py` (emit ai.decision.made),
|
||||
`adapters/terraform/policy/checkov_adapter.py` (emit policy.evaluated),
|
||||
`scripts/run_platform.sh` (invoke manifest writer + Infracost),
|
||||
`core/regression_verify.py` (CAP-023/024),
|
||||
`pyproject.toml` (addopts --junitxml),
|
||||
`tests/test_metrics_emitters.py` (NEW),
|
||||
`tests/test_decision_ledger.py` (NEW).
|
||||
|
||||
### data-engineer
|
||||
- **Domain:** data (schema, SQLite store, PowerBI export)
|
||||
- **Active:** true
|
||||
- **Phase-specific:** false
|
||||
- **Reason:** Reactivated with a new territory for v1.17: the metrics
|
||||
collector (P2) and the PowerBI export (P3). Owns the schema design
|
||||
(metrics_*.schema.json), the SQLite cold store (nova_metrics.db), the
|
||||
fact/dimension table design, the 8 deferred placeholder views, and
|
||||
the CSV/JSON export. The data-engineer's schema-first constraint
|
||||
applies: all event types and fact/dim tables have JSON Schema
|
||||
definitions before any code is written. The collector reads files +
|
||||
events → SQLite; the export reads SQLite → CSV/JSON. This is the
|
||||
heaviest data-territory work since v1.11's terraform modules.
|
||||
- **Territory:** `core/metrics/collector.py` (NEW),
|
||||
`core/metrics/powerbi_export.py` (NEW),
|
||||
`schemas/metrics_*.schema.json` (NEW — event + fact/dim schemas),
|
||||
`metrics/nova_metrics.db` (NEW — SQLite cold store),
|
||||
`metrics/powerbi/` (NEW — CSV/JSON export dir),
|
||||
`docs/METRICS_VIEWS.md` (NEW — schema doc for PowerBI views),
|
||||
`tests/test_metrics_collector.py` (NEW),
|
||||
`tests/test_powerbi_export.py` (NEW).
|
||||
|
||||
## Deactivated personas
|
||||
|
||||
### frontend-engineer
|
||||
- **Domain:** frontend
|
||||
- **Active:** false
|
||||
- **Phase-specific:** false
|
||||
- **Reason:** v1.17 has no Nova web UI. The leadership dashboards are
|
||||
PowerBI (an external tool that ingests CSV/JSON files), not a
|
||||
Nova-built frontend. The decks are markdown (lead-developer
|
||||
territory). frontend-engineer stays deactivated, consistent with
|
||||
v1.11–v1.16. Reactivates if a future milestone builds a Nova web UI.
|
||||
|
||||
### lambda-engineer, platform-engineer, security-engineer
|
||||
- **Active:** false (carried forward from v1.11)
|
||||
- **Reason:** v1.17 does not touch the Lambda (beyond emitting events
|
||||
from the existing hitl_gates/attestation_matrix), does not do IR-
|
||||
shaped module authoring, and does not touch security adapters beyond
|
||||
emitting policy.evaluated events. The existing components are
|
||||
instrumented, not rewritten.
|
||||
|
||||
## v1.17 phase assignment
|
||||
|
||||
| Phase | Primary persona | Supporting | Territory |
|
||||
|-------|----------------|------------|-----------|
|
||||
| P0 pre-execution | lead-developer | — | `.ciagent/NORTH_STAR.md`, `PROJECT.md`, `REQUIREMENTS.md`, `RESEARCH.md`, `ARCHITECTURE.md`, `PERSONAS.md`, `PLAN.md` |
|
||||
| P1 event-emitters | backend-engineer | data-engineer (schemas) | `core/metrics/event_envelope.py`, `run_manifest.py`, `decision_ledger.py`, `infracost_adapter.py`, `outbox_writer.py`, `hitl_gates.py`, `confidence_signal.py`, `checkov_adapter.py`, `run_platform.sh`, `pyproject.toml` |
|
||||
| P2 metrics-collector | data-engineer | backend-engineer (event formats) | `core/metrics/collector.py`, `schemas/metrics_*.schema.json`, `metrics/nova_metrics.db` |
|
||||
| P3 powerbi-export | data-engineer | — | `core/metrics/powerbi_export.py`, `metrics/powerbi/`, `docs/METRICS_VIEWS.md` |
|
||||
| P4 metrics-catalog + north-star-integration | lead-developer | data-engineer (metric definitions) | `docs/METRICS.md`, `docs/metrics/*.md`, `PROJECT.md`, `ARCHITECTURE.md`, `config.json` |
|
||||
| P5 deck-rebuild | lead-developer | — | `docs/presentations/nova-no-humans-platform*.md`, retire old decks |
|
||||
| P6 regression-capability | backend-engineer | data-engineer (CAP-023 schema) | `core/regression_verify.py` (CAP-023, CAP-024) |
|
||||
| P7 final-review-ship | lead-developer | all active (review) | `.ciagent/**`, review + audit + ship |
|
||||
|
||||
## v1.17 domain priority
|
||||
|
||||
`backend → data → lead` (the emitter work in P1 is the foundation;
|
||||
data-engineer's collector + export in P2–P3 depends on P1's event
|
||||
formats; lead-developer's catalog + deck in P4–P5 depends on the
|
||||
metrics being grounded).
|
||||
|
||||
## v1.17 verification toolchain
|
||||
|
||||
```
|
||||
typecheck: terraform validate && python3 -m py_compile core/**/*.py adapters/**/*.py
|
||||
test: bash scripts/run_regression.sh # 22-capability gate + CAP-023/024 (v1.17)
|
||||
build: bash scripts/run_ci.sh # full local CI reproduction
|
||||
```
|
||||
|
||||
The regression gate (22 capabilities + CAP-023 metrics collector +
|
||||
CAP-024 deck structure) must pass at P6 and P7. CAP-009 (offline pytest
|
||||
suite) must remain Verified after the `--junitxml` addopts change
|
||||
(assumption A5).
|
||||
|
||||
+407
-1169
File diff suppressed because it is too large
Load Diff
+2
-63
@@ -989,7 +989,7 @@ conversation before execution; D-108..D-112 resolved at CLARIFY.
|
||||
| D-111 | Lambda env-var defaults (`CONTRACTS_TABLE` default `"acdl-contracts"`, etc.) → `nova-contracts`. | `core/lambda/contract_ingestor.py` has hardcoded `acdl-*` default table names. These become `nova-*` in P4 (resource migration). P2 changes the env-var name (`ACDL_*`→`NOVA_*`); P4 changes the default values to `nova-*`. | P4 updates Lambda defaults. |
|
||||
| D-112 | `nova` slug: no `project:` prefix on branches (single-project mode). | `config.json` has `projects[]` with one entry (slug `acdl`) but `git.branching_strategy` is `flat` and the established convention since v1.0 is flat branches (no `<slug>/` prefix). Nova rebrand does NOT change the branch prefix convention. Commit `---ci---` blocks use `project: acdl` (the config slug, unchanged). | Branches stay `milestone/v1.15-nova`, `phase/NN-*`; no `acdl/` or `nova/` prefix. |
|
||||
|
||||
## Objective for Milestone v1.16 (complete — NFR Simplification, tag `v1.15.26`)
|
||||
## Objective for Milestone v1.16 (active — NFR Simplification)
|
||||
|
||||
A 20-phase NFR sweep (no new features) themed around five axes the user
|
||||
directed during ideation: **Simplify without regressions**, **Security**,
|
||||
@@ -1072,65 +1072,4 @@ conversation; D-117..D-119 resolved at CLARIFY.
|
||||
| D-116 | Drift fixes = P1 of v1.16 (not a hotfix to main). | User chose "P1 of v1.16." The state-bucket drift (`adapter.py:117`) and Kyverno label contradiction are correctness regressions but latent in plan-only mode (no live apply in the default path), so they are not an active outage. Fixing them as P1 keeps the milestone self-contained. | P1 fixes both; no hotfix to main. |
|
||||
| D-117 | v1.14 NFR categories are NOT re-proposed. | v1.14 already swept over-broad excepts (REQ-141), hardcoded account-ID (REQ-142), IAM `Resource:"*"` scoping (REQ-143), contractId/env validation (REQ-144), `.gitignore` catch-all (REQ-146), `--kube-version` removal (REQ-147), orphan cleanup (REQ-148), `set -euo pipefail` parity (REQ-150). v1.16 finds NEW residual signals (the v1.15 rebrand left a fresh debt layer) and does not duplicate completed work. | Wave 1–5 target only fresh debt. |
|
||||
| D-118 | Regression gate (D-091) gates Wave 2 completion and P21. | "Simplify without regressions" is only credible if the regression gate runs after the simplification wave. The gate runs after P9 (Wave 2 done) and at P21 (milestone complete); any non-Verified capability halts W3. Mid-milestone checkpoint after P14 (offline). | P9 + P21 run the gate; P14 checkpoint. |
|
||||
| D-119 | `onboard_consumer` action stores a CMDB row pending grant (not auto-provisions). | The request-path-only scope (D-113) means the Lambda accepts an onboarding request and writes a `pending` row to `nova-contracts` (or a new `nova-onboarding` partition key); the platform automation that grants the ABAC role is the P20 Terraform (offline-proven). No AWS resources are created by the Lambda action itself. | P18 writes the pending row; P20 proves the grant Terraform offline. |
|
||||
|
||||
## Objective for Milestone v1.17 (active — Strategic Direction, Leadership Metrics & Unified Story)
|
||||
|
||||
**Milestone type:** Feature (P1–P3 feat; P4 docs; P5 docs+test; P6 test;
|
||||
P7 review+audit+ship). Tags on the v1.16.x line: `v1.16.0` (P0) →
|
||||
`v1.16.1..v1.16.7` (P1–P7) → `v1.16.8` (P8 final = milestone release).
|
||||
|
||||
**Three pillars:**
|
||||
|
||||
- **Pillar A — Strategic Direction.** A durable, PO-authored
|
||||
`.ciagent/NORTH_STAR.md` encodes the platform's vision, 4 strategic
|
||||
objectives, 5 anti-goals, v1.17 non-goals, 12–18mo targets (with a
|
||||
grounding column), and success criteria. CIAgent reads it in every
|
||||
future `/ci-run` so the direction survives across milestones. The
|
||||
attestation clarification is reflected: human attestation required at
|
||||
stage gates (QA for production, SRE for operational readiness); autonomy
|
||||
in operations, not in accountability.
|
||||
|
||||
- **Pillar B — Leadership Metrics + PowerBI.** Instrument Nova to
|
||||
collect, aggregate, and surface leadership-grade metrics that prove the
|
||||
"no-humans" autonomous-infrastructure value proposition. Nova-native
|
||||
minimal tech (CloudEvents 1.0 envelope, JSONL event log, SQLite cold
|
||||
store, hash-chained Decision Ledger via `outbox_writer.py` extension)
|
||||
+ Infracost for pre-apply cost estimates. Hybrid model: existing
|
||||
file-based signals (REGRESSION_REPORT.json, pcr.json, signal.json,
|
||||
junit XML) are sources the collector reads and projects into events;
|
||||
new emitters emit CloudEvents directly. PowerBI export = CSV/JSON
|
||||
views (fact + dimension tables + 8 empty placeholder views for
|
||||
deferred metrics). **Hard constraint: DO NOT make anything up.** Every
|
||||
metric is `grounded` (cites source file + schema), `derived`
|
||||
(documented formula), or `deferred` (cites decision ID — D-096/D-083/
|
||||
D-113/D-114/D-119). The 8 deferred metrics: drift detection, GreenOps/
|
||||
carbon, predictive/reactive, live CUR reconciliation, multi-cloud,
|
||||
red-team MTTR, self-healing velocity, SLA/downtime.
|
||||
|
||||
- **Pillar C — Unified Narrative Deck.** Merge the two existing decks
|
||||
(`how-the-platform-works` + `the-developer-experience`) into one unified
|
||||
narrative deck "Nova — The No-Humans Infrastructure Platform" with a
|
||||
single arc: Problem → Vision/Direction (NORTH_STAR) → How it works →
|
||||
Proof (metrics) → Roadmap/Ask. The "tell them x3" structure applies at
|
||||
deck level AND per slide (each slide opens with what it covers,
|
||||
delivers, closes with an explicit "benefit of this stage" callout).
|
||||
Fluid transitions between slides. Both old decks retired.
|
||||
|
||||
**Key decisions resolved in the planning conversation (D-120+):**
|
||||
|
||||
| ID | Decision | Rationale | Outcome |
|
||||
|----|----------|-----------|---------|
|
||||
| D-120 | Tech stack = Nova-native + Infracost, drift deferred. | The PO's technical-direction document specifies Kafka/Prometheus/ClickHouse/QLDB/OTel — none exist in Nova today. Adopt the PRINCIPLES (events as source of truth, CloudEvents envelope, decision ledger, definition-of-success docs, dashboards-as-projections) but implement with Nova-native minimal tech (JSONL + SQLite + hash-chained ledger). No Kafka/Prometheus/ClickHouse/QLDB. Infracost adopted (runs offline on plan JSON). Drift detection deferred (D-096 + no scheduler). | P1–P3 use Nova-native tech; Infracost in P1; drift deferred. |
|
||||
| D-121 | Decision Ledger = extend outbox_writer.py → SQLite append-only hash chain. | The direction's #1 priority is the Decision Ledger. Nova already has a hash-chained outbox (outbox_writer.py). Extend it to a SQLite append-only table with hash chain; add ai.decision.made + attestation.recorded events. Honors D-083 (no S3 Object Lock/JWS). | P1 extends outbox_writer; ledger is SQLite hash-chain. |
|
||||
| D-122 | AI Planner framing = map Nova's real decision points. | The direction assumes an "AI Planner/Reasoner" (planner-v3.2). Nova's actual decision path is confidence_signal + HITL gate. Model ai.decision.made from confidence_signal (decision_id=run_id, chosen_action=band, confidence=score, alternatives=perInput, human_override=HITL block). LLM planner marked future/aspirational. | P1 emits honest decision events; no fabricated LLM. |
|
||||
| D-123 | Deferred metrics = all 8 (drift, GreenOps, predictive/reactive, live CUR, multi-cloud, red-team MTTR, self-healing, SLA/downtime). | These require live AWS (D-096) or new external systems. Ship as empty PowerBI placeholder views with documented schemas. | P3 ships 8 placeholder views; METRICS.md marks them deferred. |
|
||||
| D-124 | NORTH_STAR = strategy; tech direction = engineering input. | The PO's technical-direction document is engineering architecture, not strategy. NORTH_STAR.md captures strategic vision/objectives/anti-goals (PO-authored). The tech direction becomes the telemetry reference architecture section in RESEARCH.md/ARCHITECTURE.md, cited by NORTH_STAR's engineering objectives. | P0 writes NORTH_STAR; RESEARCH writes the telemetry reference. |
|
||||
| D-125 | Events vs files = hybrid. | Existing file-based signals (REGRESSION_REPORT.json, pcr.json, signal.json, junit) stay as files; the collector reads them and emits normalized CloudEvents into JSONL + SQLite. New emitters emit CloudEvents directly. | P2 collector reads files + events. |
|
||||
| D-126 | Hot/cold split = cold-only SQLite (hot path deferred). | Nova has no live ops dashboard (no live AWS, D-096). The SQLite store is cold-only (batch/historical). The hot path is documented as deferred. | P2 SQLite is cold-only. |
|
||||
| D-127 | Definition-of-success = per-KPI docs. | The direction's §11 requires a definition-of-success doc for every executive KPI. Adopt this standard; docs live in `docs/metrics/`. | P4 writes per-KPI docs. |
|
||||
| D-128 | Storage location = metrics/ at repo root. | metrics/runs/ (per-run manifests), metrics/nova_metrics.db (SQLite), metrics/events.jsonl (event log), metrics/powerbi/ (export). | P1–P3 use metrics/ at repo root. |
|
||||
| D-129 | PowerBI delivery = CSV/JSON files, folder connector. | Nova is offline-first; no live connector to a running service. PowerBI ingests via the folder connector. | P3 emits CSV/JSON to metrics/powerbi/. |
|
||||
| D-130 | Deck arc = Problem → Vision → How → Proof → Roadmap. | The unified narrative deck's 5-act structure. x3 arc at deck + slide level. Per-slide benefit callouts. Fluid transitions. Both old decks retired. | P5 builds the unified deck; old decks deleted. |
|
||||
| D-131 | MTTR scope = platform-run MTTR. | The <60s MTTR target refers to platform-run failures (apply.failed → successful retry), not infra-incident MTTR (no incident detection system). Infra-incident MTTR deferred. | P4 grounds platform-run MTTR. |
|
||||
| D-132 | Attestation instrumentation = emit attestation.recorded events. | The attestation system (hitl_gates.py + attestation_matrix.py + separation_of_duties.py) already exists. Instrument it: emit attestation.recorded events into the Decision Ledger + PowerBI. Attestation Coverage = 100% target grounded from outbox approver_* attributes. | P1 emits attestation events; P4 grounds Attestation Coverage. |
|
||||
| D-119 | `onboard_consumer` action stores a CMDB row pending grant (not auto-provisions). | The request-path-only scope (D-113) means the Lambda accepts an onboarding request and writes a `pending` row to `nova-contracts` (or a new `nova-onboarding` partition key); the platform automation that grants the ABAC role is the P20 Terraform (offline-proven). No AWS resources are created by the Lambda action itself. | P18 writes the pending row; P20 proves the grant Terraform offline. |
|
||||
@@ -1,13 +1,12 @@
|
||||
{
|
||||
"run_id": "regr-1785591207",
|
||||
"run_at_utc": "2026-08-01T13:33:27Z",
|
||||
"run_id": "regr-1785375318",
|
||||
"run_at_utc": "2026-07-30T01:35:18Z",
|
||||
"milestone": "v1.10",
|
||||
"phase": 52,
|
||||
"summary": {
|
||||
"Verified": 18,
|
||||
"Verified": 22,
|
||||
"Decayed": 0,
|
||||
"Broken": 0,
|
||||
"Skipped": 4
|
||||
"Broken": 0
|
||||
},
|
||||
"passed": true,
|
||||
"results": [
|
||||
@@ -17,7 +16,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; 2 sample contracts validate",
|
||||
"tier": "local",
|
||||
"duration_ms": 235
|
||||
"duration_ms": 230
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-002",
|
||||
@@ -25,7 +24,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; env schema validates",
|
||||
"tier": "local",
|
||||
"duration_ms": 201
|
||||
"duration_ms": 204
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-003",
|
||||
@@ -33,7 +32,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; ",
|
||||
"tier": "local",
|
||||
"duration_ms": 261
|
||||
"duration_ms": 247
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-004",
|
||||
@@ -41,7 +40,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; ",
|
||||
"tier": "local",
|
||||
"duration_ms": 259
|
||||
"duration_ms": 241
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-005",
|
||||
@@ -49,7 +48,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; ",
|
||||
"tier": "local",
|
||||
"duration_ms": 337
|
||||
"duration_ms": 326
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-006",
|
||||
@@ -57,7 +56,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; interpolation ok",
|
||||
"tier": "local",
|
||||
"duration_ms": 242
|
||||
"duration_ms": 216
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-007",
|
||||
@@ -65,7 +64,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; confidence band=pass",
|
||||
"tier": "local",
|
||||
"duration_ms": 91
|
||||
"duration_ms": 79
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-008",
|
||||
@@ -73,15 +72,15 @@
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; outbox hash chain ok",
|
||||
"tier": "local",
|
||||
"duration_ms": 456
|
||||
"duration_ms": 333
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-009",
|
||||
"name": "offline pytest suite passes",
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; [ 98%]\ntests/test_wiz_adapter_real_client.py ......... [100%]\n\n================= 586 passed, 2 deselected in 71.63s (0:01:11) =================",
|
||||
"detail": "exit 0; [ 98%]\ntests/test_wiz_adapter_real_client.py ......... [100%]\n\n====================== 555 passed, 2 deselected in 51.11s ======================",
|
||||
"tier": "local",
|
||||
"duration_ms": 72988
|
||||
"duration_ms": 52574
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-010",
|
||||
@@ -89,63 +88,63 @@
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; resource(s))\n\n=== PLATFORM CHECK OK ===\ncontract -> resolver -> stack -> adapter -> structure validated (offline, no AWS)\ncheck-only: OK\n\n=== CI PIPELINE OK ===\n3 stages passed: lint, test, check-only",
|
||||
"tier": "local",
|
||||
"duration_ms": 73275
|
||||
"duration_ms": 59608
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-011",
|
||||
"name": "headline E2E runs against the local emulating tier (microservice)",
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; al-emulator\",\n \"desired_count\": 1,\n \"running_count\": 1\n },\n \"outbox_dir\": \"/tmp/nova_local_e2e_6vnrnin1/outbox\",\n \"outbox_events\": 2,\n \"outbox_chain_verified\": true,\n \"lambda_status\": 200\n}",
|
||||
"detail": "exit 0; al-emulator\",\n \"desired_count\": 1,\n \"running_count\": 1\n },\n \"outbox_dir\": \"/tmp/acdl_local_e2e_0v1bpi48/outbox\",\n \"outbox_events\": 2,\n \"outbox_chain_verified\": true,\n \"lambda_status\": 200\n}",
|
||||
"tier": "local",
|
||||
"duration_ms": 634
|
||||
"duration_ms": 1072
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-012",
|
||||
"name": "local E2E on the static-assets stack (no ECS)",
|
||||
"status": "Verified",
|
||||
"detail": "exit 0; nova_local_e2e_uq4kkhze/tf\",\n \"backend\": \"local\",\n \"ecs\": null,\n \"outbox_dir\": \"/tmp/nova_local_e2e_uq4kkhze/outbox\",\n \"outbox_events\": 2,\n \"outbox_chain_verified\": true,\n \"lambda_status\": 200\n}",
|
||||
"detail": "exit 0; acdl_local_e2e_0cjcizgd/tf\",\n \"backend\": \"local\",\n \"ecs\": null,\n \"outbox_dir\": \"/tmp/acdl_local_e2e_0cjcizgd/outbox\",\n \"outbox_events\": 2,\n \"outbox_chain_verified\": true,\n \"lambda_status\": 200\n}",
|
||||
"tier": "local",
|
||||
"duration_ms": 584
|
||||
"duration_ms": 490
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-013",
|
||||
"name": "terraform init+validate+plan live AWS (microservice)",
|
||||
"status": "Skipped",
|
||||
"detail": "terraform init: state bucket absent (post-v1.11-teardown, D-096) [microservice]",
|
||||
"status": "Verified",
|
||||
"detail": "terraform init+validate+plan OK (live AWS, microservice)",
|
||||
"tier": "live-aws",
|
||||
"duration_ms": 737
|
||||
"duration_ms": 28176
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-014",
|
||||
"name": "terraform init+validate+plan live AWS (static-assets)",
|
||||
"status": "Skipped",
|
||||
"detail": "terraform init: state bucket absent (post-v1.11-teardown, D-096) [static-assets]",
|
||||
"status": "Verified",
|
||||
"detail": "terraform init+validate+plan OK (live AWS, static-assets)",
|
||||
"tier": "live-aws",
|
||||
"duration_ms": 676
|
||||
"duration_ms": 31892
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-015",
|
||||
"name": "DynamoDB outbox table exists (live AWS)",
|
||||
"status": "Skipped",
|
||||
"detail": "nova-outbox absent (post-v1.11-teardown steady state, D-096)",
|
||||
"status": "Verified",
|
||||
"detail": "acdl-outbox exists, item_count=9",
|
||||
"tier": "live-aws",
|
||||
"duration_ms": 664
|
||||
"duration_ms": 507
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-016",
|
||||
"name": "S3 state bucket exists + readable (live AWS)",
|
||||
"status": "Skipped",
|
||||
"detail": "state bucket nova-tfstate-581513795199-us-east-1 absent (post-v1.11-teardown, D-096)",
|
||||
"status": "Verified",
|
||||
"detail": "state bucket exists, keys=['platform/terraform.tfstate', 'spike/alb/dev/terraform.tfstate', 'spike/assets/dev/terraform.tfstate', 'spike/cdn/dev/terraform.tfstate', 'spike/ci-vpc/terraform.tfstate']",
|
||||
"tier": "live-aws",
|
||||
"duration_ms": 245
|
||||
"duration_ms": 329
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-017",
|
||||
"name": "DynamoDB nova-contracts table (lifecycle pipeline evidence)",
|
||||
"name": "DynamoDB acdl-contracts table (lifecycle pipeline evidence)",
|
||||
"status": "Verified",
|
||||
"detail": "terraform files present + fmt -check passes + simple/complex contracts resolve",
|
||||
"tier": "lifecycle-pipeline",
|
||||
"duration_ms": 586
|
||||
"duration_ms": 588
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-018",
|
||||
@@ -153,7 +152,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "LocalLambdaStub instantiates (local tier evidence)",
|
||||
"tier": "lifecycle-pipeline",
|
||||
"duration_ms": 138
|
||||
"duration_ms": 135
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-019",
|
||||
@@ -161,7 +160,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "L2 composition resolves (simple + complex contracts; offline proxy)",
|
||||
"tier": "lifecycle-pipeline",
|
||||
"duration_ms": 519
|
||||
"duration_ms": 498
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-020",
|
||||
@@ -169,7 +168,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "L2 composition resolves (simple + complex contracts; offline proxy)",
|
||||
"tier": "lifecycle-pipeline",
|
||||
"duration_ms": 521
|
||||
"duration_ms": 510
|
||||
},
|
||||
{
|
||||
"capability_id": "CAP-021",
|
||||
@@ -185,7 +184,7 @@
|
||||
"status": "Verified",
|
||||
"detail": "terraform files present + fmt -check passes + simple/complex contracts resolve",
|
||||
"tier": "lifecycle-pipeline",
|
||||
"duration_ms": 611
|
||||
"duration_ms": 554
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -1,51 +1,51 @@
|
||||
# Regression Report — v1.10 Phase 52
|
||||
|
||||
- **Run ID:** `regr-1785591207`
|
||||
- **Run at (UTC):** 2026-08-01T13:33:27Z
|
||||
- **Summary:** {'Verified': 18, 'Decayed': 0, 'Broken': 0, 'Skipped': 4}
|
||||
- **Run ID:** `regr-1785375318`
|
||||
- **Run at (UTC):** 2026-07-30T01:35:18Z
|
||||
- **Summary:** {'Verified': 22, 'Decayed': 0, 'Broken': 0}
|
||||
- **Passed (milestone gate):** True
|
||||
|
||||
| Capability | Name | Tier | Status | Duration (ms) | Detail |
|
||||
|-----------|------|------|--------|--------------|--------|
|
||||
| CAP-001 | contract.schema.json validates sample contracts | local | **Verified** | 235 | exit 0; 2 sample contracts validate |
|
||||
| CAP-002 | environment.schema.json validates env files | local | **Verified** | 201 | exit 0; env schema validates |
|
||||
| CAP-003 | contract_resolver resolves static-assets | local | **Verified** | 261 | exit 0; |
|
||||
| CAP-004 | contract_resolver resolves microservice | local | **Verified** | 259 | exit 0; |
|
||||
| CAP-005 | terraform adapter emits .tf files | local | **Verified** | 337 | exit 0; |
|
||||
| CAP-006 | contract interpolation expands env/contract tokens | local | **Verified** | 242 | exit 0; interpolation ok |
|
||||
| CAP-007 | confidence_signal.compute returns a band | local | **Verified** | 91 | exit 0; confidence band=pass |
|
||||
| CAP-008 | outbox_writer builds a hash-chained item | local | **Verified** | 456 | exit 0; outbox hash chain ok |
|
||||
| CAP-009 | offline pytest suite passes | local | **Verified** | 72988 | exit 0; [ 98%]
|
||||
| CAP-001 | contract.schema.json validates sample contracts | local | **Verified** | 230 | exit 0; 2 sample contracts validate |
|
||||
| CAP-002 | environment.schema.json validates env files | local | **Verified** | 204 | exit 0; env schema validates |
|
||||
| CAP-003 | contract_resolver resolves static-assets | local | **Verified** | 247 | exit 0; |
|
||||
| CAP-004 | contract_resolver resolves microservice | local | **Verified** | 241 | exit 0; |
|
||||
| CAP-005 | terraform adapter emits .tf files | local | **Verified** | 326 | exit 0; |
|
||||
| CAP-006 | contract interpolation expands env/contract tokens | local | **Verified** | 216 | exit 0; interpolation ok |
|
||||
| CAP-007 | confidence_signal.compute returns a band | local | **Verified** | 79 | exit 0; confidence band=pass |
|
||||
| CAP-008 | outbox_writer builds a hash-chained item | local | **Verified** | 333 | exit 0; outbox hash chain ok |
|
||||
| CAP-009 | offline pytest suite passes | local | **Verified** | 52574 | exit 0; [ 98%]
|
||||
tests/test_wiz_adapter_real_client.py ......... [100%]
|
||||
|
||||
================= 586 passed, 2 |
|
||||
| CAP-010 | run_ci.sh reproduces CI pipeline locally | local | **Verified** | 73275 | exit 0; resource(s))
|
||||
====================== 555 passe |
|
||||
| CAP-010 | run_ci.sh reproduces CI pipeline locally | local | **Verified** | 59608 | exit 0; resource(s))
|
||||
|
||||
=== PLATFORM CHECK OK ===
|
||||
contract -> resolver -> stack -> adapter -> structure validated (offline, no AWS)
|
||||
check-only: OK
|
||||
|
||||
=== CI PIPELIN |
|
||||
| CAP-011 | headline E2E runs against the local emulating tier (microservice) | local | **Verified** | 634 | exit 0; al-emulator",
|
||||
| CAP-011 | headline E2E runs against the local emulating tier (microservice) | local | **Verified** | 1072 | exit 0; al-emulator",
|
||||
"desired_count": 1,
|
||||
"running_count": 1
|
||||
},
|
||||
"outbox_dir": "/tmp/nova_local_e2e_6vnrnin1/outbox",
|
||||
"outbox_dir": "/tmp/acdl_local_e2e_0v1bpi48/outbox",
|
||||
"outbox_events": 2,
|
||||
"outbox |
|
||||
| CAP-012 | local E2E on the static-assets stack (no ECS) | local | **Verified** | 584 | exit 0; nova_local_e2e_uq4kkhze/tf",
|
||||
| CAP-012 | local E2E on the static-assets stack (no ECS) | local | **Verified** | 490 | exit 0; acdl_local_e2e_0cjcizgd/tf",
|
||||
"backend": "local",
|
||||
"ecs": null,
|
||||
"outbox_dir": "/tmp/nova_local_e2e_uq4kkhze/outbox",
|
||||
"outbox_dir": "/tmp/acdl_local_e2e_0cjcizgd/outbox",
|
||||
"outbox_events": 2,
|
||||
"outbox |
|
||||
| CAP-013 | terraform init+validate+plan live AWS (microservice) | live-aws | **Skipped** | 737 | terraform init: state bucket absent (post-v1.11-teardown, D-096) [microservice] |
|
||||
| CAP-014 | terraform init+validate+plan live AWS (static-assets) | live-aws | **Skipped** | 676 | terraform init: state bucket absent (post-v1.11-teardown, D-096) [static-assets] |
|
||||
| CAP-015 | DynamoDB outbox table exists (live AWS) | live-aws | **Skipped** | 664 | nova-outbox absent (post-v1.11-teardown steady state, D-096) |
|
||||
| CAP-016 | S3 state bucket exists + readable (live AWS) | live-aws | **Skipped** | 245 | state bucket nova-tfstate-581513795199-us-east-1 absent (post-v1.11-teardown, D-096) |
|
||||
| CAP-017 | DynamoDB nova-contracts table (lifecycle pipeline evidence) | lifecycle-pipeline | **Verified** | 586 | terraform files present + fmt -check passes + simple/complex contracts resolve |
|
||||
| CAP-018 | Lambda contract-ingestor (local stub + lifecycle evidence) | lifecycle-pipeline | **Verified** | 138 | LocalLambdaStub instantiates (local tier evidence) |
|
||||
| CAP-019 | ECS cluster + service (L2 microservice lifecycle evidence) | lifecycle-pipeline | **Verified** | 519 | L2 composition resolves (simple + complex contracts; offline proxy) |
|
||||
| CAP-020 | CloudFront + WAF (L2 static-assets lifecycle evidence) | lifecycle-pipeline | **Verified** | 521 | L2 composition resolves (simple + complex contracts; offline proxy) |
|
||||
| CAP-013 | terraform init+validate+plan live AWS (microservice) | live-aws | **Verified** | 28176 | terraform init+validate+plan OK (live AWS, microservice) |
|
||||
| CAP-014 | terraform init+validate+plan live AWS (static-assets) | live-aws | **Verified** | 31892 | terraform init+validate+plan OK (live AWS, static-assets) |
|
||||
| CAP-015 | DynamoDB outbox table exists (live AWS) | live-aws | **Verified** | 507 | acdl-outbox exists, item_count=9 |
|
||||
| CAP-016 | S3 state bucket exists + readable (live AWS) | live-aws | **Verified** | 329 | state bucket exists, keys=['platform/terraform.tfstate', 'spike/alb/dev/terraform.tfstate', 'spike/assets/dev/terraform.tfstate', 'spike/cdn/dev/terraform.tfsta |
|
||||
| CAP-017 | DynamoDB acdl-contracts table (lifecycle pipeline evidence) | lifecycle-pipeline | **Verified** | 588 | terraform files present + fmt -check passes + simple/complex contracts resolve |
|
||||
| CAP-018 | Lambda contract-ingestor (local stub + lifecycle evidence) | lifecycle-pipeline | **Verified** | 135 | LocalLambdaStub instantiates (local tier evidence) |
|
||||
| CAP-019 | ECS cluster + service (L2 microservice lifecycle evidence) | lifecycle-pipeline | **Verified** | 498 | L2 composition resolves (simple + complex contracts; offline proxy) |
|
||||
| CAP-020 | CloudFront + WAF (L2 static-assets lifecycle evidence) | lifecycle-pipeline | **Verified** | 510 | L2 composition resolves (simple + complex contracts; offline proxy) |
|
||||
| CAP-021 | uptime-kuma (L1 uptime lifecycle evidence) | lifecycle-pipeline | **Verified** | 562 | terraform files present + fmt -check passes + simple/complex contracts resolve |
|
||||
| CAP-022 | OIDC role (L1 iam-role lifecycle evidence) | lifecycle-pipeline | **Verified** | 611 | terraform files present + fmt -check passes + simple/complex contracts resolve |
|
||||
| CAP-022 | OIDC role (L1 iam-role lifecycle evidence) | lifecycle-pipeline | **Verified** | 554 | terraform files present + fmt -check passes + simple/complex contracts resolve |
|
||||
|
||||
+20
-244
@@ -921,26 +921,26 @@ simplification and the first self-service onboarding request path.
|
||||
|
||||
| Requirement | Phase | Status |
|
||||
|-------------|-------|--------|
|
||||
| REQ-165 | P1 | complete |
|
||||
| REQ-166 | P2 | complete |
|
||||
| REQ-167 | P3 | complete |
|
||||
| REQ-168 | P4 | complete |
|
||||
| REQ-169 | P5 | complete |
|
||||
| REQ-170 | P6 | complete |
|
||||
| REQ-171 | P7 | complete |
|
||||
| REQ-172 | P8 | complete |
|
||||
| REQ-173 | P9 | complete |
|
||||
| REQ-174 | P10 | complete |
|
||||
| REQ-175 | P11 | complete |
|
||||
| REQ-176 | P12 | complete |
|
||||
| REQ-177 | P13 | complete |
|
||||
| REQ-178 | P14 | complete |
|
||||
| REQ-179 | P15 | complete |
|
||||
| REQ-180 | P16 | complete |
|
||||
| REQ-181 | P17 | complete |
|
||||
| REQ-182 | P18 | complete |
|
||||
| REQ-183 | P19 | complete |
|
||||
| REQ-184 | P20 | complete |
|
||||
| REQ-165 | P1 | pending |
|
||||
| REQ-166 | P2 | pending |
|
||||
| REQ-167 | P3 | pending |
|
||||
| REQ-168 | P4 | pending |
|
||||
| REQ-169 | P5 | pending |
|
||||
| REQ-170 | P6 | pending |
|
||||
| REQ-171 | P7 | pending |
|
||||
| REQ-172 | P8 | pending |
|
||||
| REQ-173 | P9 | pending |
|
||||
| REQ-174 | P10 | pending |
|
||||
| REQ-175 | P11 | pending |
|
||||
| REQ-176 | P12 | pending |
|
||||
| REQ-177 | P13 | pending |
|
||||
| REQ-178 | P14 | pending |
|
||||
| REQ-179 | P15 | pending |
|
||||
| REQ-180 | P16 | pending |
|
||||
| REQ-181 | P17 | pending |
|
||||
| REQ-182 | P18 | pending |
|
||||
| REQ-183 | P19 | pending |
|
||||
| REQ-184 | P20 | pending |
|
||||
|
||||
### Out of Scope (v1.16)
|
||||
- New features (feat phases). v1.16 is NFR-only.
|
||||
@@ -956,227 +956,3 @@ simplification and the first self-service onboarding request path.
|
||||
scoping (REQ-143), contractId/env validation (REQ-144), `.gitignore`
|
||||
catch-all (REQ-146), `--kube-version` removal (REQ-147), orphan
|
||||
cleanup (REQ-148), `set -euo pipefail` parity (REQ-150).
|
||||
|
||||
## v1.17 — Strategic Direction, Leadership Metrics & Unified Story
|
||||
|
||||
**Milestone type:** Feature (P1–P3 feat; P4 docs; P5 docs+test; P6 test;
|
||||
P7 review+audit+ship). Progressive patches; the final phase's patch IS
|
||||
the milestone release. Tags run on the v1.16.x line: `v1.16.0` (P0) →
|
||||
`v1.16.1..v1.16.7` (P1–P7) → `v1.16.8` (P8 final = milestone release).
|
||||
|
||||
**Objective:** Three pillars. (A) Encode the PO's strategic direction in
|
||||
a durable `NORTH_STAR.md` read by CIAgent in every future `/ci-run`.
|
||||
(B) Instrument Nova to collect, aggregate, and surface leadership-grade
|
||||
metrics that prove the "no-humans" autonomous-infrastructure value
|
||||
proposition — grounded in signals Nova actually emits, derived via
|
||||
documented formulas, or explicitly deferred with a decision ID — flowing
|
||||
into PowerBI-ready views. (C) Merge the two existing decks into one
|
||||
unified narrative deck with the "tell them x3" arc at deck + slide level,
|
||||
per-slide benefit callouts, and fluid transitions.
|
||||
|
||||
**Hard constraint:** DO NOT make anything up. Every metric carries a
|
||||
`grounded` / `derived` / `deferred` status with a source file or
|
||||
decision ID. Deferred metrics ship as empty PowerBI placeholder views
|
||||
with documented schemas.
|
||||
|
||||
### Requirements
|
||||
|
||||
**Pillar A — Strategic Direction**
|
||||
|
||||
- **REQ-185** — `.ciagent/NORTH_STAR.md` is PO-authored with Vision,
|
||||
Strategic Objectives (4), Anti-Goals (5), Non-Goals (v1.17 scope),
|
||||
12–18mo Targets (with grounding column), and Success Criteria. The
|
||||
attestation clarification is reflected: human attestation required at
|
||||
stage gates (QA for production, SRE for operational readiness);
|
||||
autonomy in operations, not in accountability. (Phase P0)
|
||||
- **REQ-186** — CIAgent reads `NORTH_STAR.md` in context-loading for all
|
||||
future milestones; the file is referenced from PROJECT.md and
|
||||
ARCHITECTURE.md so the strategic direction survives across milestones.
|
||||
(Phase P4)
|
||||
|
||||
**Pillar B — Leadership Metrics + PowerBI**
|
||||
|
||||
- **REQ-187** — Event emitters: a CloudEvents 1.0 envelope is adopted;
|
||||
a per-run manifest writer emits structured events (run_id, contractId,
|
||||
env, stages×durations, exit, confidence, HITL block count) to
|
||||
`metrics/runs/`; existing ephemeral `$WORK/*.json` (pcr, signal,
|
||||
event, outbox, stack) are persisted as durable artifacts; pytest
|
||||
`addopts` gains `--junitxml`+`--json-report`; Infracost runs as a
|
||||
plan post-processor emitting `cost.estimated{delta_usd}` (offline).
|
||||
(Phase P1)
|
||||
- **REQ-188** — Decision Ledger: `outbox_writer.py` is extended to emit
|
||||
to a SQLite append-only table with hash chain; `ai.decision.made`
|
||||
events are modeled from Nova's real decision points (decision_id=run_id,
|
||||
chosen_action=band outcome, confidence=score, alternatives=perInput
|
||||
breakdown, human_override=HITL block) with outcome backfill from
|
||||
apply.completed; `attestation.recorded` events capture qa/prod/dr
|
||||
sign-offs (approver, env, concerns, result). Honors D-083 (no S3 Object
|
||||
Lock/JWS). (Phase P1)
|
||||
- **REQ-189** — Metrics collector: `core/metrics/collector.py` +
|
||||
`schemas/metrics_*.schema.json` read all grounded signals
|
||||
(REGRESSION_REPORT.json, per-run manifests, junit XML, pcr.json,
|
||||
signal.json, COST.md, decision ledger) → normalized SQLite cold store
|
||||
at `metrics/nova_metrics.db`; idempotent re-runs. (Phase P2)
|
||||
- **REQ-190** — PowerBI export: `core/metrics/powerbi_export.py` emits
|
||||
CSV/JSON views to `metrics/powerbi/` (fact_run, fact_capability,
|
||||
fact_policy_check, fact_confidence, fact_test, fact_decision,
|
||||
fact_cost_estimate, dim_capability, dim_milestone + 8 empty
|
||||
placeholder views for deferred metrics with documented schemas) +
|
||||
`docs/METRICS_VIEWS.md` schema doc. (Phase P3)
|
||||
- **REQ-191** — Zero-touch efficiency metrics: Autonomous Resolution
|
||||
Rate (runs without operational HITL block ÷ total; attestation gates
|
||||
excluded), Human Escalation Frequency (operational HITL blocks only),
|
||||
AI Decision Accuracy (decisions not followed by apply.failed/incident
|
||||
within 5min), MTTD/MTTR (platform-run: apply.failed → successful
|
||||
retry). (Attestation Coverage is owned by REQ-194, not here.)
|
||||
(Phase P4)
|
||||
- **REQ-192** — Velocity metrics: Provisioning Lead Time
|
||||
(apply.completed.time − intent.received.time), Deployment Frequency
|
||||
(count(apply.completed) per day). Self-Healing Velocity deferred (no
|
||||
auto-remediator). (Phase P4)
|
||||
- **REQ-193** — Financial & cost-ROI metrics: FTE Hours Saved (derived:
|
||||
run count × manual baseline), Cost Savings via Infracost estimates
|
||||
(grounded), Cost Efficiency Ratio (derived), Platform ROI (derived
|
||||
formula). Live CUR reconciliation deferred (D-096). (Phase P4)
|
||||
- **REQ-194** — Reliability, security & compliance metrics: Zero-Trust
|
||||
Policy Compliance Rate (from pcr.json), Attestation Coverage (prod/dr
|
||||
promotions attested by a human ÷ total prod/dr promotions; grounded in
|
||||
hitl_gates.py + outbox approver_* attributes; canonical owner of this
|
||||
metric). Uptime, Patch Remediation, SLA/downtime deferred (D-096).
|
||||
(Phase P4)
|
||||
- **REQ-195** — Metrics catalog doc: `docs/METRICS.md` catalogs every
|
||||
executive KPI with `grounded`/`derived`/`deferred` status, source
|
||||
file or decision ID, and a per-KPI definition-of-success doc in
|
||||
`docs/metrics/<kpi>.md`. (Phase P4)
|
||||
|
||||
**Pillar C — Unified Narrative Deck**
|
||||
|
||||
- **REQ-196** — The two existing decks (`how-the-platform-works` +
|
||||
`the-developer-experience`) are merged into one unified narrative deck
|
||||
"Nova — The No-Humans Infrastructure Platform" with a single arc:
|
||||
Problem → Vision/Direction (NORTH_STAR) → How it works → Proof
|
||||
(metrics) → Roadmap/Ask. The x3 structure ("tell them what you're
|
||||
going to tell them → tell them → tell them what you told them") applies
|
||||
at deck level (opening = arc; body = tell them; closing = recap + ask).
|
||||
Both old decks are retired (all derived artifacts deleted). (Phase P5)
|
||||
- **REQ-197** — Each slide has the x3 structure (opens with what it
|
||||
covers, delivers, closes with an explicit "benefit of this stage"
|
||||
callout) + fluid transitions between slides (no disjointed jumps).
|
||||
The 4-step deck process (source `.md` → Marp → HTML → talking-points)
|
||||
is re-run for the unified deck. (Phase P5)
|
||||
|
||||
**Cross-cutting**
|
||||
|
||||
- **REQ-198** — Regression capability: CAP-023 (metrics collector runs,
|
||||
emits expected schema) + CAP-024 (deck structure: slide count, x3
|
||||
present, per-slide benefit present) added to `core/regression_verify.py`.
|
||||
(Phase P6)
|
||||
|
||||
**Ideation enhancements (REQ-199..213 — additive, within D-120..D-132)**
|
||||
|
||||
- **REQ-199** — Metrics schema validation in CI: `run_ci.sh` validates
|
||||
`metrics/powerbi/*.json` + a sample `metrics/events.jsonl` against
|
||||
their schemas; exits 0. (Phase P3)
|
||||
- **REQ-200** — Idempotent collector re-run test: `test_metrics_collector_idempotent`
|
||||
passes (two runs → identical row counts + chain verified). (Phase P2)
|
||||
- **REQ-201** — Metrics store backup/restore doc: `metrics/README.md`
|
||||
documents regenerable vs append-only artifacts + restore procedure.
|
||||
(Phase P2)
|
||||
- **REQ-202** — Metrics glossary appendix slide: the unified deck has a
|
||||
"Metrics Glossary" appendix slide with one-line KPI definitions +
|
||||
grounding badges. (Phase P5)
|
||||
- **REQ-203** — "What's Deferred — and Why" slide: the unified deck has
|
||||
a slide pairing each of 8 deferred metrics with its blocking decision
|
||||
ID. (Phase P5)
|
||||
- **REQ-204** — NORTH_STAR diff-check in CI: `run_ci.sh` includes
|
||||
`check_north_star_diff` that fails when Vision/Objectives/Anti-Goals/
|
||||
Targets sections change without a `NORTH_STAR-CHANGE:` commit trailer.
|
||||
(Phase P4)
|
||||
- **REQ-205** — Per-module lifecycle success-rate report: each lifecycle
|
||||
run writes `metrics/lifecycle/<module>-<env>.json`; collector projects
|
||||
into `fact_lifecycle`; PowerBI "Module Lifecycle Health" view. (Phase
|
||||
P1 emitter + P2 collector + P3 view)
|
||||
- **REQ-206** — Code coverage trend emission: `pyproject.toml` addopts
|
||||
gains `--cov=core --cov=adapters --cov-report=json:metrics/coverage.json`;
|
||||
collector ingests; `fact_test` carries a coverage column. (Phase P1 +
|
||||
P2)
|
||||
- **REQ-207** — Decision Ledger CLI: `core/metrics/decision_ledger_cli.py`
|
||||
supports `query`, `verify-chain`, `stats`, `export`, `replay`;
|
||||
`verify-chain` detects broken hashes; `replay` prints ordered events;
|
||||
tests pass offline. (Phase P2)
|
||||
- **REQ-208** — PowerBI starter dashboard README: `metrics/powerbi/NOVA_DASHBOARD_README.md`
|
||||
documents folder-connector import + starter visual model + reference
|
||||
screenshot. (Phase P3)
|
||||
- **REQ-209** — PowerBI column-level data dictionary: `docs/METRICS_VIEWS.md`
|
||||
has a per-column data-dictionary table (column, type, source/formula,
|
||||
unit, grounded/derived/deferred status). (Phase P3/P4)
|
||||
- **REQ-210** — Deferred-metrics activation roadmap: `docs/METRICS_DEFERRED_ROADMAP.md`
|
||||
lists 8 deferred metrics + onboarding-grant half with {blocking
|
||||
decision, unblock requirement, candidate milestone} + a "Hot-Path
|
||||
Activation (post-D-096)" section (Nova-native only, D-120) +
|
||||
"Re-evaluation Triggers" section. (Phase P4)
|
||||
- **REQ-211** — Trust-snapshot report: `core/metrics/trust_snapshot.py`
|
||||
emits `metrics/TRUST_SNAPSHOT.md` with 5 trust metrics (Decision Ledger
|
||||
Coverage, Attestation Coverage, Capability Health, AI Decision
|
||||
Accuracy, Confidence-Gate Halt Rate) + chain-integrity verdict +
|
||||
snapshot hash; runs offline. (Phase P4)
|
||||
- **REQ-212** — Confidence-Gate Halt Rate metric: `docs/METRICS.md` +
|
||||
trust snapshot include "Confidence-Gate Halt Rate" (signal.json
|
||||
band=halt ÷ total runs); PowerBI view includes it. (Phase P4)
|
||||
- **REQ-213** — "No-humans" thesis defensibility brief: `docs/NO_HUMANS_THESIS.md`
|
||||
defines the thesis, grounded proof metrics, deferred proof metrics,
|
||||
and explicit anti-claims (incl. D-122 honesty); the unified deck's
|
||||
Vision act cites it. (Phase P4/P5)
|
||||
|
||||
### v1.17 Traceability
|
||||
|
||||
| Requirement | Phase | Status |
|
||||
|-------------|-------|--------|
|
||||
| REQ-185 | P0 | in_progress |
|
||||
| REQ-186 | P4 | pending |
|
||||
| REQ-187 | P1 | pending |
|
||||
| REQ-188 | P1 | pending |
|
||||
| REQ-189 | P2 | pending |
|
||||
| REQ-190 | P3 | pending |
|
||||
| REQ-191 | P4 | pending |
|
||||
| REQ-192 | P4 | pending |
|
||||
| REQ-193 | P4 | pending |
|
||||
| REQ-194 | P4 | pending |
|
||||
| REQ-195 | P4 | pending |
|
||||
| REQ-196 | P5 | pending |
|
||||
| REQ-197 | P5 | pending |
|
||||
| REQ-198 | P6 | pending |
|
||||
| REQ-199 | P3 | pending |
|
||||
| REQ-200 | P2 | pending |
|
||||
| REQ-201 | P2 | pending |
|
||||
| REQ-202 | P5 | pending |
|
||||
| REQ-203 | P5 | pending |
|
||||
| REQ-204 | P4 | pending |
|
||||
| REQ-205 | P1+P2+P3 | pending |
|
||||
| REQ-206 | P1+P2 | pending |
|
||||
| REQ-207 | P2 | pending |
|
||||
| REQ-208 | P3 | pending |
|
||||
| REQ-209 | P3/P4 | pending |
|
||||
| REQ-210 | P4 | pending |
|
||||
| REQ-211 | P4 | pending |
|
||||
| REQ-212 | P4 | pending |
|
||||
| REQ-213 | P4/P5 | pending |
|
||||
|
||||
### Out of Scope (v1.17)
|
||||
|
||||
- Live AWS re-provisioning (D-096) — metrics requiring live
|
||||
infrastructure ship as placeholder views.
|
||||
- Onboarding auto-grant (D-113/D-114/D-119) — only the request-path
|
||||
metric is grounded.
|
||||
- ML anomaly-forecasting / predictive remediation — no emitter today;
|
||||
Predictive-vs-Reactive metric ships as a placeholder.
|
||||
- Drift detection scheduled job (D-096 + no scheduler) — drift metrics
|
||||
ship as placeholders.
|
||||
- Live cost CUR reconciliation (D-096) — Infracost pre-apply estimates
|
||||
are grounded; actuals are not.
|
||||
- S3 Object Lock / JWS tamper-evident ledger (D-083) — Decision Ledger
|
||||
uses a local SQLite hash-chain this milestone.
|
||||
- Multi-cloud support (Azure/GCP/K8s) — Nova is AWS-only this milestone.
|
||||
- A third deck — the two existing decks merge into one; no new
|
||||
standalone metrics deck.
|
||||
- A Nova web UI — dashboards are PowerBI, not a Nova-built frontend.
|
||||
|
||||
@@ -1188,308 +1188,3 @@ stays a future feature (D-113).
|
||||
- A4 (0.85): The regression gate (D-091, D-118) at P9 and P21 confirms
|
||||
"simplify without regressions" — 22/22 capabilities must stay Verified.
|
||||
The gate is the credible control for the simplification wave.
|
||||
|
||||
---
|
||||
|
||||
# v1.17 Research — Strategic Direction, Leadership Metrics & Unified Story
|
||||
|
||||
> Phase: research (P0). Milestone: v1.17. Status: research.
|
||||
> Researcher: ci-researcher + explore agent (signal inventory).
|
||||
> Autonomy: full. Decisions D-120..D-132 locked in the planning
|
||||
> conversation (PROJECT.md). NORTH_STAR.md drafted (pending GRILL).
|
||||
|
||||
## 1. Telemetry Signal Inventory (grounding audit)
|
||||
|
||||
**Methodology:** every claim below is grounded in a concrete file path +
|
||||
line number in `/root/acdl`. No speculation. The explore agent performed
|
||||
a full sweep of the repo. The finding: **Nova has no metrics/telemetry/
|
||||
dashboard aggregation layer today.** What exists is a set of discrete,
|
||||
structured, file-based signal artifacts (JSON reports, JSONL logs,
|
||||
hash-chained outbox events, PR comments, Checkov JSON) plus unstructured
|
||||
stdout logs. A metrics milestone must aggregate these existing signals
|
||||
— it must not invent new ones without first adding emitters.
|
||||
|
||||
### (a) Signals that EXIST TODAY and are STRUCTURED (groundable)
|
||||
|
||||
| Signal | File / Emitter | Schema | Persistent? |
|
||||
|--------|---------------|--------|-------------|
|
||||
| Regression report (22 caps, status, duration_ms, gate) | `.ciagent/REGRESSION_REPORT.json` ← `core/regression_verify.py:643-667` | `regression_verify.py:82-91` | **Yes** (committed file) |
|
||||
| Regression report (markdown mirror) | `.ciagent/REGRESSION_REPORT.md` | same | Yes |
|
||||
| Checkpoint (milestone/phase/tag/regression summary) | `.ciagent/CHECKPOINT.json` (CIAgent-managed) | ad-hoc | Yes |
|
||||
| PolicyCheckResult list (per-rule pass/fail/severity/resourceRef) | `$WORK/pcr.json` ← `run_platform.sh:395` + `checkov_adapter.py:50-71` | `schemas/policy_check_result.schema.json` | **No** (ephemeral `/tmp/`) |
|
||||
| Confidence signal (score, band, perInput, reasonCodes) | `$WORK/signal.json` ← `run_platform.sh:412-426` + `confidence_signal.py:60-65` | `confidence_signal.py:60-65` | No (ephemeral) |
|
||||
| Outbox event (hash-chained, CONFIDENCE_COMPUTED) | `$WORK/event.json` + `$WORK/outbox_item.json` ← `run_platform.sh:444-459` + `outbox_writer.py:44-56` | `audit_ledger_design.md:44-45,81-97` | No (ephemeral; live DynamoDB torn down D-096) |
|
||||
| Resolved Target Stack | `$WORK/stack.json` ← `contract_resolver.py:581-603` | `schemas/stack.schema.json` | No (ephemeral) |
|
||||
| Lambda return bodies (submit/report_error/validate_cr/onboard) | `core/lambda/contract_ingestor.py:171,265,284,392,446` | ad-hoc JSON | No (Lambda not live; local stub only) |
|
||||
| DynamoDB CMDB rows (submitted/pending contracts) | `nova-contracts` table ← `contract_ingestor.py:160-170,433-445` | ad-hoc | **No** (table torn down D-096) |
|
||||
| SSM parameters (deploy outputs) | `/nova/<env>/<contractId>/<name>` ← `output_publisher.py:123-156` | ad-hoc | No (live AWS, torn down) |
|
||||
| PR stage comment (mode, runId) | GitHub PR API ← `post_stage_comment.sh:34-48` + `deploy.yml:141` | markdown table | Yes (GitHub) |
|
||||
| PR deploy-outputs comment | GitHub PR API ← `output_publisher.py:159-189` | markdown table | Yes (GitHub) |
|
||||
| GitHub issue (deploy failure alert) | GitHub API ← `contract_ingestor.py:179-290` + `deploy.yml:143-152` | issue body | Yes (GitHub) |
|
||||
| Local E2E result (stack_name, tier, outbox_events, chain_verified, lambda_status) | stdout JSON ← `core/local_emulators.py:498-508,519` | ad-hoc | No (stdout) |
|
||||
| HITL gate result | `core/hitl_gates.py:87,90` + `run_platform.sh:179-185` | stdout `HITL PASS/BLOCK` | No (stdout) |
|
||||
| Attestation matrix result | `core/attestation_matrix.py:184,187` | stdout `ATTESTATION PASS/BLOCK` | No (stdout) |
|
||||
| Cost figures | `.ciagent/COST.md` (manual Cost Explorer query) | markdown table | Yes (manual, not automated) |
|
||||
|
||||
### (b) Signals that EXIST but are UNSTRUCTURED (log-only)
|
||||
|
||||
| Signal | Source | Format |
|
||||
|--------|--------|--------|
|
||||
| CI pipeline result | `scripts/run_ci.sh:70-71` | stdout banner `=== CI PIPELINE OK ===` |
|
||||
| Platform stage banners + summaries | `scripts/run_platform.sh:222,241,258,263,315,383,411,442,463,490,496` | stdout `=== Step N: ... ===` + summary lines |
|
||||
| Terraform init/validate/plan/apply/destroy logs | `$WORK/tf-*.log` ← `run_platform.sh:320,324,328,352,375` | raw terraform stdout (via `tee`) |
|
||||
| Lifecycle test results | `scripts/run_lifecycle_test.sh` etc. | exit code only (no report file) |
|
||||
| Decommission step counts | `scripts/run_decommission.sh:40,54` | stdout `decommission step N: M resources...` |
|
||||
| Uptime endpoint count | `scripts/run_uptime.sh:72,87` | stdout `uptime: N endpoint(s) to monitor` |
|
||||
| Onboarding prompt | `core/environment_check.py:57-81` | stdout text block |
|
||||
| Pytest results | `pyproject.toml:25` (`-v --tb=short`) | stdout only (no junit/json) |
|
||||
| sync_workflows result | `scripts/sync_workflows.py:56,53` | stdout `OK: 3 workflow pairs match` / `DRIFT: ...` |
|
||||
|
||||
### (c) Proposed executive metrics with NO grounding today (DEFERRED)
|
||||
|
||||
| Proposed metric | Why no grounding | Controlling decision |
|
||||
|------------------|------------------|---------------------|
|
||||
| Live infrastructure health (ECS running count, ALB 5xx, RPS) | Live AWS torn down; CAP-013..016 Skipped | **D-096** |
|
||||
| Live outbox write rate / ledger append latency | DynamoDB outbox table absent | **D-096** |
|
||||
| Tamper-evident ledger checkpoint count / JWS signature rate | S3 Object Lock + JWS + async worker deferred | **D-083** |
|
||||
| Onboarding funnel: requested → granted conversion | Only "requested" (pending row) is emitted; no grant event | **D-113, D-114, D-119** |
|
||||
| Time-to-provision (onboarding SLA) | Real AWS provisioning deferred | **D-113** |
|
||||
| Cross-account role grant count | Offline-proven only, no live apply | **D-114** |
|
||||
| Drift detection (scheduled terraform plan -detailed-exitcode) | Needs live AWS workspaces + a scheduler Nova doesn't have | **D-096** + no scheduler |
|
||||
| GreenOps / carbon (WattTime/Electricity Maps API) | No grounding; new external API | future emitter |
|
||||
| Predictive vs Reactive ratio | Requires an ML anomaly-forecasting service | future emitter |
|
||||
| Multi-cloud normalization (Azure/GCP/K8s, FOCUS spec) | Nova is AWS-only | future |
|
||||
| Red Team MTTR | No red-team program exists | future |
|
||||
| Self-healing velocity | Nova has no auto-remediator | future emitter |
|
||||
| SLA / unplanned downtime | Needs live service uptime monitoring against SLOs | **D-096** |
|
||||
| Per-module lifecycle success rate over time | No structured report file written; only exit code | gap (no decision) |
|
||||
| Test pass rate / test count time-series | No junit/json reporter configured | gap (add `--junitxml` to addopts) |
|
||||
| Code coverage trend | `pytest-cov` installed but not in `addopts` | gap |
|
||||
| Deploy frequency / lead time / MTTR (DORA) | No deploy-event emitter; pipeline runs not counted | gap |
|
||||
| Policy pass rate time-series | `pcr.json` emitted but ephemeral; not persisted | gap (D-096 blocks live persistence) |
|
||||
| Confidence score distribution over time | `signal.json` emitted but ephemeral | gap |
|
||||
| Consumer adoption count / active consumers | `PROJECT.md:487` explicitly states "0 consumer adoption today" | honest scope |
|
||||
| Cost time-series (automated) | `COST.md` is a one-shot manual query; no automated emitter | gap |
|
||||
|
||||
**Bottom line:** the single richest existing structured signal is
|
||||
`.ciagent/REGRESSION_REPORT.json` (22 capabilities × {status, tier,
|
||||
duration_ms, detail} + summary counts + boolean gate). The next richest
|
||||
is the per-run `$WORK/*.json` family (pcr.json, signal.json, event.json,
|
||||
stack.json) — but these are **ephemeral** and **not persisted in CI**.
|
||||
The lowest-friction grounding for a "no-humans" dashboard is therefore:
|
||||
(1) regression report → capability health, (2) PR comments + GitHub
|
||||
issues → deploy/failure activity, (3) add `--junitxml` to pytest → test
|
||||
trend, (4) persist `$WORK/*.json` → policy/confidence/outbox time-series,
|
||||
(5) extend outbox_writer → Decision Ledger, (6) add Infracost →
|
||||
pre-apply cost estimates.
|
||||
|
||||
## 2. Telemetry Reference Architecture (Nova-native adaptation)
|
||||
|
||||
The PO provided a full distributed-system telemetry reference
|
||||
architecture (CloudEvents 1.0 envelope, OpenTelemetry SDK, Kafka/NATS
|
||||
event bus, Prometheus hot path, ClickHouse warehouse, QLDB decision
|
||||
ledger, Infracost, drift detection, ML anomaly forecasting). Per
|
||||
D-120, we adopt the **principles** but implement with **Nova-native
|
||||
minimal tech**. The mapping:
|
||||
|
||||
| Direction's principle | Nova-native implementation (v1.17) |
|
||||
|---|---|
|
||||
| Events are the source of truth; dashboards are projections | Hybrid (D-125): existing file signals stay as files; collector reads them and emits normalized CloudEvents into `metrics/events.jsonl` + SQLite. New emitters emit CloudEvents directly. |
|
||||
| Every AI action is logged with confidence + alternatives | Decision Ledger (D-121): `outbox_writer.py` extended → SQLite append-only hash-chain table. `ai.decision.made` modeled from confidence_signal (D-122): decision_id=run_id, chosen_action=band, confidence=score, alternatives=perInput, human_override=HITL block. |
|
||||
| Hot/cold storage split | Cold-only SQLite (D-126): `metrics/nova_metrics.db`. Hot path deferred (no live ops, D-096). |
|
||||
| Read-only external integrators | Infracost (pre-apply, offline, reads plan JSON). Cloud billing CUR deferred (D-096). Carbon APIs deferred (future). |
|
||||
| CloudEvents 1.0 envelope | Adopted. `core/metrics/event_envelope.py` defines the envelope + `platform.*` semantic conventions. |
|
||||
| Decision Ledger = append-only with hash chain + outcome backfill | SQLite append-only table with hash chain (D-121). Outcome backfilled from apply.completed via decision_id → request_id correlation. Honors D-083 (no S3 Object Lock/JWS). |
|
||||
| Cost governance: mandatory tags + Infracost pre-apply | Nova already enforces `nova:*` tags (nova_tagging.py, hard mode). Infracost added as plan post-processor (D-120). Post-apply CUR deferred (D-096). |
|
||||
| Definition-of-success docs for every KPI | Per-KPI docs in `docs/metrics/` (D-127). |
|
||||
| Replay-ability | SQLite store + JSONL event log are replayable by design. |
|
||||
|
||||
### CloudEvents envelope (Nova-native)
|
||||
|
||||
```json
|
||||
{
|
||||
"specversion": "1.0",
|
||||
"id": "<uuid>",
|
||||
"source": "nova.platform",
|
||||
"type": "nova.run.completed",
|
||||
"time": "<ISO8601>",
|
||||
"subject": "<contractId>/<env>",
|
||||
"datacontenttype": "application/json",
|
||||
"platform": {
|
||||
"tenant_id": "acdl",
|
||||
"run_id": "run-<epoch>",
|
||||
"contract_id": "<uuid>",
|
||||
"environment": "dev|qa|prod|dr",
|
||||
"actor": {"type": "confidence-gate", "id": "confidence_signal"},
|
||||
"trace_id": "<run_id>"
|
||||
},
|
||||
"data": {
|
||||
"duration_ms": 4800,
|
||||
"stages": ["resolve", "adapt", "validate", "plan", "apply"],
|
||||
"exit_code": 0,
|
||||
"confidence": {"score": 0.94, "band": "pass", "perInput": {...}},
|
||||
"policy": {"passed": 12, "failed": 0, "skipped": 0},
|
||||
"hitl": {"gate": "dev", "result": "autonomous", "block": false},
|
||||
"cost_estimate_usd": -12.40,
|
||||
"decision_id": "run-<epoch>",
|
||||
"outcome": "succeeded"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### Core event types (Nova-native minimum viable set)
|
||||
|
||||
| Event type | Emitted by | Purpose | Grounding |
|
||||
|---|---|---|---|
|
||||
| `nova.run.started` | run_platform.sh | Measures demand; provisioning lead time start | new emitter (P1) |
|
||||
| `nova.run.completed` | run_platform.sh | Run count, stage durations, exit, MTTR | new emitter (P1) |
|
||||
| `nova.run.failed` | run_platform.sh | Failure count, MTTR numerator | new emitter (P1) |
|
||||
| `nova.policy.evaluated` | checkov_adapter.py | Policy pass rate, compliance KPIs | grounded (pcr.json → P1 persists) |
|
||||
| `nova.confidence.computed` | confidence_signal.py | Confidence distribution, decision accuracy | grounded (signal.json → P1 persists) |
|
||||
| `nova.ai.decision.made` | outbox_writer.py (extended) | Decision Ledger entry | grounded (D-121, D-122) |
|
||||
| `nova.attestation.recorded` | hitl_gates.py | Attestation Coverage, human-in-the-loop audit | grounded (D-132) |
|
||||
| `nova.cost.estimated` | Infracost post-processor | Pre-apply cost estimate | new emitter (P1, Infracost) |
|
||||
| `nova.capability.verified` | regression_verify.py | Capability health, regression gate | grounded (REGRESSION_REPORT.json) |
|
||||
| `nova.test.completed` | pytest (junit XML) | Test count, pass rate | new (P1 adds --junitxml) |
|
||||
|
||||
## 3. Metric-to-Signal Scorecard (the "no fabrication" contract)
|
||||
|
||||
| Executive metric (NORTH_STAR target) | Status | Source / formula | Decision |
|
||||
|---|---|---|---|
|
||||
| Touchless Resolution Rate ≥99% | grounded (after P1) | runs without operational HITL block ÷ total runs (attestation gates excluded) | D-122, D-132 |
|
||||
| Human Escalation Frequency <0.1% | grounded (after P1) | operational HITL blocks ÷ total runs (attestation sign-offs excluded) | D-122, D-132 |
|
||||
| MTTR (p95) <60s | grounded (platform-run) | apply.failed.time → successful retry.time | D-131 |
|
||||
| Predictive vs Reactive ≥3:1 | **deferred** | requires ML forecasting (future emitter) | future |
|
||||
| AI Decision Accuracy ≥99.5% | grounded (after decision ledger) | decisions not followed by apply.failed/incident within 5min | D-121, D-122 |
|
||||
| Drift Auto-Reversal ≥95% | **deferred** | requires drift detection (D-096 + scheduler) | D-096 |
|
||||
| Cloud Spend Reduction ≥25% | partial | pre-apply estimate grounded (Infracost); actuals deferred (D-096 CUR) | D-120 |
|
||||
| L1/L2 Ops Hours Avoided ≥70% | derived | formula: run count × manual baseline minutes × blended rate | D-127 |
|
||||
| Platform ROI ≥250% | derived | formula: (labor savings + cloud savings + avoided downtime) ÷ platform op cost | D-127 |
|
||||
| Decision Ledger Coverage 100% | grounded (this milestone) | outbox_writer.py → SQLite hash-chain | D-121 |
|
||||
| Attestation Coverage 100% | grounded | hitl_gates.py + outbox approver_* attributes; prod/dr | D-132 |
|
||||
| AI-Agent Intent Share ≥40% | future | no AI-agent consumers today; placeholder view | future |
|
||||
| Capability health (18V+4S) | grounded | REGRESSION_REPORT.json | existing |
|
||||
| Confidence score distribution | grounded (after P1) | signal.json → decision ledger | D-121 |
|
||||
| Policy pass rate | grounded (after P1) | pcr.json → persisted | D-120 |
|
||||
| Test count / pass rate | grounded (after P1) | pytest --junitxml | D-120 |
|
||||
| Provisioning Lead Time | grounded (after P1) | run.started → run.completed | D-120 |
|
||||
| Deployment Frequency | grounded (after P1) | count(run.completed) per day | D-120 |
|
||||
| Deploy-failure alert count | grounded | GitHub issues via Lambda report_error (D-055) | existing |
|
||||
| Cost figures (actuals) | manual one-shot | COST.md (Cost Explorer query) | existing |
|
||||
| FTE Hours Saved / TRV | derived | formula over run count + COST.md | D-127 |
|
||||
| Self-healing velocity | **deferred** | no auto-remediator | future |
|
||||
| SLA / unplanned downtime | **deferred** | needs live service uptime (D-096) | D-096 |
|
||||
| GreenOps / carbon | **deferred** | WattTime/Electricity Maps API (future) | future |
|
||||
| Red Team MTTR | **deferred** | no red-team program | future |
|
||||
| Multi-cloud normalization | **deferred** | Nova is AWS-only | future |
|
||||
| Live CUR reconciliation | **deferred** | needs live AWS billing (D-096) | D-096 |
|
||||
|
||||
## 4. Deferred-Decision Ledger (constraints on this milestone)
|
||||
|
||||
| Decision | Scope | Grounding impact |
|
||||
|----------|-------|------------------|
|
||||
| D-096 | Live AWS torn down post-v1.11 | BLOCKS all live-AWS metrics (CAP-013..016 Skipped; live outbox; live state bucket; live CUR) |
|
||||
| D-083 | S3 Object Lock + JWS + async worker deferred | BLOCKS tamper-evident ledger; v1.17 uses SQLite hash-chain instead |
|
||||
| D-113/D-114/D-119 | Onboarding = request-path only; no auto-grant | BLOCKS onboarding funnel "granted" half |
|
||||
| D-091/D-118 | Regression gate (D-091) gates milestone completion | ENABLES the strongest metric signal (REGRESSION_REPORT.json) |
|
||||
| D-092 | Local emulating adapters | ENABLES offline E2E metrics (CAP-011/012) |
|
||||
| D-055 | report_error Lambda action creates GitHub issues | ENABLES deploy-failure alert metric |
|
||||
| D-050 | Publish deploy outputs to SSM + GitHub PR comment | ENABLES outputs-published metric |
|
||||
| D-054/D-043/D-109 | Nova tagging standard (hard mode) | ENABLES tagging-compliance metric |
|
||||
| D-084 | 8-concern attestation matrix | ENABLES attestation metrics (operator-supplied evidence artifacts) |
|
||||
| D-089 | Signature verification skipped when signing key unset (dev/CI) | Signature metrics are no-ops in dev |
|
||||
|
||||
## 5. Deck-Storytelling Research (x3 arc + per-slide benefit)
|
||||
|
||||
### The "tell them x3" structure
|
||||
|
||||
The PO's direction: "Tell them what you're going to tell them, then tell
|
||||
them, then tell them what you told them." Applied at two levels:
|
||||
|
||||
**Deck level (the 5-act arc):**
|
||||
1. **Opening slide** = "what I'm going to tell you" — the full arc
|
||||
preview: Problem → Vision → How → Proof → Roadmap.
|
||||
2. **Body** (acts 1–5) = "tell them" — each act delivers its content.
|
||||
3. **Closing slide** = "what I told you" — recap of the 5 acts + the ask.
|
||||
|
||||
**Per slide:**
|
||||
1. **Slide opens** with what it'll cover (1 line: "This slide shows X").
|
||||
2. **Slide delivers** the content (bullets, diagram, or table).
|
||||
3. **Slide closes** with an explicit **"benefit of this stage" callout**
|
||||
(1 line: "Benefit: you now know Y" or "Why this matters: Z").
|
||||
|
||||
### Fluidity conventions
|
||||
|
||||
- **Transitions are written, not hand-waved.** Each slide's opening line
|
||||
references the previous slide's close ("Having seen X, now consider Y").
|
||||
- **No disjointed jumps.** If a topic shift is needed, a bridge slide or
|
||||
a transition sentence carries the audience across.
|
||||
- **The arc is visible.** A small "act indicator" in the Marp footer
|
||||
(e.g., `Act 3/5: How it works`) keeps the audience oriented.
|
||||
|
||||
### Existing deck inventory (to be retired)
|
||||
|
||||
Two decks exist today in `docs/presentations/`:
|
||||
- `how-the-platform-works.md` (32,916 bytes) → marp → html → talking-points
|
||||
- `the-developer-experience.md` (27,509 bytes) → marp → html → talking-points
|
||||
|
||||
Both follow a 4-step process (source `.md` → Marp → HTML → talking-points)
|
||||
documented in `docs/presentations/README.md`. Per D-130, both are merged
|
||||
into one unified narrative deck and retired.
|
||||
|
||||
### Grounded metrics already cited in existing decks
|
||||
|
||||
- "22/22 auto-verifiable capabilities Verified" — **stale** vs current
|
||||
REGRESSION_REPORT.json (18V+4S post-D-096). The unified deck must
|
||||
derive this from the report, not copy the stale claim.
|
||||
- Confidence thresholds: dev ≥0.50, qa ≥0.75, prod ≥0.90, dr ≥0.95 —
|
||||
grounded in `core/confidence_signal.py:57` (THRESHOLDS).
|
||||
- RPO = 0 (evidence write synchronous) — grounded in
|
||||
`core/audit_ledger_design.md:27,103`.
|
||||
- Cost figures — `how-the-platform-works.md:461`; cites COST.md.
|
||||
- Confidence signal 6 inputs + weights — grounded in
|
||||
`core/confidence_signal.py:40-47`.
|
||||
- "~80-line stateless adapter" vs "918-line monolith" — grounded in
|
||||
ROADMAP/RESEARCH prose.
|
||||
|
||||
### Planned deck structure (for PLAN to detail)
|
||||
|
||||
The unified deck "Nova — The No-Humans Infrastructure Platform":
|
||||
|
||||
| Act | Slides | Content | Proof source |
|
||||
|---|---|---|---|
|
||||
| 1. Problem | 2–3 | The no-humans imperative; why operators are the bottleneck; the trust gap | NORTH_STAR vision |
|
||||
| 2. Vision/Direction | 2–3 | Nova's vision; 4 strategic objectives; anti-goals; the attestation model (autonomy in operations, human at stage gates) | NORTH_STAR |
|
||||
| 3. How it works | 3–4 | Contract → resolver → adapter → confidence → HITL gate; the Decision Ledger; the 8-concern attestation matrix | code grounding |
|
||||
| 4. Proof (metrics) | 3–4 | Capability health (18V+4S); confidence distribution; policy pass rate; Decision Ledger coverage; Attestation Coverage; cost estimates; the grounded/derived/deferred honesty model | metrics export |
|
||||
| 5. Roadmap/Ask | 2 | 12–18mo targets (committed); deferred metrics (honest); the ask | NORTH_STAR targets |
|
||||
|
||||
Total: ~12–16 slides. Opening = arc preview; closing = recap + ask.
|
||||
|
||||
## 6. Assumptions logged (v1.17)
|
||||
|
||||
- A1 (0.9): No live AWS access during execution (consistent with
|
||||
v1.11–v1.16). All metrics that require live AWS ship as placeholder
|
||||
views. The Infracost integration runs offline (reads plan JSON).
|
||||
- A2 (0.85): The Decision Ledger SQLite hash-chain is sufficient for
|
||||
v1.17's audit needs. The full tamper-evident ledger (S3 Object Lock +
|
||||
JWS, D-083) is a future milestone. The hash-chain provides
|
||||
append-only + integrity verification locally.
|
||||
- A3 (0.8): The "AI decision" framing (D-122) is honest: Nova's "AI" is
|
||||
the confidence-gated policy engine (confidence_signal + HITL gate),
|
||||
not an LLM planner. The deck and METRICS.md must frame this accurately
|
||||
— overclaiming "AI" would violate the "no fabrication" constraint.
|
||||
- A4 (0.85): The unified deck's "Proof" section cites only grounded
|
||||
metrics with real numbers. Deferred metrics are shown as "Planned"
|
||||
with the `<span class="badge planned">Planned</span>` badge. No
|
||||
fabricated numbers in any slide.
|
||||
- A5 (0.8): `--junitxml` + `--json-report` added to pytest addopts
|
||||
does not break the existing test suite (the flags are additive; pytest
|
||||
continues to run normally). CAP-009 (offline pytest suite passes)
|
||||
must remain Verified after the change.
|
||||
- A6 (0.75): Infracost is available as a CLI tool that can be installed
|
||||
in the CI environment and run locally. It reads `terraform plan
|
||||
-out=plan.tfplan` + `terraform show -json plan.tfplan` to produce a
|
||||
cost estimate. No live AWS access required. If Infracost is not
|
||||
available, the `cost.estimated` event is omitted (degraded mode, not
|
||||
a failure).
|
||||
|
||||
+302
-90
@@ -1,112 +1,324 @@
|
||||
# Nova v1.16 — Multi-Persona Code Review (final phase P21)
|
||||
# Nova v1.11 — Multi-Persona Code Review (P60–P65 retrofit + new work)
|
||||
|
||||
**Reviewer:** lead-developer (model: glm-5.2)
|
||||
**Scope:** v1.16 milestone — 22 tags (v1.15.5..v1.15.26), 20 execution
|
||||
phases + final. Squash-merged to main via `milestone/v1.16-nova-simplification`.
|
||||
**Date:** 2026-07-30
|
||||
**Reviewer:** ci-code-reviewer (model: glm-5.2)
|
||||
**Scope:** v1.11 milestone, branch `milestone/v1.11-restart` — 22 commits
|
||||
(e1bb214..8c09580), 25 files, +790/-142 lines
|
||||
**Date:** 2026-07-29
|
||||
|
||||
> **Historical note:** REVIEW.md was reconstructed at v1.16 P21 (the
|
||||
> v1.3–v1.15 reviews were not persisted or were overwritten per the
|
||||
> established convention). The v1.16 review overwrites prior content.
|
||||
## Commits reviewed
|
||||
|
||||
## Review approach
|
||||
|
||||
The v1.16 milestone is an NFR sweep (no new features). Each of the 20
|
||||
execution phases shipped with a 4-layer verify (structural/behavioral/
|
||||
security/quality) + `run_ci.sh` 3-stage PASS at every phase boundary.
|
||||
The final-phase review (P21) is a milestone-level cross-phase check,
|
||||
not a per-phase re-review (the per-phase verify already ran).
|
||||
| Commit | Phase | Type | Summary |
|
||||
|--------|-------|------|---------|
|
||||
| e1bb214 | 60 | docs | retrofit plan — L1 lifecycle pipeline live-run |
|
||||
| bc9058f | 60 | feat | L1 module lifecycle live run — module fixes (retrofit) |
|
||||
| bb3ac7c | 60 | fix | WAF scope case + VPC modify DependencyViolation |
|
||||
| 0c5c4d1 | 61 | docs | create phase plan — L2 lifecycle pipeline author |
|
||||
| 361fe60 | 61 | feat | L2 lifecycle pipeline — extend matrix + workflows + tests |
|
||||
| 9ac5720 | 61 | verify | 4-layer gate — PASS |
|
||||
| 6441633 | 62 | docs | create phase plan — L2 lifecycle pipeline live run |
|
||||
| 4dad967 | 60 | fix | ALB target group name_prefix — avoid orphaned conflicts |
|
||||
| adfcf86 | 63 | docs | create phase plan — regression registry + cost docs |
|
||||
| b71e63c | 63 | feat | CAP-017..022 regression registry + COST.md |
|
||||
| beac2ef | 63 | verify | 4-layer gate — PASS |
|
||||
| 06f4fc7 | 60 | fix | free disk space in lifecycle jobs |
|
||||
| 92bb03e | 64 | docs | create phase plan — pre-mortem + teardown |
|
||||
| 186cdde | 64 | feat | pre-mortem — v1.10 post-mortem + forward pre-mortem |
|
||||
| 4102950 | 64 | feat | pre-mortem + teardown plan — HITL escalation CHG0680001 |
|
||||
| 7c4fc1f | 64 | feat | teardown complete — zero live ACDL resources remain |
|
||||
| a52f8a5 | 64 | verify | 4-layer gate — PASS |
|
||||
| a03c019 | 60/62 | fix | ALB name_prefix + adapter dedup + L2 composition wiring |
|
||||
| 93a6598 | 65 | docs | create phase plan — rewrite caps + decks |
|
||||
| 6394801 | 65 | feat | rewrite caps — CAP-017..022 Verified via lifecycle pipeline |
|
||||
| fc91f24 | 65 | verify | 4-layer gate — PASS |
|
||||
| 8c09580 | 65 | docs | update v1.11 status — all phases complete |
|
||||
|
||||
## P0 issues (0)
|
||||
|
||||
No blocking issues found. The 4-layer verify at each phase boundary +
|
||||
the regression gate (D-118, 18V+4S at P9 + P21) are the structural
|
||||
controls. No P0 was auto-applied at P21.
|
||||
No blocking issues found. The targeted fixes are correct for their stated
|
||||
purposes. The 447 fast offline tests pass (485/490 collected; 5 slow
|
||||
deselected, including 2 slow regression-integration tests that exercise the
|
||||
CAPABILITY_REGISTRY against the live codebase).
|
||||
|
||||
## P1 issues (0)
|
||||
## P1 issues (5 — should fix)
|
||||
|
||||
No P1 issues flagged. The grill binding decisions (G-111..G-113) were
|
||||
incorporated into the plan before execution; the regression gate (G-111)
|
||||
passed at both checkpoints (P9 + P21).
|
||||
### P1-1: Adapter dedup silently drops resources whose module is not in the registry
|
||||
[correctness] `adapters/terraform/adapter.py:159-170`
|
||||
|
||||
## P2 issues (2 — post-hoc, non-blocking)
|
||||
The new dedup loop only adds resources to `seen` when `tf_dir` is truthy
|
||||
(in the registry). A resource whose module is missing from the registry is
|
||||
**silently dropped** from `merged` — it never reaches `_emit_module_block`,
|
||||
so no error is raised. The pre-dedup code (`parts.extend(... for r in
|
||||
resources)`) would have raised `ValueError("no terraform_dir in registry
|
||||
for module ...")` via `_emit_module_block`, surfacing the misconfiguration.
|
||||
|
||||
### P2-1: Onboarding framing (E-002, deferred from grill)
|
||||
[scope] `.ciagent/PROJECT.md`, `.ciagent/ROADMAP.md`
|
||||
Confirmed by simulation: two resources, one with `module: nonexistent@1.0.0`,
|
||||
produces a `merged` list of length 1 — the unknown-module resource vanishes
|
||||
without diagnostic.
|
||||
|
||||
The grill escalation E-002 (confidence 0.55) flagged that the PROJECT.md
|
||||
framing "first self-service onboarding request path" may over-promise
|
||||
relative to a request-*acceptance* path that writes a pending row +
|
||||
generates an env-file + proves the role Terraform offline but never
|
||||
fulfills (no live role grant). The milestone is internally consistent
|
||||
with D-113 (request-path only) — the wording is the only risk. The
|
||||
ROADMAP/PROJECT use "request path" (not "request-fulfillment"), and the
|
||||
Out-of-Scope section explicitly defers real AWS provisioning. **Accepted
|
||||
as-is** — the framing is accurate for what was delivered (a request path,
|
||||
not a fulfillment path).
|
||||
**Recommendation:** in the dedup loop, when `tf_dir` is `None`, either
|
||||
(a) raise immediately (preserving the prior contract), or (b) append the
|
||||
resource to a separate `unknown` list and extend `parts` with it so
|
||||
`_emit_module_block` raises the descriptive error. As written, a typo in
|
||||
a composition's `module` field (e.g. `iam-role@1.0.0` vs `iam_roles@1.0.0`)
|
||||
will silently omit a resource from the emitted terraform — a class of
|
||||
defect the v1.10 sweep was specifically created to catch.
|
||||
|
||||
### P2-2: REVIEW.md + AUDIT.md not updated during the run
|
||||
[maintainability] `.ciagent/REVIEW.md`, `.ciagent/AUDIT.md`
|
||||
### P1-2: L2 static-assets "modify" example is a no-op — complex ≡ simple
|
||||
[correctness] `modules/l2/static-assets/examples/complex.yml`,
|
||||
`modules/l2/static-assets/composition.json`
|
||||
|
||||
REVIEW.md still held v1.11 content during the v1.16 run (the per-phase
|
||||
verify ran but wasn't persisted to REVIEW.md until P21). AUDIT.md held
|
||||
v1.15 content. Both are reconstructed at P21 (this review + the audit
|
||||
running now). This matches the established convention (REVIEW.md is
|
||||
overwritten at milestone complete; the per-phase verify commits are the
|
||||
record). Not a defect.
|
||||
The complex.yml comment claims "Modify variant: same bucket_name as simple
|
||||
(in-place modify, adds CDN + WAF)". But resolving both examples yields
|
||||
**identical** resource sets: `['s3','cloudfront-distribution',
|
||||
'cloudfront-originaccesscontrol','waf','kms']`. The CDN and WAF are
|
||||
**always present** in the static-assets composition (they are unconditional
|
||||
children + wires); the `waf_enabled`, `default_ttl`, `max_ttl`,
|
||||
`price_class`, `viewer_protocol_policy` inputs in complex.yml have **no
|
||||
corresponding wires** in composition.json and are silently dropped at
|
||||
resolve time. So the L2 static-assets lifecycle cell's "modify" step
|
||||
applies a contract that produces the same terraform as "simple" — it
|
||||
exercises `terraform apply` twice with no change, not a true modify.
|
||||
|
||||
This is not a regression (the inputs were never wired), but the
|
||||
CAPABILITY_INVENTORY claim "CAP-020 Verified live-aws via L2 static-assets
|
||||
lifecycle pipeline (apply/modify/destroy exit 0)" overstates what the
|
||||
modify step proves: it proves idempotent re-apply, not in-place modify.
|
||||
|
||||
**Recommendation:** either (a) wire `waf_enabled`/`default_ttl`/etc. in
|
||||
composition.json so the complex contract genuinely differs, or (b) correct
|
||||
the comment + CAPABILITY_INVENTORY wording to "apply + idempotent re-apply
|
||||
+ destroy" rather than "apply/modify/destroy". The microservice complex
|
||||
example, by contrast, is a real modify (desired_count 1→2) — that one is
|
||||
fine.
|
||||
|
||||
### P1-3: L2 lifecycle scripts ignore the ci-vpc-outputs.json argument
|
||||
[correctness] `scripts/run_l2_lifecycle_test.sh:14`,
|
||||
`scripts/run_l2_lifecycle_destroy.sh:12`
|
||||
|
||||
Both L2 scripts declare `Usage: ... <module> <example> [ci-vpc-outputs.json]`
|
||||
but neither reads `$3`/`$2`. The microservice composition references the
|
||||
platform VPC via `terraform_remote_state` (data source), and the script
|
||||
sets `ACDL_REMOTE_STATE_KEY=spike/ci-vpc/terraform.tfstate` so the data
|
||||
source reads from the CI VPC state — that part is correct. But the
|
||||
`ci-vpc-outputs.json` argument is positional noise: the workflow passes
|
||||
it (`run_l2_lifecycle_test.sh ${{ matrix.module }} simple
|
||||
/tmp/ci-vpc-outputs.json`) and it is silently ignored. The L1 scripts
|
||||
(`run_lifecycle_test.sh`) inject VPC outputs by rewriting the contract in
|
||||
Python; the L2 path takes a different approach (remote state) and does not
|
||||
need the file, so the argument is vestigial, not a bug — but the usage
|
||||
string advertises a feature the script does not provide, which will
|
||||
confuse a future maintainer who assumes parity with the L1 scripts.
|
||||
|
||||
**Recommendation:** remove the `[ci-vpc-outputs.json]` token from the
|
||||
usage strings (or add a comment explaining the L2 path uses remote state
|
||||
and the arg is accepted-but-ignored for workflow-argument parity).
|
||||
|
||||
### P1-4: CAPABILITY_INVENTORY summary table is stale (says 16, body lists 22)
|
||||
[maintainability] `.ciagent/CAPABILITY_INVENTORY.md:9-16`
|
||||
|
||||
The Summary table still reads "Verified 16 / Decayed 0 / Broken 0 / Total
|
||||
16" — the v1.10 sweep count. The body (lines 93-110) now lists CAP-017..022
|
||||
as **Verified** via the lifecycle pipeline, bringing the real total to 22.
|
||||
The two counts disagree: a reader scanning the summary sees 16 Verified; a
|
||||
reader scanning the inventory body sees 22 Verified. The PRE_MORTEM
|
||||
(lines 82-83) and CAPABILITY_INVENTORY prose both assert all 22 are
|
||||
Verified, but the headline table was not updated in the P65 rewrite.
|
||||
|
||||
**Recommendation:** update the Summary table to "Verified 22 / Decayed 0
|
||||
/ Broken 0 / Total 22" and add CAP-017..022 rows to the Inventory table
|
||||
(the body section "Cloud capabilities NOT re-verified..." is now
|
||||
mis-titled — they ARE verified, just via the lifecycle-pipeline tier).
|
||||
|
||||
### P1-5: CAP-017..022 regression checks are offline proxies, not pipeline evidence
|
||||
[adversarial] `core/regression_verify.py:432-519`,
|
||||
`.ciagent/CAPABILITY_INVENTORY.md:93-110`
|
||||
|
||||
The CAP-017..022 checks (`_check_cap_017_dynamodb` etc.) call
|
||||
`_check_lifecycle_module_terraform` / `_check_lifecycle_l2_module`, which
|
||||
verify only that (a) the terraform dir + required files exist and (b) the
|
||||
example contracts **resolve** (resolver exit 0). They do **not** run
|
||||
`terraform validate`, do not run apply/modify/destroy, and do not query
|
||||
the pipeline's actual green/red status. The CAPABILITY_INVENTORY claims
|
||||
"Evidence = L1 rds module lifecycle pipeline green (terraform validate +
|
||||
contracts resolve)" — but the check does not run terraform validate, and
|
||||
"lifecycle pipeline green" is asserted, not verified by the regression
|
||||
gate.
|
||||
|
||||
This means the lifecycle-pipeline evidence CAN be faked at the regression
|
||||
tier: a module whose terraform is syntactically broken (e.g.
|
||||
`scope = upper(var.scope)` removed, or a missing required variable) would
|
||||
still pass `_check_lifecycle_module_terraform` as long as the files exist
|
||||
and the resolver runs. The real green/red evidence lives only in the
|
||||
workflow run history (Gitea/GitHub Actions), which the regression gate does
|
||||
not read.
|
||||
|
||||
**Mitigation context:** the modules-lifecycle workflow IS the live
|
||||
evidence — when it runs on a PR, the cells genuinely apply/modify/destroy
|
||||
against live AWS. The gap is that the *regression gate* (which gates
|
||||
milestone COMPLETE) trusts the workflow will be run, rather than proving it
|
||||
was run and passed. A milestone could in principle be marked COMPLETE with
|
||||
CAP-017..022 "Verified" if the regression gate runs but the workflow was
|
||||
never executed (e.g. workflow_dispatch never triggered, or the PR was
|
||||
merged without the workflow running).
|
||||
|
||||
**Recommendation:** (a) tighten the CAP-017..022 check docstrings + the
|
||||
CAPABILITY_INVENTORY wording to "terraform files present + contracts
|
||||
resolve (offline proxy; live apply/modify/destroy verified by the
|
||||
modules-lifecycle workflow run, not by this gate)"; and/or (b) add a
|
||||
`terraform validate` step to `_check_lifecycle_module_terraform` (slow but
|
||||
cheap relative to init+apply) so at least HCL syntax is verified at the
|
||||
gate. The teardown trustworthiness (P64) is good — `ci-vpc-destroy` runs
|
||||
`if: always()` and the decommission `---ci---` block is the audit trail.
|
||||
|
||||
## P2 issues (4 — post-hoc)
|
||||
|
||||
### P2-1: ALB `name_prefix = "tg-ci-"` discards `var.name` entirely
|
||||
[maintainability] `modules/l1/alb/terraform/main.tf:9`
|
||||
|
||||
The fix replaces `name = var.name` with `name_prefix = "tg-ci-"` (a
|
||||
hardcoded literal). This is the correct terraform pattern for
|
||||
create_before_destroy resources with name-uniqueness constraints, and the
|
||||
commit message explains the orphaned-resource motivation well. However
|
||||
the target group name is now non-configurable (always `tg-ci-<random>`),
|
||||
and the `var.name` variable is no longer used by the target group at all
|
||||
(it is still used by `aws_lb.this.name`). A consumer who sets `name:
|
||||
my-app` gets an LB named `my-app` but a target group named `tg-ci-...` —
|
||||
inconsistent tagging. Consider `name_prefix = "${var.name}-"` to keep the
|
||||
consumer's name as a prefix while preserving uniqueness. Post-hoc: not
|
||||
blocking; the lifecycle pipeline is the only current consumer and `tg-ci-`
|
||||
is fine for CI.
|
||||
|
||||
### P2-2: No test covers the new dedup merge behavior or `ACDL_REMOTE_STATE_KEY`
|
||||
[testing] `tests/test_adapter.py`, `tests/test_pipeline_contract.py`
|
||||
|
||||
The adapter gained (a) a dedup-merge loop for multi-resource L1s sharing a
|
||||
terraform dir and (b) `ACDL_REMOTE_STATE_KEY` env override for the remote
|
||||
state data block. Neither has a unit test:
|
||||
- No test asserts that two resources with the same `module` collapse to one
|
||||
`module "<first_id>" { ... }` block with merged inputs.
|
||||
- No test asserts that `ACDL_REMOTE_STATE_KEY` overrides the default
|
||||
`platform/terraform.tfstate` key in the emitted `data
|
||||
terraform_remote_state` block.
|
||||
- No test covers the L2 lifecycle scripts (`run_l2_lifecycle_test.sh` /
|
||||
`run_l2_lifecycle_destroy.sh`) — the L1 equivalents are also untested at
|
||||
the script level, so this is consistent with existing practice, but the
|
||||
L2 scripts are new in this session and the `ACDL_REMOTE_STATE_KEY` wiring
|
||||
is the load-bearing correctness mechanism for the microservice lifecycle.
|
||||
|
||||
The 485 offline tests adequately cover the *contract* (pipeline schema,
|
||||
byte-identical workflows, matrix membership, job needs) — the
|
||||
`TestModulesLifecyclePipeline` class is solid (89 tests pass). The gap is
|
||||
adapter *behavior* at the unit level.
|
||||
|
||||
**Recommendation:** add a `test_adapter_dedup_merges_same_module` and a
|
||||
`test_adapter_remote_state_key_override` to `tests/test_adapter.py`.
|
||||
|
||||
### P2-3: `waf` complex example uses `scope: CLOUDFRONT` but WAF scope is now `upper()`'d
|
||||
[correctness] `modules/l1/waf/examples/complex.yml:8`,
|
||||
`modules/l1/waf/terraform/locals.tf:3`
|
||||
|
||||
The `locals.tf` change `scope = upper(var.scope)` is the correct defensive
|
||||
fix (the AWS provider requires `CLOUDFRONT`/`REGIONAL` regardless of input
|
||||
case). The complex.yml was simultaneously changed from `scope: cloudfront`
|
||||
to `scope: CLOUDFRONT`. Both are now correct, but the example's uppercase
|
||||
value is now redundant with the `upper()` — a future reader may wonder
|
||||
which is authoritative. Minor; the defensive `upper()` is the right call
|
||||
and the example matching it is fine. Post-hoc only.
|
||||
|
||||
### P2-4: COST.md reproducibility snippet could leak the account ID via CloudTrail
|
||||
[security] `.ciagent/COST.md:106`
|
||||
|
||||
COST.md contains the AWS account ID `581513795199` in multiple places
|
||||
(summary, S3 bucket name, methodology). This is consistent with the rest of
|
||||
the repo (the bucket name `acdl-tfstate-581513795199-us-east-1` is hardcoded
|
||||
in `adapter.py:130` and `adapter.py:146`), so it is not new leakage and not
|
||||
a regression. No actual secret material (access keys, secret access keys)
|
||||
appears in COST.md, PRE_MORTEM.md, CAPABILITY_INVENTORY.md, or the workflow
|
||||
files — all credential references use `${{ secrets.ACDL_AWS_* }}` or env
|
||||
var names only. The `.ciagent/PROJECT.md:731` reference to a deactivated
|
||||
root key is redacted (`AKIA…ROOT-DEACTIVATED`). **No credential leakage
|
||||
found.** The P2 is only that the account ID is published; if the account
|
||||
is meant to be opaque, this is an accepted exposure (the bucket name
|
||||
already requires it).
|
||||
|
||||
## What is correct
|
||||
|
||||
- **State-bucket drift fix (P1):** `adapter.py:117` now emits
|
||||
`nova-tfstate-*` (matching the live bucket renamed in v1.15 P4). The
|
||||
new `test_adapt_emits_nova_state_bucket` regression guard asserts this.
|
||||
- **Kyverno label fix (P1):** `require-resource-labels.yml` enforces
|
||||
`nova:*` labels (consistent with `nova_tagging.py` hard-fail on
|
||||
`acdl:*`). No policy contradiction.
|
||||
- **Ingestor defense-in-depth (P10):** fail-closed on missing IAM
|
||||
identity (401, not silent pass); env enum derived from
|
||||
`core/environments/` (not hardcoded). The `NOVA_LAMBDA_LOCAL_BYPASS`
|
||||
env allows local/stub testing without blocking the fail-closed path.
|
||||
- **Payload validation (P11):** 256 KB size cap + contract.schema.json
|
||||
validation before the DynamoDB write; aligned error/stackTrace caps
|
||||
(both 10000).
|
||||
- **Regression gate (G-111):** CAP-013..016 return `Skipped` (not
|
||||
`Decayed`/`Broken`) for the post-teardown steady state (D-096).
|
||||
`passed` accepts Skipped. Gate passes at 18V+4S.
|
||||
- **Workflow generator (P8):** `sync_workflows.py` + `workflows-src/`
|
||||
single source; the byte-identity test is replaced with a generator-
|
||||
output test (`--check` exits 0). The 3 pairs are no longer hand-synced.
|
||||
- **Onboarding request path (P18-P20):** schema + Lambda action (pending
|
||||
CMDB row, no AWS resources) + env-file autogen + offline-proven
|
||||
cross-account Terraform. Self-service message (no "contact the platform
|
||||
team"). Real AWS provisioning explicitly deferred (D-113/D-114).
|
||||
- **Splits (P12/P13):** `contract_resolver` + `regression_verify` split
|
||||
with re-export shims; G-113 one-way import direction documented. All
|
||||
tests pass without modification (backwards compat preserved).
|
||||
- **DX (P15-P17):** `--help` works + documents all 9 flags; workflows
|
||||
README catalogs all 7 workflows; getting-started is offline-first.
|
||||
- **Regression gate:** 18 Verified + 4 Skipped at P9 + P21 (0 Decayed/
|
||||
Broken). The 4 Skipped are the post-v1.11-teardown live-AWS caps.
|
||||
- **WAF scope fix (`upper(var.scope)`):** correct and defensive; AWS
|
||||
provider v5 requires uppercase. The `local.scope` indirection is clean.
|
||||
- **VPC `create_before_destroy` + same-CIDR complex example:** correct
|
||||
fix for the DependencyViolation on modify. Using the same CIDR means
|
||||
terraform modifies in-place rather than replacing the VPC (which would
|
||||
cascade-fail on dependent subnets/IGW). The `create_before_destroy`
|
||||
lifecycle is the right guard.
|
||||
- **ALB `name_prefix`:** correct terraform pattern for
|
||||
create_before_destroy + name-uniqueness; well-documented commit message.
|
||||
- **Adapter dedup (for the registered-module case):** correct —
|
||||
multi-resource L1s like cloudfront (distribution + OAC) correctly merge
|
||||
into one `module "cloudfront-distribution" { ... }` block. The merge
|
||||
preserves first-resource inputs and union of outputs. (The
|
||||
unregistered-module drop is P1-1, a separate concern.)
|
||||
- **L2 composition wiring (`ecr.inputs.name`, `roles.inputs.role_name`):**
|
||||
correct. Resolving microservice complex now shows `ecr.inputs.name =
|
||||
"app-repo"` and `roles.inputs.role_name = "app-role"` (defaults applied
|
||||
since the contract doesn't set `name`). Previously these would have hit
|
||||
the "missing required arg" defect class from the v1.10 sweep.
|
||||
- **Microservice complex = real modify:** `desired_count: 2` (vs simple's
|
||||
default 1) is a genuine in-place modify — confirmed by resolving both
|
||||
and diffing `service-service.inputs.desired_count`.
|
||||
- **`ACDL_REMOTE_STATE_KEY` plumbing:** correct end-to-end — the L2 scripts
|
||||
export it, the adapter reads it with a sensible default, and the
|
||||
microservice composition's `terraform_remote_state` data block picks it
|
||||
up. This cleanly separates the short-lived CI VPC state from the
|
||||
long-lived platform VPC state.
|
||||
- **Workflow structure:** `l2-lifecycle` correctly `needs: ci-vpc-apply`;
|
||||
`ci-vpc-destroy` correctly `needs: [lifecycle, l2-lifecycle]` and
|
||||
`if: always()`. The 7 new L2 pipeline-contract tests assert all of this.
|
||||
- **Byte-identical workflows:** `.gitea` and `.github` modules-lifecycle.yml
|
||||
are byte-identical (test asserts this); the `test_workflow_has_four_jobs`
|
||||
rename from three→four is correct.
|
||||
- **Adapter line count:** 194 lines — under the 200-line ceiling, still a
|
||||
clean stateless assembler. The dedup logic added ~16 lines without
|
||||
bloating.
|
||||
- **Teardown verification (P64):** trustworthy in structure — the
|
||||
`ci-vpc-destroy` job runs unconditionally and the decommission
|
||||
`---ci---` block is the audit trail. The adversarial concern (P1-5) is
|
||||
about the regression gate trusting the workflow ran, not about the
|
||||
teardown itself being fakeable.
|
||||
- **Security:** no credential leakage in any reviewed file. All AWS auth
|
||||
in workflows uses `${{ secrets.* }}`; COST.md references only env var
|
||||
names and a redacted/deactivated root key ID.
|
||||
|
||||
## Test coverage assessment
|
||||
## Test coverage assessment (485 offline tests)
|
||||
|
||||
~635 tests pass (was ~620 at v1.15.4). New test files:
|
||||
- `tests/test_onboarding.py` (3 tests — env-file generation)
|
||||
- `tests/test_onboarding_terraform.py` (3 tests — terraform validate + tags)
|
||||
- `tests/test_docs_coverage.py` (expanded — workflows README catalog)
|
||||
- **Adequate:** pipeline contract (89 tests), schema validation, contract
|
||||
resolution, adapter emission (basic), confidence signal, outbox,
|
||||
interpolation, local emulators, module-standards file presence, design-doc
|
||||
currency.
|
||||
- **Gaps (post-hoc):**
|
||||
1. Adapter dedup merge behavior (P2-2) — no unit test.
|
||||
2. `ACDL_REMOTE_STATE_KEY` override (P2-2) — no unit test.
|
||||
3. CAP-017..022 regression checks (P1-5) — not exercised at the unit
|
||||
level; the 2 slow tests in `test_verify_regression_mode.py` run the
|
||||
full registry but are `@pytest.mark.slow` and deselected from the
|
||||
fast suite, so a CI run of the 485 fast tests does not verify
|
||||
CAP-017..022 even at the offline-proxy level.
|
||||
4. WAF `upper()` scope — no test asserts the locals transform; relies
|
||||
on the lifecycle pipeline cell to catch a regression.
|
||||
5. ALB `name_prefix` — no test asserts the target group uses
|
||||
`name_prefix` (P2-1 context).
|
||||
|
||||
New tests in existing files: `test_adapt_emits_nova_state_bucket`,
|
||||
`test_onboarding_message_says_nova_not_acdl`, `test_no_identity_fails_closed`,
|
||||
`test_no_identity_passes_with_local_bypass`, `test_oversized_contract_rejected`,
|
||||
`test_schema_invalid_contract_rejected`, `TestNarrowedException` (2 tests),
|
||||
`TestOnboardConsumer` (3 tests), `TestOnboardingMessageSelfService` (2 tests),
|
||||
`test_sync_workflows_check_passes`.
|
||||
The 485 count is honest (447 pass fast, 5 deselected slow, 485/490
|
||||
collected). The gap is behavioral coverage of the new adapter + module
|
||||
logic, not contract/schema coverage.
|
||||
|
||||
## Verdict
|
||||
|
||||
**PASS — 0 P0, 0 P1, 2 P2 (post-hoc, accepted).** The v1.16 NFR milestone
|
||||
is complete. All 20 requirements (REQ-165..184) satisfied; regression
|
||||
gate 18V+4S; CI 3-stage PASS at every phase boundary. The onboarding
|
||||
request path is self-service; real AWS provisioning deferred. The
|
||||
state-bucket drift + Kyverno label contradiction (the two correctness
|
||||
regressions from the v1.15 rebrand) are fixed with regression guards.
|
||||
**PASS with P1 flags for post-hoc review.** No P0 fixes applied. The
|
||||
milestone's structural controls (regression gate, mandatory teardown,
|
||||
byte-identical workflows, byte-identical contract↔workflow tests) are
|
||||
sound. The most material finding is P1-5 (the regression gate's
|
||||
CAP-017..022 evidence is an offline proxy, not live pipeline evidence) —
|
||||
this is a repeat of the v1.10 "VERIFY was diff-scoped" structural defect
|
||||
in a milder form: the gate trusts the workflow was run rather than proving
|
||||
it. The mitigations in PRE_MORTEM (FM-1..FM-4) acknowledge related risks;
|
||||
P1-5 is the specific instance for the lifecycle-pipeline tier.
|
||||
@@ -1630,56 +1630,3 @@ milestone release). (G-104 binding.)
|
||||
- Tag `v1.15.4` created; milestone merged to main.
|
||||
|
||||
After Phase P5: milestone COMPLETE — `v1.15.4` IS the v1.15 release.
|
||||
|
||||
---
|
||||
|
||||
## v1.16 (complete — Nova Simplification, tag `v1.15.26`)
|
||||
|
||||
A 20-phase NFR sweep (no new features) themed around five user-directed
|
||||
axes: **Simplify without regressions**, **Security**, **Maintainability**,
|
||||
**User/Developer Experience**, **No Humans Onboarding Flow**. The v1.15
|
||||
rebrand left a fresh debt layer (stale brand strings, a state-bucket
|
||||
drift, a Kyverno policy contradicting the Nova tagging standard, dead
|
||||
code) that this milestone cleared, alongside genuine simplification
|
||||
(dedup helpers, a workflow generator, file splits) and the first
|
||||
self-service onboarding request path (request-path only; real AWS
|
||||
provisioning deferred, D-113).
|
||||
|
||||
**Milestone type:** NFR (all phases fix/chore/docs/refactor/test). The
|
||||
final phase's patch IS the deliverable. Tags on the v1.15.x line:
|
||||
`v1.15.5` (P0) → `v1.15.6..v1.15.25` (P1–P20) → `v1.15.26` (P21 final =
|
||||
milestone release).
|
||||
|
||||
**Regression gate (D-118, G-111):** 18 Verified + 4 Skipped (CAP-013..016
|
||||
live-AWS caps are the post-v1.11-teardown steady state, D-096; re-
|
||||
provisioning is a future feature). 0 Decayed/Broken at P9 + P21.
|
||||
|
||||
**Grill:** PASS-with-binding (G-111..G-113, E-002 deferred to P21).
|
||||
G-111: gate criterion restated 18V+4S + Skipped logic. G-112: P9 source
|
||||
model pinned. G-113: P12/P13 import direction documented.
|
||||
|
||||
**Wave outcomes:**
|
||||
- Wave 1 (P1–P4): state-bucket + Kyverno rebrand fix (correctness
|
||||
regression), user-facing ACDL→Nova sweep, dead-code cleanup, except
|
||||
narrowing.
|
||||
- Wave 2 (P5–P9): regression-verify dedup (~70 lines), run-platform
|
||||
HITL fn + config, contract-resolver envloader + registry kind, workflow
|
||||
generator (sync_workflows.py + workflows-src/), run-platform split
|
||||
(decommission + uptime helpers). Gate PASS at P9.
|
||||
- Wave 3 (P10–P14): ingestor defense-in-depth (fail closed on missing
|
||||
IAM), payload validation (size cap + schema), split contract-resolver
|
||||
(decommission + CLI modules), split regression-verify (CLI module),
|
||||
schema-driven outputs + schema cache. Mid-milestone checkpoint clean.
|
||||
- Wave 4 (P15–P17): run-platform --help + flags doc, workflows README
|
||||
catalog (7 workflows), getting-started consolidation (offline-first).
|
||||
- Wave 5 (P18–P20): onboarding schema + onboard_consumer Lambda action,
|
||||
env-file autogen (core/onboarding.py), cross-account role Terraform
|
||||
(offline-proven, D-114).
|
||||
|
||||
**Outcome:** 20 requirements (REQ-165..184) satisfied; ~630 tests pass;
|
||||
regression gate 18V+4S; the onboarding request path is self-service (no
|
||||
"contact the platform team" handoff); real AWS provisioning explicitly
|
||||
deferred (D-113/D-114).
|
||||
|
||||
Ship tag at milestone COMPLETE: `v1.15.26` (NFR milestone; final patch IS
|
||||
the release). **DONE.**
|
||||
|
||||
@@ -8,7 +8,7 @@
|
||||
],
|
||||
"active_project": "acdl",
|
||||
"active_projects": ["acdl"],
|
||||
"active_milestone": "v1.17",
|
||||
"active_milestone": "v1.16",
|
||||
"autonomy": {
|
||||
"level": "full",
|
||||
"escalation_hooks": ["deploy", "delete_data", "merge_to_main"],
|
||||
|
||||
@@ -1,50 +0,0 @@
|
||||
# GitHub Workflows — Nova Platform CI/CD Catalog
|
||||
|
||||
This directory contains the 7 GitHub Actions workflows for the Nova
|
||||
platform. 3 are byte-identical Gitea mirrors (generated from
|
||||
`workflows-src/` by `scripts/sync_workflows.py`, P8/REQ-172); 4 are
|
||||
GitHub-only (Gitea act_runner feature gaps).
|
||||
|
||||
## Shared workflows (byte-identical Gitea + GitHub)
|
||||
|
||||
These 3 are generated from `workflows-src/<name>` by
|
||||
`scripts/sync_workflows.py`; the `.gitea/workflows/<name>` mirror is kept
|
||||
byte-identical. Run `python3 scripts/sync_workflows.py --check` to verify
|
||||
no drift.
|
||||
|
||||
| Workflow | Trigger | Inputs | Required Secrets | Purpose |
|
||||
|----------|---------|--------|------------------|---------|
|
||||
| `ci.yml` | `pull_request: [main]` | — | — | Lint + test + check-only (runs on every PR) |
|
||||
| `deploy.yml` | `workflow_call` (reusable) + `push: [main]` | `contract` (string, required), `mode` (string, default `deploy`), `changeRequestId` (string), `environment` (string) | `NOVA_AWS_ACCESS_KEY_ID`, `NOVA_AWS_SECRET_ACCESS_KEY`, `NOVA_AWS_DEFAULT_REGION`, `NOVA_KMS_KEY_ID`, `NOVA_LAMBDA_URL` | Reusable deploy workflow (invoked by consumer repos via `uses: acdl/.github/workflows/deploy.yml@v1.15`) |
|
||||
| `modules-lifecycle.yml` | `pull_request: [main]` + `workflow_dispatch` | `lifecycle_mode` (string, default `plan` — `plan` or `full`) | `NOVA_AWS_ACCESS_KEY_ID`, `NOVA_AWS_SECRET_ACCESS_KEY`, `NOVA_AWS_DEFAULT_REGION`, `NOVA_AWS_ACCOUNT_ID` | L1 + L2 module lifecycle pipeline (plan-only default; full apply/modify/destroy on override) |
|
||||
|
||||
## GitHub-only workflows (no Gitea mirror)
|
||||
|
||||
These 4 have no Gitea counterpart (Gitea act_runner lacks the features
|
||||
they require — reusable workflows, matrix `needs`, release API). See
|
||||
`.gitea/workflows/README.md` for the limitation rationale.
|
||||
|
||||
| Workflow | Trigger | Inputs | Required Secrets | Purpose |
|
||||
|----------|---------|--------|------------------|---------|
|
||||
| `platform-test.yml` | `pull_request: [main]` | — | — | Lint + unit + integration + schema-validation (replaces `ci.yml` for PRs) |
|
||||
| `primitives-plan.yml` | `pull_request: [main]` | — | `NOVA_AWS_*` | Plan-only for all L1 primitives (matrix) |
|
||||
| `patterns-plan.yml` | `pull_request: [main]` | — | `NOVA_AWS_*` | Plan-only for all L2 modules (matrix) |
|
||||
| `release.yml` | `push: [main]` | — | `NOVA_GITEA_TOKEN` (for Gitea release API) | Semver tag + MAJOR.MINOR/MAJOR floating-tag maintenance + release creation on merge to main |
|
||||
|
||||
## Reusable deploy workflow (`deploy.yml`)
|
||||
|
||||
Consumer repos invoke the deploy workflow via a versioned tag:
|
||||
|
||||
```yaml
|
||||
jobs:
|
||||
deploy:
|
||||
uses: acdl/.github/workflows/deploy.yml@v1.15
|
||||
with:
|
||||
contract: .nova/contract.yml
|
||||
environment: dev
|
||||
secrets: inherit
|
||||
```
|
||||
|
||||
The workflow checks out the consumer repo + the Nova platform repo, runs
|
||||
`scripts/run_platform.sh`, and posts deploy outputs as a PR comment +
|
||||
to SSM Parameter Store.
|
||||
-12
@@ -14,18 +14,6 @@ terraform/bootstrap/.bootstrap_state.json
|
||||
# CIAgent runtime artifacts
|
||||
.ciagent/logs/
|
||||
|
||||
# Nova metrics runtime artifacts (REQ-187, D-128)
|
||||
# Generated: nova_metrics.db, decision_ledger.db, events.jsonl, runs/, test-results.xml, coverage.json, test-report.json
|
||||
# NOT ignored: metrics/README.md, metrics/powerbi/ (export views), schemas/metrics_*.schema.json
|
||||
metrics/nova_metrics.db
|
||||
metrics/decision_ledger.db
|
||||
metrics/events.jsonl
|
||||
metrics/test-results.xml
|
||||
metrics/test-report.json
|
||||
metrics/coverage.json
|
||||
metrics/runs/
|
||||
metrics/lifecycle/
|
||||
|
||||
# Terraform — recursively ignore .terraform dirs, lock files, plans, and state
|
||||
**/.terraform/
|
||||
**/.terraform.lock.hcl
|
||||
|
||||
@@ -126,46 +126,20 @@ engine-specific code. `modules/`, `schemas/`, `contracts/`,
|
||||
|
||||
## How to run
|
||||
|
||||
### Quick start (offline, no AWS required)
|
||||
### Prerequisites
|
||||
|
||||
The fastest way to verify the platform works — no AWS credentials, no
|
||||
bootstrap, no cost. See the [Consumer guide](docs/consumer-guide.md)
|
||||
for the consumer happy path (a consumer owns only a contract + app code).
|
||||
> These prerequisites are for running the **platform repo** locally. A
|
||||
> consumer does not need any of these — see the
|
||||
> [Consumer guide](docs/consumer-guide.md) for the consumer happy path.
|
||||
|
||||
```bash
|
||||
# Install test dependencies
|
||||
pip install -r requirements-test.txt
|
||||
- A platform-managed environment (see [docs/environments/](docs/environments/)).
|
||||
For local testing, `core/environments/dev.json` is provided as the sample.
|
||||
- AWS credentials for the dev environment (in `.env.secrets`, gitignored;
|
||||
see [Credentials & zero-trust](#credentials--zero-trust)).
|
||||
- `terraform` (pin `1.9.*`), `checkov` (pin `>=3.2,<4`), `python3` + `boto3`
|
||||
+ `jsonschema`.
|
||||
|
||||
# 1. Run the test suite (all offline — uses moto for DynamoDB mocking)
|
||||
python3 -m pytest tests/ -v
|
||||
|
||||
# 2. Run the platform in check-only mode (offline — contract -> resolver ->
|
||||
# adapter -> structure validation). Uses the default sample contract
|
||||
# (contracts/static-assets.yaml) + sample dev environment.
|
||||
bash scripts/run_platform.sh --check-only
|
||||
# Expected: "=== PLATFORM CHECK OK ==="
|
||||
|
||||
# 3. Run the headline E2E against the local emulating tier (emulates ECS,
|
||||
# outbox, S3 state, Lambda in-process; D-092).
|
||||
bash scripts/run_platform.sh --local
|
||||
# Expected: "=== LOCAL E2E OK ==="
|
||||
|
||||
# 4. Reproduce the full CI pipeline locally (lint -> test -> check-only)
|
||||
bash scripts/run_ci.sh
|
||||
# Expected: "=== CI PIPELINE OK ==="
|
||||
|
||||
# Show all run_platform.sh flags:
|
||||
bash scripts/run_platform.sh --help
|
||||
```
|
||||
|
||||
### Run against live AWS (requires credentials + bootstrap)
|
||||
|
||||
> Prerequisites: a platform-managed environment (see
|
||||
> [docs/environments/](docs/environments/); `core/environments/dev.json`
|
||||
> is the sample), AWS credentials for dev (in `.env.secrets`, gitignored;
|
||||
> see [Credentials & zero-trust](#credentials--zero-trust)), `terraform`
|
||||
> (pin `1.9.*`), `checkov` (pin `>=3.2,<4`), `python3` + `boto3` +
|
||||
> `jsonschema`.
|
||||
### Run the platform pipeline end-to-end
|
||||
|
||||
```bash
|
||||
# 1. Bootstrap the AWS state backend + runner IAM user (one-time, idempotent)
|
||||
@@ -194,6 +168,26 @@ bash scripts/run_platform.sh --plan-only contracts/static-assets.yaml
|
||||
bash scripts/run_platform.sh --quiet contracts/static-assets.yaml
|
||||
```
|
||||
|
||||
### Test the platform (offline, no AWS required)
|
||||
|
||||
```bash
|
||||
# Install test dependencies
|
||||
pip install -r requirements-test.txt
|
||||
|
||||
# Run the test suite (all offline — uses moto for DynamoDB mocking)
|
||||
python3 -m pytest tests/ -v
|
||||
|
||||
# Run the platform in check-only mode (offline — no AWS, no policy checks,
|
||||
# no outbox). Uses the default sample contract (contracts/static-assets.yaml)
|
||||
# and the sample dev environment (core/environments/dev.json).
|
||||
bash scripts/run_platform.sh --check-only
|
||||
# Expected: "=== PLATFORM CHECK OK ==="
|
||||
|
||||
# Reproduce the full CI pipeline locally (lint -> test -> check-only)
|
||||
bash scripts/run_ci.sh
|
||||
# Expected: "=== CI PIPELINE OK ==="
|
||||
```
|
||||
|
||||
### CI/CD pipelines
|
||||
|
||||
The CI/CD pipeline is defined by a **central pipeline contract** — a
|
||||
|
||||
@@ -17,12 +17,8 @@ ACDL_TAG_NAMING in P2 (REQ-158); the rule is in hard mode as of P3
|
||||
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))))
|
||||
from core.metrics.event_envelope import emit
|
||||
|
||||
|
||||
RULE_MAP = {
|
||||
"CKV_AWS_41": ("secrets-in-plaintext", "high"),
|
||||
@@ -75,7 +71,7 @@ def _to_pcr(checkov_record, contract_id, result_str):
|
||||
}
|
||||
|
||||
|
||||
def adapt(checkov_json_path, contract_id, run_id=None, environment="dev"):
|
||||
def adapt(checkov_json_path, contract_id):
|
||||
with open(checkov_json_path, "r", encoding="utf-8") as fh:
|
||||
data = json.load(fh)
|
||||
out = []
|
||||
@@ -89,25 +85,6 @@ def adapt(checkov_json_path, contract_id, run_id=None, environment="dev"):
|
||||
out.append(_to_pcr(rec, contract_id, "FAILED"))
|
||||
for rec in results.get("skipped_checks", []):
|
||||
out.append(_to_pcr(rec, contract_id, "SKIPPED"))
|
||||
|
||||
# Emit nova.policy.evaluated event (REQ-187).
|
||||
if run_id:
|
||||
passed = sum(1 for p in out if p["result"] == "pass")
|
||||
failed = sum(1 for p in out if p["result"] == "fail")
|
||||
skipped = sum(1 for p in out if p["result"] == "skipped")
|
||||
severity_breakdown = {}
|
||||
for p in out:
|
||||
sev = p.get("severity", "info")
|
||||
severity_breakdown[sev] = severity_breakdown.get(sev, 0) + 1
|
||||
try:
|
||||
emit("nova.policy.evaluated", run_id, environment, {
|
||||
"passed": passed, "failed": failed, "skipped": skipped,
|
||||
"severity_breakdown": severity_breakdown,
|
||||
"rule_count": len(out),
|
||||
}, contract_id=contract_id)
|
||||
except Exception:
|
||||
pass # metrics emission must never break the policy adapter
|
||||
|
||||
return out
|
||||
|
||||
|
||||
|
||||
@@ -34,13 +34,8 @@ per-input scores.
|
||||
from dataclasses import dataclass, asdict
|
||||
from typing import List, Literal, Optional, Dict, Any
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||||
from core.metrics.event_envelope import emit, make_event, append_event
|
||||
from core.metrics.decision_ledger import append as ledger_append
|
||||
|
||||
|
||||
WEIGHTS = {
|
||||
"policy": 0.30,
|
||||
@@ -166,33 +161,7 @@ def compute(contract_id: str, environment: str,
|
||||
band = "warn"
|
||||
if environment == "dev" and band == "warn":
|
||||
band = "block"
|
||||
signal = Signal(score, band, per_input, reasons)
|
||||
|
||||
# Emit nova.confidence.computed + nova.ai.decision.made events (D-122).
|
||||
# The "AI decision" is the confidence-gated policy engine, not an LLM.
|
||||
# decision_id = run_id (or "cli-<ts>" when called from CLI without a run).
|
||||
try:
|
||||
run_id = os.environ.get("NOVA_RUN_ID", f"cli-{int(__import__('time').time())}")
|
||||
conf_data = {"score": score, "band": band, "perInput": per_input, "reasonCodes": reasons}
|
||||
emit("nova.confidence.computed", run_id, environment, conf_data, contract_id=contract_id)
|
||||
|
||||
decision_data = {
|
||||
"decision_id": run_id,
|
||||
"chosen_action": band,
|
||||
"confidence": score,
|
||||
"alternatives": per_input,
|
||||
"human_override": band == "block",
|
||||
"threshold": THRESHOLDS[environment],
|
||||
}
|
||||
decision_event = make_event("nova.ai.decision.made", run_id, environment, decision_data,
|
||||
contract_id=contract_id, actor_type="confidence-gate",
|
||||
actor_id="confidence_signal")
|
||||
append_event(decision_event)
|
||||
ledger_append(decision_event)
|
||||
except Exception:
|
||||
pass # metrics emission must never break the confidence gate
|
||||
|
||||
return signal
|
||||
return Signal(score, band, per_input, reasons)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
+37
-23
@@ -64,21 +64,6 @@ def _load_json(path):
|
||||
return json.load(fh)
|
||||
|
||||
|
||||
# P14 (REQ-178): cache loaded JSON schemas so resolve() doesn't re-read
|
||||
# from disk on every call.
|
||||
_SCHEMA_CACHE: dict = {}
|
||||
|
||||
|
||||
def _load_schema(path):
|
||||
"""Load a JSON schema with caching (P14, REQ-178)."""
|
||||
cached = _SCHEMA_CACHE.get(path)
|
||||
if cached is not None:
|
||||
return cached
|
||||
schema = _load_json(path)
|
||||
_SCHEMA_CACHE[path] = schema
|
||||
return schema
|
||||
|
||||
|
||||
def _load_yaml(path):
|
||||
with open(path, "r") as fh:
|
||||
return yaml.safe_load(fh)
|
||||
@@ -452,9 +437,24 @@ def _namespace_resources(resources, module_name):
|
||||
|
||||
|
||||
def decommission_transform(stack_instance):
|
||||
"""REQ-92: re-export from core.decommission_transform (P12, REQ-176)."""
|
||||
from core.decommission_transform import decommission_transform as _dt
|
||||
return _dt(stack_instance)
|
||||
"""REQ-92: Transform a resolved stack instance for decommission.
|
||||
|
||||
Sets all scalable counts to 0 and deletion_protection to false on
|
||||
every resource. Used by the decommission pipeline mode after the
|
||||
first step (disable deletion protection) has been applied.
|
||||
"""
|
||||
for res in stack_instance.get("resources", []):
|
||||
if "nfrs" not in res:
|
||||
res["nfrs"] = {}
|
||||
res["nfrs"]["deletion_protection"] = False
|
||||
inputs = res.get("inputs", {})
|
||||
if "desired_count" in inputs:
|
||||
inputs["desired_count"] = 0
|
||||
if "min_capacity" in inputs:
|
||||
inputs["min_capacity"] = 0
|
||||
if "max_capacity" in inputs:
|
||||
inputs["max_capacity"] = 0
|
||||
return stack_instance
|
||||
|
||||
|
||||
def resolve(contract_path, repo_root=None, environment_override=None):
|
||||
@@ -483,7 +483,7 @@ def resolve(contract_path, repo_root=None, environment_override=None):
|
||||
contract["environment"] = environment_override
|
||||
|
||||
# Load schemas
|
||||
contract_schema = _load_schema(os.path.join(repo_root, "schemas", "contract.schema.json"))
|
||||
contract_schema = _load_json(os.path.join(repo_root, "schemas", "contract.schema.json"))
|
||||
|
||||
# Validate contract against schema
|
||||
jsonschema.validate(contract, contract_schema)
|
||||
@@ -603,13 +603,27 @@ def resolve(contract_path, repo_root=None, environment_override=None):
|
||||
stack_instance["outputs"] = merged_outputs
|
||||
|
||||
# Validate against stack schema
|
||||
stack_schema = _load_schema(os.path.join(repo_root, "schemas", "stack.schema.json"))
|
||||
stack_schema = _load_json(os.path.join(repo_root, "schemas", "stack.schema.json"))
|
||||
jsonschema.validate(stack_instance, stack_schema)
|
||||
|
||||
return stack_instance
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
# P12 (REQ-176): CLI extracted to core/contract_resolver_cli.py.
|
||||
from core.contract_resolver_cli import main
|
||||
sys.exit(main())
|
||||
if len(sys.argv) < 3:
|
||||
print("usage: contract_resolver.py <contract.yml> <out.json> [--environment <name>]", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
contract_path = sys.argv[1]
|
||||
out_path = sys.argv[2]
|
||||
env_override = None
|
||||
if "--environment" in sys.argv:
|
||||
idx = sys.argv.index("--environment")
|
||||
if idx + 1 < len(sys.argv):
|
||||
env_override = sys.argv[idx + 1]
|
||||
# Also honor the NOVA_ENVIRONMENT_OVERRIDE env var (used by run_platform.sh).
|
||||
# Dual-read via core/env.py: NOVA_* preferred, ACDL_* fallback until P5.
|
||||
if env_override is None and env.get_env("ENVIRONMENT_OVERRIDE"):
|
||||
env_override = env.get_env("ENVIRONMENT_OVERRIDE")
|
||||
result = resolve(contract_path, environment_override=env_override)
|
||||
with open(out_path, "w") as fh:
|
||||
json.dump(result, fh, indent=2)
|
||||
@@ -1,41 +0,0 @@
|
||||
"""Nova Contract Resolver CLI — command-line entry point.
|
||||
|
||||
Extracted from core/contract_resolver.py (P12, REQ-176).
|
||||
|
||||
G-113 import direction: this module imports core.contract_resolver (the
|
||||
re-export shim) for the resolve function. The shim imports the split
|
||||
modules. Nothing imports this CLI module except direct invocation.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import sys
|
||||
|
||||
from core.contract_resolver import resolve
|
||||
from core import env
|
||||
|
||||
|
||||
def main(argv=None):
|
||||
"""CLI: resolve a contract YAML to a Target Stack JSON."""
|
||||
argv = argv if argv is not None else sys.argv[1:]
|
||||
if len(argv) < 2:
|
||||
print("usage: contract_resolver.py <contract.yml> <out.json> [--environment <name>", file=sys.stderr)
|
||||
return 2
|
||||
contract_path = argv[0]
|
||||
out_path = argv[1]
|
||||
env_override = None
|
||||
if "--environment" in argv:
|
||||
idx = argv.index("--environment")
|
||||
if idx + 1 < len(argv):
|
||||
env_override = argv[idx + 1]
|
||||
# Also honor the NOVA_ENVIRONMENT_OVERRIDE env var (used by run_platform.sh).
|
||||
if env_override is None and env.get_env("ENVIRONMENT_OVERRIDE"):
|
||||
env_override = env.get_env("ENVIRONMENT_OVERRIDE")
|
||||
result = resolve(contract_path, environment_override=env_override)
|
||||
with open(out_path, "w") as fh:
|
||||
json.dump(result, fh, indent=2)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -1,31 +0,0 @@
|
||||
"""Nova Decommission Transform — zero counts + disable deletion protection (REQ-92).
|
||||
|
||||
Extracted from core/contract_resolver.py (P12, REQ-176).
|
||||
|
||||
G-113 import direction: this module imports only stdlib. The re-export
|
||||
shim core/contract_resolver.py imports this module. Nothing imports the
|
||||
shim except external callers.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
|
||||
def decommission_transform(stack_instance):
|
||||
"""REQ-92: Transform a resolved stack instance for decommission.
|
||||
|
||||
Sets all scalable counts to 0 and deletion_protection to false on
|
||||
every resource. Used by the decommission pipeline mode after the
|
||||
first step (disable deletion protection) has been applied.
|
||||
"""
|
||||
for res in stack_instance.get("resources", []):
|
||||
if "nfrs" not in res:
|
||||
res["nfrs"] = {}
|
||||
res["nfrs"]["deletion_protection"] = False
|
||||
inputs = res.get("inputs", {})
|
||||
if "desired_count" in inputs:
|
||||
inputs["desired_count"] = 0
|
||||
if "min_capacity" in inputs:
|
||||
inputs["min_capacity"] = 0
|
||||
if "max_capacity" in inputs:
|
||||
inputs["max_capacity"] = 0
|
||||
return stack_instance
|
||||
@@ -55,8 +55,6 @@ def load(env_name, root=None):
|
||||
|
||||
|
||||
def _onboarding_message(env_name):
|
||||
# P19 (REQ-183): rebranded Nova self-service request path — no longer
|
||||
# routes to "contact the platform team" for the request step.
|
||||
return (
|
||||
"=== Nova Environment Onboarding ===\n"
|
||||
f"No environment named '{env_name}' is bound to this repository.\n\n"
|
||||
@@ -68,15 +66,13 @@ def _onboarding_message(env_name):
|
||||
" - an IAM role surfaced to your repo via attribute-based\n"
|
||||
" authorization (ABAC)\n\n"
|
||||
"You do not provide an AWS account, VPC, subnet, or state bucket.\n\n"
|
||||
"To request an environment (self-service):\n"
|
||||
" 1. Submit an onboarding request to the Nova Lambda\n"
|
||||
" (action: onboard_consumer) with your repo name + the\n"
|
||||
"To request an environment:\n"
|
||||
" 1. Contact the platform team with your repo name + the\n"
|
||||
" environment name you need (e.g. 'dev').\n"
|
||||
" 2. The platform generates an environment binding + opens a PR.\n"
|
||||
" 3. The platform provisions the account/network/state/role and\n"
|
||||
" grants the ABAC role. Your next pipeline run proceeds.\n\n"
|
||||
"Run: python3 core/onboarding.py --request '{...}' to generate a\n"
|
||||
"binding file locally, or POST to the Lambda onboard_consumer action.\n"
|
||||
" 2. The platform team provisions the account/network/state/role\n"
|
||||
" and binds the environment to your repo.\n"
|
||||
" 3. Your next pipeline run will proceed normally.\n\n"
|
||||
"Expected turnaround: contact the platform team for current SLA.\n"
|
||||
"===================================\n"
|
||||
)
|
||||
|
||||
|
||||
@@ -33,13 +33,5 @@ halting the pipeline before any work is done.
|
||||
|
||||
A new environment is a platform-team action: provision the AWS account /
|
||||
network / state backend / IAM role, then add a `<name>.json` here and bind
|
||||
it to the consumer repo.
|
||||
|
||||
**P19 (REQ-183):** the *request* step is now self-service. A consumer
|
||||
submits an onboarding request (POST to the Nova Lambda `onboard_consumer`
|
||||
action, or `python3 core/onboarding.py --request '{...}'`) and the
|
||||
platform generates a `<name>.json` binding file from the request + opens
|
||||
a PR. The actual AWS account/network/state provisioning + cross-account
|
||||
role grant remains a platform-team action (a future feature milestone
|
||||
will automate the provisioning; the cross-account role Terraform is
|
||||
offline-proven in P20/REQ-184).
|
||||
it to the consumer repo. Self-service environment provisioning is on the
|
||||
roadmap; today it is a platform-team action.
|
||||
@@ -12,10 +12,6 @@ import os
|
||||
import sys
|
||||
from typing import Optional, Tuple
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||||
from core.metrics.event_envelope import make_event, append_event
|
||||
from core.metrics.decision_ledger import append as ledger_append
|
||||
|
||||
|
||||
def _approver_attr(env: str) -> str:
|
||||
return {"qa": "approver_qa", "prod": "approver_prod", "dr": "approver_dr"}.get(env, "")
|
||||
@@ -65,24 +61,6 @@ def attest(contract_id: str, env: str, approver: str,
|
||||
if not ok:
|
||||
return (False, reason)
|
||||
|
||||
# Emit attestation.recorded event to the Decision Ledger (D-132).
|
||||
try:
|
||||
run_id = os.environ.get("NOVA_RUN_ID", f"attest-{contract_id[:8]}")
|
||||
attestation_data = {
|
||||
"approver": approver,
|
||||
"environment": env,
|
||||
"concerns": reason,
|
||||
"result": "pass",
|
||||
"contract_id": contract_id,
|
||||
}
|
||||
attestation_event = make_event("nova.attestation.recorded", run_id, env, attestation_data,
|
||||
contract_id=contract_id, actor_type="human-attestation",
|
||||
actor_id=approver)
|
||||
append_event(attestation_event)
|
||||
ledger_append(attestation_event)
|
||||
except Exception:
|
||||
pass # metrics emission must never break the attestation gate
|
||||
|
||||
return (True, f"{env} attested by {approver}")
|
||||
|
||||
|
||||
|
||||
@@ -30,53 +30,10 @@ PLATFORM_REPO = os.environ.get("PLATFORM_REPO", "nova/acdl")
|
||||
# to a Gitea API root (e.g. https://git.cloudinit.dev/api/v1) for Gitea.
|
||||
GITHUB_API_BASE = os.environ.get("GITHUB_API_BASE", "https://api.github.com")
|
||||
|
||||
# P11 (REQ-175): consistent cap for error/stackTrace fields (was 10k vs 2k).
|
||||
MAX_ERROR_FIELD_CHARS = 10000
|
||||
# P11 (REQ-175): max contract blob size before the DynamoDB write (256 KB).
|
||||
MAX_CONTRACT_BYTES = 256 * 1024
|
||||
|
||||
_dynamodb = None
|
||||
_secrets_client = None
|
||||
|
||||
|
||||
def _discover_environments():
|
||||
"""P10 (REQ-174): derive the valid environment names from
|
||||
core/environments/*.json (the directory is the single source of truth,
|
||||
not a hardcoded set). Falls back to {'dev','qa','prod','dr'} if the
|
||||
directory is not readable (e.g. packaged Lambda without the dir).
|
||||
"""
|
||||
env_dir = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(
|
||||
os.path.abspath(__file__)))), "core", "environments")
|
||||
try:
|
||||
names = {f[:-5] for f in os.listdir(env_dir) if f.endswith(".json")}
|
||||
return names or {"dev", "qa", "prod", "dr"}
|
||||
except OSError:
|
||||
return {"dev", "qa", "prod", "dr"}
|
||||
|
||||
|
||||
def _validate_contract_schema(contract):
|
||||
"""P11 (REQ-175): validate the contract blob against
|
||||
schemas/contract.schema.json before the DynamoDB write. Raises
|
||||
ValueError on invalid. Falls back to a no-op if the schema or
|
||||
jsonschema is unavailable (e.g. packaged Lambda without the schema).
|
||||
"""
|
||||
try:
|
||||
import json as _json
|
||||
import jsonschema
|
||||
schema_path = os.path.join(os.path.dirname(os.path.dirname(
|
||||
os.path.dirname(os.path.abspath(__file__)))),
|
||||
"schemas", "contract.schema.json")
|
||||
with open(schema_path) as f:
|
||||
schema = _json.load(f)
|
||||
jsonschema.validate(instance=contract, schema=schema)
|
||||
except (OSError, ImportError):
|
||||
# Schema or jsonschema unavailable — no-op (the contract is
|
||||
# validated upstream by run_platform.sh in the normal path).
|
||||
pass
|
||||
except jsonschema.ValidationError as e:
|
||||
raise ValueError(f"contract schema validation failed: {e.message}")
|
||||
|
||||
|
||||
def _get_dynamodb():
|
||||
global _dynamodb
|
||||
if _dynamodb is None:
|
||||
@@ -137,25 +94,6 @@ def _submit_contract(payload):
|
||||
contract_id = payload["contractId"]
|
||||
contract = payload["contract"]
|
||||
environment = payload["environment"]
|
||||
|
||||
# P11 (REQ-175): size-cap the contract blob before the DynamoDB write
|
||||
# (unbounded payload → write amplification). 256 KB matches DynamoDB
|
||||
# item limit headroom; reject oversized with a clear error.
|
||||
import json as _json
|
||||
contract_json = _json.dumps(contract).encode()
|
||||
if len(contract_json) > MAX_CONTRACT_BYTES:
|
||||
raise ValueError(
|
||||
f"contract payload too large: {len(contract_json)} bytes "
|
||||
f"(max {MAX_CONTRACT_BYTES} bytes / 256 KB)"
|
||||
)
|
||||
|
||||
# P11 (REQ-175): schema-validate the contract blob against
|
||||
# schemas/contract.schema.json before the write. Reject invalid with 400.
|
||||
# The local Lambda stub (NOVA_LAMBDA_LOCAL_BYPASS) skips schema validation
|
||||
# — it tests the invoke path, not real contract submission.
|
||||
if not os.environ.get("NOVA_LAMBDA_LOCAL_BYPASS"):
|
||||
_validate_contract_schema(contract)
|
||||
|
||||
submitted_at = _iso8601_now()
|
||||
table = _get_dynamodb().Table(TABLE_NAME)
|
||||
item = {
|
||||
@@ -193,7 +131,7 @@ def _report_error(payload):
|
||||
contract_id = payload["contractId"]
|
||||
error = payload.get("error", "unknown error")
|
||||
run_url = payload.get("runUrl", "")
|
||||
stack_trace = payload.get("stackTrace", "")[:MAX_ERROR_FIELD_CHARS] # P11: aligned cap
|
||||
stack_trace = payload.get("stackTrace", "")[:2000] # truncate
|
||||
|
||||
# Get the GitHub token from Secrets Manager
|
||||
secrets = _get_secrets_client()
|
||||
@@ -298,33 +236,20 @@ def _validate_caller_identity(event, payload):
|
||||
in the payload matches the principal's ARN-derived source identity, preventing
|
||||
one consumer from impersonating another.
|
||||
|
||||
P10 (REQ-174): if the IAM identity is absent (no callerArn), the function
|
||||
FAILS CLOSED (raises ValueError) rather than silently passing. The ABAC
|
||||
policy at the IAM layer is the primary enforcement; this is defense-in-
|
||||
depth so a misconfigured Function URL (no IAM auth) does not allow
|
||||
unauthenticated contract submission. Local testing must set a test ARN
|
||||
via the event requestContext or the LOCAL_LAMBDA_STUB env bypass.
|
||||
If the identity is not available (e.g. local testing or non-IAM auth), the
|
||||
check is skipped (the ABAC policy at the IAM layer enforces the scope).
|
||||
|
||||
v1.14 (REQ-144): also validates contractId format, environment enum, and
|
||||
error length. P10 (REQ-174): the environment enum is derived from the
|
||||
core/environments/ directory (not hardcoded), so a new env JSON is the
|
||||
single source of truth. The ABAC reliance is documented here: the
|
||||
Function URL IAM identity does not expose principal tags in the event,
|
||||
so full enforcement of consumerRepo ownership is at the IAM layer (ABAC
|
||||
via aws:PrincipalTag/nova:owner). This function validates format only,
|
||||
not ownership.
|
||||
error length. The ABAC reliance is documented here: the Function URL IAM
|
||||
identity does not expose principal tags in the event, so full enforcement
|
||||
of consumerRepo ownership is at the IAM layer (ABAC via
|
||||
aws:PrincipalTag/nova:owner). This function validates format only, not
|
||||
ownership.
|
||||
"""
|
||||
identity = event.get("requestContext", {}).get("identity", {})
|
||||
caller_arn = identity.get("userArn", "")
|
||||
if not caller_arn:
|
||||
# P10 (REQ-174): fail closed. A local-test bypass is allowed via
|
||||
# the NOVA_LAMBDA_LOCAL_BYPASS env var (set by the LocalLambdaStub).
|
||||
import os as _os
|
||||
if not _os.environ.get("NOVA_LAMBDA_LOCAL_BYPASS"):
|
||||
raise ValueError(
|
||||
"missing IAM caller identity (requestContext.identity.userArn) — "
|
||||
"the Function URL must use IAM auth; refusing unauthenticated submission"
|
||||
)
|
||||
pass # no identity available — rely on IAM ABAC enforcement
|
||||
payload_repo = payload.get("consumerRepo", "")
|
||||
if payload_repo:
|
||||
# consumerRepo must be org/repo format, <=128 chars
|
||||
@@ -338,18 +263,17 @@ def _validate_caller_identity(event, payload):
|
||||
if not re.match(r'^[a-zA-Z0-9][a-zA-Z0-9_-]{0,63}$', contract_id):
|
||||
raise ValueError(f"invalid contractId format: {contract_id!r} (alphanumeric, hyphen, underscore; max 64 chars)")
|
||||
|
||||
# P10 (REQ-174): environment enum derived from core/environments/ (not
|
||||
# hardcoded) — the directory is the single source of truth.
|
||||
# v1.14 (REQ-144): environment enum validation
|
||||
environment = payload.get("environment", "")
|
||||
if environment:
|
||||
valid_envs = _discover_environments()
|
||||
valid_envs = {"dev", "qa", "prod", "dr"}
|
||||
if environment not in valid_envs:
|
||||
raise ValueError(f"invalid environment: {environment!r} (must be one of {sorted(valid_envs)})")
|
||||
raise ValueError(f"invalid environment: {environment!r} (must be one of {valid_envs})")
|
||||
|
||||
# v1.14 (REQ-144): error length cap (for report_error action)
|
||||
error_msg = payload.get("error", "")
|
||||
if error_msg and len(str(error_msg)) > MAX_ERROR_FIELD_CHARS:
|
||||
payload["error"] = str(error_msg)[:MAX_ERROR_FIELD_CHARS]
|
||||
if error_msg and len(str(error_msg)) > 10000:
|
||||
payload["error"] = str(error_msg)[:10000]
|
||||
|
||||
|
||||
def _validate_change_request(payload):
|
||||
@@ -398,65 +322,6 @@ def _validate_change_request(payload):
|
||||
}
|
||||
|
||||
|
||||
def _onboard_consumer(payload):
|
||||
"""P18 (REQ-182): accept a self-service onboarding request.
|
||||
|
||||
Validates the payload against schemas/onboarding.schema.json, then
|
||||
writes a 'pending' row to nova-contracts (D-119). No AWS resources
|
||||
are created by this action (D-113); the cross-account role + ABAC
|
||||
tag grant is offline-proven Terraform (P20/REQ-184).
|
||||
"""
|
||||
import jsonschema
|
||||
schema_path = os.path.join(os.path.dirname(os.path.dirname(
|
||||
os.path.dirname(os.path.abspath(__file__)))),
|
||||
"schemas", "onboarding.schema.json")
|
||||
try:
|
||||
with open(schema_path) as f:
|
||||
schema = json.load(f)
|
||||
# Strip the Lambda dispatch envelope (action) before validating
|
||||
# against the onboarding schema (the schema is about the request,
|
||||
# not the Lambda wrapper).
|
||||
onboarding_payload = {k: v for k, v in payload.items() if k != "action"}
|
||||
jsonschema.validate(instance=onboarding_payload, schema=schema)
|
||||
except OSError:
|
||||
raise ValueError("onboarding schema unavailable")
|
||||
except jsonschema.ValidationError as e:
|
||||
raise ValueError(f"onboarding payload invalid: {e.message}")
|
||||
|
||||
consumer_repo = payload["consumerRepo"]
|
||||
requested_env = payload["requestedEnvironment"]
|
||||
owner_id = payload["ownerId"]
|
||||
billing_tag = payload["billingTag"]
|
||||
submitted_at = _iso8601_now()
|
||||
|
||||
# Write a pending CMDB row (PK consumerRepo, SK onboarding#env#timestamp).
|
||||
table = _get_dynamodb().Table(TABLE_NAME)
|
||||
item = {
|
||||
"consumerRepo": consumer_repo,
|
||||
"contractId#submittedAt": f"onboarding#{requested_env}#{submitted_at}",
|
||||
"contractId": f"onboarding-{requested_env}",
|
||||
"environment": requested_env,
|
||||
"status": "pending",
|
||||
"ownerId": owner_id,
|
||||
"billingTag": billing_tag,
|
||||
"notes": payload.get("notes", ""),
|
||||
"submittedAt": submitted_at,
|
||||
}
|
||||
table.put_item(TableName=TABLE_NAME, Item=item)
|
||||
return {
|
||||
"status": "pending",
|
||||
"consumerRepo": consumer_repo,
|
||||
"requestedEnvironment": requested_env,
|
||||
"action": "onboard_consumer",
|
||||
"submittedAt": submitted_at,
|
||||
"message": (
|
||||
"Onboarding request received. The platform team will provision "
|
||||
"the environment binding + cross-account role. Track the status "
|
||||
"via the nova-contracts table (status=pending → granted)."
|
||||
),
|
||||
}
|
||||
|
||||
|
||||
def lambda_handler(event, context):
|
||||
"""AWS Lambda handler entry point.
|
||||
|
||||
@@ -485,8 +350,6 @@ def lambda_handler(event, context):
|
||||
result = _report_error(payload)
|
||||
elif action == "validate_change_request":
|
||||
result = _validate_change_request(payload)
|
||||
elif action == "onboard_consumer":
|
||||
result = _onboard_consumer(payload)
|
||||
else:
|
||||
return {
|
||||
"statusCode": 400,
|
||||
@@ -494,9 +357,6 @@ def lambda_handler(event, context):
|
||||
}
|
||||
return {"statusCode": 200, "body": json.dumps(result)}
|
||||
except ValueError as e:
|
||||
# P10 (REQ-174): identity failures are 401, field validation is 400.
|
||||
if "missing IAM caller identity" in str(e):
|
||||
return {"statusCode": 401, "body": json.dumps({"error": str(e)})}
|
||||
return {"statusCode": 400, "body": json.dumps({"error": str(e)})}
|
||||
except Exception as e: # pragma: no cover - defensive top-level guard
|
||||
return {"statusCode": 500, "body": json.dumps({"error": str(e)})}
|
||||
@@ -392,24 +392,12 @@ class LocalLambdaStub:
|
||||
"httpContext": {"authorizer": {"iam": {"userId": "local-stub"}}}
|
||||
},
|
||||
}
|
||||
# P10 (REQ-174): the local stub has no real IAM identity; set
|
||||
# the bypass so the fail-closed identity check passes for local
|
||||
# tier testing. The ABAC layer is the primary enforcement in
|
||||
# real AWS; the stub is defense-in-depth-testable via the
|
||||
# explicit TestCallerIdentityValidation tests.
|
||||
import os as _os
|
||||
_prev_bypass = _os.environ.get("NOVA_LAMBDA_LOCAL_BYPASS")
|
||||
_os.environ["NOVA_LAMBDA_LOCAL_BYPASS"] = "1"
|
||||
result = ci.lambda_handler(event, None)
|
||||
finally:
|
||||
ci._get_dynamodb = original_get
|
||||
if original_urlopen is not None:
|
||||
import urllib.request
|
||||
urllib.request.urlopen = original_urlopen
|
||||
if _prev_bypass is None:
|
||||
_os.environ.pop("NOVA_LAMBDA_LOCAL_BYPASS", None)
|
||||
else:
|
||||
_os.environ["NOVA_LAMBDA_LOCAL_BYPASS"] = _prev_bypass
|
||||
return result
|
||||
|
||||
|
||||
|
||||
@@ -1,364 +0,0 @@
|
||||
"""Nova Metrics Collector (REQ-189, P2).
|
||||
|
||||
Reads all grounded signals (REGRESSION_REPORT.json, per-run manifests,
|
||||
junit XML, pcr.json, signal.json, COST.md, decision ledger, coverage.json)
|
||||
and normalizes them into a SQLite cold store at metrics/nova_metrics.db.
|
||||
|
||||
D-120: Nova-native (SQLite, no ClickHouse/BigQuery).
|
||||
D-125: hybrid model — reads files + events → SQLite.
|
||||
D-126: cold-only (no hot path; hot path deferred D-096).
|
||||
D-128: metrics/ at repo root.
|
||||
|
||||
Idempotent: re-running the collector against the same inputs produces
|
||||
identical row counts (REQ-200). The collector uses INSERT OR REPLACE
|
||||
on fact tables keyed by natural keys.
|
||||
"""
|
||||
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import sys
|
||||
import xml.etree.ElementTree as ET
|
||||
|
||||
_METRICS_DIR = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), "metrics")
|
||||
_STORE_PATH = os.path.join(_METRICS_DIR, "nova_metrics.db")
|
||||
_REPO_ROOT = os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
|
||||
_REGRESSION_REPORT = os.path.join(_REPO_ROOT, ".ciagent", "REGRESSION_REPORT.json")
|
||||
_RUNS_DIR = os.path.join(_METRICS_DIR, "runs")
|
||||
_LEDGER_DB = os.path.join(_METRICS_DIR, "decision_ledger.db")
|
||||
_COVERAGE_JSON = os.path.join(_METRICS_DIR, "coverage.json")
|
||||
_TEST_RESULTS_XML = os.path.join(_METRICS_DIR, "test-results.xml")
|
||||
|
||||
|
||||
def _iso8601_now():
|
||||
return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
||||
|
||||
|
||||
def _init_store(db_path=None):
|
||||
"""Create the fact/dim tables in the SQLite cold store."""
|
||||
if db_path is None:
|
||||
db_path = _STORE_PATH
|
||||
os.makedirs(os.path.dirname(db_path), exist_ok=True)
|
||||
conn = sqlite3.connect(db_path)
|
||||
conn.executescript("""
|
||||
CREATE TABLE IF NOT EXISTS fact_run (
|
||||
run_id TEXT PRIMARY KEY,
|
||||
contract_id TEXT,
|
||||
environment TEXT,
|
||||
started_at TEXT,
|
||||
completed_at TEXT,
|
||||
exit_code INTEGER,
|
||||
outcome TEXT,
|
||||
confidence_score REAL,
|
||||
confidence_band TEXT,
|
||||
hitl_block INTEGER,
|
||||
cost_estimate_usd REAL,
|
||||
decision_id TEXT
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS fact_capability (
|
||||
capability_id TEXT,
|
||||
run_id TEXT,
|
||||
name TEXT,
|
||||
status TEXT,
|
||||
tier TEXT,
|
||||
duration_ms REAL,
|
||||
detail TEXT,
|
||||
run_at_utc TEXT,
|
||||
PRIMARY KEY (capability_id, run_id)
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS fact_policy_check (
|
||||
run_id TEXT,
|
||||
rule_id TEXT,
|
||||
severity TEXT,
|
||||
result TEXT,
|
||||
resource_ref TEXT,
|
||||
evaluated_at TEXT,
|
||||
PRIMARY KEY (run_id, rule_id, resource_ref)
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS fact_confidence (
|
||||
run_id TEXT,
|
||||
score REAL,
|
||||
band TEXT,
|
||||
per_input TEXT,
|
||||
reason_codes TEXT,
|
||||
environment TEXT,
|
||||
computed_at TEXT,
|
||||
PRIMARY KEY (run_id)
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS fact_test (
|
||||
run_id TEXT,
|
||||
total_tests INTEGER,
|
||||
passed INTEGER,
|
||||
failed INTEGER,
|
||||
errors INTEGER,
|
||||
skipped INTEGER,
|
||||
duration_s REAL,
|
||||
coverage_pct REAL,
|
||||
collected_at TEXT,
|
||||
PRIMARY KEY (run_id)
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS fact_decision (
|
||||
decision_id TEXT,
|
||||
run_id TEXT,
|
||||
chosen_action TEXT,
|
||||
confidence REAL,
|
||||
alternatives TEXT,
|
||||
human_override INTEGER,
|
||||
outcome TEXT,
|
||||
event_time TEXT,
|
||||
PRIMARY KEY (decision_id)
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS fact_cost_estimate (
|
||||
run_id TEXT,
|
||||
delta_usd REAL,
|
||||
total_monthly_usd REAL,
|
||||
available INTEGER,
|
||||
estimated_at TEXT,
|
||||
PRIMARY KEY (run_id)
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS fact_lifecycle (
|
||||
module TEXT,
|
||||
environment TEXT,
|
||||
phase TEXT,
|
||||
result TEXT,
|
||||
duration_ms REAL,
|
||||
run_at TEXT,
|
||||
PRIMARY KEY (module, environment, phase, run_at)
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS dim_capability (
|
||||
capability_id TEXT PRIMARY KEY,
|
||||
name TEXT,
|
||||
tier TEXT,
|
||||
source_milestone TEXT
|
||||
);
|
||||
|
||||
CREATE TABLE IF NOT EXISTS dim_milestone (
|
||||
milestone TEXT PRIMARY KEY,
|
||||
phase INTEGER,
|
||||
tag TEXT,
|
||||
completed_at TEXT
|
||||
);
|
||||
""")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
|
||||
|
||||
def collect_regression_report(db_path=None, report_path=None):
|
||||
"""Read REGRESSION_REPORT.json → fact_capability + dim_capability."""
|
||||
if db_path is None:
|
||||
db_path = _STORE_PATH
|
||||
if report_path is None:
|
||||
report_path = _REGRESSION_REPORT
|
||||
if not os.path.isfile(report_path):
|
||||
return 0
|
||||
_init_store(db_path)
|
||||
with open(report_path) as f:
|
||||
report = json.load(f)
|
||||
run_id = report.get("run_id", f"regr-{report.get('run_at_utc','')}")
|
||||
run_at = report.get("run_at_utc", _iso8601_now())
|
||||
milestone = report.get("milestone", "")
|
||||
conn = sqlite3.connect(db_path)
|
||||
for result in report.get("results", []):
|
||||
cap_id = result.get("capability_id", "")
|
||||
conn.execute("""
|
||||
INSERT OR REPLACE INTO fact_capability
|
||||
(capability_id, run_id, name, status, tier, duration_ms, detail, run_at_utc)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?)
|
||||
""", (cap_id, run_id, result.get("name", ""), result.get("status", ""),
|
||||
result.get("tier", ""), result.get("duration_ms", 0),
|
||||
result.get("detail", ""), run_at))
|
||||
conn.execute("""
|
||||
INSERT OR REPLACE INTO dim_capability
|
||||
(capability_id, name, tier, source_milestone)
|
||||
VALUES (?, ?, ?, ?)
|
||||
""", (cap_id, result.get("name", ""), result.get("tier", ""), milestone))
|
||||
conn.execute("""
|
||||
INSERT OR REPLACE INTO dim_milestone
|
||||
(milestone, phase, tag, completed_at)
|
||||
VALUES (?, ?, ?, ?)
|
||||
""", (milestone, report.get("phase", 0), "", run_at))
|
||||
conn.commit()
|
||||
conn.close()
|
||||
return len(report.get("results", []))
|
||||
|
||||
|
||||
def collect_run_manifests(db_path=None, runs_dir=None):
|
||||
"""Read per-run manifests from metrics/runs/*.json → fact_run."""
|
||||
if db_path is None:
|
||||
db_path = _STORE_PATH
|
||||
if runs_dir is None:
|
||||
runs_dir = _RUNS_DIR
|
||||
if not os.path.isdir(runs_dir):
|
||||
return 0
|
||||
_init_store(db_path)
|
||||
count = 0
|
||||
conn = sqlite3.connect(db_path)
|
||||
for fname in sorted(os.listdir(runs_dir)):
|
||||
if not fname.endswith(".json"):
|
||||
continue
|
||||
fpath = os.path.join(runs_dir, fname)
|
||||
if os.path.isdir(fpath):
|
||||
continue
|
||||
with open(fpath) as f:
|
||||
manifest = json.load(f)
|
||||
run_id = manifest.get("run_id", fname.replace(".json", ""))
|
||||
conf = manifest.get("confidence", {})
|
||||
hitl = manifest.get("hitl", {})
|
||||
conn.execute("""
|
||||
INSERT OR REPLACE INTO fact_run
|
||||
(run_id, contract_id, environment, started_at, completed_at,
|
||||
exit_code, outcome, confidence_score, confidence_band,
|
||||
hitl_block, cost_estimate_usd, decision_id)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
""", (run_id, manifest.get("contract_id", ""), manifest.get("environment", ""),
|
||||
manifest.get("started_at", ""), manifest.get("completed_at", ""),
|
||||
manifest.get("exit_code", 0), manifest.get("outcome", ""),
|
||||
conf.get("score", 0), conf.get("band", ""),
|
||||
1 if hitl.get("block") else 0,
|
||||
manifest.get("cost_estimate_usd", 0), manifest.get("decision_id", "")))
|
||||
count += 1
|
||||
conn.commit()
|
||||
conn.close()
|
||||
return count
|
||||
|
||||
|
||||
def collect_decision_ledger(db_path=None, ledger_db=None):
|
||||
"""Read the Decision Ledger SQLite → fact_decision."""
|
||||
if db_path is None:
|
||||
db_path = _STORE_PATH
|
||||
if ledger_db is None:
|
||||
ledger_db = _LEDGER_DB
|
||||
if not os.path.isfile(ledger_db):
|
||||
return 0
|
||||
_init_store(db_path)
|
||||
ledger_conn = sqlite3.connect(ledger_db)
|
||||
rows = ledger_conn.execute(
|
||||
"SELECT event_type, run_id, event_time, payload FROM decision_ledger WHERE event_type = 'nova.ai.decision.made' ORDER BY seq"
|
||||
).fetchall()
|
||||
ledger_conn.close()
|
||||
conn = sqlite3.connect(db_path)
|
||||
count = 0
|
||||
for etype, run_id, event_time, payload_json in rows:
|
||||
payload = json.loads(payload_json)
|
||||
data = payload.get("data", {})
|
||||
decision_id = data.get("decision_id", run_id)
|
||||
conn.execute("""
|
||||
INSERT OR REPLACE INTO fact_decision
|
||||
(decision_id, run_id, chosen_action, confidence, alternatives,
|
||||
human_override, outcome, event_time)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?)
|
||||
""", (decision_id, run_id, data.get("chosen_action", ""),
|
||||
data.get("confidence", 0), json.dumps(data.get("alternatives", {})),
|
||||
1 if data.get("human_override") else 0,
|
||||
data.get("outcome", "pending"), event_time))
|
||||
count += 1
|
||||
conn.commit()
|
||||
conn.close()
|
||||
return count
|
||||
|
||||
|
||||
def collect_test_results(db_path=None, junit_path=None, coverage_path=None):
|
||||
"""Read junit XML + coverage.json → fact_test."""
|
||||
if db_path is None:
|
||||
db_path = _STORE_PATH
|
||||
if junit_path is None:
|
||||
junit_path = _TEST_RESULTS_XML
|
||||
if coverage_path is None:
|
||||
coverage_path = _COVERAGE_JSON
|
||||
if not os.path.isfile(junit_path):
|
||||
return 0
|
||||
_init_store(db_path)
|
||||
run_id = f"test-{_iso8601_now()}"
|
||||
total = passed = failed = errors = skipped = 0
|
||||
duration = 0.0
|
||||
try:
|
||||
tree = ET.parse(junit_path)
|
||||
root = tree.getroot()
|
||||
for suite in root.iter("testsuite"):
|
||||
total += int(suite.get("tests", 0))
|
||||
failed += int(suite.get("failures", 0))
|
||||
errors += int(suite.get("errors", 0))
|
||||
skipped += int(suite.get("skipped", 0))
|
||||
duration += float(suite.get("time", 0))
|
||||
passed = total - failed - errors - skipped
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
coverage_pct = 0.0
|
||||
if os.path.isfile(coverage_path):
|
||||
try:
|
||||
with open(coverage_path) as f:
|
||||
cov = json.load(f)
|
||||
coverage_pct = cov.get("totals", {}).get("percent_covered", 0.0)
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
conn = sqlite3.connect(db_path)
|
||||
conn.execute("""
|
||||
INSERT OR REPLACE INTO fact_test
|
||||
(run_id, total_tests, passed, failed, errors, skipped, duration_s, coverage_pct, collected_at)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
""", (run_id, total, passed, failed, errors, skipped, duration, coverage_pct, _iso8601_now()))
|
||||
conn.commit()
|
||||
conn.close()
|
||||
return 1
|
||||
|
||||
|
||||
def collect_lifecycle_reports(db_path=None, lifecycle_dir=None):
|
||||
"""Read metrics/lifecycle/*.json → fact_lifecycle."""
|
||||
if db_path is None:
|
||||
db_path = _STORE_PATH
|
||||
if lifecycle_dir is None:
|
||||
lifecycle_dir = os.path.join(_METRICS_DIR, "lifecycle")
|
||||
if not os.path.isdir(lifecycle_dir):
|
||||
return 0
|
||||
_init_store(db_path)
|
||||
count = 0
|
||||
conn = sqlite3.connect(db_path)
|
||||
for fname in sorted(os.listdir(lifecycle_dir)):
|
||||
if not fname.endswith(".json"):
|
||||
continue
|
||||
fpath = os.path.join(lifecycle_dir, fname)
|
||||
with open(fpath) as f:
|
||||
report = json.load(f)
|
||||
conn.execute("""
|
||||
INSERT OR REPLACE INTO fact_lifecycle
|
||||
(module, environment, phase, result, duration_ms, run_at)
|
||||
VALUES (?, ?, ?, ?, ?, ?)
|
||||
""", (report.get("module", ""), report.get("environment", ""),
|
||||
report.get("phase", ""), report.get("result", ""),
|
||||
report.get("duration_ms", 0), report.get("run_at", _iso8601_now())))
|
||||
count += 1
|
||||
conn.commit()
|
||||
conn.close()
|
||||
return count
|
||||
|
||||
|
||||
def collect_all(db_path=None):
|
||||
"""Run all collectors. Returns a summary dict."""
|
||||
if db_path is None:
|
||||
db_path = _STORE_PATH
|
||||
_init_store(db_path)
|
||||
summary = {
|
||||
"capabilities": collect_regression_report(db_path),
|
||||
"runs": collect_run_manifests(db_path),
|
||||
"decisions": collect_decision_ledger(db_path),
|
||||
"tests": collect_test_results(db_path),
|
||||
"lifecycle": collect_lifecycle_reports(db_path),
|
||||
"collected_at": _iso8601_now(),
|
||||
}
|
||||
return summary
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
result = collect_all()
|
||||
print(json.dumps(result, indent=2))
|
||||
@@ -1,257 +0,0 @@
|
||||
"""Nova Decision Ledger — SQLite append-only hash-chain (REQ-188, D-121).
|
||||
|
||||
Extends outbox_writer.py to emit to a SQLite append-only table with a hash
|
||||
chain (prev_hash + own hash, SHA-256). Stores ai.decision.made events
|
||||
(decision_id=run_id, chosen_action=band, confidence=score,
|
||||
alternatives=perInput, human_override=HITL block) with outcome backfill
|
||||
from apply.completed. Also stores attestation.recorded events (D-132).
|
||||
|
||||
Honors D-083 (no S3 Object Lock/JWS — local SQLite hash-chain only).
|
||||
D-120: Nova-native (SQLite, no QLDB).
|
||||
D-128: metrics/ at repo root.
|
||||
"""
|
||||
|
||||
import datetime
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import sys
|
||||
|
||||
_LEDGER_PATH = os.path.join(
|
||||
os.path.dirname(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))),
|
||||
"metrics", "decision_ledger.db",
|
||||
)
|
||||
|
||||
_GENESIS_HASH = "GENESIS"
|
||||
|
||||
|
||||
def _iso8601_now():
|
||||
return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
||||
|
||||
|
||||
def _canonical_hash(event):
|
||||
"""SHA-256 over canonical JSON (sort_keys, compact separators)."""
|
||||
canonical = json.dumps(event, sort_keys=True, separators=(",", ":"))
|
||||
return hashlib.sha256(canonical.encode("utf-8")).hexdigest()
|
||||
|
||||
|
||||
def _init_db(db_path=None):
|
||||
"""Create the ledger table if it doesn't exist."""
|
||||
if db_path is None:
|
||||
db_path = _LEDGER_PATH
|
||||
os.makedirs(os.path.dirname(db_path), exist_ok=True)
|
||||
conn = sqlite3.connect(db_path)
|
||||
conn.execute("""
|
||||
CREATE TABLE IF NOT EXISTS decision_ledger (
|
||||
seq INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
event_id TEXT NOT NULL,
|
||||
event_type TEXT NOT NULL,
|
||||
run_id TEXT NOT NULL,
|
||||
contract_id TEXT,
|
||||
environment TEXT,
|
||||
event_time TEXT NOT NULL,
|
||||
payload TEXT NOT NULL,
|
||||
prev_hash TEXT NOT NULL,
|
||||
hash TEXT NOT NULL
|
||||
)
|
||||
""")
|
||||
conn.execute("CREATE INDEX IF NOT EXISTS idx_run_id ON decision_ledger(run_id)")
|
||||
conn.execute("CREATE INDEX IF NOT EXISTS idx_event_type ON decision_ledger(event_type)")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
|
||||
|
||||
def _get_last_hash(db_path=None):
|
||||
"""Get the hash of the last row in the ledger (or GENESIS if empty)."""
|
||||
if db_path is None:
|
||||
db_path = _LEDGER_PATH
|
||||
conn = sqlite3.connect(db_path)
|
||||
row = conn.execute("SELECT hash FROM decision_ledger ORDER BY seq DESC LIMIT 1").fetchone()
|
||||
conn.close()
|
||||
return row[0] if row else _GENESIS_HASH
|
||||
|
||||
|
||||
def append(event, db_path=None):
|
||||
"""Append an event to the Decision Ledger with hash-chain integrity.
|
||||
|
||||
Args:
|
||||
event: a CloudEvents 1.0 envelope dict (from event_envelope.make_event)
|
||||
db_path: path to the SQLite ledger
|
||||
|
||||
Returns:
|
||||
The row dict (seq, event_id, event_type, run_id, hash, prev_hash).
|
||||
"""
|
||||
if db_path is None:
|
||||
db_path = _LEDGER_PATH
|
||||
_init_db(db_path)
|
||||
prev_hash = _get_last_hash(db_path)
|
||||
event_hash = _canonical_hash(event)
|
||||
platform = event.get("platform", {})
|
||||
data = event.get("data", {})
|
||||
|
||||
conn = sqlite3.connect(db_path)
|
||||
conn.execute("BEGIN IMMEDIATE")
|
||||
cursor = conn.execute(
|
||||
"""INSERT INTO decision_ledger
|
||||
(event_id, event_type, run_id, contract_id, environment, event_time, payload, prev_hash, hash)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)""",
|
||||
(
|
||||
event.get("id", ""),
|
||||
event.get("type", ""),
|
||||
platform.get("run_id", ""),
|
||||
platform.get("contract_id", ""),
|
||||
platform.get("environment", ""),
|
||||
event.get("time", _iso8601_now()),
|
||||
json.dumps(event, sort_keys=True),
|
||||
prev_hash,
|
||||
event_hash,
|
||||
),
|
||||
)
|
||||
seq = cursor.lastrowid
|
||||
conn.commit()
|
||||
conn.close()
|
||||
return {"seq": seq, "event_id": event.get("id", ""), "event_type": event.get("type", ""),
|
||||
"run_id": platform.get("run_id", ""), "hash": event_hash, "prev_hash": prev_hash}
|
||||
|
||||
|
||||
def verify_chain(db_path=None):
|
||||
"""Verify the hash chain integrity. Returns (ok, broken_count, details).
|
||||
|
||||
Recomputes each row's hash from its payload and checks:
|
||||
1. The stored hash matches the recomputed hash.
|
||||
2. The prev_hash matches the previous row's hash.
|
||||
"""
|
||||
if db_path is None:
|
||||
db_path = _LEDGER_PATH
|
||||
_init_db(db_path)
|
||||
conn = sqlite3.connect(db_path)
|
||||
rows = conn.execute("SELECT seq, hash, prev_hash, payload FROM decision_ledger ORDER BY seq").fetchall()
|
||||
conn.close()
|
||||
if not rows:
|
||||
return True, 0, "empty ledger"
|
||||
|
||||
broken = 0
|
||||
details = []
|
||||
prev_hash = _GENESIS_HASH
|
||||
for seq, stored_hash, stored_prev, payload_json in rows:
|
||||
event = json.loads(payload_json)
|
||||
recomputed = _canonical_hash(event)
|
||||
if recomputed != stored_hash:
|
||||
broken += 1
|
||||
details.append(f"seq={seq}: hash mismatch (stored={stored_hash[:12]}... recomputed={recomputed[:12]}...)")
|
||||
if stored_prev != prev_hash:
|
||||
broken += 1
|
||||
details.append(f"seq={seq}: prev_hash mismatch (expected={prev_hash[:12]}... got={stored_prev[:12]}...)")
|
||||
prev_hash = stored_hash
|
||||
return broken == 0, broken, "; ".join(details) if details else "chain intact"
|
||||
|
||||
|
||||
def query_by_run(run_id, db_path=None):
|
||||
"""Query all ledger entries for a given run_id."""
|
||||
if db_path is None:
|
||||
db_path = _LEDGER_PATH
|
||||
_init_db(db_path)
|
||||
conn = sqlite3.connect(db_path)
|
||||
rows = conn.execute(
|
||||
"SELECT seq, event_type, event_time, payload FROM decision_ledger WHERE run_id = ? ORDER BY seq",
|
||||
(run_id,),
|
||||
).fetchall()
|
||||
conn.close()
|
||||
return [{"seq": r[0], "event_type": r[1], "event_time": r[2], "payload": json.loads(r[3])} for r in rows]
|
||||
|
||||
|
||||
def stats(db_path=None):
|
||||
"""Return ledger statistics."""
|
||||
if db_path is None:
|
||||
db_path = _LEDGER_PATH
|
||||
_init_db(db_path)
|
||||
conn = sqlite3.connect(db_path)
|
||||
total = conn.execute("SELECT COUNT(*) FROM decision_ledger").fetchone()[0]
|
||||
by_type = conn.execute("SELECT event_type, COUNT(*) FROM decision_ledger GROUP BY event_type").fetchall()
|
||||
by_env = conn.execute("SELECT environment, COUNT(*) FROM decision_ledger GROUP BY environment").fetchall()
|
||||
conn.close()
|
||||
return {
|
||||
"total": total,
|
||||
"by_event_type": dict(by_type),
|
||||
"by_environment": dict(by_env),
|
||||
}
|
||||
|
||||
|
||||
def export_since(since_iso, fmt="json", db_path=None):
|
||||
"""Export ledger entries since a given ISO8601 timestamp."""
|
||||
if db_path is None:
|
||||
db_path = _LEDGER_PATH
|
||||
_init_db(db_path)
|
||||
conn = sqlite3.connect(db_path)
|
||||
rows = conn.execute(
|
||||
"SELECT seq, event_type, run_id, event_time, payload FROM decision_ledger WHERE event_time >= ? ORDER BY seq",
|
||||
(since_iso,),
|
||||
).fetchall()
|
||||
conn.close()
|
||||
entries = [{"seq": r[0], "event_type": r[1], "run_id": r[2], "event_time": r[3], "payload": json.loads(r[4])} for r in rows]
|
||||
if fmt == "csv":
|
||||
import csv
|
||||
import io
|
||||
buf = io.StringIO()
|
||||
writer = csv.DictWriter(buf, fieldnames=["seq", "event_type", "run_id", "event_time", "payload"])
|
||||
writer.writeheader()
|
||||
for e in entries:
|
||||
e["payload"] = json.dumps(e["payload"])
|
||||
writer.writerow(e)
|
||||
return buf.getvalue()
|
||||
return json.dumps(entries, indent=2)
|
||||
|
||||
|
||||
def replay_run(run_id, db_path=None):
|
||||
"""Reconstruct a run's full event sequence from the ledger.
|
||||
|
||||
Prints the ordered event sequence (run.started -> policy.evaluated ->
|
||||
confidence.computed -> ai.decision.made -> attestation.recorded ->
|
||||
run.completed/failed) with the decision's confidence, alternatives,
|
||||
and outcome.
|
||||
"""
|
||||
if db_path is None:
|
||||
db_path = _LEDGER_PATH
|
||||
entries = query_by_run(run_id, db_path)
|
||||
if not entries:
|
||||
return f"no events found for run_id={run_id}"
|
||||
lines = [f"=== Replay: run_id={run_id} ({len(entries)} events) ==="]
|
||||
for e in entries:
|
||||
payload = e["payload"]
|
||||
data = payload.get("data", {})
|
||||
etype = e["event_type"]
|
||||
line = f" [{e['seq']}] {e['event_time']} {etype}"
|
||||
if etype == "nova.ai.decision.made":
|
||||
line += f" confidence={data.get('confidence', '?')} band={data.get('chosen_action', '?')} override={data.get('human_override', '?')}"
|
||||
elif etype == "nova.attestation.recorded":
|
||||
line += f" env={data.get('environment', '?')} approver={data.get('approver', '?')} result={data.get('result', '?')}"
|
||||
elif etype == "nova.run.completed":
|
||||
line += f" exit={data.get('exit_code', '?')} outcome={data.get('outcome', '?')}"
|
||||
elif etype == "nova.run.failed":
|
||||
line += f" exit={data.get('exit_code', '?')} outcome=failed"
|
||||
lines.append(line)
|
||||
lines.append("=== End replay ===")
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
if len(sys.argv) < 2:
|
||||
print("usage: decision_ledger.py <verify-chain|stats|query|export|replay> [args]", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
cmd = sys.argv[1]
|
||||
if cmd == "verify-chain":
|
||||
ok, broken, details = verify_chain()
|
||||
print(f"chain_ok={ok} broken={broken} details={details}")
|
||||
sys.exit(0 if ok else 1)
|
||||
elif cmd == "stats":
|
||||
print(json.dumps(stats(), indent=2))
|
||||
elif cmd == "query" and len(sys.argv) >= 3:
|
||||
print(json.dumps(query_by_run(sys.argv[2]), indent=2))
|
||||
elif cmd == "export" and len(sys.argv) >= 3:
|
||||
print(export_since(sys.argv[2]))
|
||||
elif cmd == "replay" and len(sys.argv) >= 3:
|
||||
print(replay_run(sys.argv[2]))
|
||||
else:
|
||||
print(f"unknown command: {cmd}", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
@@ -1,39 +0,0 @@
|
||||
"""Nova Decision Ledger CLI (REQ-207).
|
||||
|
||||
Subcommands: query, verify-chain, stats, export, replay.
|
||||
Read-only CLI for the Decision Ledger SQLite hash-chain.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))))
|
||||
from core.metrics.decision_ledger import query_by_run, verify_chain, stats, export_since, replay_run
|
||||
|
||||
|
||||
def main():
|
||||
if len(sys.argv) < 2:
|
||||
print("usage: decision_ledger_cli.py <query|verify-chain|stats|export|replay> [args]", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
cmd = sys.argv[1]
|
||||
if cmd == "query" and len(sys.argv) >= 3:
|
||||
print(json.dumps(query_by_run(sys.argv[2]), indent=2))
|
||||
elif cmd == "verify-chain":
|
||||
ok, broken, details = verify_chain()
|
||||
print(f"chain_ok={ok} broken={broken} details={details}")
|
||||
sys.exit(0 if ok else 1)
|
||||
elif cmd == "stats":
|
||||
print(json.dumps(stats(), indent=2))
|
||||
elif cmd == "export" and len(sys.argv) >= 3:
|
||||
fmt = sys.argv[3] if len(sys.argv) >= 4 else "json"
|
||||
print(export_since(sys.argv[2], fmt=fmt))
|
||||
elif cmd == "replay" and len(sys.argv) >= 3:
|
||||
print(replay_run(sys.argv[2]))
|
||||
else:
|
||||
print(f"unknown command: {cmd}", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -1,98 +0,0 @@
|
||||
"""Nova CloudEvents 1.0 envelope + platform.* semantic conventions (REQ-187).
|
||||
|
||||
Defines the standard event envelope for all Nova metrics events. Every
|
||||
emitter (run_manifest, decision_ledger, confidence_signal, checkov_adapter,
|
||||
hitl_gates, regression_verify) uses `make_event()` to produce a valid
|
||||
CloudEvents 1.0 envelope. Events are appended to `metrics/events.jsonl`.
|
||||
|
||||
D-120: Nova-native minimal tech (no Kafka/OTel SDK — JSONL + SQLite).
|
||||
D-125: hybrid model — existing file signals stay as files; the collector
|
||||
reads them and emits normalized CloudEvents. New emitters emit directly.
|
||||
"""
|
||||
|
||||
import datetime
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import uuid
|
||||
|
||||
METRICS_DIR = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), "metrics")
|
||||
EVENTS_LOG = os.path.join(METRICS_DIR, "events.jsonl")
|
||||
|
||||
|
||||
def _iso8601_now():
|
||||
return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
||||
|
||||
|
||||
def make_event(event_type, run_id, environment, data, contract_id="", source="nova.platform", subject="", actor_type="confidence-gate", actor_id="confidence_signal"):
|
||||
"""Build a CloudEvents 1.0 envelope with Nova platform.* conventions.
|
||||
|
||||
Args:
|
||||
event_type: e.g. "nova.run.completed", "nova.ai.decision.made"
|
||||
run_id: the run identifier (e.g. "run-<epoch>")
|
||||
environment: dev|qa|prod|dr
|
||||
data: the event payload dict
|
||||
contract_id: the contract UUID (optional)
|
||||
source: the event source (default "nova.platform")
|
||||
subject: the event subject (default "<contract_id>/<env>")
|
||||
actor_type: the actor type (default "confidence-gate")
|
||||
actor_id: the actor id (default "confidence_signal")
|
||||
|
||||
Returns:
|
||||
A CloudEvents 1.0 envelope dict.
|
||||
"""
|
||||
if not subject:
|
||||
subject = f"{contract_id}/{environment}" if contract_id else environment
|
||||
return {
|
||||
"specversion": "1.0",
|
||||
"id": str(uuid.uuid4()),
|
||||
"source": source,
|
||||
"type": event_type,
|
||||
"time": _iso8601_now(),
|
||||
"subject": subject,
|
||||
"datacontenttype": "application/json",
|
||||
"platform": {
|
||||
"tenant_id": "acdl",
|
||||
"run_id": run_id,
|
||||
"contract_id": contract_id,
|
||||
"environment": environment,
|
||||
"actor": {"type": actor_type, "id": actor_id},
|
||||
"trace_id": run_id,
|
||||
},
|
||||
"data": data,
|
||||
}
|
||||
|
||||
|
||||
def append_event(event, events_log=None):
|
||||
"""Append a CloudEvents envelope to the JSONL event log.
|
||||
|
||||
Creates the metrics/ directory if it doesn't exist.
|
||||
"""
|
||||
if events_log is None:
|
||||
events_log = EVENTS_LOG
|
||||
os.makedirs(os.path.dirname(events_log), exist_ok=True)
|
||||
with open(events_log, "a", encoding="utf-8") as fh:
|
||||
fh.write(json.dumps(event, sort_keys=True, separators=(",", ":")) + "\n")
|
||||
|
||||
|
||||
def emit(event_type, run_id, environment, data, **kwargs):
|
||||
"""Make an event + append it to the JSONL log. Convenience wrapper."""
|
||||
event = make_event(event_type, run_id, environment, data, **kwargs)
|
||||
append_event(event)
|
||||
return event
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
if len(sys.argv) < 4:
|
||||
print("usage: event_envelope.py <event_type> <run_id> <environment> [data.json]", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
_type = sys.argv[1]
|
||||
_run_id = sys.argv[2]
|
||||
_env = sys.argv[3]
|
||||
_data = {}
|
||||
if len(sys.argv) >= 5 and os.path.isfile(sys.argv[4]):
|
||||
with open(sys.argv[4]) as f:
|
||||
_data = json.load(f)
|
||||
ev = emit(_type, _run_id, _env, _data)
|
||||
print(json.dumps(ev, indent=2))
|
||||
@@ -1,73 +0,0 @@
|
||||
"""Nova Infracost Post-Processor (REQ-187, D-120).
|
||||
|
||||
Runs Infracost on `terraform show -json plan.tfplan` (offline, reads plan
|
||||
JSON, no live AWS). Emits nova.cost.estimated{delta_usd} events. Degrades
|
||||
gracefully (omits the event, logs a warning) when Infracost CLI is absent
|
||||
(assumption A6).
|
||||
|
||||
run_platform.sh invokes it after the plan stage.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))))
|
||||
from core.metrics.event_envelope import emit
|
||||
|
||||
|
||||
def _is_infracost_available():
|
||||
"""Check if the Infracost CLI is on PATH."""
|
||||
return shutil.which("infracost") is not None
|
||||
|
||||
|
||||
def estimate(plan_json_path, run_id, contract_id, environment):
|
||||
"""Run Infracost on a terraform plan JSON. Returns the cost estimate dict.
|
||||
|
||||
Args:
|
||||
plan_json_path: path to `terraform show -json plan.tfplan` output
|
||||
run_id: the run identifier
|
||||
contract_id: the contract UUID
|
||||
environment: dev|qa|prod|dr
|
||||
|
||||
Returns:
|
||||
{"delta_usd": float, "total_monthly_usd": float, "available": bool}
|
||||
or {"available": False} if Infracost is not installed.
|
||||
"""
|
||||
if not _is_infracost_available():
|
||||
sys.stderr.write("[infracost] CLI not found — cost.estimated event omitted (A6 degraded mode)\n")
|
||||
return {"available": False, "delta_usd": 0.0, "total_monthly_usd": 0.0}
|
||||
|
||||
if not os.path.isfile(plan_json_path):
|
||||
sys.stderr.write(f"[infracost] plan JSON not found: {plan_json_path}\n")
|
||||
return {"available": False, "delta_usd": 0.0, "total_monthly_usd": 0.0}
|
||||
|
||||
try:
|
||||
result = subprocess.run(
|
||||
["infracost", "breakdown", "--path", plan_json_path, "--format", "json"],
|
||||
capture_output=True, text=True, timeout=30,
|
||||
)
|
||||
if result.returncode != 0:
|
||||
sys.stderr.write(f"[infracost] CLI failed: {result.stderr[:200]}\n")
|
||||
return {"available": False, "delta_usd": 0.0, "total_monthly_usd": 0.0}
|
||||
|
||||
breakdown = json.loads(result.stdout)
|
||||
delta = float(breakdown.get("diffTotalMonthlyCost", 0.0))
|
||||
total = float(breakdown.get("totalMonthlyCost", 0.0))
|
||||
estimate_data = {"available": True, "delta_usd": delta, "total_monthly_usd": total}
|
||||
|
||||
emit("nova.cost.estimated", run_id, environment, estimate_data, contract_id=contract_id)
|
||||
return estimate_data
|
||||
except Exception as exc:
|
||||
sys.stderr.write(f"[infracost] error: {exc}\n")
|
||||
return {"available": False, "delta_usd": 0.0, "total_monthly_usd": 0.0}
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
if len(sys.argv) < 5:
|
||||
print("usage: infracost_adapter.py <plan_json_path> <run_id> <contract_id> <environment>", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
est = estimate(sys.argv[1], sys.argv[2], sys.argv[3], sys.argv[4])
|
||||
print(json.dumps(est, indent=2))
|
||||
@@ -1,198 +0,0 @@
|
||||
"""Nova PowerBI Export (REQ-190, P3).
|
||||
|
||||
Emits CSV/JSON views to metrics/powerbi/ from the SQLite cold store.
|
||||
Fact + dimension tables + 8 empty placeholder views for deferred metrics
|
||||
(with documented schemas ready to fill when their blocking decisions lift).
|
||||
|
||||
D-120: Nova-native (CSV/JSON files, no live connector)
|
||||
D-129: PowerBI ingests via the folder connector
|
||||
D-128: metrics/ at repo root
|
||||
"""
|
||||
|
||||
import csv
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import sys
|
||||
|
||||
_METRICS_DIR = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), "metrics")
|
||||
_STORE_PATH = os.path.join(_METRICS_DIR, "nova_metrics.db")
|
||||
_EXPORT_DIR = os.path.join(_METRICS_DIR, "powerbi")
|
||||
|
||||
FACT_VIEWS = [
|
||||
"fact_run",
|
||||
"fact_capability",
|
||||
"fact_policy_check",
|
||||
"fact_confidence",
|
||||
"fact_test",
|
||||
"fact_decision",
|
||||
"fact_cost_estimate",
|
||||
"fact_lifecycle",
|
||||
]
|
||||
|
||||
DIM_VIEWS = [
|
||||
"dim_capability",
|
||||
"dim_milestone",
|
||||
]
|
||||
|
||||
PLACEHOLDER_VIEWS = {
|
||||
"placeholder_live_infra_health": {
|
||||
"columns": ["timestamp", "resource_id", "resource_type", "running_count", "healthy", "downtime_seconds"],
|
||||
"blocking_decision": "D-096",
|
||||
"description": "Live infrastructure health (ECS running count, ALB 5xx, RPS). Blocked: live AWS torn down.",
|
||||
},
|
||||
"placeholder_live_outbox_rate": {
|
||||
"columns": ["timestamp", "contract_id", "write_latency_ms", "append_count"],
|
||||
"blocking_decision": "D-096",
|
||||
"description": "Live outbox write rate / ledger append latency. Blocked: DynamoDB outbox table absent.",
|
||||
},
|
||||
"placeholder_tamper_evident_checkpoints": {
|
||||
"columns": ["timestamp", "checkpoint_id", "jws_signed", "object_lock_enabled"],
|
||||
"blocking_decision": "D-083",
|
||||
"description": "Tamper-evident ledger checkpoints / JWS signature rate. Blocked: S3 Object Lock + JWS deferred.",
|
||||
},
|
||||
"placeholder_onboarding_funnel": {
|
||||
"columns": ["timestamp", "consumer_repo", "requested_environment", "status", "granted_at"],
|
||||
"blocking_decision": "D-113/D-114/D-119",
|
||||
"description": "Onboarding funnel: requested → granted conversion. Blocked: no auto-grant event.",
|
||||
},
|
||||
"placeholder_drift_detection": {
|
||||
"columns": ["timestamp", "workspace_id", "drift_count", "auto_reverted", "detection_cycle"],
|
||||
"blocking_decision": "D-096 + no scheduler",
|
||||
"description": "Drift detection (scheduled terraform plan -detailed-exitcode). Blocked: live AWS + scheduler.",
|
||||
},
|
||||
"placeholder_live_cur_reconciliation": {
|
||||
"columns": ["timestamp", "resource_address", "actual_usd", "baseline_usd", "saved_usd"],
|
||||
"blocking_decision": "D-096",
|
||||
"description": "Live cost CUR reconciliation. Blocked: live AWS billing. Infracost pre-apply estimates are in fact_cost_estimate.",
|
||||
},
|
||||
"placeholder_sla_downtime": {
|
||||
"columns": ["timestamp", "service", "uptime_pct", "downtime_minutes", "slo_target"],
|
||||
"blocking_decision": "D-096",
|
||||
"description": "SLA / unplanned downtime. Blocked: needs live service uptime monitoring.",
|
||||
},
|
||||
"placeholder_predictive_reactive": {
|
||||
"columns": ["timestamp", "action_id", "label", "trigger", "count"],
|
||||
"blocking_decision": "future emitter",
|
||||
"description": "Predictive vs Reactive ratio. Blocked: requires ML anomaly-forecasting service.",
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def _iso8601_now():
|
||||
return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
||||
|
||||
|
||||
def _export_table_csv(conn, table_name, export_dir):
|
||||
"""Export a SQLite table to a CSV file."""
|
||||
rows = conn.execute(f"SELECT * FROM {table_name}").fetchall()
|
||||
if not rows:
|
||||
return 0
|
||||
columns = [desc[0] for desc in conn.execute(f"SELECT * FROM {table_name} LIMIT 0").description]
|
||||
csv_path = os.path.join(export_dir, f"{table_name}.csv")
|
||||
with open(csv_path, "w", newline="", encoding="utf-8") as f:
|
||||
writer = csv.writer(f)
|
||||
writer.writerow(columns)
|
||||
writer.writerows(rows)
|
||||
return len(rows)
|
||||
|
||||
|
||||
def _export_table_json(conn, table_name, export_dir):
|
||||
"""Export a SQLite table to a JSON file."""
|
||||
rows = conn.execute(f"SELECT * FROM {table_name}").fetchall()
|
||||
if not rows:
|
||||
return 0
|
||||
columns = [desc[0] for desc in conn.execute(f"SELECT * FROM {table_name} LIMIT 0").description]
|
||||
records = [dict(zip(columns, row)) for row in rows]
|
||||
json_path = os.path.join(export_dir, f"{table_name}.json")
|
||||
with open(json_path, "w", encoding="utf-8") as f:
|
||||
json.dump(records, f, indent=2, default=str)
|
||||
return len(rows)
|
||||
|
||||
|
||||
def _export_placeholder_csv(view_name, schema, export_dir):
|
||||
"""Export a placeholder CSV with headers only (no data rows)."""
|
||||
csv_path = os.path.join(export_dir, f"{view_name}.csv")
|
||||
with open(csv_path, "w", newline="", encoding="utf-8") as f:
|
||||
writer = csv.writer(f)
|
||||
writer.writerow(schema["columns"])
|
||||
return 0
|
||||
|
||||
|
||||
def _export_placeholder_json(view_name, schema, export_dir):
|
||||
"""Export a placeholder JSON with schema metadata (no data rows)."""
|
||||
json_path = os.path.join(export_dir, f"{view_name}.json")
|
||||
with open(json_path, "w", encoding="utf-8") as f:
|
||||
json.dump({"schema": schema, "data": []}, f, indent=2)
|
||||
return 0
|
||||
|
||||
|
||||
def export_all(store_path=None, export_dir=None, fmt="both"):
|
||||
"""Export all fact/dim tables + placeholder views to CSV and/or JSON.
|
||||
|
||||
Args:
|
||||
store_path: path to the SQLite cold store
|
||||
export_dir: directory for exported files
|
||||
fmt: "csv", "json", or "both"
|
||||
|
||||
Returns:
|
||||
Summary dict with export counts.
|
||||
"""
|
||||
if store_path is None:
|
||||
store_path = _STORE_PATH
|
||||
if export_dir is None:
|
||||
export_dir = _EXPORT_DIR
|
||||
os.makedirs(export_dir, exist_ok=True)
|
||||
|
||||
summary = {"exported_at": _iso8601_now(), "fact_tables": {}, "dim_tables": {}, "placeholder_views": {}}
|
||||
|
||||
if not os.path.isfile(store_path):
|
||||
summary["error"] = f"SQLite store not found: {store_path}"
|
||||
for view_name, schema in PLACEHOLDER_VIEWS.items():
|
||||
if fmt in ("csv", "both"):
|
||||
_export_placeholder_csv(view_name, schema, export_dir)
|
||||
if fmt in ("json", "both"):
|
||||
_export_placeholder_json(view_name, schema, export_dir)
|
||||
summary["placeholder_views"][view_name] = 0
|
||||
return summary
|
||||
|
||||
conn = sqlite3.connect(store_path)
|
||||
|
||||
for table in FACT_VIEWS:
|
||||
count = 0
|
||||
try:
|
||||
if fmt in ("csv", "both"):
|
||||
count = _export_table_csv(conn, table, export_dir)
|
||||
if fmt in ("json", "both"):
|
||||
count = _export_table_json(conn, table, export_dir)
|
||||
except sqlite3.OperationalError:
|
||||
count = 0
|
||||
summary["fact_tables"][table] = count
|
||||
|
||||
for table in DIM_VIEWS:
|
||||
count = 0
|
||||
try:
|
||||
if fmt in ("csv", "both"):
|
||||
count = _export_table_csv(conn, table, export_dir)
|
||||
if fmt in ("json", "both"):
|
||||
count = _export_table_json(conn, table, export_dir)
|
||||
except sqlite3.OperationalError:
|
||||
count = 0
|
||||
summary["dim_tables"][table] = count
|
||||
|
||||
conn.close()
|
||||
|
||||
for view_name, schema in PLACEHOLDER_VIEWS.items():
|
||||
if fmt in ("csv", "both"):
|
||||
_export_placeholder_csv(view_name, schema, export_dir)
|
||||
if fmt in ("json", "both"):
|
||||
_export_placeholder_json(view_name, schema, export_dir)
|
||||
summary["placeholder_views"][view_name] = 0
|
||||
|
||||
return summary
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
result = export_all()
|
||||
print(json.dumps(result, indent=2))
|
||||
@@ -1,137 +0,0 @@
|
||||
"""Nova Per-Run Manifest Writer (REQ-187).
|
||||
|
||||
Emits nova.run.started, nova.run.completed, nova.run.failed events with
|
||||
(run_id, contractId, env, stages x durations, exit, confidence, HITL block
|
||||
count). Writes metrics/runs/<run_id>.json. scripts/run_platform.sh invokes
|
||||
the writer at run start + run end.
|
||||
|
||||
D-120: Nova-native (JSONL events + JSON manifest file, no Kafka).
|
||||
D-128: metrics/ at repo root.
|
||||
"""
|
||||
|
||||
import datetime
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import time
|
||||
import uuid
|
||||
|
||||
_METRICS_DIR = os.path.join(os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))), "metrics")
|
||||
_RUNS_DIR = os.path.join(_METRICS_DIR, "runs")
|
||||
|
||||
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))))
|
||||
from core.metrics.event_envelope import emit, make_event, append_event
|
||||
|
||||
|
||||
def _iso8601_now():
|
||||
return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
||||
|
||||
|
||||
def _run_id():
|
||||
return f"run-{int(time.time())}-{uuid.uuid4().hex[:8]}"
|
||||
|
||||
|
||||
def start_run(contract_id, environment, stages=None):
|
||||
"""Emit nova.run.started + return the run_id."""
|
||||
run_id = _run_id()
|
||||
data = {
|
||||
"contract_id": contract_id,
|
||||
"environment": environment,
|
||||
"started_at": _iso8601_now(),
|
||||
"stages": stages or [],
|
||||
}
|
||||
emit("nova.run.started", run_id, environment, data, contract_id=contract_id)
|
||||
return run_id
|
||||
|
||||
|
||||
def complete_run(run_id, contract_id, environment, stages, exit_code, confidence=None, hitl=None, policy=None, cost_estimate_usd=None, decision_id=None):
|
||||
"""Emit nova.run.completed + write the per-run manifest JSON.
|
||||
|
||||
Args:
|
||||
run_id: the run identifier from start_run()
|
||||
contract_id: the contract UUID
|
||||
environment: dev|qa|prod|dr
|
||||
stages: list of {name, duration_ms, exit_code, error?}
|
||||
exit_code: the overall run exit code
|
||||
confidence: optional {score, band, perInput}
|
||||
hitl: optional {gate, result, block}
|
||||
policy: optional {passed, failed, skipped}
|
||||
cost_estimate_usd: optional float
|
||||
decision_id: optional string (links to the Decision Ledger)
|
||||
"""
|
||||
started_at = stages[0].get("started_at", _iso8601_now()) if stages else _iso8601_now()
|
||||
completed_at = _iso8601_now()
|
||||
outcome = "succeeded" if exit_code == 0 else "failed"
|
||||
|
||||
manifest = {
|
||||
"run_id": run_id,
|
||||
"contract_id": contract_id,
|
||||
"environment": environment,
|
||||
"started_at": started_at,
|
||||
"completed_at": completed_at,
|
||||
"exit_code": exit_code,
|
||||
"stages": stages,
|
||||
"outcome": outcome,
|
||||
}
|
||||
if confidence:
|
||||
manifest["confidence"] = confidence
|
||||
if hitl:
|
||||
manifest["hitl"] = hitl
|
||||
if policy:
|
||||
manifest["policy"] = policy
|
||||
if cost_estimate_usd is not None:
|
||||
manifest["cost_estimate_usd"] = cost_estimate_usd
|
||||
if decision_id:
|
||||
manifest["decision_id"] = decision_id
|
||||
|
||||
os.makedirs(_RUNS_DIR, exist_ok=True)
|
||||
manifest_path = os.path.join(_RUNS_DIR, f"{run_id}.json")
|
||||
with open(manifest_path, "w", encoding="utf-8") as fh:
|
||||
json.dump(manifest, fh, indent=2, sort_keys=True)
|
||||
|
||||
event_type = "nova.run.completed" if exit_code == 0 else "nova.run.failed"
|
||||
emit(event_type, run_id, environment, manifest, contract_id=contract_id)
|
||||
|
||||
return manifest
|
||||
|
||||
|
||||
def persist_run_artifacts(run_id, work_dir):
|
||||
"""Copy ephemeral $WORK/*.json to metrics/runs/<run_id>/ as durable artifacts.
|
||||
|
||||
Args:
|
||||
run_id: the run identifier
|
||||
work_dir: the $WORK directory (e.g. /tmp/nova_platform_run)
|
||||
"""
|
||||
if not work_dir or not os.path.isdir(work_dir):
|
||||
return []
|
||||
dest = os.path.join(_RUNS_DIR, run_id)
|
||||
os.makedirs(dest, exist_ok=True)
|
||||
copied = []
|
||||
for fname in ("pcr.json", "signal.json", "event.json", "outbox_item.json", "stack.json", "checkov.json"):
|
||||
src = os.path.join(work_dir, fname)
|
||||
if os.path.isfile(src):
|
||||
import shutil
|
||||
shutil.copy2(src, os.path.join(dest, fname))
|
||||
copied.append(fname)
|
||||
return copied
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
if len(sys.argv) < 4:
|
||||
print("usage: run_manifest.py <start|complete|persist> <contract_id> <environment> [run_id] [work_dir]", file=sys.stderr)
|
||||
sys.exit(2)
|
||||
action = sys.argv[1]
|
||||
cid = sys.argv[2]
|
||||
env = sys.argv[3]
|
||||
if action == "start":
|
||||
rid = start_run(cid, env)
|
||||
print(rid)
|
||||
elif action == "complete":
|
||||
rid = sys.argv[4] if len(sys.argv) >= 5 else _run_id()
|
||||
m = complete_run(rid, cid, env, [], 0)
|
||||
print(json.dumps(m, indent=2))
|
||||
elif action == "persist":
|
||||
rid = sys.argv[4] if len(sys.argv) >= 5 else ""
|
||||
wd = sys.argv[5] if len(sys.argv) >= 6 else ""
|
||||
copied = persist_run_artifacts(rid, wd)
|
||||
print(json.dumps({"copied": copied}))
|
||||
@@ -1,131 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Nova Onboarding — auto-generate an environment binding file (P19, REQ-183).
|
||||
|
||||
Given a consumer onboarding request (validated against
|
||||
schemas/onboarding.schema.json), generate a ``<env>.json`` environment
|
||||
binding file from the dev template, filling in the consumer's ownerId +
|
||||
billingTag. The generated file is a starting point for the platform team
|
||||
(or a future automation) to bind to a real AWS account.
|
||||
|
||||
This is the "request path" half of the no-humans onboarding flow (D-113).
|
||||
Real AWS account/network/state provisioning is a future feature milestone;
|
||||
this module removes the human handoff from the *request* step by
|
||||
generating the binding file + emitting a git patch / PR-branch instruction.
|
||||
|
||||
Usage:
|
||||
python3 core/onboarding.py <request.json> [--out <env.json>]
|
||||
python3 core/onboarding.py --request '{"consumerRepo":"acdl/c","requestedEnvironment":"qa","ownerId":"team-a","billingTag":"cc-a"}'
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any, Dict
|
||||
|
||||
|
||||
def _repo_root() -> Path:
|
||||
return Path(__file__).resolve().parent.parent
|
||||
|
||||
|
||||
def _load_template_env(template_env: str = "dev", root: Path | None = None) -> Dict[str, Any]:
|
||||
"""Load the template environment JSON (defaults to dev.json)."""
|
||||
root = root or _repo_root()
|
||||
env_path = root / "core" / "environments" / f"{template_env}.json"
|
||||
if not env_path.is_file():
|
||||
raise FileNotFoundError(f"template environment {env_path} not found")
|
||||
return json.loads(env_path.read_text())
|
||||
|
||||
|
||||
def generate_env_file(
|
||||
request: Dict[str, Any],
|
||||
template_env: str = "dev",
|
||||
root: Path | None = None,
|
||||
) -> Dict[str, Any]:
|
||||
"""Generate an environment binding dict from a consumer onboarding request.
|
||||
|
||||
The generated dict is a copy of the template env with:
|
||||
- ``name`` → the requested environment
|
||||
- ``description`` → notes the consumer + owner
|
||||
- ``account_id`` → placeholder (000000000000) for the platform team
|
||||
to fill with the real account
|
||||
- ``ownerId`` + ``billingTag`` → from the request (for ABAC + cost)
|
||||
|
||||
The dict validates against schemas/environment.schema.json.
|
||||
|
||||
Returns the generated env dict.
|
||||
"""
|
||||
template = _load_template_env(template_env, root)
|
||||
requested = request["requestedEnvironment"]
|
||||
owner = request["ownerId"]
|
||||
billing = request["billingTag"]
|
||||
consumer = request["consumerRepo"]
|
||||
|
||||
env = dict(template)
|
||||
env["name"] = requested
|
||||
env["description"] = (
|
||||
f"Auto-generated binding for {consumer} (owner={owner}, "
|
||||
f"billing={billing}). Replace account_id with the real "
|
||||
f"{requested} account before deploying."
|
||||
)
|
||||
env["account_id"] = "000000000000" # placeholder — platform team fills
|
||||
env["ownerId"] = owner
|
||||
env["billingTag"] = billing
|
||||
return env
|
||||
|
||||
|
||||
def _onboarding_request_message(env_name: str) -> str:
|
||||
"""P19 (REQ-183): the rebranded Nova onboarding message — self-service
|
||||
request path, no longer routes to 'contact the platform team'."""
|
||||
return (
|
||||
"=== Nova Environment Onboarding ===\n"
|
||||
f"No environment named '{env_name}' is bound to this repository.\n\n"
|
||||
"Nova environments are platform-managed. The platform provisions on\n"
|
||||
"your behalf:\n"
|
||||
" - an AWS account (or a scoped partition of one)\n"
|
||||
" - a network (VPC + subnets)\n"
|
||||
" - a state backend (an S3 bucket + DynamoDB lock table)\n"
|
||||
" - an IAM role surfaced to your repo via attribute-based\n"
|
||||
" authorization (ABAC)\n\n"
|
||||
"You do not provide an AWS account, VPC, subnet, or state bucket.\n\n"
|
||||
"To request an environment (self-service):\n"
|
||||
" 1. Submit an onboarding request to the Nova Lambda\n"
|
||||
" (action: onboard_consumer) with your repo name + the\n"
|
||||
" environment name you need (e.g. 'dev').\n"
|
||||
" 2. The platform generates an environment binding + opens a PR.\n"
|
||||
" 3. The platform provisions the account/network/state/role and\n"
|
||||
" grants the ABAC role. Your next pipeline run proceeds.\n\n"
|
||||
"Run: python3 core/onboarding.py --request '{...}' to generate a\n"
|
||||
"binding file locally, or POST to the Lambda onboard_consumer action.\n"
|
||||
"===================================\n"
|
||||
)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
parser = argparse.ArgumentParser(description="Generate an env binding from an onboarding request.")
|
||||
group = parser.add_mutually_exclusive_group(required=True)
|
||||
group.add_argument("request_file", nargs="?", help="path to a request JSON file")
|
||||
group.add_argument("--request", help="inline request JSON string")
|
||||
parser.add_argument("--out", help="output path for the generated env JSON (default: stdout)")
|
||||
parser.add_argument("--template-env", default="dev", help="template environment (default: dev)")
|
||||
args = parser.parse_args(argv)
|
||||
|
||||
if args.request:
|
||||
request = json.loads(args.request)
|
||||
else:
|
||||
request = json.loads(Path(args.request_file).read_text())
|
||||
|
||||
env = generate_env_file(request, template_env=args.template_env)
|
||||
env_json = json.dumps(env, indent=2) + "\n"
|
||||
if args.out:
|
||||
Path(args.out).write_text(env_json)
|
||||
print(f"wrote: {args.out}")
|
||||
else:
|
||||
print(env_json)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -38,10 +38,8 @@ from core import env as _envhelper
|
||||
SSM_PREFIX = "/nova"
|
||||
KMS_KEY_ID_ENV = "NOVA_KMS_KEY_ID"
|
||||
|
||||
# P14 (REQ-178): SAFE_OUTPUT_NAMES is schema-driven (derived from
|
||||
# modules/l1/*/interface.json outputs that don't have sensitive:true).
|
||||
# Falls back to the hardcoded set if the interfaces can't be read.
|
||||
_HARDCODED_SAFE_OUTPUTS = {
|
||||
# Outputs that are safe to display in a PR comment (no secrets).
|
||||
SAFE_OUTPUT_NAMES = {
|
||||
"distribution_domain_name",
|
||||
"bucket_arn",
|
||||
"bucket_name",
|
||||
@@ -61,37 +59,6 @@ _HARDCODED_SAFE_OUTPUTS = {
|
||||
}
|
||||
|
||||
|
||||
def _load_safe_output_names():
|
||||
"""Derive the safe-output allowlist from interface.json outputs.
|
||||
|
||||
P14 (REQ-178): scan modules/l1/*/interface.json; an output is safe if
|
||||
its spec does not set sensitive:true. Falls back to the hardcoded set
|
||||
if no interfaces are readable.
|
||||
"""
|
||||
import json
|
||||
from pathlib import Path
|
||||
root = Path(__file__).resolve().parent.parent
|
||||
safe = set()
|
||||
try:
|
||||
for iface in (root / "modules" / "l1").glob("*/interface.json"):
|
||||
d = json.loads(iface.read_text())
|
||||
outs = d.get("outputs", {})
|
||||
if isinstance(outs, dict):
|
||||
for name, spec in outs.items():
|
||||
if not (isinstance(spec, dict) and spec.get("sensitive")):
|
||||
safe.add(name)
|
||||
elif isinstance(outs, list):
|
||||
for out in outs:
|
||||
if isinstance(out, dict) and not out.get("sensitive"):
|
||||
safe.add(out.get("name", ""))
|
||||
except (OSError, ValueError):
|
||||
pass
|
||||
return safe or _HARDCODED_SAFE_OUTPUTS
|
||||
|
||||
|
||||
SAFE_OUTPUT_NAMES = _load_safe_output_names()
|
||||
|
||||
|
||||
def _ssm_client():
|
||||
if boto3 is None:
|
||||
raise RuntimeError("boto3 is required for SSM publishing")
|
||||
|
||||
+16
-40
@@ -74,10 +74,7 @@ class RegressionReport:
|
||||
|
||||
@property
|
||||
def passed(self) -> bool:
|
||||
# G-111: Skipped is the post-teardown steady state (D-096) for the
|
||||
# live-AWS tier caps (CAP-013..016). The gate passes when every
|
||||
# capability is Verified OR Skipped (no Decayed/Broken).
|
||||
return all(r.status in ("Verified", "Skipped") for r in self.results)
|
||||
return all(r.status == "Verified" for r in self.results)
|
||||
|
||||
def to_dict(self) -> dict:
|
||||
return {
|
||||
@@ -358,11 +355,6 @@ def _check_live_terraform_plan(contract_path: str, label: str) -> Tuple[Status,
|
||||
cwd=tf_dir, timeout=120, env=env,
|
||||
)
|
||||
if rc != 0:
|
||||
# G-111: the state bucket was torn down in v1.11 (D-096) and not
|
||||
# re-provisioned. A NoSuchBucket on init is the known post-teardown
|
||||
# steady state → Skipped (not Broken).
|
||||
if "NoSuchBucket" in err or "NoSuchBucket" in out:
|
||||
return "Skipped", f"terraform init: state bucket absent (post-v1.11-teardown, D-096) [{label}]"
|
||||
return "Broken", f"terraform init failed: {err.strip()[-200:]}"
|
||||
rc, out, err = _run_subprocess(
|
||||
["terraform", "validate"], cwd=tf_dir, timeout=60, env=env,
|
||||
@@ -391,16 +383,8 @@ def _check_live_terraform_plan_static_assets() -> Tuple[Status, str]:
|
||||
|
||||
|
||||
def _check_dynamodb_outbox_table() -> Tuple[Status, str]:
|
||||
"""CAP-015: DynamoDB outbox table exists + is describable (live AWS).
|
||||
|
||||
G-111: the live AWS resources were torn down in v1.11 (D-096) and not
|
||||
re-provisioned (v1.15 P4 was plan-only). A ResourceNotFoundException
|
||||
is the known post-teardown steady state → Skipped (not Decayed), so
|
||||
the gate's strict-`all` `passed` doesn't block on a known absence.
|
||||
Re-provisioning is a future feature milestone, not an NFR regression.
|
||||
"""
|
||||
"""CAP-015: DynamoDB outbox table exists + is describable (live AWS)."""
|
||||
import boto3
|
||||
from botocore.exceptions import ClientError
|
||||
env = _load_aws_env()
|
||||
try:
|
||||
dyn = boto3.client("dynamodb", region_name=env.get("AWS_DEFAULT_REGION", "us-east-1"),
|
||||
@@ -409,40 +393,24 @@ def _check_dynamodb_outbox_table() -> Tuple[Status, str]:
|
||||
r = dyn.describe_table(TableName="nova-outbox")
|
||||
count = r["Table"].get("ItemCount", "unknown")
|
||||
return "Verified", f"nova-outbox exists, item_count={count}"
|
||||
except ClientError as e:
|
||||
code = e.response.get("Error", {}).get("Code", "")
|
||||
if code == "ResourceNotFoundException":
|
||||
return "Skipped", "nova-outbox absent (post-v1.11-teardown steady state, D-096)"
|
||||
return "Decayed", f"describe_table failed: {type(e).__name__}: {str(e)[:150]}"
|
||||
except Exception as e:
|
||||
return "Decayed", f"describe_table failed: {type(e).__name__}: {str(e)[:150]}"
|
||||
|
||||
|
||||
def _check_s3_state_bucket() -> Tuple[Status, str]:
|
||||
"""CAP-016: S3 state bucket exists + readable (live AWS).
|
||||
|
||||
G-111: the live state bucket was torn down in v1.11 (D-096) and not
|
||||
re-provisioned. A 404 on head_bucket is the known post-teardown steady
|
||||
state → Skipped (not Decayed). Re-provisioning is a future feature.
|
||||
"""
|
||||
"""CAP-016: S3 state bucket exists + readable (live AWS)."""
|
||||
import boto3
|
||||
from botocore.exceptions import ClientError
|
||||
env = _load_aws_env()
|
||||
account_id = _envhelper.get_env("AWS_ACCOUNT_ID", "581513795199")
|
||||
state_bucket = f"nova-tfstate-{account_id}-us-east-1"
|
||||
try:
|
||||
s3 = boto3.client("s3", region_name=env.get("AWS_DEFAULT_REGION", "us-east-1"),
|
||||
aws_access_key_id=env.get("AWS_ACCESS_KEY_ID"),
|
||||
aws_secret_access_key=env.get("AWS_SECRET_ACCESS_KEY"))
|
||||
account_id = _envhelper.get_env("AWS_ACCOUNT_ID", "581513795199")
|
||||
state_bucket = f"nova-tfstate-{account_id}-us-east-1"
|
||||
s3.head_bucket(Bucket=state_bucket)
|
||||
r = s3.list_objects_v2(Bucket=state_bucket, MaxKeys=5)
|
||||
keys = [o["Key"] for o in r.get("Contents", [])]
|
||||
return "Verified", f"state bucket exists, keys={keys}"
|
||||
except ClientError as e:
|
||||
code = e.response.get("Error", {}).get("Code", "")
|
||||
if code in ("404", "NoSuchBucket", "NotFound"):
|
||||
return "Skipped", f"state bucket {state_bucket} absent (post-v1.11-teardown, D-096)"
|
||||
return "Decayed", f"head_bucket failed: {type(e).__name__}: {str(e)[:150]}"
|
||||
except Exception as e:
|
||||
return "Decayed", f"head_bucket failed: {type(e).__name__}: {str(e)[:150]}"
|
||||
|
||||
@@ -668,9 +636,17 @@ def write_report(report: RegressionReport,
|
||||
|
||||
|
||||
def main() -> int:
|
||||
"""P13 (REQ-177): re-export from core.regression_verify_cli."""
|
||||
from core.regression_verify_cli import main as _cli_main
|
||||
return _cli_main()
|
||||
milestone = _envhelper.get_env("REGRESSION_MILESTONE", "v1.10") or "v1.10"
|
||||
phase = int(_envhelper.get_env("REGRESSION_PHASE", "52") or "52")
|
||||
report = run_regression(milestone=milestone, phase=phase)
|
||||
md, js = write_report(report)
|
||||
print(f"regression: {report.summary} -> {md}")
|
||||
if not report.passed:
|
||||
print("FAIL: regression surfaced non-Verified capabilities "
|
||||
"(milestone gate blocks)", file=sys.stderr)
|
||||
return 1
|
||||
print("regression: all capabilities Verified (milestone gate passes)")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
@@ -1,33 +0,0 @@
|
||||
"""Nova Regression Verify CLI — command-line entry point.
|
||||
|
||||
Extracted from core/regression_verify.py (P13, REQ-177).
|
||||
|
||||
G-113 import direction: this module imports core.regression_verify (the
|
||||
library) for run_regression + write_report. The library does not import
|
||||
this CLI module. Nothing imports this CLI except direct invocation.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
|
||||
from core import env as _envhelper
|
||||
from core.regression_verify import run_regression, write_report
|
||||
|
||||
|
||||
def main(argv=None):
|
||||
"""CLI: run the regression gate and write the report."""
|
||||
milestone = _envhelper.get_env("REGRESSION_MILESTONE", "v1.10") or "v1.10"
|
||||
phase = int(_envhelper.get_env("REGRESSION_PHASE", "52") or "52")
|
||||
report = run_regression(milestone=milestone, phase=phase)
|
||||
md, js = write_report(report)
|
||||
print(f"regression: {report.summary} -> {md}")
|
||||
if not report.passed:
|
||||
print("FAIL: regression surfaced non-Verified/non-Skipped capabilities "
|
||||
"(milestone gate blocks)", file=sys.stderr)
|
||||
return 1
|
||||
print(f"regression: gate passes (summary={report.summary})")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -1,143 +0,0 @@
|
||||
# Nova Metrics Views — PowerBI Data Dictionary
|
||||
|
||||
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (REQ-190, REQ-209)
|
||||
> Generated: 2026-08-04
|
||||
|
||||
This document is the column-level data dictionary for the PowerBI export
|
||||
views in `metrics/powerbi/`. Each fact/dimension table and placeholder
|
||||
view is documented with: column, type, source/formula, unit, and
|
||||
grounded/derived/deferred status.
|
||||
|
||||
## Fact tables (grounded)
|
||||
|
||||
### fact_run
|
||||
| Column | Type | Source | Unit | Status |
|
||||
|--------|------|--------|------|--------|
|
||||
| run_id | TEXT | run_manifest.py | — | grounded |
|
||||
| contract_id | TEXT | run_manifest.py | — | grounded |
|
||||
| environment | TEXT | run_manifest.py | dev/qa/prod/dr | grounded |
|
||||
| started_at | TEXT | run_manifest.py | ISO8601 | grounded |
|
||||
| completed_at | TEXT | run_manifest.py | ISO8601 | grounded |
|
||||
| exit_code | INTEGER | run_manifest.py | — | grounded |
|
||||
| outcome | TEXT | run_manifest.py | succeeded/failed | grounded |
|
||||
| confidence_score | REAL | confidence_signal.py | 0.0–1.0 | grounded |
|
||||
| confidence_band | TEXT | confidence_signal.py | pass/warn/block | grounded |
|
||||
| hitl_block | INTEGER | hitl_gates.py | 0/1 | grounded |
|
||||
| cost_estimate_usd | REAL | infracost_adapter.py | USD | grounded (Infracost) |
|
||||
| decision_id | TEXT | decision_ledger.py | — | grounded |
|
||||
|
||||
### fact_capability
|
||||
| Column | Type | Source | Unit | Status |
|
||||
|--------|------|--------|------|--------|
|
||||
| capability_id | TEXT | REGRESSION_REPORT.json | CAP-NNN | grounded |
|
||||
| run_id | TEXT | REGRESSION_REPORT.json | — | grounded |
|
||||
| name | TEXT | REGRESSION_REPORT.json | — | grounded |
|
||||
| status | TEXT | REGRESSION_REPORT.json | Verified/Decayed/Broken/Skipped | grounded |
|
||||
| tier | TEXT | REGRESSION_REPORT.json | local/live-aws/lifecycle-pipeline | grounded |
|
||||
| duration_ms | REAL | REGRESSION_REPORT.json | milliseconds | grounded |
|
||||
| detail | TEXT | REGRESSION_REPORT.json | — | grounded |
|
||||
| run_at_utc | TEXT | REGRESSION_REPORT.json | ISO8601 | grounded |
|
||||
|
||||
### fact_decision
|
||||
| Column | Type | Source | Unit | Status |
|
||||
|--------|------|--------|------|--------|
|
||||
| decision_id | TEXT | decision_ledger.py | = run_id | grounded |
|
||||
| run_id | TEXT | decision_ledger.py | — | grounded |
|
||||
| chosen_action | TEXT | confidence_signal.py | pass/warn/block | grounded |
|
||||
| confidence | REAL | confidence_signal.py | 0.0–1.0 | grounded |
|
||||
| alternatives | TEXT (JSON) | confidence_signal.py | perInput breakdown | grounded |
|
||||
| human_override | INTEGER | hitl_gates.py | 0/1 | grounded |
|
||||
| outcome | TEXT | decision_ledger.py | succeeded/failed/pending | grounded |
|
||||
| event_time | TEXT | decision_ledger.py | ISO8601 | grounded |
|
||||
|
||||
### fact_test
|
||||
| Column | Type | Source | Unit | Status |
|
||||
|--------|------|--------|------|--------|
|
||||
| run_id | TEXT | junit XML | — | grounded |
|
||||
| total_tests | INTEGER | junit XML | count | grounded |
|
||||
| passed | INTEGER | junit XML | count | grounded |
|
||||
| failed | INTEGER | junit XML | count | grounded |
|
||||
| errors | INTEGER | junit XML | count | grounded |
|
||||
| skipped | INTEGER | junit XML | count | grounded |
|
||||
| duration_s | REAL | junit XML | seconds | grounded |
|
||||
| coverage_pct | REAL | coverage.json | % | grounded |
|
||||
| collected_at | TEXT | collector.py | ISO8601 | grounded |
|
||||
|
||||
### fact_cost_estimate
|
||||
| Column | Type | Source | Unit | Status |
|
||||
|--------|------|--------|------|--------|
|
||||
| run_id | TEXT | infracost_adapter.py | — | grounded |
|
||||
| delta_usd | REAL | Infracost | USD/month | grounded (pre-apply) |
|
||||
| total_monthly_usd | REAL | Infracost | USD/month | grounded (pre-apply) |
|
||||
| available | INTEGER | infracost_adapter.py | 0/1 | grounded |
|
||||
| estimated_at | TEXT | infracost_adapter.py | ISO8601 | grounded |
|
||||
|
||||
### fact_lifecycle
|
||||
| Column | Type | Source | Unit | Status |
|
||||
|--------|------|--------|------|--------|
|
||||
| module | TEXT | lifecycle report | — | grounded |
|
||||
| environment | TEXT | lifecycle report | — | grounded |
|
||||
| phase | TEXT | lifecycle report | apply/modify/destroy | grounded |
|
||||
| result | TEXT | lifecycle report | pass/fail | grounded |
|
||||
| duration_ms | REAL | lifecycle report | milliseconds | grounded |
|
||||
| run_at | TEXT | lifecycle report | ISO8601 | grounded |
|
||||
|
||||
## Dimension tables
|
||||
|
||||
### dim_capability
|
||||
| Column | Type | Source | Status |
|
||||
|--------|------|--------|--------|
|
||||
| capability_id | TEXT | REGRESSION_REPORT.json | grounded |
|
||||
| name | TEXT | REGRESSION_REPORT.json | grounded |
|
||||
| tier | TEXT | REGRESSION_REPORT.json | grounded |
|
||||
| source_milestone | TEXT | REGRESSION_REPORT.json | grounded |
|
||||
|
||||
### dim_milestone
|
||||
| Column | Type | Source | Status |
|
||||
|--------|------|--------|--------|
|
||||
| milestone | TEXT | REGRESSION_REPORT.json | grounded |
|
||||
| phase | INTEGER | REGRESSION_REPORT.json | grounded |
|
||||
| tag | TEXT | — | grounded |
|
||||
| completed_at | TEXT | REGRESSION_REPORT.json | grounded |
|
||||
|
||||
## Placeholder views (deferred — 8 views, headers only, no data)
|
||||
|
||||
### placeholder_live_infra_health
|
||||
- **Blocking decision:** D-096
|
||||
- **Description:** Live infrastructure health (ECS running count, ALB 5xx, RPS)
|
||||
- **Columns:** timestamp, resource_id, resource_type, running_count, healthy, downtime_seconds
|
||||
|
||||
### placeholder_live_outbox_rate
|
||||
- **Blocking decision:** D-096
|
||||
- **Description:** Live outbox write rate / ledger append latency
|
||||
- **Columns:** timestamp, contract_id, write_latency_ms, append_count
|
||||
|
||||
### placeholder_tamper_evident_checkpoints
|
||||
- **Blocking decision:** D-083
|
||||
- **Description:** Tamper-evident ledger checkpoints / JWS signature rate
|
||||
- **Columns:** timestamp, checkpoint_id, jws_signed, object_lock_enabled
|
||||
|
||||
### placeholder_onboarding_funnel
|
||||
- **Blocking decision:** D-113/D-114/D-119
|
||||
- **Description:** Onboarding funnel: requested → granted conversion
|
||||
- **Columns:** timestamp, consumer_repo, requested_environment, status, granted_at
|
||||
|
||||
### placeholder_drift_detection
|
||||
- **Blocking decision:** D-096 + no scheduler
|
||||
- **Description:** Drift detection (scheduled terraform plan -detailed-exitcode)
|
||||
- **Columns:** timestamp, workspace_id, drift_count, auto_reverted, detection_cycle
|
||||
|
||||
### placeholder_live_cur_reconciliation
|
||||
- **Blocking decision:** D-096
|
||||
- **Description:** Live cost CUR reconciliation
|
||||
- **Columns:** timestamp, resource_address, actual_usd, baseline_usd, saved_usd
|
||||
|
||||
### placeholder_sla_downtime
|
||||
- **Blocking decision:** D-096
|
||||
- **Description:** SLA / unplanned downtime
|
||||
- **Columns:** timestamp, service, uptime_pct, downtime_minutes, slo_target
|
||||
|
||||
### placeholder_predictive_reactive
|
||||
- **Blocking decision:** future emitter
|
||||
- **Description:** Predictive vs Reactive ratio
|
||||
- **Columns:** timestamp, action_id, label, trigger, count
|
||||
@@ -1,87 +0,0 @@
|
||||
# Nova Onboarding — No-Humans Request Path (v1.16, REQ-182..184)
|
||||
|
||||
The v1.16 milestone implements the **request path** of the no-humans
|
||||
onboarding flow (D-113). A consumer can submit an onboarding request
|
||||
without contacting the platform team; the platform generates an
|
||||
environment binding + (in a future milestone) provisions the AWS resources.
|
||||
|
||||
## The 3-step request path
|
||||
|
||||
### Step 1 — Submit an onboarding request (P18, REQ-182)
|
||||
|
||||
A consumer submits an onboarding request to the Nova platform Lambda:
|
||||
|
||||
```bash
|
||||
# Via the Lambda Function URL (IAM auth):
|
||||
curl -X POST "$NOVA_LAMBDA_URL" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"action": "onboard_consumer",
|
||||
"consumerRepo": "acdl/my-app",
|
||||
"requestedEnvironment": "dev",
|
||||
"ownerId": "team-x",
|
||||
"billingTag": "cost-center-x"
|
||||
}'
|
||||
```
|
||||
|
||||
The Lambda validates the payload against
|
||||
[`schemas/onboarding.schema.json`](../schemas/onboarding.schema.json),
|
||||
then writes a `pending` row to the `nova-contracts` DynamoDB table
|
||||
(D-119). No AWS resources are created by this action (D-113).
|
||||
|
||||
### Step 2 — Generate an environment binding (P19, REQ-183)
|
||||
|
||||
The platform (or the consumer locally) generates an environment binding
|
||||
file from the request:
|
||||
|
||||
```bash
|
||||
python3 core/onboarding.py --request '{
|
||||
"consumerRepo": "acdl/my-app",
|
||||
"requestedEnvironment": "qa",
|
||||
"ownerId": "team-x",
|
||||
"billingTag": "cost-center-x"
|
||||
}' --out core/environments/qa.json
|
||||
```
|
||||
|
||||
This produces a `<env>.json` from the `dev.json` template, filling in
|
||||
the `ownerId` + `billingTag` + a description. The `account_id` is a
|
||||
placeholder (`000000000000`) for the platform team to fill with the real
|
||||
account. The generated file validates against
|
||||
[`schemas/environment.schema.json`](../schemas/environment.schema.json).
|
||||
|
||||
### Step 3 — Cross-account role + ABAC tag grant (P20, REQ-184)
|
||||
|
||||
The platform authors the consumer deploy-role + `nova:owner` ABAC tag
|
||||
grant via Terraform:
|
||||
|
||||
```bash
|
||||
cd terraform/onboarding
|
||||
terraform init -backend=false
|
||||
terraform validate
|
||||
NOVA_AWS_ACCOUNT_ID=123456789012 terraform plan \
|
||||
-var consumer_repo=acdl/my-app \
|
||||
-var owner_id=team-x
|
||||
```
|
||||
|
||||
**Offline-proven only (D-114):** `terraform validate` + `terraform plan`
|
||||
pass; **no live apply** in v1.16. The live apply (creating the real
|
||||
cross-account role + OIDC trust) is deferred to a future feature
|
||||
milestone (D-113).
|
||||
|
||||
## What is NOT automated (deferred)
|
||||
|
||||
- **Real AWS account/network/state provisioning** — the request path
|
||||
generates a binding file with a placeholder `account_id`; the actual
|
||||
AWS account creation + VPC + state backend is a future feature (D-113).
|
||||
- **Live cross-account role apply** — the Terraform is offline-proven
|
||||
only (D-114); live apply is deferred.
|
||||
- **OIDC trust policy** — the onboarding Terraform uses a placeholder
|
||||
OIDC provider; real OIDC federation is blocked on
|
||||
go-gitea/gitea#36988 (carries forward from v1.1).
|
||||
|
||||
## See also
|
||||
|
||||
- [`schemas/onboarding.schema.json`](../schemas/onboarding.schema.json) — the request schema
|
||||
- [`core/onboarding.py`](../core/onboarding.py) — the env-file generator
|
||||
- [`terraform/onboarding/`](../terraform/onboarding/) — the role-grant Terraform
|
||||
- [`core/environments/README.md`](../core/environments/README.md) — environment binding docs
|
||||
@@ -1,51 +0,0 @@
|
||||
# Nova Metrics Directory
|
||||
|
||||
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (D-128)
|
||||
|
||||
This directory holds Nova's telemetry/observability artifacts. The
|
||||
metrics layer is **Nova-native** (D-120): JSONL event log + SQLite cold
|
||||
store + hash-chained Decision Ledger. No Kafka, Prometheus, ClickHouse,
|
||||
or QLDB.
|
||||
|
||||
## Artifact inventory
|
||||
|
||||
| Artifact | Type | Regenerable? | Description |
|
||||
|----------|------|-------------|-------------|
|
||||
| `events.jsonl` | Append-only event log | No (append-only state) | CloudEvents 1.0 envelopes from all emitters (REQ-187) |
|
||||
| `decision_ledger.db` | SQLite append-only hash-chain | No (append-only state) | Decision Ledger: `ai.decision.made` + `attestation.recorded` events (REQ-188, D-121) |
|
||||
| `nova_metrics.db` | SQLite cold store | Yes (regenerate via collector) | Normalized fact/dimension tables (REQ-189, P2) |
|
||||
| `runs/<run_id>.json` | Per-run manifest | Yes (regenerate from events) | Run lifecycle: stages, durations, exit, confidence, HITL (REQ-187) |
|
||||
| `runs/<run_id>/` | Durable run artifacts | Yes (regenerate from $WORK) | Persisted copies of pcr.json, signal.json, event.json, etc. (REQ-187) |
|
||||
| `lifecycle/<module>-<env>.json` | Lifecycle report | Yes (regenerate from lifecycle runs) | Per-module apply/modify/destroy results (REQ-205) |
|
||||
| `test-results.xml` | JUnit XML | Yes (regenerate via pytest) | Test results (REQ-187, P1 addopts) |
|
||||
| `test-report.json` | JSON test report | Yes (regenerate via pytest) | Test results in JSON (REQ-187, P1 addopts) |
|
||||
| `coverage.json` | Coverage report | Yes (regenerate via pytest) | Code coverage (REQ-206, P1 addopts) |
|
||||
| `powerbi/` | PowerBI export | Yes (regenerate via powerbi_export) | CSV/JSON views for PowerBI ingestion (REQ-190, P3) |
|
||||
| `TRUST_SNAPSHOT.md` | Trust snapshot report | Yes (regenerate via trust_snapshot) | 5 trust metrics + chain-integrity verdict (REQ-211, P4) |
|
||||
|
||||
## Backup + restore
|
||||
|
||||
**Append-only state** (`events.jsonl`, `decision_ledger.db`): these are
|
||||
the source of truth. They should be committed to git (events.jsonl) or
|
||||
snapshotted (decision_ledger.db). If lost, they CANNOT be regenerated —
|
||||
the events they captured are gone.
|
||||
|
||||
**Regenerable artifacts** (`nova_metrics.db`, `runs/`, `lifecycle/`,
|
||||
`test-results.xml`, `coverage.json`, `powerbi/`): these are derived from
|
||||
the append-only state + the source signals (REGRESSION_REPORT.json,
|
||||
$WORK/*.json, junit XML). If lost, re-run the collector
|
||||
(`core/metrics/collector.py`, P2) to rebuild `nova_metrics.db`, then
|
||||
re-run the PowerBI export (`core/metrics/powerbi_export.py`, P3) to
|
||||
rebuild `powerbi/`.
|
||||
|
||||
**Restore procedure:**
|
||||
1. Recover `events.jsonl` + `decision_ledger.db` from git/snapshot.
|
||||
2. `python3 core/metrics/collector.py` → rebuilds `nova_metrics.db`.
|
||||
3. `python3 core/metrics/powerbi_export.py` → rebuilds `powerbi/`.
|
||||
4. `python3 core/metrics/trust_snapshot.py` → rebuilds `TRUST_SNAPSHOT.md`.
|
||||
|
||||
## Concurrency model
|
||||
|
||||
Single-writer per run: the run manifest writer is the only writer per
|
||||
run. SQLite WAL mode + `BEGIN IMMEDIATE` prevents concurrent-write
|
||||
corruption on the Decision Ledger (P1 risk mitigation).
|
||||
@@ -1,66 +0,0 @@
|
||||
# Nova PowerBI Dashboard — Import Guide
|
||||
|
||||
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (REQ-208)
|
||||
> Generated: 2026-08-04
|
||||
|
||||
This guide documents how to import Nova's metrics views into PowerBI
|
||||
via the folder connector, and suggests a starter visual model.
|
||||
|
||||
## Import via folder connector
|
||||
|
||||
1. Open PowerBI Desktop.
|
||||
2. **Get Data** → **Folder** → navigate to `metrics/powerbi/`.
|
||||
3. PowerBI discovers all CSV/JSON files in the folder.
|
||||
4. Combine the files — PowerBI creates a single query per file.
|
||||
|
||||
## Starter visual model
|
||||
|
||||
### Suggested joins
|
||||
- `fact_run` ←→ `fact_decision` on `run_id` (run-level decision path)
|
||||
- `fact_run` ←→ `fact_cost_estimate` on `run_id` (run-level cost)
|
||||
- `fact_capability` ←→ `dim_capability` on `capability_id` (capability lookup)
|
||||
- `fact_capability` ←→ `dim_milestone` on `milestone` (milestone lookup)
|
||||
|
||||
### Suggested visuals (4 starter visuals)
|
||||
|
||||
1. **Capability Health over Time** — bar chart: `fact_capability.status`
|
||||
grouped by `run_at_utc`. Shows Verified/Skipped/Broken/Decayed trend.
|
||||
Source: `fact_capability.csv`.
|
||||
|
||||
2. **Confidence Distribution** — histogram: `fact_confidence.score`.
|
||||
Shows the distribution of confidence scores across all runs.
|
||||
Source: `fact_confidence.csv`.
|
||||
|
||||
3. **Decision Accuracy** — KPI card: count of `fact_decision` where
|
||||
`outcome = 'succeeded'` ÷ total `fact_decision` rows. Shows AI
|
||||
Decision Accuracy (NORTH_STAR target ≥99.5%).
|
||||
Source: `fact_decision.csv`.
|
||||
|
||||
4. **Cost Trend** — line chart: `fact_cost_estimate.delta_usd` over
|
||||
`estimated_at`. Shows pre-apply cost estimate trend (Infracost).
|
||||
Source: `fact_cost_estimate.csv`.
|
||||
|
||||
## Placeholder views (deferred metrics)
|
||||
|
||||
The 8 `placeholder_*.csv` files contain headers only (no data rows).
|
||||
Each has a companion `placeholder_*.json` with the schema metadata
|
||||
(columns, blocking decision, description). When the blocking decision
|
||||
lifts (e.g., D-096 for live AWS), the collector will populate these
|
||||
views and PowerBI will automatically pick up the data.
|
||||
|
||||
## Data refresh
|
||||
|
||||
The export is regenerated by running:
|
||||
```bash
|
||||
python3 core/metrics/collector.py # rebuilds nova_metrics.db
|
||||
python3 core/metrics/powerbi_export.py # exports to metrics/powerbi/
|
||||
```
|
||||
|
||||
In PowerBI, click **Refresh** to pick up the updated CSV/JSON files.
|
||||
|
||||
## Honesty model
|
||||
|
||||
Every metric in the export is `grounded` (cites a source file), `derived`
|
||||
(documented formula), or `deferred` (cites a blocking decision ID). See
|
||||
`docs/METRICS.md` (P4) for the canonical catalog and `docs/METRICS_VIEWS.md`
|
||||
for the column-level data dictionary.
|
||||
+1
-2
@@ -13,7 +13,6 @@ dependencies = [
|
||||
test = [
|
||||
"pytest>=8.0",
|
||||
"pytest-cov>=4.0",
|
||||
"pytest-json-report>=1.5",
|
||||
"moto[dynamodb]>=5.0",
|
||||
]
|
||||
|
||||
@@ -23,7 +22,7 @@ markers = [
|
||||
"offline: tests that run without AWS/Checkov/DynamoDB",
|
||||
"slow: tests that invoke the full platform pipeline (long-running)",
|
||||
]
|
||||
addopts = "-v --tb=short --junitxml=metrics/test-results.xml --json-report --cov=core --cov=adapters --cov-report=json:metrics/coverage.json --json-report-file=metrics/test-report.json"
|
||||
addopts = "-v --tb=short"
|
||||
filterwarnings = [
|
||||
"ignore::DeprecationWarning:botocore.*",
|
||||
]
|
||||
|
||||
@@ -1,36 +0,0 @@
|
||||
{
|
||||
"$schema": "http://json-schema.org/draft-07/schema#",
|
||||
"title": "Nova CloudEvents 1.0 Envelope",
|
||||
"description": "CloudEvents 1.0 envelope with Nova platform.* semantic conventions. Used for all metrics events (REQ-187).",
|
||||
"type": "object",
|
||||
"required": ["specversion", "id", "source", "type", "time", "datacontenttype", "platform", "data"],
|
||||
"properties": {
|
||||
"specversion": {"type": "string", "const": "1.0"},
|
||||
"id": {"type": "string", "minLength": 1},
|
||||
"source": {"type": "string", "minLength": 1},
|
||||
"type": {"type": "string", "minLength": 1, "pattern": "^nova\\."},
|
||||
"time": {"type": "string", "format": "date-time"},
|
||||
"subject": {"type": "string"},
|
||||
"datacontenttype": {"type": "string", "const": "application/json"},
|
||||
"platform": {
|
||||
"type": "object",
|
||||
"required": ["run_id", "environment"],
|
||||
"properties": {
|
||||
"tenant_id": {"type": "string"},
|
||||
"run_id": {"type": "string", "minLength": 1},
|
||||
"contract_id": {"type": "string"},
|
||||
"environment": {"type": "string", "enum": ["dev", "qa", "prod", "dr"]},
|
||||
"actor": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"type": {"type": "string"},
|
||||
"id": {"type": "string"}
|
||||
}
|
||||
},
|
||||
"trace_id": {"type": "string"}
|
||||
}
|
||||
},
|
||||
"data": {"type": "object"}
|
||||
},
|
||||
"additionalProperties": true
|
||||
}
|
||||
@@ -1,56 +0,0 @@
|
||||
{
|
||||
"$schema": "http://json-schema.org/draft-07/schema#",
|
||||
"title": "Nova Per-Run Manifest",
|
||||
"description": "Per-run manifest written to metrics/runs/<run_id>.json (REQ-187). Captures the full run lifecycle.",
|
||||
"type": "object",
|
||||
"required": ["run_id", "contract_id", "environment", "started_at", "completed_at", "exit_code", "stages"],
|
||||
"properties": {
|
||||
"run_id": {"type": "string", "minLength": 1},
|
||||
"contract_id": {"type": "string"},
|
||||
"environment": {"type": "string", "enum": ["dev", "qa", "prod", "dr"]},
|
||||
"started_at": {"type": "string", "format": "date-time"},
|
||||
"completed_at": {"type": "string", "format": "date-time"},
|
||||
"exit_code": {"type": "integer"},
|
||||
"stages": {
|
||||
"type": "array",
|
||||
"items": {
|
||||
"type": "object",
|
||||
"required": ["name", "duration_ms", "exit_code"],
|
||||
"properties": {
|
||||
"name": {"type": "string"},
|
||||
"duration_ms": {"type": "number"},
|
||||
"exit_code": {"type": "integer"},
|
||||
"error": {"type": "string"}
|
||||
}
|
||||
}
|
||||
},
|
||||
"confidence": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"score": {"type": "number"},
|
||||
"band": {"type": "string", "enum": ["pass", "warn", "block"]},
|
||||
"perInput": {"type": "object"}
|
||||
}
|
||||
},
|
||||
"hitl": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"gate": {"type": "string"},
|
||||
"result": {"type": "string"},
|
||||
"block": {"type": "boolean"}
|
||||
}
|
||||
},
|
||||
"policy": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"passed": {"type": "integer"},
|
||||
"failed": {"type": "integer"},
|
||||
"skipped": {"type": "integer"}
|
||||
}
|
||||
},
|
||||
"cost_estimate_usd": {"type": "number"},
|
||||
"decision_id": {"type": "string"},
|
||||
"outcome": {"type": "string", "enum": ["succeeded", "failed", "pending"]}
|
||||
},
|
||||
"additionalProperties": true
|
||||
}
|
||||
@@ -1,39 +0,0 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"$id": "https://nova.cloudinit.dev/schemas/onboarding.schema.json",
|
||||
"title": "Nova Consumer Onboarding Request",
|
||||
"description": "A self-service onboarding request from a consumer repo. Submitted to the contract_ingestor Lambda 'onboard_consumer' action (D-113, P18/REQ-182). The Lambda validates the payload against this schema, then writes a 'pending' CMDB row to nova-contracts. No AWS resources are created by this action (D-119); the cross-account role + ABAC tag grant is offline-proven Terraform (P20/REQ-184).",
|
||||
"type": "object",
|
||||
"required": ["consumerRepo", "requestedEnvironment", "ownerId", "billingTag"],
|
||||
"additionalProperties": false,
|
||||
"properties": {
|
||||
"consumerRepo": {
|
||||
"type": "string",
|
||||
"description": "The consumer repository in org/repo format.",
|
||||
"pattern": "^[a-zA-Z0-9_.-]+/[a-zA-Z0-9_.-]+$",
|
||||
"maxLength": 128
|
||||
},
|
||||
"requestedEnvironment": {
|
||||
"type": "string",
|
||||
"description": "The environment the consumer requests (must exist as a core/environments/<name>.json).",
|
||||
"enum": ["dev", "qa", "prod", "dr"]
|
||||
},
|
||||
"ownerId": {
|
||||
"type": "string",
|
||||
"description": "The owning team or individual (for ABAC nova:owner tag + CMDB).",
|
||||
"minLength": 1,
|
||||
"maxLength": 64
|
||||
},
|
||||
"billingTag": {
|
||||
"type": "string",
|
||||
"description": "The cost-center / billing tag for the consumer's resources.",
|
||||
"minLength": 1,
|
||||
"maxLength": 64
|
||||
},
|
||||
"notes": {
|
||||
"type": "string",
|
||||
"description": "Optional free-form notes for the platform team.",
|
||||
"maxLength": 500
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,60 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# scripts/run_decommission.sh — decommission mode (extracted from run_platform.sh, P9/REQ-173).
|
||||
# Sourced by run_platform.sh (G-112: source, not invoke — shares CONTRACT/WORK/ROOT env).
|
||||
# Exits 0 on completion; caller exits after sourcing.
|
||||
|
||||
echo ""
|
||||
echo "=== Decommission Step 1: validate change request against CMDB ==="
|
||||
[ -n "$CHANGE_REQUEST_ID" ] || fail "change request ID required for decommission mode"
|
||||
CONSUMER_REPO="${GITHUB_REPOSITORY:-$(python3 -c "import yaml; c=yaml.safe_load(open('$CONTRACT')); print(c.get('id','unknown'))" 2>/dev/null || echo 'unknown')}"
|
||||
python3 -c "
|
||||
import json, sys
|
||||
sys.path.insert(0, '$ROOT')
|
||||
# In a real deployment, this invokes the Lambda. For local/CI, we simulate.
|
||||
cr_id = '$CHANGE_REQUEST_ID'
|
||||
repo = '$CONSUMER_REPO'
|
||||
print(f'validate_change_request: crId={cr_id} repo={repo}')
|
||||
# The Lambda action would be:
|
||||
# payload = {'action': 'validate_change_request', 'changeRequestId': cr_id, 'consumerRepo': repo}
|
||||
# result = invoke_lambda(payload)
|
||||
# For now, just print the intent (the actual validation happens via the Lambda in CI/prod)
|
||||
print('change request validation: PASS (simulated for local mode)')
|
||||
"
|
||||
echo ""
|
||||
echo "=== Decommission Step 2: disable deletion protection (HITL SRE gate) ==="
|
||||
echo "This step requires SRE approval via GitHub environment 'decommission-gate-sre'."
|
||||
echo "The contract is resolved with deletion_protection=false injected."
|
||||
python3 core/contract_resolver.py "$CONTRACT" "$WORK/stack.json" 2>/dev/null || fail "resolver failed"
|
||||
python3 -c "
|
||||
import json, sys
|
||||
sys.path.insert(0, '$ROOT')
|
||||
from core.contract_resolver import resolve, decommission_transform
|
||||
stack = resolve('$CONTRACT', '$ROOT')
|
||||
# Step 2: disable deletion protection only (counts still as-is)
|
||||
for res in stack['resources']:
|
||||
if 'nfrs' not in res:
|
||||
res['nfrs'] = {}
|
||||
res['nfrs']['deletion_protection'] = False
|
||||
with open('$WORK/stack-decommission-step1.json', 'w') as f:
|
||||
json.dump(stack, f, indent=2)
|
||||
print(f'decommission step 1: {len(stack[\"resources\"])} resources with deletion_protection=false')
|
||||
"
|
||||
echo ""
|
||||
echo "=== Decommission Step 3: zero counts (HITL SRE gate) ==="
|
||||
echo "This step requires a second SRE approval via GitHub environment 'decommission-destroy-sre'."
|
||||
python3 -c "
|
||||
import json, sys
|
||||
sys.path.insert(0, '$ROOT')
|
||||
from core.contract_resolver import resolve, decommission_transform
|
||||
stack = resolve('$CONTRACT', '$ROOT')
|
||||
stack = decommission_transform(stack)
|
||||
with open('$WORK/stack-decommission-step2.json', 'w') as f:
|
||||
json.dump(stack, f, indent=2)
|
||||
zeroed = sum(1 for r in stack['resources'] if r.get('nfrs',{}).get('deletion_protection') is False)
|
||||
print(f'decommission step 2: {zeroed} resources with deletion_protection=false + counts=0')
|
||||
"
|
||||
echo ""
|
||||
echo "=== Decommission Step 4: confirm ==="
|
||||
echo "The terraform apply for step 2 + step 3 would now destroy all resources."
|
||||
echo "=== DECOMMISSION READY ==="
|
||||
exit 0
|
||||
+144
-48
@@ -7,9 +7,6 @@
|
||||
# run_platform.sh --plan-only <contract.yml> (AWS plan only, no Checkov/outbox)
|
||||
# run_platform.sh --apply <contract.yml> (AWS apply: init/validate/plan/apply)
|
||||
# run_platform.sh --destroy <contract.yml> (AWS destroy: init/validate/destroy)
|
||||
# run_platform.sh --local [contract.yml] (local emulating tier, no AWS)
|
||||
# run_platform.sh --decommission <CR> <contract.yml> (gated teardown)
|
||||
# run_platform.sh --help (show all flags)
|
||||
#
|
||||
# Modes:
|
||||
# --check-only (offline, no AWS/Checkov/DynamoDB — for CI)
|
||||
@@ -20,19 +17,13 @@
|
||||
# contract -> resolver -> stack -> adapter -> terraform init/validate/plan/apply -> exit 0
|
||||
# --destroy (requires AWS creds; use --decommission <CR> for gated production teardown)
|
||||
# contract -> resolver -> stack -> adapter -> terraform init/validate/destroy -> exit 0
|
||||
# --local (no AWS creds; local emulating tier D-092)
|
||||
# contract -> resolver -> adapter -> local S3/ECS/outbox/Lambda stubs -> exit 0
|
||||
# (default) (requires AWS creds + Checkov + DynamoDB)
|
||||
# contract -> resolver -> stack -> adapter -> terraform plan -> Checkov ->
|
||||
# confidence -> outbox
|
||||
#
|
||||
# Flags:
|
||||
# --quiet suppress terraform/checkov streaming (output to log only)
|
||||
# --decommission gate --destroy with D-070 two-step CR validation (requires <CR>)
|
||||
# --deploy-uptime deploy the uptime monitoring stack (separate state)
|
||||
# --local run the headline E2E against the local emulating tier (D-092)
|
||||
# --environment <name> override the contract's environment at load time (D-088)
|
||||
# --help, -h show all flags + a one-line description
|
||||
# --quiet suppress terraform/checkov streaming (output to log only)
|
||||
# --decommission gate --destroy with D-070 two-step CR validation (requires <CR>)
|
||||
#
|
||||
# The contract file is a YAML file validated against schemas/contract.schema.json.
|
||||
# The resolver (core/contract_resolver.py) resolves it to a Target Stack
|
||||
@@ -68,36 +59,6 @@ CHANGE_REQUEST_ID=""
|
||||
ENVIRONMENT_OVERRIDE=""
|
||||
CONTRACT=""
|
||||
|
||||
# P15 (REQ-179): --help / -h prints all flags + a one-line description.
|
||||
_print_help() {
|
||||
cat <<'HELP'
|
||||
Nova platform pipeline — run_platform.sh
|
||||
|
||||
Usage:
|
||||
run_platform.sh <contract.yml> (full e2e with AWS)
|
||||
run_platform.sh --check-only [contract.yml] (offline, no AWS)
|
||||
run_platform.sh --plan-only <contract.yml> (AWS plan only)
|
||||
run_platform.sh --apply <contract.yml> (AWS apply)
|
||||
run_platform.sh --destroy <contract.yml> (AWS destroy)
|
||||
run_platform.sh --local [contract.yml] (local emulating tier)
|
||||
run_platform.sh --decommission <CR> <contract.yml> (gated teardown)
|
||||
|
||||
Flags:
|
||||
--check-only Offline validation (no AWS/Checkov/DynamoDB) — for CI
|
||||
--plan-only AWS plan only (requires AWS creds, no Checkov/outbox)
|
||||
--apply AWS apply: init/validate/plan/apply (HITL gate for qa/prod/dr)
|
||||
--destroy AWS destroy: init/validate/destroy
|
||||
--decommission Gate --destroy with D-070 two-step CR validation (requires <CR>)
|
||||
--local Run the headline E2E against the local emulating tier (D-092, no AWS)
|
||||
--quiet Suppress terraform/checkov streaming (log only)
|
||||
--deploy-uptime Deploy the uptime monitoring stack (separate state)
|
||||
--environment <name> Override the contract's environment at load time (D-088)
|
||||
--help, -h Show this help
|
||||
|
||||
The contract file is a YAML file validated against schemas/contract.schema.json.
|
||||
HELP
|
||||
}
|
||||
|
||||
# Parse args; --environment takes a value (either --environment=VALUE or
|
||||
# --environment VALUE). The contract / changeRequestId are the remaining
|
||||
# positional args.
|
||||
@@ -108,7 +69,6 @@ for arg in "$@"; do
|
||||
continue
|
||||
fi
|
||||
case "$arg" in
|
||||
--help|-h) _print_help; exit 0 ;;
|
||||
--check-only) CHECK_ONLY=1 ;;
|
||||
--plan-only) PLAN_ONLY=1 ;;
|
||||
--apply) APPLY_ONLY=1 ;;
|
||||
@@ -248,10 +208,63 @@ jsonschema.validate(contract, schema)
|
||||
print(f'contract: id={contract[\"id\"]} env={contract[\"environment\"]} modules={list(contract.get(\"infrastructure\",{}).keys())}')
|
||||
"
|
||||
|
||||
# Decommission mode: extracted to scripts/run_decommission.sh (P9, REQ-173).
|
||||
# G-112: sourced (shared env) — the block references CONTRACT/WORK/ROOT.
|
||||
# Decommission mode: validate change request, disable deletion protection, zero counts
|
||||
if [ "$DECOMMISSION" = "1" ]; then
|
||||
source "$ROOT/scripts/run_decommission.sh"
|
||||
echo ""
|
||||
echo "=== Decommission Step 1: validate change request against CMDB ==="
|
||||
[ -n "$CHANGE_REQUEST_ID" ] || fail "change request ID required for decommission mode"
|
||||
CONSUMER_REPO="${GITHUB_REPOSITORY:-$(python3 -c "import yaml; c=yaml.safe_load(open('$CONTRACT')); print(c.get('id','unknown'))" 2>/dev/null || echo 'unknown')}"
|
||||
python3 -c "
|
||||
import json, sys
|
||||
sys.path.insert(0, '$ROOT')
|
||||
# In a real deployment, this invokes the Lambda. For local/CI, we simulate.
|
||||
cr_id = '$CHANGE_REQUEST_ID'
|
||||
repo = '$CONSUMER_REPO'
|
||||
print(f'validate_change_request: crId={cr_id} repo={repo}')
|
||||
# The Lambda action would be:
|
||||
# payload = {'action': 'validate_change_request', 'changeRequestId': cr_id, 'consumerRepo': repo}
|
||||
# result = invoke_lambda(payload)
|
||||
# For now, just print the intent (the actual validation happens via the Lambda in CI/prod)
|
||||
print('change request validation: PASS (simulated for local mode)')
|
||||
"
|
||||
echo ""
|
||||
echo "=== Decommission Step 2: disable deletion protection (HITL SRE gate) ==="
|
||||
echo "This step requires SRE approval via GitHub environment 'decommission-gate-sre'."
|
||||
echo "The contract is resolved with deletion_protection=false injected."
|
||||
python3 core/contract_resolver.py "$CONTRACT" "$WORK/stack.json" 2>/dev/null || fail "resolver failed"
|
||||
python3 -c "
|
||||
import json, sys
|
||||
sys.path.insert(0, '$ROOT')
|
||||
from core.contract_resolver import resolve, decommission_transform
|
||||
stack = resolve('$CONTRACT', '$ROOT')
|
||||
# Step 2: disable deletion protection only (counts still as-is)
|
||||
for res in stack['resources']:
|
||||
if 'nfrs' not in res:
|
||||
res['nfrs'] = {}
|
||||
res['nfrs']['deletion_protection'] = False
|
||||
with open('$WORK/stack-decommission-step1.json', 'w') as f:
|
||||
json.dump(stack, f, indent=2)
|
||||
print(f'decommission step 1: {len(stack[\"resources\"])} resources with deletion_protection=false')
|
||||
"
|
||||
echo ""
|
||||
echo "=== Decommission Step 3: zero counts (HITL SRE gate) ==="
|
||||
echo "This step requires a second SRE approval via GitHub environment 'decommission-destroy-sre'."
|
||||
python3 -c "
|
||||
import json, sys
|
||||
sys.path.insert(0, '$ROOT')
|
||||
from core.contract_resolver import resolve, decommission_transform
|
||||
stack = resolve('$CONTRACT', '$ROOT')
|
||||
stack = decommission_transform(stack)
|
||||
with open('$WORK/stack-decommission-step2.json', 'w') as f:
|
||||
json.dump(stack, f, indent=2)
|
||||
zeroed = sum(1 for r in stack['resources'] if r.get('nfrs',{}).get('deletion_protection') is False)
|
||||
print(f'decommission step 2: {zeroed} resources with deletion_protection=false + counts=0')
|
||||
"
|
||||
echo ""
|
||||
echo "=== Decommission Step 4: confirm ==="
|
||||
echo "The terraform apply for step 2 + step 3 would now destroy all resources."
|
||||
echo "=== DECOMMISSION READY ==="
|
||||
exit 0
|
||||
fi
|
||||
|
||||
echo ""
|
||||
@@ -488,9 +501,92 @@ PY
|
||||
fi
|
||||
|
||||
echo ""
|
||||
# Uptime monitoring: extracted to scripts/run_uptime.sh (P9, REQ-173).
|
||||
# G-112: sourced (shared env) — the block references CONTRACT/WORK/DEPLOY_UPTIME.
|
||||
source "$ROOT/scripts/run_uptime.sh"
|
||||
echo "=== Step 9b: deploy uptime monitoring (separate state) ==="
|
||||
# The uptime stack is deployed by default after the L2 module. It uses a
|
||||
# separate terraform state ($WORK/uptime-tf). Endpoints from the L2 outputs
|
||||
# are passed as monitored_endpoints. The feature flag (uptime_enabled,
|
||||
# default true) controls whether this step runs.
|
||||
#
|
||||
# P57 contract shape: uptime_enabled is a per-module input under
|
||||
# infrastructure.<module>.inputs.uptime_enabled (the old top-level
|
||||
# contract.inputs.uptime_enabled was removed). Scan every module's inputs;
|
||||
# any module setting uptime_enabled=false disables the uptime step (one
|
||||
# contract = one logical stack, so a single false wins).
|
||||
if [ "$DEPLOY_UPTIME" = "1" ] || ( [ "$CHECK_ONLY" = "0" ] && [ "$PLAN_ONLY" = "0" ] ); then
|
||||
UPTIME_ENABLED=$(python3 -c "
|
||||
import yaml
|
||||
c = yaml.safe_load(open('$CONTRACT'))
|
||||
infra = c.get('infrastructure', {})
|
||||
# Default true; a module may override to false.
|
||||
for m, entry in infra.items():
|
||||
if isinstance(entry, dict) and entry.get('inputs', {}).get('uptime_enabled') is False:
|
||||
print('False'); break
|
||||
else:
|
||||
print('True')
|
||||
" 2>/dev/null || echo "True")
|
||||
if [ "$UPTIME_ENABLED" = "True" ] || [ "$UPTIME_ENABLED" = "true" ]; then
|
||||
echo "uptime: feature flag enabled — constructing uptime contract"
|
||||
UPTIME_DIR="$WORK/uptime-tf"
|
||||
mkdir -p "$UPTIME_DIR"
|
||||
# Build the uptime stack from the L2 outputs
|
||||
python3 "$ROOT/core/contract_resolver.py" "$CONTRACT" "$WORK/stack.json" 2>/dev/null || true
|
||||
python3 -c "
|
||||
import json, sys, yaml
|
||||
sys.path.insert(0, '$ROOT')
|
||||
from core.contract_resolver import resolve
|
||||
stack = resolve('$CONTRACT', '$ROOT')
|
||||
# Extract HTTP/DNS/TCP endpoints from the stack outputs
|
||||
endpoints = []
|
||||
outputs = stack.get('outputs', {})
|
||||
for name, spec in outputs.items():
|
||||
src_rid = spec.get('from', '')
|
||||
src_output = spec.get('output', name)
|
||||
if 'domain' in name.lower() or 'url' in name.lower() or 'endpoint' in name.lower():
|
||||
endpoints.append({
|
||||
'name': name,
|
||||
'url': f'ref:{src_rid}.{src_output}',
|
||||
'type': 'http',
|
||||
'interval_seconds': 60,
|
||||
'timeout_seconds': 30
|
||||
})
|
||||
# Build the uptime contract
|
||||
uptime_contract = {
|
||||
'id': 'uptime',
|
||||
'name': 'uptime-monitoring',
|
||||
'environment': 'dev',
|
||||
'infrastructure': {
|
||||
'uptime': {
|
||||
'version': '1.0.0',
|
||||
'inputs': {
|
||||
'region': 'us-east-1',
|
||||
'feature_flag_enabled': True,
|
||||
'monitored_endpoints': endpoints,
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
with open('$WORK/uptime-contract.yml', 'w') as f:
|
||||
yaml.dump(uptime_contract, f)
|
||||
print(f'uptime: {len(endpoints)} endpoint(s) to monitor')
|
||||
" 2>/dev/null || echo "uptime: no endpoints found (skipping monitor config)"
|
||||
|
||||
# Resolve + adapt the uptime contract to a separate TF dir
|
||||
python3 "$ROOT/core/contract_resolver.py" "$WORK/uptime-contract.yml" "$WORK/uptime-stack.json" 2>/dev/null || true
|
||||
python3 "$ROOT/adapters/terraform/adapter.py" "$WORK/uptime-stack.json" "$UPTIME_DIR" 2>/dev/null || true
|
||||
|
||||
if [ "$DEPLOY_UPTIME" = "1" ] && [ -f "$UPTIME_DIR/main.tf" ]; then
|
||||
echo "uptime: emitted Terraform to $UPTIME_DIR"
|
||||
if [ "$QUIET" = "0" ]; then
|
||||
echo "--- uptime main.tf ---"
|
||||
cat "$UPTIME_DIR/main.tf"
|
||||
echo "--- end uptime main.tf ---"
|
||||
fi
|
||||
fi
|
||||
echo "uptime: monitoring stack ready (separate state: $UPTIME_DIR)"
|
||||
else
|
||||
echo "uptime: feature flag disabled (inputs.uptime_enabled=false) — skipping"
|
||||
fi
|
||||
fi
|
||||
|
||||
echo ""
|
||||
echo "=== PLATFORM E2E OK ==="
|
||||
|
||||
@@ -1,91 +0,0 @@
|
||||
#!/usr/bin/env bash
|
||||
# scripts/run_uptime.sh — uptime monitoring deploy (extracted from run_platform.sh, P9/REQ-173).
|
||||
# Sourced by run_platform.sh (G-112: source, not invoke — shares CONTRACT/WORK/ROOT env).
|
||||
# Uses: $ROOT, $CONTRACT, $WORK, $DEPLOY_UPTIME, $CHECK_ONLY, $PLAN_ONLY, $QUIET.
|
||||
|
||||
echo "=== Step 9b: deploy uptime monitoring (separate state) ==="
|
||||
# The uptime stack is deployed by default after the L2 module. It uses a
|
||||
# separate terraform state ($WORK/uptime-tf). Endpoints from the L2 outputs
|
||||
# are passed as monitored_endpoints. The feature flag (uptime_enabled,
|
||||
# default true) controls whether this step runs.
|
||||
#
|
||||
# P57 contract shape: uptime_enabled is a per-module input under
|
||||
# infrastructure.<module>.inputs.uptime_enabled (the old top-level
|
||||
# contract.inputs.uptime_enabled was removed). Scan every module's inputs;
|
||||
# any module setting uptime_enabled=false disables the uptime step (one
|
||||
# contract = one logical stack, so a single false wins).
|
||||
if [ "$DEPLOY_UPTIME" = "1" ] || ( [ "$CHECK_ONLY" = "0" ] && [ "$PLAN_ONLY" = "0" ] ); then
|
||||
UPTIME_ENABLED=$(python3 -c "
|
||||
import yaml
|
||||
c = yaml.safe_load(open('$CONTRACT'))
|
||||
infra = c.get('infrastructure', {})
|
||||
# Default true; a module may override to false.
|
||||
for m, entry in infra.items():
|
||||
if isinstance(entry, dict) and entry.get('inputs', {}).get('uptime_enabled') is False:
|
||||
print('False'); break
|
||||
else:
|
||||
print('True')
|
||||
" 2>/dev/null || echo "True")
|
||||
if [ "$UPTIME_ENABLED" = "True" ] || [ "$UPTIME_ENABLED" = "true" ]; then
|
||||
echo "uptime: feature flag enabled — constructing uptime contract"
|
||||
UPTIME_DIR="$WORK/uptime-tf"
|
||||
mkdir -p "$UPTIME_DIR"
|
||||
# Build the uptime stack from the L2 outputs
|
||||
python3 "$ROOT/core/contract_resolver.py" "$CONTRACT" "$WORK/stack.json" 2>/dev/null || true
|
||||
python3 -c "
|
||||
import json, sys, yaml
|
||||
sys.path.insert(0, '$ROOT')
|
||||
from core.contract_resolver import resolve
|
||||
stack = resolve('$CONTRACT', '$ROOT')
|
||||
# Extract HTTP/DNS/TCP endpoints from the stack outputs
|
||||
endpoints = []
|
||||
outputs = stack.get('outputs', {})
|
||||
for name, spec in outputs.items():
|
||||
src_rid = spec.get('from', '')
|
||||
src_output = spec.get('output', name)
|
||||
if 'domain' in name.lower() or 'url' in name.lower() or 'endpoint' in name.lower():
|
||||
endpoints.append({
|
||||
'name': name,
|
||||
'url': f'ref:{src_rid}.{src_output}',
|
||||
'type': 'http',
|
||||
'interval_seconds': 60,
|
||||
'timeout_seconds': 30
|
||||
})
|
||||
# Build the uptime contract
|
||||
uptime_contract = {
|
||||
'id': 'uptime',
|
||||
'name': 'uptime-monitoring',
|
||||
'environment': 'dev',
|
||||
'infrastructure': {
|
||||
'uptime': {
|
||||
'version': '1.0.0',
|
||||
'inputs': {
|
||||
'region': 'us-east-1',
|
||||
'feature_flag_enabled': True,
|
||||
'monitored_endpoints': endpoints,
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
with open('$WORK/uptime-contract.yml', 'w') as f:
|
||||
yaml.dump(uptime_contract, f)
|
||||
print(f'uptime: {len(endpoints)} endpoint(s) to monitor')
|
||||
" 2>/dev/null || echo "uptime: no endpoints found (skipping monitor config)"
|
||||
|
||||
# Resolve + adapt the uptime contract to a separate TF dir
|
||||
python3 "$ROOT/core/contract_resolver.py" "$WORK/uptime-contract.yml" "$WORK/uptime-stack.json" 2>/dev/null || true
|
||||
python3 "$ROOT/adapters/terraform/adapter.py" "$WORK/uptime-stack.json" "$UPTIME_DIR" 2>/dev/null || true
|
||||
|
||||
if [ "$DEPLOY_UPTIME" = "1" ] && [ -f "$UPTIME_DIR/main.tf" ]; then
|
||||
echo "uptime: emitted Terraform to $UPTIME_DIR"
|
||||
if [ "$QUIET" = "0" ]; then
|
||||
echo "--- uptime main.tf ---"
|
||||
cat "$UPTIME_DIR/main.tf"
|
||||
echo "--- end uptime main.tf ---"
|
||||
fi
|
||||
fi
|
||||
echo "uptime: monitoring stack ready (separate state: $UPTIME_DIR)"
|
||||
else
|
||||
echo "uptime: feature flag disabled (inputs.uptime_enabled=false) — skipping"
|
||||
fi
|
||||
fi
|
||||
@@ -1,42 +0,0 @@
|
||||
# terraform/onboarding/ — Consumer deploy-role + ABAC tag grant (P20, REQ-184)
|
||||
|
||||
Offline-proven Terraform for the cross-account consumer deploy-role +
|
||||
`nova:owner` ABAC tag grant. This is the "role grant" half of the
|
||||
no-humans onboarding flow (D-113); the "request" half is P18 (Lambda
|
||||
action) + P19 (env-file autogen).
|
||||
|
||||
## Scope (D-114)
|
||||
|
||||
This Terraform is **offline-proven only** in v1.16:
|
||||
- `terraform validate` passes.
|
||||
- `terraform plan` (with `NOVA_AWS_ACCOUNT_ID` set) produces the expected
|
||||
role + policy.
|
||||
- **No live apply** — `NOVA_LIFECYCLE_MODE=plan` default. Live apply is
|
||||
deferred to a future feature milestone (D-113/D-114).
|
||||
|
||||
## Variables
|
||||
|
||||
| Variable | Description | Default |
|
||||
|----------|-------------|---------|
|
||||
| `consumer_repo` | The consumer repository (org/repo) | `acdl/consumer-a` |
|
||||
| `owner_id` | The owning team (for `nova:owner` tag) | `team-a` |
|
||||
| `account_id` | The consumer's AWS account ID | `000000000000` |
|
||||
| `region` | AWS region | `us-east-1` |
|
||||
|
||||
## Resources
|
||||
|
||||
- `aws_iam_role.consumer_deploy` — the consumer's deploy role with a
|
||||
trust policy (assumed by the consumer's CI runner).
|
||||
- `aws_iam_role_policy.consumer_invoke` — inline policy granting
|
||||
`lambda:InvokeFunctionUrl` on the platform Lambda, scoped via
|
||||
`aws:PrincipalTag/nova:owner == var.owner_id` (ABAC).
|
||||
- `aws_iam_tag.owner` — tags the role with `nova:owner` + `nova:contract`.
|
||||
|
||||
## Usage (offline)
|
||||
|
||||
```bash
|
||||
cd terraform/onboarding
|
||||
terraform init -backend=false
|
||||
terraform validate
|
||||
NOVA_AWS_ACCOUNT_ID=123456789012 terraform plan -var consumer_repo=acdl/my-app -var owner_id=team-x
|
||||
```
|
||||
@@ -1,120 +0,0 @@
|
||||
terraform {
|
||||
required_version = ">= 1.9, < 1.10"
|
||||
required_providers {
|
||||
aws = {
|
||||
source = "hashicorp/aws"
|
||||
version = "~> 5.0"
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
variable "consumer_repo" {
|
||||
description = "The consumer repository (org/repo) — for the nova:contract tag."
|
||||
type = string
|
||||
default = "acdl/consumer-a"
|
||||
}
|
||||
|
||||
variable "owner_id" {
|
||||
description = "The owning team (for the nova:owner ABAC tag)."
|
||||
type = string
|
||||
default = "team-a"
|
||||
}
|
||||
|
||||
variable "account_id" {
|
||||
description = "The consumer's AWS account ID (where the deploy role is created)."
|
||||
type = string
|
||||
default = "000000000000"
|
||||
}
|
||||
|
||||
variable "region" {
|
||||
description = "AWS region."
|
||||
type = string
|
||||
default = "us-east-1"
|
||||
}
|
||||
|
||||
provider "aws" {
|
||||
region = var.region
|
||||
}
|
||||
|
||||
# P20 (REQ-184): consumer deploy role — the role the consumer's CI runner
|
||||
# assumes to invoke the platform Lambda + deploy via the reusable workflow.
|
||||
# The trust policy allows the consumer's CI runner (GitHub Actions /
|
||||
# Gitea act_runner) to assume this role. In a real deployment, the trust
|
||||
# policy is scoped to the consumer's OIDC provider; for offline-proven
|
||||
# mode, a placeholder trust is used.
|
||||
resource "aws_iam_role" "consumer_deploy" {
|
||||
name = "nova-${replace(var.consumer_repo, "/", "-")}-deploy"
|
||||
|
||||
assume_role_policy = jsonencode({
|
||||
Version = "2012-10-17"
|
||||
Statement = [
|
||||
{
|
||||
Effect = "Allow"
|
||||
Principal = {
|
||||
# Placeholder: in a real deployment, this is the consumer's
|
||||
# OIDC provider ARN. Offline-proven mode uses a wildcard.
|
||||
Federated = "arn:aws:iam::${var.account_id}:oidc-provider/token.actions.githubusercontent.com"
|
||||
}
|
||||
Action = "sts:AssumeRoleWithWebIdentity"
|
||||
Condition = {
|
||||
StringEquals = {
|
||||
"token.actions.githubusercontent.com:aud" = "sts.amazonaws.com"
|
||||
}
|
||||
StringLike = {
|
||||
"token.actions.githubusercontent.com:sub" = "repo:${var.consumer_repo}:*"
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
})
|
||||
|
||||
tags = {
|
||||
"nova:owner" = var.owner_id
|
||||
"nova:contract" = var.consumer_repo
|
||||
"nova:environment" = "dev"
|
||||
}
|
||||
}
|
||||
|
||||
# P20 (REQ-184): inline policy granting the consumer's deploy role the
|
||||
# right to invoke the platform Lambda's Function URL, scoped via ABAC
|
||||
# (aws:PrincipalTag/nova:owner == var.owner_id). The platform Lambda's
|
||||
# resource-based policy + the consumer_invoke_policy.json template
|
||||
# enforce the ABAC scope at the Lambda side; this policy grants the
|
||||
# invoke permission on the consumer side.
|
||||
resource "aws_iam_role_policy" "consumer_invoke" {
|
||||
name = "nova-consumer-invoke"
|
||||
role = aws_iam_role.consumer_deploy.id
|
||||
|
||||
policy = jsonencode({
|
||||
Version = "2012-10-17"
|
||||
Statement = [
|
||||
{
|
||||
Effect = "Allow"
|
||||
Action = [
|
||||
"lambda:InvokeFunctionUrl",
|
||||
]
|
||||
Resource = [
|
||||
# The platform Lambda ARN (cross-account). The account_id is
|
||||
# the platform account, not the consumer account. For offline-
|
||||
# proven mode, a placeholder ARN is used.
|
||||
"arn:aws:lambda:${var.region}:000000000000:function:nova-contract-ingestor"
|
||||
]
|
||||
Condition = {
|
||||
StringEquals = {
|
||||
"aws:PrincipalTag/nova:owner" = var.owner_id
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
})
|
||||
}
|
||||
|
||||
output "consumer_deploy_role_arn" {
|
||||
description = "The ARN of the consumer deploy role."
|
||||
value = aws_iam_role.consumer_deploy.arn
|
||||
}
|
||||
|
||||
output "consumer_deploy_role_name" {
|
||||
description = "The name of the consumer deploy role."
|
||||
value = aws_iam_role.consumer_deploy.name
|
||||
}
|
||||
+27
-129
@@ -37,29 +37,12 @@ _spec.loader.exec_module(ingestor)
|
||||
# Fixtures
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def _local_lambda_bypass(monkeypatch):
|
||||
"""P10 (REQ-174): set NOVA_LAMBDA_LOCAL_BYPASS for all ingestor tests
|
||||
so the fail-closed identity check doesn't block handler-routing tests.
|
||||
Tests that explicitly exercise the identity check (TestCallerIdentity
|
||||
Validation) override this per-test."""
|
||||
monkeypatch.setenv("NOVA_LAMBDA_LOCAL_BYPASS", "1")
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def sample_payload():
|
||||
# P11 (REQ-175): the contract blob must validate against
|
||||
# contract.schema.json (requires id/name/environment/infrastructure;
|
||||
# id matches ^[a-z][a-z0-9-]{2,5}$).
|
||||
return {
|
||||
"consumerRepo": "acdl/consumer-a",
|
||||
"contractId": "contract-001",
|
||||
"contract": {
|
||||
"id": "test",
|
||||
"name": "test-contract",
|
||||
"environment": "dev",
|
||||
"infrastructure": {"s3": {"version": "1.0.0", "inputs": {}}},
|
||||
},
|
||||
"contract": {"stack": "s3", "environment": "dev"},
|
||||
"environment": "dev",
|
||||
"action": "submit_contract",
|
||||
}
|
||||
@@ -67,8 +50,7 @@ def sample_payload():
|
||||
|
||||
@pytest.fixture
|
||||
def function_url_event(sample_payload):
|
||||
# P10 (REQ-174): include a test IAM identity so the fail-closed check passes.
|
||||
return {"body": json.dumps(sample_payload), "requestContext": {"identity": {"userArn": "arn:aws:sts::000:assumed-role/nova-deploy/test"}}}
|
||||
return {"body": json.dumps(sample_payload)}
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
@@ -141,15 +123,16 @@ class TestSubmitContract:
|
||||
assert item["submittedAt"]["S"] == result["submittedAt"]
|
||||
# The contract attribute holds the full contract object. boto3's
|
||||
# resource API serializes a dict as a DynamoDB Map (type "M"); each
|
||||
# leaf scalar is wrapped in its own type tag. P11 (REQ-175): the
|
||||
# fixture contract has a nested infrastructure map; assert the
|
||||
# top-level keys are present (full deep-equality is fragile with
|
||||
# moto's recursive type wrapping).
|
||||
actual_contract = item["contract"]["M"]
|
||||
assert set(actual_contract.keys()) == set(sample_payload["contract"].keys())
|
||||
assert actual_contract["id"]["S"] == sample_payload["contract"]["id"]
|
||||
assert actual_contract["name"]["S"] == sample_payload["contract"]["name"]
|
||||
assert actual_contract["environment"]["S"] == sample_payload["contract"]["environment"]
|
||||
# leaf scalar is wrapped in its own type tag.
|
||||
expected_contract = sample_payload["contract"]
|
||||
actual_contract = item["contract"]
|
||||
# The resource API stores scalars inside the map with their own type
|
||||
# tags (e.g. {"S": ...}); unwrap one level for the two known leaves.
|
||||
unwrapped = {
|
||||
k: list(v.values())[0] if isinstance(v, dict) and len(v) == 1 else v
|
||||
for k, v in actual_contract["M"].items()
|
||||
}
|
||||
assert unwrapped == expected_contract
|
||||
|
||||
def test_submit_contract_sk_contains_contract_id_and_timestamp(self, moto_contracts_table, sample_payload):
|
||||
result = ingestor._submit_contract(sample_payload)
|
||||
@@ -160,24 +143,6 @@ class TestSubmitContract:
|
||||
ts = sk.split("#", 1)[1]
|
||||
datetime.datetime.strptime(ts, "%Y-%m-%dT%H:%M:%SZ")
|
||||
|
||||
def test_oversized_contract_rejected(self, moto_contracts_table, sample_payload):
|
||||
"""P11 (REQ-175): a contract blob > 256 KB is rejected."""
|
||||
sample_payload["contract"] = {"blob": "x" * (300 * 1024)}
|
||||
with pytest.raises(ValueError, match="contract payload too large"):
|
||||
ingestor._submit_contract(sample_payload)
|
||||
|
||||
def test_schema_invalid_contract_rejected(self, moto_contracts_table, sample_payload, monkeypatch):
|
||||
"""P11 (REQ-175): a contract that fails contract.schema.json
|
||||
validation is rejected with a clear error."""
|
||||
# The autouse fixture sets NOVA_LAMBDA_LOCAL_BYPASS; unset it so
|
||||
# the schema validation runs (the bypass skips schema validation).
|
||||
monkeypatch.delenv("NOVA_LAMBDA_LOCAL_BYPASS", raising=False)
|
||||
# The contract schema requires id/name/environment/infrastructure;
|
||||
# an empty dict fails validation.
|
||||
sample_payload["contract"] = {}
|
||||
with pytest.raises(ValueError, match="contract schema validation failed"):
|
||||
ingestor._submit_contract(sample_payload)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# report_error (D-055) — GitHub issue creation via the GitHub API
|
||||
@@ -299,21 +264,20 @@ class TestReportError:
|
||||
ingestor._report_error(error_payload)
|
||||
|
||||
def test_report_error_truncates_stack_trace(self, monkeypatch, error_payload, patched_secrets):
|
||||
# P11 (REQ-175): a very long stack trace is truncated to
|
||||
# MAX_ERROR_FIELD_CHARS (10000) in the body (was 2000; aligned).
|
||||
error_payload["stackTrace"] = "x" * 20000
|
||||
# A very long stack trace should be truncated to 2000 chars in the body.
|
||||
error_payload["stackTrace"] = "x" * 5000
|
||||
calls = self._mock_urlopen(monkeypatch, [
|
||||
(200, json.dumps({"items": []})),
|
||||
(201, json.dumps({"number": 1, "html_url": "u"})),
|
||||
])
|
||||
result = ingestor._report_error(error_payload)
|
||||
assert result["status"] == "issue_created"
|
||||
# The create request body should contain exactly 10000 'x' chars.
|
||||
# The create request body should contain exactly 2000 'x' chars.
|
||||
create_req = calls[1]
|
||||
body = json.loads(create_req.data.decode())
|
||||
# The body markdown contains the (truncated) stack trace.
|
||||
assert "x" * 10000 in body["body"]
|
||||
assert "x" * 10001 not in body["body"]
|
||||
assert "x" * 2000 in body["body"]
|
||||
assert "x" * 2001 not in body["body"]
|
||||
|
||||
def test_lambda_handler_routes_report_error(self, monkeypatch, error_payload, patched_secrets):
|
||||
# End-to-end via lambda_handler: action=report_error → 200.
|
||||
@@ -397,21 +361,9 @@ class TestLambdaHandler:
|
||||
class TestCallerIdentityValidation:
|
||||
"""P1-2: the Lambda validates consumerRepo against the invoking principal."""
|
||||
|
||||
def test_no_identity_fails_closed(self, moto_contracts_table, sample_payload, monkeypatch):
|
||||
# P10 (REQ-174): no requestContext.identity → fail closed (defense-in-
|
||||
# depth). The old behavior (silent pass) is replaced with a 401.
|
||||
monkeypatch.delenv("NOVA_LAMBDA_LOCAL_BYPASS", raising=False)
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {}}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
assert resp["statusCode"] == 401
|
||||
assert "missing IAM caller identity" in json.loads(resp["body"])["error"]
|
||||
|
||||
def test_no_identity_passes_with_local_bypass(self, moto_contracts_table, sample_payload, monkeypatch):
|
||||
# P10 (REQ-174): the NOVA_LAMBDA_LOCAL_BYPASS env allows local/stub
|
||||
# testing without an IAM identity (the LocalLambdaStub sets it).
|
||||
monkeypatch.setenv("NOVA_LAMBDA_LOCAL_BYPASS", "1")
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {}}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
def test_no_identity_skips_check(self, moto_contracts_table, function_url_event):
|
||||
# No requestContext.identity in the event — check is skipped (relies on IAM ABAC).
|
||||
resp = ingestor.lambda_handler(function_url_event, None)
|
||||
assert resp["statusCode"] == 200
|
||||
|
||||
def test_invalid_consumer_repo_format_rejected(self, moto_contracts_table, sample_payload):
|
||||
@@ -544,7 +496,7 @@ class TestValidateChangeRequest:
|
||||
"action": "validate_change_request",
|
||||
"changeRequestId": "CHG0678912",
|
||||
"consumerRepo": "acdl/consumer-a",
|
||||
}), "requestContext": {"identity": {"userArn": "arn:aws:sts::000:assumed-role/nova-deploy/test"}}}
|
||||
})}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
assert resp["statusCode"] == 200
|
||||
body = json.loads(resp["body"])
|
||||
@@ -553,39 +505,33 @@ class TestValidateChangeRequest:
|
||||
|
||||
class TestV14IdentityValidation:
|
||||
"""v1.14 (REQ-144): contractId format, environment enum, error length
|
||||
validation + spoofing resistance.
|
||||
|
||||
P10 (REQ-174): these tests supply a valid userArn so the fail-closed
|
||||
identity check passes and the field validation is reached."""
|
||||
|
||||
_ARN = "arn:aws:sts::000:assumed-role/nova-deploy/test-session"
|
||||
validation + spoofing resistance."""
|
||||
|
||||
def test_invalid_contract_id_rejected(self, moto_contracts_table, sample_payload):
|
||||
sample_payload["contractId"] = "bad contract!@#"
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {"identity": {"userArn": self._ARN}}}
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {}}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
assert resp["statusCode"] == 400
|
||||
assert "invalid contractId" in resp["body"]
|
||||
|
||||
def test_contract_id_too_long_rejected(self, moto_contracts_table, sample_payload):
|
||||
sample_payload["contractId"] = "a" * 65
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {"identity": {"userArn": self._ARN}}}
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {}}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
assert resp["statusCode"] == 400
|
||||
assert "invalid contractId" in resp["body"]
|
||||
|
||||
def test_invalid_environment_rejected(self, moto_contracts_table, sample_payload):
|
||||
sample_payload["environment"] = "staging"
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {"identity": {"userArn": self._ARN}}}
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {}}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
assert resp["statusCode"] == 400
|
||||
assert "invalid environment" in resp["body"]
|
||||
|
||||
def test_valid_environments_accepted(self, moto_contracts_table, sample_payload):
|
||||
arn = "arn:aws:sts::000:assumed-role/nova-deploy/test"
|
||||
for env in ["dev", "qa", "prod", "dr"]:
|
||||
sample_payload["environment"] = env
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {"identity": {"userArn": arn}}}
|
||||
event = {"body": json.dumps(sample_payload), "requestContext": {}}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
assert resp["statusCode"] == 200
|
||||
|
||||
@@ -593,52 +539,4 @@ class TestV14IdentityValidation:
|
||||
"""The _validate_caller_identity docstring documents the ABAC reliance."""
|
||||
docstring = ingestor._validate_caller_identity.__doc__
|
||||
assert "ABAC" in docstring
|
||||
assert "PrincipalTag" in docstring
|
||||
|
||||
class TestOnboardConsumer:
|
||||
"""P18 (REQ-182): the onboard_consumer action writes a pending CMDB row."""
|
||||
|
||||
_ARN = "arn:aws:sts::000:assumed-role/nova-deploy/test"
|
||||
|
||||
def test_valid_onboarding_writes_pending_row(self, moto_contracts_table):
|
||||
payload = {
|
||||
"action": "onboard_consumer",
|
||||
"consumerRepo": "acdl/consumer-b",
|
||||
"requestedEnvironment": "dev",
|
||||
"ownerId": "team-b",
|
||||
"billingTag": "cost-center-b",
|
||||
}
|
||||
event = {"body": json.dumps(payload), "requestContext": {"identity": {"userArn": self._ARN}}}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
assert resp["statusCode"] == 200
|
||||
body = json.loads(resp["body"])
|
||||
assert body["status"] == "pending"
|
||||
assert body["action"] == "onboard_consumer"
|
||||
assert body["requestedEnvironment"] == "dev"
|
||||
|
||||
def test_invalid_onboarding_rejected(self, moto_contracts_table):
|
||||
# An invalid consumerRepo (no /) fails the identity format check
|
||||
# (which runs for all actions) before the onboarding schema.
|
||||
payload = {
|
||||
"action": "onboard_consumer",
|
||||
"consumerRepo": "not-a-repo-format",
|
||||
"requestedEnvironment": "dev",
|
||||
"ownerId": "team-b",
|
||||
"billingTag": "cost-center-b",
|
||||
}
|
||||
event = {"body": json.dumps(payload), "requestContext": {"identity": {"userArn": self._ARN}}}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
assert resp["statusCode"] == 400
|
||||
assert "invalid consumerRepo" in json.loads(resp["body"])["error"]
|
||||
|
||||
def test_missing_onboarding_field_rejected(self, moto_contracts_table):
|
||||
payload = {
|
||||
"action": "onboard_consumer",
|
||||
"consumerRepo": "acdl/consumer-b",
|
||||
"requestedEnvironment": "dev",
|
||||
# ownerId + billingTag missing
|
||||
}
|
||||
event = {"body": json.dumps(payload), "requestContext": {"identity": {"userArn": self._ARN}}}
|
||||
resp = ingestor.lambda_handler(event, None)
|
||||
assert resp["statusCode"] == 400
|
||||
assert "onboarding payload invalid" in json.loads(resp["body"])["error"]
|
||||
assert "PrincipalTag" in docstring
|
||||
@@ -34,15 +34,4 @@ class TestDocsCoverage:
|
||||
assert "How to Write an Adapter" in content
|
||||
assert "How to Wire" in content
|
||||
assert "How to Test" in content
|
||||
assert "Existing Adapters" in content
|
||||
|
||||
def test_github_workflows_readme_catalogs_all_workflows():
|
||||
"""P16 (REQ-180): .github/workflows/README.md catalogs all 7 workflows."""
|
||||
from pathlib import Path
|
||||
readme = Path(__file__).resolve().parent.parent / ".github" / "workflows" / "README.md"
|
||||
assert readme.is_file(), ".github/workflows/README.md missing"
|
||||
text = readme.read_text()
|
||||
for wf in ["ci.yml", "deploy.yml", "modules-lifecycle.yml",
|
||||
"platform-test.yml", "primitives-plan.yml", "patterns-plan.yml",
|
||||
"release.yml"]:
|
||||
assert wf in text, f"{wf} not cataloged in .github/workflows/README.md"
|
||||
assert "Existing Adapters" in content
|
||||
@@ -21,10 +21,7 @@ class TestEnvironmentCheck:
|
||||
assert ok is False
|
||||
assert "nonexistent-env" in msg
|
||||
assert "onboarding" in msg.lower() or "Environment Onboarding" in msg
|
||||
# P19 (REQ-183): the message now routes to the self-service
|
||||
# request path (onboard_consumer), not "contact the platform team".
|
||||
assert "platform team" not in msg.lower()
|
||||
assert "onboard_consumer" in msg or "self-service" in msg.lower()
|
||||
assert "platform team" in msg.lower()
|
||||
|
||||
def test_onboarding_message_lists_platform_provisions(self):
|
||||
msg = _onboarding_message("qa")
|
||||
@@ -105,16 +102,4 @@ class TestRunPlatformWireIn:
|
||||
)
|
||||
assert result.returncode == 0, f"stdout: {result.stdout}\nstderr: {result.stderr}"
|
||||
assert "PLATFORM CHECK OK" in result.stdout
|
||||
assert "environment" in result.stdout.lower() or "Step 0" in result.stdout
|
||||
|
||||
class TestOnboardingMessageSelfService:
|
||||
"""P19 (REQ-183): the onboarding message is self-service, not 'contact
|
||||
the platform team'."""
|
||||
|
||||
def test_no_contact_platform_team(self):
|
||||
msg = _onboarding_message("qa")
|
||||
assert "contact the platform team" not in msg.lower()
|
||||
|
||||
def test_mentions_self_service_request(self):
|
||||
msg = _onboarding_message("qa")
|
||||
assert "self-service" in msg.lower() or "onboard_consumer" in msg
|
||||
assert "environment" in result.stdout.lower() or "Step 0" in result.stdout
|
||||
@@ -1,200 +0,0 @@
|
||||
"""Tests for Nova metrics collector (P2, REQ-189/200).
|
||||
|
||||
Tests the collector's idempotent re-run property (REQ-200) and the
|
||||
SQLite cold store schema.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
ROOT = Path(__file__).resolve().parent.parent
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def tmp_store(tmp_path, monkeypatch):
|
||||
"""Redirect metrics/ to a tmp dir for isolated testing."""
|
||||
metrics_dir = tmp_path / "metrics"
|
||||
metrics_dir.mkdir()
|
||||
runs_dir = metrics_dir / "runs"
|
||||
runs_dir.mkdir()
|
||||
lifecycle_dir = metrics_dir / "lifecycle"
|
||||
lifecycle_dir.mkdir()
|
||||
store_db = metrics_dir / "nova_metrics.db"
|
||||
ledger_db = metrics_dir / "decision_ledger.db"
|
||||
|
||||
monkeypatch.setattr("core.metrics.collector._METRICS_DIR", str(metrics_dir))
|
||||
monkeypatch.setattr("core.metrics.collector._STORE_PATH", str(store_db))
|
||||
monkeypatch.setattr("core.metrics.collector._RUNS_DIR", str(runs_dir))
|
||||
monkeypatch.setattr("core.metrics.collector._LEDGER_DB", str(ledger_db))
|
||||
monkeypatch.setattr("core.metrics.collector._REPO_ROOT", str(tmp_path))
|
||||
monkeypatch.setattr("core.metrics.collector._REGRESSION_REPORT", str(tmp_path / "REGRESSION_REPORT.json"))
|
||||
monkeypatch.setattr("core.metrics.collector._COVERAGE_JSON", str(metrics_dir / "coverage.json"))
|
||||
monkeypatch.setattr("core.metrics.collector._TEST_RESULTS_XML", str(metrics_dir / "test-results.xml"))
|
||||
monkeypatch.setattr("core.metrics.decision_ledger._LEDGER_PATH", str(ledger_db))
|
||||
return {"metrics_dir": metrics_dir, "store_db": store_db, "ledger_db": ledger_db, "runs_dir": runs_dir}
|
||||
|
||||
|
||||
def _write_regression_report(path, run_id="regr-test-1"):
|
||||
report = {
|
||||
"run_id": run_id,
|
||||
"run_at_utc": "2026-08-04T12:00:00Z",
|
||||
"milestone": "v1.17",
|
||||
"phase": 0,
|
||||
"summary": {"Verified": 18, "Decayed": 0, "Broken": 0, "Skipped": 4},
|
||||
"passed": True,
|
||||
"results": [
|
||||
{"capability_id": "CAP-001", "name": "test cap", "status": "Verified", "tier": "local", "duration_ms": 100, "detail": "ok"},
|
||||
{"capability_id": "CAP-002", "name": "test cap 2", "status": "Skipped", "tier": "live-aws", "duration_ms": 50, "detail": "D-096"},
|
||||
],
|
||||
}
|
||||
with open(path, "w") as f:
|
||||
json.dump(report, f)
|
||||
|
||||
|
||||
def _write_run_manifest(runs_dir, run_id="run-test-1"):
|
||||
manifest = {
|
||||
"run_id": run_id,
|
||||
"contract_id": "cid-1",
|
||||
"environment": "dev",
|
||||
"started_at": "2026-08-04T12:00:00Z",
|
||||
"completed_at": "2026-08-04T12:01:00Z",
|
||||
"exit_code": 0,
|
||||
"stages": [{"name": "resolve", "duration_ms": 100, "exit_code": 0}],
|
||||
"outcome": "succeeded",
|
||||
"confidence": {"score": 0.9, "band": "pass", "perInput": {"policy": 1.0}},
|
||||
"hitl": {"gate": "dev", "result": "autonomous", "block": False},
|
||||
"cost_estimate_usd": -12.5,
|
||||
"decision_id": run_id,
|
||||
}
|
||||
with open(runs_dir / f"{run_id}.json", "w") as f:
|
||||
json.dump(manifest, f)
|
||||
|
||||
|
||||
def _write_junit(path):
|
||||
xml = """<?xml version="1.0" encoding="utf-8"?>
|
||||
<testsuites>
|
||||
<testsuite name="test_metrics" tests="10" failures="0" errors="0" skipped="0" time="1.5">
|
||||
<testcase name="test_one" time="0.1"/>
|
||||
</testsuite>
|
||||
</testsuites>"""
|
||||
path.write_text(xml)
|
||||
|
||||
|
||||
def _write_coverage(path):
|
||||
with open(path, "w") as f:
|
||||
json.dump({"totals": {"percent_covered": 85.5}}, f)
|
||||
|
||||
|
||||
def test_collector_init(tmp_store):
|
||||
from core.metrics.collector import _init_store
|
||||
_init_store()
|
||||
assert tmp_store["store_db"].exists()
|
||||
conn = sqlite3.connect(str(tmp_store["store_db"]))
|
||||
tables = conn.execute("SELECT name FROM sqlite_master WHERE type='table'").fetchall()
|
||||
conn.close()
|
||||
table_names = [t[0] for t in tables]
|
||||
assert "fact_run" in table_names
|
||||
assert "fact_capability" in table_names
|
||||
assert "fact_decision" in table_names
|
||||
assert "dim_capability" in table_names
|
||||
assert "dim_milestone" in table_names
|
||||
|
||||
|
||||
def test_collector_regression_report(tmp_store):
|
||||
from core.metrics.collector import collect_regression_report
|
||||
_write_regression_report(tmp_store["metrics_dir"].parent / "REGRESSION_REPORT.json")
|
||||
count = collect_regression_report()
|
||||
assert count == 2
|
||||
conn = sqlite3.connect(str(tmp_store["store_db"]))
|
||||
rows = conn.execute("SELECT capability_id, status FROM fact_capability").fetchall()
|
||||
conn.close()
|
||||
assert len(rows) == 2
|
||||
assert rows[0][0] == "CAP-001"
|
||||
|
||||
|
||||
def test_collector_run_manifests(tmp_store):
|
||||
from core.metrics.collector import collect_run_manifests
|
||||
_write_run_manifest(tmp_store["runs_dir"])
|
||||
count = collect_run_manifests()
|
||||
assert count == 1
|
||||
conn = sqlite3.connect(str(tmp_store["store_db"]))
|
||||
row = conn.execute("SELECT run_id, confidence_score, cost_estimate_usd FROM fact_run").fetchone()
|
||||
conn.close()
|
||||
assert row[0] == "run-test-1"
|
||||
assert row[1] == 0.9
|
||||
assert row[2] == -12.5
|
||||
|
||||
|
||||
def test_collector_idempotent(tmp_store):
|
||||
"""REQ-200: re-running the collector produces identical row counts."""
|
||||
from core.metrics.collector import collect_all
|
||||
_write_regression_report(tmp_store["metrics_dir"].parent / "REGRESSION_REPORT.json")
|
||||
_write_run_manifest(tmp_store["runs_dir"])
|
||||
_write_junit(tmp_store["metrics_dir"] / "test-results.xml")
|
||||
_write_coverage(tmp_store["metrics_dir"] / "coverage.json")
|
||||
|
||||
result1 = collect_all()
|
||||
conn = sqlite3.connect(str(tmp_store["store_db"]))
|
||||
cap_count_1 = conn.execute("SELECT COUNT(*) FROM fact_capability").fetchone()[0]
|
||||
run_count_1 = conn.execute("SELECT COUNT(*) FROM fact_run").fetchone()[0]
|
||||
conn.close()
|
||||
|
||||
result2 = collect_all()
|
||||
conn = sqlite3.connect(str(tmp_store["store_db"]))
|
||||
cap_count_2 = conn.execute("SELECT COUNT(*) FROM fact_capability").fetchone()[0]
|
||||
run_count_2 = conn.execute("SELECT COUNT(*) FROM fact_run").fetchone()[0]
|
||||
conn.close()
|
||||
|
||||
assert cap_count_1 == cap_count_2
|
||||
assert run_count_1 == run_count_2
|
||||
|
||||
|
||||
def test_collector_decision_ledger(tmp_store):
|
||||
from core.metrics.event_envelope import make_event
|
||||
from core.metrics.decision_ledger import append
|
||||
from core.metrics.collector import collect_decision_ledger
|
||||
ev = make_event("nova.ai.decision.made", "run-dl-collect-1", "dev",
|
||||
{"decision_id": "run-dl-collect-1", "chosen_action": "pass",
|
||||
"confidence": 0.94, "alternatives": {"policy": 1.0},
|
||||
"human_override": False, "outcome": "succeeded"})
|
||||
append(ev)
|
||||
count = collect_decision_ledger()
|
||||
assert count == 1
|
||||
conn = sqlite3.connect(str(tmp_store["store_db"]))
|
||||
row = conn.execute("SELECT decision_id, confidence, chosen_action FROM fact_decision").fetchone()
|
||||
conn.close()
|
||||
assert row[0] == "run-dl-collect-1"
|
||||
assert row[1] == 0.94
|
||||
assert row[2] == "pass"
|
||||
|
||||
|
||||
def test_collector_test_results(tmp_store):
|
||||
from core.metrics.collector import collect_test_results
|
||||
_write_junit(tmp_store["metrics_dir"] / "test-results.xml")
|
||||
_write_coverage(tmp_store["metrics_dir"] / "coverage.json")
|
||||
count = collect_test_results()
|
||||
assert count == 1
|
||||
conn = sqlite3.connect(str(tmp_store["store_db"]))
|
||||
row = conn.execute("SELECT total_tests, passed, coverage_pct FROM fact_test").fetchone()
|
||||
conn.close()
|
||||
assert row[0] == 10
|
||||
assert row[1] == 10
|
||||
assert row[2] == 85.5
|
||||
|
||||
|
||||
def test_collector_all(tmp_store):
|
||||
from core.metrics.collector import collect_all
|
||||
_write_regression_report(tmp_store["metrics_dir"].parent / "REGRESSION_REPORT.json")
|
||||
_write_run_manifest(tmp_store["runs_dir"])
|
||||
_write_junit(tmp_store["metrics_dir"] / "test-results.xml")
|
||||
_write_coverage(tmp_store["metrics_dir"] / "coverage.json")
|
||||
result = collect_all()
|
||||
assert result["capabilities"] == 2
|
||||
assert result["runs"] == 1
|
||||
assert result["tests"] == 1
|
||||
@@ -1,270 +0,0 @@
|
||||
"""Tests for Nova metrics event emitters (P1, REQ-187/188).
|
||||
|
||||
Tests the CloudEvents envelope, per-run manifest writer, Decision Ledger
|
||||
(hash-chain integrity + verify-chain), and the event emission from
|
||||
confidence_signal, hitl_gates, and checkov_adapter.
|
||||
"""
|
||||
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
ROOT = Path(__file__).resolve().parent.parent
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def tmp_metrics(tmp_path, monkeypatch):
|
||||
"""Redirect metrics/ to a tmp dir for isolated testing."""
|
||||
metrics_dir = tmp_path / "metrics"
|
||||
metrics_dir.mkdir()
|
||||
runs_dir = metrics_dir / "runs"
|
||||
runs_dir.mkdir()
|
||||
events_log = metrics_dir / "events.jsonl"
|
||||
ledger_db = metrics_dir / "decision_ledger.db"
|
||||
|
||||
monkeypatch.setattr("core.metrics.event_envelope.METRICS_DIR", str(metrics_dir))
|
||||
monkeypatch.setattr("core.metrics.event_envelope.EVENTS_LOG", str(events_log))
|
||||
monkeypatch.setattr("core.metrics.run_manifest._METRICS_DIR", str(metrics_dir))
|
||||
monkeypatch.setattr("core.metrics.run_manifest._RUNS_DIR", str(runs_dir))
|
||||
monkeypatch.setattr("core.metrics.decision_ledger._LEDGER_PATH", str(ledger_db))
|
||||
return {"metrics_dir": metrics_dir, "events_log": events_log, "ledger_db": ledger_db, "runs_dir": runs_dir}
|
||||
|
||||
|
||||
# --- Task 1: CloudEvents envelope ---
|
||||
|
||||
def test_envelope_valid(tmp_metrics):
|
||||
from core.metrics.event_envelope import make_event
|
||||
ev = make_event("nova.run.completed", "run-test-1", "dev", {"exit_code": 0}, contract_id="cid-123")
|
||||
assert ev["specversion"] == "1.0"
|
||||
assert ev["type"] == "nova.run.completed"
|
||||
assert ev["platform"]["run_id"] == "run-test-1"
|
||||
assert ev["platform"]["environment"] == "dev"
|
||||
assert ev["platform"]["contract_id"] == "cid-123"
|
||||
assert ev["data"]["exit_code"] == 0
|
||||
assert ev["datacontenttype"] == "application/json"
|
||||
assert "id" in ev and len(ev["id"]) > 0
|
||||
assert "time" in ev
|
||||
|
||||
|
||||
def test_envelope_append(tmp_metrics):
|
||||
from core.metrics.event_envelope import make_event, append_event
|
||||
ev = make_event("nova.test.event", "run-test-2", "dev", {"key": "value"})
|
||||
append_event(ev)
|
||||
assert tmp_metrics["events_log"].exists()
|
||||
lines = tmp_metrics["events_log"].read_text().strip().split("\n")
|
||||
assert len(lines) == 1
|
||||
parsed = json.loads(lines[0])
|
||||
assert parsed["type"] == "nova.test.event"
|
||||
|
||||
|
||||
# --- Task 2: Per-run manifest writer ---
|
||||
|
||||
def test_run_manifest_start(tmp_metrics):
|
||||
from core.metrics.run_manifest import start_run
|
||||
run_id = start_run("cid-123", "dev")
|
||||
assert run_id.startswith("run-")
|
||||
assert tmp_metrics["events_log"].exists()
|
||||
|
||||
|
||||
def test_run_manifest_complete(tmp_metrics):
|
||||
from core.metrics.run_manifest import start_run, complete_run
|
||||
run_id = start_run("cid-123", "dev")
|
||||
stages = [{"name": "resolve", "duration_ms": 100, "exit_code": 0}]
|
||||
manifest = complete_run(run_id, "cid-123", "dev", stages, 0)
|
||||
assert manifest["run_id"] == run_id
|
||||
assert manifest["exit_code"] == 0
|
||||
assert manifest["outcome"] == "succeeded"
|
||||
manifest_path = tmp_metrics["runs_dir"] / f"{run_id}.json"
|
||||
assert manifest_path.exists()
|
||||
saved = json.loads(manifest_path.read_text())
|
||||
assert saved["run_id"] == run_id
|
||||
|
||||
|
||||
def test_run_manifest_failed(tmp_metrics):
|
||||
from core.metrics.run_manifest import complete_run
|
||||
manifest = complete_run("run-fail-1", "cid-123", "dev", [], 1)
|
||||
assert manifest["outcome"] == "failed"
|
||||
|
||||
|
||||
# --- Task 6: Decision Ledger (SQLite hash-chain) ---
|
||||
|
||||
def test_decision_ledger_append(tmp_metrics):
|
||||
from core.metrics.event_envelope import make_event
|
||||
from core.metrics.decision_ledger import append, verify_chain
|
||||
ev = make_event("nova.ai.decision.made", "run-dl-1", "dev",
|
||||
{"decision_id": "run-dl-1", "chosen_action": "pass", "confidence": 0.9,
|
||||
"alternatives": {"policy": 1.0}, "human_override": False})
|
||||
row = append(ev)
|
||||
assert row["seq"] == 1
|
||||
assert row["prev_hash"] == "GENESIS"
|
||||
ok, broken, _ = verify_chain()
|
||||
assert ok
|
||||
assert broken == 0
|
||||
|
||||
|
||||
def test_decision_ledger_chain_integrity(tmp_metrics):
|
||||
from core.metrics.event_envelope import make_event
|
||||
from core.metrics.decision_ledger import append, verify_chain
|
||||
for i in range(5):
|
||||
ev = make_event("nova.ai.decision.made", f"run-dl-{i}", "dev",
|
||||
{"decision_id": f"run-dl-{i}", "confidence": 0.9 + i * 0.01})
|
||||
append(ev)
|
||||
ok, broken, details = verify_chain()
|
||||
assert ok, f"chain broken: {details}"
|
||||
assert broken == 0
|
||||
|
||||
|
||||
def test_decision_ledger_tamper_detection(tmp_metrics):
|
||||
from core.metrics.event_envelope import make_event
|
||||
from core.metrics.decision_ledger import append, verify_chain
|
||||
ev = make_event("nova.ai.decision.made", "run-tamper-1", "dev", {"confidence": 0.9})
|
||||
append(ev)
|
||||
# Tamper: directly modify the payload in the DB
|
||||
conn = sqlite3.connect(str(tmp_metrics["ledger_db"]))
|
||||
conn.execute("UPDATE decision_ledger SET payload = '{}' WHERE seq = 1")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
ok, broken, details = verify_chain()
|
||||
assert not ok
|
||||
assert broken > 0
|
||||
|
||||
|
||||
def test_decision_ledger_query_by_run(tmp_metrics):
|
||||
from core.metrics.event_envelope import make_event
|
||||
from core.metrics.decision_ledger import append, query_by_run
|
||||
ev = make_event("nova.ai.decision.made", "run-query-1", "dev", {"confidence": 0.9})
|
||||
append(ev)
|
||||
entries = query_by_run("run-query-1")
|
||||
assert len(entries) == 1
|
||||
assert entries[0]["event_type"] == "nova.ai.decision.made"
|
||||
|
||||
|
||||
def test_decision_ledger_stats(tmp_metrics):
|
||||
from core.metrics.event_envelope import make_event
|
||||
from core.metrics.decision_ledger import append, stats
|
||||
for env in ("dev", "qa", "dev"):
|
||||
ev = make_event("nova.ai.decision.made", f"run-stats-{env}", env, {"confidence": 0.9})
|
||||
append(ev)
|
||||
s = stats()
|
||||
assert s["total"] == 3
|
||||
assert s["by_environment"].get("dev", 0) == 2
|
||||
assert s["by_environment"].get("qa", 0) == 1
|
||||
|
||||
|
||||
def test_decision_ledger_replay(tmp_metrics):
|
||||
from core.metrics.event_envelope import make_event
|
||||
from core.metrics.decision_ledger import append, replay_run
|
||||
ev = make_event("nova.ai.decision.made", "run-replay-1", "dev",
|
||||
{"decision_id": "run-replay-1", "chosen_action": "pass",
|
||||
"confidence": 0.94, "human_override": False})
|
||||
append(ev)
|
||||
replay = replay_run("run-replay-1")
|
||||
assert "run-replay-1" in replay
|
||||
assert "nova.ai.decision.made" in replay
|
||||
|
||||
|
||||
# --- Task 8: Confidence signal event emission ---
|
||||
|
||||
def test_confidence_event_emission(tmp_metrics):
|
||||
from core.confidence_signal import compute
|
||||
inputs = {
|
||||
"policy": [{"result": "pass", "severity": "info"}],
|
||||
"validation": {"schema": True, "stack_resolved": True, "tf_validated": True, "tf_planned": True},
|
||||
"freshness": {"age_days": 0, "max_age_days": 1},
|
||||
"source": {"submitter": "test", "commit_sha": "abc"},
|
||||
"history": {"prior_rollbacks": 0, "prior_policy_fails": 0},
|
||||
"nfrs": {"conformance": 1.0},
|
||||
}
|
||||
sig = compute("cid-conf-1", "dev", inputs)
|
||||
assert sig.band == "pass"
|
||||
# Check that events were emitted
|
||||
assert tmp_metrics["events_log"].exists()
|
||||
lines = tmp_metrics["events_log"].read_text().strip().split("\n")
|
||||
types = [json.loads(l)["type"] for l in lines]
|
||||
assert "nova.confidence.computed" in types
|
||||
assert "nova.ai.decision.made" in types
|
||||
# Check the decision ledger has the entry
|
||||
from core.metrics.decision_ledger import query_by_run
|
||||
entries = query_by_run(lines[0].split('"run_id":"')[1].split('"')[0] if '"run_id":"' in lines[0] else "")
|
||||
# The run_id is dynamic; just verify the ledger has entries
|
||||
from core.metrics.decision_ledger import stats
|
||||
s = stats()
|
||||
assert s["total"] > 0
|
||||
|
||||
|
||||
# --- Task 7: Attestation event emission ---
|
||||
|
||||
def test_attestation_event_emission(tmp_metrics):
|
||||
from core.hitl_gates import attest
|
||||
# Dev skips (autonomous) — no event
|
||||
ok, reason = attest("cid-attest-1", "dev", "testuser")
|
||||
assert ok
|
||||
# QA requires approver + attestation matrix — mock evidence
|
||||
ok, reason = attest("cid-attest-2", "qa", "testuser",
|
||||
evidence={"functional_correctness": {"timestamp": "2026-08-04T12:00:00Z", "type": "test", "payload": {}},
|
||||
"performance_baseline": {"timestamp": "2026-08-04T12:00:00Z", "type": "test", "payload": {}},
|
||||
"security_posture": {"timestamp": "2026-08-04T12:00:00Z", "type": "test", "payload": {}},
|
||||
"contract_nfrs": {"valid": True}})
|
||||
assert ok
|
||||
# Check the attestation event was emitted
|
||||
from core.metrics.decision_ledger import stats
|
||||
s = stats()
|
||||
assert s["total"] > 0
|
||||
|
||||
|
||||
# --- Task 9: Policy event emission ---
|
||||
|
||||
def test_policy_event_emission(tmp_metrics, tmp_path):
|
||||
"""Test that checkov_adapter emits nova.policy.evaluated when given a run_id."""
|
||||
checkov_json = tmp_path / "checkov.json"
|
||||
checkov_json.write_text(json.dumps({
|
||||
"terraform_plan": {
|
||||
"results": {
|
||||
"passed_checks": [{"check_id": "CKV_AWS_1", "check_name": "test", "file_path": "main.tf"}],
|
||||
"failed_checks": [],
|
||||
"skipped_checks": [],
|
||||
}
|
||||
}
|
||||
}))
|
||||
from adapters.terraform.policy.checkov_adapter import adapt
|
||||
pcrs = adapt(str(checkov_json), "cid-policy-1", run_id="run-policy-1", environment="dev")
|
||||
assert len(pcrs) == 1
|
||||
assert pcrs[0]["result"] == "pass"
|
||||
# Check the event was emitted
|
||||
assert tmp_metrics["events_log"].exists()
|
||||
lines = tmp_metrics["events_log"].read_text().strip().split("\n")
|
||||
types = [json.loads(l)["type"] for l in lines]
|
||||
assert "nova.policy.evaluated" in types
|
||||
|
||||
|
||||
# --- Task 5: Infracost adapter (degraded mode) ---
|
||||
|
||||
def test_infracost_degraded_mode(tmp_metrics):
|
||||
"""When Infracost CLI is absent, the adapter degrades gracefully (A6)."""
|
||||
from core.metrics.infracost_adapter import estimate
|
||||
# Infracost is not installed in the test env — degraded mode
|
||||
result = estimate("/nonexistent/plan.json", "run-infracost-1", "cid-1", "dev")
|
||||
assert result["available"] is False
|
||||
assert result["delta_usd"] == 0.0
|
||||
|
||||
|
||||
# --- Task 3: Persist ephemeral $WORK/*.json ---
|
||||
|
||||
def test_persist_run_artifacts(tmp_metrics, tmp_path):
|
||||
from core.metrics.run_manifest import persist_run_artifacts
|
||||
work_dir = tmp_path / "work"
|
||||
work_dir.mkdir()
|
||||
(work_dir / "pcr.json").write_text('{"test": true}')
|
||||
(work_dir / "signal.json").write_text('{"score": 0.9}')
|
||||
copied = persist_run_artifacts("run-persist-1", str(work_dir))
|
||||
assert "pcr.json" in copied
|
||||
assert "signal.json" in copied
|
||||
dest = tmp_metrics["runs_dir"] / "run-persist-1"
|
||||
assert (dest / "pcr.json").exists()
|
||||
assert (dest / "signal.json").exists()
|
||||
@@ -1,50 +0,0 @@
|
||||
"""Unit tests for core/onboarding.py (P19, REQ-183)."""
|
||||
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
|
||||
from core.onboarding import generate_env_file, _onboarding_request_message
|
||||
|
||||
|
||||
class TestGenerateEnvFile:
|
||||
"""P19 (REQ-183): generate_env_file produces a valid env JSON."""
|
||||
|
||||
def test_generates_env_with_request_fields(self):
|
||||
request = {
|
||||
"consumerRepo": "acdl/consumer-b",
|
||||
"requestedEnvironment": "qa",
|
||||
"ownerId": "team-b",
|
||||
"billingTag": "cost-center-b",
|
||||
}
|
||||
env = generate_env_file(request, template_env="dev")
|
||||
assert env["name"] == "qa"
|
||||
assert env["ownerId"] == "team-b"
|
||||
assert env["billingTag"] == "cost-center-b"
|
||||
assert env["account_id"] == "000000000000" # placeholder
|
||||
assert "consumer-b" in env["description"]
|
||||
|
||||
def test_preserves_template_network_and_state(self):
|
||||
request = {
|
||||
"consumerRepo": "acdl/c",
|
||||
"requestedEnvironment": "prod",
|
||||
"ownerId": "team-a",
|
||||
"billingTag": "cc-a",
|
||||
}
|
||||
env = generate_env_file(request, template_env="dev")
|
||||
assert "vpc_cidr" in env["network"]
|
||||
assert "bucket" in env["state_backend"]
|
||||
assert env["region"] == "us-east-1"
|
||||
|
||||
|
||||
class TestOnboardingRequestMessage:
|
||||
"""P19 (REQ-183): the request message is self-service."""
|
||||
|
||||
def test_message_mentions_onboard_consumer(self):
|
||||
msg = _onboarding_request_message("dev")
|
||||
assert "onboard_consumer" in msg
|
||||
assert "Nova" in msg
|
||||
@@ -1,38 +0,0 @@
|
||||
"""Unit tests for terraform/onboarding (P20, REQ-184)."""
|
||||
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
ROOT = Path(__file__).resolve().parent.parent
|
||||
ONBOARDING_DIR = ROOT / "terraform" / "onboarding"
|
||||
|
||||
|
||||
def test_onboarding_terraform_dir_exists():
|
||||
"""P20 (REQ-184): terraform/onboarding/ exists with main.tf + README."""
|
||||
assert ONBOARDING_DIR.is_dir()
|
||||
assert (ONBOARDING_DIR / "main.tf").is_file()
|
||||
assert (ONBOARDING_DIR / "README.md").is_file()
|
||||
|
||||
|
||||
def test_onboarding_terraform_validates():
|
||||
"""P20 (REQ-184): terraform validate passes for the onboarding module
|
||||
(offline-proven, D-114). Skipped if terraform is not installed."""
|
||||
if not subprocess.call(["which", "terraform"], stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) == 0:
|
||||
pytest.skip("terraform not installed")
|
||||
rc = subprocess.call(
|
||||
["terraform", "validate"],
|
||||
cwd=str(ONBOARDING_DIR),
|
||||
stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
|
||||
)
|
||||
assert rc == 0, "terraform validate failed for terraform/onboarding/"
|
||||
|
||||
|
||||
def test_onboarding_main_tf_has_nova_tags():
|
||||
"""P20 (REQ-184): the deploy role is tagged with nova:owner + nova:contract."""
|
||||
main_tf = (ONBOARDING_DIR / "main.tf").read_text()
|
||||
assert '"nova:owner"' in main_tf
|
||||
assert '"nova:contract"' in main_tf
|
||||
assert "aws_iam_role" in main_tf
|
||||
assert "lambda:InvokeFunctionUrl" in main_tf
|
||||
@@ -1,91 +0,0 @@
|
||||
"""Tests for Nova PowerBI export (P3, REQ-190)."""
|
||||
|
||||
import json
|
||||
import os
|
||||
import sqlite3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
ROOT = Path(__file__).resolve().parent.parent
|
||||
sys.path.insert(0, str(ROOT))
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def tmp_export(tmp_path, monkeypatch):
|
||||
metrics_dir = tmp_path / "metrics"
|
||||
metrics_dir.mkdir()
|
||||
export_dir = metrics_dir / "powerbi"
|
||||
store_db = metrics_dir / "nova_metrics.db"
|
||||
monkeypatch.setattr("core.metrics.powerbi_export._METRICS_DIR", str(metrics_dir))
|
||||
monkeypatch.setattr("core.metrics.powerbi_export._STORE_PATH", str(store_db))
|
||||
monkeypatch.setattr("core.metrics.powerbi_export._EXPORT_DIR", str(export_dir))
|
||||
return {"metrics_dir": metrics_dir, "store_db": store_db, "export_dir": export_dir}
|
||||
|
||||
|
||||
def _init_store_with_data(db_path):
|
||||
conn = sqlite3.connect(str(db_path))
|
||||
conn.executescript("""
|
||||
CREATE TABLE fact_run (run_id TEXT PRIMARY KEY, contract_id TEXT, environment TEXT, exit_code INTEGER);
|
||||
CREATE TABLE dim_capability (capability_id TEXT PRIMARY KEY, name TEXT, tier TEXT);
|
||||
INSERT INTO fact_run VALUES ('run-1', 'cid-1', 'dev', 0);
|
||||
INSERT INTO dim_capability VALUES ('CAP-001', 'test cap', 'local');
|
||||
""")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
|
||||
|
||||
def test_export_csv(tmp_export):
|
||||
from core.metrics.powerbi_export import export_all
|
||||
_init_store_with_data(tmp_export["store_db"])
|
||||
result = export_all(fmt="csv")
|
||||
assert (tmp_export["export_dir"] / "fact_run.csv").exists()
|
||||
assert (tmp_export["export_dir"] / "dim_capability.csv").exists()
|
||||
assert result["fact_tables"]["fact_run"] == 1
|
||||
|
||||
|
||||
def test_export_json(tmp_export):
|
||||
from core.metrics.powerbi_export import export_all
|
||||
_init_store_with_data(tmp_export["store_db"])
|
||||
result = export_all(fmt="json")
|
||||
assert (tmp_export["export_dir"] / "fact_run.json").exists()
|
||||
data = json.loads((tmp_export["export_dir"] / "fact_run.json").read_text())
|
||||
assert len(data) == 1
|
||||
assert data[0]["run_id"] == "run-1"
|
||||
|
||||
|
||||
def test_export_placeholder_views(tmp_export):
|
||||
from core.metrics.powerbi_export import export_all, PLACEHOLDER_VIEWS
|
||||
result = export_all(fmt="both")
|
||||
for view_name in PLACEHOLDER_VIEWS:
|
||||
assert (tmp_export["export_dir"] / f"{view_name}.csv").exists()
|
||||
assert (tmp_export["export_dir"] / f"{view_name}.json").exists()
|
||||
assert len(PLACEHOLDER_VIEWS) == 8
|
||||
|
||||
|
||||
def test_export_placeholder_csv_headers_only(tmp_export):
|
||||
from core.metrics.powerbi_export import export_all
|
||||
export_all(fmt="csv")
|
||||
csv_path = tmp_export["export_dir"] / "placeholder_drift_detection.csv"
|
||||
lines = csv_path.read_text().strip().split("\n")
|
||||
assert len(lines) == 1 # headers only, no data
|
||||
assert "timestamp" in lines[0]
|
||||
|
||||
|
||||
def test_export_placeholder_json_schema(tmp_export):
|
||||
from core.metrics.powerbi_export import export_all
|
||||
export_all(fmt="json")
|
||||
json_path = tmp_export["export_dir"] / "placeholder_sla_downtime.json"
|
||||
data = json.loads(json_path.read_text())
|
||||
assert "schema" in data
|
||||
assert data["schema"]["blocking_decision"] == "D-096"
|
||||
assert data["data"] == []
|
||||
|
||||
|
||||
def test_export_no_store(tmp_export):
|
||||
from core.metrics.powerbi_export import export_all
|
||||
result = export_all(fmt="csv")
|
||||
assert "error" in result
|
||||
# Placeholders still exported
|
||||
assert (tmp_export["export_dir"] / "placeholder_drift_detection.csv").exists()
|
||||
Reference in New Issue
Block a user