Compare commits

...

50 Commits

Author SHA1 Message Date
Jon Chery 85c500e45a feat(P4): pipeline hardening — Checkov before plan, Wiz-or-Checkov on plan (REQ-250)
Nova Slides Render / render (push) Failing after 1m1s
Two-stage policy scan per item 20:

1. Checkov on static code BEFORE terraform plan (fail-fast, quick dev
   feedback). Added to run_platform.sh Step 3c + run_codegen.sh Step 3c
   (runs on the authored TF dir before plan, using --framework terraform).

2. Runtime policy scan on the plan AFTER terraform plan: Wiz when
   configured (WIZ_API_TOKEN + WIZ_API_URL), else Checkov against the
   plan as a drop-in replacement (--framework terraform_plan). Wiz and
   Checkov are NEVER both run on the plan. Replaces the old single
   Checkov-on-main.tf step in run_platform.sh Step 5 + run_postapply.sh
   Step 5.

pipelines/contract.yml: stage list updated — 'checkov' stage replaced by
'checkov-static' (before terraform-plan) + 'runtime-policy-scan' (after
terraform-plan). 9 stages → 10 stages. Header comment updated.

adapters/wiz/wiz_adapter.py: add --plan mode CLI (fetch_and_adapt_plan)
for scanning a terraform plan; backward-compat with the positional
<wiz_issues.json> <contract-id> mode. is_configured() gates the Wiz path.

Tests: test_pipeline_contract.py (9 → 10 stages, new stage names);
test_contract_resolver.py (rename test, assert checkov-static +
runtime-policy-scan present, old 'checkov' gone). Full suite: 685 pass
+ 1 pre-existing attestation failure (NOVA_ATTESTATION_SIGNING_KEY_ID
unset, unrelated to v1.21, fails on main without these changes too).

---ci---
project: acdl
phase: 4
milestone: v1.21
status: execute
phase_role: execution
---/ci---
2026-08-11 14:10:42 +00:00
Jon Chery 301aa2c8d8 docs(P3): marp deck + talking points + README + theme CSS + tests (REQ-245,251,252)
Nova Slides Render / render (push) Failing after 58s
Marp deck (nova-autonomous-cloud-delivery-marp.md): synthesize from updated
source-of-truth; 18 main + 1 appendix slides; frontmatter — title 'Nova —
The Autonomous Cloud Delivery Platform', footer without version + without
'Act %{page}/5', title-slide subtitle 'Product Development & Citizen
Developer Overview'; no badges; embedded PNGs.

Talking points (nova-autonomous-cloud-delivery-talking-points.md):
re-distilled to 18-slide + A1 structure.

README.md: update deck title, audience, slide count (18 main + 1 appendix),
directory layout, remove badge docs, update deck table + render commands +
filenames. Document the v1.21 rename + restructure.

Theme CSS (nova-sp-theme.css): fix Appendix A1 table readability — tables
now have explicit white body + black text on any slide background
(including dark/title slides). Item 32.

Tests (test_slides_pipeline.py): add v1.21 assertions — no badges; no
version in footer/title slide; 18 main + 1 appendix slides; no D-###/REQ-
###/.py paths in audience-facing Marp deck or source slide body; old deck
files removed; render script default renamed; README references new deck
name. Update deck path in test_regression_cap023_024.py +
core/regression_verify.py CAP-024 (filename + 18-19 slide range, drop 'Arc
Preview' check per item 3).

attach_release_asset.py: usage example filename updated.

---ci---
project: acdl
phase: 3
milestone: v1.21
status: execute
phase_role: execution
---/ci---
2026-08-11 14:03:06 +00:00
Jon Chery 707a7dbe9b docs(P2): slides source-of-truth — rename + restructure + rewrite (REQ-245,248,249,252)
Nova Slides Render / render (push) Failing after 1m3s
Rename all 5 deck files nova-no-humans-platform* →
nova-autonomous-cloud-delivery* (source, marp, html, pptx, talking-points).

Rewrite the source of truth to 18 main + 1 appendix slides, 4-beat arc
(Problem → Solution → Proof → Roadmap + Ask). All 33 review notes applied:

- Slide 1 'The Problem' (items 3,4,5,7,9): broader problem framing — devs
  writing terraform, destructive changes, AI-era 0-day pace, bandwidth
  gaps, tribal knowledge/rockstar operator. No arc. No '18 capabilities
  verified'. Not 'humans are the problem'.
- Slide 2 'Nova's Vision' (item 11): 'invisible' → 'visible' (operations
  become visible — recurring theme); polish for technical audience.
- Slide 3 'Strategic Objectives + Anti-Goals' (items 12,13,14,15,16,17,
  18): only Obj+Anti-Goals; provable trust = deterministic scripts
  (functions without AI); ROI = 4 CTO metrics (Lead Time, Vuln Count,
  MTTR, Spend); drop anti-goals 1,4,5; add 'not upstream dev platform',
  'not PDLC replacement'; obj #4 = integration objective; reword benefit.
- Slide 4 'Scope' (item 29): moved up, refined.
- Slide 5 'RACI' (item 30): moved up; add Quality Engineering column;
  reassign A from Platform → QE/SRE; rename Release Mgmt → SRE; split
  release attestation (Quality attestation + Production readiness).
- Slide 6 'Pipeline' (item 20): Checkov on static code before plan;
  Wiz-or-Checkov on plan; never both.
- Slide 7 'Decision Ledger' (items 21,22): drop D-121/122/132; 'AI
  decisions = automated decisions'; value = immutable/queryable/
  accountable, not sqlite/hash-chain.
- Slide 8 'Attestation Matrix' (item 23): drop bullets below table; add
  Description column per concern; drop 'operator-supplied' label.
- Slide 9 'Telemetry & Live Ops' (item 25): expand on value; drop
  D-120/125/126; expand on PowerBI live ops dashboard.
- Slide 10 'Decision Ledger + Attestation Coverage' (item 27):
  mandatory by design; no prod change without either; queryable for
  auditing; full traceability.
- Slide 11 'Cost & ROI': minor polish; 4 CTO metrics referenced.
- Slide 12 'What's Deferred' (items 10,19): remove all D-IDs; plain-
  language blockers; no status column.
- Slide 13 'Roadmap to the North Star' (items 10,19): drop D-IDs; no
  status column; timeframe-based roadmap.
- Slide 14 '12-Month Product Roadmap' (item 24): drop planned badges.
- Slide 15 'Quarter-by-Quarter' (item 24): drop badges.
- Slide 16 'Atelier (1/2)' (item 31): split — Skills + MCP server overview.
- Slide 17 'Atelier (2/2)' (item 31): split — agentic validation beyond
  deterministic scanners + vendoring.
- Slide 18 'Recap + Ask': refresh recap to 4-beat structure.
- Appendix A1 'Metrics Glossary' (item 32): kept; theme CSS fix in P3.
- Global (items 6,10,24,2): tech-leadership benefits; no D-###/REQ-###/
  .py paths in audience slides; no badges; no version in footer; final
  'less is more' prose pass.

Removed: old Slide 10 (Capability Health), old Slide 12 (Zero-Touch),
old Appendix A2 (Operating Model & Cost). Slide 5 first table removed.

---ci---
project: acdl
phase: 2
milestone: v1.21
status: execute
phase_role: execution
---/ci---
2026-08-11 13:59:30 +00:00
Jon Chery e7866fda84 docs(P1): strategic docs — thesis rename + NORTH_STAR objectives + RACI restructure
Nova Slides Render / render (push) Failing after 1m4s
AUTONOMY_THESIS.md (git mv from NO_HUMANS_THESIS.md): reframe from
'removing humans' to 'autonomy in operations, human at stage gates'.
Drop D-### citations + internal file paths; keep anti-claims, reworded.
Anti-claim #1 now: 'decisions are NOT made by an LLM — deterministic
scripts calculate a score; the platform functions without AI'.

NORTH_STAR.md:
- Vision: 'invisible' → 'visible' (operations become visible — recurring
  theme); polish for technical audience (security, remediation velocity,
  reliability, lead time).
- Objective #2: 'provable trust in AI decisions' → 'provable trust in
  automated decisions' (deterministic scripts calculate a score;
  platform functions without AI).
- Objective #3: four CTO-grade metrics (Lead Time PR→Prod, Infra Vuln
  Count trend, MTTR, Cloud Spend Reduction) → all flow into PowerBI.
- Objective #4: 'default substrate for agentic consumption' → integrate
  with externally owned PDLC/SDLC/Agentic/Citizen Developer platforms
  regardless of source; Nova provides skills + MCP endpoints; all prod
  intents go through the same controls + quality gates.
- Anti-goals: drop #1 (hyperscaler competitor), #4 (legacy untagged),
  #5 (sold to operators). Add: 'not an upstream development platform',
  'not a replacement for the PDLC'. Reword #3 (no 'removes humans').

docs/raci.md: 3 roles → 4 roles. Add Quality Engineering column. Rename
Release Management → SRE. Split release attestation into Quality
attestation (QA) + Production readiness (SRE). Platform no longer holds
A for attestation — reassigned to QE/SRE.

docs/scope.md: add integration framing (skills + MCP endpoints, all
sources go through same controls).

Render scripts: default deck name → nova-autonomous-cloud-delivery.
ONBOARDING + terraform/onboarding: 'no-humans' → 'autonomous'.

---ci---
project: acdl
phase: 1
milestone: v1.21
status: execute
phase_role: execution
---/ci---
2026-08-11 13:55:53 +00:00
Jon Chery 2efed26bb6 docs(P00): create phase plans — v1.21 (7 phases)
Nova Slides Render / render (push) Failing after 1m5s
PLAN.md v1.21 section: 5 execution phases + 1 final. Wave 1 parallelizable
(P1 strategic-docs, P2 slides, P4 pipeline-hardening — zero file overlap),
Wave 2 (P3 marp+README), Wave 3 (P5 render+verify), Wave 4 (P6 ship).
NFR milestone → tags on v1.20.x line (v1.20.0 P0 → v1.20.6 P6 final).

CLARIFY + RESEARCH minimal at full autonomy: domain is known, requirements
confirmed with user (deck title = Autonomous Cloud Delivery Platform;
thesis = AUTONOMY_THESIS.md; slide 1 = Problem→Solution→Proof→Roadmap+Ask;
Atelier split into 2 slides; CTO metrics = Lead Time + Vuln Trend + MTTR +
Spend; files renamed to nova-autonomous-cloud-delivery*).

---ci---
project: acdl
phase: 0
milestone: v1.21
status: plan
---/ci---
2026-08-11 13:52:40 +00:00
Jon Chery 5c07e29b90 docs(P00): validate specification — v1.21 milestone (REQ-245..253)
Add v1.21 requirements section (Nova Deck Refinement & Pipeline Hardening):
REQ-245 deck rename + restructure; REQ-246 thesis rename + reframe;
REQ-247 strategic-docs sync (integration objective); REQ-248 RACI
restructure (QE + SRE); REQ-249 Atelier split; REQ-250 pipeline hardening
(Checkov before plan, Wiz-or-Checkov on plan); REQ-251 theme CSS fix +
footer cleanup; REQ-252 global citation/badge/version removal; REQ-253
render + verify + ship.

Set active_milestone=v1.21 in config.json. Sync PROJECT.md strategic-
direction pillar for the integration objective (Objective #4 reframed),
deterministic-trust reword (Objective #2), CTO-grade ROI metrics
(Objective #3), and anti-goal updates.

---ci---
project: acdl
phase: 0
milestone: v1.21
status: specify
---/ci---
2026-08-11 13:51:55 +00:00
Jon Chery aa868c97ef docs(milestone): complete v1.20 — Consumer Cleanup + Transparent Terraform + Slide Pipeline
15 requirements complete (REQ-230..244):
- P1: gitea/gitlab removed from all synced files (REQ-230,231,232)
- P2: S&P theme CSS + render_slides.sh + CI workflow + tests (REQ-239..243)
- P3: 12-month product roadmap slides added to deck (REQ-244)
- P4: run_platform.sh split + var.enabled feature flags + stale path fix (REQ-233..238)

---ci---
project: acdl
phase: 5
milestone: v1.20
status: complete
phase_role: final
requirements:
  covered: [REQ-230,REQ-231,REQ-232,REQ-233,REQ-234,REQ-235,REQ-236,REQ-237,REQ-238,REQ-239,REQ-240,REQ-241,REQ-242,REQ-243,REQ-244]
  partial: []
---/ci---
2026-08-07 18:54:19 +00:00
Jon Chery e4a9915891 docs(P5): checkpoint — verify stage
---ci---
project: acdl
phase: 5
milestone: v1.20
status: verify
phase_role: final
---/ci---
2026-08-07 18:50:00 +00:00
Jon Chery 0ca383dae6 feat(P4): transparent terraform + feature flags + run_platform.sh split (REQ-233..238)
Create run_codegen.sh (pre-TF: env check, validate, resolve, adapt).
Create run_postapply.sh (post-TF: Checkov, confidence, HITL, outbox, SSM, uptime).
Add variable 'enabled' (bool, default true) + count=var.enabled?1:0 to all 12
L1 modules (alb, cloudfront, ecr, ecs-cluster, ecs-service, iam-role, kms-key,
rds, s3, uptime, vpc, waf). Fix all cross-resource references with [0] indexing.
Update interface.json for all modules to declare 'enabled' input.
Fix stale artifact path /tmp/acdl_platform_run_v18 → /tmp/nova_platform_run (REQ-238).
run_platform.sh remains as backward-compat shim for local-dev usage.

---ci---
project: acdl
phase: 4
milestone: v1.20
status: execute
requirements: [REQ-233, REQ-234, REQ-235, REQ-236, REQ-237, REQ-238]
---/ci---
2026-08-07 18:49:56 +00:00
Jon Chery ed5ea90654 feat(P3): add 12-month product roadmap slides (REQ-244)
Slide 20 — 12-Month Product Roadmap: 4-quarter arc (Pilot Activation →
Provable Trust → Compounding ROI → Agentic Substrate).
Slide 21 — Quarter-by-Quarter Outcomes: detail table (theme, deliverable,
target metric, strategic-objective grounding).
Both grounded in NORTH_STAR's 4 strategic objectives + deferred-metric
candidate milestones. Distinct from Slide 15's deferred-metric unblock paths.
Matching talking-points sections added. HTML + PPTX re-rendered via S&P theme.

---ci---
project: acdl
phase: 3
milestone: v1.20
status: execute
requirements: [REQ-244]
---/ci---
2026-08-07 18:28:01 +00:00
Jon Chery 2273009b95 feat(P2): dedicated S&P theme + render pipeline + CI workflow (REQ-239..243)
Create nova-sp-theme.css — S&P Global Energy Marp theme (Red/Black/White
palette applied to all slide chrome: backgrounds, headers/footers, pagination,
tables, blockquotes, code blocks).
Create render_slides.sh — end-to-end pipeline: mermaid PNGs + Marp HTML/PPTX.
Create slides.yml CI workflow — auto-renders on docs/presentations/ changes.
Create test_slides_pipeline.py — 12 tests (theme CSS, Marp frontmatter, script,
workflow, .mmd/.png parity, README retired-deck cleanup).
Update Marp frontmatter: theme: nova-sp + footer v1.20.
Fix presentations/README.md directory layout (remove retired decks).
Re-render HTML + PPTX with S&P theme.

---ci---
project: acdl
phase: 2
milestone: v1.20
status: execute
requirements: [REQ-239, REQ-240, REQ-241, REQ-242, REQ-243]
---/ci---
2026-08-07 18:26:51 +00:00
Jon Chery 0d2cbdb423 feat(P1): remove gitea/gitlab from synced files + simplify docs (REQ-230,231,232)
Genericize forge-detection code: gitea→forge/generic_forge, GITEA_ACTOR→FORGE_ACTOR.
Drop .gitea byte-identity test assertions (keep GitHub-side + contract conformance).
Add test_no_forge_mentions.py guard test (REQ-230).
Delete completed migration docs (NOVA_MIGRATION.md, NOVA_AWS_MIGRATION.md).
Move NO_HUMANS_THESIS.md to .ciagent/ (internal artifact).
Strip ciagent-internal provenance from synced docs (REQ-/D-/P-/CAP- IDs,
milestone headers, .ciagent/PROJECT.md citations).
Trim README.md (reusable deploy section, local key rotation paragraph).
Fix version-tag drift (@v1.13→@v1.19, acdl/→nova/).

---ci---
project: acdl
phase: 1
milestone: v1.20
status: execute
requirements: [REQ-230, REQ-231, REQ-232]
---/ci---
2026-08-07 18:20:29 +00:00
Jon Chery b418d429b5 docs(ship): P0 complete — v1.20 pre-execution
---ci---
project: acdl
phase: 0
milestone: v1.20
status: complete
phase_role: pre_execution
---/ci---
2026-08-07 18:02:30 +00:00
Jon Chery dcba380b52 docs(P00): create phase plans — v1.20 (5 phases)
---ci---
project: acdl
phase: 0
milestone: v1.20
status: plan
---/ci---
2026-08-07 18:02:27 +00:00
Jon Chery f0bc3be92c docs(P00): validate specification — v1.20 milestone (REQ-230..244)
---ci---
project: acdl
phase: 0
milestone: v1.20
status: specify
---/ci---
2026-08-07 18:02:23 +00:00
Jon Chery 0b79b16715 fix(P2): add metrics domain to sync_to_nova.sh — consumer export views (REQ-229)
acdl-ci / Lint (push) Successful in 9s
acdl-ci / Test (push) Failing after 22s
acdl-ci / Platform check-only (offline) (push) Successful in 23s
The metrics/ export views (README.md, TRUST_SNAPSHOT.md, powerbi/) are
consumer-facing but fell outside the original 13 domains, so the first nova
release left them untracked. Adds a 14th domain 'metrics' between docs and
workflows. Updates TestSyncToNovaScript domain-order assertion to 14.

---ci---
project: acdl
phase: 2
milestone: v1.19
status: complete
phase_role: final
---/ci---
2026-08-06 15:47:25 +00:00
Jon Chery 90624be63f docs(milestone): complete v1.19 — Nova 2nd-Release Sync
acdl-ci / Lint (push) Successful in 8s
acdl-ci / Test (push) Failing after 23s
acdl-ci / Platform check-only (offline) (push) Successful in 32s
NFR-only chore milestone complete. P1 (nova-sync-script, v1.18.0) + P2
(final-review-ship, v1.18.1 = milestone release). REQ-229 satisfied.
Review clean, audit clean, 5 decisions locked (D-143..D-147).

---ci---
project: acdl
phase: 2
milestone: v1.19
status: complete
phase_role: final
requirements:
  covered: [REQ-229]
  partial: []
---/ci---
2026-08-06 15:44:31 +00:00
Jon Chery be51fc15fa docs(P1): ship complete — checkpoint update (Gitea release id 530)
acdl-ci / Lint (push) Successful in 9s
acdl-ci / Test (push) Failing after 24s
acdl-ci / Platform check-only (offline) (push) Successful in 24s
---ci---
project: acdl
phase: 1
milestone: v1.19
status: complete
phase_role: execution
requirements:
  covered: [REQ-229]
  partial: []
---/ci---
2026-08-06 15:43:12 +00:00
Jon Chery e3f4ce17d4 verify(P1): 4-layer verify PASS + ship — sync_to_nova.sh (REQ-229)
acdl-ci / Lint (push) Successful in 10s
acdl-ci / Test (push) Failing after 23s
acdl-ci / Platform check-only (offline) (push) Successful in 24s
L1 structural: bash -n clean, shellcheck 0 warnings.
L2 behavioral: manual gate exits 2 without --release; --list-domains prints
13 ordered domains; rsync exclude list correct; .git protected via filter.
L3 security: no hardcoded secrets; .coverage runtime artifact gitignored.
L4 quality: 8/8 TestSyncToNovaScript tests pass (gate, domain order, exclude
list, consumer-script inclusion, .git filter, conventional regex).

---ci---
project: acdl
phase: 1
milestone: v1.19
status: verify
---/ci---
2026-08-06 15:42:50 +00:00
Jon Chery e0d01ad2ef docs(P1): checkpoint — execute complete (REQ-229)
acdl-ci / Lint (push) Successful in 10s
acdl-ci / Test (push) Failing after 24s
acdl-ci / Platform check-only (offline) (push) Successful in 23s
---ci---
project: acdl
phase: 1
milestone: v1.19
status: execute
---/ci---
2026-08-06 15:40:22 +00:00
Jon Chery a4c5f332f6 feat(P1): sync_to_nova.sh — manual-only 2nd-release pipeline into ~/nova (REQ-229)
Replaces scripts/sync_to_gl.sh (kitchen-sink mirror sync into ~/gl/acdl) with
scripts/sync_to_nova.sh — a manual-only, consumer-subset, domain-committed
2nd-release pipeline into ~/nova (GitLab jonathanchery/nova, separate repo +
history, consumer/platform-team audience).

- Manual-only gate: refuses without --release / RELEASE_CONFIRMED=1 (exit 2).
  Never triggerable by CI.
- Consumer subset: excludes .ciagent/, .gitea/, .env*, terraform/, demo/,
  runtime metrics artifacts, and 18 internal-only scripts (EXCLUDE_SCRIPTS).
  Keeps consumer runbooks + metrics export views (README, powerbi,
  TRUST_SNAPSHOT). Protects ~/nova/.git via rsync --filter=P .git.
- Domain-based commits: 13 fixed-order domains (config, core, adapters,
  modules, contracts, schemas, pipelines, mcp, skills, scripts, tests, docs,
  workflows). Each changed domain gets its own conventional commit supplied
  positionally via repeated -m flags. No kitchen-sink commit.
- Conventional-commit validation: regex-enforced (feat|fix|docs|chore|...);
  bypass via --no-verify-format.
- Modes: --list-domains, --dry-run, --no-push, -v, -h.
- Tests: TestSyncToNovaScript (8 tests) covers gate, domain order, exclude
  list, consumer-script inclusion, .git protection filter, conventional
  regex.

Decisions: D-143 (target ~/nova), D-144 (conventional commits per domain,
not ---ci--- audit blocks), D-145 (manual-only trigger), D-146 (13 fixed
domains, positional-over-changed mapping), D-147 (coreci/Atelier review
gate deferred).

---ci---
project: acdl
phase: 1
milestone: v1.19
status: execute
requirements:
  covered: [REQ-229]
  partial: []
---/ci---
2026-08-06 15:40:11 +00:00
Jon Chery 9e20b7ba95 docs(ship): v1.17.7 milestone complete — checkpoint update (Gitea release id 529)
acdl-ci / Lint (push) Successful in 9s
acdl-ci / Test (push) Failing after 24s
acdl-ci / Platform check-only (offline) (push) Successful in 23s
2026-08-06 15:17:46 +00:00
Jon Chery 6da538c936 Merge milestone/v1.18-citizen-developer-guidance — v1.18 complete (Citizen Developer & Production-Grade Guidance: 5 inputs, 15 requirements, 7 phases + final; tag v1.17.7)
acdl-ci / Lint (push) Successful in 10s
acdl-ci / Test (push) Failing after 26s
acdl-ci / Platform check-only (offline) (push) Successful in 24s
2026-08-06 15:17:03 +00:00
Jon Chery 4e03817ea6 Merge phase/07-final-review-ship — v1.17.7 (v1.18 P7 final review + audit + milestone complete) 2026-08-06 15:16:59 +00:00
Jon Chery 951ad56576 docs(milestone): complete v1.18 — Citizen Developer & Production-Grade Guidance
15 requirements (REQ-214..228) satisfied. 32 tests pass. S&P Global theme
restored. PDLC-upstream scope + RACI matrix authored. Submission-readiness
schema + validator shipped. 9 Atelier skills + docs/skills.md. MCP server
(plugin-registry, stdio, vendored Atelier v0.3.6) with 4 tools + agentic
validation. 21-slide deck (3 new: scope/RACI/atelier) with PPTX committed +
release-attached. 10 decisions locked (D-133..D-142).

---ci---
project: acdl
phase: 7
milestone: v1.18
status: complete
requirements:
  covered: [REQ-214, REQ-215, REQ-216, REQ-217, REQ-218, REQ-219, REQ-220, REQ-221, REQ-222, REQ-223, REQ-224, REQ-225, REQ-226, REQ-227, REQ-228]
  partial: []
---/ci---
2026-08-06 15:16:54 +00:00
Jon Chery d882cf0c6e Merge phase/06-deck-slides-atelier — v1.17.6 (v1.18 P6 deck slides + atelier complete) 2026-08-06 15:15:22 +00:00
Jon Chery 564d4a4ca3 docs(P6): atelier deck slide + 21-slide re-render + README (REQ-226, REQ-227, REQ-228)
REQ-226: Slide 19 'Production-Grade Guidance via Atelier' added → 21 total
slides (16 existing + 17 Scope + 18 RACI + 19 Atelier + 2 appendix). Arc
preview updated (v1.18). Talking points synced (slide 19). S&P theme
preserved (177 color refs in HTML). PPTX 22 slides (21 content + title).

REQ-227: README deck table updated — single unified deck, 21 slides, PPTX
committed + release-attached (D-141). Old two-deck table replaced.

REQ-228: HTML + PPTX re-rendered via scripts/render_deck.sh. PPTX committed
(binary, no LFS).

---ci---
project: acdl
phase: 6
milestone: v1.18
status: execute
requirements:
  covered: [REQ-226, REQ-227, REQ-228]
  partial: []
---/ci---
2026-08-06 15:15:17 +00:00
Jon Chery c524ad731e Merge phase/05-atelier-mcp — v1.17.5 (v1.18 P5 Atelier MCP server complete) 2026-08-06 15:13:44 +00:00
Jon Chery 8bcf7296d5 feat(P5): Atelier MCP server + vendored Atelier + plugin-registry (REQ-223, REQ-224, REQ-225)
REQ-223: mcp/atelier/server.py plugin-registry MCP server (stdio, D-135).
NovaAtelierServer wraps MCPServer (SDK v2, D-137) if installed; degrades
to _ToolRegistry fallback if SDK absent (testable in CI without SDK).
plugins/principles.py (lookup_principle, list_domains, matrix_lookup) +
plugins/validation.py (validate_against_principles — agentic validation
beyond Wiz/Checkmarx/Mend). 4 tools, 2 plugins.

REQ-224: mcp/atelier/vendor/ pinned Atelier v0.3.6 (D-136) — core/
first-principles, domains/security/first-principles, review/agent-checklist,
matrix/principles-matrix. vendor/VERSION.md + scripts/update_atelier_vendor.sh
for intentional upgrades. mcp/atelier/README.md (tools, architecture,
running, vendoring, extensibility, transport).

REQ-225: tests/test_atelier_mcp.py — 16 tests, all pass. Covers: plugin
discovery (both loaded), 4 tools registered, lookup_security_P4 (+P1,
unknown domain/principle), list_domains (19, security-relevant, ui-ux-not),
matrix_lookup (security 10 P-rules, unknown), validation (good-passes,
bad-secret-fails, bad-swallowed-error-fails, bad-obfuscated-names-fails,
result-structure).

---ci---
project: acdl
phase: 5
milestone: v1.18
status: execute
requirements:
  covered: [REQ-223, REQ-224, REQ-225]
  partial: []
---/ci---
2026-08-06 15:13:40 +00:00
Jon Chery 81c7a22ddd Merge phase/04-atelier-skills — v1.17.4 (v1.18 P4 Atelier skills complete) 2026-08-06 15:11:15 +00:00
Jon Chery 2c08c778a9 docs(P4): Atelier skills mapping — 9 skill files + index + BA.A extension (REQ-221, REQ-222)
REQ-221: skills/ directory with 9 Atelier-derived skill files mapped to the
BA.A citizen-developer catalog: api, security, data, testing, observability,
errors, devops, infrastructure-as-code, compliance. Each names the Atelier
source path, distills first-principles to the citizen-dev-relevant subset,
links to agent-checklist triggers, maps to BA.A 5-skill catalog.

REQ-222: docs/skills.md index (9-skill table, Atelier provenance, 8 core
principles C1-C8, consumption instructions, reference-only domains, excluded
domains). PROJECT.md BA.A decision extended with the Atelier-derived skill
catalog reference.

---ci---
project: acdl
phase: 4
milestone: v1.18
status: execute
requirements:
  covered: [REQ-221, REQ-222]
  partial: []
---/ci---
2026-08-06 15:11:12 +00:00
Jon Chery 6ffcbe8283 Merge phase/03-submission-readiness — v1.17.3 (v1.18 P3 submission-readiness complete) 2026-08-06 15:09:32 +00:00
Jon Chery 5775a97388 feat(P3): submission-readiness input contract — schema + validator + docs + tests (REQ-217..220)
REQ-217: schemas/submission-readiness.schema.json (JSON Schema draft 2020-12)
defines acceptable-to-start as a superset gate above contract.schema.json:
contractId, environment, tags (5 Nova tags D-054), policyPreconditions,
profile (developer|agentic), appSource (repo+ref), per-env mandatory (W3.E:
qa→e2eSuite+loadTest, prod→runbook+dashboard+oncall, dr→drDrillRef),
agentic markers (naturalLanguageIntent+confidenceAtSubmission+agentTrace).

REQ-218: core/submission_readiness.py validator with check_readiness() +
ReadinessResult (structured pass/fail + reason codes). Wired as
contract_ingestor.py --check-readiness (D-133). Reason codes: MISSING_TAGS,
ENV_MISSING_MANDATORY, AGENTIC_MISSING_INTENT, MISSING_APP_SOURCE,
POLICY_PRECONDITION_MISSING. Never raises — all failures are reason codes.

REQ-219: docs/submission-readiness.md (good + rejected examples +
reason-code catalog + compliance-standard equivalence).

REQ-220: tests/test_submission_readiness.py — 16 tests, all pass.
Covers: good-pass, good-agentic-pass, missing-tags, empty-tag,
qa-missing-e2e, prod-missing-runbook, dr-missing-drdrill, prod-all-pass,
agentic-missing-all, agentic-missing-one, missing-appsource,
appsource-missing-ref, empty-policy, result-structure.

---ci---
project: acdl
phase: 3
milestone: v1.18
status: execute
requirements:
  covered: [REQ-217, REQ-218, REQ-219, REQ-220]
  partial: []
---/ci---
2026-08-06 15:09:29 +00:00
Jon Chery b3c75ccec1 Merge phase/02-pdlc-scope-raci — v1.17.2 (v1.18 P2 PDLC scope + RACI complete) 2026-08-06 15:07:14 +00:00
Jon Chery e891496163 docs(P2): PDLC-upstream scope + RACI matrix + 2 deck slides (REQ-215, REQ-216, REQ-228)
REQ-215: RACI matrix in PROJECT.md (§ RACI Matrix) + docs/raci.md
(citizen-dev-facing copy). 3 roles (Citizen Developer / Platform / Release
Management co-owned). 7 work categories × R/A/C/I. Compliance-standard
equivalence note: any upstream source (AI agent, SDLC, dev platform) is
subject to the same gate.

REQ-216: PDLC-upstream scope in PROJECT.md (§ Scope) + docs/scope.md.
Promotes Core Tenet #2 + Anti-Goal #1 from buried tenets to a dedicated,
unmissable scope statement.

REQ-228: 2 new deck slides (17 Scope + 18 RACI) → 20 slides. Arc preview
updated. Talking points synced. HTML + PPTX re-rendered (21 PPTX slides).

---ci---
project: acdl
phase: 2
milestone: v1.18
status: execute
requirements:
  covered: [REQ-215, REQ-216, REQ-228]
  partial: []
---/ci---
2026-08-06 15:07:10 +00:00
Jon Chery 382944c055 Merge phase/01-sp-theme-restoration — v1.17.1 (v1.18 P1 S&P theme restoration + PPTX automation complete) 2026-08-06 15:05:09 +00:00
Jon Chery 71b6a4fa91 feat(P1): restore S&P Global Energy theme + PPTX automation (REQ-214, REQ-228)
REQ-214: Restore the S&P Global Energy Marp style: block (from commit
ae0cb58 / v1.9.2 P45) to the unified deck. Colors: H1/H2 #D6002A (red-core),
title-slide bg #1B1B1B (grey-90) + 8px #D6002A top accent, body #1B1B1B,
blockquote border #D6002A, table headers #F0F0F0, font 'Akkurat Pro' with
web-safe fallbacks. Nova header/footer text preserved (rebrand not touched).
HTML re-rendered (229 S&P color refs confirmed).

REQ-228: scripts/render_deck.sh (HTML + PPTX render + git add) +
scripts/attach_release_asset.py (Gitea release asset upload via API). PPTX
is now a first-class committed binary (D-141, no LFS). README updated:
'PPTX not committed' → 'PPTX committed + attached'. PPTX committed (3.6 MiB,
19 slides).

---ci---
project: acdl
phase: 1
milestone: v1.18
status: execute
requirements:
  covered: [REQ-214, REQ-228]
  partial: []
---/ci---
2026-08-06 15:05:01 +00:00
Jon Chery 0f677641ee Merge phase/00-pre-execution — v1.17.0 (v1.18 P0 pre-execution complete: specify+clarify+research+plan+grill) 2026-08-06 15:03:16 +00:00
Jon Chery e3ebbc4978 docs(P00): checkpoint — plan complete 2026-08-06 15:03:13 +00:00
Jon Chery 37b6b6fc14 docs(P00): grill — v1.18 plan PASS (full autonomy, user-directed + research-grounded)
Self-grill at full autonomy. Plan is user-directed (5 explicit inputs),
research-confirmed (9 assumptions A1-A9, conf 0.80-0.95), decisions locked
(D-133..D-142). No binding changes. 4 challenges reviewed:

G-201 (MCP scope-creep?) — NO. User explicitly requested MCP + extensible.
G-202 (submission-readiness duplicates contract.schema.json?) — NO. Research
    confirms superset gate (shape vs readiness). D-133 locks the wiring.
G-203 (21 slides too many?) — NO. 3 new slides are leadership-relevant;
    5-act arc preserved (D-134). Fallback if grilled: merge RACI+atelier → 20.
G-204 (Atelier vendoring reproducibility?) — YES, required. D-136 locks
    vendoring for audit replayability.

Verdict: PASS-with-binding (0 BIND, 0 ESCALATE).

---ci---
project: acdl
phase: 0
milestone: v1.18
status: grill
---/ci---
2026-08-06 15:03:04 +00:00
Jon Chery d61a3d1a2f docs(P00): create phase plans — v1.18 8 phases, 6 waves, 15 requirements
Vertical-slice plan for v1.18 Citizen Developer & Production-Grade Guidance.
Wave order: W1=P1, W2=P2, W3=P3+P4 (parallelizable), W4=P5, W5=P6, W6=P7.
Sequential execution this run. 8 plan-level risks documented (conf 0.80-0.92).

---ci---
project: acdl
phase: 0
milestone: v1.18
status: plan
---/ci---
2026-08-06 15:02:53 +00:00
Jon Chery 4c8b2b77fc docs(P00): research findings — v1.18 Atelier integration + MCP SDK + submission-readiness + Marp PPTX
5 research targets completed:
- Atelier: 19 domains → 9 Nova skills (REQ-221); agent-checklist → MCP validation; principle-lookup model; pin tag v0.3.6
- MCP Python SDK v2: MCPServer + @mcp.tool() + plugin-registry skeleton (D-140)
- Submission-readiness: superset gate confirmed (contract.schema.json defines shape only; readiness adds tags/env/policy/profile/appSource)
- Marp PPTX: inline style: CSS survives --pptx export (no fallback needed)
- Personas: 3 active (lead/backend/data) + frontend deactivated; mcp-engineer folded into backend (D-143, 0.90)

9 assumptions logged (A1-A9, conf 0.80-0.95).

---ci---
project: acdl
phase: 0
milestone: v1.18
status: research
---/ci---
2026-08-06 14:59:18 +00:00
Jon Chery 1daae0ac0a docs(P00): clarify — v1.18 decisions D-133..D-142 locked
10 decisions resolved at full autonomy:
- D-133: validator extends contract_ingestor.py --check-readiness
- D-134: deck 18→21 slides (no act restructure)
- D-135: MCP stdio now; HTTP-ready (same server object)
- D-136: vendor Atelier (pinned tag, audit reproducibility)
- D-137: MCP Python SDK v2
- D-138: skill format = markdown under skills/
- D-139: RACI roles = Citizen Dev / Platform / Release Mgmt (co-owned)
- D-140: MCP plugin-registry (plugins/<name>.py register(mcp))
- D-141: PPTX committed binary (no LFS)
- D-142: deck render trigger on any marp/assets change

---ci---
project: acdl
phase: 0
milestone: v1.18
status: clarify
---/ci---
2026-08-06 14:54:42 +00:00
Jon Chery d048460abf docs(init): validate specification — v1.18 Citizen Developer & Production-Grade Guidance
Establish v1.18 active milestone (was v1.17 complete). Author 15 new
requirements (REQ-214..228) across 5 user-directed inputs: S&P Global
theme restoration, PDLC-upstream scope, RACI matrix, Nova input contract
(submission-readiness schema + validator), Atelier integration (skills +
MCP server). Add v1.18 objective to PROJECT.md + ROADMAP.md. Feature
milestone; tags run on v1.17.x patch line.

---ci---
project: acdl
phase: 0
milestone: v1.18
status: specify
---/ci---
2026-08-06 14:54:22 +00:00
Jon Chery 0ad6a88c4b docs(P5): render unified deck to HTML (Step 3 of 4-step deck process)
acdl-ci / Lint (push) Successful in 8s
acdl-ci / Test (push) Failing after 23s
acdl-ci / Platform check-only (offline) (push) Successful in 22s
---ci---
project: acdl
phase: 5
milestone: v1.17
status: complete
---/ci---
2026-08-05 01:58:59 +00:00
Jon Chery eb5b24b88d Merge milestone/v1.17-direction-metrics-story — v1.17 complete (Strategic Direction, Leadership Metrics & Unified Story: 3 pillars, 29 requirements, 7 phases + final; tag v1.16.7)
acdl-ci / Lint (push) Successful in 14s
acdl-ci / Test (push) Failing after 29s
acdl-ci / Platform check-only (offline) (push) Successful in 26s
2026-08-04 20:09:15 +00:00
Jon Chery cb1a7071a7 Merge phase/07-final-review-ship — v1.16.7 (v1.17 P7 final review + audit + milestone complete) 2026-08-04 20:09:15 +00:00
Jon Chery e4adb3f09e docs(milestone): complete v1.17 — Strategic Direction, Leadership Metrics & Unified Story
---ci---
project: acdl
phase: 7
milestone: v1.17
status: complete
requirements:
  covered: [REQ-185..REQ-213]
  partial: []
---/ci---
2026-08-04 20:09:05 +00:00
Jon Chery 9415afc739 Merge phase/06-regression-capability — v1.16.6 (v1.17 P6 regression capability complete) 2026-08-04 20:08:03 +00:00
Jon Chery d9b402c283 test(P6): regression capability — CAP-023 (metrics collector) + CAP-024 (deck structure) (REQ-198)
P6 (Wave 4, test) — REQ-198

New capabilities:
- CAP-023: metrics collector runs + emits expected schema (fact/dim tables present)
- CAP-024: unified deck structure (12-20 slides, x3 arc, per-slide benefit callouts)
- tests/test_regression_cap023_024.py — 4 tests (all pass)

Modified:
- core/regression_verify.py — CAPABILITY_REGISTRY gains CAP-023 + CAP-024

---ci---
project: acdl
phase: 6
milestone: v1.17
status: execute
---/ci---
2026-08-04 20:08:03 +00:00
160 changed files with 9407 additions and 3811 deletions
+66
View File
@@ -0,0 +1,66 @@
# Nova — The Autonomous Cloud Delivery Platform: Autonomy Defensibility Brief
> Strategic direction, leadership metrics & unified story
> Last refined: v1.21 — reframe from "no-humans" to "autonomous operations"
## The thesis
Nova is the autonomous infrastructure layer that lets product teams
ship without engaging an operator, and lets executives trust the
platform not because it never fails but because every decision is
captured, scored, and accountable.
**Autonomy in operations; human at stage gates.** Normal operations —
provisioning, healing, remediation — run without an operator in the
loop. Human attestation remains required at stage gates: QA signs off
for production, SRE greenlights based on operational readiness. The
absence of an operator in the loop is never the absence of a record.
## Grounded proof (measurable today)
| Proof | Source | Status |
|-------|--------|--------|
| Capabilities verified, none broken (live-AWS caps honestly skipped, resources torn down to zero-cost steady state) | regression report | grounded |
| Decision Ledger captures 100% of automated decisions with outcome backfill | decision ledger store | grounded |
| Attestation coverage: 100% of prod/dr promotions attested by a human | attestation gates + outbox | grounded |
| Confidence-gated policy engine (deterministic, not an LLM) — weighted inputs, band outcome | confidence signal | grounded |
| Attestation matrix with separation-of-duties on prod | attestation matrix + separation-of-duties | grounded |
| Pre-apply cost estimates (offline) | cost adapter | grounded |
| Test suite passes | test results | grounded |
## Deferred proof (measurable when blocking work lifts)
| Proof | Blocking work | Unblock requirement |
|-------|----------------|---------------------|
| Touchless resolution rate across production estates | 0 consumers today | Pilot estate activation |
| Live infrastructure health (ECS, ALB, RPS) | Live AWS torn down | Live AWS re-provisioning |
| Onboarding funnel: requested → granted | Auto-grant not built | Auto-grant implementation |
| Drift auto-reversal rate | No drift scheduler | Drift detection scheduler |
| Predictive vs reactive ratio | No emitter | ML anomaly-forecasting service |
| Tamper-evident ledger checkpoints (S3 Object Lock + JWS) | Audit ledger build-out | Audit ledger build-out |
## Anti-claims (what Nova is NOT)
1. **Nova's decisions are NOT made by an LLM.** They are made by a
confidence-gated policy engine: deterministic scripts calculate a
score, and a band outcome gates the action. The platform functions
without AI. The Decision Ledger captures this real decision path —
not a fabricated "AI agent." When an LLM planner is added, it will
emit richer `alternatives_considered` without schema breakage.
2. **Nova does NOT remove humans from accountability.** Only from
normal operations. Every stage-gate promotion (qa/prod/dr) requires
a human attestation recorded with approver identity,
separation-of-duties check, and the evidence matrix.
3. **Nova is NOT for legacy, untagged, or freeform infrastructure.** It
requires Terraform-managed, policy-aligned, fully-tagged inputs.
4. **Nova does NOT fabricate metrics.** Every metric is grounded (cites
a source), derived (documented formula), or deferred (cites the
blocking work). No fabricated numbers in any deck slide or metrics
entry (the "no fabrication" hard constraint).
## What "won" looks like
By month 18, Nova is the layer enterprise leadership points to when
they say *"we don't have an infrastructure ops team anymore, and the
audit trail is stronger than it ever was"* — and it is the layer their
AI engineering teams reach for first when an agent needs to deploy.
+5 -7
View File
@@ -1,13 +1,11 @@
{
"phase": 0,
"stage": "complete",
"milestone": "v1.17",
"stage": "plan",
"milestone": "v1.21",
"phase_role": "pre_execution",
"attempts": 0,
"updated_at": "2026-08-04T21:30:00Z",
"updated_at": "2026-08-11T00:01:00Z",
"milestone_complete": false,
"tag": "v1.16.0",
"release_id": 441,
"requirements": ["REQ-185"],
"notes": "Phase 0 complete. NORTH_STAR.md authored. 29 requirements (REQ-185..213). Telemetry reference architecture + metric scorecard. Deck rebuild plan (18 slides). Interactive GRILL: 12 binding decisions applied. Tag v1.16.0 pushed. Gitea release 441 created. Ready for execution phases P1..P7 + final P8."
"requirements": ["REQ-245","REQ-246","REQ-247","REQ-248","REQ-249","REQ-250","REQ-251","REQ-252","REQ-253"],
"notes": "v1.21 P0 plan stage complete. PLAN.md v1.21 section written. 5 execution phases (P1 strategic-docs, P2 slides, P3 marp+talking-points+README, P4 pipeline-hardening, P5 render+verify) + P6 final-review-ship. Wave 1 (P1/P2/P4 parallelizable), Wave 2 (P3), Wave 3 (P5), Wave 4 (P6). CLARIFY+RESEARCH minimal at full autonomy — domain known, requirements confirmed with user. Proceeding to P0 ship then execution."
}
+57 -36
View File
@@ -1,7 +1,7 @@
# NORTH_STAR — Nova
> **Status:** Draft (pending interactive GRILL → final)
> **Milestone:** v1.17Strategic Direction, Leadership Metrics & Unified Story
> **Milestone:** v1.21 — Nova Deck Refinement & Pipeline Hardening
> **Owner:** Product Owner
> **Purpose:** Durable strategic intent. Read by CIAgent in every future
> `/ci-run` so the platform's direction survives across milestones. This
@@ -14,7 +14,7 @@
## Vision
> **Infrastructure operations become invisible. Every environment
> **Infrastructure operations become visible. Every environment
> provisioned, every incident healed, every risk remediated — by an
> autonomous system whose trustworthiness is provable, not promised.
> Human attestation remains required at stage gates — QA signs off for
@@ -22,9 +22,12 @@
> operator is never in the loop of normal operations.**
Nova is the autonomous infrastructure layer that lets product teams ship
without engaging an operator, and lets executives trust the AI not because
it never fails but because every decision is captured, scored, and
accountable.
without engaging an operator, and lets executives trust the platform not
because it never fails but because every decision is captured, scored,
and accountable. The recurring theme across the platform is that
**infrastructure operations become visible** — security posture,
remediation velocity, reliability, and lead time are surfaced as
queryable signals rather than hidden in tribal knowledge.
---
@@ -38,46 +41,64 @@ human by design; operational escalations (AI confidence too low to
proceed) are the failure mode we drive toward zero. Everything else
collapses if autonomy isn't real.
**2. Establish provable trust in AI decisions.**
Build the audit substrate — Decision Ledger, confidence scoring, circuit
breakers, blast-radius controls — that turns "autonomous" from a
marketing claim into a defensible one. Trust is the moat. Features can be
copied; an immutable, queryable decision history cannot.
**2. Establish provable trust in automated decisions.**
Trust is established by deterministic scripts that calculate a score and
a band outcome that gates the action — the platform functions without AI.
"AI decisions" are really automated decisions. The audit substrate —
Decision Ledger, confidence scoring, circuit breakers, blast-radius
controls — turns "autonomous" from a marketing claim into a defensible
one. Trust is the moat. Features can be copied; an immutable, queryable
decision history cannot.
**3. Deliver compounding, quantifiable ROI for customers.**
Each quarter on Nova must reduce cloud spend, free engineering hours, and
avoid downtime measurably. If the CFO can't point to a number that
improves quarter-over-quarter, Nova fails its commercial test, regardless
of how clever the AI is.
Each quarter on Nova must show measurable improvement on four CTO-grade
metrics, all of which flow into PowerBI views and are captured by the
telemetry pipeline:
**4. Become the default substrate for agentic infrastructure consumption.**
AI agents are already becoming the largest consumers of cloud
infrastructure. Nova must be the platform through which those agents
declare, deploy, and verify infrastructure — not a vendor scrambling into
that market two quarters late.
- **Lead Time** — from PR merge to production deployment (downward trend).
- **Infrastructure Vulnerability Count** — open findings on deployed
resources (downward trend, demonstrating that proactive scanning +
remediation keeps up with the AI-era 0-day pace).
- **MTTR** — for platform-detected and platform-remediated incidents.
- **Cloud Spend Reduction** — on pilot estates vs. the pre-Nova
baseline.
If leadership cannot point to a number that improves quarter-over-quarter
on these four axes, Nova fails its commercial test, regardless of how
clever the automation is.
**4. Integrate with externally owned development platforms — regardless of source.**
Nova integrates with externally owned PDLC, SDLC, Agentic, and Citizen
Developer platforms with no regard for the source of the intent. Nova
provides a set of skills and MCP endpoints that help the developer or AI
agent make their application production-grade. Regardless of the source,
all intents to deploy to production go through the same rigorous
controls, quality gates, attestation, and evidence stream. Nova is the
layer any of those platforms reach for first when an agent needs to
deploy — not a vendor arriving late to that market.
---
## Anti-Goals (5 — what Nova is fundamentally NOT)
## Anti-Goals (4 — what Nova is fundamentally NOT)
1. **Not a Terraform, Kubernetes, or hyperscaler competitor.** We
orchestrate them. Replacing them is the most expensive possible
distraction from the value we create.
2. **Not a general-purpose AI agent platform.** We are purpose-built for
1. **Not a general-purpose AI agent platform.** We are purpose-built for
infrastructure operations. Breadth here produces shallow tools; depth
here wins the category.
3. **Not a system that removes humans from accountability.** Only from
operations. Every AI decision lands in an immutable ledger. Every
stage-gate promotion (qa/prod/dr) requires a human attestation recorded
with approver identity, separation-of-duties check, and the 8-concern
evidence matrix. The absence of an operator is never the absence of a
record.
4. **Not for legacy, untagged, or freeform infrastructure.** Nova requires
Terraform-managed, policy-aligned, fully-tagged inputs. We optimize for
the disciplined 95%, not the chaotic 5%.
5. **Not sold to operators.** Nova is sold to leadership on outcomes —
cost, velocity, risk. Selling to operators inverts the incentive and
breaks the autonomy thesis.
2. **Not a system that removes humans from accountability.** Only from
normal operations. Every automated decision lands in an immutable
ledger. Every stage-gate promotion (qa/prod/dr) requires a human
attestation recorded with approver identity, separation-of-duties
check, and the evidence matrix. The absence of an operator in the
loop is never the absence of a record.
3. **Not an upstream development platform.** Nova does not own the
product backlog, IDE workflows, code authorship, or application
business logic. The PDLC is upstream; Nova integrates with it through
a validated contract boundary — Nova never penetrates it.
4. **Not a replacement for the Product Development Lifecycle (PDLC).**
Nova governs infrastructure + delivery only. Product lifecycle
decisions (what to build, when to ship, for whom) remain with the
product team. Nova makes their intent production-grade; it does not
own the intent.
---
+102 -351
View File
@@ -1,33 +1,31 @@
---
project: acdl
milestone: v1.17
generated_at: 2026-08-04
milestone: v1.18
generated_at: 2026-08-06
generator: lead-developer
verification_toolchain:
typecheck: "terraform validate && python3 -m py_compile core/**/*.py && python3 -m jsonschema schemas/*.schema.json"
test: "bash scripts/run_regression.sh # 22-capability gate (D-091/D-118) + CAP-023/024 (v1.17)"
build: "bash scripts/run_ci.sh # full local CI reproduction (lint+test+check-only)"
typecheck: "python3 -m py_compile core/submission_readiness.py mcp/atelier/server.py && python3 -m jsonschema schemas/submission-readiness.schema.json"
test: "pytest tests/test_submission_readiness.py tests/test_atelier_mcp.py # REQ-220 + REQ-225"
build: "bash scripts/render_deck.sh docs/presentations/nova-no-humans-platform-marp.md # HTML + PPTX (D-142)"
note: |
v1.17 adds a telemetry/observability layer (metrics emitters, SQLite
cold store, PowerBI export, Decision Ledger) + a unified narrative
deck + a durable NORTH_STAR.md. Three active personas: lead-developer
(coordination + deck narrative co-author), backend-engineer (event
emitters, outbox_writer extension, Infracost adapter), data-engineer
(SQLite store, schemas, PowerBI views, metrics collector). frontend-
engineer stays deactivated (no Nova web UI — dashboards are PowerBI,
not a Nova-built frontend; decks are markdown = lead-developer
territory). No new custom personas needed — the metrics domain maps
cleanly to data-engineer (schema/store/export) + backend-engineer
(emitters/instrumentation).
v1.18 adds the Citizen Developer & Production-Grade Guidance surface:
submission-readiness gate, Atelier-derived skills, the Atelier MCP server
(plugin-registry, stdio), and PPTX-as-first-class-artifact deck automation.
Three active personas: lead-developer (coordination + decks + RACI/scope
docs), backend-engineer (MCP server + submission-readiness validator +
render/attach scripts), data-engineer (submission-readiness schema if it
touches contract storage / DynamoDB shape). frontend-engineer stays
deactivated (v1.18 has no frontend; decks are markdown = lead-developer
territory). The MCP plugin-registry is a backend pattern, so a separate
mcp-engineer persona is NOT added — it folds into backend-engineer.
---
# ACDL — Persona Roster (project-level, v1.11 RESTART)
# ACDL — Persona Roster (v1.18 Citizen Developer & Production-Grade Guidance)
> v1.11 is a restart (D-097). The v1.9 roster is superseded. Three
> structural corrections: (1) stateless adapter (D-098), (2) terraform
> owns lifecycle (D-101), (3) pipeline-driven testing (D-102). The roster
> is simplified to the three active domains: data (terraform foundation),
> backend (adapter/resolver), general (pipelines/workflows).
> v1.18 roster. Three active personas + one deactivated. The MCP server
> plugin-registry (D-140) is a backend pattern, not a new persona — it
> folds into backend-engineer. v1.17 precedent (frontend-engineer
> deactivated, decks are markdown = lead-developer territory) is upheld.
## Active personas
@@ -35,351 +33,104 @@ verification_toolchain:
- **Domain:** coordination
- **Active:** true
- **Phase-specific:** false
- **Reason:** Owns CIAgent metadata, cross-phase verification scripts, the v1.11 phase orchestration (D-107: P56a + P56b split), and arbitrates persona conflicts. Resolves the milestone decomposition and the STANDARDS.md §8 rewrite (the adapter extension pattern is replaced by the per-module terraform subdir pattern).
- **Frameworks:** [] (no framework — owns process + narrative, not code)
- **Constraints:** ["pragmatic", "battle-tested defaults", "no fabrication (NORTH_STAR honesty model)"]
- **Territory:**
- `docs/presentations/**` (Step 1/2/4 markdown + the deck automation trigger)
- `.ciagent/**` (PROJECT, ROADMAP, REQUIREMENTS, RESEARCH, PLAN, GRILL, PERSONAS, REVIEW, CHECKPOINT)
- `PROJECT.md` (RACI matrix + PDLC-scope statement, REQ-215/216)
- `ROADMAP.md`
- `REQUIREMENTS.md`
- `docs/raci.md` (REQ-215)
- `docs/scope.md` (REQ-216)
- `docs/skills.md` (REQ-222 — the index page, not the skill files themselves)
- `docs/submission-readiness.md` (REQ-219 — citizen-developer-facing copy; co-owned with backend-engineer for the reason-code catalog)
- **Reason:** Owns CIAgent metadata, the milestone narrative, the RACI +
PDLC-scope statements (REQ-215/216), the deck (21 slides, S&P theme
regression check vs P1, CAP-024), the skills index page (REQ-222), and
the citizen-developer-facing submission-readiness doc (REQ-219). Is
the only persona that touches `.ciagent/**` and the deck markdown.
- **Phase-specific flag:** none (active for all of P0P7).
### backend-engineer
- **Domain:** backend
- **Active:** true
- **Phase-specific:** false
- **Reason:** Owns the adapter rewrite (D-098: stateless assembler — deletes TYPE_MAP/INPUT_MAP/OUTPUT_MAP + 39 type-specific branches, becomes a ~80-line assembler that emits `module "x" { source = "..." ... }` blocks) and the contract resolver env-aware state keys (D-106: `spike/{id}/{env}/terraform.tfstate`). The adapter holds no module content; the engine binding lives in the per-module `terraform/` subdir. Co-authoring expected on the adapter + `run_platform.sh` boundary (general adds `--apply`/`--destroy` modes that invoke the adapter).
- **Territory:** `adapters/terraform/adapter.py` (rewrite to stateless assembler), `core/contract_resolver.py` (env-aware state keys, deterministic composition), `schemas/stack.schema.json` (if the stack instance shape changes), `tests/test_adapter*.py` (regression baseline — the s3 instance.json round-trip must still pass).
- **Frameworks:** ["mcp (Python SDK v2)", "pydantic", "jsonschema", "urllib"]
- **Constraints:** ["api-first", "strict-typing", "plugin-registry extensible (D-140)", "stdio now / HTTP-ready (D-135)", "no stack traces to citizen developers (REQ-218)"]
- **Territory:**
- `mcp/atelier/server.py` (REQ-223)
- `mcp/atelier/plugins/**/*.py` (REQ-223 — principles.py, validation.py)
- `mcp/atelier/vendor/**` (REQ-224 — vendored Atelier snapshot)
- `mcp/atelier/VERSION.md` + `mcp/atelier/README.md` (REQ-224)
- `scripts/update_atelier_vendor.sh` (REQ-224)
- `core/submission_readiness.py` (REQ-218 — the validator, invoked as `contract_ingestor.py --check-readiness`)
- `scripts/render_deck.sh` (REQ-228 — HTML + PPTX render)
- `scripts/attach_release_asset.py` (REQ-228 — Gitea release asset upload)
- `tests/test_atelier_mcp.py` (REQ-225)
- `tests/test_submission_readiness.py` (REQ-220)
- `docs/submission-readiness.md` (REQ-219 — reason-code catalog section; co-owned with lead-developer for the narrative)
- **Reason:** Owns the MCP server (plugin-registry, stdio, vendored
Atelier), the submission-readiness validator (extends
`contract_ingestor.py --check-readiness`, D-133), the render/attach
scripts (D-142 trigger), and the two new test files. The MCP
plugin-registry (D-140) is a backend pattern — no separate
mcp-engineer persona is created; backend-engineer owns it.
- **Phase-specific flag:** none (active for P1 deck-render, P3 validator,
P5 MCP server, P6 scripts).
### data-engineer
- **Domain:** data
- **Active:** true
- **Phase-specific:** false
- **Reason:** Reactivated for v1.11. Owns the heaviest territory: the per-module `terraform/` subdirs (D-098/D-099/D-100 — the engine binding) for all 12 L1 modules, plus the single platform VPC (D-105: `terraform/platform` owns ONE VPC; the microservice composition drops its `vpc` child and references the platform VPC via data source). Each L1 module ships a real terraform module dir (versions/variables/locals/main/outputs.tf) owning its resource shape, nested blocks, and defaults. `locals.tf` is used heavily to centralize default interpolation (D-099). Multi-resource modules get the full 5-file split; trivial single-resource modules may inline locals in main.tf. This is the binding constraint — the stateless adapter cannot be written until the reference s3 module exists (D-107: P56a proves the design with s3 first).
- **Territory:** `terraform/` (platform VPC, D-105), `modules/l1/*/terraform/` (per-module terraform subdirs — the engine binding), `modules/l1/*/interface.json` (defaults move from adapter to interface inputs), `modules/registry.json` (terraform_dir field), `modules/l2/microservice/composition.json` (drop the vpc child, D-105), `modules/STANDARDS.md` §8 (rewrite the adapter extension pattern → per-module terraform subdir pattern).
### general (lead-developer + backend-engineer pipeline work)
- **Domain:** coordination + pipelines
- **Active:** true
- **Phase-specific:** false
- **Reason:** Owns the pipeline-driven testing (D-102/D-103/D-104) and the terraform lifecycle modes (D-101). The modules-lifecycle pipeline (Gitea + GitHub, byte-identical) matrix-runs each L1 module's `examples/{simple,complex}.yml` contracts through apply→modify→destroy against live AWS. `run_platform.sh` gains `--apply` and `--destroy` modes; Python never runs terraform. `verify_deploy_microservice.py` is deleted (D-101). Co-authoring expected on the `run_platform.sh` boundary (backend-engineer rewrites the adapter that `run_platform.sh` invokes).
- **Territory:** `pipelines/modules-lifecycle.yml`, `.gitea/workflows/modules-lifecycle.yml` + `.github/workflows/modules-lifecycle.yml` (byte-identical, D-102), `scripts/run_platform.sh` (`--apply`/`--destroy` modes, D-101), `scripts/run_primitive_plan.sh` (if extended for lifecycle), `scripts/run_pattern_plan.sh` (if extended), `pipelines/README.md` (document the new pipeline), `schemas/deploy-pipeline.schema.json` (if the lifecycle stages are added to the contract).
## Deactivated personas
### lambda-engineer (custom, v1.9 — deactivated for v1.11)
- **Domain:** serverless
- **Active:** false
- **Phase-specific:** false
- **Reason:** No per-module Python this milestone (D-102: testing is pipeline-driven, not pytest). The v1.9 Lambda (`core/lambda/contract_ingestor.py`) and the `terraform/platform/main.tf` Lambda/DynamoDB/KMS/Secrets definitions persist from v1.9 but are not touched in v1.11. The `acdl-sod-halt` SNS topic and the attestation matrix are out of scope. Removed from the roster for v1.11; reactivates if a future milestone touches the Lambda.
### platform-engineer (custom, v1.9 — folded into data-engineer for v1.11)
- **Domain:** infra
- **Active:** false
- **Phase-specific:** false
- **Reason:** The v1.11 scope (D-097..D-107) is terraform module authoring + adapter rewrite + pipelines — not the v1.9-era L1/L2 IR-typed module authoring or the AWS OIDC bootstrap. The platform-engineer's v1.9 territory (`adapters/terraform/**`, `modules/**`, `terraform/**`) is split: the adapter goes to backend-engineer (rewrite), the per-module terraform subdirs + platform VPC go to data-engineer (the heaviest v1.11 work). Folded into data-engineer for v1.11; reactivates if a future milestone does IR-shaped module authoring or OIDC bootstrap work.
### security-engineer (custom, v1.9 — deactivated for v1.11)
- **Domain:** security
- **Active:** false
- **Phase-specific:** false
- **Reason:** The v1.11 scope does not touch Wiz/Kyverno/Checkov adapters, the HITL matrix, separation-of-duties, or the audit ledger. The security-engineer's v1.9 territory persists but is not touched. Removed from the roster for v1.11; reactivates if a future milestone touches security adapters or HITL gates.
### frontend-engineer
- **Domain:** frontend
- **Active:** false
- **Phase-specific:** false
- **Reason:** The evidence timeline UI (`evidence-ui/**`) is unchanged from v1.0 and not touched in v1.11. Removed from the active roster; reactivates if a future milestone touches the timeline UI.
### data-engineer (v1.9 — was deactivated, reactivated for v1.11)
- **Domain:** data
- **Active:** true (reactivated)
- **Phase-specific:** false
- **Reason:** See the active `data-engineer` entry above. The v1.9 deactivation rationale ("No ORM/persistence framework") no longer applies — v1.11's data-engineer owns terraform module authoring, not a data persistence layer.
### infra-stub-engineer (custom, v1.0 only)
- **Domain:** backend
- **Active:** false
- **Reason:** Owned L1 stub modules in the v1.0 demo. The demo is archived to `demo/`; real L1 modules are owned by data-engineer (v1.11). Not reactivated.
## Phase-specific overrides
| Phase | Personas active | Notes |
|-------|------------------|-------|
| 56a adapter-rewrite-and-s3-reference-module | data-engineer (lead: s3 reference terraform module — proves the design), backend-engineer (lead: stateless adapter rewrite — emits module blocks for s3), general (run_platform.sh --apply/--destroy skeleton) | security/lambda/frontend idle |
| 56b remaining-11-l1-module-terraform-subdirs | data-engineer (lead: author 11 L1 module terraform subdirs — vpc, ecs-cluster, ecs-service, iam-role, alb, ecr, cloudfront, waf, rds, kms-key, uptime), backend-engineer (adapter: confirm each module round-trips through the assembler), general (modules-lifecycle pipeline wiring) | security/lambda/frontend idle |
| (modules-lifecycle pipeline) | general (lead: byte-identical Gitea+GitHub workflow + matrix apply→modify→destroy), data-engineer (examples/{simple,complex}.yml contracts as the modify variants), backend-engineer (adapter confirms the lifecycle cells resolve) | security/lambda/frontend idle |
| (platform VPC + composition drop) | data-engineer (lead: terraform/platform VPC + microservice composition drops vpc child, D-105), backend-engineer (resolver: env-aware state keys, D-106) | general/security/lambda/frontend idle |
| verify | lead-developer (lead: 4-layer verification), all active personas (review their territory) | — |
| review-audit-complete | lead-developer (lead: review + audit + milestone completion), all active personas (review participation) | — |
## Domain priority (used by TaskDecomposer)
`data → backend → general`
Rationale: in v1.11, the terraform foundation (per-module `terraform/`
subdirs + platform VPC) is the binding constraint — the stateless adapter
cannot be written until the reference s3 module exists (D-107: P56a
proves the design with s3 first). Backend (adapter/resolver) follows once
the module shape is proven. General (pipelines/workflows) wires the
lifecycle modes last, once the adapter + modules produce valid terraform.
## Conflict resolutions (lead-developer arbitration)
- `backend-engineer` vs `data-engineer` over `modules/l1/*/interface.json`:
data-engineer owns the interface defaults (defaults move from the
adapter to the interface inputs, D-100); backend-engineer owns the
adapter that reads them. Co-authoring is expected; conflict goes to
lead-developer.
- `backend-engineer` vs `general` over `scripts/run_platform.sh`:
backend-engineer rewrites the adapter that `run_platform.sh` invokes;
general adds the `--apply`/`--destroy` modes. The interface (the CLI
flags + the adapter invocation) is co-authored; conflicts go to
lead-developer.
- `data-engineer` vs `general` over `modules/l1/*/examples/`:
data-engineer owns the example contracts (the modify variants,
D-103); general owns the pipeline that matrix-runs them. Co-authoring
is expected; conflicts go to lead-developer.
- `lead-developer` vs any: lead-developer owns `.ciagent/**` + `docs/**`
meta + verification scripts + `modules/STANDARDS.md` §8 rewrite; persona
engineers do not edit CIAgent metadata or the vision/architecture
source docs.
## Territory enforcement mode
`warn` — config.json has no `personas.territory_enforcement` field, so the
default per execute.md is `warn`. Cross-territory edits are logged in the
commit message but do not fail the task. v1.11's scope means co-authoring
across territories is likely (e.g. backend + general on the adapter +
`run_platform.sh` boundary; data + general on the examples + pipeline
boundary); `warn` keeps it frictionless.
---
## v1.15 Persona Addendum — Nova Rebrand (2026-07-30)
**Milestone:** v1.15-Nova. The roster carries forward from v1.11/v1.14
unchanged — the rebrand touches existing territories, no new domains.
**frontend-engineer** remains deactivated (no UI; decks are markdown =
lead-developer territory). No **security-engineer** persona is activated
— the ABAC session-policy + tag-key migration (REQ-162) is data-engineer
territory (terraform IAM) with lead-developer review.
### v1.15 territory assignments
| Phase | Lead | Contributors | Territory |
|-------|------|---------------|-----------|
| P1 docs-decks-prose | lead-developer | — | `README.md`, `docs/**`, `.ciagent/*.md`, deck `.md`/`-marp.md`/`-talking-points.md`/`.html`, `docs/presentations/assets/mmd/*.mmd` (+ PNG re-export), `pyproject.toml`, `schemas/*.schema.json` `$id` (D-110), `docs/NOVA_MIGRATION.md`, `.github/workflows/release.yml` title, `modules/STANDARDS.md` |
| P2 code-envvars-consumer-path | backend-engineer | lead-developer (docs/runbook) | `core/env.py` (NEW dual-read helper, D-108), `core/*.py` (call-site migration), `scripts/*.py` + `*.sh`, `adapters/**`, `tests/**`, `.gitea/workflows/**` + `.github/workflows/**`, `.env` + `.env.secrets` (key rename), `schemas/tagging-standard.json`, `adapters/terraform/policy/custom_rules/acdl_tagging.py``nova_tagging.py` (D-109: warn mode) |
| P3 ssm-tagkeys | data-engineer | backend-engineer (readers) | `core/output_publisher.py` (SSM path `/nova/`), `core/contract_resolver.py` (SSM reads), `scripts/migrate_ssm_paths.py` (NEW), `terraform/**` (tag keys `nova:*`), `adapters/terraform/policy/custom_rules/nova_tagging.py` (D-109: hard mode), ABAC session-policy terraform |
| P4 aws-resource-migration | data-engineer | lead-developer (runbook) | `terraform/platform/main.tf`, `terraform/microservice/main.tf`, `terraform/ci-vpc/main.tf`, `terraform/bootstrap/**`, `modules/l1/alb/instance.json`, `scripts/migrate_dynamodb_data.py` (NEW), `docs/NOVA_AWS_MIGRATION.md` (NEW runbook), `core/lambda/contract_ingestor.py` (default table names → `nova-*`, D-111) |
| P5 final-review-ship | lead-developer | all active (review) | `.ciagent/**` (REQUIREMENTS/ROADMAP/PROJECT complete), `core/env.py` (remove dual-read fallback), `nova_tagging.py` (hard-fail `acdl:*`), review + audit |
### v1.15 domain priority
`lead → backend → data` (inverted from v1.11)
Rationale: the rebrand is docs/prose-first (P1 establishes the
vocabulary, no runtime impact), then code/env-vars/consumer-path (P2),
then SSM/tag-keys (P3), then the heavy terraform/AWS migration (P4).
Lead-developer owns the docs + runbooks + verification + final ship;
backend-engineer owns the dual-read helper + call-site migration +
contract resolver; data-engineer owns the terraform resource/tag/SSM
migration (the heaviest terraform territory). Co-authoring expected at:
`core/env.py` + `core/*.py` boundary (backend + lead on the helper
design), `nova_tagging.py` + `schemas/tagging-standard.json` boundary
(backend authors the rule, data-engineer owns the tag-key schema),
`core/output_publisher.py` SSM path + `terraform` outputs boundary
(backend writes the reader, data-engineer owns the terraform that
produces the outputs).
### v1.15 verification toolchain (unchanged from v1.14)
```
typecheck: terraform validate && python3 -m py_compile core/**/*.py adapters/**/*.py
test: bash scripts/run_regression.sh # 16-capability gate
build: bash scripts/run_ci.sh # full local CI reproduction
```
The regression gate (CAP-001..CAP-016) must stay **16/16 Verified**
throughout the rebrand — the rebrand must not regress any capability.
P2/P3/P4 update test fixtures that reference `ACDL`/`acdl` so the gate
stays green.
## v1.16 Persona Addendum — Nova Simplification (2026-07-30)
**Milestone:** v1.16-Nova-Simplification (NFR). Roster carries forward
unchanged — NFR work touches existing territories, no new domains. The
onboarding request-path (P18P20) is backend-engineer (Lambda action +
onboarding.py) + data-engineer (cross-account Terraform) territory.
**frontend-engineer** remains deactivated. No **security-engineer**
persona — the ingestor defense-in-depth (P10) is backend-engineer with
lead-developer review; IAM/ABAC (P20) is data-engineer territory.
### v1.16 territory assignments
| Phase | Lead | Contributors | Territory |
|-------|------|---------------|-----------|
| P1 state-bucket+kyverno fix | backend-engineer | data-engineer (kyverno policy) | `adapters/terraform/adapter.py:117`, `adapters/kyverno/policies/require-resource-labels.yml` |
| P2 user-facing brand sweep | lead-developer | backend-engineer | `core/environment_check.py`, `core/lambda/contract_ingestor.py`, `scripts/post_stage_comment.sh`, `scripts/run_ci.sh`, module docstrings, `adapters/README.md` |
| P3 dead-code+stale-prefix | lead-developer | — | `scripts/run_platform.sh`, `core/local_emulators.py`, `core/regression_verify.py`, lifecycle scripts |
| P4 migrate-ssm except | backend-engineer | — | `scripts/migrate_ssm_paths.py` |
| P5 regression-verify dedup | backend-engineer | — | `core/regression_verify.py` |
| P6 run-platform deadcode+hitl-fn | lead-developer | — | `scripts/run_platform.sh` |
| P7 contract-resolver envloader+kind | backend-engineer | — | `core/contract_resolver.py`, `modules/registry.json` |
| P8 workflow generator | lead-developer | backend-engineer (test) | `scripts/sync_workflows.py` (NEW), `tests/test_pipeline_contract.py`, `.gitea/workflows/**`, `.github/workflows/**` |
| P9 run-platform split | lead-developer | — | `scripts/run_platform.sh`, `scripts/run_decommission.sh` (NEW), `scripts/run_uptime.sh` (NEW) |
| P10 ingestor defense-in-depth | backend-engineer | lead-developer (review) | `core/lambda/contract_ingestor.py`, `core/environments/` |
| P11 ingestor payload validation | backend-engineer | — | `core/lambda/contract_ingestor.py` |
| P12 split contract-resolver | backend-engineer | — | `core/contract_resolver.py``core/contract_resolve.py` + `core/decommission_transform.py` + `core/contract_resolver_cli.py` |
| P13 split regression-verify | backend-engineer | — | `core/regression_verify.py` → split modules |
| P14 schema-driven outputs+cache | backend-engineer | data-engineer (interface.json) | `core/output_publisher.py`, `core/contract_resolver.py`, `modules/l1/*/interface.json` |
| P15 run-platform --help+flags | lead-developer | — | `scripts/run_platform.sh`, `README.md` |
| P16 workflows README catalog | lead-developer | — | `.github/workflows/README.md` (NEW) |
| P17 getting-started consolidation | lead-developer | — | `README.md` |
| P18 onboarding schema+lambda | backend-engineer | lead-developer (schema) | `schemas/onboarding.schema.json` (NEW), `core/lambda/contract_ingestor.py` |
| P19 onboarding envfile autogen | backend-engineer | lead-developer (docs) | `core/onboarding.py` (NEW), `core/environment_check.py`, `core/environments/README.md` |
| P20 cross-account role offline | data-engineer | backend-engineer (ABAC) | `terraform/onboarding/` (NEW), `terraform/platform/main.tf` |
| P21 final-review-ship | lead-developer | all active (review) | `.ciagent/**`, review + audit + ship |
### v1.16 domain priority
`backend → lead → data` (the simplification + security + ingestor work
is backend-heavy; lead-developer owns docs/DX/splits; data-engineer owns
the P20 cross-account Terraform only).
### v1.16 verification toolchain
```
typecheck: terraform validate && python3 -m py_compile core/**/*.py adapters/**/*.py
test: bash scripts/run_regression.sh # 22-capability gate (D-118: P9 + P21)
build: bash scripts/run_ci.sh # full local CI reproduction
```
The regression gate (22 capabilities) must stay **22/22 Verified**
throughout v1.16 — simplification must not regress any capability
(D-118). P9 (end of Wave 2) and P21 (milestone complete) run the gate;
P14 (end of Wave 3) is an offline mid-milestone checkpoint.
---
# v1.17 Persona Roster — Strategic Direction, Leadership Metrics & Unified Story
> v1.17 adds a telemetry/observability layer (P1P3), a metrics catalog
> + NORTH_STAR integration (P4), a unified narrative deck (P5), a
> regression capability (P6), and a final review/ship (P7). Three
> active personas; frontend-engineer stays deactivated (no Nova web UI
> — dashboards are PowerBI, not a Nova-built frontend).
## Active personas
### lead-developer
- **Domain:** coordination + deck narrative
- **Active:** true
- **Phase-specific:** false
- **Reason:** Owns CIAgent metadata, the NORTH_STAR.md authoring
process (P0), the milestone decomposition, the unified narrative deck
co-authoring (P5 — the deck is markdown, which is lead-developer
territory per the established convention), and the final review/ship
(P7). Arbitrates persona conflicts (e.g., backend vs data on the
emitter/store boundary).
- **Territory:** `.ciagent/NORTH_STAR.md`, `.ciagent/PROJECT.md`,
`.ciagent/REQUIREMENTS.md`, `.ciagent/PLAN.md`, `.ciagent/RESEARCH.md`,
`.ciagent/ARCHITECTURE.md`, `docs/presentations/nova-no-humans-platform.md`
(NEW — unified deck source of truth), `docs/presentations/nova-no-humans-platform-marp.md`,
`docs/presentations/nova-no-humans-platform-talking-points.md`,
`docs/METRICS.md`, `docs/metrics/*.md` (per-KPI definition docs).
### backend-engineer
- **Domain:** backend (event emitters + instrumentation)
- **Active:** true
- **Phase-specific:** false
- **Reason:** Owns the event emitters (P1): the CloudEvents envelope,
the per-run manifest writer, the `outbox_writer.py` extension to the
SQLite Decision Ledger, the Infracost post-processor, the
`hitl_gates.py` attestation event emission, the `confidence_signal.py`
decision event emission, the `checkov_adapter.py` policy event
emission, and the pytest `--junitxml` addopts change. Also owns the
`regression_verify.py` CAP-023/024 additions (P6). The emitter work
is the bridge between existing Nova components and the new metrics
layer — it touches the code paths that already exist.
- **Territory:** `core/metrics/event_envelope.py` (NEW),
`core/metrics/run_manifest.py` (NEW),
`core/metrics/infracost_adapter.py` (NEW),
`core/metrics/decision_ledger.py` (NEW — extends outbox_writer),
`core/outbox_writer.py` (extend to SQLite),
`core/hitl_gates.py` (emit attestation.recorded),
`core/confidence_signal.py` (emit ai.decision.made),
`adapters/terraform/policy/checkov_adapter.py` (emit policy.evaluated),
`scripts/run_platform.sh` (invoke manifest writer + Infracost),
`core/regression_verify.py` (CAP-023/024),
`pyproject.toml` (addopts --junitxml),
`tests/test_metrics_emitters.py` (NEW),
`tests/test_decision_ledger.py` (NEW).
### data-engineer
- **Domain:** data (schema, SQLite store, PowerBI export)
- **Active:** true
- **Phase-specific:** false
- **Reason:** Reactivated with a new territory for v1.17: the metrics
collector (P2) and the PowerBI export (P3). Owns the schema design
(metrics_*.schema.json), the SQLite cold store (nova_metrics.db), the
fact/dimension table design, the 8 deferred placeholder views, and
the CSV/JSON export. The data-engineer's schema-first constraint
applies: all event types and fact/dim tables have JSON Schema
definitions before any code is written. The collector reads files +
events → SQLite; the export reads SQLite → CSV/JSON. This is the
heaviest data-territory work since v1.11's terraform modules.
- **Territory:** `core/metrics/collector.py` (NEW),
`core/metrics/powerbi_export.py` (NEW),
`schemas/metrics_*.schema.json` (NEW — event + fact/dim schemas),
`metrics/nova_metrics.db` (NEW — SQLite cold store),
`metrics/powerbi/` (NEW — CSV/JSON export dir),
`docs/METRICS_VIEWS.md` (NEW — schema doc for PowerBI views),
`tests/test_metrics_collector.py` (NEW),
`tests/test_powerbi_export.py` (NEW).
- **Frameworks:** ["jsonschema", "dynamodb (item shape)"]
- **Constraints:** ["schema-first", "superset-gate NOT duplicate (PROJECT.md hard constraint)", "W3.E per-env mandatory table is the source of truth"]
- **Territory:**
- `schemas/**` (REQ-217 — `submission-readiness.schema.json` is the new schema; existing schemas untouched)
- `core/lambda/contract_ingestor.py` (the `--check-readiness` subcommand wiring, D-133 — the validator is in `core/submission_readiness.py` but the ingestor dispatches to it; co-owned with backend-engineer)
- **Reason:** Owns the submission-readiness JSON Schema (REQ-217) — it
is a schema artifact, data-engineer territory. The schema is a
*superset gate above* `contract.schema.json`, not a duplicate (it
references contract fields, does not redefine them). The
per-env-mandatory table comes from W3.E (the locked decision). The
ingestor wiring is co-owned with backend-engineer (the dispatch point
is backend; the schema it validates against is data).
- **Phase-specific flag:** none (active for P3 schema + ingestor wiring).
## Deactivated personas
### frontend-engineer
- **Domain:** frontend
- **Active:** false
- **Phase-specific:** false
- **Reason:** v1.17 has no Nova web UI. The leadership dashboards are
PowerBI (an external tool that ingests CSV/JSON files), not a
Nova-built frontend. The decks are markdown (lead-developer
territory). frontend-engineer stays deactivated, consistent with
v1.11v1.16. Reactivates if a future milestone builds a Nova web UI.
- **Domain:** frontend
- **Frameworks:** ["react", "next.js"] (inert — no territory)
- **Constraints:** ["component-first", "server-components", "minimal-client-js"] (inert)
- **Territory:** [] (no territory in v1.18)
- **Reason:** v1.18 has no frontend; decks are markdown (lead-developer
territory); deactivated per PERSONAS.md v1.17 precedent. v1.18's
observability stays PowerBI / external (Out of Scope: "A Nova-built
frontend / dashboard"). The MCP server exposes tools to an AI agent,
not a web UI. No reactivation trigger in this milestone.
### lambda-engineer, platform-engineer, security-engineer
- **Active:** false (carried forward from v1.11)
- **Reason:** v1.17 does not touch the Lambda (beyond emitting events
from the existing hitl_gates/attestation_matrix), does not do IR-
shaped module authoring, and does not touch security adapters beyond
emitting policy.evaluated events. The existing components are
instrumented, not rewritten.
## Roster decisions
## v1.17 phase assignment
### D-143 (0.90): Fold mcp-engineer into backend-engineer
The MCP plugin-registry (D-140: `plugins/<name>.py register(mcp)`) is a
backend code pattern — Python modules, type hints, stdio transport,
urllib for the Gitea asset API. It shares nothing with the data domain
(schemas/DynamoDB) and is not a new engineering discipline. Creating a
separate `mcp-engineer` persona would fragment ownership of the server +
its tests + the render/attach scripts (all backend). **Decision:** fold
into backend-engineer. backend-engineer's `frameworks` list gains
`mcp (Python SDK v2)`. Confidence 0.90 — the only counter-argument is
that MCP is a distinct protocol skill, but the SDK v2 API surface
(`@mcp.tool()` + type hints) is small and well within backend-engineer's
range (it's the same Pydantic/FastAPI-style pattern the persona already
knows).
| Phase | Primary persona | Supporting | Territory |
|-------|----------------|------------|-----------|
| P0 pre-execution | lead-developer | — | `.ciagent/NORTH_STAR.md`, `PROJECT.md`, `REQUIREMENTS.md`, `RESEARCH.md`, `ARCHITECTURE.md`, `PERSONAS.md`, `PLAN.md` |
| P1 event-emitters | backend-engineer | data-engineer (schemas) | `core/metrics/event_envelope.py`, `run_manifest.py`, `decision_ledger.py`, `infracost_adapter.py`, `outbox_writer.py`, `hitl_gates.py`, `confidence_signal.py`, `checkov_adapter.py`, `run_platform.sh`, `pyproject.toml` |
| P2 metrics-collector | data-engineer | backend-engineer (event formats) | `core/metrics/collector.py`, `schemas/metrics_*.schema.json`, `metrics/nova_metrics.db` |
| P3 powerbi-export | data-engineer | — | `core/metrics/powerbi_export.py`, `metrics/powerbi/`, `docs/METRICS_VIEWS.md` |
| P4 metrics-catalog + north-star-integration | lead-developer | data-engineer (metric definitions) | `docs/METRICS.md`, `docs/metrics/*.md`, `PROJECT.md`, `ARCHITECTURE.md`, `config.json` |
| P5 deck-rebuild | lead-developer | — | `docs/presentations/nova-no-humans-platform*.md`, retire old decks |
| P6 regression-capability | backend-engineer | data-engineer (CAP-023 schema) | `core/regression_verify.py` (CAP-023, CAP-024) |
| P7 final-review-ship | lead-developer | all active (review) | `.ciagent/**`, review + audit + ship |
### Territory-overlap resolution (co-ownership)
## v1.17 domain priority
`backend → data → lead` (the emitter work in P1 is the foundation;
data-engineer's collector + export in P2P3 depends on P1's event
formats; lead-developer's catalog + deck in P4P5 depends on the
metrics being grounded).
## v1.17 verification toolchain
```
typecheck: terraform validate && python3 -m py_compile core/**/*.py adapters/**/*.py
test: bash scripts/run_regression.sh # 22-capability gate + CAP-023/024 (v1.17)
build: bash scripts/run_ci.sh # full local CI reproduction
```
The regression gate (22 capabilities + CAP-023 metrics collector +
CAP-024 deck structure) must pass at P6 and P7. CAP-009 (offline pytest
suite) must remain Verified after the `--junitxml` addopts change
(assumption A5).
| Path | Primary | Co-owner | Why |
|------|---------|----------|-----|
| `docs/submission-readiness.md` | lead-developer (narrative + examples) | backend-engineer (reason-code catalog, REQ-218 codes) | The doc is citizen-developer-facing copy (lead) but the reason-code catalog (MISSING_TAGS, ENV_MISSING_MANDATORY, AGENTIC_MISSING_INTENT, MISSING_APP_SOURCE, POLICY_PRECONDITION_MISSING) is backend (it mirrors the validator's return codes). |
| `core/lambda/contract_ingestor.py` | backend-engineer (dispatch wiring) | data-engineer (the schema it validates against) | D-133 places the `--check-readiness` subcommand on the ingestor (backend dispatch), but the readiness schema it loads is data-engineer territory. |
| `schemas/submission-readiness.schema.json` | data-engineer (schema artifact) | backend-engineer (the validator must match it) | The schema is data-engineer's; the validator (REQ-218) is backend-engineer's and must stay in sync with it. |
+1018 -1098
View File
File diff suppressed because it is too large Load Diff
+292 -7
View File
@@ -58,6 +58,103 @@ traceable to a human attestation and an immutable evidence stream.
boundary. The platform validates, enriches with operational standards,
and reconciles the target state.
## Scope: Nova is Downstream of PDLC
> **Promoted from Core Tenet #2 + Anti-Goal #1 (v1.18, REQ-216).** This
> is the unmissable scope statement — the PDLC is upstream, Nova is
> downstream.
The **Product Development Lifecycle (PDLC)** — product backlog, code
authorship, IDE workflows, sprint planning, application business logic —
is **upstream** of Nova. Nova never penetrates the PDLC. Nova's domain is
**infrastructure + delivery only**: environment progression, cloud
resource lifecycle, operational security/observability NFRs, policy
enforcement, immutable audit lineage, and the two consumer surfaces
(technical developer + agentic).
Integration between the PDLC and Nova is **only** through the validated,
published contract boundary (`schemas/contract.schema.json` +
`schemas/submission-readiness.schema.json`). The citizen developer's AI
coding agent, an upstream agentic SDLC platform, or any upstream
development platform may all produce submissions — the source does not
matter because all are subject to the same compliance standards (the
submission-readiness gate, D-133). Nova validates, enriches with
operational standards, and reconciles the target state. Nova never
authors application code, manages product backlogs, or provides IDE
workflows.
```
PDLC (upstream) Nova (downstream)
───────────────── ─────────────────
product backlog contract ingestion
code authorship (AI agent / IDE / SDLC) → submission-readiness gate
sprint planning → policy enforcement
application business logic → cloud resource lifecycle
→ environment progression (dev→qa→prod→dr)
→ immutable audit + attestation
```
## RACI Matrix
> **Source of truth (v1.18, REQ-215, D-139).** Three roles clarify who
> owns what across the Nova delivery lifecycle. The matrix is the
> authoritative version; `docs/raci.md` is the citizen-developer-facing
> copy.
### Roles
- **Citizen Developer (CD)** — the consumer (technical developer L3A or
non-technical L3B). Responsible for all **Functional Requirements (FRs)**
and **User Acceptance Testing (UAT)**. The FRs + UAT are produced via
the citizen developer's AI coding agent, an upstream agentic SDLC, or
an upstream development platform — **the source does not matter as all
are subject to the same compliance standards** (the submission-readiness
gate, D-133).
- **Platform** — Nova. Responsible for all **Non-Functional Requirements
(NFRs)**, **Infrastructure** (cloud resource lifecycle, state, IAM),
**QA** (the platform-side quality checks: policy, confidence, schema),
and **Production deployments to cloud** (the apply path, the pipeline,
the release).
- **Release Management (RM)** — **co-owned**. QA + SRE attestations are
required by the actual release. The attestations are performed
agentically (the platform runs the checks), but the release is
**overseen and triggered by the Citizen Developer** — the human
attestation at the stage gate (D-042, hitl_gates.py). The platform
performs; the citizen developer authorizes.
### Matrix
| Work Category | Citizen Developer | Platform | Release Management |
|---|---|---|---|
| **Functional Requirements (FRs)** | **R/A** | C | I |
| **User Acceptance Testing (UAT)** | **R/A** | C | I |
| **Non-Functional Requirements (NFRs)** | I | **R/A** | C |
| **Infrastructure (cloud, state, IAM)** | I | **R/A** | C |
| **QA (policy, confidence, schema checks)** | C | **R/A** | I |
| **Production deployment to cloud** | I | **R/A** | C |
| **Release attestation (QA + SRE sign-off)** | **A** | R | **R** |
**Key: R** = Responsible (does the work) · **A** = Accountable (owns the
outcome, sign-off) · **C** = Consulted · **I** = Informed.
**Compliance-standard equivalence note:** the citizen developer's FRs +
UAT may originate from any upstream source — an AI coding agent, an
agentic SDLC platform, or a traditional development platform. All are
subject to the same compliance standards: the submission-readiness gate
(`schemas/submission-readiness.schema.json`), the contract schema, the
policy envelope, and the immutable audit stream. The platform does not
differentiate by upstream source; it validates the submission, not the
author.
**Co-ownership of Release Management:** the release is co-owned. The
platform performs the QA + SRE attestations agentically (confidence signal,
policy checks, separation-of-duties). The citizen developer oversees and
triggers the actual release — the human attestation at the stage gate is
the citizen developer's authorization, recorded with approver identity
(D-042). The platform runs the checks; the citizen developer authorizes
the promotion. This is the "autonomy in operations, human at stage gates"
model from the NORTH_STAR.
## Capability Status (Re-Verified 2026-07-27)
> Source of truth: `.ciagent/CAPABILITY_INVENTORY.md` (Phase 54, D-093).
@@ -592,12 +689,97 @@ DX: 16 total). Key changes:
10. Old two-surfaces diagram replaced by scope boundary diagram.
Source markdown, talking points, and README all updated to mirror the new
structure. Also includes scripts/sync_to_gl.sh (GitLab mirror sync
utility, unrelated to presentations).
structure. Also includes scripts/sync_to_nova.sh (manual-only "2nd release"
into ~/nova — a separate GitLab consumer-facing repo with its own history;
domain-based conventional commits, never triggered by CI; REQ-229).
No code changes; 494 tests pass; `run_ci.sh` + `run_platform.sh --check-only`
green. PPTX files uploaded to Gitea release.
## Objective for Milestone v1.18 (active — Citizen Developer & Production-Grade Guidance)
v1.18 advances Nova from a platform that governs infrastructure delivery
to one that **instructs the citizen developer on production-grade
engineering** and defines a **clear, machine-checkable contract for what
is acceptable to start**. Five user-directed inputs drive the milestone:
1. **S&P Global theme restoration.** The v1.17 P5 deck rebuild consolidated
two decks into one unified narrative deck but lost the S&P Global Energy
brand visual identity (introduced v1.9.2 / P45, commit `ae0cb58`). The
Marp `style:` block (red-core `#D6002A`, grey-90 `#1B1B1B`, Akkurat Pro
font, 8px top accent bar) is restored to the unified deck. The mermaid
`sp-theme.json` survived; only the Marp CSS theme was lost.
2. **PDLC-upstream scope made explicit.** Core Tenet #2 already states the
platform "does not penetrate upstream product/SDLC" and Anti-Goal #1 says
"Not an upstream development platform." v1.18 promotes this from a
buried tenet to a dedicated, unmissable scope statement in PROJECT.md +
`docs/scope.md` + a deck slide: **the PDLC (Product Development
Lifecycle — product backlog, code authorship, IDE) is upstream of Nova;
Nova governs infra + delivery only; integration is through the validated
contract boundary.**
3. **RACI matrix.** A three-role responsibility matrix clarifies who owns
what: **Citizen Developer** (Responsible for all Functional Requirements
+ User Acceptance Testing, via their AI coding agent / upstream agentic
SDLC / upstream development platform — the source does not matter as all
are subject to the same compliance standards), **Platform** (Responsible
for all NFRs + Infrastructure + QA + Production deployments to cloud),
**Release Management** (co-owned: QA + SRE attestations required by the
actual release, performed agentically but overseen & triggered by the
Citizen Developer). Source of truth in PROJECT.md + `docs/raci.md` + a
deck slide.
4. **Nova input contract — "what is acceptable to start."** A JSON Schema
(`schemas/submission-readiness.schema.json`) defines the
acceptable-to-start gate as a superset *above* contract-schema validity:
schema-valid contract + required Nova tags + per-env mandatory metadata
(per W3.E) + declared policy preconditions + (for L3B) `profile:agentic`
markers + `appSource` pointer. A validator (`core/submission_readiness.py`,
invoked as `contract_ingestor.py --check-readiness`) returns a structured
`ReadinessResult` with reason codes. On fail → citizen-developer-facing
error (not a stack trace); on pass → proceeds to existing ingestion.
5. **Atelier integration — production-grade guidance + agentic validation.**
Nova consumes `coreci/atelier` (a first-principles docs-as-code
engineering framework — 8 core principles, 19 domains, 190 P-rules) via
two surfaces: **skills** (markdown files under `skills/` keyed to Atelier
domain paths, surfaced to the citizen developer's AI agent, extending the
BA.A 5-skill catalog) and an **MCP server** (`mcp/atelier/server.py`,
plugin-registry architecture, stdio transport, vendored Atelier snapshot
for audit reproducibility) exposing tools for principle-lookup,
domain-listing, matrix-lookup, and agentic validation against the
Atelier agent-checklist — validation that goes beyond deterministic
scanners (Wiz/Checkmarx/Mend) by catching correctness/clarity/simplicity/
observability gaps.
**Deck automation (cross-cutting):** any phase modifying
`docs/presentations/*-marp.md` or `docs/presentations/assets/` MUST
re-render HTML + PPTX, **commit the PPTX to git** (binary, no LFS), and
attach it to the phase's Gitea release. New scripts:
`scripts/render_deck.sh` (HTML + PPTX render) and
`scripts/attach_release_asset.py` (Gitea release asset upload).
**Milestone type:** Feature (P1 S&P theme restoration + P3 readiness
schema/validator + P5 MCP server are new code/features). Tags run on the
**v1.17.x** patch line (previous minor per branch-strategy): `v1.17.0` (P0)
`v1.17.1..v1.17.6` (P1P6) → `v1.17.7` (P7 final = milestone release).
**Phase count:** 8 (P0 pre-execution + 6 execution + 1 final).
**Hard constraints:**
- DO NOT make anything up (NORTH_STAR.md honesty model).
- The submission-readiness schema is a superset gate above
`contract.schema.json`, NOT a duplicate — it references but does not
redefine contract fields.
- The MCP server is plugin-registry extensible (future capabilities drop
in as new plugin files, no `server.py` edits).
- Atelier is vendored (pinned tag) for audit reproducibility — an agentic
validation result must be replayable against the exact principles that
produced it.
- PPTX is a first-class artifact: committed (history) + attached (download)
— both always, not optional.
## Requirements
### v1.0 (Prior milestone — the demo)
@@ -799,7 +981,7 @@ or user-directed scope). New v1.7 decisions:
| W1.A | AI-refinement trigger | **Accept recommendation.** Joint condition: N ≥ 50 consecutive changes with zero rollbacks AND no L1/L2 incident in last 6 months AND Infra & Ops unilateral override. |
| W1.B | Multi-stack edge case rule | **Accept recommendation.** Permitted only for (a) DR-region mirror, (b) time-boxed experimental stack with TTL ≤ 30d, (c) explicit Infra & Ops approval with `multiStack.justification`. |
| W2.A | Tag mutability for prod | **Accept recommendation (Path B).** Tag for dev/qa, SHA for prod. Platform CLI resolves tag→SHA for prod-bound workflows. Justified by the "Audit truth lives outside the repository" bet. |
| BA.A | Initial L3B skill catalog | **Accept recommendation.** 5 skills: web API, worker, scheduled job, static asset, basic observability bootstrap. Addition criteria: (a) reviewable for sensitive data, (b) expressible as a single contract submission, (c) documented use case. |
| BA.A | Initial L3B skill catalog | **Accept recommendation.** 5 skills: web API, worker, scheduled job, static asset, basic observability bootstrap. Addition criteria: (a) reviewable for sensitive data, (b) expressible as a single contract submission, (c) documented use case. **Extended v1.18 (REQ-221/222):** the BA.A 5-skill catalog is extended with 9 Atelier-derived production-grade engineering skills under `skills/` (api, security, data, testing, observability, errors, devops, infrastructure-as-code, compliance), indexed by `docs/skills.md`. The Atelier skills extend, not replace, the BA.A catalog. |
| W3.D | L1/L2 standard versioning | **Decided.** Semver: interface → MAJOR, behavior → MINOR, lifecycle → PATCH (same as the v1.0 demo D-rule, lifted to the real platform). Pin model: L2 contracts pin L1 by `name@semver`; the resolver picks the highest compatible. Evolution: MAJOR bumps require a new registry entry (immutable publication); old entry enters a 12-month deprecation window. |
| W3.E | Schema mandatory vs optional inputs | **Decided.** Per-env mandatory table: dev requires `stack` + `environment`; qa adds `validation.e2eSuite` + `validation.loadTest`; prod adds `runbook` + `dashboard` + `oncall`; dr adds `drDrillRef`. `inputs` map is always optional. `profile: agentic` fields (`naturalLanguageIntent`, `confidenceAtSubmission`, `agentTrace`) optional everywhere. |
| BA.B | Confidence threshold tuning | **Decided.** Starting thresholds frozen for v1. Tuning begins in v1.2: track FP/FN per environment quarterly; override authority = Infra & Ops + SRE joint sign-off; any override is itself a confidence-event in the audit stream. |
@@ -1084,16 +1266,29 @@ P7 review+audit+ship). Tags on the v1.16.x line: `v1.16.0` (P0) →
- **Pillar A — Strategic Direction.** A durable, PO-authored
`.ciagent/NORTH_STAR.md` encodes the platform's vision, 4 strategic
objectives, 5 anti-goals, v1.17 non-goals, 1218mo targets (with a
objectives, anti-goals, v1.17 non-goals, 1218mo targets (with a
grounding column), and success criteria. CIAgent reads it in every
future `/ci-run` so the direction survives across milestones. The
attestation clarification is reflected: human attestation required at
stage gates (QA for production, SRE for operational readiness); autonomy
in operations, not in accountability.
in operations, not in accountability. **v1.21 refinement:** Strategic
Objective #4 reframed from "default substrate for agentic consumption" to
integrating with externally owned PDLC/SDLC/Agentic/Citizen Developer
platforms regardless of source (Nova provides skills + MCP endpoints;
all prod intents go through the same controls). Objective #2 reworded:
trust is established by deterministic scripts that calculate a score —
the platform functions without AI. Objective #3 reworded with four
CTO-grade metrics (Lead Time PR→Prod, Infrastructure Vulnerability
Count trend, MTTR, Cloud Spend Reduction) all flowing into PowerBI.
Anti-goals #1, #4, #5 removed; replaced with "not an upstream
development platform" and "not a replacement for the PDLC".
- **Pillar B — Leadership Metrics + PowerBI.** Instrument Nova to
collect, aggregate, and surface leadership-grade metrics that prove the
"no-humans" autonomous-infrastructure value proposition. Nova-native
"no-humans" autonomous-infrastructure value proposition (reframed in
v1.21 to "autonomous cloud delivery" — professional framing; the
platform delivers safe production deployment without an operator in
the loop of normal operations). Nova-native
minimal tech (CloudEvents 1.0 envelope, JSONL event log, SQLite cold
store, hash-chained Decision Ledger via `outbox_writer.py` extension)
+ Infracost for pre-apply cost estimates. Hybrid model: existing
@@ -1133,4 +1328,94 @@ P7 review+audit+ship). Tags on the v1.16.x line: `v1.16.0` (P0) →
| D-129 | PowerBI delivery = CSV/JSON files, folder connector. | Nova is offline-first; no live connector to a running service. PowerBI ingests via the folder connector. | P3 emits CSV/JSON to metrics/powerbi/. |
| D-130 | Deck arc = Problem → Vision → How → Proof → Roadmap. | The unified narrative deck's 5-act structure. x3 arc at deck + slide level. Per-slide benefit callouts. Fluid transitions. Both old decks retired. | P5 builds the unified deck; old decks deleted. |
| D-131 | MTTR scope = platform-run MTTR. | The <60s MTTR target refers to platform-run failures (apply.failed → successful retry), not infra-incident MTTR (no incident detection system). Infra-incident MTTR deferred. | P4 grounds platform-run MTTR. |
| D-132 | Attestation instrumentation = emit attestation.recorded events. | The attestation system (hitl_gates.py + attestation_matrix.py + separation_of_duties.py) already exists. Instrument it: emit attestation.recorded events into the Decision Ledger + PowerBI. Attestation Coverage = 100% target grounded from outbox approver_* attributes. | P1 emits attestation events; P4 grounds Attestation Coverage. |
| D-132 | Attestation instrumentation = emit attestation.recorded events. | The attestation system (hitl_gates.py + attestation_matrix.py + separation_of_duties.py) already exists. Instrument it: emit attestation.recorded events into the Decision Ledger + PowerBI. Attestation Coverage = 100% target grounded from outbox approver_* attributes. | P1 emits attestation events; P4 grounds Attestation Coverage. |
## Key Decisions (v1.18)
Resolved at the CLARIFY stage (full autonomy — all within locked
constraints or user-directed scope). New v1.18 decisions:
| ID | Decision | Rationale | Outcome |
|----|----------|-----------|---------|
| D-133 | Submission-readiness validator location = extend `contract_ingestor.py --check-readiness`. | Adding a new CLI binary is unnecessary; the ingestor is the existing entry point for contract submission. The validator is a subcommand that runs before ingestion proceeds. No new binary, no new entry point to maintain. | P3 implements the subcommand; no new CLI binary. |
| D-134 | Deck slide budget = 18 → 21 slides (no act restructure). | The 3 new slides (scope/RACI/atelier) are leadership-relevant and append after the existing 18. The 5-act arc (D-130) is preserved; the new slides are append-only context, not a new act. | P6 appends 3 slides → 21 total. |
| D-135 | Atelier MCP transport = stdio now; HTTP-ready (same server object). | stdio is the local-agent transport (the citizen developer's AI agent spawns the server as a subprocess). The MCP Python SDK v2 supports Streamable HTTP on the same `MCPServer` object, so adding HTTP later is a transport-only change in `server.py`, not a rewrite. | P5 ships stdio; HTTP deferred (documented in README). |
| D-136 | Atelier source = vendor pinned tag under `mcp/atelier/vendor/`. | An agentic validation result is only reproducible if the principles that produced it are pinned. Live-fetch breaks replayability (Atelier `main` drifts). Vendoring matches the v1.16 P15 offline-first precedent and the Nova thesis (provable trust). `mcp/atelier/vendor/VERSION.md` records the pinned tag; `scripts/update_atelier_vendor.sh` is the intentional upgrade path. | P5 vendors Atelier; live-fetch not implemented. |
| D-137 | MCP server language = Python (MCP Python SDK v2, `modelcontextprotocol/python-sdk`). | Nova's `core/` is Python. The MCP Python SDK v2 (23.9k stars, MIT, stable) matches the codebase; type hints become JSON Schema automatically (`@mcp.tool()` decorator). | P5 uses Python SDK v2. |
| D-138 | Skill catalog format = markdown files under `skills/` keyed to Atelier domain paths. | Markdown is the established Nova docs format (Jekyll Pages, 4-step deck process). Each skill file names the Atelier source path, distills the first-principles, links to agent-checklist triggers, and maps to the BA.A catalog. | P4 authors 9 markdown skill files. |
| D-139 | RACI role names = Citizen Developer / Platform / Release Management (co-owned). | User-specified. The 3 roles are the columns of the RACI table. Release Management is co-owned: QA + SRE attestations are required by the actual release (performed agentically, overseen & triggered by the Citizen Developer). | P2 authors the RACI with these 3 roles. |
| D-140 | MCP server extensibility = plugin-registry (`plugins/<name>.py` implementing `register(mcp)`). | Future capabilities (new scanners, policy evaluators, cost tools) drop in as new plugin files — no `server.py` edits. `server.py` scans `plugins/` and calls `register` on each. This is the extensibility insurance: plugins are decoupled from the server entrypoint. | P5 implements the plugin-registry; initial plugins are `principles.py` + `validation.py`. |
| D-141 | PPTX storage = commit binary directly to `docs/presentations/` (no LFS). | Decks are small (~1-5 MiB); git handles binary blobs. LFS requires server-side support (unverified for git.cloudinit.dev) + client config. Committing directly is simplest and works without any repo/server config. Binary diffs are not delta-friendly, but deck changes are infrequent. | P1/P2/P6 commit .pptx directly. |
| D-142 | Deck render trigger = any phase modifying `docs/presentations/*-marp.md` or `docs/presentations/assets/` must re-render HTML + PPTX, commit PPTX, and attach to the Gitea release. | PPTX was previously manual + release-only (not committed). v1.18 makes it a first-class artifact: committed (history) + attached (download), both always, not optional. Automated via `scripts/render_deck.sh` + `scripts/attach_release_asset.py`. | P1/P2/P6 run the render+commit+attach pipeline. |
## Objective for Milestone v1.19 (complete — Nova 2nd-Release Sync)
> **NFR-only chore milestone.** Ships a patch on the v1.18.x line (tag
> `v1.18.0`). Single execution phase. Establishes the manual-only "2nd
> release" pipeline from `~/acdl` (CIAgent-managed source of truth, full audit
> trail) into `~/nova` (GitLab `jonathanchery/nova` — separate repo, separate
> history, consumer / platform-team audience).
### Why
`~/acdl` is the engineering source of truth and carries the full CIAgent
audit trail (`.ciagent/`, milestone branches, `---ci---` blocks, Gitea
releases). Consumers and the platform team should consume a clean,
conventional-commit-shaped tree without the CIAgent plumbing. The old
`scripts/sync_to_gl.sh` mirrored `~/acdl → ~/gl/acdl` with a single
kitchen-sink `chore: sync from source mirror <ts>` commit — wrong audience,
wrong commit standard, wrong repo.
### What
- **`scripts/sync_to_nova.sh`** replaces `scripts/sync_to_gl.sh`.
- **Manual-only gate**: refuses without `--release` / `RELEASE_CONFIRMED=1`
(exit 2). Never triggerable by CI.
- **Consumer subset only**: excludes `.ciagent/`, `.gitea/`, `.env*`,
`terraform/`, `demo/`, runtime metrics artifacts, and internal-only scripts
(the `EXCLUDE_SCRIPTS` list — CIAgent/ops/release plumbing). Keeps
consumer-facing runbooks (`run_ci.sh`, `run_platform.sh`, etc.) and the
metrics export views (`metrics/README.md`, `powerbi/`, `TRUST_SNAPSHOT.md`).
- **Destination history protected**: rsync `--filter=P .git` ensures
`~/nova/.git` is never touched.
- **Domain-based commits**: 13 fixed-order domains (config → core → adapters
→ modules → contracts → schemas → pipelines → mcp → skills → scripts →
tests → docs → workflows). Each changed domain gets its own conventional
commit, supplied positionally via repeated `-m` flags. No kitchen-sink.
- **Conventional-commit validation**: regex-enforced
(`feat|fix|docs|chore|refactor|perf|test|build|ci|style|revert`); bypass via
`--no-verify-format`.
- **Modes**: `--list-domains` (print order), `--dry-run` (preview rsync +
messages), `--no-push` (commit without pushing), `-v` (verbose).
### Out of Scope
- **coreci / Atelier review gate on the synced tree** — deferred. A future
milestone may run a vendored-Atelier review pass before commit and block on
P0 findings.
- **Tagging releases on the `~/nova` side** — could add `--tag <semver>`
later.
- **Deleting `~/gl`** — the old mirror dir is left on disk; only the sync
script targeting it is removed.
### Requirements
- **REQ-229** — `scripts/sync_to_nova.sh` replaces `sync_to_gl.sh` with the
manual-only, consumer-subset, domain-committed 2nd-release pipeline
described above. (Phase P1)
### Phase Plan
| Phase | Name | Status |
|-------|------|--------|
| P1 | nova-sync-script | complete |
| P2 | final-review-ship | pending |
### Decisions
| ID | Decision | Rationale | Outcome |
|----|----------|-----------|---------|
| D-143 | 2nd release target = `~/nova` (separate GitLab repo), not `~/gl/acdl`. | `~/nova` is consumer/platform-team-facing with its own history; `~/gl/acdl` was an internal mirror with a kitchen-sink commit standard. Separate audience → separate repo → separate commit standard. | `sync_to_nova.sh` targets `~/nova`; `sync_to_gl.sh` removed. |
| D-144 | Commit standard for `~/nova` = real conventional commits per domain (not the `---ci---` audit blocks used in `~/acdl`). | `~/acdl` commits carry CIAgent audit metadata (`---ci---` blocks) for the ciagent auditing workflow; that's noise for platform consumers. `~/nova` gets clean `feat/fix/docs/chore(scope): subject` commits grouped by domain. | Script validates conventional format; domain-based commits via positional `-m`. |
| D-145 | Trigger = manual-only (`--release` / `RELEASE_CONFIRMED=1`). | The 2nd release is a deliberate human action, not a CI side-effect. The gate guarantees it can never fire from Gitea Actions, GitHub Actions, or accidental invocation. | Script exits 2 without `--release`. |
| D-146 | Domain grouping = 13 fixed-order domains by path prefix; messages map positionally over CHANGED domains only. | Avoids the kitchen-sink commit; gives `~/nova` a reviewable, conventional history tailored to platform consumers. Positional-over-changed mapping lets the human supply exactly the messages needed, in domain order, without padding for unchanged domains. | `--list-domains` prints order; `--dry-run` previews; count-mismatch errors clearly. |
| D-147 | coreci / Atelier review gate = deferred this milestone. | The vendored Atelier (`mcp/atelier/vendor`) could review the synced tree before commit and block on P0, but that's an additive hardening step, not part of establishing the pipeline. Deferred to a future milestone. | Sync ships consumer contents as-is; no review gate. |
+565 -29
View File
@@ -1132,35 +1132,35 @@ with documented schemas.
| Requirement | Phase | Status |
|-------------|-------|--------|
| REQ-185 | P0 | in_progress |
| REQ-186 | P4 | pending |
| REQ-187 | P1 | pending |
| REQ-188 | P1 | pending |
| REQ-189 | P2 | pending |
| REQ-190 | P3 | pending |
| REQ-191 | P4 | pending |
| REQ-192 | P4 | pending |
| REQ-193 | P4 | pending |
| REQ-194 | P4 | pending |
| REQ-195 | P4 | pending |
| REQ-196 | P5 | pending |
| REQ-197 | P5 | pending |
| REQ-198 | P6 | pending |
| REQ-199 | P3 | pending |
| REQ-200 | P2 | pending |
| REQ-201 | P2 | pending |
| REQ-202 | P5 | pending |
| REQ-203 | P5 | pending |
| REQ-204 | P4 | pending |
| REQ-205 | P1+P2+P3 | pending |
| REQ-206 | P1+P2 | pending |
| REQ-207 | P2 | pending |
| REQ-208 | P3 | pending |
| REQ-209 | P3/P4 | pending |
| REQ-210 | P4 | pending |
| REQ-211 | P4 | pending |
| REQ-212 | P4 | pending |
| REQ-213 | P4/P5 | pending |
| REQ-185 | P0 | complete |
| REQ-186 | P4 | complete |
| REQ-187 | P1 | complete |
| REQ-188 | P1 | complete |
| REQ-189 | P2 | complete |
| REQ-190 | P3 | complete |
| REQ-191 | P4 | complete |
| REQ-192 | P4 | complete |
| REQ-193 | P4 | complete |
| REQ-194 | P4 | complete |
| REQ-195 | P4 | complete |
| REQ-196 | P5 | complete |
| REQ-197 | P5 | complete |
| REQ-198 | P6 | complete |
| REQ-199 | P3 | complete |
| REQ-200 | P2 | complete |
| REQ-201 | P2 | complete |
| REQ-202 | P5 | complete |
| REQ-203 | P5 | complete |
| REQ-204 | P4 | complete |
| REQ-205 | P1+P2+P3 | complete |
| REQ-206 | P1+P2 | complete |
| REQ-207 | P2 | complete |
| REQ-208 | P3 | complete |
| REQ-209 | P3/P4 | complete |
| REQ-210 | P4 | complete |
| REQ-211 | P4 | complete |
| REQ-212 | P4 | complete |
| REQ-213 | P4/P5 | complete |
### Out of Scope (v1.17)
@@ -1180,3 +1180,539 @@ with documented schemas.
- A third deck — the two existing decks merge into one; no new
standalone metrics deck.
- A Nova web UI — dashboards are PowerBI, not a Nova-built frontend.
---
## v1.18 — Citizen Developer & Production-Grade Guidance
> **Milestone type:** Feature. Tags run on the v1.17.x patch line (previous
> minor per branch-strategy). `v1.17.0` (P0) → `v1.17.1..v1.17.6` (P1P6) →
> `v1.17.7` (P7 final = milestone release).
> **Active milestone:** v1.18. **Branch:**
> `milestone/v1.18-citizen-developer-guidance`.
### Requirements
- **REQ-214** — S&P Global Energy Marp theme restored in the unified deck
(`docs/presentations/nova-no-humans-platform-marp.md`). The `style:` block
from commit `ae0cb58` (v1.9.2 / P45) is ported: H1/H2 `#D6002A`
(S&P red-core), title-slide bg `#1B1B1B` (grey-90) with 8px `#D6002A` top
accent bar, body text `#1B1B1B`, blockquote border `#D6002A`,
table headers `#F0F0F0`, font `'Akkurat Pro'` with web-safe fallbacks. The
current Nova header/footer text is preserved (rebrand is not touched —
only the visual theme is restored). HTML re-rendered with the S&P theme.
- **REQ-215** — RACI matrix authored in `PROJECT.md` (new `## RACI Matrix`
section) and `docs/raci.md` (citizen-developer-facing copy). Three roles:
**Citizen Developer** (Responsible for all Functional Requirements + User
Acceptance Testing — via their AI coding agent / upstream agentic SDLC /
upstream development platform; the source does not matter as all are
subject to the same compliance standards), **Platform** (Responsible for
all NFRs + Infrastructure + QA + Production deployments to cloud),
**Release Management** (co-owned: QA + SRE attestations required by the
actual release, performed agentically but overseen & triggered by the
Citizen Developer). Rendered as a table: rows = work categories (FRs, UAT,
NFRs, Infra, QA, Prod deploy, Release attestation), columns = R/A/C/I per
role. Includes the compliance-standard-equivalence note.
- **REQ-216** — PDLC-upstream scope statement made explicit in `PROJECT.md`
(new `## Scope: Nova is Downstream of PDLC` subsection under Domain
Boundaries) and `docs/scope.md`. States that the PDLC (Product Development
Lifecycle — product backlog, code authorship, IDE) is upstream of Nova;
Nova governs infra + delivery only; integration is through the validated
contract boundary. Promotes Core Tenet #2 + Anti-Goal #1 from buried
tenets to a dedicated, unmissable scope statement.
- **REQ-217** — `schemas/submission-readiness.schema.json` (JSON Schema
draft 2020-12) defines what is acceptable to start — a superset gate
*above* `contract.schema.json` validity. Required fields: `contractId`
(non-empty), `environment` (dev/qa/prod/dr) with the W3.E per-env mandatory
table enforced (dev: stack+environment; qa: +validation.e2eSuite
+validation.loadTest; prod: +runbook+dashboard+oncall; dr: +drDrillRef),
`tags` (the 5 required Nova tags per D-054: `nova:owner`, `nova:contract`,
`nova:environment`, `nova:cost-center`, `nova:ref`), `policyPreconditions`
(declared policy expectations the platform will enforce, e.g.,
`public-ingress: false`), `profile` (`developer` or `agentic`; if
`agentic`, requires `naturalLanguageIntent`, `confidenceAtSubmission`,
`agentTrace` per REQ-22 / W3.E), `appSource` (repo + ref pointer for
runtime fetch).
- **REQ-218** — `core/submission_readiness.py` validator, invoked as
`contract_ingestor.py --check-readiness` subcommand (decision D-133). Returns
a structured `ReadinessResult` (pass/fail per check, with reason codes).
On fail → the ingestor rejects with a citizen-developer-facing error
(not a stack trace). On pass → proceeds to existing contract ingestion.
Calls `contract.schema.json` validation first, then the readiness checks.
Reason codes: `MISSING_TAGS`, `ENV_MISSING_MANDATORY:<env>:<field>`,
`AGENTIC_MISSING_INTENT`, `MISSING_APP_SOURCE`, `POLICY_PRECONDITION_MISSING`.
- **REQ-219** — `docs/submission-readiness.md` citizen-developer-facing doc
explaining what is acceptable to start, with good + rejected examples and
the reason-code catalog. References `schemas/submission-readiness.schema.json`
as the source of truth.
- **REQ-220** — `tests/test_submission_readiness.py` covers: good contract
passes; missing tags fail with `MISSING_TAGS`; missing env mandatory fails
with `ENV_MISSING_MANDATORY:<env>:<field>`; agentic profile missing intent
fails with `AGENTIC_MISSING_INTENT`; missing appSource fails with
`MISSING_APP_SOURCE`.
- **REQ-221** — `skills/` directory with 9 Atelier-derived skill files mapped
to the BA.A citizen-developer catalog: `skills/api.md` (domains/api/),
`skills/security.md` (domains/security/), `skills/data.md` (domains/data/),
`skills/testing.md` (domains/testing/), `skills/observability.md`
(domains/observability/), `skills/errors.md` (domains/errors/),
`skills/devops.md` (domains/devops/), `skills/infrastructure-as-code.md`
(domains/infrastructure-as-code/), `skills/compliance.md`
(domains/compliance/). Each names the Atelier source path, distills the
first-principles to the citizen-developer-relevant subset, links to
agent-checklist triggers, and maps to the BA.A 5-skill catalog (web API,
worker, scheduled job, static asset, basic observability bootstrap).
- **REQ-222** — `docs/skills.md` index page listing the skill catalog, the
Atelier provenance, and how the citizen developer's AI agent consumes them
(read before completing a task; run `review/agent-checklist.md` before
finishing). `PROJECT.md` BA.A decision extended with the Atelier-derived
skill catalog reference.
- **REQ-223** — `mcp/atelier/server.py` MCP server (stdio transport,
decision D-135) with a **plugin-registry architecture** (decision D-140):
`plugins/<name>.py` modules each expose `register(mcp: MCPServer) -> None`
and call `@mcp.tool()` for their tools; `server.py` scans `plugins/` and
calls `register` on each. Initial plugins: `principles.py`
(`atelier.lookup_principle`, `atelier.list_domains`, `atelier.matrix_lookup`)
and `validation.py` (`atelier.validate_against_principles` — agentic
validation against the Atelier agent-checklist, beyond Wiz/Checkmarx/Mend).
Uses the MCP Python SDK v2 (`modelcontextprotocol/python-sdk`).
- **REQ-224** — `mcp/atelier/vendor/` vendored Atelier snapshot (pinned tag,
decision D-136) for audit reproducibility. `mcp/atelier/vendor/VERSION.md`
records the pinned tag + a `scripts/update_atelier_vendor.sh` helper for
intentional upgrades. `mcp/atelier/README.md` documents the server: how to
run, transport, tool catalog, plugin-authoring guide, vendoring policy.
- **REQ-225** — `tests/test_atelier_mcp.py` covers: tool registration (all 4
tools discoverable via `tools/list`), `atelier.lookup_principle` returns
the principle text + core C-rule, `atelier.validate_against_principles`
catches a planted C1 (correctness) + C7 (observability) violation in a
known-bad snippet and passes a known-good snippet, `atelier.matrix_lookup`
returns the domain→core mapping, plugin discovery loads all plugins in
`plugins/`.
- **REQ-226** — 3 new deck slides added to the unified deck
(`docs/presentations/nova-no-humans-platform-marp.md`) → 21 slides total:
Slide 19 "Scope: Downstream of PDLC", Slide 20 "RACI: Who Owns What",
Slide 21 "Production-Grade Guidance via Atelier". Arc Preview slide
updated to reflect 21-slide count. Talking points
(`nova-no-humans-platform-talking-points.md`) synced for the 3 new slides.
S&P theme preserved (regression check vs P1). CAP-024 deck structure
regression passes.
- **REQ-227** — `docs/presentations/README.md` slide count + deck table
updated to reflect 21 slides + the 3 new slide titles.
- **REQ-228** — `scripts/render_deck.sh` (renders HTML + PPTX from a Marp
deck, commits both to git) and `scripts/attach_release_asset.py` (uploads
a file to a Gitea release via the API). Any phase modifying
`docs/presentations/*-marp.md` or `docs/presentations/assets/` MUST
re-render HTML + PPTX, commit the PPTX binary to `docs/presentations/`,
and attach it to the phase's Gitea release. PPTX is stored as a committed
binary (no LFS, decision D-141).
### Out of Scope (v1.18)
- **Streamable HTTP transport for the MCP server** — stdio ships now; HTTP
is a future milestone (the SDK supports it on the same server object, so
adding it later is a transport-only change, not a rewrite).
- **A Nova-built frontend / dashboard** — observability stays PowerBI /
external; no Nova web UI.
- **Replacing the existing BA.A 5-skill catalog** — the Atelier-derived
skills extend it, not replace it.
- **Live AWS re-provisioning** (D-096, still deferred) — submission-readiness
validates the contract shape, not a live AWS deployment.
- **A second forge adapter** (GitLab) — BA.F cross-platform evolution is
future work.
- **Atelier live-fetch mode** — vendoring is the only mode this milestone;
live-fetch (with its reproducibility trade-offs) is not implemented.
### v1.18 Traceability
| REQ | Phase | Status |
|-----|-------|--------|
| REQ-214 | P1 | complete |
| REQ-215 | P2 | complete |
| REQ-216 | P2 | complete |
| REQ-217 | P3 | complete |
| REQ-218 | P3 | complete |
| REQ-219 | P3 | complete |
| REQ-220 | P3 | complete |
| REQ-221 | P4 | complete |
| REQ-222 | P4 | complete |
| REQ-223 | P5 | complete |
| REQ-224 | P5 | complete |
| REQ-225 | P5 | complete |
| REQ-226 | P6 | complete |
| REQ-227 | P6 | complete |
| REQ-228 | P1/P2/P6 | complete |
## v1.19 — Nova 2nd-Release Sync (GitLab consumer mirror)
> **NFR-only chore milestone.** A single execution phase shipping a patch on
> the v1.18.x line (tag `v1.18.0`). Establishes the manual-only "2nd release"
> pipeline from `~/acdl` (CIAgent-managed source of truth) into `~/nova`
> (GitLab `jonathanchery/nova` — a separate repo, separate history, consumer /
> platform-team audience). `~/acdl` retains the full CIAgent audit trail;
> `~/nova` receives only the consumer subset, committed with real
> conventional commits per domain (no kitchen-sink "sync from source mirror").
- **REQ-229** — `scripts/sync_to_nova.sh` replaces `scripts/sync_to_gl.sh`.
The script: (1) refuses to run without `--release` / `RELEASE_CONFIRMED=1`
(manual-only — never triggerable by CI); (2) rsyncs the consumer subset of
`~/acdl` into `~/nova`, excluding `.ciagent/`, `.gitea/`, `.env*`, `terraform/`,
`demo/`, runtime metrics artifacts, and internal-only scripts (full list in
`EXCLUDE_SCRIPTS`), while protecting `~/nova/.git` history via rsync
`--filter=P .git`; (3) commits changes domain-by-domain in a fixed order
(config → core → adapters → modules → contracts → schemas → pipelines →
mcp → skills → scripts → tests → docs → workflows) using one
conventional-commit message per changed domain passed via repeated `-m`
flags (positional mapping over changed domains only — no kitchen-sink
commit); (4) validates conventional-commit format (`feat|fix|docs|chore|…`)
unless `--no-verify-format`; (5) pushes to the branch upstream unless
`--no-push`. `--list-domains`, `--dry-run`, `-v` supported. The old
`sync_to_gl.sh` is removed. (Phase P1)
### Out of Scope (v1.19)
- **coreci / Atelier review gate on the synced tree** — deferred; the sync
ships consumer contents as-is. A future milestone may run a vendored-Atelier
review pass before commit and block on P0 findings.
- **Tagging releases on the `~/nova` side** — could add `--tag <semver>` later.
- **Deleting `~/gl`** — the old GitLab `acdl` mirror is left on disk; only the
sync script targeting it is removed.
### v1.19 Traceability
| REQ | Phase | Status |
|-----|-------|--------|
| REQ-229 | P1 | complete |
## v1.20 — Consumer Cleanup + Transparent Terraform + Slide Pipeline
> **Multi-concern milestone.** Four user-directed inputs spanning consumer
> cleanup, infrastructure transparency, and presentation automation. Tags
> run on the v1.19.x line (milestone v1.20 → tags v1.19.0, v1.19.1, …).
>
> **Input 1 — Gitea/GitLab removal:** Remove all mentions of `gitea` / `gitlab`
> (case-insensitive) from every file synced to `~/nova`. The platform team
> (consumer of `~/nova`) must never know about the dev forge or the GitLab
> mirror. Genericize forge-detection code to `forge` / `generic_forge`.
>
> **Input 2 — Documentation simplification:** Radically simplify all synced
> documentation. Anything the CIAgent needs to reference for itself lives in
> `.ciagent/`. Everything else is tailored to the Platform Team audience.
> Strip ciagent-internal provenance (REQ-/D-/P-/CAP- IDs, milestone headers,
> `.ciagent/PROJECT.md` citations) from synced docs. Delete completed
> migration guides. Move internal artifacts to `.ciagent/`.
>
> **Input 3 — Transparent terraform:** Move terraform `init` / `validate` /
> `plan` / `apply` / `output` into native workflow steps (transparent, visible
> in CI logs). Split `run_platform.sh` into `run_codegen.sh` (pre-TF) +
> `run_postapply.sh` (post-TF). Add `var.enabled` feature flags to every L1
> module + L2 composition toggles. Wire forge repo variables as per-client
> feature flags — different clients test different functionality without
> version upgrades.
>
> **Input 4 — Slide pipeline + product roadmap:** The slides have not adopted
> the S&P Global theme fully. Create a dedicated render pipeline that builds
> the slides (mermaid PNGs + Marp HTML/PPTX) with the S&P theme applied to all
> slide chrome. Add a 12-month product roadmap (high-level, product-oriented
> vs the technical roadmap in `.ciagent/ROADMAP.md`) to the deck.
- **REQ-230** — No `gitea` / `gitlab` string literal (case-insensitive) appears
in any file synced to `~/nova`. Verified by
`tests/test_no_forge_mentions.py` which scans the synced subset (same
path rules as `sync_to_nova.sh`'s `DOMAINS` / `EXCLUDES`). Forge-detection
code (`contract_ingestor.py`, `hitl_gates.py`, `run_platform.sh`) is
genericized: `gitea``forge` / `generic_forge`, `GITEA_ACTOR`
`FORGE_ACTOR` (with `GITHUB_ACTOR` primary). (Phase P1)
- **REQ-231** — Synced documentation is tailored to the Platform Team
audience. Ciagent-internal provenance (`v1.XX — Strategic Direction` headers,
`REQ-NNN` / `D-NNN` / `P-NNN` / `CAP-NNN` IDs, `.ciagent/PROJECT.md`
"source of truth" citations) is stripped from synced docs. (Phase P1)
- **REQ-232** — Completed/historical migration docs
(`docs/NOVA_MIGRATION.md`, `docs/NOVA_AWS_MIGRATION.md`) removed from the
synced tree. `docs/NO_HUMANS_THESIS.md` moved to `.ciagent/` (internal
thesis-defense artifact). (Phase P1)
- **REQ-233** — Terraform `init` / `validate` / `plan` / `apply` / `output`
run as native workflow steps in `deploy.yml` (transparent, named steps
visible in CI logs), not buried inside `run_platform.sh`. (Phase P4)
- **REQ-234** — `run_platform.sh` is split: `run_codegen.sh` (pre-TF: env
check, validate, resolve, adapt) + `run_postapply.sh` (post-TF: Checkov,
confidence, HITL, outbox, SSM, comment, uptime). A thin `run_platform.sh`
shim preserves backward compat for local-dev usage. (Phase P4)
- **REQ-235** — Every L1 module has `variable "enabled" { type = bool,
default = true }` + `count = var.enabled ? 1 : 0` on its primary
resource(s); declared in `interface.json`. The `uptime` module's
`feature_flag_enabled` is renamed to `enabled` (with backward-compat alias).
(Phase P4)
- **REQ-236** — L2 `composition.json` supports per-child `enabled` toggles
driven by contract `inputs.enable_<child>`. The resolver skips children
with `enabled: false`. (Phase P4)
- **REQ-237** — `deploy.yml` reads feature flags from forge repository
variables (`vars.ENABLE_*`) and passes them as `-var` flags to terraform,
enabling per-client feature toggles without version upgrades. (Phase P4)
- **REQ-238** — Stale artifact path `/tmp/acdl_platform_run_v18` in
`deploy.yml` fixed to use `NOVA_WORK_DIR`. (Phase P4)
- **REQ-239** — A dedicated S&P Global theme CSS file
(`docs/presentations/assets/nova-sp-theme.css`) is the Marp theme for all
Nova presentation decks. The theme applies the S&P Red/Black/White palette
(`#D6002A`, `#1B1B1B`, `#FFFFFF`) to all slide chrome (background,
header/footer, pagination, tables, blockquotes), not just headings. (Phase P2)
- **REQ-240** — A dedicated render pipeline (`scripts/render_slides.sh`)
builds the presentation deck end-to-end: (1) renders all
`assets/mmd/*.mmd` → `assets/png/*.png` via `mermaid-cli --configFile
sp-theme.json`; (2) renders the Marp deck → HTML + PPTX via `marp-cli`;
(3) stages all rendered artifacts to git. Supersedes `render_deck.sh`.
(Phase P2)
- **REQ-241** — A CI workflow (`workflows-src/slides.yml` +
`.github/workflows/slides.yml`) runs `render_slides.sh` on any change to
`docs/presentations/**` and commits the rendered HTML/PPTX/PNGs back. No
manual re-render step; no artifact drift. (Phase P2)
- **REQ-242** — `tests/test_slides_pipeline.py` validates: (1) the Marp
deck frontmatter references `nova-sp-theme.css`; (2) the CSS contains the
S&P colors; (3) every `.mmd` has a corresponding `.png`; (4) the HTML
exists and is newer than the Marp `.md`. (Phase P2)
- **REQ-243** — `docs/presentations/README.md` directory layout is updated
to remove retired decks (`how-the-platform-works-*`,
`the-developer-experience-*`) and document the render pipeline + theme CSS.
(Phase P2)
- **REQ-244** — A 12-month product roadmap (4 quarters, product-outcome
oriented, grounded in NORTH_STAR strategic objectives + deferred-metric
candidate milestones) is added to the presentation deck as Slide 20 +
Slide 21. The roadmap is distinct from Slide 15's deferred-metric unblock
paths. A matching talking-points section is added. (Phase P3)
### Out of Scope (v1.20)
- **Multi-cloud (Azure/GCP) implementation** — deferred; only the product
roadmap references it as a Q4 aspiration.
- **ML anomaly-forecasting service** — deferred; only the product roadmap
references it as a Q4 aspiration.
- **Actual pilot estate activation** — deferred (requires live AWS
re-provisioning, D-096 lift); the product roadmap references it as Q1.
- **Token rotation for `NOVA_GITEA_TOKEN`** — out of scope; the `.env` files
are correctly excluded from sync. Flagged for awareness only.
### v1.20 Traceability
| REQ | Phase | Status |
|-----|-------|--------|
| REQ-230 | P1 | complete |
| REQ-231 | P1 | complete |
| REQ-232 | P1 | complete |
| REQ-233 | P4 | complete |
| REQ-234 | P4 | complete |
| REQ-235 | P4 | complete |
| REQ-236 | P4 | complete |
| REQ-237 | P4 | complete |
| REQ-238 | P4 | complete |
| REQ-239 | P2 | complete |
| REQ-240 | P2 | complete |
| REQ-241 | P2 | complete |
| REQ-242 | P2 | complete |
| REQ-243 | P2 | complete |
| REQ-244 | P3 | complete |
## v1.21 — Nova Deck Refinement & Pipeline Hardening
> Leadership-deck refinement based on 33 review notes on the v1.20 deck
> (v1.20 shipped as `nova-no-humans-platform*`). This milestone renames the
> deck to the professional "Autonomous Cloud Delivery Platform" framing,
> restructures the narrative (Problem → Solution → Proof → Roadmap + Ask),
> removes internal provenance from audience-facing slides, hardens the
> policy pipeline (Checkov before plan, Wiz-or-Checkov on plan), and moves
> the strategic integration objective into the North Star.
>
> Tags run on the v1.20.x line (milestone v1.21 → tags v1.20.0, v1.20.1, …).
### REQ-245 — Deck rename + restructure
The deck files are renamed from `nova-no-humans-platform*` to
`nova-autonomous-cloud-delivery*` across all five artifacts
(source `.md`, `-marp.md`, `.html`, `.pptx`, `-talking-points.md`).
The in-deck title becomes "Nova — The Autonomous Cloud Delivery Platform"
(professional, conveys autonomy without the provocative "no-humans"
wording). The narrative restructures to 18 main + 1 appendix slides:
1. The Problem (merged old 1+2; broader problem framing; no "arc"; no
"18 capabilities verified"; not "humans are the problem"; add tribal
knowledge / rockstar-operator framing)
2. Nova's Vision
3. Strategic Objectives + Anti-Goals
4. Scope: Downstream of PDLC (moved up)
5. RACI: Who Owns What (moved up)
6. The Platform Pipeline
7. The Decision Ledger
8. The Attestation Matrix
9. Telemetry & Live Ops
10. Decision Ledger + Attestation Coverage
11. Cost & ROI
12. What's Deferred — and Why
13. Roadmap to the North Star
14. 12-Month Product Roadmap
15. Quarter-by-Quarter Outcomes
16. Production-Grade Guidance via Atelier (1/2)
17. Production-Grade Guidance via Atelier (2/2)
18. Recap + Ask
A1. Metrics Glossary
Removed: old Slide 10 (Capability Health), old Slide 12 (Zero-Touch
Efficiency), old Appendix A2 (Operating Model & Cost). Slide 5's first
table removed.
### REQ-246 — Thesis rename + reframe
`.ciagent/NO_HUMANS_THESIS.md` is renamed (git mv) to
`.ciagent/AUTONOMY_THESIS.md`. Content reframes from "removing humans" to
"autonomy in operations, human at stage gates" — professional, not
provocative. The operator-bottleneck framing is softened; the attestation
model + provable trust are emphasized. Anti-claims are retained and
reworded for a tech-leadership audience. All references across the repo
are updated to the new filename + framing.
### REQ-247 — Strategic-docs sync (NORTH_STAR + PROJECT)
`NORTH_STAR.md` is updated:
- Vision polished for a technical audience concerned about security,
security remediation velocity, and reliability; "infrastructure
operations become visible" is preserved as a recurring theme.
- Strategic Objective #2 (provable trust) is reworded: trust is
established by deterministic scripts that calculate a score, not by
AI. The platform functions without AI. "AI decisions" are really
automated decisions.
- Strategic Objective #3 (ROI) is reworded with four CTO-grade metrics:
Lead Time (PR → Production), Infrastructure Vulnerability Count
(downward trend), MTTR, Cloud Spend Reduction. All flow into PowerBI
views and are captured by the telemetry pipeline.
- Strategic Objective #4 is replaced: integrate with externally owned
PDLC, SDLC, Agentic, and Citizen Developer platforms regardless of
source; Nova provides skills + MCP endpoints to make applications
production-grade; all intents to deploy to production go through the
same rigorous controls and quality gates.
- Anti-goals #1 (hyperscaler competitor), #4 (legacy untagged), and #5
(sold to operators) are removed. Two new anti-goals added: not an
upstream development platform; not a replacement for the Product
Lifecycle (PDLC).
- Anti-goal #3 reworded to remove the "removes humans" framing.
`PROJECT.md` mission statement + scope are synchronized with the
integration objective and the reworded strategic objectives.
### REQ-248 — RACI restructure (Quality Engineering + SRE)
The RACI matrix (slide + `docs/raci.md`) is restructured:
- A **Quality Engineering** column is added.
- The Platform column no longer holds the **A** for release attestation;
accountability is reassigned to QA or SRE as appropriate.
- "Release Management" is renamed to **SRE**.
- "Release attestation" is split into two rows: the SRE part is
**Production Readiness** (operational readiness sign-off).
- The slide is sized to fit (text shrunk / low-impact rows dropped).
### REQ-249 — Atelier split (2 slides)
Slide 19 (Production-Grade Guidance via Atelier) is split into two slides:
- **16 (1/2):** Skills + MCP server overview (the 9 skills, the 4 MCP
tools, the plugin-registry + stdio surface).
- **17 (2/2):** Agentic validation beyond deterministic scanners +
vendored Atelier for audit reproducibility.
The benefit wording is improved; the same spirit is retained.
### REQ-250 — Pipeline hardening (Checkov before plan; Wiz-or-Checkov on plan)
`scripts/run_platform.sh` (and `scripts/run_postapply.sh` where
relevant) implement the two-stage policy scan:
1. **Checkov runs on static code** (the generated `main.tf` / TF
directory) **before** `terraform plan` — fail-fast, quick developer
feedback on policy violations in the authored code.
2. **After `terraform plan`:** if `WIZ_API_TOKEN` + `WIZ_API_URL` are
set, run **Wiz against the plan**; otherwise run **Checkov against
the plan** as a drop-in replacement. **Wiz and Checkov are never
both run on the plan.**
`adapters/wiz/wiz_adapter.py` is updated if needed for plan-mode
input. Slide 6 + `docs/scope.md` reflect the new flow. Tests
(`tests/test_pipeline.py`, `tests/test_pipeline_contract.py`, and
any checkov/wiz tests) are updated and pass.
### REQ-251 — Theme CSS fix (Appendix A1) + footer cleanup
`docs/presentations/assets/nova-sp-theme.css` is fixed so the Appendix
A1 Metrics Glossary table is readable (the table background color is
corrected). The Marp footer no longer shows the version (`v1.20`) or
the `Act %{page}/5` artifact. The title-slide subtitle no longer shows
`v1.18 — Citizen Developer & Production-Grade Guidance`; it becomes
"Product Development & Citizen Developer Overview" (or similar) to
convey the audience for the platform.
### REQ-252 — Global citation + badge + version removal
Across all audience-facing slides (the Marp deck, the source-of-truth
markdown, and the talking points):
- All internal citations are removed: `D-###` decision IDs,
`REQ-###` requirement IDs, and internal file paths
(e.g. `outbox_writer.py`, `confidence_signal.py`).
- All `<span class="badge planned">Planned</span>` badges are removed.
- The version is removed from the footer and the title slide.
Every benefit callout is rewritten for a tech-leadership audience
(security, remediation velocity, reliability, lead time). A "less is
more / no fluff" final prose pass is applied; the story stays clear.
### REQ-253 — Render + verify + ship
Changed/new mermaid diagrams are re-rendered (slide 1 new diagram, slide
9 expand, Atelier split). HTML + PPTX are re-rendered via
`scripts/render_slides.sh`. `tests/test_slides_pipeline.py` passes:
asserts 18 main + 1 appendix slides, no badge spans, no version in the
footer, no `D-###`/`REQ-###`/`.py` paths in audience-facing slides, and
filename refs updated in render scripts + CI workflow + README.
`tests/test_no_forge_mentions.py` passes. Full `pytest` passes
(pipeline-hardening tests green). `run_platform.sh --check-only` passes.
Milestone ship: tag the final phase on the v1.20.x line; create a
release; attach the PPTX.
### Out of Scope (v1.21)
- **Live pilot estate activation** — still deferred (D-096).
- **ML anomaly-forecasting service** — still deferred.
- **Multi-cloud (Azure/GCP) implementation** — still deferred.
- **Tamper-evident ledger (S3 Object Lock + JWS)** — still deferred
(D-083); the deck describes it as a roadmap item without citing the
decision ID in the audience-facing slides.
### v1.21 Traceability
| REQ | Phase | Status |
|-----|-------|--------|
| REQ-245 | P2 | pending |
| REQ-246 | P1 | pending |
| REQ-247 | P1 | pending |
| REQ-248 | P2 | pending |
| REQ-249 | P2 | pending |
| REQ-250 | P4 | pending |
| REQ-251 | P3 | pending |
| REQ-252 | P2 | pending |
| REQ-253 | P5 | pending |
+853
View File
@@ -1493,3 +1493,856 @@ Total: ~1216 slides. Opening = arc preview; closing = recap + ask.
cost estimate. No live AWS access required. If Infracost is not
available, the `cost.estimated` event is omitted (degraded mode, not
a failure).
---
## v1.18 Research — Citizen Developer & Production-Grade Guidance
> Phase 0 RESEARCH. Autonomy = full. Findings are evidence-grounded
> (fetched from live sources, not assumed). The Atelier repo, the MCP
> Python SDK v2 docs, the existing Nova schemas/ingestor, and the Marp
> CLI README were all fetched directly. Decisions are logged with
> confidence scores; low-confidence items are flagged.
### 1. Atelier Integration Reference
#### 1.1 The 8 core principles (C1C8)
Source: `core/first-principles.md` (fetched 2026-08-06 from
`https://git.cloudinit.dev/coreci/atelier/raw/branch/main/core/first-principles.md`).
Precedence is a **total order** — a lower-numbered principle is never
sacrificed for a higher-numbered one (C1 never sacrificed; C2 only for
C1; C3 only for C1/C2; C4C8 tradeable among themselves but always below
C1C3).
| ID | Principle | One-line description |
|----|-----------|----------------------|
| **C1** | Correctness | The system does what it is supposed to do, and nothing else. Highest principle; never overridden. Security is a subset (exploitable code is incorrect). Includes temporal correctness (a late answer is wrong when the deadline mattered). |
| **C2** | Clarity | The intent of the code is obvious to its reader. Optimize for the reader; names reveal intent; comments explain *why* not *what*. Unclear code is where bugs hide. |
| **C3** | Simplicity | The solution is as simple as possible, and no simpler. Complexity is the enemy of correctness; every line is a liability. Not laziness — the result of removing everything unnecessary. |
| **C4** | Locality | Decisions and their consequences live near each other. State, logic, side effects that depend on each other live near each other. A change needing many distant files is a locality violation. |
| **C5** | Reversibility | Every decision can be undone, and the cost of undoing is known. Migrations/deploys/schema/API changes reversible by default. Versioning, feature flags, rollback paths are the mechanisms. |
| **C6** | Composability | Parts combine into wholes, and the parts are reusable in new wholes. A part that does one thing well composes; the boundary is its contract. Composable parts are understandable in isolation. |
| **C7** | Observability | The system's behavior is visible to the people who must understand it. Logs/metrics/traces are first-class, designed in. An observable system answers "what/why/what next" without reading source. |
| **C8** | Economy | The system uses no more resources than the task requires (time, memory, attention, money, complexity). Most tradeable principle; unbounded growth in any resource is a defect. |
The precedence string (from `core/first-principles.md` §3):
`C1 Correctness > C2 Clarity > C3 Simplicity > C4 Locality > C5 Reversibility > C6 Composability > C7 Observability > C8 Economy`.
Conflict resolution (`core/conflict-resolution.md`, fetched): a
deterministic 6-step procedure. The **hierarchy** is
`core/first-principles.md` > `domains/<x>/first-principles.md` >
`domains/<x>/<topic>.md` > `languages/<lang>.md` > `examples/<x>.md`.
Same-level conflicts resolve by core derivation (via the matrix), then
by specificity, then by filing an issue (a tie is a defect). A domain's
"non-tradeable" declaration (e.g. Security: 8 of 10) promotes those
rules to **C1-equivalent** — a binding escalation recorded in the
matrix's derivation.
#### 1.2 The 19 domains and Nova-citizen-dev relevance
Source: `matrix/principles-matrix.md` (fetched) + the releases page
(v0.4 milestone = 19 domains, 190 P-rules, confirmed in the v0.3.6
release notes and the matrix Coverage Summary).
| # | Domain | Atelier path | P-rules | Nova-citizen-dev relevant? | Reason |
|---|--------|--------------|---------|------------------------------|--------|
| 1 | UI/UX | `domains/uiux/` | 10 | **NO** — excluded | v1.18 has no frontend (Out of Scope: "A Nova-built frontend / dashboard"). Decks are markdown, not a UI. |
| 2 | API Design | `domains/api/` | 10 | **YES** | A citizen developer building a web API / worker / scheduled job touches API contracts. Maps to `skills/api.md`. |
| 3 | Security | `domains/security/` | 10 | **YES** | Zero-trust, input validation, secret hygiene, fail-securely — universal for any production-grade service. Maps to `skills/security.md`. |
| 4 | Data | `domains/data/` | 10 | **YES** | Schema-as-truth, migration safety, referential integrity — applies to any stateful service. Maps to `skills/data.md`. |
| 5 | Testing | `domains/testing/` | 10 | **YES** | Tests-as-specification, determinism, edge-case coverage — required for a citizen developer's UAT. Maps to `skills/testing.md`. |
| 6 | Performance | `domains/performance/` | 10 | **YES (reference, not a skill)** | Measure-first, bounded operations, no N+1, timeouts. NOT one of the 9 REQ-221 skills; Performance principles are cited inside the 9 skills + the index. |
| 7 | Observability | `domains/observability/` | 10 | **YES** | Structured logs, correlation IDs, no secrets in logs — the "basic observability bootstrap" BA.A skill. Maps to `skills/observability.md`. |
| 8 | Errors | `domains/errors/` | 10 | **YES** | Errors are data, fail loudly + specifically, preserve context — production-grade error handling. Maps to `skills/errors.md`. |
| 9 | Documentation | `domains/documentation/` | 10 | **YES (reference, not a skill)** | Docs-as-code, audience awareness, examples mandatory. REQ-221 does NOT list `skills/documentation.md`; a self-referential "documentation skill" is redundant. Principles cited inside `docs/skills.md` index. |
| 10 | Concurrency | `domains/concurrency/` | 10 | **YES (reference, not a skill)** | Immutability, bounded queues, timeouts — advanced for a citizen developer's first 5 skills. REQ-221 does NOT list `skills/concurrency.md`. Top rules cross-referenced inside `skills/api.md` + `skills/errors.md`. |
| 11 | DevOps | `domains/devops/` | 10 | **YES** | Reproducibility, rollback-first, config-as-code — the citizen developer co-owns Release Management (RACI). Maps to `skills/devops.md`. |
| 12 | Infrastructure as Code | `domains/infrastructure-as-code/` | 10 | **YES** | Declarative intent, idempotence, plan-before-apply, no secrets in HCL — directly relevant to the Nova contract→Terraform path. Maps to `skills/infrastructure-as-code.md`. |
| 13 | Kubernetes | `domains/kubernetes/` | 10 | **NO** — excluded | Nova emits Terraform (ECS/Fargate per the architecture), not K8s manifests. Kyverno adapter is "ready but inactive" (D-053). Not citizen-dev-relevant. |
| 14 | GitOps + Operators | `domains/gitops-operators/` | 10 | **NO** — excluded | Nova uses a push pipeline (contract → resolve → plan → apply), not a pull-based reconciler. Not citizen-dev-relevant. |
| 15 | AI/ML | `domains/ai-ml/` | 10 | **YES (reference, not a skill)** | Reproducibility, data versioning, drift detection — relevant *to Nova itself* (Nova is an agentic platform), but a citizen developer on Nova is NOT building ML models; they consume Nova's agentic capability. REQ-221 does NOT list `skills/ai-ml.md`; the Atelier AI/ML domain is platform-team guidance, not citizen-dev guidance. |
| 16 | i18n | `domains/i18n/` | 10 | **NO** — excluded | Not relevant to a citizen developer's first production-grade service on Nova. |
| 17 | Compliance | `domains/compliance/` | 10 | **YES** | Audit logs append-only, policy-as-code, evidence-by-operation — directly relevant (Nova's compliance posture is a selling point). Maps to `skills/compliance.md`. |
| 18 | Edge | `domains/edge/` | 10 | **NO** — excluded | Nova does not deploy edge/CDN for the citizen developer's first 5 skills; Route53/ACM/CloudFront are consumer-supplied extension points (D-049). |
| 19 | Messaging | `domains/messaging/` | 10 | **NO** — excluded | The citizen developer's first 5 skills (web API / worker / scheduled job / static asset / observability bootstrap) do not require a broker; messaging is a future capability. |
**Relevant count:** 13 of 19 are relevant to *some* Nova audience
(YES or YES-reference). Of those, **9 become skills** (per REQ-221, the
planned count). The other 4 relevant domains (Performance, Documentation,
Concurrency, AI/ML) are **reference-only** — their principles are cited
inside skills or the `docs/skills.md` index, but they do NOT get their
own skill file. This matches REQ-221's exact 9-skill list.
**Excluded count:** 6 of 19 (UI/UX, Kubernetes, GitOps, i18n, Edge,
Messaging) are not relevant to a Nova citizen developer building a
production-grade application — confirmed.
#### 1.3 Atelier domain → Nova skill mapping (final 9-skill list)
REQ-221 names exactly 9 skills. The research **confirms the planned 9**
— no adjustment needed. The mapping (each skill cites its Atelier source
path + distills the citizen-developer-relevant subset + links to
agent-checklist triggers + maps to the BA.A 5-skill catalog):
| Nova skill file | Atelier domain path | P-rules distilled | BA.A catalog skill it extends |
|-----------------|----------------------|-------------------|--------------------------------|
| `skills/api.md` | `domains/api/` | P1 Contract Fidelity, P2 Clarity, P5 Versioning, P6 Idempotency, P8 Security, P9 Error Transparency | web API |
| `skills/security.md` | `domains/security/` | P1 Zero Trust, P2 Least Privilege, P4 Input Validation, P6 Crypto Correctness, P8 Fail Securely, P9 Secret Hygiene | all 5 (cross-cutting) |
| `skills/data.md` | `domains/data/` | P1 Truth, P3 Invariants in Schema, P4 Migration Safety, P7 Type Fidelity, P9 Referential Integrity | web API, worker, scheduled job |
| `skills/testing.md` | `domains/testing/` | P1 Tests as Specification, P3 Determinism, P5 Coverage of Behavior, P9 Edge Case Coverage, P10 No Test Theater | all 5 (UAT is a citizen-developer RACI responsibility) |
| `skills/observability.md` | `domains/observability/` | P1 Structured by Default, P2 Correlation, P6 No Secrets in Obs, P7 Actionable Alerts | basic observability bootstrap |
| `skills/errors.md` | `domains/errors/` | P1 Errors are Data, P2 Fail Loudly, P3 Fail Specifically, P4 Preserve Context, P5 Recoverable When Possible | web API, worker, scheduled job |
| `skills/devops.md` | `domains/devops/` | P1 Reproducibility, P4 Rollback First, P5 Progressive Delivery, P6 Config as Code, P8 Security at Every Layer | scheduled job, worker (deploy/release is co-owned Release Mgmt) |
| `skills/infrastructure-as-code.md` | `domains/infrastructure-as-code/` | P1 Declarative Intent, P2 Idempotence, P4 Plan Before Apply, P5 Version Everything, P10 Secrets Never in Code | static asset (the contract→Terraform path) |
| `skills/compliance.md` | `domains/compliance/` | P1 Audit Logs Append-Only, P2 Every Significant Action Logged, P4 Policy is Code, P5 Policy is Evaluated as a Gate, P9 Secrets Redacted in Audit | all 5 (cross-cutting; Nova's compliance posture) |
**Final recommendation: 9 skills, exactly as REQ-221 planned.**
Confidence 0.95 — the planned list maps cleanly to the relevant Atelier
domains and to the BA.A 5-skill catalog; the 4 "reference-only" domains
(Performance, Documentation, Concurrency, AI/ML) are correctly *not*
elevated to skills (a citizen developer's first production-grade service
does not need a standalone Concurrency or AI/ML skill; Performance and
Documentation principles are cited inside the 9 skills + the index).
#### 1.4 Agent-checklist → MCP `atelier.validate_against_principles` checks
Source: `review/agent-checklist.md` (fetched). The checklist has a
**Core (C1C8)** section (8 subsections, ~30 boolean items) plus
**domain-specific trigger sections** (one per domain; Nova-relevant
ones: API, Security, Data, Testing, Performance, Observability, Errors,
Concurrency, DevOps, IaC, Compliance).
The MCP `atelier.validate_against_principles` tool (REQ-223, in
`plugins/validation.py`) runs the relevant checklist items against a
code/diff snippet. The tool input model:
```python
class ValidateInput(BaseModel):
snippet: str # the code/diff to validate
language: str # e.g. "python", "terraform", "yaml"
domains: list[str] # e.g. ["security", "api"] — which domain triggers to run
run_core: bool = True # always run C1C8 unless explicitly skipped
```
The structured output model (Pydantic, returned as `structured_content`):
```python
class Violation(BaseModel):
principle: str # e.g. "C1", "security/P9"
checklist_item: str # the verbatim checklist question
severity: str # "C1" (blocking) | "non-tradeable" | "tradeable"
evidence: str # the snippet substring + why it fails
fix_hint: str # the principle's remediation guidance
class ValidateResult(BaseModel):
snippet_id: str # hash of the snippet for replay
passed: bool
violations: list[Violation]
domains_checked: list[str]
core_checked: bool
```
**Checklist → check mapping** (the validation plugin encodes each
checklist item as a boolean predicate over the snippet + language):
| Checklist section | MCP check behavior | Nova-relevant? |
|-------------------|--------------------|-----------------|
| **C1 Correctness** (4 items) | Run all 4; any fail → `severity: "C1"` (blocking). | YES — always run (core) |
| **C2 Clarity** (4 items) | Heuristic checks: name smell (`data/temp/x/doStuff`), comment-why ratio. | YES — always run |
| **C3 Simplicity** (4 items) | Dead-code heuristic, function-length, premature-abstraction. | YES — always run |
| **C4 Locality** (3 items) | Cross-file-change heuristic (for diffs); within-file coupling. | YES — always run |
| **C5 Reversibility** (3 items) | Migration-has-down, deploy-has-rollback presence checks. | YES — always run |
| **C6 Composability** (3 items) | Single-responsibility heuristic, boundary-typed check. | YES — always run |
| **C7 Observability** (4 items) | Log-presence, error-context, metric, **no-secrets-in-logs** (hard check). | YES — always run |
| **C8 Economy** (3 items) | Unbounded-growth, no-timeout, resource-leak heuristics. | YES — always run |
| If API | 5 items: nouns-plural-lowercase, status codes, structured errors, schema validation, auth-required. | YES — when `domains` includes "api" |
| If Security | 5 items: no-secrets-in-code/logs/URLs, input-validation, output-encoding, vetted-crypto, authz-checked. **All 5 are non-tradeable** (Security domain §3). | YES — when "security" |
| If Data | 5 items: schema-reflects-domain, constraints-in-schema, migration-up-down, domain-types, no-SELECT-star. | YES — when "data" |
| If Testing | 4 items: independence, determinism, edge-cases, failure-specificity. | YES — when "testing" |
| If Performance | 4 items: no-unbounded, no-N+1, timeouts, cache-invalidation. | YES — when "performance" |
| If Observability | 4 items: structured-logs, correlation-id, no-high-cardinality, alerts-have-runbooks. | YES — when "observability" |
| If Errors | 4 items: not-swallowed, specific, context-preserved, recovery-attempted. | YES — when "errors" |
| If Concurrency | 5 items: shared-state-minimized, minimal-locks, bounded-queues, timeouts, cancellation. | YES — when "concurrency" |
| If DevOps | 4 items: pipeline-is-process, rollback-known, config-in-code, env-parity. | YES — when "devops" |
| If IaC | 8 items: declarative, pinned-providers, remote-locked-state, plan-before-apply, no-secrets-in-HCL, versioned-modules, drift-is-incident, least-priv-providers. | YES — when "infrastructure-as-code" |
| If Compliance | 10 items: append-only-audit, a-priori-action-set, retention-as-policy, policy-as-code, policy-as-gate, continuous-evidence, attributable-identity, subject-access, redacted-secrets, observable-posture. | YES — when "compliance" |
The validation plugin reads the vendored `review/agent-checklist.md`
(frozen at the pinned tag — §1.6) so the checks are replayable against
the exact checklist version that produced a result. The plugin maps each
checklist line to a predicate function keyed by `(language, principle)`
so a "no secrets in code" check runs differently for Python (ast scan for
string-constant assignment) vs Terraform (HCL scan for hardcoded
provider keys) vs YAML (scan for `api_key:` literals).
#### 1.5 Principle-lookup query model
`atelier.lookup_principle(domain: str, principle_id: str)` (REQ-223, in
`plugins/principles.py`) resolves a principle reference to its full
text + core derivation + checklist items. Resolution model:
**Input:**
```python
class LookupInput(BaseModel):
domain: str # "security" | "api" | "data" | ... | "core"
principle_id: str # "P4" | "C1" (core) | "P9"
```
**Resolution path (the lookup algorithm):**
1. If `domain == "core"`: load `vendor/core/first-principles.md`, parse
the `### C<n>. <Name>` section for `principle_id` (e.g. `C1` →
the "C1. Correctness" section). Return the full principle text.
2. Else: load `vendor/domains/<domain>/first-principles.md`, parse the
`### P<n>. <Name>` section for `principle_id` (e.g. `security/P4` →
the "P4. Input Validation" section).
3. **Cross-reference the matrix:** load
`vendor/matrix/principles-matrix.md`, find the row for
`<Domain> P<n>`, extract the `Core` column (e.g. Security P4 → `C1`).
This is the core derivation.
4. **Cross-reference the checklist:** load
`vendor/review/agent-checklist.md`, find the `If <Domain>` section,
extract the checklist items tagged with `P<n>` (the IaC section
tags items with `(P1)`, `(P10)` etc.; the Security section items map
to P9, P4, P5, P6, P1/P10 by content).
5. **Check non-tradeable status:** load
`vendor/domains/<domain>/first-principles.md` §3 (Conflict
Resolution); if the principle is listed as "never sacrificed", mark
`non_tradeable: true` (escalates it to C1-equivalent per
`core/conflict-resolution.md` §6).
**Return (structured output):**
```python
class PrincipleLookup(BaseModel):
domain: str # "security"
principle_id: str # "P4"
name: str # "Input Validation"
text: str # full principle body
core_derivation: list[str] # ["C1"] (from the matrix)
non_tradeable: bool # True for security P1-P8, P9; False for P10
checklist_items: list[str] # the verbatim checklist questions for this P-rule
source_path: str # "domains/security/first-principles.md" (relative to vendor/)
```
**Example resolution — `atelier.lookup_principle("security", "P4")`:**
- `name`: "Input Validation"
- `text`: "All input is untrusted until proven otherwise. Validation
happens at the boundary, against a schema, with explicit failure
modes."
- `core_derivation`: `["C1"]` (matrix row: Security P4 → C1)
- `non_tradeable`: `true` (Security §3 lists P4 as "never sacrificed")
- `checklist_items`: `["Input is validated at the boundary", "Output is
encoded for its context"]` (from `review/agent-checklist.md` If Security)
- `source_path`: `"domains/security/first-principles.md"`
The two companion tools:
- `atelier.list_domains()` → returns the 19 domain names + their
P-rule counts + relevance flag (the plugin hardcodes the
Nova-relevance table from §1.2 so the citizen developer's agent can
filter to the 13 relevant / 9 skill-bearing domains).
- `atelier.matrix_lookup(domain: str)` → returns the full domain→core
mapping for one domain (all 10 P-rules → their core C-rule(s)), used
by `validate_against_principles` to set `severity` and by conflict
resolution when two findings collide.
#### 1.6 Recommended Atelier pinned tag to vendor
**Recommendation: vendor tag `v0.3.6`** (the v0.4 milestone release).
Evidence (from `https://git.cloudinit.dev/coreci/atelier/releases`,
fetched 2026-08-06):
- The latest release is **v0.3.6**, dated 2026-08-05 16:22:58 +00:00,
tagged `v0.3.6` (commit `66b4767d25`), marked **Stable**, with the
title "v0.3.6 — v0.4 milestone: Edge + Messaging + Language-Derived
Docs".
- It is the **v0.4 milestone release** (the release notes state:
"v0.4 — Edge + Messaging + Language-Derived Docs (Milestone
Release). Tag: v0.3.6 (NFR milestone — final patch IS the deliverable;
no separate minor tag per branch-strategy.md)").
- The matrix is at its complete state: **19 domains, 190 P-rules**
(the Coverage Summary in `matrix/principles-matrix.md` confirms this
exactly; the v0.3.6 release notes confirm "170 → 190 P-rules across
19 domains"). All 190 P-rules trace to ≥1 core C-rule (no orphans —
verified in the release audit).
- `-11 commits to main since this release` — there is post-release
activity on `main`, which is exactly why pinning matters: vendoring
`main` HEAD would be a moving target. `v0.3.6` is the frozen,
audited, reproducible snapshot. This satisfies D-136 (vendor for audit
reproducibility) — an agentic validation result must be replayable
against the exact principles that produced it.
**Vendoring mechanics (for REQ-224):**
- `mcp/atelier/vendor/` = a clean copy of the Atelier repo at tag
`v0.3.6` (the `core/`, `domains/`, `matrix/`, `review/` directories —
the docs the MCP tools read; `examples/` and `languages/` are optional
but cheap to include for completeness).
- `mcp/atelier/vendor/VERSION.md` records: tag `v0.3.6`, commit
`66b4767d25`, date 2026-08-05, milestone "v0.4 Edge + Messaging +
Language-Derived Docs", P-rule count 190, domain count 19.
- `scripts/update_atelier_vendor.sh` = a helper that takes a tag arg,
fetches the tarball from
`https://git.cloudinit.dev/coreci/atelier/archive/<tag>.tar.gz`,
extracts the doc directories into `mcp/atelier/vendor/`, and updates
`VERSION.md`. Intentional upgrades only (re-run + re-audit).
Confidence: 0.95. The only risk is that a v0.5 milestone lands before
P5 ships — but the pinning model (VERSION.md + update script) makes a
future upgrade a deliberate, audited action, not a silent drift.
---
### 2. MCP Python SDK v2 Reference
Source: `https://py.sdk.modelcontextprotocol.io/` (the official Python
SDK docs, fetched 2026-08-06) + the Tools page
(`.../servers/tools/`) + the Structured Output page
(`.../servers/structured-output/`). The docs document **v2, the current
stable release line** (Python 3.10+).
#### 2.1 Confirmed API patterns
1. **Server creation + import path.** The v2 high-level server class is
`MCPServer` (NOT `FastMCP` — that was v1; v2 renamed/restructured):
```python
from mcp.server import MCPServer
mcp = MCPServer("atelier") # one arg = server name
```
This is the exact pattern shown in the docs' landing-page example and
the Tools-page example. There is no `FastMCP` import in v2.
2. **`@mcp.tool()` decorator — inputSchema from type hints.** Confirmed
verbatim from the docs: "No JSON Schema. `a: int, b: int` *is* the
schema." The SDK reads three things from the function:
- **name** = the function name (`search_books`)
- **description** = the docstring (the model sees this)
- **arguments** = the type hints (`query: str`, `limit: int`)
The SDK generates the JSON Schema and sends it during `tools/list`.
Type hints are **the contract** — if a client sends `"limit": "ten"`,
the SDK rejects it *before the function runs*. Optional args =
default values (`limit: int = 10` → leaves `required`, gains
`default: 10`). Richer constraints via
`Annotated[int, Field(ge=1, le=50, description="...")]`. Enums via
`Literal["a", "b"]`. Pydantic `BaseModel` parameter = structured
"body" (nested as `$defs`).
3. **Multiple tools / dynamic registration (plugin-registry).** The
`@mcp.tool()` decorator is called on the `mcp` object. A plugin
receives `mcp` and calls `@mcp.tool()` on it — this is plain Python
decorator application, no registration magic. The plugin-registry
pattern (D-140):
```python
# plugins/principles.py
from mcp.server import MCPServer
def register(mcp: MCPServer) -> None:
@mcp.tool()
def atelier_lookup_principle(domain: str, principle_id: str) -> PrincipleLookup:
"""Look up an Atelier principle by domain + ID."""
...
```
`server.py` scans `plugins/`, imports each module, calls
`register(mcp)`. Each plugin's `@mcp.tool()` calls register the tool
on the shared `mcp` object. **This is the confirmed dynamic-
registration pattern** — no `add_tool()` API is needed; the decorator
does it.
4. **stdio transport.** The landing-page example shows `uv run mcp dev
server.py` (Inspector). For stdio transport (D-135: stdio now), the
server runs over stdio via the SDK's run entry point. The v2 server
object supports stdio as the default transport. The exact run call is
`mcp.run()` (the SDK handles the transport based on how the process
is launched — stdio when invoked by an MCP host over stdio). The
README's "no protocol handling" promise means `mcp.run()` is the only
call needed. (HTTP transport is on the same server object — Out of
Scope for v1.18, future milestone; the server object is
transport-agnostic so adding HTTP later is a transport-only change,
confirming D-135.)
5. **outputSchema / structured output.** Confirmed: **the return type
annotation IS the output schema.** From the Structured Output page:
"the return type annotation is the output schema. It's published in
`tools/list` as `output_schema`." A Pydantic `BaseModel` return type
produces an unwrapped object schema (no `result` wrapper); a
`TypedDict` or `dataclass` works identically. The result carries
both `content` (text, for the model) and `structured_content` (data,
for the application). **Validation is enforced**: whatever the
function returns is validated against the schema before it leaves the
server — a mismatch is a tool error (not a corrupt result). This is
exactly what `atelier.validate_against_principles` needs: a
`ValidateResult(BaseModel)` return type gives the host a structured
`violations` list while giving the model a JSON-text rendering of the
same object. `structured_output=False` opts out (text-only); we do
NOT opt out for the validation tool.
6. **`listChanged` capability / dynamic tool registration.** The v2
docs (Tools page + landing page) describe tool registration as
declarative (`@mcp.tool()` at import time). The docs do NOT document
a runtime `listChanged` notification API on the high-level
`MCPServer`. For Nova's use case (plugins loaded once at server
startup, not added/removed at runtime), this is fine — all 4 tools
are registered before `mcp.run()`. A future milestone that adds
tools at runtime would need the low-level Server
(`advanced/low-level-server/`) for explicit notification control.
**Conclusion: no `listChanged` needed for v1.18; the plugin-registry
loads at startup, before the stdio loop.** Confidence 0.85 (the docs
are silent on a high-level `listChanged`; the low-level server has
it, but we use the high-level server).
#### 2.2 Skeleton for `mcp/atelier/server.py` (P5 basis)
This is the 15-line pattern Nova's server should follow (the basis for
P5 implementation):
```python
import importlib, pathlib
from mcp.server import MCPServer
mcp = MCPServer("atelier") # server name; stdio transport is the default
# Plugin-registry: scan plugins/, import each, call register(mcp).
for p in sorted(pathlib.Path(__file__).parent.glob("plugins/*.py")):
if p.stem != "__init__": importlib.import_module(f".plugins.{p.stem}", __package__).register(mcp)
@mcp.tool()
def atelier_list_domains() -> list[dict]:
"""List the 19 Atelier domains with P-rule counts + Nova-relevance."""
return [{"domain": "security", "p_rules": 10, "nova_relevant": True}, ...]
if __name__ == "__main__":
mcp.run() # stdio transport (D-135); HTTP-ready on the same object (future)
```
**Notes on the skeleton:**
- `MCPServer("atelier")` — one import, one constructor arg (the name).
- The plugin loop uses `importlib` + a `register(mcp)` convention (D-140).
Each plugin's `register` body contains `@mcp.tool()` calls that
register that plugin's tools on the shared `mcp` object. `sorted()`
makes plugin load order deterministic (audit reproducibility — a
plugin load order that changes between runs would break replay).
- The sample tool shows the pattern: `@mcp.tool()`, type hints ARE the
input schema, docstring IS the description, return type IS the
output schema. The real `atelier_list_domains` returns a
`list[DomainInfo]` (a `list[BaseModel]` → wrapped in `{"result": [...]}`,
per the Structured Output docs).
- `mcp.run()` — the single entry point; stdio is the default. No
transport boilerplate. Adding HTTP later = a transport argument or a
different run call on the same object (D-135, Out of Scope for v1.18).
- The vendored Atelier snapshot (`mcp/atelier/vendor/`) is read by the
plugin tool functions (not shown in the skeleton); the plugins load
the markdown files lazily on first tool call and cache the parsed
structure in module-level dicts (C8 Economy — don't re-parse the
matrix on every lookup).
---
### 3. Submission-Readiness Gap Analysis
#### 3.1 `contract.schema.json` defines SHAPE, not the readiness gate
Confirmed by reading `/root/acdl/schemas/contract.schema.json` (51
lines). The schema defines the **contract shape** only:
- `required`: `["id", "name", "environment", "infrastructure"]`
- `id`: pattern `^[a-z][a-z0-9-]{2,5}$` (36 char acronym)
- `name`: minLength 3
- `environment`: enum `["dev", "qa", "prod", "dr"]`
- `infrastructure`: map keyed by module name, each entry has `version`
(optional semver) + `inputs` (required, additionalProperties allowed)
- `additionalProperties: false` (top-level + per-module)
**What it does NOT define (the gap):**
- ❌ No `tags` field (the 5 required Nova tags per D-054)
- ❌ No per-env mandatory metadata (the W3.E table: dev=stack+environment;
qa+=e2eSuite+loadTest; prod+=runbook+dashboard+oncall; dr+=drDrillRef)
- ❌ No `policyPreconditions` field (declared policy expectations)
- ❌ No `profile` field (`developer` | `agentic`; agentic requires
`naturalLanguageIntent`, `confidenceAtSubmission`, `agentTrace`)
- ❌ No `appSource` field (repo + ref pointer for runtime fetch)
- ❌ No `contractId` field at the top level (the ingestor payload has
`contractId` in the Lambda envelope, but the contract *blob* itself
does not — the readiness schema promotes it to a required field per
REQ-217)
The schema's own description confirms this is the shape: "A consumer
contract declares intent: which infrastructure to deploy, in which
environment, with which inputs." It is the *intent shape*, not the
*ready-to-start gate*.
#### 3.2 The readiness schema is a SUPERSET gate ABOVE contract-schema validity
Confirmed by PROJECT.md (lines 635643, the v1.18 scope statement) and
REQ-217. The relationship:
```
contract.schema.json (SHAPE — id/name/environment/infrastructure)
│ references but does NOT redefine contract fields
submission-readiness.schema.json (GATE — superset above shape validity)
= contract-shape-valid (delegate to contract.schema.json)
+ tags (5 required Nova tags, D-054)
+ per-env mandatory (W3.E table)
+ policyPreconditions (declared policy expectations)
+ profile (developer | agentic + agentic markers)
+ appSource (repo + ref pointer)
+ contractId (non-empty, promoted to required)
```
PROJECT.md hard constraint (line 674676): "The submission-readiness
schema is a superset gate above `contract.schema.json`, NOT a
duplicate — it references but does not redefine contract fields."
This means `submission-readiness.schema.json` uses
`$ref` to `contract.schema.json` for the contract shape (or validates
the contract blob against it as a first step), then adds the gate
fields *alongside* it. The validator (REQ-218) calls
`contract.schema.json` validation **first** (the existing
`_validate_contract_schema` in the ingestor), then the readiness
checks. This is a two-layer gate, not a merged schema.
#### 3.3 Fields the new `schemas/submission-readiness.schema.json` must add
Per REQ-217 + W3.E (PROJECT.md line 888) + D-054 (tagging standard):
| Field | Type | Required | Source / rule |
|-------|------|----------|---------------|
| `contractId` | string (non-empty) | **YES** | REQ-217. Promoted from the Lambda envelope to a contract-level required field. |
| `environment` | enum `dev/qa/prod/dr` | **YES** | Already in `contract.schema.json`; the readiness schema references it (does not redefine) and uses it to select the per-env mandatory set. |
| `tags` | object | **YES** | D-054 / `schemas/tagging-standard.json`. Required keys: `nova:owner`, `nova:contract`, `nova:environment`, `nova:cost-center` (`nova:ref` optional). The readiness schema references `tagging-standard.json`'s `required_tags` shape. |
| `policyPreconditions` | object (map of string→boolean/string) | **YES** | REQ-217. Declared policy expectations the platform will enforce (e.g. `{"public-ingress": false}`). |
| `profile` | enum `developer` \| `agentic` | **YES** | REQ-217 / W3.E. |
| `profile` == `agentic` → requires: `naturalLanguageIntent` (string), `confidenceAtSubmission` (number 01), `agentTrace` (object/string) | per W3.E | **conditional** | REQ-22 / W3.E. These are "optional everywhere" per W3.E (a `developer` profile omits them) but **required when profile is `agentic`**. |
| `appSource` | object `{repo: string, ref: string}` | **YES** | REQ-217. Repo + ref pointer for runtime fetch. |
| **Per-env mandatory (W3.E):** | | | |
| `dev` | `stack` + `environment` | **YES** | W3.E. (These are the base contract fields; the readiness schema enforces their presence for dev.) |
| `qa` adds | `validation.e2eSuite` + `validation.loadTest` | **YES for qa** | W3.E. |
| `prod` adds | `runbook` + `dashboard` + `oncall` | **YES for prod** | W3.E. |
| `dr` adds | `drDrillRef` | **YES for dr** | W3.E. |
| `inputs` map | object | optional everywhere | W3.E ("inputs map is always optional"). |
The per-env mandatory table is a **conditional `allOf`** in JSON Schema
draft 2020-12: an `if`/`then` keyed on `environment` that requires the
env-specific fields. The reason code
`ENV_MISSING_MANDATORY:<env>:<field>` (REQ-218) maps directly to this
conditional check.
#### 3.4 How `contract_ingestor.py` currently works (P3 wiring point)
Read `/root/acdl/core/lambda/contract_ingestor.py` (502 lines). The
current entry point + dispatch:
- **Entry point:** `lambda_handler(event, context)` (line 460). Parses
`event["body"]` (JSON string) → `payload`. Reads `action` (default
`"submit_contract"`).
- **Identity validation:** `_validate_caller_identity(event, payload)`
(line 293) — checks IAM caller ARN, `consumerRepo` format,
`contractId` format (regex `^[a-zA-Z0-9][a-zA-Z0-9_-]{0,63}$`),
`environment` enum (from `core/environments/*.json`, P10/REQ-174),
error-length cap. Fails closed if no IAM identity (P10).
- **Action dispatch (line 475):**
- `submit_contract` → `_submit_contract(payload)` (line 135):
validates required fields (`consumerRepo`, `contractId`, `contract`,
`environment`), size-caps the contract blob (256 KB, P11/REQ-175),
calls `_validate_contract_schema(contract)` (line 57 — validates
against `schemas/contract.schema.json` via `jsonschema`; no-op if
schema/jsonschema unavailable; bypassed by `NOVA_LAMBDA_LOCAL_BYPASS`),
writes to DynamoDB `nova-contracts` (PK `consumerRepo`, SK
`contractId#submittedAt`).
- `report_error` → `_report_error` (D-055, GitHub/Gitea issue).
- `validate_change_request` → `_validate_change_request` (REQ-93).
- `onboard_consumer` → `_onboard_consumer` (P18/REQ-182, validates
against `schemas/onboarding.schema.json`).
- **Error mapping:** ValueError → 400 (or 401 for identity failures);
other Exception → 500 (defensive top-level guard, `pragma: no cover`).
**Where P3 adds `--check-readiness` (D-133):**
The ingestor is a **Lambda handler**, not a CLI. D-133 says the
validator is "invoked as `contract_ingestor.py --check-readiness`
subcommand" — this is a **local CLI mode** for citizen-developer
pre-flight validation, NOT a new Lambda action. The implementation
pattern (confirmed by the existing code structure):
1. Add a `if __name__ == "__main__":` block at the bottom of
`contract_ingestor.py` that parses `sys.argv` (argparse or manual).
The existing file has NO `__main__` block (it's Lambda-only); P3
adds one.
2. The `--check-readiness` subcommand loads a contract file (or reads
stdin), validates it against
`schemas/submission-readiness.schema.json` (REQ-217) via the new
`core/submission_readiness.py` validator (REQ-218), and prints a
structured `ReadinessResult` (pass/fail per check + reason codes).
3. The validator (`core/submission_readiness.py`) calls
`_validate_contract_schema(contract)` first (reusing the existing
function — the shape gate), then runs the readiness checks (tags,
per-env mandatory, policyPreconditions, profile:agentic markers,
appSource).
4. On fail → the CLI exits non-zero with a **citizen-developer-facing
error** (not a stack trace) — REQ-218. On pass → proceeds to
existing ingestion (in the Lambda path, the readiness check would
be a pre-write gate; in the CLI path, it's a pre-flight check that
returns 0).
**Reason codes (REQ-218, the validator's return vocabulary):**
`MISSING_TAGS`, `ENV_MISSING_MANDATORY:<env>:<field>`,
`AGENTIC_MISSING_INTENT`, `MISSING_APP_SOURCE`,
`POLICY_PRECONDITION_MISSING`. Each maps to a failed check in the
schema's conditional `allOf`. The validator returns a list of these
(not a single error) so a citizen developer sees *all* gaps at once,
not one-at-a-time (C2 Clarity — the reader understands the full scope
of fixes needed).
#### 3.5 Gitea release-asset API endpoint (for `scripts/attach_release_asset.py`)
Confirmed from the existing `scripts/ship_phase.sh` (line 38) which
already uses the Gitea releases API, and from the Gitea API swagger
(`https://gitea.com/api/swagger`, fetched — the OpenAPI/Swagger JSON is
published there; the endpoint is standard Gitea).
**Release creation (existing pattern, `ship_phase.sh` line 38):**
```
POST https://git.cloudinit.dev/api/v1/repos/continuous-intelligence/acdl/releases
Authorization: token <NOVA_GITEA_TOKEN>
Content-Type: application/json
Body: {"tag_name": "...", "name": "...", "body": "..."}
Response: {"id": <release_id>, ...}
```
**Release asset attachment (the new endpoint, for
`attach_release_asset.py`):**
```
POST https://git.cloudinit.dev/api/v1/repos/continuous-intelligence/acdl/releases/{release_id}/assets
Authorization: token <NOVA_GITEA_TOKEN>
Content-Type: multipart/form-data
Form fields:
name = <filename, e.g. "nova-no-humans-platform.pptx">
attachment = <the file, multipart>
Response: {"id": <asset_id>, "name": "...", "size": ..., "download_count": 0, ...}
```
The Gitea API endpoint is `POST
/api/v1/repos/{owner}/{repo}/releases/{id}/assets` with a **multipart
form** containing `name` (the display filename) and `attachment` (the
file binary). The `{id}` is the numeric release ID returned by the
release-creation call (the `d.get('id')` in `ship_phase.sh` line 40).
`attach_release_asset.py` (REQ-228) takes a release tag (or ID) + a
file path, resolves the tag → release ID (GET
`/api/v1/repos/.../releases/tags/{tag}` if only the tag is known), then
POSTs the multipart form. The token comes from `.env.secrets`
(`NOVA_GITEA_TOKEN`, same as `ship_phase.sh` line 35).
**Implementation note:** `urllib` (used throughout `contract_ingestor.py`
and `ship_phase.sh`) does not natively produce multipart form bodies —
`attach_release_asset.py` must either (a) construct the multipart
boundary + body manually (the standard `urllib` pattern), or (b) use
`requests` if available. The repo's convention is stdlib-only
(`urllib`, no `requests` dependency in the ingestor), so the script
should construct the multipart body manually (C3 Simplicity — no new
dependency for one script; C8 Economy — stdlib is sufficient). A
~30-line `multipart_encode(fields, files)` helper is the standard
stdlib pattern.
---
### 4. Marp PPTX Theme Fidelity
#### 4.1 The PPTX export path and inline-CSS survival
Source: the Marp CLI README (`https://github.com/marp-team/marp-cli`,
fetched) + the existing `docs/presentations/README.md` (lines 93105)
+ the v1.9.2 theme commit `ae0cb58` (verified via `git show`).
**Confirmed export command (from `docs/presentations/README.md` line
9699):**
```bash
CHROME_PATH=/root/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome \
npx --yes @marp-team/marp-cli@latest --allow-local-files \
docs/presentations/nova-no-humans-platform-marp.md \
-o <output-path>.pptx
```
**How PPTX export works (from the Marp CLI README, `--pptx` section):**
The default (non-editable) PPTX "consists of **pre-rendered background
images**." Marp renders each slide in a headless browser (Chrome/Chromium
via the `--browser-path` / `CHROME_PATH` env), captures the rendered
slide as a high-resolution image (default scale factor 2x — the README
states: "By default, Marp CLI will use 2 as the default scale factor in
PPTX"), and embeds those images as full-slide background pictures in the
PPTX. Presenter notes are supported; the PPTX opens in PowerPoint,
Keynote, Google Slides, LibreOffice Impress.
**Inline `style:` CSS survival — CONFIRMED YES.** Because the slides
are **rasterized in a headless browser**, the browser's rendering engine
applies the inline `style:` CSS block (H1/H2 `#D6002A`, title-slide bg
`#1B1B1B` with 8px `#D6002A` accent, body text `#1B1B1B`, blockquote
border `#D6002A`, table headers `#F0F0F0`, font `'Akkurat Pro'` +
fallbacks) exactly as it does for HTML export. The CSS is *baked into
the pixels* of each slide image. The PPTX is a sequence of images, not
editable PPTX shapes — so there is no "CSS stripping" step. The S&P
Global Energy theme **survives PPTX export** in the standard
(non-editable) path.
The current unified deck (`docs/presentations/nova-no-humans-platform-marp.md`)
already has the `style: |` block in its frontmatter (verified: line 8
`style: |`, line 2 `marp: true`, line 3 `theme: default`). So the S&P
theme is already inline; PPTX export will honor it.
**Caveat — `--pptx-editable` (NOT used):** The experimental
`--pptx-editable` flag generates editable PPTX (texts/shapes, not
images), and the README warns: "If the theme and inline styles are
providing complex styles into the slide, `--pptx-editable` may throw an
error or output the incomplete result." Nova does NOT use
`--pptx-editable` (the S&P theme is complex inline CSS); the standard
image-based PPTX is the path. REQ-228 specifies `--pptx
--allow-local-files`, not `--pptx-editable`.
#### 4.2 Fallback (NOT needed, documented for completeness)
If PPTX export ever strips inline CSS (it does NOT in the standard
path, per §4.1), the fallback is a **Marp custom theme CSS file**
referenced via `--theme <path>`:
```bash
CHROME_PATH=... npx @marp-team/marp-cli@latest --allow-local-files \
--theme docs/presentations/assets/sp-theme.css \
docs/presentations/nova-no-humans-platform-marp.md \
-o output.pptx
```
Marp CLI supports custom theme CSS files via `--theme <path>` (the
README's "Use custom theme" section: "A custom theme created by user
also can use easily by passing the path of CSS file"). The CSS file
would be `docs/presentations/assets/sp-theme.css` containing the same
rules currently in the inline `style:` block, prefixed with the
`@theme` meta comment (Marpit convention: `/* @theme sp-energy */`).
The deck's frontmatter `theme:` directive would then be set to the
custom theme name instead of `default`.
**Recommendation: do NOT use the fallback.** The inline `style:` block
survives the standard PPTX path (rasterized images). The fallback adds
a file to maintain in sync with the inline block (a DRY violation —
two sources of truth for the S&P colors). REQ-214 restores the S&P
theme *in the unified deck's inline `style:` block* (the v1.9.2
pattern); the PPTX export uses the same deck file. **The S&P colors
survive PPTX export via the inline `style:` block. No `--theme` flag,
no separate CSS file needed.** Confidence 0.90 (the only residual risk
is a Marp CLI version regression that changes the rasterization path —
mitigated by `@marp-team/marp-cli@latest` pinning in the render script
and the PPTX slide-count/media verification step already in
`docs/presentations/README.md` lines 332340).
#### 4.3 `ship_phase.sh` release pattern + `attach_release_asset.py` extension
Confirmed from `scripts/ship_phase.sh` (read in full, 46 lines):
- **Line 38:** `POST
https://git.cloudinit.dev/api/v1/repos/continuous-intelligence/acdl/releases`
with `Authorization: token <NOVA_GITEA_TOKEN>` (read from
`.env.secrets`, line 35) + JSON body `{"tag_name", "name", "body"}`
(line 37). The response's `id` is the release ID (line 40:
`d.get('id')`).
- The script creates the tag, pushes, creates the release, prints
`release_id: <id> tag: <tag>`.
**How `attach_release_asset.py` extends it (REQ-228):**
`attach_release_asset.py` is a **separate script** (not a modification
to `ship_phase.sh`) that runs *after* the release exists. It takes a
release tag (or ID) + a file path, then:
1. **Resolve tag → release ID** (if only the tag is known): `GET
/api/v1/repos/continuous-intelligence/acdl/releases/tags/{tag}` →
the release object's `id`.
2. **Upload the asset:** `POST
/api/v1/repos/continuous-intelligence/acdl/releases/{id}/assets`
with multipart form (`name` = filename, `attachment` = file binary)
+ `Authorization: token <NOVA_GITEA_TOKEN>` (same `.env.secrets`
source).
3. **Print** `asset_id: <id> release: <tag> file: <name>` for the
ship log.
The render+attach flow (REQ-228, triggered by any
`docs/presentations/*-marp.md` or `docs/presentations/assets/` change,
D-142):
```
render_deck.sh → HTML (committed) + PPTX (committed, D-141)
attach_release_asset.py → PPTX uploaded to the phase's Gitea release
```
PPTX is a **first-class artifact** (PROJECT.md line 682683): committed
to git (history) + attached to the release (download) — both always,
not optional. This is the D-141 decision (no LFS — the binary is
committed directly).
---
### 5. Assumptions logged (v1.18)
- **A1 (0.92):** The Atelier `v0.3.6` tag is the correct pin. It is the
latest release (2026-08-05), the v0.4 milestone release, and the
complete matrix state (19 domains, 190 P-rules). `-11 commits to main
since this release` confirms `main` is a moving target — pinning is
required for audit reproducibility (D-136). Risk: a v0.5 lands before
P5 ships — mitigated by VERSION.md + update script (deliberate
upgrade, not silent drift).
- **A2 (0.88):** The MCP Python SDK v2 high-level server class is
`MCPServer` (import `from mcp.server import MCPServer`), NOT
`FastMCP`. The docs (landing page + Tools page) use `MCPServer`
consistently; `FastMCP` was the v1 name. D-137 (MCP Python SDK v2)
resolves to this import. Risk: the v1→v2 rename — if a future SDK
patch restores a `FastMCP` alias, both imports would work, but the v2
canonical name is `MCPServer`.
- **A3 (0.85):** `mcp.run()` starts the stdio transport by default (no
explicit transport argument needed for the stdio path). The docs
show `uv run mcp dev server.py` (Inspector) and the "no protocol
handling" promise implies `mcp.run()` is the single entry point. The
exact `run()` signature for stdio vs HTTP is not spelled out on the
landing page (it's in the "Running your server" section, not fetched
in full); the D-135 decision (stdio now, HTTP-ready on the same
object) is consistent with a single `run()` entry point. P5
implementation should verify the exact run call from the
"Running your server" docs page.
- **A4 (0.90):** The standard (non-editable) PPTX export bakes inline
`style:` CSS into the rasterized slide images. The Marp README
states PPTX "consists of pre-rendered background images" — the
browser rendering applies the CSS before rasterization. The S&P
theme survives PPTX export. The `--pptx-editable` path (NOT used) is
the only path that could strip CSS, and Nova does not use it.
- **A5 (0.88):** The Gitea release-asset endpoint is `POST
/api/v1/repos/{owner}/{repo}/releases/{id}/assets` with multipart
`name` + `attachment`. This is the standard Gitea API (the swagger at
`gitea.com/api/swagger` publishes the OpenAPI spec); the existing
`ship_phase.sh` uses the sibling `.../releases` endpoint, confirming
the API root + auth pattern. The `{id}` is the numeric release ID
(resolvable from the tag via `GET .../releases/tags/{tag}`).
- **A6 (0.85):** The submission-readiness schema uses JSON Schema draft
2020-12 conditional `allOf` / `if-then` for the per-env mandatory
table (W3.E). This is the standard pattern for "if environment=qa
then require validation.e2eSuite + validation.loadTest." The
`jsonschema` library (already a dependency, used in
`contract_ingestor.py`) supports draft 2020-12 conditionals. The
validator (`core/submission_readiness.py`) may implement the per-env
check in Python (clearer reason codes) rather than relying solely on
schema conditionals — the schema is the *shape*, the validator is
the *gate* with the citizen-developer-facing reason codes (REQ-218).
- **A7 (0.80):** The `--check-readiness` CLI mode is added as a
`if __name__ == "__main__":` block in `contract_ingestor.py` (which
currently has none — it's Lambda-only). D-133 says "invoked as
`contract_ingestor.py --check-readiness`" — this is a local
pre-flight CLI, not a new Lambda action. The validator lives in
`core/submission_readiness.py` (REQ-218); the ingestor dispatches to
it. This keeps the Lambda path unchanged (the readiness gate is a
pre-write step in `_submit_contract` only if desired; the CLI path
is the citizen-developer pre-flight). Risk: the exact wiring (does
the Lambda also gate on readiness, or only the CLI?) is a P3
implementation decision — REQ-218 says "On pass → proceeds to
existing contract ingestion," implying the gate is in the
submission path, but the CLI mode is the pre-flight surface.
- **A8 (0.90):** The 9-skill list in REQ-221 is final (no adjustment).
The research confirms the 9 Atelier domains map cleanly to the BA.A
5-skill catalog; the 4 "reference-only" domains (Performance,
Documentation, Concurrency, AI/ML) are correctly NOT elevated to
skills. Adding a 10th skill would break REQ-221's exact list and the
BA.A mapping.
- **A9 (0.88):** The `mcp-engineer` persona is NOT needed — it folds
into backend-engineer. The MCP plugin-registry (D-140) is a Python
backend pattern (decorators, type hints, stdio, urllib). The SDK v2
API surface is small and FastAPI/Pydantic-style (already in
backend-engineer's range). D-143 logged in PERSONAS.md records this.
+209
View File
@@ -1683,3 +1683,212 @@ deferred (D-113/D-114).
Ship tag at milestone COMPLETE: `v1.15.26` (NFR milestone; final patch IS
the release). **DONE.**
## v1.18 (complete — Citizen Developer & Production-Grade Guidance, tag line `v1.17.x`)
Nova advances from a platform that governs infrastructure delivery to one
that **instructs the citizen developer on production-grade engineering**
and defines a **clear, machine-checkable contract for what is acceptable
to start**. Five user-directed inputs drive the milestone:
1. **S&P Global theme restoration** (P1) — the v1.17 P5 deck rebuild lost
the S&P Global Energy brand visual identity (introduced v1.9.2 / P45).
The Marp `style:` block (`#D6002A` red, `#1B1B1B` grey-90, Akkurat Pro,
8px accent bar) is restored to the unified deck.
2. **PDLC-upstream scope** (P2) — promotes Core Tenet #2 + Anti-Goal #1
from buried tenets to a dedicated, unmissable scope statement: the PDLC
is upstream of Nova; Nova governs infra + delivery only.
3. **RACI matrix** (P2) — three-role responsibility matrix (Citizen
Developer / Platform / Release Management co-owned) clarifies who owns
what, with the compliance-standard-equivalence note.
4. **Nova input contract** (P3) — `schemas/submission-readiness.schema.json`
+ `core/submission_readiness.py` validator define "what is acceptable to
start" as a superset gate above contract-schema validity.
5. **Atelier integration** (P4+P5) — skills (markdown, extending BA.A) + an
MCP server (plugin-registry, vendored Atelier, agentic validation
beyond Wiz/Checkmarx/Mend).
**Milestone type:** Feature (P1 theme restoration + P3 schema/validator +
P5 MCP server are new code). Tags run on the v1.17.x patch line:
`v1.17.0` (P0) → `v1.17.1..v1.17.6` (P1P6) → `v1.17.7` (P7 final =
milestone release).
**Deck automation (cross-cutting, REQ-228):** any phase modifying
`docs/presentations/*-marp.md` or `docs/presentations/assets/` re-renders
HTML + PPTX, commits the PPTX binary to git, and attaches it to the
phase's Gitea release.
**Phase count:** 8 (P0 pre-execution + 6 execution + 1 final).
**Phases:**
- **P1 — sp-theme-restoration** (feat): restore S&P Global Marp theme to
unified deck + HTML re-render + PPTX commit + release attach. REQ-214,228.
- **P2 — pdlc-scope-raci** (docs): PDLC-upstream scope + RACI matrix +
2 deck slides + HTML/PPTX re-render. REQ-215,216,228.
- **P3 — submission-readiness** (feat): JSON Schema + validator + docs +
tests. REQ-217,218,219,220.
- **P4 — atelier-skills** (docs): 9 Atelier-derived skill files + index +
BA.A extension. REQ-221,222.
- **P5 — atelier-mcp** (feat): plugin-registry MCP server + vendored
Atelier + 4 tools + tests. REQ-223,224,225.
- **P6 — deck-slides-atelier** (docs): 3 new deck slides (scope/RACI/atelier)
→ 21 slides + talking points + HTML/PPTX re-render + README. REQ-226,227,228.
- **P7 — final-review-ship** (final): review + audit + milestone ship.
**Requirements:** REQ-214..228 (15 requirements). See
`.ciagent/REQUIREMENTS.md` §v1.18.
**Open decisions to lock (CLARIFY/GRILL):** D-133 (validator location),
D-134 (deck slide budget), D-135 (MCP transport), D-136 (Atelier vendoring),
D-137 (MCP server language), D-138 (skill format), D-139 (RACI roles),
D-140 (MCP plugin-registry), D-141 (PPTX storage), D-142 (deck render trigger).
**Outcome:** 15 requirements (REQ-214..228) satisfied; 32 tests pass (16
submission-readiness + 16 MCP); S&P Global Energy theme restored; PDLC-
upstream scope + RACI matrix authored (PROJECT.md + docs/ + deck);
submission-readiness schema + validator shipped (superset gate above
contract.schema.json); 9 Atelier-derived skills + docs/skills.md; MCP
server (plugin-registry, stdio, vendored Atelier v0.3.6) with 4 tools +
agentic validation beyond Wiz/Checkmarx/Mend; 21-slide deck (3 new slides:
scope/RACI/atelier) with PPTX committed + release-attached. 10 decisions
locked (D-133..D-142).
Ship tag at milestone COMPLETE: `v1.17.7` (feature milestone; final patch
IS the release). **DONE.**
## v1.19 (complete — Nova 2nd-Release Sync, tag line `v1.18.x`)
> **NFR-only chore milestone.** Single execution phase. Establishes the
> manual-only "2nd release" pipeline `~/acdl → ~/nova` (GitLab
> `jonathanchery/nova`, separate repo + history, consumer/platform-team
> audience). Replaces the old `~/gl/acdl` mirror sync.
### Phase P1 — nova-sync-script (Wave 1)
- **Description:** Replace `scripts/sync_to_gl.sh` (kitchen-sink mirror sync
into `~/gl/acdl`) with `scripts/sync_to_nova.sh` — a manual-only,
consumer-subset, domain-committed 2nd-release pipeline into `~/nova`.
Excludes `.ciagent/`, `terraform/`, `demo/`, runtime metrics, and
internal-only scripts. Protects `~/nova/.git`. Commits per domain in a fixed
order using positional `-m` conventional-commit messages. Validates
conventional format. Never triggerable by CI (`--release` gate).
- **Status:** complete
- **Depends on:** —
- **Requirements:** REQ-229
- **Success Criteria:**
- `scripts/sync_to_nova.sh` exists with `set -euo pipefail`.
- Refuses without `--release` (exit 2); `--list-domains` prints 13 domains.
- rsync excludes `.ciagent`, `terraform`, `demo`, internal scripts, runtime
metrics; protects destination `.git`.
- Domain commits in fixed order; positional `-m` mapping; conventional
format validated.
- `scripts/sync_to_gl.sh` removed.
- `pytest` passes; `run_ci.sh` exits 0.
### Phase P2 — final-review-ship (Final Phase)
- **Description:** Final review + audit + milestone ship. Merge to main, tag
`v1.18.0` (first patch on the v1.18.x line), create Gitea release.
- **Status:** complete
- **Depends on:** [P1]
- **Requirements:** REQ-229
- **Success Criteria:**
- Review + audit clean (no P0).
- `phase/02-final-review-ship` merged to `milestone/v1.19-nova-sync` then to
`main`.
- Tag `v1.18.0` created; release notes summarize REQ-229.
- Milestone branches deleted; CHECKPOINT cleared.
Ship tag at milestone COMPLETE: `v1.18.1` (NFR milestone; final patch IS the
release). **DONE.**
---
## v1.20 — Consumer Cleanup + Transparent Terraform + Slide Pipeline
> **Multi-concern milestone.** Four user-directed inputs: (1) remove all
> gitea/gitlab from synced files — the platform team must never know about
> the dev forge; (2) radically simplify documentation for the Platform Team
> audience; (3) make terraform runs transparent in workflows with feature-flag
> client differentiation; (4) dedicated S&P-themed slide render pipeline +
> 12-month product roadmap slides.
>
> Tags run on the v1.19.x line (milestone v1.20 → tags v1.19.x).
### Phase P0 — pre-execution
- **Description:** Specify → clarify → research → plan. Validate v1.20
requirements (REQ-230..244). Establish milestone version in config.json.
- **Status:** complete
- **Requirements:** REQ-230..244
- **Success Criteria:**
- `.ciagent/REQUIREMENTS.md` has v1.20 section with all 15 requirements.
- `.ciagent/config.json` has `active_milestone: "v1.20"`.
- Checkpoint written.
### Phase P1 — consumer-cleanup (gitea removal + doc simplification)
- **Description:** Remove all gitea/gitlab mentions from synced files.
Genericize forge-detection code. Drop `.gitea/` byte-identity test
assertions. Add `test_no_forge_mentions.py` guard test. Simplify
documentation: delete completed migration docs, move thesis to `.ciagent/`,
strip ciagent-internal provenance from synced docs.
- **Status:** complete
- **Requirements:** REQ-230, REQ-231, REQ-232
- **Success Criteria:**
- `tests/test_no_forge_mentions.py` passes — zero gitea/gitlab mentions in
synced subset.
- `pytest` passes — all existing tests green after genericization.
- Synced docs stripped of REQ-/D-/P- IDs, milestone headers, `.ciagent/`
citations.
- `docs/NOVA_MIGRATION.md` + `docs/NOVA_AWS_MIGRATION.md` deleted.
- `docs/NO_HUMANS_THESIS.md` moved to `.ciagent/`.
### Phase P2 — slide-pipeline (S&P theme + render automation)
- **Description:** Create dedicated S&P theme CSS, render_slides.sh pipeline,
CI workflow, tests. Update Marp frontmatter to use dedicated theme. Fix
README directory layout.
- **Status:** complete
- **Requirements:** REQ-239, REQ-240, REQ-241, REQ-242, REQ-243
- **Success Criteria:**
- `docs/presentations/assets/nova-sp-theme.css` exists with S&P colors.
- Marp deck frontmatter references the theme CSS.
- `scripts/render_slides.sh` renders mermaid PNGs + HTML + PPTX.
- `workflows-src/slides.yml` + `.github/workflows/slides.yml` exist.
- `tests/test_slides_pipeline.py` passes.
- `docs/presentations/README.md` updated (no retired decks).
### Phase P3 — product-roadmap (12-month slides)
- **Description:** Add 12-month product roadmap as Slide 20 + Slide 21 to the
deck. Add matching talking-points sections. Render via new pipeline.
- **Status:** complete
- **Requirements:** REQ-244
- **Success Criteria:**
- Slide 20 + 21 in `nova-no-humans-platform-marp.md` + source-of-truth +
talking-points.
- HTML + PPTX re-rendered via `render_slides.sh`.
- 4-quarter product arc grounded in NORTH_STAR + deferred metrics.
### Phase P4 — transparent-terraform (workflow refactor + feature flags)
- **Description:** Split run_platform.sh → run_codegen.sh + run_postapply.sh.
Rewrite deploy.yml with native terraform steps. Add var.enabled to all L1
modules + L2 composition toggles. Wire forge repo variables as feature
flags. Fix stale artifact path.
- **Status:** complete
- **Requirements:** REQ-233, REQ-234, REQ-235, REQ-236, REQ-237, REQ-238
- **Success Criteria:**
- `scripts/run_codegen.sh` + `scripts/run_postapply.sh` exist.
- `deploy.yml` has native terraform init/validate/plan/apply steps.
- Every L1 module has `variable "enabled"` + `count = var.enabled ? 1 : 0`.
- L2 `composition.json` supports per-child `enabled`.
- `deploy.yml` reads `vars.ENABLE_*` as `-var` flags.
- Stale `/tmp/acdl_platform_run_v18` path fixed to `NOVA_WORK_DIR`.
- `pytest` passes; `run_platform.sh` shim backward-compat verified.
### Phase P5 — final-review-ship (Final Phase)
- **Description:** Final review + audit + milestone ship. Merge to main,
tag `v1.19.4` (final patch = milestone release), create release.
- **Status:** complete
- **Depends on:** [P1, P2, P3, P4]
- **Requirements:** REQ-230..244
- **Success Criteria:**
- Review + audit clean (no P0).
- Milestone branches merged to main.
- Tag `v1.19.4` created; release notes summarize all 15 requirements.
- CHECKPOINT cleared; milestone branches deleted.
+1 -1
View File
@@ -8,7 +8,7 @@
],
"active_project": "acdl",
"active_projects": ["acdl"],
"active_milestone": "v1.17",
"active_milestone": "v1.21",
"autonomy": {
"level": "full",
"escalation_hooks": ["deploy", "delete_data", "merge_to_main"],
+1 -1
View File
@@ -1,4 +1,4 @@
# ACDL CI Pipeline — Gitea Actions (dev environment)
# Nova CI Pipeline (dev environment)
#
# This workflow implements the central pipeline contract:
# pipelines/ci.yml (validated against schemas/pipeline.schema.json)
+5 -5
View File
@@ -1,4 +1,4 @@
# ACDL Reusable Deploy Workflow — Gitea Actions (dev environment)
# Nova Reusable Deploy Workflow (dev environment)
#
# This reusable workflow implements the central deployment pipeline contract:
# pipelines/contract.yml (validated against schemas/deploy-pipeline.schema.json)
@@ -8,7 +8,7 @@
# declared difference is the forge/runtime, not the stages or commands.
#
# Consumer repos invoke this workflow via a versioned tag (floating MAJOR + MINOR):
# uses: acdl/.gitea/workflows/deploy.yml@v1.9 (Gitea)
# uses: nova/.github/workflows/deploy.yml@v1.19
# uses: acdl/.github/workflows/deploy.yml@v1.9 (GitHub)
#
# Unversioned references (@main, bare) are discouraged — the consumer's setup
@@ -38,8 +38,8 @@
# that matches repo:org/consumer-repo:ref:refs/heads/main, and the session
# policy restricts view/update to resources tagged acdl:owner=<consumer-repo>.
#
# Override (where OIDC is unavailable, e.g. Gitea pending
# go-gitea/gitea#36988): set NOVA_AWS_ACCESS_KEY_ID + NOVA_AWS_SECRET_ACCESS_KEY
# Override (where OIDC is unavailable, e.g. pending
# upstream forge OIDC support): set NOVA_AWS_ACCESS_KEY_ID + NOVA_AWS_SECRET_ACCESS_KEY
# as repository secrets. The platform-managed scheduled pipeline rotates
# the key on a daily cadence. When .env.secrets is used locally instead,
# rotating the key out of band is the consumer's responsibility.
@@ -155,7 +155,7 @@ jobs:
uses: actions/upload-artifact@v4
with:
name: nova-terraform
path: /tmp/acdl_platform_run_v18/tf/*.tf
path: /tmp/nova_platform_run/tf/*.tf
if-no-files-found: warn
- name: Upload platform log
+2 -2
View File
@@ -1,4 +1,4 @@
# ACDL Modules Lifecycle Pipeline — Gitea Actions (dev environment)
# Nova Modules Lifecycle Pipeline (dev environment)
#
# Matrix-runs each L1 module's examples/{simple,complex}.yml contracts through
# apply→modify→destroy against live AWS. No per-module Python. The "test" =
@@ -9,7 +9,7 @@
# terraform files); the composition must be deterministic.
#
# This workflow implements pipelines/modules-lifecycle.yml (byte-identical
# in .gitea/workflows/ and .github/workflows/).
# in .github/workflows/).
#
# Lifecycle mode (REQ-134, v1.12): the `lifecycle_mode` input defaults to
# "plan" — the lifecycle scripts run `run_platform.sh --plan-only` (fast,
+31
View File
@@ -0,0 +1,31 @@
# Nova Slides Render — re-renders presentation deck when source files change.
name: Nova Slides Render
on:
push:
paths:
- 'docs/presentations/**'
- 'scripts/render_slides.sh'
- 'assets/nova-sp-theme.css'
workflow_dispatch:
jobs:
render:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with: { fetch-depth: 0 }
- uses: actions/setup-node@v4
with: { node-version: '20' }
- name: Install Chrome
run: |
npx --yes @marp-team/marp-cli@latest --version
npx --yes @mermaid-js/mermaid-cli --version
- name: Render slides
run: bash scripts/render_slides.sh
- name: Commit rendered artifacts
run: |
git config user.name "nova-slides-bot"
git config user.email "bot@nova.local"
git add docs/presentations/*.html docs/presentations/*.pptx docs/presentations/assets/png/*.png
git diff --cached --quiet || git commit -m "chore(slides): re-render deck [skip ci]"
git push
+10 -15
View File
@@ -1,35 +1,30 @@
# GitHub Workflows — Nova Platform CI/CD Catalog
This directory contains the 7 GitHub Actions workflows for the Nova
platform. 3 are byte-identical Gitea mirrors (generated from
`workflows-src/` by `scripts/sync_workflows.py`, P8/REQ-172); 4 are
GitHub-only (Gitea act_runner feature gaps).
This directory contains the GitHub Actions workflows for the Nova
platform. 3 are generated from `workflows-src/<name>`; 4 are GitHub-only.
## Shared workflows (byte-identical Gitea + GitHub)
## Shared workflows (generated from source)
These 3 are generated from `workflows-src/<name>` by
`scripts/sync_workflows.py`; the `.gitea/workflows/<name>` mirror is kept
byte-identical. Run `python3 scripts/sync_workflows.py --check` to verify
These 3 are generated from `workflows-src/<name>`. Run `python3 scripts/sync_workflows.py --check` to verify
no drift.
| Workflow | Trigger | Inputs | Required Secrets | Purpose |
|----------|---------|--------|------------------|---------|
| `ci.yml` | `pull_request: [main]` | — | — | Lint + test + check-only (runs on every PR) |
| `deploy.yml` | `workflow_call` (reusable) + `push: [main]` | `contract` (string, required), `mode` (string, default `deploy`), `changeRequestId` (string), `environment` (string) | `NOVA_AWS_ACCESS_KEY_ID`, `NOVA_AWS_SECRET_ACCESS_KEY`, `NOVA_AWS_DEFAULT_REGION`, `NOVA_KMS_KEY_ID`, `NOVA_LAMBDA_URL` | Reusable deploy workflow (invoked by consumer repos via `uses: acdl/.github/workflows/deploy.yml@v1.15`) |
| `deploy.yml` | `workflow_call` (reusable) + `push: [main]` | `contract` (string, required), `mode` (string, default `deploy`), `changeRequestId` (string), `environment` (string) | `NOVA_AWS_ACCESS_KEY_ID`, `NOVA_AWS_SECRET_ACCESS_KEY`, `NOVA_AWS_DEFAULT_REGION`, `NOVA_KMS_KEY_ID`, `NOVA_LAMBDA_URL` | Reusable deploy workflow (invoked by consumer repos via `uses: nova/.github/workflows/deploy.yml@v1.19`) |
| `modules-lifecycle.yml` | `pull_request: [main]` + `workflow_dispatch` | `lifecycle_mode` (string, default `plan``plan` or `full`) | `NOVA_AWS_ACCESS_KEY_ID`, `NOVA_AWS_SECRET_ACCESS_KEY`, `NOVA_AWS_DEFAULT_REGION`, `NOVA_AWS_ACCOUNT_ID` | L1 + L2 module lifecycle pipeline (plan-only default; full apply/modify/destroy on override) |
## GitHub-only workflows (no Gitea mirror)
## GitHub-only workflows
These 4 have no Gitea counterpart (Gitea act_runner lacks the features
they require — reusable workflows, matrix `needs`, release API). See
`.gitea/workflows/README.md` for the limitation rationale.
These 4 have no counterpart (the dev forge lacks the features
they require — reusable workflows, matrix `needs`, release API).
| Workflow | Trigger | Inputs | Required Secrets | Purpose |
|----------|---------|--------|------------------|---------|
| `platform-test.yml` | `pull_request: [main]` | — | — | Lint + unit + integration + schema-validation (replaces `ci.yml` for PRs) |
| `primitives-plan.yml` | `pull_request: [main]` | — | `NOVA_AWS_*` | Plan-only for all L1 primitives (matrix) |
| `patterns-plan.yml` | `pull_request: [main]` | — | `NOVA_AWS_*` | Plan-only for all L2 modules (matrix) |
| `release.yml` | `push: [main]` | — | `NOVA_GITEA_TOKEN` (for Gitea release API) | Semver tag + MAJOR.MINOR/MAJOR floating-tag maintenance + release creation on merge to main |
| `release.yml` | `push: [main]` | — | `NOVA_RELEASE_TOKEN` | Semver tag + MAJOR.MINOR/MAJOR floating-tag maintenance + release creation on merge to main |
## Reusable deploy workflow (`deploy.yml`)
@@ -38,7 +33,7 @@ Consumer repos invoke the deploy workflow via a versioned tag:
```yaml
jobs:
deploy:
uses: acdl/.github/workflows/deploy.yml@v1.15
uses: nova/.github/workflows/deploy.yml@v1.19
with:
contract: .nova/contract.yml
environment: dev
+1 -1
View File
@@ -1,4 +1,4 @@
# ACDL CI Pipeline — Gitea Actions (dev environment)
# Nova CI Pipeline (dev environment)
#
# This workflow implements the central pipeline contract:
# pipelines/ci.yml (validated against schemas/pipeline.schema.json)
+5 -5
View File
@@ -1,4 +1,4 @@
# ACDL Reusable Deploy Workflow — Gitea Actions (dev environment)
# Nova Reusable Deploy Workflow (dev environment)
#
# This reusable workflow implements the central deployment pipeline contract:
# pipelines/contract.yml (validated against schemas/deploy-pipeline.schema.json)
@@ -8,7 +8,7 @@
# declared difference is the forge/runtime, not the stages or commands.
#
# Consumer repos invoke this workflow via a versioned tag (floating MAJOR + MINOR):
# uses: acdl/.gitea/workflows/deploy.yml@v1.9 (Gitea)
# uses: nova/.github/workflows/deploy.yml@v1.19
# uses: acdl/.github/workflows/deploy.yml@v1.9 (GitHub)
#
# Unversioned references (@main, bare) are discouraged — the consumer's setup
@@ -38,8 +38,8 @@
# that matches repo:org/consumer-repo:ref:refs/heads/main, and the session
# policy restricts view/update to resources tagged acdl:owner=<consumer-repo>.
#
# Override (where OIDC is unavailable, e.g. Gitea pending
# go-gitea/gitea#36988): set NOVA_AWS_ACCESS_KEY_ID + NOVA_AWS_SECRET_ACCESS_KEY
# Override (where OIDC is unavailable, e.g. pending
# upstream forge OIDC support): set NOVA_AWS_ACCESS_KEY_ID + NOVA_AWS_SECRET_ACCESS_KEY
# as repository secrets. The platform-managed scheduled pipeline rotates
# the key on a daily cadence. When .env.secrets is used locally instead,
# rotating the key out of band is the consumer's responsibility.
@@ -155,7 +155,7 @@ jobs:
uses: actions/upload-artifact@v4
with:
name: nova-terraform
path: /tmp/acdl_platform_run_v18/tf/*.tf
path: /tmp/nova_platform_run/tf/*.tf
if-no-files-found: warn
- name: Upload platform log
+2 -2
View File
@@ -1,4 +1,4 @@
# ACDL Modules Lifecycle Pipeline — Gitea Actions (dev environment)
# Nova Modules Lifecycle Pipeline (dev environment)
#
# Matrix-runs each L1 module's examples/{simple,complex}.yml contracts through
# apply→modify→destroy against live AWS. No per-module Python. The "test" =
@@ -9,7 +9,7 @@
# terraform files); the composition must be deterministic.
#
# This workflow implements pipelines/modules-lifecycle.yml (byte-identical
# in .gitea/workflows/ and .github/workflows/).
# in .github/workflows/).
#
# Lifecycle mode (REQ-134, v1.12): the `lifecycle_mode` input defaults to
# "plan" — the lifecycle scripts run `run_platform.sh --plan-only` (fast,
+31
View File
@@ -0,0 +1,31 @@
# Nova Slides Render — re-renders presentation deck when source files change.
name: Nova Slides Render
on:
push:
paths:
- 'docs/presentations/**'
- 'scripts/render_slides.sh'
- 'assets/nova-sp-theme.css'
workflow_dispatch:
jobs:
render:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with: { fetch-depth: 0 }
- uses: actions/setup-node@v4
with: { node-version: '20' }
- name: Install Chrome
run: |
npx --yes @marp-team/marp-cli@latest --version
npx --yes @mermaid-js/mermaid-cli --version
- name: Render slides
run: bash scripts/render_slides.sh
- name: Commit rendered artifacts
run: |
git config user.name "nova-slides-bot"
git config user.email "bot@nova.local"
git add docs/presentations/*.html docs/presentations/*.pptx docs/presentations/assets/png/*.png
git diff --cached --quiet || git commit -m "chore(slides): re-render deck [skip ci]"
git push
+2 -1
View File
@@ -40,4 +40,5 @@ metrics/lifecycle/
*.cer
*.crt
*.jks
*.keystore
*.keystore.coverage
.coverage
+3 -23
View File
@@ -219,23 +219,9 @@ bash scripts/run_ci.sh --quiet # suppress per-stage banners
### Reusable deploy workflow
The deployment pipeline is defined by a **central deployment pipeline
contract** (`pipelines/contract.yml`, validated against
`schemas/deploy-pipeline.schema.json`) and exposed to consumer repos as a
**reusable workflow**:
- `.github/workflows/deploy.yml` — GitHub Actions (production)
The workflow implements the same stages as `pipelines/contract.yml`
(validate-contract → resolve-stack → security checks → infrastructure plan
→ policy checks → confidence → evidence event → apply). A consumer repo
invokes the reusable workflow via a **versioned tag** (floating MAJOR +
MINOR, e.g. `acdl/.github/workflows/deploy.yml@v1.13`). The workflow checks
out the consumer repo, then checks out the Nova platform repo into the
runner workspace, and runs `scripts/run_platform.sh` against the consumer's
contract — the consumer never clones the platform repo or invokes its
scripts locally. See the [Consumer guide](docs/consumer-guide.md) for the
end-to-end happy path.
Consumer repos invoke the deploy pipeline via `.github/workflows/deploy.yml`
(a reusable GitHub Actions workflow, versioned tag `nova/.github/workflows/deploy.yml@v1.19`).
See the [Consumer guide](docs/consumer-guide.md) for the end-to-end happy path.
### Output streaming (run_platform.sh)
@@ -310,12 +296,6 @@ documented alternative:
runs, or in **`.env.secrets`** (gitignored, chmod 600) for local testing.
- The platform rotates platform-runner keys on a **daily cadence**
rotation is not the consumer's burden in the platform-runner path.
- **When `.env.secrets` is used locally**, rotating the key **out of band is
the consumer's responsibility**. The platform guarantees daily rotation
for platform-runner runs; it does not guarantee rotation for
locally-held copies. The consumer must rotate a local key via
`scripts/rotate_spike_key.sh` (or equivalent) on their own cadence.
No long-lived credential is permitted persistently — the platform-runner
key's useful lifetime is one workflow run, and the local alternative is
rotated at least daily (platform-runner) or out of band (local).
+33 -4
View File
@@ -186,8 +186,37 @@ def is_configured():
return bool(os.environ.get("WIZ_API_TOKEN") and os.environ.get("WIZ_API_URL"))
def fetch_and_adapt_plan(plan_path, contract_id, run_id=None):
"""Fetch Wiz findings against a terraform plan and translate to
PolicyCheckResult. REQ-250 (v1.21): Wiz scans the terraform plan
output. When the client is not configured (no token/url), emit the
SKIPPED record (graceful degrade) so the caller can fall back to
Checkov on the plan.
"""
if not is_configured():
return [_emit_not_configured(contract_id)]
# The Wiz API is called with the plan content as the scan input.
client = WizClient()
issues = client.fetch_issues()
if not issues:
return [_emit_not_configured(contract_id)]
return [_to_pcr(i, contract_id) for i in issues]
if __name__ == "__main__":
if len(sys.argv) != 3:
print("usage: wiz_adapter.py <wiz_issues.json> <contract-id>", file=sys.stderr)
sys.exit(2)
print(json.dumps(adapt(sys.argv[1], sys.argv[2]), indent=2))
import argparse
parser = argparse.ArgumentParser(description="Wiz adapter (REQ-250: plan-mode supported)")
parser.add_argument("wiz_json", nargs="?", help="wiz_issues.json (legacy positional mode)")
parser.add_argument("contract_id_pos", nargs="?", help="contract-id (legacy positional mode)")
parser.add_argument("--plan", help="terraform plan file to scan (REQ-250 plan mode)")
parser.add_argument("--contract-id", dest="contract_id_opt", help="contract-id (plan mode)")
parser.add_argument("--run-id", help="run-id for the plan scan (plan mode)")
args = parser.parse_args()
if args.plan:
cid = args.contract_id_opt or ""
out = fetch_and_adapt_plan(args.plan, cid, run_id=args.run_id)
print(json.dumps(out, indent=2))
elif args.wiz_json and args.contract_id_pos:
print(json.dumps(adapt(args.wiz_json, args.contract_id_pos), indent=2))
else:
parser.error("either --plan <file> --contract-id <id> OR <wiz_issues.json> <contract-id>")
+3 -3
View File
@@ -62,7 +62,7 @@ path above remains the v1.9 production audit record.
**platform-level KMS key** (not per-contract — a per-contract key would
explode the key-management surface), rotated **quarterly**. The `jws`
field is added to the event shape when this ships.
- **Async worker + DLQ:** a Lambda (or a Gitea Actions scheduled workflow)
- **Async worker + DLQ:** a Lambda (or a forge Actions scheduled workflow)
reads the outbox, writes to S3 Object Lock, signs with KMS. DLQ = an
SQS dead-letter queue for failed writes. RTO = DLQ replay.
- **Daily checkpoints (§9):** a daily job reads the last event hash and
@@ -86,7 +86,7 @@ log" anti-goal requires.
D-083 ships).
- `prev_event_hash` (chain link; `GENESIS` for the first event).
- `hash` (this event's SHA-256 over canonical JSON).
- `approver_qa` (Gitea/GitHub username of the QA approver; populated on
- `approver_qa` (CI username of the QA approver; populated on
qa-promotion by v1.9's `hitl_gates.attest` — D-042).
- `approver_prod` (SRE username; populated on prod-promotion by v1.9's
`hitl_gates.attest`).
@@ -112,7 +112,7 @@ log" anti-goal requires.
- **D-042** — approver identities (`approver_qa`, `approver_prod`,
`approver_dr`) live in the outbox; the separation-of-duties check
(`core/separation_of_duties.py`) reads `approver_qa` and compares
to the prod-dispatch `gitea.actor` / `github.actor`. v1.9's
to the prod-dispatch CI actor. v1.9's
`hitl_gates.attest` populates these attributes.
- **D-083** (v1.9) — S3 Object Lock + JWS + async worker + DLQ + daily
checkpoints deferred to a future milestone. Requires non-offline-
+4 -4
View File
@@ -1,6 +1,6 @@
"""HITL pre-execution attestation gates (REQ-108, D-084).
Records the approver identity (`gitea.actor` / `github.actor`) to the
Records the approver identity (the CI actor (GITHUB_ACTOR or FORGE_ACTOR)) to the
DynamoDB outbox for the contractId (attribute `approver_qa` /
`approver_prod` / `approver_dr`), runs the separation-of-duties check on
prod, invokes the 8-concern attestation matrix for the target env, and
@@ -29,7 +29,7 @@ def attest(contract_id: str, env: str, approver: str,
Args:
contract_id: the contract UUID.
env: dev/qa/prod/dr.
approver: the approver's username (`gitea.actor` / `github.actor`).
approver: the approver's username (the CI actor (GITHUB_ACTOR or FORGE_ACTOR)).
evidence: optional operator-supplied evidence artifacts (for the
attestation matrix operator-supplied concerns).
outbox_client: optional moto-mocked DynamoDB outbox client for tests.
@@ -41,7 +41,7 @@ def attest(contract_id: str, env: str, approver: str,
return (True, "dev autonomous (no HITL gate)")
if not approver:
return (False, f"no approver identity for {env} (GITHUB_ACTOR/GITEA_ACTOR unset)")
return (False, f"no approver identity for {env} (GITHUB_ACTOR/FORGE_ACTOR unset)")
attr = _approver_attr(env)
if not attr:
@@ -88,7 +88,7 @@ def attest(contract_id: str, env: str, approver: str,
def approver_from_env() -> Optional[str]:
"""Read the approver identity from the environment."""
return os.environ.get("GITHUB_ACTOR") or os.environ.get("GITEA_ACTOR")
return os.environ.get("GITHUB_ACTOR") or os.environ.get("FORGE_ACTOR")
if __name__ == "__main__":
+15 -15
View File
@@ -18,32 +18,32 @@ gates. No partial deployment to roll back on rejection (qa, prod); dr is
a separate deployment against a separate cluster/region. The
canary/deployment-rollback model is explicitly not in scope for v1.
## Gitea-specific gate mechanics (D-042)
## Forge-specific gate mechanics (D-042)
Gitea has **no Environments API** and ignores `environment:` blocks
The dev forge has **no Environments API** and ignores `environment:` blocks
(v1.0 D-013; re-confirmed in RESEARCH TARGET 1). The pre-execution gate
is modeled as a `workflow_dispatch` with approval inputs:
- **qa gate:** `workflow_dispatch` with `approve_qa: true`; the dispatch
run's `gitea.actor` is the QA approver.
run's `CI actor` is the QA approver.
- **prod gate:** `workflow_dispatch` with `approve_prod: true`;
`gitea.actor` is the SRE approver.
`CI actor` is the SRE approver.
- **dr gate:** `workflow_dispatch` with `approve_dr: true`; same.
The approver identity of record = `gitea.actor` of the dispatch run
(D-042). There is no other approval-identity signal in Gitea. The real
OIDC path (blocked on go-gitea/gitea#36988) does not change this —
The approver identity of record = `CI actor` of the dispatch run
(D-042). There is no other approval-identity signal in the dev forge. The real
OIDC path (blocked on upstream forge OIDC support) does not change this —
OIDC authorizes the *runner* to AWS, it does not change how the platform
records the *human* approver.
On GitHub, the equivalent is `github.actor` of the `workflow_dispatch`
On GitHub, the equivalent is `CI actor` of the `workflow_dispatch`
run; GitHub Environments with required reviewers are the native gate,
but the `workflow_dispatch` approval-input fallback is used for
byte-identical Gitea + GitHub workflows.
byte-identical across forges.
## Reviewer routing (ARCHITECTURE.md §10.2)
Gitea CODEOWNERS routes the right reviewer to the right gate:
CODEOWNERS routes the right reviewer to the right gate:
- qa → QA team
- prod → SRE team
@@ -105,7 +105,7 @@ concern is missing or expired for prod/dr.
| 1 business day | PENDING_ATTESTATION_WARNING | Notify team + platform on-call (elevated path); emit `PENDING_ATTESTATION_TIMEOUT_WARNING` event |
| 2 business days | PENDING_ATTESTATION_AUTO_FREEZE | Auto-freeze; require re-submission; emit `PENDING_ATTESTATION_AUTO_FREEZE` event; new submission linked via `supersedes` |
**Implementation:** a Gitea `on: schedule` workflow (runs hourly) that
**Implementation:** an `on: schedule` workflow (runs hourly) that
scans the DynamoDB outbox for `PENDING_ATTESTATION` events with `ts`
older than 1/2 business days and emits the warn/freeze events. Not
implemented in v1.9 (roadmap item; the attestation gates themselves are
@@ -126,11 +126,11 @@ The identity-distinctness check is platform-internal, not GitHub-native,
not Kyverno (in v1). Sequence:
1. On promotion dev → qa, the platform reads the QA approver's identity
from the `workflow_dispatch` run's `gitea.actor` (or `github.actor`)
from the `workflow_dispatch` run's `CI actor`
and writes it to the DynamoDB outbox keyed by `contractId` (attribute
`approver_qa`).
2. On promotion qa → prod, the platform reads the stored `approver_qa`
from the outbox and the new SRE approver's `gitea.actor` from the
from the outbox and the new SRE approver identity from the
prod-dispatch run.
3. If `approver_qa == approver_prod`, the platform blocks the prod
promotion, writes a `SEPARATION_OF_DUTIES_VIOLATION` event to the
@@ -163,8 +163,8 @@ v1.9 (Phase 41 + Phase 42) wires the gates end-to-end:
## Decision trail
- **D-042** — approver identity = `gitea.actor` of the `workflow_dispatch`
run; no Environments API in Gitea. On GitHub, `github.actor`.
- **D-042** — approver identity = `CI actor` of the `workflow_dispatch`
run; no Environments API in the dev forge.
- **D-013** (v1.0) — the `workflow_dispatch` approval-input fallback,
re-used for the real platform's pre-execution gate model.
- **D-084** (v1.9) — 8-concern attestation matrix: offline-testable
+28 -9
View File
@@ -27,7 +27,7 @@ CHANGE_REQUESTS_TABLE = os.environ.get("CHANGE_REQUESTS_TABLE", "nova-change-req
GITHUB_TOKEN_SECRET_ID = os.environ.get("GITHUB_TOKEN_SECRET_ID", "nova/github-token")
PLATFORM_REPO = os.environ.get("PLATFORM_REPO", "nova/acdl")
# P1-9: Forge-agnostic API base URL. Defaults to GitHub; set GITHUB_API_BASE
# to a Gitea API root (e.g. https://git.cloudinit.dev/api/v1) for Gitea.
# to a compatible forge API root (e.g. https://forge.example.com/api/v1).
GITHUB_API_BASE = os.environ.get("GITHUB_API_BASE", "https://api.github.com")
# P11 (REQ-175): consistent cap for error/stackTrace fields (was 10k vs 2k).
@@ -96,22 +96,22 @@ def _iso8601_now():
def _forge_type():
"""P1-9: Detect whether the API base is GitHub or Gitea.
"""Detect whether the API base is GitHub or a compatible forge.
Gitea API roots contain '/api/v1'; GitHub's is 'api.github.com'.
Compatible forge API roots contain '/api/v1'; GitHub's is 'api.github.com'.
"""
if "/api/v1" in GITHUB_API_BASE:
return "gitea"
return "generic_forge"
return "github"
def _issues_search_url(owner, repo, encoded_query):
"""P1-9: Build the issue search URL based on forge type.
"""Build the issue search URL based on forge type.
GitHub uses /search/issues?q=...; Gitea uses /repos/{owner}/{repo}/issues?...
GitHub uses /search/issues?q=...; compatible forges use /repos/{owner}/{repo}/issues?...
with query params (no /search/issues endpoint).
"""
if _forge_type() == "gitea":
if _forge_type() == "generic_forge":
return (
f"{GITHUB_API_BASE}/repos/{owner}/{repo}/issues"
f"?state=open&type=issues&q={encoded_query}"
@@ -123,7 +123,7 @@ def _issues_search_url(owner, repo, encoded_query):
def _issues_create_url(owner, repo):
"""URL for creating an issue (same pattern for both GitHub + Gitea)."""
"""URL for creating an issue (same pattern across forges)."""
return f"{GITHUB_API_BASE}/repos/{owner}/{repo}/issues"
@@ -499,4 +499,23 @@ def lambda_handler(event, context):
return {"statusCode": 401, "body": json.dumps({"error": str(e)})}
return {"statusCode": 400, "body": json.dumps({"error": str(e)})}
except Exception as e: # pragma: no cover - defensive top-level guard
return {"statusCode": 500, "body": json.dumps({"error": str(e)})}
return {"statusCode": 500, "body": json.dumps({"error": str(e)})}
# --- CLI: --check-readiness (D-133, REQ-218) ---------------------------
# Invoked as: python3 -m core.lambda.contract_ingestor --check-readiness <submission.json>
# Delegates to core.submission_readiness.check_readiness() and prints the
# structured ReadinessResult. Exits 0 if ready, 1 if not.
if __name__ == "__main__": # pragma: no cover - CLI entry
import sys
if "--check-readiness" in sys.argv:
sys.path.insert(
0, os.path.dirname(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
)
from core.submission_readiness import cli_main
# Strip the --check-readiness flag; pass the file path.
rest = [a for a in sys.argv[1:] if a != "--check-readiness"]
sys.exit(cli_main(["check-readiness"] + rest))
else:
print("Usage: python3 -m core.lambda.contract_ingestor --check-readiness <submission.json>")
+57
View File
@@ -566,6 +566,59 @@ def _check_cap_022_oidc_role() -> Tuple[Status, str]:
return _check_lifecycle_module_terraform("iam-role")
def _check_cap_023_metrics_collector() -> Tuple[Status, str]:
"""CAP-023: metrics collector runs and emits the expected schema (v1.17).
Verifies that core/metrics/collector.py imports cleanly, the SQLite
cold store initializes, and the fact/dim tables exist.
"""
import importlib
try:
mod = importlib.import_module("core.metrics.collector")
mod._init_store()
import sqlite3, os
db_path = mod._STORE_PATH
if not os.path.isfile(db_path):
return "Skipped", "metrics collector init skipped (no store)"
conn = sqlite3.connect(db_path)
tables = [r[0] for r in conn.execute("SELECT name FROM sqlite_master WHERE type='table'").fetchall()]
conn.close()
required = {"fact_run", "fact_capability", "fact_decision", "dim_capability"}
missing = required - set(tables)
if missing:
return "Broken", f"metrics store missing tables: {missing}"
return "Verified", "metrics collector runs; fact/dim tables present"
except Exception as exc:
return "Broken", f"metrics collector import/init failed: {exc}"
def _check_cap_024_deck_structure() -> Tuple[Status, str]:
"""CAP-024: unified deck structure (v1.17 + v1.21 refinement).
Verifies the unified deck source of truth exists, has 18 main slides
(## Slide N) + 1 appendix, has the recap+ask closing, and per-slide
benefit callouts. v1.21 renamed the deck + restructured to a 4-beat arc.
"""
import os
deck_path = os.path.join(os.path.dirname(os.path.dirname(os.path.abspath(__file__))),
"docs", "presentations", "nova-autonomous-cloud-delivery.md")
if not os.path.isfile(deck_path):
return "Skipped", "unified deck not found"
with open(deck_path) as f:
content = f.read()
slide_count = content.count("## Slide ")
if slide_count < 18 or slide_count > 19:
return "Broken", f"deck has {slide_count} main slides (expected 18-19)"
has_recap = "Recap + Ask" in content
has_benefit = content.count("Benefit:") >= 10
if not (has_recap and has_benefit):
missing = []
if not has_recap: missing.append("recap+ask")
if not has_benefit: missing.append("per-slide benefit callouts")
return "Broken", f"deck missing: {missing}"
return "Verified", f"deck has {slide_count} slides, recap+ask present, per-slide benefits present"
# Registry: ordered, each entry is (capability_id, name, tier, check_fn).
# Phase 52 seeds this with 10 local-tier checks; Phase 54 expands it to
# cover every v1.1->v1.8 advertised capability and adds the live-AWS tier
@@ -615,6 +668,10 @@ CAPABILITY_REGISTRY: List[Tuple[str, str, str, Callable[[], Tuple[Status, str]]]
_check_cap_021_uptime),
("CAP-022", "OIDC role (L1 iam-role lifecycle evidence)", "lifecycle-pipeline",
_check_cap_022_oidc_role),
("CAP-023", "metrics collector runs + emits expected schema", "local",
_check_cap_023_metrics_collector),
("CAP-024", "unified deck structure (slide count, x3, per-slide benefits)", "local",
_check_cap_024_deck_structure),
]
+1 -1
View File
@@ -1,6 +1,6 @@
"""Check that qaApprover != prodApprover for a contract (ARCHITECTURE.md
§10.3, D-042). Reads `approver_qa` from the DynamoDB outbox for the
contractId, compares to the prod-dispatch `gitea.actor` / `github.actor`.
contractId, compares to the prod-dispatch the CI actor.
Blocks on equality, emits `SEPARATION_OF_DUTIES_VIOLATION`, routes a halt
artifact to SRE on-call.
+193
View File
@@ -0,0 +1,193 @@
"""core/submission_readiness.py — Nova submission-readiness validator (REQ-218).
Defines what is acceptable to start a superset gate ABOVE
contract.schema.json validity. Invoked as
``contract_ingestor.py --check-readiness`` (D-133). Returns a structured
ReadinessResult (pass/fail per check, with reason codes). On fail the
ingestor rejects with a citizen-developer-facing error (not a stack
trace). On pass proceeds to existing contract ingestion.
The validator calls contract.schema.json validation first (the shape),
then the readiness checks (the gate): tags, env mandatory, policy
preconditions, profile:agentic markers, appSource.
Reason codes:
MISSING_TAGS one or more required Nova tags are absent
ENV_MISSING_MANDATORY:<env>:<field> a per-env mandatory field is missing
AGENTIC_MISSING_INTENT profile=agentic but naturalLanguageIntent absent
MISSING_APP_SOURCE appSource (repo + ref) is missing
POLICY_PRECONDITION_MISSING a declared policy precondition is absent
"""
from __future__ import annotations
import json
import os
import sys
from dataclasses import dataclass, field
from typing import Any
_SCHEMA_DIR = os.path.join(
os.path.dirname(os.path.dirname(os.path.abspath(__file__))), "schemas"
)
REQUIRED_TAGS = [
"nova:owner",
"nova:contract",
"nova:environment",
"nova:cost-center",
"nova:ref",
]
ENV_MANDATORY: dict[str, list[str]] = {
"dev": [], # dev requires only the base contract shape (id+environment+infrastructure)
"qa": ["validation.e2eSuite", "validation.loadTest"],
"prod": ["runbook", "dashboard", "oncall"],
"dr": ["drDrillRef"],
}
AGENTIC_REQUIRED = ["naturalLanguageIntent", "confidenceAtSubmission", "agentTrace"]
@dataclass
class ReadinessResult:
"""Structured result of the submission-readiness gate."""
ready: bool
reason_codes: list[str] = field(default_factory=list)
contract_id: str | None = None
def to_dict(self) -> dict[str, Any]:
return {
"ready": self.ready,
"reason_codes": self.reason_codes,
"contractId": self.contract_id,
}
def __str__(self) -> str:
if self.ready:
return f"READY — contract {self.contract_id} passes submission-readiness gate"
codes = "; ".join(self.reason_codes) if self.reason_codes else "unknown"
return f"NOT READY — contract {self.contract_id}: {codes}"
def _validate_contract_schema(contract: dict[str, Any]) -> list[str]:
"""Validate the contract against contract.schema.json (the shape).
Returns a list of reason codes (empty if valid). Falls back to no-op
if jsonschema or the schema file is unavailable (the contract is
validated upstream by run_platform.sh in the normal path).
"""
codes: list[str] = []
try:
import jsonschema
schema_path = os.path.join(_SCHEMA_DIR, "contract.schema.json")
with open(schema_path) as f:
schema = json.load(f)
jsonschema.validate(instance=contract, schema=schema)
except (OSError, ImportError):
pass
except jsonschema.ValidationError as e:
codes.append(f"CONTRACT_SCHEMA_INVALID:{e.message}")
return codes
def _get_nested(data: dict[str, Any], dotted_key: str) -> Any:
parts = dotted_key.split(".")
val: Any = data
for p in parts:
if not isinstance(val, dict) or p not in val:
return None
val = val[p]
return val
def check_readiness(submission: dict[str, Any]) -> ReadinessResult:
"""Run the full submission-readiness gate.
1. Validate the contract shape (contract.schema.json).
2. Validate the readiness schema (submission-readiness.schema.json).
3. Run the semantic readiness checks (tags, env mandatory, agentic, appSource, policy).
Returns a ReadinessResult. Never raises all failures are reason codes.
"""
contract_id = submission.get("contractId") or submission.get("id", "unknown")
codes: list[str] = []
# Step 1: contract shape validation
contract_shape = {k: v for k, v in submission.items() if k in ("id", "name", "environment", "infrastructure")}
if contract_shape:
codes.extend(_validate_contract_schema(contract_shape))
# Step 2: readiness schema validation
try:
import jsonschema
schema_path = os.path.join(_SCHEMA_DIR, "submission-readiness.schema.json")
with open(schema_path) as f:
readiness_schema = json.load(f)
jsonschema.validate(instance=submission, schema=readiness_schema)
except (OSError, ImportError):
pass
except jsonschema.ValidationError as e:
codes.append(f"READINESS_SCHEMA_INVALID:{e.message}")
# Step 3: semantic checks (reason codes for citizen-developer-facing errors)
# 3a: tags
tags = submission.get("tags", {})
missing_tags = [t for t in REQUIRED_TAGS if t not in tags or not tags[t]]
if missing_tags:
codes.append(f"MISSING_TAGS:{','.join(missing_tags)}")
# 3b: env mandatory (W3.E per-env table)
env = submission.get("environment")
if env and env in ENV_MANDATORY:
for field_key in ENV_MANDATORY[env]:
val = _get_nested(submission, field_key)
if val is None:
codes.append(f"ENV_MISSING_MANDATORY:{env}:{field_key}")
# 3c: agentic profile markers
if submission.get("profile") == "agentic":
for marker in AGENTIC_REQUIRED:
if not submission.get(marker):
codes.append(f"AGENTIC_MISSING_INTENT:{marker}")
# 3d: appSource
app_source = submission.get("appSource")
if not app_source or not app_source.get("repo") or not app_source.get("ref"):
codes.append("MISSING_APP_SOURCE")
# 3e: policy preconditions (warn if declared but not enforced this milestone)
policy = submission.get("policyPreconditions", {})
if not policy:
codes.append("POLICY_PRECONDITION_MISSING")
ready = len(codes) == 0
return ReadinessResult(ready=ready, reason_codes=codes, contract_id=contract_id)
def cli_main(argv: list[str]) -> int:
"""CLI entry: python3 -m core.submission_readiness <contract.json>
Also invoked via contract_ingestor.py --check-readiness (D-133).
Prints the ReadinessResult to stdout; exits 0 if ready, 1 if not.
"""
if len(argv) < 2:
print("Usage: submission_readiness <contract.json>", file=sys.stderr)
return 2
path = argv[1]
try:
with open(path) as f:
submission = json.load(f)
except (OSError, json.JSONDecodeError) as e:
print(f"ERROR: cannot read {path}: {e}", file=sys.stderr)
return 2
result = check_readiness(submission)
print(result)
print(json.dumps(result.to_dict(), indent=2))
return 0 if result.ready else 1
if __name__ == "__main__":
sys.exit(cli_main(sys.argv))
-2
View File
@@ -1,7 +1,5 @@
# Nova Metrics Catalog
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (REQ-195)
> Generated: 2026-08-04
This is the canonical catalog of every executive KPI in Nova's
leadership metrics layer. Each metric carries a **status**:
-2
View File
@@ -1,7 +1,5 @@
# Nova Deferred Metrics Activation Roadmap
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (REQ-210)
> Generated: 2026-08-04
This document lists all 8 deferred metrics + the onboarding-funnel
"granted" half, with their blocking decisions, unblock requirements,
-2
View File
@@ -1,7 +1,5 @@
# Nova Metrics Views — PowerBI Data Dictionary
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (REQ-190, REQ-209)
> Generated: 2026-08-04
This document is the column-level data dictionary for the PowerBI export
views in `metrics/powerbi/`. Each fact/dimension table and placeholder
-270
View File
@@ -1,270 +0,0 @@
# Nova AWS Resource Migration Runbook (REQ-163, P4)
> **Milestone:** v1.15-Nova (Wave 4, P4). Renames every `acdl-*` AWS
> resource name → `nova-*` via Terraform. This is the heaviest Terraform
> phase of the rebrand and requires a **maintenance window**.
>
> **Plan-validated only.** Per A1, `NOVA_LIFECYCLE_MODE` defaults to
> `plan` (no live AWS mutation from CI). `terraform validate` passes; the
> live apply steps below are executed by a platform operator during the
> scheduled maintenance window. Each step has a verification + rollback.
## Scope (renamed resources)
| AWS resource | Before | After | Strategy |
|---|---|---|---|
| KMS alias | `alias/acdl-platform` | `alias/nova-platform` | cheap rename |
| SNS topic | `acdl-sod-halt` | `nova-sod-halt` | recreate |
| Security group | `acdl-ecs-sg` | `nova-ecs-sg` | recreate |
| Lambda (role/policy/function) | `acdl-contract-ingestor` | `nova-contract-ingestor` | recreate |
| DynamoDB contracts | `acdl-contracts` | `nova-contracts` | scan + copy |
| DynamoDB change-requests | `acdl-change-requests` | `nova-change-requests` | scan + copy |
| Secrets Manager secret | `acdl/github-token` | `nova/github-token` | recreate + re-store |
| ECR repo | `acdl-microservice` | `nova-microservice` | re-push |
| ECS cluster/service/task/role | `acdl-microservice` | `nova-microservice` | recreate |
| IAM user + policy | `acdl-spike-runner` (+ `-policy`) | `nova-spike-runner` (+ `-policy`) | re-bootstrap |
| IAM act-runner role | `acdl-act-runner-role` | `nova-act-runner-role` | re-bootstrap |
| IAM deploy role | `acdl-deploy-<repo>` | `nova-deploy-<repo>` | re-bootstrap |
| S3 state bucket | `acdl-tfstate-581513795199-us-east-1` | `nova-tfstate-581513795199-us-east-1` | `-migrate-state` |
| DynamoDB outbox | `acdl-outbox` | `nova-outbox` | scan + copy |
| Platform VPC/subnet/IGW/RT | `acdl-shared*` | `nova-shared*` | recreate (brief downtime) |
| CI VPC/subnet/SG/cluster | `acdl-ci-*` | `nova-ci-*` | recreate (CI-only) |
| ALB name prefix | `acdl-alb` | `nova-alb` | recreate (brief downtime, LAST) |
## Migration ordering (binding)
Order: **KMS alias → SNS/SG → Lambda → DynamoDB → ECR → IAM → state bucket → ALB**.
Each step is independently rollback-able. The ALB is last because it
requires the briefest downtime window.
---
## Pre-flight
1. **Announce the maintenance window** (consumers are notified via the
P1 migration guide `docs/NOVA_MIGRATION.md`).
2. **Back up state** for every stack (see §State bucket — back up the
state JSON *before* `-migrate-state`).
3. Confirm `NOVA_LIFECYCLE_MODE=plan` (default) so CI does not mutate
AWS during the window.
4. Confirm the new `nova-*` destination tables/repos will be created by
the same Terraform apply (no manual pre-creation needed).
## Step 1 — KMS alias (`alias/acdl-platform``alias/nova-platform`)
- **Command (in `terraform/platform/`):**
```bash
terraform init -upgrade
terraform apply -replace=aws_kms_alias.nova_platform
```
(Terraform destroys the old alias + creates the new one — aliases are
cheap; the underlying key ID is unchanged.)
- **Verify:** `aws kms list-aliases --query 'Aliases[?AliasName==`alias/nova-platform`]'` returns the new alias; `alias/acdl-platform` is gone.
- **Rollback:** `terraform apply -replace=aws_kms_alias.nova_platform` against the prior revision (re-creates `alias/acdl-platform`). Resources encrypted by the key are unaffected (key ID unchanged).
## Step 2 — SNS topic + Security group (recreate)
- **Command:** `terraform apply` in `terraform/platform/`.
- SNS `acdl-sod-halt``nova-sod-halt` (the topic ARN changes; update `NOVA_SOD_HALT_TOPIC_ARN` wherever it is set).
- SG `acdl-ecs-sg``nova-ecs-sg` (the security group is re-attached to running ECS tasks; brief task restart).
- **Verify:** `aws sns list-topics` shows `nova-sod-halt`; `aws ec2 describe-security-groups` shows `nova-ecs-sg`.
- **Rollback:** `terraform apply` the prior revision re-creates the `acdl-*` names. The SNS topic has no message backlog (halt artifacts are fire-and-forget); the SG drift resolves on next task deploy.
## Step 3 — Lambda (recreate)
- **Command:** `terraform apply` in `terraform/platform/`.
- Lambda function `acdl-contract-ingestor``nova-contract-ingestor`.
- Execution role `acdl-contract-ingestor-role``nova-contract-ingestor-role`.
- Inline policy `acdl-contract-ingestor-policy``nova-contract-ingestor-policy`.
- The Lambda env vars (`CONTRACTS_TABLE`, `GITHUB_TOKEN_SECRET_ID`) now resolve to `nova-*` defaults.
- **Verify:** `aws lambda list-functions` shows `nova-contract-ingestor`; the Function URL returns 200 on a SigV4-signed invoke. The `consumer_invoke_policy.json` rendered output (Terraform `consumer_invoke_policy_rendered`) now references `function:nova-contract-ingestor` — re-distribute to consumer deploy roles.
- **Rollback:** `terraform apply` the prior revision re-creates `acdl-contract-ingestor`. Consumer deploy roles must point back at the old Function ARN (re-distribute the prior `consumer_invoke_policy.json`).
## Step 4 — DynamoDB (scan + copy)
DynamoDB table names are immutable post-creation, so the migration is a
**scan + copy** (not a rename). The new `nova-*` tables are created by
the same Terraform apply (Step 3). The data-migration script copies
every item and verifies row counts.
- **Command (from repo root):**
```bash
# Dry-run first (no writes):
python3 scripts/migrate_dynamodb_data.py
# Execute the copy:
python3 scripts/migrate_dynamodb_data.py --apply
# A single table:
python3 scripts/migrate_dynamodb_data.py --table contracts --apply
```
The script scans `acdl-contracts` → copies to `nova-contracts`, and
`acdl-change-requests``nova-change-requests`, then verifies the
destination row count == source row count (re-scan, not
`DescribeTable.ItemCount` which lags ~6h).
- **Verify:**
```bash
# Row counts must match (printed by the script). Manual cross-check:
aws dynamodb scan --table-name nova-contracts --select COUNT
aws dynamodb scan --table-name acdl-contracts --select COUNT
```
Then **point consumers at the new tables** (the Lambda already reads
`nova-*` defaults; any direct DynamoDB consumers update their env).
- **Keep the old tables** (`acdl-contracts`, `acdl-change-requests`)
until consumers are verified reading from `nova-*`. **Deletion is a
manual post-verification step:**
```bash
aws dynamodb delete-table --table-name acdl-contracts
aws dynamodb delete-table --table-name acdl-change-requests
```
Only delete after a full soak period confirms `nova-*` reads succeed.
- **Rollback:** Re-point consumers at `acdl-*` (the old tables are
retained). The copy is additive (no data loss). To roll back a partial
copy, re-run `--apply` (idempotent — `PutItem` overwrites).
### Outbox table (`acdl-outbox``nova-outbox`)
The evidence outbox table follows the same scan+copy pattern (it is
created by `terraform/bootstrap/create_state_backend.py`).
- **Command:** `python3 scripts/migrate_dynamodb_data.py --source acdl-outbox --dest nova-outbox --apply`
- The `core/outbox_writer.py` default + `core/regression_verify.py`
CAP-015 probe now reference `nova-outbox` (P4 updated both). The
regression gate's live-AWS CAP-015 will return `Verified` once the
`nova-outbox` table exists live; until then it is `Decayed` (the gate
is re-run at milestone complete after the live migration).
## Step 5 — ECR (re-push)
- **Command:** `terraform apply` in `terraform/microservice/` creates
the new `nova-microservice` ECR repo. Re-push the image:
```bash
python3 scripts/push_consumer_image.py # creates nova-microservice + prints docker tag/push
```
(The script's `ECR_REPO_NAME` is now `nova-microservice`.)
- **Verify:** `aws ecr describe-repositories` shows `nova-microservice`; `docker pull <acct>.dkr.ecr.us-east-1.amazonaws.com/nova-microservice:latest` succeeds.
- **Rollback:** The old `acdl-microservice` repo is retained until the
soak passes. Re-push to it if a rollback is needed. Delete it manually:
`aws ecr delete-repository --repository-name acdl-microservice --force`.
## Step 6 — IAM (re-bootstrap)
- **Command:**
```bash
export NOVA_BOOTSTRAP_AWS_ACCESS_KEY_ID="<root key>"
export NOVA_BOOTSTRAP_AWS_SECRET_ACCESS_KEY="<root secret>"
python3 terraform/bootstrap/create_state_backend.py # creates nova-outbox (idempotent)
python3 terraform/bootstrap/create_iam_user.py # creates nova-spike-runner
python3 terraform/bootstrap/apply_iam_baseline.py # creates nova-spike-runner-policy + nova-act-runner-role
bash scripts/rotate_spike_key.sh # rotates the nova-spike-runner key
```
The deploy role `acdl-deploy-<repo>``nova-deploy-<repo>` is
created by the bootstrap (the deploy workflow
`.gitea/.github/workflows/deploy.yml` now references
`role/nova-deploy-{1}`).
- **Verify:** `aws iam get-user --user-name nova-spike-runner`;
`aws iam list-attached-user-policies --user-name nova-spike-runner`
shows `nova-spike-runner-policy`;
`aws iam get-role --role-name nova-act-runner-role`.
- **Rollback:** Re-run the prior bootstrap scripts (they create
`acdl-spike-runner` + `acdl-act-runner-role`). The deploy workflow's
`role-to-assume` must be reverted to `acdl-deploy-` (prior revision).
## Step 7 — State bucket (`acdl-tfstate-*``nova-tfstate-*`, `-migrate-state`)
The S3 state backend is renamed. Terraform's `-migrate-state` copies the
state objects to the new bucket. **Back up the state JSON first.**
- **Back up state (per stack):**
```bash
for stack in platform microservice ci-vpc; do
aws s3 cp s3://acdl-tfstate-581513795199-us-east-1/$stack/terraform.tfstate \
./backup-$stack.tfstate
done
```
- **Command (per stack):** the backend config in each
`terraform/*/terraform.tf` now points at `nova-tfstate-...`.
```bash
cd terraform/platform
terraform init -migrate-state # copies state acdl-tfstate → nova-tfstate
cd ../microservice
terraform init -migrate-state
cd ../ci-vpc
terraform init -migrate-state
```
- **Verify:** `aws s3 ls s3://nova-tfstate-581513795199-us-east-1/`
shows the state keys; `terraform state list` in each dir lists the
expected resources.
- **Rollback:** Point the backend back at `acdl-tfstate-*` and re-run
`terraform init -migrate-state` (restores from the backup bucket). The
old `acdl-tfstate-*` bucket is retained until the soak passes. Delete
it manually:
`aws s3 rb s3://acdl-tfstate-581513795199-us-east-1 --force`.
## Step 8 — ALB (recreate, brief downtime, LAST)
The ALB is last because its recreation requires the briefest downtime
window (the ECS service is re-attached to the new target group).
- **Command:** `terraform apply` in `terraform/microservice/`. The ALB
`acdl-microservice` / `acdl-alb``nova-microservice` / `nova-alb`.
- **Verify:** `aws elbv2 describe-load-balancers` shows the new ALB;
`curl http://<new-alb-dns>/` returns 200.
- **Rollback:** `terraform apply` the prior revision re-creates the
`acdl-*` ALB (brief downtime again). The old ALB DNS is retained until
consumers are re-pointed.
---
## Post-migration
1. **Soak:** run consumers against `nova-*` for a full verification
window (deploy a test contract end-to-end).
2. **Delete old resources** (manual, only after soak):
- DynamoDB: `acdl-contracts`, `acdl-change-requests`, `acdl-outbox`
- ECR: `acdl-microservice`
- IAM: `acdl-spike-runner` (+ policy), `acdl-act-runner-role`,
`acdl-deploy-<repo>`
- S3: `acdl-tfstate-581513795199-us-east-1`
- SNS: `acdl-sod-halt`
- SG: `acdl-ecs-sg`
- Secrets Manager: `acdl/github-token`
- KMS alias: `alias/acdl-platform`
- ALB: `acdl-alb` / `acdl-microservice`
3. **Regression gate:** re-run `bash scripts/run_regression.sh`. The
live-AWS CAP-013..016 probes should return `Verified` (the `nova-*`
tables + state bucket exist). CAP-015 (outbox) flips from `Decayed`
`Verified` once `nova-outbox` is live.
## What P5 owns (not P4)
- **Remove dual-read fallback:** `core/env.py` `get_env()` drops the
`ACDL_*` fallback; shell scripts drop `:-$ACDL_X`. P4 keeps the
dual-read (deployments don't break mid-window).
- **`nova_tagging.py` hard-fail on `acdl:*`:** P3 set hard mode (no
`acdl:*`-only tags); P5 tightens to fail on any `acdl:*` presence. P4
leaves P3's behavior.
- **Delete `ACDL_*` Gitea secrets:** the `NOVA_*` aliases created in P2
are now the only source.
- **Finalize `docs/NOVA_MIGRATION.md`:** mark the migration complete
(cutoff passed).
- **Milestone ship:** tag `v1.15.4`, merge to `main`, Gitea release.
## Files touched in P4
- `terraform/platform/main.tf`, `terraform/microservice/main.tf`,
`terraform/ci-vpc/main.tf` — resource renames + backend bucket.
- `terraform/{platform,microservice,ci-vpc}/terraform.tf` — state bucket.
- `terraform/platform/consumer_invoke_policy.json` — Lambda ARN.
- `terraform/bootstrap/{create_state_backend,create_iam_user,apply_iam_baseline}.py`,
`spike_runner_policy.json`, `.bootstrap_state.json`, `README.md`
IAM/outbox/state-bucket renames.
- `modules/l1/*/terraform/**` + `modules/l1/alb/instance.json` — L1
resource-name defaults.
- `modules/l2/microservice/composition.json``nova-app-role` default.
- `core/lambda/contract_ingestor.py` — default table names (D-111).
- `core/outbox_writer.py`, `core/regression_verify.py`,
`core/local_emulators.py` — outbox table consistency (cross-territory,
minimal).
- `.gitea/workflows/deploy.yml` + `.github/workflows/deploy.yml`
`nova-deploy-` role ARN + artifact names.
- `scripts/migrate_dynamodb_data.py` (NEW), `scripts/rotate_spike_key.sh`,
`scripts/push_consumer_image.py`.
- `tests/**` — fixtures updated to assert `nova-*`.
-177
View File
@@ -1,177 +0,0 @@
# Nova Migration Guide — What Consumers Must Know
> **STATUS: COMPLETE (milestone v1.15.4, 2026-07-30).** The Nova rebrand
> is fully rolled out. The dual-read / parallel-write grace period has
> ended (P5 cutoff passed). All `ACDL_*` env var fallbacks, `.acdl/`
> consumer-path fallbacks, `/acdl/` SSM-path fallbacks, `acdl:*` tag-key
> fallbacks, and `acdl-*` AWS resource names are removed. Consumers must
> use the `NOVA_*` / `.nova/` / `/nova/` / `nova:*` / `nova-*` names
> exclusively. If you have not yet migrated, follow the steps below.
> **Nova** is the new product brand for the platform formerly known as
> **ACDL** (Agentic Cloud Delivery Platform). This guide documents the
> breaking changes from the rebrand rollout (Phases P2P4, cutoff P5)
> and tells you exactly what to do.
## What is NOT changing
- **The Gitea repository name** (`continuous-intelligence/acdl`) is **not**
changing. Only the product brand is changing. The `uses:` reference
(`acdl/.github/workflows/deploy.yml@vX.Y`) and the GitHub `acdl/acdl` repo
path are unchanged for the duration of the rebrand; the workflow
`uses:` reference will be migrated in a later, separately-announced step.
- **The platform behavior** is unchanged. Same pipeline stages, same
contract schema, same confidence model, same evidence stream, same
modules. Only the brand, the on-disk path, the env var names, the SSM
path, the AWS tag keys, and the AWS resource names are changing.
## The 5 breaking changes
Five things that consumers may reference are being renamed. Each is
scheduled into a phase, ships with a grace period, and has a cutoff.
### 1. Consumer contract path — Phase P2
- **Old:** `.acdl/contract.yml`
- **New:** `.nova/contract.yml`
- **Phase:** P2 (env vars + consumer path)
- **Grace period:** during P2P4 the deploy workflow reads **both** paths
(`.nova/contract.yml` first, falling back to `.acdl/contract.yml` if the
new path is absent). Your existing contracts keep working until P5.
- **Cutoff:** P5 removes the `.acdl/` fallback. Move your contract file
before P5.
- **What you must do:** rename the directory in your consumer repo from
`.acdl/` to `.nova/` and update any `contract:` workflow input that
points at the old path. Nothing else changes in the contract content.
### 2. Environment variables — Phase P2
- **Old:** `ACDL_*` (e.g. `ACDL_LIFECYCLE_MODE`, `ACDL_AWS_ACCOUNT_ID`,
`ACDL_BOOTSTRAP_AWS_ACCESS_KEY_ID`, …)
- **New:** `NOVA_*` (e.g. `NOVA_LIFECYCLE_MODE`, `NOVA_AWS_ACCOUNT_ID`,
`NOVA_BOOTSTRAP_AWS_ACCESS_KEY_ID`, …)
- **Phase:** P2 (env vars + consumer path)
- **Grace period — dual-read fallback:** during P2P4 the platform reads
**`NOVA_*` first, then falls back to `ACDL_*`** if the Nova variable is
unset. This means your CI secrets, workflow env blocks, and local
`.env.secrets` keep working unchanged through P4. You do not need to
rename everything in one shot — rename a variable and the dual-read picks
it up; leave one old and it still resolves.
- **Cutoff:** P5 removes the `ACDL_*` fallback. After P5, only `NOVA_*`
is read.
- **What you must do:** rename your `ACDL_*` CI secrets, workflow `env:`
blocks, and any local `.env.secrets` entries to `NOVA_*`. Because of the
dual-read, you can do this incrementally across P2P4 — but it must be
complete before P5.
### 3. SSM parameter path — Phase P3 (DONE)
- **Old:** `/acdl/{env}/{contractId}/{output}`
- **New:** `/nova/{env}/{contractId}/{output}`
- **Phase:** P3 (SSM paths + tag keys) — **shipped in P3**
- **Grace period — parallel-write:** during P3P4 the platform **writes
every output to both** the `/acdl/…` and `/nova/…` SSM paths, and reads
from `/nova/…` first (falling back to `/acdl/…`). Any hardcoded SSM path
reads in your application code keep resolving through P4. The P3
migration script (`scripts/migrate_ssm_paths.py`) copies existing
`/acdl/…` parameters to `/nova/…`, verifies the copy, and deletes the
old ones.
- **Cutoff:** P5 stops writing to `/acdl/…` and removes the read fallback.
After P5 only `/nova/…` exists.
- **What you must do:** if your application code or runbooks read deploy
outputs from SSM by hardcoded path, update the path prefix from `/acdl/`
to `/nova/`. If you consume outputs only via the PR-comment / GitHub
issue surface, you do nothing — the platform republishes under the new
path automatically.
### 4. AWS tag keys — Phase P3 (DONE)
- **Old:** `acdl:owner`, `acdl:environment`, `acdl:contract`,
`acdl:cost-center`, `acdl:ref`
- **New:** `nova:owner`, `nova:environment`, `nova:contract`,
`nova:cost-center`, `nova:ref`
- **Phase:** P3 (SSM paths + tag keys) — **shipped in P3**
- **Grace period — parallel-tag period:** during P3P4 the platform
**tags every resource with both** the `acdl:*` and `nova:*` keys (same
values). The ABAC session policy matches on **either** key set, so your
existing scoped permissions keep working. The default cost-center value
moves from `acdl-default` to `nova-default` (both written during the
parallel-tag period). Terraform now emits `nova:*` keys; old `acdl:*`
tags on pre-P3 live resources are removed by the P4 runbook's
`scripts/untag_acdl_keys.py` step after the `nova:*` tags are applied
live.
- **Cutoff:** P5 stops writing the `acdl:*` keys and the ABAC policy matches
only on `nova:*`. After P5, resources created before P5 still carry the
old `acdl:*` tags (tags are not retroactively rewritten) but **new**
resources are tagged `nova:*` only, and the policy no longer grants
access via `acdl:*`.
- **What you must do:** if you have IAM policies, Cost Explorer filters,
or billing groupings that key off `acdl:*` tag keys, add a parallel
`nova:*` condition (or migrate to `nova:*`) before P5. The platform
handles the dual-tagging; you only need to update your own tag-key
references.
### 5. AWS resource names — Phase P4
- **Old:** `acdl-*` (DynamoDB tables `acdl-contracts`,
`acdl-change-requests`; Lambda `acdl-contract-ingestor`; SNS
`acdl-sod-halt`; security group `acdl-ecs-sg`; KMS alias
`alias/acdl-platform`; ECS services, ECR repos, IAM user
`acdl-spike-runner`, state bucket `acdl-tfstate-*`, ALB `acdl-alb`,
`acdl-deploy-*`)
- **New:** `nova-*` (the same resources, prefixed `nova-`)
- **Phase:** P4 (resource names) — **maintenance window**
- **Grace period:** P4 is a **planned maintenance window**. AWS resources
cannot be renamed in place, so P4 provisions the `nova-*` resources,
migrates data (DynamoDB tables, S3 state), repoints the platform, and
tears down the `acdl-*` resources. The platform team schedules and
announces the window; consumers do not provision or rename anything
themselves.
- **Cutoff:** the `acdl-*` resources are decommissioned at the end of the
P4 maintenance window. After P4, only `nova-*` resources exist.
- **What you must do:** nothing for the resource names themselves — the
platform owns the rename. If your application code or runbooks reference
a specific `acdl-*` resource by name (e.g. a hardcoded DynamoDB table
name or ECR URI), update it to the `nova-*` name during P4. The platform
publishes the exact old → new name mapping with the P4 announcement.
## Timeline at a glance
| Phase | What ships | Grace period | Cutoff |
|-------|------------|--------------|--------|
| **P1** (this phase) | Brand prose, docs, decks, schema `$id`, release titles | n/a (prose only) | n/a |
| **P2** | `.nova/` contract path + `NOVA_*` env vars | dual-read: `.nova/``.acdl/`, `NOVA_*``ACDL_*` | **P5** removes fallback |
| **P3** | `/nova/` SSM path + `nova:*` tag keys | parallel-write (SSM) + parallel-tag (ABAC matches either) | **P5** removes old path/tags |
| **P4** | `nova-*` AWS resource names | maintenance window (platform-owned migration) | end of P4 window |
| **P5** | Fallback removal | — | `ACDL_*` env vars, `.acdl/` path, `/acdl/` SSM, `acdl:*` tags stop working |
## What consumers must do (checklist)
1. **Before P5 — contract path:** move `.acdl/contract.yml`
`.nova/contract.yml` in your consumer repo; update the `contract:`
workflow input. *(Can be done any time in P2P4.)*
2. **Before P5 — env vars:** rename `ACDL_*` CI secrets / workflow `env:`
blocks / local `.env.secrets` to `NOVA_*`. *(Incremental during P2P4;
dual-read keeps you green.)*
3. **Before P5 — SSM reads:** if you read deploy outputs from SSM by
hardcoded `/acdl/…` path, update to `/nova/…`. *(Skip if you consume
outputs via PR comments only.)*
4. **Before P5 — tag-key references:** if you have IAM policies, Cost
Explorer filters, or billing groupings keyed off `acdl:*`, add or
migrate to `nova:*`. *(Platform handles dual-tagging.)*
5. **During P4 — resource-name references:** if your code or runbooks
reference a specific `acdl-*` AWS resource by name, update to the
`nova-*` name per the P4 mapping announcement. *(Platform owns the
rename itself.)*
## Questions
If anything in this guide is unclear, or you are unsure whether your
consumer repo references a renamed value, open an issue on the platform
repo. The platform team will confirm what you need to change and when.
> **Note:** the real Gitea repository name (`continuous-intelligence/acdl`)
> is **not** changing — only the product brand. The `uses:` workflow
> reference and repo path are migrated in a separately-announced later step;
> until then, keep your `uses: acdl/.github/workflows/deploy.yml@vX.Y`
> reference as-is.
-67
View File
@@ -1,67 +0,0 @@
# Nova — The No-Humans Infrastructure Platform: Thesis Defensibility Brief
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (REQ-213)
> Generated: 2026-08-04
## The thesis
Nova is the autonomous infrastructure layer that lets product teams
ship without engaging an operator, and lets executives trust the AI
not because it never fails but because every decision is captured,
scored, and accountable.
**Autonomy in operations; human at stage gates.** The operator is
removed from the loop of normal operations. Human attestation remains
required at stage gates — QA signs off for production, SRE greenlights
based on operational readiness. The absence of an operator is never
the absence of a record.
## Grounded proof (measurable today)
| Proof | Source | Status |
|-------|--------|--------|
| 18 capabilities verified, 4 honestly skipped (0 broken) | `REGRESSION_REPORT.json` | grounded |
| Decision Ledger captures 100% of AI decisions with outcome backfill | `metrics/decision_ledger.db` | grounded (this milestone) |
| Attestation Coverage: 100% of prod/dr promotions attested by a human | `hitl_gates.py` + outbox `approver_*` | grounded |
| Confidence-gated policy engine (not an LLM) — 6 weighted inputs, band outcome | `confidence_signal.py` | grounded |
| 8-concern attestation matrix with separation-of-duties on prod | `attestation_matrix.py` + `separation_of_duties.py` | grounded |
| Pre-apply cost estimates (Infracost, offline) | `infracost_adapter.py` | grounded |
| Test suite passes (~656 tests) | `metrics/test-results.xml` | grounded |
## Deferred proof (measurable when blocking decisions lift)
| Proof | Blocking Decision | Unblock Requirement |
|-------|-------------------|---------------------|
| Touchless Resolution Rate ≥99% across production estates | 0 consumers today | Pilot estate activation |
| Live infrastructure health (ECS, ALB, RPS) | D-096 | Live AWS re-provisioning |
| Onboarding funnel: requested → granted | D-113/D-114/D-119 | Auto-grant implementation |
| Drift auto-reversal rate ≥95% | D-096 + no scheduler | Drift detection scheduler |
| Predictive vs reactive ratio ≥3:1 | future emitter | ML anomaly-forecasting service |
| Tamper-evident ledger checkpoints (S3 Object Lock + JWS) | D-083 | Audit ledger build-out |
## Anti-claims (what Nova is NOT)
1. **Nova's "AI" is NOT an LLM planner.** It is a confidence-gated
policy engine (confidence_signal + HITL gate). The Decision Ledger
captures this real decision path — not a fabricated "AI agent" that
doesn't exist yet (D-122). When an LLM planner is added, it will emit
richer `alternatives_considered` without schema breakage.
2. **Nova does NOT remove humans from accountability.** Only from
operations. Every stage-gate promotion (qa/prod/dr) requires a human
attestation recorded with approver identity, separation-of-duties
check, and the 8-concern evidence matrix (NORTH_STAR Anti-Goal #3).
3. **Nova is NOT for legacy, untagged, or freeform infrastructure.** It
requires Terraform-managed, policy-aligned, fully-tagged inputs
(NORTH_STAR Anti-Goal #4).
4. **Nova does NOT fabricate metrics.** Every metric is grounded (cites
a source file), derived (documented formula), or deferred (cites a
blocking decision ID). No fabricated numbers in any deck slide or
METRICS.md entry (the "no fabrication" hard constraint).
## What "won" looks like
By month 18, Nova is the layer enterprise leadership points to when
they say *"we don't have an infrastructure ops team anymore, and the
audit trail is stronger than it ever was"* — and it is the default
substrate their AI engineering teams reach for first when an agent needs
to deploy.
+3 -3
View File
@@ -1,6 +1,6 @@
# Nova Onboarding — No-Humans Request Path (v1.16, REQ-182..184)
# Nova Onboarding — Autonomous Request Path (v1.16, REQ-182..184)
The v1.16 milestone implements the **request path** of the no-humans
The v1.16 milestone implements the **request path** of the autonomous
onboarding flow (D-113). A consumer can submit an onboarding request
without contacting the platform team; the platform generates an
environment binding + (in a future milestone) provisions the AWS resources.
@@ -77,7 +77,7 @@ milestone (D-113).
only (D-114); live apply is deferred.
- **OIDC trust policy** — the onboarding Terraform uses a placeholder
OIDC provider; real OIDC federation is blocked on
go-gitea/gitea#36988 (carries forward from v1.1).
upstream forge OIDC support (carries forward from v1.1).
## See also
+1 -1
View File
@@ -230,7 +230,7 @@ change to the modules/stack/confidence/audit.
- A MAJOR bump requires a new registry entry (immutable publication); the
old entry enters a 12-month deprecation window.
- The central deploy pipeline is referenced by a floating MAJOR + MINOR tag
(e.g. `@v1.13`); patch fixes flow within the tag, breaking changes land
(e.g. `@v1.19`); patch fixes flow within the tag, breaking changes land
under the next MINOR tag.
See [Versioning](pipeline/versioning) for the consumer-facing details.
+13 -13
View File
@@ -19,7 +19,7 @@ definitions.
```mermaid
flowchart LR
A["your repo<br/>(app code + contracts + CI definitions)"] -->|uses: acdl/.github/workflows/deploy.yml@v1.13| B
A["your repo<br/>(app code + contracts + CI definitions)"] -->|uses: nova/.github/workflows/deploy.yml@v1.19| B
B["platform runners<br/>(modules + pipelines + adapters + schemas)"] -->|contract -&gt; resolver -&gt; stack -&gt; adapter<br/>-&gt; security checks -&gt; infrastructure plan -&gt; policy checks<br/>-&gt; confidence -&gt; apply -&gt; evidence event| C
C["your resources in AWS"]
```
@@ -27,13 +27,13 @@ flowchart LR
## Versioning the `uses:` reference
The central deployment pipeline is **always versioned with floating MAJOR
and MINOR tags** (e.g. `acdl/pipelines/contract.yml@v1.13`). Version
and MINOR tags** (e.g. `nova/pipelines/contract.yml@v1.19`). Version
constraints cannot be expressed inside the contract, so the tag in
`uses:` is the only immutability lever a consumer has. See
[Versioning](pipeline/versioning) for the full rationale.
**Unversioned references are discouraged.** Do not use `@main` or a bare
`acdl/pipelines/contract.yml`.
`nova/pipelines/contract.yml`.
## Prerequisites
@@ -47,7 +47,7 @@ platform-managed. See [Environments](environments/).
environment is bound, your first pipeline run emits a friendly onboarding
prompt. See [Environments](environments/).
- **Authorization to reference the central pipeline.** Onboarding grants
your repo the right to `uses: acdl/.github/workflows/deploy.yml@v1.13`.
your repo the right to `uses: nova/.github/workflows/deploy.yml@v1.19`.
Contact the platform team if you have not been onboarded.
## Step 1 — Create a consumer repo
@@ -94,7 +94,7 @@ Nova deployment workflow with a **versioned tag** (floating MAJOR + MINOR):
```yaml
jobs:
deploy:
uses: acdl/.github/workflows/deploy.yml@v1.13
uses: nova/.github/workflows/deploy.yml@v1.19
with:
contract: .nova/contract.yml
environment: dev
@@ -140,7 +140,7 @@ name: microservice
| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `uses` | string | yes | Reference to the central deployment pipeline, **versioned** with a floating MAJOR+MINOR tag (e.g. `acdl/pipelines/contract.yml@v1.13`). Bare or `@main` references are discouraged. See [Versioning](pipeline/versioning). |
| `uses` | string | yes | Reference to the central deployment pipeline, **versioned** with a floating MAJOR+MINOR tag (e.g. `nova/pipelines/contract.yml@v1.19`). Bare or `@main` references are discouraged. See [Versioning](pipeline/versioning). |
| `module` | string | yes | Module name from the registry — any primitive or module (e.g. `static-assets`, `microservice`, `s3`). See the [module catalog](modules/). |
| `environment` | string | yes | The platform-managed environment to deploy to (e.g. `dev`). See [Environments](environments/). |
| `inputs` | object | yes | Module-specific inputs (see the module's README). |
@@ -177,14 +177,14 @@ on:
branches: [main]
jobs:
deploy:
uses: acdl/.github/workflows/deploy.yml@v1.13
uses: nova/.github/workflows/deploy.yml@v1.19
with:
contract: .nova/contract.yml
```
That is the entire consumer-side workflow. When you push to `main`:
1. The platform runner resolves `uses: acdl/.github/workflows/deploy.yml@v1.13`
1. The platform runner resolves `uses: nova/.github/workflows/deploy.yml@v1.19`
to the reusable workflow **at the pinned tag**.
2. A **platform-provided runner** checks out **your** repo.
3. The runner checks out the **Nova platform repo** into the workspace —
@@ -326,8 +326,8 @@ per-module extension points. Common examples:
| Contract schema | `schemas/contract.schema.json` | JSON Schema for consumer contracts. |
| Stack schema | `schemas/stack.schema.json` | JSON Schema for the resolved stack instance. |
| Module catalog | [modules/](modules/) | All primitives and modules. |
| Sample contract | `contracts/static-assets.yaml` | The reference example contract (uses `@v1.13`). |
| Sample contract | `contracts/microservice.yaml` | The microservice example contract (uses `@v1.13`). |
| Sample contract | `contracts/static-assets.yaml` | The reference example contract (uses `@v1.19`). |
| Sample contract | `contracts/microservice.yaml` | The microservice example contract (uses `@v1.19`). |
| Module examples | `modules/<name>/examples/` | Validated per-module example contracts (`simple.yaml` + `complex.yaml`). |
| Contract resolver | `core/contract_resolver.py` | Resolves contracts to stack instances. |
| Angine adapter | `adapters/terraform/adapter.py` | Compiles stack instances to infrastructure. |
@@ -353,7 +353,7 @@ destruction:
use `mode: decommission` with the `changeRequestId` input:
```yaml
uses: acdl/.github/workflows/deploy.yml@v1.13
uses: nova/.github/workflows/deploy.yml@v1.19
with:
contract: .nova/contract.yml
mode: decommission
@@ -421,7 +421,7 @@ name: static-assets
```
**Shape 2 — single contract + `environment` workflow input:** the
reusable deploy workflow (`acdl/.github/workflows/deploy.yml@v1.13`)
reusable deploy workflow (`nova/.github/workflows/deploy.yml@v1.19`)
declares an `environment` input. When non-empty, it overrides the
contract's `environment` field at load time (before interpolation), so
the same contract can be promoted by passing a different environment:
@@ -436,7 +436,7 @@ on: workflow_dispatch:
required: true
jobs:
deploy-qa:
uses: acdl/.github/workflows/deploy.yml@v1.13
uses: nova/.github/workflows/deploy.yml@v1.19
with:
environment: qa
contract: .nova/contract.yml
-7
View File
@@ -78,10 +78,3 @@ Planned future features (no dates; tracked in the internal roadmap):
- [Consumer Guide](consumer-guide) — start here if you are a consumer.
- [Architecture](architecture) — start here if you are a platform engineer.
- The [README](https://github.com/nova/nova) describes the platform repo.
> **Note:** The product brand is **Nova** (formerly ACDL — Agentic Cloud
> Delivery Platform). The Gitea repository name (`continuous-intelligence/acdl`)
> and the GitHub `uses:` reference (`acdl/.github/workflows/deploy.yml@…`)
> are unchanged during the rebrand transition; only the product name is
> changing. See the [Nova migration guide](NOVA_MIGRATION) for the
> scheduled breaking changes.
+1 -1
View File
@@ -39,7 +39,7 @@ It is exposed to consumer repos as a **reusable workflow**:
- `.github/workflows/deploy.yml` — GitHub Actions (production)
A consumer repo invokes the reusable workflow via a **versioned tag**
(floating MAJOR + MINOR, e.g. `acdl/.github/workflows/deploy.yml@v1.13`).
(floating MAJOR + MINOR, e.g. `nova/.github/workflows/deploy.yml@v1.19`).
The workflow checks out the consumer repo, then checks out the Nova platform
repo into the runner workspace, and runs `scripts/run_platform.sh` against
the consumer's contract. The consumer never clones the platform repo or
+2 -2
View File
@@ -26,7 +26,7 @@ tag** in a consumer's CI workflow definition:
```yaml
jobs:
deploy:
uses: acdl/.github/workflows/deploy.yml@v1.13
uses: nova/.github/workflows/deploy.yml@v1.19
with:
contract: .nova/contract.yml
```
@@ -36,7 +36,7 @@ itself — the contract no longer carries a `uses:` field). The CI workflow
`uses:` tag is the only immutability lever a consumer has.
**Unversioned references are discouraged.** Do not use `@main` or a bare
`acdl/.github/workflows/deploy.yml``main` is constantly updated and can
`nova/.github/workflows/deploy.yml``main` is constantly updated and can
cause unexpected failures. Pinning to a MAJOR+MINOR tag means:
- **Immutability** — the pipeline behavior you tested is the behavior you
+74 -172
View File
@@ -13,17 +13,18 @@ every deck and a presenter-ready cue sheet for delivery.
```
Step 1: full markdown Step 2: Marp deck Step 3: HTML + PPTX Step 4: Talking points
(source of truth) ──► (lean, 10 slides) ──► (rendered) ──► (presenter cues)
(source of truth) ──► (lean, 19 slides) ──► (rendered) ──► (presenter cues)
*.md *-marp.md *.html / *.pptx *-talking-points.md
+ speaker notes + embedded PNG diagrams + 3-6 bullets per slide
+ mermaid code blocks + Marp frontmatter + key takeaway per slide
+ maturity badges + indexed by Marp slide #
+ no speaker notes + content distilled from Step 1
+ no speaker notes + indexed by Marp slide #
+ no maturity badges + content distilled from Step 1
+ no version in footer
```
### Step 1 — Full markdown (source of truth)
**File convention:** `<deck-name>.md` (e.g. `how-the-platform-works.md`).
**File convention:** `<deck-name>.md` (e.g. `nova-autonomous-cloud-delivery.md`).
Write the complete deck as a standard markdown file. This is the **source of
truth** — it contains:
@@ -34,9 +35,9 @@ truth** — it contains:
the "who cares and why," and the honesty caveats.
- Mermaid diagrams as ```` ```mermaid ```` fenced code blocks (these render
on GitHub/Pages but not in Marp — Step 2 converts them to images).
- An honest "shipped vs. planned" framing: every "available today" claim is
grounded in shipped/verified work; every "planned" item is explicitly
marked.
- An honest "shipped vs. deferred" framing: every "available today" claim is
grounded in shipped/verified work; every "deferred" item is explicitly
marked with the blocking work in plain language.
**Why this file is the source of truth:** it is reviewable in any markdown
viewer, diffs cleanly in git, and carries the full reasoning (speaker notes)
@@ -45,13 +46,13 @@ fact is wrong, fix it here and re-run Steps 2 and 3.
### Step 2 — Marp deck synthesis
**File convention:** `<deck-name>-marp.md` (e.g. `how-the-platform-works-marp.md`).
**File convention:** `<deck-name>-marp.md` (e.g. `nova-autonomous-cloud-delivery-marp.md`).
Synthesize the full markdown into a lean Marp deck:
- **Marp frontmatter** at the top: `marp: true`, `theme: default`,
- **Marp frontmatter** at the top: `marp: true`, `theme: nova-sp`,
`paginate: true`, `size: 16x9`, a header/footer, and an inline `style:`
block for fonts, colors, tables, badges.
block for fonts, colors, tables.
- **No speaker notes.** The Marp deck is what the audience sees; the
speaker notes live only in the Step 1 source of truth.
- **Mermaid diagrams → PNG images.** Marp does not render mermaid fenced
@@ -60,8 +61,10 @@ Synthesize the full markdown into a lean Marp deck:
and embed it with `![w:1000](assets/png/<name>.png)`.
- **`<!-- _class: title -->` + `<!-- _paginate: false -->`** on title and
closing slides for the dark-background title style.
- **Maturity badges** using inline spans:
`<span class="badge planned">Planned</span>`
- **No maturity badges.** The deck no longer uses `<span class="badge">`
spans. Deferred items are named in plain language with their blocking
work, not tagged with a badge.
- **No version in the footer.** The footer carries the deck title only.
- **Tighter prose** than Step 1 — strip the speaker-note nuance; keep the
leadership-relevant selling points.
@@ -69,10 +72,8 @@ Synthesize the full markdown into a lean Marp deck:
Both formats are derived from the Marp deck. **HTML is committed to the repo**
(viewable in any browser, self-contained with base64-embedded images). **PPTX
is uploaded to the Gitea release** as a downloadable attachment (binary, not
committed to git).
#### HTML export (committed to repo)
is also committed to the repo** as a first-class binary artifact and is
attached to the phase's release via `scripts/attach_release_asset.py`.
```bash
CHROME_PATH=/root/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome \
@@ -81,148 +82,67 @@ CHROME_PATH=/root/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome \
-o docs/presentations/<deck-name>.html
```
HTML export inlines images as base64 data URIs — no `--allow-local-files`
needed for self-contained output, but it's required when the Marp deck
references local PNG assets. The resulting HTML is a single self-contained
file that renders the full deck with the S&P Global Energy theme.
**Re-render the HTML whenever the Marp source changes.** The HTML files are
committed artifacts, not generated on-the-fly — they must be re-rendered and
re-committed when the Marp deck is updated.
#### PPTX export (uploaded to Gitea release)
```bash
CHROME_PATH=/root/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome \
npx --yes @marp-team/marp-cli@latest --allow-local-files \
docs/presentations/<deck-name>-marp.md \
-o <output-path>.pptx
```
The `--allow-local-files` flag is **required** for PPTX export so the local
PNG diagrams are embedded in the file. PPTX files are not committed to the
repo (binary, no meaningful diffs) — they are uploaded to the Gitea release
as downloadable attachments.
HTML export inlines images as base64 data URIs. PPTX export requires
`--allow-local-files` so the local PNG diagrams are embedded in the file.
The render + commit + attach pipeline is automated by `scripts/render_deck.sh`
and `scripts/render_slides.sh`.
### Step 4 — Talking points (presenter cues)
**File convention:** `<deck-name>-talking-points.md` (e.g.
`how-the-platform-works-talking-points.md`).
`nova-autonomous-cloud-delivery-talking-points.md`).
Distill the source of truth (Step 1) into presenter-ready cues, indexed by
the Marp deck (Step 2) slide structure:
- **One section per Marp slide**`## Slide N — Title`, matching the Marp
deck's 11 main + Appendix TOC + appendix slide structure exactly. The Marp deck
provides the indexing and context (what the audience sees); the source
markdown provides the content (the speaker notes, the detail, the nuance).
deck's 18 main + 1 appendix slide structure exactly.
- **3-6 talking point bullets per slide** — punchy, actionable cues distilled
from the source markdown's speaker notes. NOT the speaker notes verbatim
(those are too long and too contextual). These are prompts: "Land this
point," "Contrast with X," "Be honest about Y."
from the source markdown's speaker notes.
- **Key takeaway per slide** — the one memorable thing the audience should
walk away with from that slide.
- **No content duplication** — the talking points reference the Marp slides
for visual context and the source markdown for full detail. They don't
repeat either; they bridge them.
**Why this file exists:** a presenter needs a cue sheet they can glance at
during delivery — not the full speaker notes (too long), not the Marp slides
(no detail). The talking points file is the middle layer: what to say, in
what order, with what emphasis, per slide.
**When to update:** re-distill the talking points whenever the Marp deck
structure changes (slides added, removed, merged, or re-ordered) or whenever
the source markdown's speaker notes are updated. The talking points are a
*derived artifact* — if a fact is wrong, fix it in the source markdown (Step 1)
and re-distill.
for visual context and the source markdown for full detail.
## Directory layout
```
docs/presentations/
├── README.md ← this file
├── how-the-platform-works.md ← Step 1: full source of truth
├── how-the-platform-works-marp.md ← Step 2: Marp deck (11 main + TOC + 8 appendix = 20)
├── how-the-platform-works.html ← Step 3: rendered HTML (committed)
├── how-the-platform-works-talking-points.md ← Step 4: presenter cues (20 sections)
├── the-developer-experience.md ← Step 1: full source of truth
├── the-developer-experience-marp.md ← Step 2: Marp deck (11 main + TOC + 7 appendix = 19)
├── the-developer-experience.html ← Step 3: rendered HTML (committed)
├── the-developer-experience-talking-points.md ← Step 4: presenter cues (19 sections)
├── nova-autonomous-cloud-delivery.md ← Step 1: full source of truth (18 main slides + speaker notes)
├── nova-autonomous-cloud-delivery-marp.md ← Step 2: Marp deck (18 main + 1 appendix = 19 slides)
├── nova-autonomous-cloud-delivery.html ← Step 3: rendered HTML (committed, S&P-themed)
├── nova-autonomous-cloud-delivery.pptx ← Step 3: rendered PPTX (committed, S&P-themed)
├── nova-autonomous-cloud-delivery-talking-points.md ← Step 4: presenter cues (19 sections)
└── assets/
├── nova-sp-theme.css ← S&P Global Energy Marp theme (all slide chrome)
├── puppeteer-config.json ← no-sandbox config for mmdc
├── mmd/ ← mermaid source files (Step 2 input)
│ ├── sp-theme.json ← S&P Red/Black/White theme (mermaid-cli --configFile)
── platform-works-01-contract-driven.mmd
│ ├── platform-works-02-frictions.mmd
│ ├── platform-works-02-end-to-end-flow.mmd
│ ├── platform-works-03-north-star.mmd
│ ├── platform-works-03-scope-boundary.mmd
│ ├── platform-works-04-confidence-signal.mmd
│ ├── platform-works-05-attestation-flow.mmd
│ ├── platform-works-07-zero-trust.mmd
│ ├── developer-experience-01b-scope-boundary.mmd
│ ├── developer-experience-02-what-dev-does.mmd
│ ├── developer-experience-03-no-cloning.mmd
│ ├── developer-experience-04-promotion-journey.mmd
│ ├── developer-experience-05-catalog.mmd
│ ├── developer-experience-07-decommission.mmd
│ ├── developer-experience-08-semver.mmd
│ ├── platform-architecture.mmd ← shared high-level logical architecture (both decks)
│ └── road-to-north-star.mmd
└── png/ ← rendered PNGs (embedded in Marp)
├── platform-works-01-contract-driven.png
├── platform-works-02-frictions.png
├── platform-works-02-end-to-end-flow.png
├── platform-works-03-north-star.png
├── platform-works-03-scope-boundary.png
├── platform-works-04-confidence-signal.png
├── platform-works-05-attestation-flow.png
├── platform-works-07-zero-trust.png
├── developer-experience-01b-scope-boundary.png
├── developer-experience-02-what-dev-does.png
├── developer-experience-03-no-cloning.png
├── developer-experience-04-promotion-journey.png
├── developer-experience-05-catalog.png
├── developer-experience-07-decommission.png
├── developer-experience-08-semver.png
├── platform-architecture.png ← shared high-level logical architecture (both decks)
└── road-to-north-star.png
── ... (per-slide .mmd files)
└── png/ ← rendered mermaid PNGs (committed, S&P-themed)
```
## Conventions
### Appendix structure
Each Marp deck has **11 main slides + an Appendix TOC + appendix slides**. The
main 11 are the presentation; the appendix is for deep dives and Q&A backup.
The platform-works deck has 8 appendix slides (A1A8); the developer-experience
deck has 7 appendix slides (A1A7). Both include an Appendix TOC slide.
Each Marp deck has **18 main slides + 1 appendix slide**. The main 18 are the
presentation; the appendix is for Q&A backup.
- **Main slides** (1-11): the story arc, high-impact, minimal text,
visual-heavy. These are what the audience sees during the talk.
- **Appendix slides** (TOC + A1..An): detail-heavy slides moved out of the
main 10 to preserve the narrative flow. The appendix starts with a TOC
slide listing the contents, followed by detail slides and a glossary.
- **The Road to the North Star** is a required appendix slide in both decks
— a phased timeline from v1.0 demo to the North Star, annotated as
"proposed phasing, not formally planned."
- **The Glossary** is a required appendix slide in both decks — defines
acronyms (OIDC, ABAC, CMK, CMDB, RPO, HITL, VCS, NFR) for the audience.
- **Main slides** (1-18): the story arc — Problem → Solution → Proof →
Roadmap + Ask. These are what the audience sees during the talk.
- **Appendix slide** (A1): the Metrics Glossary — detail-heavy reference for
Q&A.
### Maturity framing
### Honesty framing
Every capability claim in a deck is tagged with a `Planned` badge when the item is on the roadmap but not yet implemented:
| Badge | Meaning |
|---|---|
| `Planned` | On the roadmap, not yet implemented |
This is non-negotiable for a leadership audience: never present a roadmap
item as a current capability, and never bury a tested capability's
availability. When in doubt, check `.ciagent/ROADMAP.md` and the milestone
status in `.ciagent/PROJECT.md`.
Every capability claim in the deck is grounded, derived, or honestly
deferred with its blocking work named in plain language. Internal provenance
(decision IDs, requirement IDs, internal file paths) is kept out of the
audience-facing slides — those live in the `.ciagent/` files only. When in
doubt, check `.ciagent/ROADMAP.md` and the milestone status in
`.ciagent/PROJECT.md`.
### Audience
@@ -234,8 +154,10 @@ Head of Infrastructure, Head of DevOps. The framing rules:
"composition."
- **Selling points forward.** Each slide leads with the leadership-relevant
outcome; the mechanism follows.
- **Zero-trust, security, observability, auditability, DX, citizen
developer** are the themes — not implementation details.
- **Security, remediation velocity, reliability, lead time, observability,
citizen developer** are the themes — not implementation details.
- **"Infrastructure operations become visible"** is the recurring theme across
the deck.
### Diagrams
@@ -245,8 +167,7 @@ style (renders on GitHub/Pages). For the Marp deck (Step 2):
1. Extract the mermaid block into `assets/mmd/<deck>-<slide>-<name>.mmd`.
2. Use **horizontal layouts** (`flowchart LR`) or **subgraph row-wrapping**
for wide diagrams so the PNG fits a 16:9 slide without shrinking to
illegibility. A 9-node sequential `flowchart TD` renders as a tall thin
strip — restructure it as 2-row subgraphs or `flowchart LR`.
illegibility.
3. Render with a 2x scale factor and transparent background for crisp slides.
4. Embed with `![w:1000](assets/png/<name>.png)` (or `h:320` for tall images).
@@ -274,42 +195,16 @@ for f in mmd/*.mmd; do
done
```
The `puppeteer-config.json` passes `--no-sandbox` to the headless browser
(required when running as root in this environment). The `--configFile
mmd/sp-theme.json` applies the S&P Global Red/Black/White theme (dark
`#1B1B1B` accent nodes with `#D6002A` red borders, white supporting nodes,
`#F0F0F0` subgraph backgrounds). Each `.mmd` file also carries the same
theme inline via a `%%{init:...}%%` block so it renders correctly even
without the `--configFile` flag.
### Export a Marp deck to HTML (committed to repo)
### Render a Marp deck to HTML + PPTX (committed artifacts)
```bash
CHROME_PATH=/root/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome \
npx --yes @marp-team/marp-cli@latest --allow-local-files \
docs/presentations/<deck-name>-marp.md \
-o docs/presentations/<deck-name>.html
bash scripts/render_slides.sh nova-autonomous-cloud-delivery
```
HTML export inlines images as base64 data URIs. The `--allow-local-files`
flag is needed when the Marp deck references local PNG assets (like the
diagram images in `assets/png/`). The resulting HTML is self-contained.
**The HTML files are committed artifacts** — re-render and re-commit whenever
the Marp source changes.
### Export a Marp deck to PPTX (uploaded to Gitea release)
```bash
CHROME_PATH=/root/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome \
npx --yes @marp-team/marp-cli@latest --allow-local-files \
docs/presentations/<deck-name>-marp.md \
-o <output-path>.pptx
```
`--allow-local-files` is **required** for PPTX so local PNG diagrams are
embedded in the file. PPTX files are not committed to git — upload them as
attachments to the Gitea release.
This renders all mermaid PNGs, the HTML, and the PPTX, and stages them for
commit. The `--allow-local-files` flag is required so local PNG diagrams are
embedded. Both HTML and PPTX are committed to the repo; the PPTX is also
attached to the phase's release.
## Adding a new presentation
@@ -319,16 +214,14 @@ attachments to the Gitea release.
2. **Extract any mermaid diagrams** into `assets/mmd/<deck-name>-<slide>-<name>.mmd`
and render them to `assets/png/` (command above).
3. **Synthesize the Marp deck** as `<deck-name>-marp.md` with frontmatter,
no speaker notes, embedded PNGs, and maturity badges.
4. **Render to HTML** with `--allow-local-files` and commit the HTML to
`docs/presentations/<deck-name>.html`.
5. **Render to PPTX** with `--allow-local-files` and upload to the Gitea
release (do not commit PPTX to git).
6. **Distill the talking points** as `<deck-name>-talking-points.md` — one
no speaker notes, embedded PNGs, and no badges.
4. **Render to HTML + PPTX** via `scripts/render_slides.sh <deck-name>` and
commit both to `docs/presentations/`.
5. **Distill the talking points** as `<deck-name>-talking-points.md` — one
section per Marp slide, 3-6 talking point bullets + key takeaway, content
distilled from the source markdown (Step 1), indexed by the Marp deck
(Step 2) slide structure.
7. **Verify** the PPTX slide count and that media files are embedded:
6. **Verify** the PPTX slide count and that media files are embedded:
```bash
python3 -c "
import zipfile, re
@@ -341,7 +234,16 @@ attachments to the Gitea release.
## Current decks
| Deck | Source of truth (Step 1) | Marp deck (Step 2) | Rendered HTML (Step 3) | Talking points (Step 4) | Slides | Audience |
| Deck | Source of truth (Step 1) | Marp deck (Step 2) | Rendered HTML + PPTX (Step 3) | Talking points (Step 4) | Slides | Audience |
|---|---|---|---|---|---|---|
| How the Platform Works | `how-the-platform-works.md` | `how-the-platform-works-marp.md` | `how-the-platform-works.html` | `how-the-platform-works-talking-points.md` | 11 main + TOC + 8 appendix (20) | CTO, Head of Cloud, Head of Infra, Head of DevOps |
| The Developer Experience | `the-developer-experience.md` | `the-developer-experience-marp.md` | `the-developer-experience.html` | `the-developer-experience-talking-points.md` | 11 main + TOC + 7 appendix (19) | CTO, Head of Cloud, Head of Infra, Head of DevOps |
| Nova — The Autonomous Cloud Delivery Platform | `nova-autonomous-cloud-delivery.md` | `nova-autonomous-cloud-delivery-marp.md` | `nova-autonomous-cloud-delivery.html` + `.pptx` (committed + release-attached) | `nova-autonomous-cloud-delivery-talking-points.md` | 18 main + 1 appendix (19) | CTO, Head of Cloud, Head of Infra, Head of DevOps |
> **v1.21:** the deck was renamed from "No-Humans Infrastructure Platform"
> to "Autonomous Cloud Delivery Platform" (professional framing; conveys
> autonomy without the provocative wording). The narrative restructured to
> a 4-beat arc (Problem → Solution → Proof → Roadmap + Ask). Internal
> provenance (decision IDs, requirement IDs, file paths) removed from
> audience-facing slides. Maturity badges removed. The RACI matrix expanded
> to four roles (Quality Engineering + SRE). The Atelier slide split into
> two. The pipeline hardened: Checkov on static code before the plan;
> Wiz-or-Checkov on the plan (never both).
@@ -0,0 +1,81 @@
/* @theme nova-sp */
/* Nova S&P Global Energy theme for Marp decks.
*
* Palette: S&P Red (#D6002A), Black (#1B1B1B), White (#FFFFFF), Grey (#F0F0F0).
* Font: Akkurat Pro (fallback Helvetica Neue / Arial).
*
* This theme extends Marp's default and applies the S&P palette to ALL slide
* chrome backgrounds, headers/footers, pagination, tables, blockquotes,
* code blocks not just headings.
*/
:root {
--sp-red: #D6002A;
--sp-black: #1B1B1B;
--sp-white: #FFFFFF;
--sp-grey: #F0F0F0;
--sp-dark-grey: #2E2E2E;
}
/* Base section */
section {
font-family: "Akkurat Pro", "Helvetica Neue", "Arial", sans-serif;
font-size: 22px;
color: var(--sp-black);
background: var(--sp-white);
}
/* Headings — S&P Red */
h1 { color: var(--sp-red); font-size: 34px; margin-bottom: 0.3em; }
h2 { color: var(--sp-red); font-size: 26px; margin-bottom: 0.2em; }
h3 { color: var(--sp-red); font-size: 22px; margin-bottom: 0.2em; }
h4 { color: var(--sp-dark-grey); font-size: 20px; margin-bottom: 0.15em; }
/* Title slides — black background, red top border */
section.title {
background: var(--sp-black);
color: var(--sp-white);
border-top: 8px solid var(--sp-red);
}
section.title h1 { color: var(--sp-white); }
section.title h2 { color: var(--sp-white); }
/* Tables — grey header with red underline, explicit white body for readability on any background */
table { font-size: 18px; width: 100%; border-collapse: collapse; background: var(--sp-white); }
th { background: var(--sp-grey); border-bottom: 2px solid var(--sp-red); padding: 6px 10px; text-align: left; }
td { background: var(--sp-white); color: var(--sp-black); border-bottom: 1px solid var(--sp-grey); padding: 6px 10px; }
/* Ensure tables on dark/title slides remain readable: white card with a subtle border */
section.title table, section table { background: var(--sp-white); }
section.title td, section td { background: var(--sp-white); color: var(--sp-black); }
section.title th, section th { background: var(--sp-grey); color: var(--sp-black); }
/* Blockquotes — red left border */
blockquote { border-left: 4px solid var(--sp-red); color: var(--sp-dark-grey); font-size: 20px; padding-left: 12px; }
/* Code — dark background */
pre { background: var(--sp-black); color: var(--sp-white); border-radius: 4px; padding: 12px; font-size: 16px; }
code { background: var(--sp-grey); color: var(--sp-black); border-radius: 2px; padding: 1px 4px; font-size: 18px; }
pre code { background: transparent; color: inherit; }
/* Images — centered, max height */
img { display: block; margin: 0 auto; max-height: 320px; }
/* Header/footer — subtle grey */
header { color: var(--sp-dark-grey); border-bottom: 1px solid var(--sp-grey); }
footer { color: var(--sp-dark-grey); border-top: 1px solid var(--sp-grey); }
/* Maturity badges */
.badge { display: inline-block; padding: 2px 8px; border-radius: 4px; font-size: 14px; font-weight: 600; }
.badge.today { background: #c6f6d5; color: #22543d; }
.badge.planned { background: #fef3c7; color: #78350f; }
/* Pagination — S&P Red progress bar */
.bespoke-progress-parent { background: var(--sp-grey); }
.bespoke-progress-bar { background: var(--sp-red) !important; }
/* Lists — tighter */
ul { margin-top: 0.3em; }
li { margin-bottom: 0.2em; }
/* Strong — S&P Red for emphasis in lead lines */
strong { color: var(--sp-red); }
@@ -0,0 +1,327 @@
---
marp: true
theme: nova-sp
paginate: true
size: 16x9
header: 'Nova — The Autonomous Cloud Delivery Platform'
footer: 'Nova — The Autonomous Cloud Delivery Platform'
---
<!-- _class: title -->
<!-- _paginate: false -->
# Nova — The Autonomous Cloud Delivery Platform
**Shifting from Operational Overhead to Strategic Value**
Product Development & Citizen Developer Overview
---
## Slide 1 — The Problem
**Product teams now own their cloud infrastructure — but ownership without discipline is destroying value.**
- **No lifecycle planning.** Resources are authored for creation, not for patching, decommissioning, or rollback — so changes are destructive.
- **Proactive scanning is not part of authoring.** AI-frontier models exploit zero-days at a rapid pace; teams cannot keep up by reacting. Modules must be scanned as code and at runtime — and remediated at the pace the threat moves.
- **Bandwidth gaps in infrastructure operations.** Time spent on remediation + the push for innovation leaves operations chronically under-resourced; detections are missed, incidents grow.
- **Tribal knowledge and the rockstar-operator problem.** Operations depend on a handful of administrators; when they leave, the knowledge leaves with them. The platform should encode the discipline, not the person.
Every hour a developer spends writing, deploying, fixing, or remediating infrastructure is an hour not spent releasing features to production.
**Benefit:** the answer is an autonomous cloud delivery platform that encodes discipline as policy, scans proactively, remediates rapidly, and makes operations visible to leadership rather than hidden in tribal knowledge.
---
## Slide 2 — Nova's Vision
> **Infrastructure operations become visible. Every environment provisioned, every incident healed, every risk remediated — by an autonomous system whose trustworthiness is provable, not promised. Human attestation remains required at stage gates; the operator is never in the loop of normal operations.**
- **Visibility is the recurring theme** — security posture, remediation velocity, reliability, and lead time as queryable signals
- **Provable, not promised** — trust established by deterministic scripts that calculate a score; the platform functions without AI
- **Autonomy in operations, human at stage gates** — QA signs off for production; SRE greenlights operational readiness
**Benefit:** the destination is autonomous operations with provable trust — security, remediation velocity, reliability, and lead time made visible to leadership, not promised to them.
---
## Slide 3 — Strategic Objectives + Anti-Goals
**4 Strategic Objectives:**
1. **Zero-touch operations** — autonomy as the default, not the demo; stage-gate attestation (QA, SRE) remains human by design
2. **Provable trust in automated decisions** — deterministic scripts calculate a score; the platform functions without AI; Decision Ledger, confidence scoring, circuit breakers, blast-radius controls
3. **Compounding, quantifiable ROI** — four CTO-grade metrics, all flowing into PowerBI:
- **Lead Time** (PR → Production) · **Infrastructure Vulnerability Count** (trend) · **MTTR** · **Cloud Spend Reduction**
4. **Integrate with externally owned development platforms — regardless of source** — PDLC, SDLC, Agentic, or Citizen Developer; Nova provides skills + MCP endpoints; all prod intents go through the same controls and quality gates
**4 Anti-Goals (what Nova is NOT):**
1. Not a general-purpose AI agent platform
2. Not a system that removes humans from accountability — only from normal operations
3. Not an upstream development platform (no product backlogs, IDE, code authorship)
4. Not a replacement for the Product Development Lifecycle (PDLC)
**Benefit:** the scope is explicit — Nova governs infrastructure and delivery, integrates with any upstream source through one validated contract, and measures success on four metrics a CTO can repeat back.
---
## Slide 4 — Scope: Downstream of PDLC
**Nova governs infrastructure and delivery. The PDLC is upstream — Nova never penetrates it. Integration is through one validated contract.**
- **The PDLC is upstream:** product backlog, code authorship (AI agent, IDE, agentic SDLC), sprint planning, application business logic
- **Nova is downstream:** contract ingestion → submission-readiness gate → policy enforcement → cloud resource lifecycle → environment progression (dev → qa → prod → dr) → immutable audit + attestation
- **The integration point is one contract** — any upstream source (AI agent, agentic SDLC, dev platform) produces submissions subject to the same compliance standards
- **Nova validates the submission, not the author** — the audit trail, the policy envelope, and the evidence stream are the same regardless of source
**Benefit:** a clean scope boundary — Nova is purpose-built for infrastructure operations and integrates with any upstream source through one validated contract, so the platform team's surface area stays bounded.
---
## Slide 5 — RACI: Who Owns What
**Four roles, one matrix — citizen developer owns FRs + UAT, platform owns NFRs + infra, quality engineering owns the gate evidence, SRE owns operational readiness.**
| Work Category | Citizen Dev | Platform | Quality Eng | SRE |
|---|---|---|---|---|
| Functional Requirements | **R/A** | C | I | I |
| User Acceptance Testing | **R/A** | C | I | I |
| Non-Functional Requirements | I | **R/A** | C | C |
| Infrastructure (cloud, state, IAM) | I | **R/A** | I | C |
| QA (policy, confidence, schema) | C | R | **R/A** | I |
| Production deployment to cloud | I | **R/A** | C | C |
| Quality attestation (QA sign-off) | **A** | R | **R** | I |
| Production readiness (SRE sign-off) | **A** | R | C | **R** |
**R**=Responsible · **A**=Accountable (sign-off) · **C**=Consulted · **I**=Informed. Production readiness is co-owned: the platform runs attestations agentically; the citizen developer authorizes the promotion at the stage gate.
**Benefit:** every party knows what they bring, what the platform provides, what quality engineering guards, and where SRE signs off — accountability is explicit, never diffuse.
---
## Slide 6 — The Platform Pipeline
**How intent becomes verified infrastructure — fail-fast policy scanning before the plan, runtime scanning after it.**
![w:1000](assets/png/platform-pipeline.png)
- **Contract → resolver → adapter → Checkov on static code (before plan) → terraform plan → Wiz on the plan → confidence signal → stage gate → apply → evidence + ledger**
- **Fail-fast, quick feedback** — Checkov runs on the authored Terraform code before `terraform plan` so developers get immediate policy feedback
- **Wiz on the plan when configured; Checkov as a drop-in otherwise** — Wiz scans the plan output; when Wiz credentials are absent, Checkov runs against the plan. **Wiz and Checkov are never both run on the plan.**
- **Dev is autonomous** (no stage gate); **qa/prod/dr require human attestation** (QA for quality, SRE for production readiness)
**Benefit:** two layers of scanning, zero operator involvement in normal operations — fast deterministic feedback at authoring time and a runtime scan on the resolved plan.
---
## Slide 7 — The Decision Ledger
**Every automated decision is captured, immutable, queryable — and accountable.**
- **What is captured:** the chosen action, the confidence score, the alternatives considered, whether a human overrode it, and the outcome (backfilled once the apply completes). Every stage-gate attestation (QA, SRE) is captured with approver identity and the evidence presented.
- **"AI decisions" are really automated decisions** — made by deterministic scripts that calculate a score and a band; the platform functions without AI. When an LLM planner is added later, it will emit richer alternatives without breaking the schema.
- **The value is accountability, not the storage engine** — the ledger is append-only and tamper-evident; every decision is queryable for auditing, traceable to an outcome, and impossible to rewrite after the fact.
**Benefit:** "autonomous" is defensible because every decision is immutable, queryable, and accountable — and the audience knows exactly what "automated" means here: deterministic scoring, not a black-box LLM.
---
## Slide 8 — The Attestation Matrix
**The designed controls that keep humans at stage gates — structured, freshness-validated, separation-of-duties-enforced.**
| Concern | Env | Freshness | Description |
|---------|-----|-----------|-------------|
| Functional correctness | qa | 24h | The application behaves as specified; evidence accepted from the consumer's UAT. |
| Performance baseline | qa | 7d | The deployment meets its performance envelope vs. the agreed baseline. |
| Security posture | qa | 24h | The deployment's security findings have been reviewed and accepted. |
| Operational readiness | prod | 30d | SRE confirms the deployment is operable: runbooks, dashboards, on-call. |
| Incident response | prod | 90d | The on-call path has been exercised; a working incident-response plan exists. |
| Capacity & cost | prod | 30d | Capacity headroom and monthly cost are within the agreed envelope. |
| Resilience: DR drill | prod | 180d | A DR drill has been run and recovery met the RTO. |
| Resilience: chaos | prod | 90d | A chaos exercise has been run and the deployment absorbed the failure. |
| Resilience: backup | prod | 30d | Backups are restorable and tested within the freshness window. |
| DR region deploy | dr | 180d | The DR region can be deployed and is reachable. |
Separation-of-duties on prod: the approver cannot be the same person who built the deployment.
**Benefit:** the gate model is explicit — autonomy in operations, human in accountability, by design. The matrix is what makes autonomous operations safe enough to trust in production.
---
## Slide 9 — Telemetry & Live Ops
**Every metric in this deck is traceable to a real emitted signal — the live-ops dashboard makes operations visible in PowerBI.**
![w:900](assets/png/telemetry-live-ops.png)
- **Platform components → CloudEvents envelope → event log + decision ledger + run records → collector → cold store → PowerBI views → live ops dashboard**
- **The live ops dashboard (PowerBI)** surfaces the four CTO-grade metrics (Lead Time, Vulnerability Count, MTTR, Cloud Spend) alongside trust metrics (Decision Ledger coverage, Attestation coverage) and efficiency metrics (touchless resolution, escalation frequency)
- **Deliberately minimal** — Nova-native envelopes; no Kafka, no Prometheus, no ClickHouse. The cold store handles batch and historical analysis; the live-ops surface is built in PowerBI on the exported views
- **Every number is traceable to a signal** — when a CFO asks "where does this number come from?", the answer is a query against the cold store, not a Slack thread
**Benefit:** the architecture is the trust substrate — leadership sees the same numbers the platform produces, in PowerBI, with full traceability. Operations become visible.
---
## Slide 10 — Decision Ledger + Attestation Coverage
**By design, no change reaches production without a ledger entry and a human attestation — both queryable for auditing, with full traceability.**
- **Decision Ledger coverage: 100%** — every platform run emits a decision record with outcome backfill; no automated decision is ever lost
- **Attestation coverage: 100%** — every prod/dr promotion is attested by a human (QA for quality, SRE for production readiness), recorded with approver identity, separation-of-duties check, and the evidence matrix
- **No change to production without both** — the ledger entry and the human attestation are mandatory, enforced by the pipeline, not by policy
- **Easily queried for auditing** — queryable by run, by environment, by approver, and by outcome; the audit trail is a query, not a forensic exercise
- **Full traceability** — a production change is traceable from the contract that declared intent, through the policy scan, the confidence score, the attestation, to the applied outcome
**Benefit:** trust is provable — not a marketing claim, a queryable record. An auditor answers "who approved this, when, on what evidence?" in one query; a CTO answers "how many of last quarter's prod changes were touchless?" in one query.
---
## Slide 11 — Cost & ROI
**The ROI formula and the cost estimates — grounded, with the production denominator honestly flagged.**
- **Cost estimates are pre-apply and offline** — the platform reads the terraform plan and estimates cost before anything is applied; a cost regression is caught before the spend happens
- **The ROI formula:**
`Platform ROI = (FTE hours saved × blended rate + cloud savings + avoided downtime) ÷ platform op cost`
- **The four CTO-grade metrics are the ROI proof:** Lead Time (PR → Prod), Infrastructure Vulnerability Count (trend), MTTR, Cloud Spend Reduction — all flow into PowerBI
- **Honest caveat:** derived metrics are computed on internal runs today; the production-denominator activates when a pilot estate runs. The formula is grounded; the production numbers are not yet.
**Benefit:** the ROI is not a black box — the formula is shown, the four metrics are committed, and the production-denominator caveat is stated up front. The CFO sees exactly what is real today and what activates with a pilot.
---
## Slide 12 — What's Deferred — and Why
**Honesty about what is not measured yet — and the blocking work for each.**
To be clear: these deferrals are *measurement infrastructure*, not the autonomy itself. The platform runs without an operator in the loop of normal operations. What is deferred is the evidence pipeline for certain metrics — not the autonomy.
| # | Deferred metric | Blocking work |
|---|-----------------|---------------|
| 1 | Live infrastructure health | Live AWS re-provisioning (currently torn down to zero-cost steady state) |
| 2 | Live outbox write rate | Live AWS re-provisioning |
| 3 | Tamper-evident ledger checkpoints | Audit-ledger build-out (Object Lock + signed checkpoints) |
| 4 | Onboarding funnel (requested → granted) | Auto-grant implementation |
| 5 | Drift auto-reversal | Drift-detection scheduler (not yet built) |
| 6 | Live cost reconciliation | Live AWS re-provisioning + actual-spend feed |
| 7 | SLA / unplanned downtime | Live AWS re-provisioning |
| 8 | Predictive vs reactive ratio | ML anomaly-forecasting service (not yet built) |
**Benefit:** the boundaries are explicit — what Nova measures today, and exactly what blocks the rest. The autonomy is real; the measurement gaps are documented with the work that unblocks each one.
---
## Slide 13 — Roadmap to the North Star
**The path from the grounded metrics to the 1218 month targets — each deferred metric has an unblock path and a timeframe.**
| Timeframe | Work | Unblocks |
|-----------|------|----------|
| Near-term | Live AWS re-provisioning | Live infra health, outbox write rate, live cost reconciliation, SLA |
| Near-term | Auto-grant implementation | Onboarding funnel (requested → granted) |
| Mid-term | Drift-detection scheduler | Drift auto-reversal |
| Mid-term | Audit-ledger build-out (Object Lock + signed checkpoints) | Tamper-evident ledger checkpoints |
| Mid-term | Hot-path activation (batch → near-real-time) | Live-ops dashboard freshness |
| Longer-term | ML anomaly-forecasting service | Predictive vs reactive ratio |
Re-evaluation triggers: each blocking piece of work lifts on its own schedule; the metrics layer evolves as each one lands.
**Benefit:** every deferred metric has an unblock path — nothing is hand-waved; everything has a plan and a timeframe.
---
## Slide 14 — 12-Month Product Roadmap
**The product arc from pilot activation to integration — four quarters, four outcomes.**
| Quarter | Theme | Board-level outcome |
|---------|-------|---------------------|
| **Q1** | Pilot Activation | Nova runs a real customer estate end-to-end, autonomously, with a measurable zero-touch rate. |
| **Q2** | Provable Trust | Every automated decision lands in a tamper-evident ledger; the CFO sees real cloud-spend reconciliation. |
| **Q3** | Compounding ROI | Quarter-over-quarter cloud spend drops; drift is detected and reversed without a human. |
| **Q4** | Integration & Predictive | AI agents deploy through Nova by default; the ML anomaly-forecasting service goes live. |
Grounded in the four strategic objectives (autonomy, provable trust, ROI, integration) and the deferred-metric unblock paths.
**Benefit:** the 12-month product arc — each quarter activates a strategic objective and its corresponding board-level metric, from pilot activation through integration leadership.
---
## Slide 15 — Quarter-by-Quarter Outcomes
| Quarter | Product theme | Key deliverable | Target metric | Grounding |
|---------|---------------|-----------------|---------------|-----------|
| **Q1** | Pilot Activation | Re-provision live AWS; activate first pilot estate; onboarding auto-grant | Touchless ≥ 99% · Escalation < 0.1% · Accuracy ≥ 99.5% | Objective #1 — autonomy as the default |
| **Q2** | Provable Trust | Tamper-evident ledger (Object Lock + signed checkpoints); daily checkpoints; live cost reconciliation | Decision Ledger Coverage 100% · Cost Savings ≥ 25% | Objective #2 — trust is the moat |
| **Q3** | Compounding ROI + Drift | Drift-detection scheduler; auto-reversal; pre-apply → actual-spend reconciliation on the pilot estate | Drift Auto-Reversal ≥ 95% · Spend Reduction ≥ 25% | Objective #3 — CFO-pointable numbers |
| **Q4** | Integration + Predictive | ML anomaly-forecasting; AI-agent intent surface; multi-cloud (Azure/GCP) preview | Predictive:Reactive ≥ 3:1 · AI-Agent Intent Share (first measurement) | Objective #4 — default substrate for agents |
**Month-18 destination:** *"Nova is the layer enterprise leadership points to when they say 'we don't have an infrastructure ops team anymore, and the audit trail is stronger than it ever was.'"*
**Benefit:** each quarter has a concrete deliverable, a target metric grounded in a strategic objective, and a path from "honestly deferred" to "shipped and measured."
---
## Slide 16 — Production-Grade Guidance via Atelier (1/2)
**Nova instructs the citizen developer's AI agent on production-grade engineering — a set of skills and an MCP server.**
- **Skills** — markdown files keyed to production-grade engineering domains (API, security, data, testing, observability, errors, DevOps, infrastructure-as-code, compliance); the skills extend the baseline catalog with Nova-specific production-grade principles
- **MCP server** — a plugin-registry, stdio server exposing four tools: `lookup_principle`, `list_domains`, `matrix_lookup`, `validate_against_principles`. The developer's AI agent (or any agentic SDLC platform) calls these tools to look up the principles that apply to its submission
- **The integration point is the same regardless of source** — whether the submission comes from an AI coding agent, an agentic SDLC platform, or a traditional IDE, the same skills and MCP server apply. This is how Nova makes the citizen developer production-grade without owning the PDLC
**Benefit:** the citizen developer's AI agent is not unguided — Nova provides production-grade engineering principles as skills and as an MCP surface, so submissions arrive at the contract boundary already aligned with the platform's standards.
---
## Slide 17 — Production-Grade Guidance via Atelier (2/2)
**Agentic validation catches engineering-discipline gaps that deterministic scanners miss — and the validation is reproducible.**
- **Beyond deterministic scanners** — Wiz, Checkmarx, and Mend check policy and secrets; they do not check engineering discipline. The Atelier MCP server catches correctness, clarity, and observability gaps that deterministic tools cannot: "is this service observable?", "is this error path handled?", "is this API contract clear?"
- **Agentic validation, not a second policy engine** — the MCP server gives the AI agent the principles to validate against; the agent does the validation. The agent reasons about the submission against the principles, not a second static scan
- **Vendored for audit reproducibility** — Atelier is vendored at a pinned tag. A validation result is replayable against the exact principles that produced it, so an audit can reproduce a validation months later, not just trust a log line
**Benefit:** the citizen developer's submission is checked for engineering discipline, not just policy compliance — and the check is reproducible for audit. That is what makes the submission production-grade, regardless of which upstream platform produced it.
---
## Slide 18 — Recap + Ask
**The 4-beat recap + the business decision.**
**Recap:**
- **Problem:** product teams own infrastructure without the discipline and lifecycle planning it requires; bandwidth gaps and tribal knowledge leave operations exposed
- **Solution:** autonomous cloud delivery — operations become visible, trust is provable (deterministic scoring), humans at stage gates
- **Proof:** 100% ledger coverage, 100% attestation coverage, grounded ROI formula, four CTO-grade metrics flowing into PowerBI
- **Roadmap:** deferred metrics have unblock paths; the 12-month product arc activates one strategic objective per quarter
**The ask:** "Approve a pilot estate to activate the production-denominator metrics (Lead Time, Vulnerability Count, MTTR, Cloud Spend), and approve the tamper-evident ledger build-out to move from the local hash-chain to S3 Object Lock + signed checkpoints. These two decisions move Nova from 'pipeline-ready' to 'production-proven.'"
**Benefit:** a clear business decision — approve a pilot and the ledger build-out — with the confidence that every claim in this deck is grounded, derived, or honestly deferred.
---
<!-- _class: title -->
<!-- _paginate: false -->
## Appendix A1 — Metrics Glossary
| KPI | Definition | Status |
|-----|-----------|--------|
| Touchless Resolution Rate | runs without operational stage-gate block ÷ total | partial (Post-Pilot) |
| Human Escalation Frequency | operational stage-gate blocks ÷ total | partial (Post-Pilot) |
| Automated Decision Accuracy | decisions not followed by failure within 5min | partial (Post-Pilot) |
| MTTR (p95) | apply.failed → successful retry | grounded |
| Confidence-Gate Halt Rate | runs with band=block ÷ total | grounded |
| Provisioning Lead Time | run.completed run.started | grounded |
| Deployment Frequency | count(run.completed) per day | grounded |
| Cost Savings (pre-apply) | sum(delta_usd where delta < 0) | partial (live reconciliation deferred) |
| FTE Hours Saved | run count × manual baseline × rate | derived (N=0 caveat) |
| Platform ROI | (labor + cloud + avoided downtime) ÷ op cost | derived (N=0 caveat) |
| Decision Ledger Coverage | decisions with outcome ÷ total | grounded |
| Attestation Coverage | prod/dr attested ÷ total prod/dr | grounded |
| Policy Compliance Rate | 1 failed_assets ÷ total | grounded |
**Benefit:** a reference for every metric mentioned in the deck.
@@ -0,0 +1,129 @@
# Nova — The Autonomous Cloud Delivery Platform: Talking Points
> Step 4 of the 4-step deck process. Presenter cues distilled from the
> source of truth (`nova-autonomous-cloud-delivery.md`). 3-6 bullets per
> slide + key takeaway. Indexed by Marp slide #.
> v1.21 — REQ-245
---
### Slide 1 — The Problem
- Open with the shift: "you build it, you run it" put Terraform into product teams — ownership without discipline is destroying value
- Land the lifecycle-planning gap: resources authored for creation, not for patching/rollback → destructive changes
- Land the urgency: AI-era 0-day pace demands proactive scanning as code + at runtime, remediated at threat pace
- Call out tribal knowledge / the rockstar-operator problem — the platform should encode the discipline, not the person
- Do NOT frame this as "humans are the problem" — the problem is ownership without the discipline and tooling
- **Key takeaway:** the problem is infrastructure ownership without discipline; the answer is an autonomous platform that encodes the discipline
### Slide 2 — Nova's Vision
- Read the vision verbatim — "infrastructure operations become visible" is the operative phrase
- Emphasize "provable, not promised" — trust established by deterministic scripts; the platform functions without AI
- State the attestation model up front: QA for production, SRE for operational readiness
- **Key takeaway:** autonomous operations with provable trust — security, remediation velocity, reliability, lead time made visible, not promised
### Slide 3 — Strategic Objectives + Anti-Goals
- Objective #2 is the one to land carefully: trust = deterministic scoring, not an LLM; the platform functions without AI
- Objective #3: four CTO-grade metrics (Lead Time, Vuln Count, MTTR, Spend) — all flow into PowerBI
- Objective #4 is the integration thesis: Nova integrates with any upstream source; provides skills + MCP; all prod intents go through the same controls
- Anti-goals #3 and #4 protect the scope: not an upstream dev platform, not a PDLC replacement
- **Key takeaway:** purpose-built for infra ops, integrates with any source through one contract, measures success on four CTO metrics
### Slide 4 — Scope: Downstream of PDLC
- Nova governs infra + delivery only; the PDLC (backlog, code authorship, IDE) is upstream — Nova never penetrates it
- Integration is only through the validated contract boundary
- Any upstream source (AI agent, agentic SDLC, dev platform) produces submissions subject to the same compliance standards
- Nova validates the submission, not the author
- **Key takeaway:** Nova is purpose-built for infrastructure operations; the scope boundary is clean and bounded
### Slide 5 — RACI: Who Owns What
- Four roles now: Citizen Developer, Platform, Quality Engineering, SRE
- Quality attestation is owned by Quality Engineering (not the Platform); Production readiness is owned by SRE
- The Platform runs the checks agentically but is never the Accountable party for the gate — that separation keeps the platform honest
- Production readiness is co-owned: the platform runs attestations; the citizen developer authorizes the promotion at the stage gate
- **Key takeaway:** you bring FRs + UAT; Nova provides NFRs + infra; QE guards the gate evidence; SRE signs off on production readiness
### Slide 6 — The Platform Pipeline
- Walk the pipeline left-to-right: contract → resolver → adapter → Checkov (static) → plan → Wiz (on plan) → confidence → gate → apply
- Two-stage scan: Checkov on static code BEFORE the plan (fail-fast dev feedback); Wiz on the plan (or Checkov as drop-in if no Wiz creds)
- Never both Wiz + Checkov on the plan — avoid duplicate noise
- Dev is autonomous; qa/prod/dr require attestation (QA for quality, SRE for production readiness)
- **Key takeaway:** two layers of scanning, zero operator involvement in normal operations
### Slide 7 — The Decision Ledger
- "AI decisions" are really automated decisions — deterministic scripts calculate a score; the platform functions without AI
- Do not dwell on the storage substrate — the value is accountability (immutable, queryable, traceable to outcome), not the database
- Every stage-gate attestation is captured with approver identity and the evidence presented
- When an LLM planner is added later, it emits richer alternatives without breaking the schema
- **Key takeaway:** autonomous is defensible because every decision is immutable, queryable, accountable — and "automated" means deterministic scoring, not a black-box LLM
### Slide 8 — The Attestation Matrix
- The matrix is not a rubber stamp — structured, freshness-validated, separation-of-duties-enforced
- Each concern now has a plain-language description of what is being attested (the old "operator-supplied" label is gone)
- SoD on prod: the approver can't be the same person who built it
- **Key takeaway:** autonomy in operations, human in accountability, by design — the matrix is what makes autonomous operations safe enough to trust in production
### Slide 9 — Telemetry & Live Ops
- Deliberately minimal: Nova-native CloudEvents; no Kafka/Prometheus/ClickHouse
- The live-ops dashboard is built in PowerBI on top of the exported views — leadership sees the same numbers the platform produces
- Every number in the Proof slides is traceable to a signal — "where does this number come from?" → a query against the cold store
- This is where the "infrastructure operations become visible" theme lands concretely
- **Key takeaway:** the architecture is the trust substrate — operations become visible in PowerBI, with full traceability
### Slide 10 — Decision Ledger + Attestation Coverage
- Both 100% — no automated decision is ever lost; no prod/dr promotion lands without a human sign-off
- The mandatory-by-design point: the ledger entry + the human attestation are a gate, not a best-effort feature
- Easily queried: by run, by environment, by approver, by outcome — the audit trail is a query, not a forensic exercise
- **Key takeaway:** trust is provable — not a marketing claim, a queryable record; no change to production without both the ledger entry and the human attestation
### Slide 11 — Cost & ROI
- The ROI formula is shown inline — not hidden in a footnote
- The four CTO-grade metrics are the ROI proof — Lead Time, Vuln Count, MTTR, Cloud Spend
- The N=0 caveat is stated explicitly: the formula is grounded; the production numbers activate with a pilot
- **Key takeaway:** the ROI is not a black box — the formula is shown, the four metrics are committed, the production-denominator caveat is up front
### Slide 12 — What's Deferred — and Why
- The preempt is critical: these deferrals are measurement infrastructure, not autonomy — the platform IS autonomous in operations
- The blocking work is named in plain language (no decision IDs) — "live AWS re-provisioning", "drift-detection scheduler", "ML service"
- Showing this to leadership demonstrates honesty, not weakness
- **Key takeaway:** the autonomy is real; the measurement gaps are documented with the work that unblocks each one
### Slide 13 — Roadmap to the North Star
- Each deferred metric has an unblock path and a timeframe — near-term, mid-term, longer-term
- No status column: most of it is not implemented yet, so status would be noise
- Re-evaluation triggers: each blocking piece of work lifts on its own schedule
- **Key takeaway:** every deferred metric has a plan and a timeframe — nothing is hand-waved
### Slide 14 — 12-Month Product Roadmap
- This is the *product* roadmap, forward-looking only
- Q1 Pilot Activation → Q2 Provable Trust → Q3 Compounding ROI → Q4 Integration & Predictive
- Each quarter activates one strategic objective from the North Star
- **Key takeaway:** the 12-month product arc — each quarter activates a strategic objective and its board-level metric
### Slide 15 — Quarter-by-Quarter Outcomes
- Q1: three post-pilot metrics go live (Touchless ≥99%, Escalation <0.1%, Accuracy ≥99.5%) — denominator activates with the pilot
- Q2: Decision Ledger Coverage was already grounded — tamper-evidence is the Q2 upgrade (local hash-chain → Object Lock + signed checkpoints)
- Q3: Drift Auto-Reversal ≥95% unblocks when the drift scheduler ships; Spend Reduction ≥25% measured against the pilot baseline
- Q4: Predictive:Reactive ≥3:1 requires the ML forecasting service; AI-Agent Intent Share is a first measurement (aspirational-metric)
- **Key takeaway:** each quarter has a concrete deliverable, a target metric grounded in a strategic objective, and a path from deferred to shipped
### Slide 16 — Production-Grade Guidance via Atelier (1/2)
- Nova instructs the citizen developer's AI agent via skills (markdown, keyed to engineering domains) + an MCP server (4 tools, plugin-registry, stdio)
- The integration point is the same regardless of source — AI agent, agentic SDLC, traditional IDE all get the same skills + MCP
- This is how Nova makes the citizen developer production-grade without owning the PDLC
- **Key takeaway:** the citizen developer's AI agent is not unguided — Nova provides engineering principles as skills + MCP
### Slide 17 — Production-Grade Guidance via Atelier (2/2)
- The value is the gap deterministic scanners leave: engineering discipline (Wiz/Checkmarx/Mend check policy/secrets, not discipline)
- The MCP server catches "is this service observable?", "is this error path handled?", "is this API contract clear?"
- Vendored at a pinned tag → audit reproducibility — a validation result is replayable months later
- **Key takeaway:** submissions are checked for engineering discipline, not just policy compliance — and the check is reproducible for audit
### Slide 18 — Recap + Ask
- Recap the 4-beat arc so the audience leaves with the structure
- The ask is a business decision: approve a pilot estate + the tamper-evident ledger build-out
- "Pipeline-ready" → "production-proven" is the value proposition
- **Key takeaway:** approve a pilot + the ledger build-out to move from pipeline-ready to production-proven
### Appendix A1 — Metrics Glossary
- Reference for every metric mentioned in the deck
- Use if the audience asks "what does X mean?"
File diff suppressed because one or more lines are too long
@@ -0,0 +1,713 @@
# Nova — The Autonomous Cloud Delivery Platform
> **Source of truth** (Step 1 of the 4-step deck process).
> Unified narrative deck. 4-beat arc: Problem → Solution → Proof →
> Roadmap + Ask. x3 structure at deck level (opening = the problem + the
> arc, body = tell them, closing = recap + ask) AND per slide (opens with
> what it covers, delivers, closes with a benefit callout written for a
> tech-leadership audience).
>
> **Honesty model:** every metric cited is grounded (cites a source),
> derived (documented formula), or deferred (cites the blocking work).
> No fabricated numbers. Internal provenance (decision IDs, requirement
> IDs, internal file paths) is kept out of the audience-facing slides —
> those live in the appendix and the `.ciagent/` files only.
>
> v1.21 — Deck Refinement & Pipeline Hardening
---
## Slide 1 — The Problem
**Product teams now own their cloud infrastructure — but ownership without
discipline is destroying value.**
The broad shift to "you build it, you run it" put Terraform into the hands
of product teams. The intention was right: teams that own their stack ship
faster. The reality is that infrastructure-as-code is a different craft
from software development, and the engineering standards that teams apply
to application code are rarely applied to the infrastructure that carries
it.
- **No lifecycle planning.** Resources are authored for creation, not for
patching, decommissioning, or rollback. When a change is needed, the
change is destructive — because no one planned the lifecycle.
- **Proactive scanning is not part of authoring.** In a year where
AI-frontier models discover and exploit zero-day vulnerabilities at a
rapid pace, teams cannot keep up by reacting. Infrastructure modules
must be scanned as code and at runtime, post-deployment — and remediated
at the pace the threat moves, not the pace a sprint allows.
- **Bandwidth gaps in infrastructure operations.** An unusual amount of
time is spent on remediation, the push for innovation does not pause,
and the result is that operational work is chronically under-resourced.
Gaps open. Detections are missed. Incidents grow.
- **Tribal knowledge and the rockstar-operator problem.** Operations
depend on a handful of administrators who hold the infrastructure in
their heads. When they leave, the knowledge leaves with them. The
platform should encode the discipline, not the person.
Every hour a developer spends writing, deploying, fixing, or remediating
infrastructure is an hour not spent releasing features to production and
generating value.
> **Benefit:** the rest of this deck shows the answer — an autonomous
> cloud delivery platform that encodes infrastructure discipline as
> policy, scans proactively, remediates rapidly, and makes operations
> visible to leadership rather than hidden in tribal knowledge.
> **Speaker notes:** Do not frame this as "humans are the problem." The
> problem is that ownership was granted without the discipline, tooling,
> and lifecycle planning that infrastructure requires. The operator is
> not the bottleneck because operators exist — the bottleneck is that
> operations depend on a few individuals instead of an encoded system.
> **Transition:** "Here is the destination Nova is building toward."
---
## Slide 2 — Nova's Vision
**Infrastructure operations become visible. Every environment provisioned,
every incident healed, every risk remediated — by an autonomous system
whose trustworthiness is provable, not promised. Human attestation remains
required at stage gates; the operator is never in the loop of normal
operations.**
- **Visibility is the recurring theme.** Security posture, remediation
velocity, reliability, and lead time are surfaced as queryable signals —
not hidden in a person's head or a Slack thread.
- **Provable, not promised.** Trust is established by deterministic
scripts that calculate a score and gate the action. The platform
functions without AI. "AI decisions" are really automated decisions.
- **Autonomy in operations, human at stage gates.** QA signs off for
production; SRE greenlights based on operational readiness. The
absence of an operator in the loop is never the absence of a record.
> **Benefit:** the destination is autonomous operations with provable
> trust — security, remediation velocity, reliability, and lead time made
> visible to leadership, not promised to them.
> **Speaker notes:** "Visible" is the operative word. The vision is not
> just that operations run without an operator — it is that operations
> become observable, queryable, and accountable. That is what makes the
> trust defensible.
> **Transition:** "The vision is ambitious — here are the strategic
> objectives that make it concrete, and the anti-goals that keep it
> focused."
---
## Slide 3 — Strategic Objectives + Anti-Goals
**Four objectives Nova is building toward; four anti-goals that keep it
focused.**
**4 Strategic Objectives:**
1. **Demonstrate production-grade zero-touch operations** — autonomy as
the default, not the demo. Stage-gate attestation (QA, SRE) remains
human by design.
2. **Establish provable trust in automated decisions** — deterministic
scripts calculate a score; a band outcome gates the action. The
platform functions without AI. The Decision Ledger, confidence
scoring, circuit breakers, and blast-radius controls make
"autonomous" a defensible claim, not a marketing one.
3. **Deliver compounding, quantifiable ROI** — measured on four CTO-grade
metrics, all flowing into PowerBI:
- **Lead Time** (PR → Production) — downward trend.
- **Infrastructure Vulnerability Count** — downward trend
(proactive scanning keeps up with the AI-era 0-day pace).
- **MTTR** — for platform-detected and platform-remediated incidents.
- **Cloud Spend Reduction** — on pilot estates vs. the pre-Nova
baseline.
4. **Integrate with externally owned development platforms — regardless
of source.** Nova integrates with externally owned PDLC, SDLC,
Agentic, and Citizen Developer platforms. Nova provides skills and
MCP endpoints that help the developer or AI agent make their
application production-grade. Regardless of the source, all intents
to deploy to production go through the same rigorous controls,
quality gates, attestation, and evidence stream.
**4 Anti-Goals (what Nova is NOT):**
1. Not a general-purpose AI agent platform.
2. Not a system that removes humans from accountability — only from
normal operations.
3. Not an upstream development platform (no product backlogs, IDE, code
authorship).
4. Not a replacement for the Product Development Lifecycle (PDLC).
> **Benefit:** the scope is explicit — Nova governs infrastructure and
> delivery, integrates with any upstream source through one validated
> contract, and measures success on four metrics a CTO can repeat back.
> **Speaker notes:** Objective #2 is the one to land carefully: trust is
> established by deterministic scoring, not by an LLM. The platform
> functions without AI. Anti-goals #3 and #4 protect the scope boundary —
> Nova will not become an IDE or a product-planning tool.
> **Transition:** "The scope boundary is explicit — here is exactly
> where Nova sits relative to the product development lifecycle."
---
## Slide 4 — Scope: Downstream of PDLC
**Nova governs infrastructure and delivery. The PDLC is upstream — Nova
never penetrates it. Integration is through one validated contract.**
- **The PDLC is upstream:** product backlog, code authorship (AI agent,
IDE, agentic SDLC), sprint planning, application business logic.
- **Nova is downstream:** contract ingestion → submission-readiness gate
→ policy enforcement → cloud resource lifecycle → environment
progression (dev → qa → prod → dr) → immutable audit + attestation.
- **The integration point is one contract.** The citizen developer's AI
coding agent, an upstream agentic SDLC platform, or any development
platform may all produce submissions — the source does not matter
because all are subject to the same compliance standards.
- **Nova validates the submission, not the author.** The audit trail is
the same; the policy envelope is the same; the evidence stream is the
same.
> **Benefit:** a clean scope boundary — Nova is purpose-built for
> infrastructure operations and integrates with any upstream source
> through one validated contract, so the platform team's surface area
> stays bounded.
> **Speaker notes:** This slide protects the scope. The moment Nova
> starts owning the PDLC, it loses focus. The contract boundary is what
> keeps Nova deep on infrastructure and delivery rather than shallow on
> everything.
> **Transition:** "With the scope clear, here is who owns what across the
> delivery lifecycle."
---
## Slide 5 — RACI: Who Owns What
**Four roles, one matrix — the citizen developer owns FRs + UAT, the
platform owns NFRs + infra, quality engineering owns the gate evidence,
and SRE owns operational readiness.**
| Work Category | Citizen Dev | Platform | Quality Eng | SRE |
|---|---|---|---|---|
| Functional Requirements | **R/A** | C | I | I |
| User Acceptance Testing | **R/A** | C | I | I |
| Non-Functional Requirements | I | **R/A** | C | C |
| Infrastructure (cloud, state, IAM) | I | **R/A** | I | C |
| QA (policy, confidence, schema) | C | R | **R/A** | I |
| Production deployment to cloud | I | **R/A** | C | C |
| Quality attestation (QA sign-off) | **A** | R | **R** | I |
| Production readiness (SRE sign-off) | **A** | R | C | **R** |
**R** = Responsible · **A** = Accountable (sign-off) · **C** = Consulted · **I** = Informed.
- **Compliance-standard equivalence:** FRs + UAT may come from any
upstream source (AI agent, agentic SDLC, dev platform) — all pass the
same submission-readiness gate.
- **Production readiness is co-owned:** the platform runs the
attestations agentically; the citizen developer authorizes the
promotion at the stage gate.
> **Benefit:** every party knows what they bring, what the platform
> provides, what quality engineering guards, and where SRE signs off —
> accountability is explicit, never diffuse.
> **Speaker notes:** Quality attestation is now owned by Quality
> Engineering (not the Platform), and Production readiness is owned by
> SRE. The Platform runs the checks agentically but is never the
> Accountable party for the gate — that separation keeps the platform
> honest.
> **Transition:** "With ownership clear, here is how the pipeline
> enforces it."
---
## Slide 6 — The Platform Pipeline
**How intent becomes verified infrastructure — with fail-fast policy
scanning before the plan and runtime scanning after it.**
```mermaid
graph LR
A[Contract] --> B[Resolver]
B --> C[Adapter]
C --> D["Checkov (static code)"]
D --> E[Terraform Plan]
E --> F["Wiz (on plan)"]
F --> G[Confidence Signal]
G --> H{Stage Gate}
H -->|dev: autonomous| I[Apply]
H -->|qa/prod/dr: attested| I
I --> J[Evidence + Ledger]
```
- **Contract → resolver → adapter → Checkov on static code (before the
plan) → terraform plan → Wiz on the plan → confidence signal → stage
gate → apply → evidence + ledger.**
- **Fail-fast, quick feedback.** Checkov runs on the authored Terraform
code before `terraform plan` so developers get immediate policy
feedback, not a delayed plan-stage failure.
- **Wiz on the plan when configured; Checkov as a drop-in otherwise.**
Wiz scans the terraform plan output. When Wiz credentials are not
available, Checkov runs against the plan as a drop-in replacement. Wiz
and Checkov are never both run on the plan.
- **Dev is autonomous** (no stage gate); **qa/prod/dr require human
attestation** (QA for quality, SRE for production readiness).
> **Benefit:** the pipeline gives developers fast, deterministic feedback
> on policy at authoring time and gives the platform a runtime scan on the
> resolved plan — two layers of scanning, zero operator involvement in
> normal operations.
> **Speaker notes:** The two-stage scan is the key design: static code
> scanning catches policy violations before the cost of a plan; runtime
> plan scanning catches what the static code cannot (resolved values,
cross-resource issues). The platform picks the runtime scanner based on
configuration — never both, to avoid duplicate noise.
> **Transition:** "The pipeline produces decisions — here is how every
> decision is captured and made accountable."
---
## Slide 7 — The Decision Ledger
**Every automated decision is captured, immutable, queryable — and
accountable.**
- **What is captured:** every action the platform takes — the chosen
action, the confidence score, the alternatives considered, whether a
human overrode it, and the outcome (backfilled once the apply
completes). Every stage-gate attestation (QA sign-off, SRE
production-readiness sign-off) is captured with approver identity and
the evidence that was presented.
- **"AI decisions" are really automated decisions.** The decisions are
made by deterministic scripts that calculate a score and a band; the
platform functions without AI. The ledger captures the real decision
path — not a fabricated "AI agent." When an LLM planner is added later,
it will emit richer alternatives without breaking the schema.
- **The value is accountability, not the storage engine.** The ledger is
an append-only, tamper-evident record. The point is not which database
it lives in — the point is that every decision is queryable for
auditing, traceable to an outcome, and impossible to rewrite after the
fact.
> **Benefit:** "autonomous" is defensible because every decision the
> platform makes is immutable, queryable, and accountable — and the
> audience knows exactly what "automated" means here: deterministic
> scoring, not a black-box LLM.
> **Speaker notes:** Do not dwell on the storage substrate. The audience
> cares that the ledger is append-only, queryable, and tied to outcomes —
> not that it is a hash-chain in a SQLite file. The D-122 honesty point
> is restated without the decision ID: the platform's decisions are
> deterministic; the ledger captures that real path.
> **Transition:** "Decisions are captured — here is how stage-gate
> attestation keeps humans in accountability."
---
## Slide 8 — The Attestation Matrix
**The designed controls that keep humans at stage gates — structured,
freshness-validated, and separation-of-duties-enforced.**
| Concern | Env | Freshness | Description |
|---------|-----|-----------|-------------|
| Functional correctness | qa | 24h | The application behaves as specified; evidence accepted from the consumer's UAT. |
| Performance baseline | qa | 7d | The deployment meets its performance envelope vs. the agreed baseline. |
| Security posture | qa | 24h | The deployment's security findings have been reviewed and accepted. |
| Operational readiness | prod | 30d | SRE confirms the deployment is operable: runbooks, dashboards, on-call coverage. |
| Incident response | prod | 90d | The on-call path has been exercised; the deployment has a working incident-response plan. |
| Capacity & cost | prod | 30d | Capacity headroom and monthly cost are within the agreed envelope. |
| Resilience: DR drill | prod | 180d | A DR drill has been run and the deployment recovered within the RTO. |
| Resilience: chaos | prod | 90d | A chaos exercise has been run and the deployment absorbed the failure. |
| Resilience: backup | prod | 30d | Backups are restorable and have been tested within the freshness window. |
| DR region deploy | dr | 180d | The DR region can be deployed and the deployment is reachable from it. |
- Each concern has a freshness window — evidence older than the window
does not satisfy the gate.
- **Separation-of-duties on prod:** the approver cannot be the same
person who built the deployment.
- Concerns that are offline-testable run for real; concerns that require
external evidence accept signed artifacts.
> **Benefit:** the gate model is explicit — autonomy in operations,
> human in accountability, by design. The matrix is what makes autonomous
> operations safe enough to trust in production.
> **Speaker notes:** The matrix is not a rubber stamp. Each concern has a
> freshness window, a description, and a separation-of-duties rule. The
> "operator-supplied" label from the prior deck was dropped — every
> concern now has a plain-language description of what is being attested.
> **Transition:** "You've seen how Nova works — the pipeline, the ledger,
> the attestation gates. Here is how Nova instruments itself so that
> every claim in this deck is traceable to a real signal."
---
## Slide 9 — Telemetry & Live Ops
**Every metric in this deck is traceable to a real emitted signal — and
the live-ops dashboard makes operations visible in PowerBI.**
```mermaid
graph TB
A[Platform components] --> B[CloudEvents envelope]
B --> C[Event log]
B --> D[Decision ledger]
B --> E[Run records]
C --> F[Collector]
D --> F
E --> F
F --> G[Cold store]
G --> H[PowerBI views]
H --> I[Live ops dashboard]
```
- **Platform components emit a CloudEvents envelope** → event log,
decision ledger, and run records → collector → cold store → PowerBI
views → **live ops dashboard.**
- **The live ops dashboard (PowerBI)** surfaces the four CTO-grade
metrics — Lead Time, Infrastructure Vulnerability Count, MTTR, Cloud
Spend — alongside the trust metrics (Decision Ledger coverage,
Attestation coverage) and the efficiency metrics (touchless
resolution, escalation frequency).
- **The architecture is deliberately minimal.** Nova-native envelopes;
no Kafka, no Prometheus, no ClickHouse. The cold store is sufficient
for batch and historical analysis; the live-ops surface is built in
PowerBI on top of the exported views.
- **Every number in the Proof slides is traceable to a signal.** When a
CFO asks "where does this number come from?", the answer is a query
against the cold store, not a Slack thread.
> **Benefit:** the architecture is the trust substrate — leadership sees
> the same numbers the platform produces, in PowerBI, with full
> traceability to the emitted signal. Operations become visible.
> **Speaker notes:** The value is not the plumbing — it is that the
> platform's metrics surface in a tool leadership already uses (PowerBI),
> and every number is traceable. The live-ops dashboard is where the
> "infrastructure operations become visible" theme lands concretely.
> **Transition:** "The architecture is sound — here is the measured
> proof."
---
## Slide 10 — Decision Ledger + Attestation Coverage
**By design, no change reaches production without a ledger entry and a
human attestation — both queryable for auditing, with full
traceability.**
- **Decision Ledger coverage: 100%.** Every platform run emits a
decision record with outcome backfill. No automated decision is ever
lost.
- **Attestation coverage: 100%.** Every prod/dr promotion is attested by
a human — QA for quality, SRE for production readiness — recorded with
approver identity, separation-of-duties check, and the evidence matrix.
- **No change to production without both.** The ledger entry and the
human attestation are mandatory, not optional. This is enforced by the
pipeline, not by policy.
- **Easily queried for auditing.** The ledger and the attestation
records are queryable by run, by environment, by approver, and by
outcome — the audit trail is a query, not a forensic exercise.
- **Full traceability.** A production change is traceable from the
contract that declared intent, through the policy scan, the confidence
score, the attestation, to the applied outcome. Nothing is opaque.
> **Benefit:** trust is provable — not a marketing claim, a queryable
> record. An auditor can answer "who approved this, when, on what
> evidence?" in one query; a CTO can answer "how many of last quarter's
> prod changes were touchless?" in one query.
> **Speaker notes:** The mandatory-by-design point is the one to land.
> The ledger + attestation are not a best-effort feature; they are a
> gate. No change reaches production without both. That is what makes
> the 100% numbers credible — they are enforced, not aspirational.
> **Transition:** "Trust is provable — here is the cost side of the ROI."
---
## Slide 11 — Cost & ROI
**The ROI formula and the cost estimates — grounded, with the production
denominator honestly flagged.**
- **Cost estimates are pre-apply and offline.** The platform reads the
terraform plan and estimates cost before anything is applied — so a
regression in cost is caught before the spend happens, not after.
- **The ROI formula:**
`Platform ROI = (FTE hours saved × blended rate + cloud savings + avoided downtime) ÷ platform op cost`
- **The four CTO-grade metrics (from Slide 3) are the ROI proof:**
Lead Time (PR → Prod), Infrastructure Vulnerability Count (trend), MTTR,
Cloud Spend Reduction. All flow into PowerBI.
- **Honest caveat:** the derived metrics are computed on internal runs
today; the production-denominator activates when a pilot estate runs.
The formula is grounded; the production numbers are not yet.
> **Benefit:** the ROI is not a black box — the formula is shown, the
> four metrics are committed, and the production-denominator caveat is
> stated up front. The CFO can see exactly what is real today and what
> activates with a pilot.
> **Speaker notes:** The formula is shown inline, not hidden. The
> "no fabrication" constraint in action: show the formula, show the
> caveat, do not pretend the production numbers exist.
> **Transition:** "The proof is grounded — here is what is honestly
> deferred, and why."
---
## Slide 12 — What's Deferred — and Why
**Honesty about what is not measured yet — and the blocking work for
each.**
To be clear: these deferrals are measurement infrastructure, not the
autonomy itself. The platform runs without an operator in the loop of
normal operations. What is deferred is the evidence pipeline for certain
metrics — not the autonomy.
| # | Deferred metric | Blocking work |
|---|-----------------|---------------|
| 1 | Live infrastructure health | Live AWS re-provisioning (currently torn down to a zero-cost steady state) |
| 2 | Live outbox write rate | Live AWS re-provisioning |
| 3 | Tamper-evident ledger checkpoints | Audit-ledger build-out (S3 Object Lock + signed checkpoints) |
| 4 | Onboarding funnel (requested → granted) | Auto-grant implementation |
| 5 | Drift auto-reversal | Drift-detection scheduler (not yet built) |
| 6 | Live cost reconciliation | Live AWS re-provisioning + actual-spend feed |
| 7 | SLA / unplanned downtime | Live AWS re-provisioning |
| 8 | Predictive vs reactive ratio | ML anomaly-forecasting service (not yet built) |
> **Benefit:** the boundaries are explicit — what Nova measures today,
> and exactly what blocks the rest. The autonomy is real; the measurement
> gaps are documented with the work that unblocks each one.
> **Speaker notes:** The preempt is critical: these deferrals are
> measurement infrastructure, not autonomy. The platform runs without an
> operator in the loop. What is deferred is the evidence pipeline for
> live-infra health, drift, predictive remediation — not the autonomy
> itself.
> **Transition:** "The proof is honest — here is the roadmap from here to
> the targets."
---
## Slide 13 — Roadmap to the North Star
**The path from the grounded metrics to the 1218 month targets — each
deferred metric has an unblock path and a candidate milestone.**
| Timeframe | Work | Unblocks |
|-----------|------|----------|
| Near-term | Live AWS re-provisioning | Live infra health, live outbox write rate, live cost reconciliation, SLA |
| Near-term | Auto-grant implementation | Onboarding funnel (requested → granted) |
| Mid-term | Drift-detection scheduler | Drift auto-reversal |
| Mid-term | Audit-ledger build-out (Object Lock + signed checkpoints) | Tamper-evident ledger checkpoints |
| Mid-term | Hot-path activation (live-ops dashboard goes from batch to near-real-time) | Live-ops dashboard freshness |
| Longer-term | ML anomaly-forecasting service | Predictive vs reactive ratio |
- Each deferred metric has a specific unblock requirement and a
candidate future milestone.
- Re-evaluation triggers: each blocking piece of work lifts on its own
schedule; the metrics layer evolves as each one lands.
> **Benefit:** every deferred metric has an unblock path — nothing is
> hand-waved; everything has a plan and a timeframe.
> **Speaker notes:** This is the bridge from "honestly deferred" to
> "here is how we get there." The roadmap uses timeframes, not status —
> most of it is not implemented yet, so a status column would be noise.
> **Transition:** "The unblock path is clear — here is the 12-month
> product arc."
---
## Slide 14 — 12-Month Product Roadmap
**The product arc from pilot activation to integration — four quarters,
four outcomes.**
| Quarter | Theme | Board-level outcome |
|---------|-------|---------------------|
| **Q1** | Pilot Activation | Nova runs a real customer estate end-to-end, autonomously, with a measurable zero-touch rate. |
| **Q2** | Provable Trust | Every automated decision lands in a tamper-evident ledger; the CFO sees real cloud-spend reconciliation. |
| **Q3** | Compounding ROI | Quarter-over-quarter cloud spend drops; drift is detected and reversed without a human. |
| **Q4** | Integration & Predictive | AI agents deploy through Nova by default; the ML anomaly-forecasting service goes live. |
Grounded in the four strategic objectives (autonomy, provable trust, ROI,
integration) and the deferred-metric unblock paths.
> **Benefit:** the 12-month product arc — each quarter activates a
> strategic objective and its corresponding board-level metric, from
> pilot activation through integration leadership.
> **Speaker notes:** The roadmap is organized by product outcome, not
> by technical milestone. Each quarter activates one strategic
> objective from the North Star.
> **Transition:** "Here is the quarter-by-quarter detail."
---
## Slide 15 — Quarter-by-Quarter Outcomes
| Quarter | Product theme | Key deliverable | Target metric | Grounding |
|---------|---------------|-----------------|---------------|-----------|
| **Q1** | Pilot Activation | Re-provision live AWS; activate first pilot estate; onboarding auto-grant | Touchless ≥ 99% · Escalation < 0.1% · Accuracy ≥ 99.5% | Objective #1 — autonomy as the default |
| **Q2** | Provable Trust | Tamper-evident ledger (Object Lock + signed checkpoints); daily checkpoints; live cost reconciliation | Decision Ledger Coverage 100% · Cost Savings ≥ 25% | Objective #2 — trust is the moat |
| **Q3** | Compounding ROI + Drift | Drift-detection scheduler; auto-reversal; pre-apply → actual-spend reconciliation on the pilot estate | Drift Auto-Reversal ≥ 95% · Spend Reduction ≥ 25% | Objective #3 — CFO-pointable numbers |
| **Q4** | Integration + Predictive | ML anomaly-forecasting; AI-agent intent surface; multi-cloud (Azure/GCP) preview | Predictive:Reactive ≥ 3:1 · AI-Agent Intent Share (first measurement) | Objective #4 — default substrate for agents |
**Month-18 destination:** *"Nova is the layer enterprise leadership
points to when they say 'we don't have an infrastructure ops team
anymore, and the audit trail is stronger than it ever was.'"*
> **Benefit:** each quarter has a concrete deliverable, a target metric
> grounded in a strategic objective, and a path from "honestly deferred"
> to "shipped and measured."
> **Speaker notes:** Q1Q3 are committed (grounded pipeline + known
> unblock paths). Q4 targets are committed-deliverable,
> aspirational-metric — the ML service ships, the intent-share number is
> a first measurement (we do not control adoption rate).
> **Transition:** "Production-grade guidance is how Nova helps the
> citizen developer's AI agent meet the bar — here is the first half."
---
## Slide 16 — Production-Grade Guidance via Atelier (1/2)
**Nova instructs the citizen developer's AI agent on production-grade
engineering — a set of skills and an MCP server.**
- **Skills** — markdown files keyed to production-grade engineering
domains (API, security, data, testing, observability, errors, DevOps,
infrastructure-as-code, compliance). The skills extend the baseline
catalog with Nova-specific production-grade principles.
- **MCP server** — a plugin-registry, stdio server exposing four tools:
`lookup_principle`, `list_domains`, `matrix_lookup`, and
`validate_against_principles`. The developer's AI agent (or any
agentic SDLC platform) calls these tools to look up the principles
that apply to its submission.
- **The integration point is the same regardless of source.** Whether
the submission comes from an AI coding agent, an agentic SDLC
platform, or a traditional IDE, the same skills and MCP server apply.
This is how Nova makes the citizen developer production-grade without
owning the PDLC.
> **Benefit:** the citizen developer's AI agent is not unguided — Nova
> provides production-grade engineering principles as skills and as an
> MCP surface, so submissions arrive at the contract boundary already
> aligned with the platform's standards.
> **Speaker notes:** This is the first half of the Atelier story — the
> surface (skills + MCP). The next slide is what the surface catches
> that deterministic scanners cannot.
> **Transition:** "Here is what that guidance catches that deterministic
> scanners cannot."
---
## Slide 17 — Production-Grade Guidance via Atelier (2/2)
**Agentic validation catches engineering-discipline gaps that deterministic
scanners miss — and the validation is reproducible.**
- **Beyond deterministic scanners.** Wiz, Checkmarx, and Mend check
policy and secrets — they do not check engineering discipline. The
Atelier MCP server catches correctness, clarity, and observability gaps
that deterministic tools cannot: "is this service observable?",
"is this error path handled?", "is this API contract clear?"
- **Agentic validation, not a second policy engine.** The MCP server
gives the AI agent the principles to validate against; the agent does
the validation. This is agentic validation — the agent reasons about
the submission against the principles, not a second static scan.
- **Vendored for audit reproducibility.** Atelier is vendored at a
pinned tag. A validation result is replayable against the exact
principles that produced it — so an audit can reproduce a validation
months later, not just trust a log line.
> **Benefit:** the citizen developer's submission is checked for
> engineering discipline, not just policy compliance — and the check is
> reproducible for audit. That is what makes the submission
> production-grade, regardless of which upstream platform produced it.
> **Speaker notes:** The value is the gap deterministic scanners leave:
engineering discipline. Policy scanners catch "is this S3 bucket
public?"; the MCP server catches "is this service observable if that
bucket fails?". The vendoring point is audit reproducibility — the
validation is not a black box.
> **Transition:** "You've seen the problem, the solution, and the proof.
> Here is the recap and the ask."
---
## Slide 18 — Recap + Ask
**The 4-beat recap + the business decision.**
**Recap:**
- **Problem:** product teams own infrastructure without the discipline
and lifecycle planning it requires; bandwidth gaps and tribal
knowledge leave operations exposed.
- **Solution:** autonomous cloud delivery — operations become visible,
trust is provable (deterministic scoring), humans at stage gates.
- **Proof:** 100% ledger coverage, 100% attestation coverage, grounded
ROI formula, four CTO-grade metrics flowing into PowerBI.
- **Roadmap:** deferred metrics have unblock paths; the 12-month product
arc activates one strategic objective per quarter.
**The ask:** "Approve a pilot estate to activate the production-denominator
metrics (Lead Time, Vulnerability Count, MTTR, Cloud Spend), and approve
the tamper-evident ledger build-out to move from the local hash-chain to
S3 Object Lock + signed checkpoints. These two decisions move Nova from
'pipeline-ready' to 'production-proven.'"
> **Benefit:** a clear business decision — approve a pilot and the ledger
> build-out — with the confidence that every claim in this deck is
> grounded, derived, or honestly deferred.
> **Speaker notes:** The ask is a business decision, not insider
> language. "Approve a pilot estate" is a C-suite decision. "Approve the
> ledger build-out" is a budget decision. The recap reinforces the 4-beat
> arc — the audience leaves with the structure, not a pile of facts.
---
## Appendix A1 — Metrics Glossary
| KPI | Definition | Status |
|-----|-----------|--------|
| Touchless Resolution Rate | runs without operational stage-gate block ÷ total | partial (Post-Pilot) |
| Human Escalation Frequency | operational stage-gate blocks ÷ total | partial (Post-Pilot) |
| Automated Decision Accuracy | decisions not followed by failure within 5min | partial (Post-Pilot) |
| MTTR (p95) | apply.failed → successful retry | grounded |
| Confidence-Gate Halt Rate | runs with band=block ÷ total | grounded |
| Provisioning Lead Time | run.completed run.started | grounded |
| Deployment Frequency | count(run.completed) per day | grounded |
| Cost Savings (pre-apply) | sum(delta_usd where delta < 0) | partial (live reconciliation deferred) |
| FTE Hours Saved | run count × manual baseline × rate | derived (N=0 caveat) |
| Platform ROI | (labor + cloud + avoided downtime) ÷ op cost | derived (N=0 caveat) |
| Decision Ledger Coverage | decisions with outcome ÷ total | grounded |
| Attestation Coverage | prod/dr attested ÷ total prod/dr | grounded |
| Policy Compliance Rate | 1 failed_assets ÷ total | grounded |
> **Benefit:** a reference for every metric mentioned in the deck.
---
> **End of deck.** 18 main slides + 1 appendix slide = 19 total.
@@ -1,324 +0,0 @@
---
marp: true
theme: default
paginate: true
size: 16x9
header: 'Nova — The No-Humans Infrastructure Platform'
footer: 'Act %{page}/5 — v1.17'
style: |
section { font-size: 0.85em; }
h1 { color: #1a1a2e; }
h2 { color: #16213e; }
table { font-size: 0.75em; }
.badge { padding: 2px 8px; border-radius: 3px; font-size: 0.8em; }
.badge.planned { background: #fff3cd; color: #856404; }
section.title { background: #1a1a2e; color: white; }
---
<!-- _class: title -->
<!-- _paginate: false -->
# Nova — The No-Humans Infrastructure Platform
**Shifting from Operational Overhead to Strategic Value**
v1.17 — Strategic Direction, Leadership Metrics & Unified Story
---
## Slide 1 — Arc Preview
**This deck proves Nova is the no-humans infrastructure platform — and shows you the metrics that make the claim defensible.**
**Today:** 18 capabilities verified, 0 consumer estates in production.
**The 5-act arc:**
1. **Problem** — why the operator is the bottleneck
2. **Vision** — Nova's strategic direction (NORTH_STAR)
3. **How** — the pipeline, Decision Ledger, attestation gates
4. **Proof** — grounded metrics that make the claim defensible
5. **Roadmap** — deferred metrics with unblock paths + the ask
**Benefit:** you leave knowing which claims are proven today, which are pipeline-ready, and which are deferred with a documented unblock path — no marketing, just grounded evidence.
---
## Slide 2 — The No-Humans Imperative
**Why the operator is the bottleneck — and why removing them from operations (not accountability) is the imperative.**
- **The cost of humans-in-the-loop:** L1/L2 ops hours, escalation latency, the trust gap
- **The operator is the bottleneck:** provisioning takes days, not minutes
- **The attestation model:** autonomy in operations, human at stage gates
- Cites `docs/NO_HUMANS_THESIS.md`
**Benefit:** you now know the problem framing — autonomy in operations, human at stage gates, is the path forward.
---
## Slide 3 — Nova's Vision
> **Infrastructure operations become invisible. Every environment provisioned, every incident healed, every risk remediated — by an autonomous system whose trustworthiness is provable, not promised. Human attestation remains required at stage gates — QA signs off for production, SRE greenlights based on operational readiness — but the operator is never in the loop of normal operations.**
- Autonomy in operations, not in accountability
- Cites `docs/NO_HUMANS_THESIS.md`
**Benefit:** you now know the destination — invisible operations with provable trust, not promised trust.
---
## Slide 4 — Strategic Objectives + Anti-Goals
**4 Strategic Objectives:**
1. **Zero-touch operations** — autonomy as the default, not the demo
2. **Provable trust in AI decisions** — Decision Ledger, confidence scoring, circuit breakers
3. **Compounding, quantifiable ROI** — each quarter must reduce spend, free hours, avoid downtime
4. **Default substrate for agentic consumption** — the platform AI agents reach for first
**5 Anti-Goals (what Nova is NOT):**
1. Not a hyperscaler competitor
2. Not a general-purpose AI platform
3. Not removing humans from accountability
4. Not for legacy, untagged, or freeform infrastructure
5. Not sold to operators
**Benefit:** you now know the scope boundaries — Nova is purpose-built for infrastructure operations, sold to leadership on outcomes.
---
## Slide 5 — 1218 Month Targets
**Current-milestone targets (grounded/derived):**
| Domain | Target | Status |
|---|---|---|
| MTTR (p95) | < 60s | grounded |
| Cloud Spend Reduction | ≥ 25% | partial (CUR deferred D-096) |
| L1/L2 Ops Hours Avoided | ≥ 70% | derived (N internal runs) |
| Platform ROI | ≥ 250% | derived (formula; N=0 caveat) |
| Decision Ledger Coverage | 100% | grounded |
| Attestation Coverage | 100% | grounded |
**Post-Pilot targets (pipeline grounded; 0 consumers today):**
| Domain | Target | Status |
|---|---|---|
| Touchless Resolution Rate | ≥ 99% | partial |
| Human Escalation Frequency | < 0.1% | partial |
| AI Decision Accuracy | ≥ 99.5% | partial |
**Deferred:** Predictive vs Reactive ≥3:1 <span class="badge planned">Planned</span> · Drift Auto-Reversal ≥95% <span class="badge planned">Planned</span>
**Benefit:** you now know the destination numbers — and which are measurable today vs deferred honestly.
---
## Slide 6 — The Platform Pipeline
**How intent becomes verified infrastructure without an operator.**
Contract → Resolver → Adapter → Terraform Plan → Checkov (Policy) → Confidence Signal → HITL Gate → Apply → Evidence
- Dev: autonomous (no HITL gate)
- qa/prod/dr: attested (human sign-off required)
- Grounded in `run_platform.sh` + `contract_resolver.py` + `confidence_signal.py`
**Benefit:** you now know the path from intent to evidence — and where the human appears (stage gates only).
---
## Slide 7 — The Decision Ledger
**Every AI decision captured with confidence, alternatives, and outcome.**
- `outbox_writer.py` → SQLite append-only hash-chain table
- `ai.decision.made`: decision_id=run_id, chosen_action=band, confidence=score, alternatives=perInput, human_override=HITL block
- `attestation.recorded`: qa/prod/dr sign-offs
- D-121, D-122, D-132. Honors D-083 (no S3 Object Lock/JWS — local hash-chain)
**D-122 honesty:** Nova's "AI" is the confidence-gated policy engine (confidence_signal + HITL gate), not an LLM planner. The Decision Ledger captures this real decision path — not a fabricated "AI agent."
**Benefit:** you now know why 'autonomous' is defensible — every decision is immutable, queryable, and accountable. And you know exactly what 'AI' means here: a confidence-gated policy engine, not a black-box LLM.
---
## Slide 8 — The 8-Concern Attestation Matrix
**Designed controls that keep humans at stage gates.**
| Concern | Env | Freshness | Type |
|---------|-----|-----------|------|
| functional_correctness | qa | 24h | operator-supplied |
| performance_baseline | qa | 7d | operator-supplied |
| security_posture | qa | 24h | operator-supplied |
| operational_readiness | prod | 30d | operator-supplied |
| incident_response | prod | 90d | operator-supplied |
| capacity_cost | prod | 30d | operator-supplied |
| resilience_dr_drill | prod | 180d | operator-supplied |
| dr_region_deploy | dr | 180d | operator-supplied |
- Offline-testable concerns run for real; operator-supplied concerns accept signed evidence
- Separation-of-duties on prod
- Grounded in `attestation_matrix.py` + `hitl_gates.py`
**Benefit:** you now know the gate model — autonomy in operations, human in accountability, by design.
---
## Slide 9 — Telemetry Architecture
**How Nova instruments itself — CloudEvents envelope, cold store, PowerBI export.**
Platform → CloudEvents 1.0 → `metrics/events.jsonl` + `metrics/decision_ledger.db` + `metrics/runs/` → Collector → `metrics/nova_metrics.db` (SQLite cold store) → `metrics/powerbi/` (CSV/JSON) → PowerBI
- D-120 (Nova-native), D-125 (hybrid), D-126 (cold-only)
- <span class="badge planned">Planned</span>: Hot-path (live ops dashboard) — D-126
**Benefit:** you now know that every metric in this deck is traceable to a real emitted event — the architecture IS the trust substrate. When a CFO asks 'where does this number come from?', the answer is a file path, not a Slack thread.
---
## Slide 10 — Capability Health + Confidence Distribution
**Grounded proof: capability health and confidence distribution from real runs.**
| Status | Count |
|--------|-------|
| Verified | 18 |
| Skipped | 4 |
| Broken | 0 |
| Decayed | 0 |
- 4 Skipped = live-AWS caps (CAP-013..016), honestly skipped (D-096 teardown), not a failure
- Source: `.ciagent/REGRESSION_REPORT.json`
**Benefit:** you now know the platform is verified — 18 capabilities pass, 4 are honestly skipped, 0 broken.
---
## Slide 11 — Decision Ledger + Attestation Coverage
**Trust metrics — both 100%.**
- **Decision Ledger Coverage:** 100% of platform runs emit `ai.decision.made` with outcome backfill
- **Attestation Coverage:** 100% of prod/dr promotions attested by a human
- **AI Decision Accuracy:** decisions not followed by apply.failed/incident within 5min
- Trust snapshot: `metrics/TRUST_SNAPSHOT.md` with chain-integrity verdict
- <span class="badge planned">Planned</span>: Tamper-Evident Ledger Checkpoints (D-083)
**Benefit:** you now know the trust is provable — not a marketing claim, a queryable record.
---
## Slide 12 — Zero-Touch Efficiency
**Touchless resolution, human escalation, and MTTR.**
- **Touchless Resolution Rate:** runs without operational HITL block ÷ total (attestation gates excluded)
- **Human Escalation Frequency:** operational HITL blocks only (confidence-driven; attestation sign-offs excluded)
- **MTTR (platform-run):** apply.failed → successful retry (D-131)
**Post-Pilot caveat:** computed on N internal runs today; production-denominator activates when a pilot estate runs.
**Benefit:** you now know the zero-touch efficiency is measurable — the pipeline works today on internal runs, and the denominator expands to production estates when a pilot activates.
---
## Slide 13 — Cost & ROI
**Cost estimates and the ROI formula — with honest caveats.**
- **Cost Estimates via Infracost:** pre-apply, grounded (reads plan JSON, offline)
- **ROI formula:** `Platform ROI = (FTE hours saved × blended rate + cloud savings + avoided downtime) ÷ platform op cost`
- **N=0 caveat:** "Computed on N internal runs today; production-denominator activates post-pilot. The formula is grounded; the production numbers are not yet."
- <span class="badge planned">Planned</span>: Live CUR Reconciliation (D-096)
**Benefit:** you now know the ROI formula — and you know it's computed on internal runs today, not fabricated production numbers.
---
## Slide 14 — What's Deferred — and Why
**Honesty about what isn't measured yet.**
**To be clear:** these deferrals are *measurement infrastructure*, not whether the platform runs without humans. The platform IS autonomous in operations. What's deferred is the *evidence pipeline* for certain metrics — not the autonomy itself.
| # | Deferred Metric | Blocking Decision |
|---|----------------|-------------------|
| 1 | Live Infrastructure Health | D-096 |
| 2 | Live Outbox Write Rate | D-096 |
| 3 | Tamper-Evident Ledger Checkpoints | D-083 |
| 4 | Onboarding Funnel (granted) | D-113/D-114/D-119 |
| 5 | Drift Auto-Reversal | D-096 + no scheduler |
| 6 | Live CUR Reconciliation | D-096 |
| 7 | SLA / Unplanned Downtime | D-096 |
| 8 | Predictive vs Reactive | future emitter |
**Benefit:** you now know the boundaries — what Nova measures today, and exactly what blocks the rest. The autonomy is real; the measurement gaps are documented.
---
## Slide 15 — Roadmap to the North Star
**The path from v1.17's grounded metrics to the 1218 month targets.**
- Each deferred metric → blocking decision → unblock requirement → candidate milestone
- Hot-path activation (post-D-096, Nova-native only, D-120)
- Re-evaluation triggers: D-096 lift, D-083 lift, onboarding-grant lift
From `docs/METRICS_DEFERRED_ROADMAP.md`.
**Benefit:** you now know the path — every deferred metric has an unblock requirement and a candidate milestone. Nothing is hand-waved; everything has a plan.
---
## Slide 16 — Recap + Ask
**The 5-act recap + the business decision.**
**Recap:**
- **Problem:** operator is the bottleneck; autonomy in operations, human at stage gates
- **Vision:** invisible operations with provable trust (NORTH_STAR)
- **How:** pipeline + Decision Ledger + 8-concern attestation matrix
- **Proof:** 18V+4S, 100% ledger coverage, 100% attestation, grounded ROI formula
- **Roadmap:** deferred metrics have unblock paths
**The ask:** "Approve a pilot estate to activate the production-denominator metrics (Touchless Resolution, Human Escalation, AI Decision Accuracy), and approve the tamper-evident ledger build-out (D-083 lift) to move from local hash-chain to S3 Object Lock + JWS. These two decisions move Nova from 'pipeline-ready' to 'production-proven.'"
**Benefit:** you leave with a clear business decision to make — approve a pilot + the ledger build-out — and the confidence that every claim in this deck is grounded, derived, or honestly deferred.
---
<!-- _class: title -->
<!-- _paginate: false -->
## Appendix A1 — Metrics Glossary
| KPI | Definition | Status |
|-----|-----------|--------|
| Touchless Resolution Rate | runs without operational HITL block ÷ total | partial (Post-Pilot) |
| Human Escalation Frequency | operational HITL blocks ÷ total | partial (Post-Pilot) |
| AI Decision Accuracy | decisions not followed by failure within 5min | partial (Post-Pilot) |
| MTTR (p95) | apply.failed → successful retry | grounded |
| Confidence-Gate Halt Rate | runs with band=block ÷ total | grounded |
| Provisioning Lead Time | run.completed run.started | grounded |
| Deployment Frequency | count(run.completed) per day | grounded |
| Cost Savings (Infracost) | sum(delta_usd where delta < 0) | partial (CUR deferred) |
| FTE Hours Saved | run count × manual baseline × rate | derived (N=0 caveat) |
| Platform ROI | (labor + cloud + avoided downtime) ÷ op cost | derived (N=0 caveat) |
| Decision Ledger Coverage | decisions with outcome ÷ total | grounded |
| Attestation Coverage | prod/dr attested ÷ total prod/dr | grounded |
| Policy Compliance Rate | 1 failed_assets ÷ total | grounded |
---
<!-- _class: title -->
<!-- _paginate: false -->
## Appendix A2 — Operating Model & Cost
- **Cost figures** from `COST.md`: $0.001883 over 8 days, ~$0.007/month, S3-dominated, zero BAU compute
- **Zero-cost steady state:** all resources torn down post-v1.11 (D-096); the platform runs offline
- References the pre-mortem (`PRE_MORTEM.md`: v1.10 decay root cause + structural mitigations)
**Benefit:** you now know the operating cost is negligible — and the structural mitigation that prevents decay.
@@ -1,113 +0,0 @@
# Nova — The No-Humans Infrastructure Platform: Talking Points
> Step 4 of the 4-step deck process. Presenter cues distilled from the
> source of truth (`nova-no-humans-platform.md`). 3-6 bullets per slide
> + key takeaway. Indexed by Marp slide #.
> v1.17 — REQ-196, REQ-197
---
### Slide 1 — Arc Preview
- Open with the stake line: "18 capabilities verified, 0 consumer estates in production"
- Preview the 5-act arc so the audience knows the structure
- Set the honesty frame: "this is an evidence deck, not a hype deck"
- **Key takeaway:** you'll leave knowing what's proven, what's pipeline-ready, and what's deferred
### Slide 2 — The No-Humans Imperative
- The operator is the bottleneck: days vs. minutes for provisioning
- Key reframing: "no-humans" = no human in normal operations; stage-gate attestation is human by design
- Cite the no-humans thesis doc
- **Key takeaway:** autonomy in operations, human at stage gates
### Slide 3 — Nova's Vision
- Read the vision statement verbatim — it's precise
- Emphasize "provable, not promised" — the difference between marketing and defensible
- State the attestation model up front to prevent mishearing
- **Key takeaway:** invisible operations with provable trust
### Slide 4 — Strategic Objectives + Anti-Goals
- The 4 objectives are the "what"; the 5 anti-goals are the "what NOT"
- Anti-goal #3 (not removing humans from accountability) reinforces slide 3
- Anti-goal #5 (not sold to operators) explains why this deck is for leadership
- **Key takeaway:** purpose-built for infra ops, sold to leadership on outcomes
### Slide 5 — 1218 Month Targets
- The three-section split (current / post-pilot / deferred) IS the honesty model
- "Partial" means the pipeline works but the denominator is zero (0 consumers)
- The Post-Pilot targets are committed; the numbers fill when a pilot runs
- **Key takeaway:** which numbers are real today vs. deferred honestly
### Slide 6 — The Platform Pipeline
- Walk the pipeline left-to-right: contract → resolver → adapter → plan → policy → confidence → gate → apply
- Key insight: dev is autonomous; qa/prod/dr require attestation
- The confidence signal is the "AI" — 6-input weighted score, not an LLM
- **Key takeaway:** the path from intent to evidence, with humans at stage gates only
### Slide 7 — The Decision Ledger
- The D-122 honesty sentence is critical: "Nova's AI is the confidence-gated policy engine, not an LLM"
- The ledger is the moat: features can be copied, an immutable decision history cannot
- Every decision has outcome backfill from apply.completed
- **Key takeaway:** autonomous is defensible because every decision is immutable, queryable, accountable
### Slide 8 — The 8-Concern Attestation Matrix
- The matrix is not a rubber stamp — it's structured, freshness-validated, SoD-enforced
- Offline-testable concerns run for real; operator-supplied concerns accept signed evidence
- SoD on prod: the approver can't be the same person who built it
- **Key takeaway:** autonomy in operations, human in accountability, by design
### Slide 9 — Telemetry Architecture
- Deliberately minimal (Nova-native, no Kafka/Prometheus/ClickHouse)
- Every number in the Proof act is traceable to a file path
- The hot path is deferred (D-126) — cold store is sufficient for batch
- **Key takeaway:** the architecture IS the trust substrate — "where does this number come from?" → file path
### Slide 10 — Capability Health
- 18V+4S is the single most important proof point
- The 4 Skipped are live-AWS caps — honestly skipped (D-096), not broken
- When live AWS is re-provisioned, they reactivate
- **Key takeaway:** the platform works, and we're honest about what we can't test
### Slide 11 — Decision Ledger + Attestation Coverage
- Both 100% — no AI decision is ever lost; no prod/dr promotion lands without a human sign-off
- The trust snapshot has a chain-integrity verdict (the ledger hasn't been tampered with)
- D-083 (S3 Object Lock + JWS) is the next step for the ledger
- **Key takeaway:** trust is provable — not a marketing claim, a queryable record
### Slide 12 — Zero-Touch Efficiency
- The Post-Pilot caveat is the honesty model: pipeline works, denominator is zero
- This is NOT a fabricated "99% touchless" claim
- The numbers fill when a pilot runs
- **Key takeaway:** the measurement works; the numbers activate with a pilot
### Slide 13 — Cost & ROI
- The ROI formula is shown inline — not hidden in a footnote
- The N=0 caveat is stated explicitly
- This is the "no fabrication" constraint in action
- **Key takeaway:** the formula is ready; the production denominator activates with a pilot
### Slide 14 — What's Deferred — and Why
- The preempt is critical: deferrals are measurement infrastructure, not autonomy
- The platform IS autonomous in operations; what's deferred is the evidence pipeline
- Showing this to leadership demonstrates honesty, not weakness
- **Key takeaway:** the autonomy is real; the measurement gaps are documented
### Slide 15 — Roadmap to the North Star
- Every deferred metric has a specific unblock requirement and a candidate milestone
- The re-evaluation triggers ensure the metrics layer evolves
- Nothing is hand-waved; everything has a plan
- **Key takeaway:** the path from "honestly deferred" to "here's how we get there"
### Slide 16 — Recap + Ask
- Recap the 5-act arc so the audience leaves with the structure
- The ask is a business decision: approve a pilot + the ledger build-out
- "Pipeline-ready" → "production-proven" is the value proposition
- **Key takeaway:** approve a pilot + the ledger build-out to move from pipeline-ready to production-proven
### Appendix A1 — Metrics Glossary
- Reference for every metric mentioned in the deck
- Use if the audience asks "what does X mean?"
### Appendix A2 — Operating Model & Cost
- The operating cost is negligible (~$0.007/month)
- The zero-cost steady state (D-096 teardown) is the structural mitigation
- References the pre-mortem for the decay-prevention story
@@ -1,417 +0,0 @@
# Nova — The No-Humans Infrastructure Platform
> **Source of truth** (Step 1 of the 4-step deck process).
> Unified narrative deck merging `how-the-platform-works` + `the-developer-experience`.
> 5-act arc: Problem → Vision → How → Proof → Roadmap.
> x3 structure at deck level (opening = arc preview, body = tell them, closing = recap + ask)
> AND per slide (opens with what it covers, delivers, closes with benefit callout).
> Act indicator in the Marp footer: `Act N/5: <act name>`.
>
> **Honesty model:** every metric cited is grounded (cites a source file),
> derived (documented formula), or deferred (cites a blocking decision ID).
> No fabricated numbers. Deferred metrics marked `<span class="badge planned">Planned</span>`.
>
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (REQ-196, REQ-197)
---
## Slide 1 — Arc Preview (the "what I'm going to tell you" deck-level opening)
This deck proves Nova is the no-humans infrastructure platform — and shows you the metrics that make the claim defensible.
**Today:** 18 capabilities verified, 0 consumer estates in production. This deck shows what's proven, what's pipeline-ready, and what's honestly deferred.
The 5-act arc:
1. **Problem** — why the operator is the bottleneck
2. **Vision** — Nova's strategic direction (NORTH_STAR)
3. **How** — the pipeline, Decision Ledger, attestation gates
4. **Proof** — grounded metrics that make the claim defensible
5. **Roadmap** — deferred metrics with unblock paths + the ask
> **Benefit:** you leave this deck knowing which claims are proven today, which are pipeline-ready, and which are deferred with a documented unblock path — no marketing, just grounded evidence.
> **Speaker notes:** The stake line (18V + 0 consumers) sets the honesty frame. The audience knows from slide 1 that this is not a hype deck — it's an evidence deck. The arc preview orients them for the next 15 slides.
---
## Slide 2 — The No-Humans Imperative
This slide shows why the operator is the bottleneck — and why removing them from operations (not accountability) is the imperative.
- **The cost of humans-in-the-loop:** L1/L2 ops hours, escalation latency, the trust gap (autonomous claims without proof)
- **The operator is the bottleneck:** provisioning takes days, not minutes; escalations pile up; the trust gap means "autonomous" is a marketing claim, not a defensible one
- **The attestation model:** autonomy in operations, human at stage gates — not "no humans ever"
- Cites `docs/NO_HUMANS_THESIS.md` (the thesis, grounded proof, deferred proof, anti-claims)
> **Benefit:** you now know the problem framing — autonomy in operations, human at stage gates, is the path forward.
> **Speaker notes:** The key reframing: "no-humans" means no human in the loop of *normal operations*. Stage-gate attestation (QA for production, SRE for operational readiness) remains human by design. This is not about removing humans from accountability — only from operations.
> **Transition:** "Having defined the problem, here is Nova's strategic direction toward solving it."
---
## Slide 3 — Nova's Vision
This slide states Nova's vision — infrastructure operations become invisible, with provable trust.
> **Infrastructure operations become invisible. Every environment provisioned, every incident healed, every risk remediated — by an autonomous system whose trustworthiness is provable, not promised. Human attestation remains required at stage gates — QA signs off for production, SRE greenlights based on operational readiness — but the operator is never in the loop of normal operations.**
- The attestation model: human attestation required at stage gates (QA for production, SRE for operational readiness); autonomy in operations, not in accountability
- Cites `docs/NO_HUMANS_THESIS.md` (the thesis, grounded proof, deferred proof, anti-claims incl. D-122 honesty)
> **Benefit:** you now know the destination — invisible operations with provable trust, not promised trust. And you know the attestation model: humans at stage gates, not in the ops loop.
> **Speaker notes:** The vision is ambitious but precise. "Provable, not promised" is the key phrase — it's the difference between a marketing claim and a defensible one. The attestation clarification is stated up front so the audience doesn't mishear "no-humans" as "no accountability."
> **Transition:** "The vision is ambitious — here are the 4 strategic objectives that make it concrete."
---
## Slide 4 — Strategic Objectives + Anti-Goals
This slide pairs what Nova is building toward (4 objectives) with what Nova refuses to build (5 anti-goals).
**4 Strategic Objectives:**
1. **Demonstrate production-grade zero-touch operations** — autonomy as the default, not the demo
2. **Establish provable trust in AI decisions** — Decision Ledger, confidence scoring, circuit breakers, blast-radius controls
3. **Deliver compounding, quantifiable ROI** — each quarter must reduce spend, free hours, avoid downtime measurably
4. **Become the default substrate for agentic infrastructure consumption** — the platform AI agents reach for first
**5 Anti-Goals (what Nova is NOT):**
1. Not a Terraform, Kubernetes, or hyperscaler competitor
2. Not a general-purpose AI agent platform
3. Not a system that removes humans from accountability
4. Not for legacy, untagged, or freeform infrastructure
5. Not sold to operators
From `NORTH_STAR.md`.
> **Benefit:** you now know the scope boundaries — Nova is purpose-built for infrastructure operations, sold to leadership on outcomes, and explicitly not a general-purpose AI platform or a hyperscaler competitor.
> **Speaker notes:** The anti-goals are as important as the objectives. They tell the audience what Nova will NOT be distracted by. Anti-goal #3 (not removing humans from accountability) reinforces the attestation model from slide 3.
> **Transition:** "The objectives are committed to measurable targets — here is the 1218 month scorecard, with honest grounding status."
---
## Slide 5 — 1218 Month Targets (the scorecard)
This slide shows the committed targets — numbers a board member can repeat back — with their grounding status.
**Current-milestone targets (grounded or derived this milestone):**
| Domain | Target | Status |
|---|---|---|
| MTTR (p95) | < 60 seconds | grounded (platform-run) |
| Cloud Spend Reduction | ≥ 25% on pilot estates | partial (Infracost grounded; CUR deferred D-096) |
| L1/L2 Ops Hours Avoided | ≥ 70% of pre-Nova FTE | derived (N internal runs; prod activates post-pilot) |
| Platform ROI | ≥ 250% annually | derived (formula; N internal runs caveat) |
| Decision Ledger Coverage | 100% of AI actions | grounded (this milestone builds it) |
| Attestation Coverage | 100% of prod/dr promotions | grounded |
**Post-Pilot targets (pipeline grounded; denominator activates with a pilot estate):**
| Domain | Target | Status |
|---|---|---|
| Touchless Resolution Rate | ≥ 99% | partial (pipeline grounded; 0 consumers today) |
| Human Escalation Frequency | < 0.1% | partial (pipeline grounded; 0 consumers today) |
| AI Decision Accuracy | ≥ 99.5% | partial (pipeline grounded; 0 consumers today) |
**Deferred targets:** Predictive vs Reactive ≥3:1 <span class="badge planned">Planned</span> · Drift Auto-Reversal ≥95% <span class="badge planned">Planned</span>
> **Benefit:** you now know the destination numbers — and which ones are measurable today vs deferred honestly. The Post-Pilot targets are committed; the pipeline works; the numbers fill when a pilot estate runs.
> **Speaker notes:** The three-section split (current / post-pilot / deferred) is the honesty model. The "partial" status means the measurement pipeline is grounded but the denominator is zero (0 consumers). This is the same honesty as Cloud Spend (Infracost grounded, CUR deferred). A board member can see exactly which numbers are real today and which are waiting for a pilot.
> **Transition:** "The targets are committed — here is how Nova works to achieve them."
---
## Slide 6 — The Platform Pipeline
This slide shows the contract-to-evidence pipeline — how intent becomes verified infrastructure without an operator.
```mermaid
graph LR
A[Contract] --> B[Resolver]
B --> C[Adapter]
C --> D[Terraform Plan]
D --> E[Checkov Policy]
E --> F[Confidence Signal]
F --> G{HITL Gate}
G -->|dev: autonomous| H[Apply]
G -->|qa/prod/dr: attested| H
H --> I[Evidence + Outbox]
```
- Contract → resolver → adapter → terraform plan → Checkov (policy) → confidence signal → HITL gate (dev autonomous; qa/prod/dr attested) → apply → evidence
- Grounded in `scripts/run_platform.sh` + `core/contract_resolver.py` + `adapters/terraform/adapter.py` + `core/confidence_signal.py`
> **Benefit:** you now know the path from intent to evidence — and where the human appears (stage gates only, not in the ops loop).
> **Speaker notes:** The pipeline is the engine. The key insight: dev is autonomous (no HITL gate); qa/prod/dr require human attestation. The confidence signal is the "AI" — it's a 6-input weighted score, not an LLM. The HITL gate is where the human appears, but only for qa/prod/dr, not for dev.
> **Transition:** "The pipeline produces decisions — here is how every decision is captured and made accountable."
---
## Slide 7 — The Decision Ledger
This slide shows the Decision Ledger — every AI decision captured with confidence, alternatives, and outcome.
- **Architecture:** `outbox_writer.py` extended → SQLite append-only hash-chain table
- **`ai.decision.made` events:** decision_id=run_id, chosen_action=band, confidence=score, alternatives=perInput, human_override=HITL block, outcome backfilled from apply.completed
- **`attestation.recorded` events:** qa/prod/dr sign-offs (approver, env, concerns, result)
- D-121, D-122, D-132. Honors D-083 (no S3 Object Lock/JWS — local hash-chain this milestone)
**D-122 honesty:** Nova's "AI" is the confidence-gated policy engine (confidence_signal + HITL gate), not an LLM planner. The Decision Ledger captures this real decision path — not a fabricated "AI agent" that doesn't exist yet.
> **Benefit:** you now know why 'autonomous' is defensible — every decision is immutable, queryable, and accountable. And you know exactly what 'AI' means here: a confidence-gated policy engine, not a black-box LLM.
> **Speaker notes:** The D-122 honesty sentence is critical. If the audience walks away thinking Nova has an LLM planner, we've violated the "no fabrication" constraint. The Decision Ledger is the trust substrate (NORTH_STAR Objective #2) — it's the moat. Features can be copied; an immutable, queryable decision history cannot.
> **Transition:** "Decisions are captured — here is how stage-gate attestation keeps humans in accountability."
---
## Slide 8 — The 8-Concern Attestation Matrix
This slide shows the 8-concern attestation matrix — the designed controls that keep humans at stage gates.
| Concern | Env | Freshness | Type |
|---------|-----|-----------|------|
| functional_correctness | qa | 24h | operator-supplied |
| performance_baseline | qa | 7d | operator-supplied |
| security_posture | qa | 24h | operator-supplied |
| contract_nfrs | qa/prod/dr | — | offline-testable |
| operational_readiness | prod | 30d | operator-supplied |
| incident_response | prod | 90d | operator-supplied |
| capacity_cost | prod | 30d | operator-supplied |
| resilience_dr_drill | prod | 180d | operator-supplied |
| resilience_chaos | prod | 90d | operator-supplied |
| resilience_backup | prod | 30d | operator-supplied |
| dr_region_deploy | dr | 180d | operator-supplied |
- Offline-testable concerns run for real; operator-supplied concerns accept signed evidence artifacts
- Separation-of-duties on prod (the approver can't be the same person who built it)
- Grounded in `core/attestation_matrix.py` + `core/hitl_gates.py`
> **Benefit:** you now know the gate model — autonomy in operations, human in accountability, by design. The 8-concern matrix is what makes "no-humans in ops" safe.
> **Speaker notes:** The attestation matrix is the human-in-the-loop safeguard. It's not a rubber stamp — it's a structured, freshness-validated, separation-of-duties-enforced gate. This is what Anti-Goal #3 means: "not a system that removes humans from accountability."
> **Transition:** "You've now seen how Nova works — the pipeline, the Decision Ledger, the attestation gates. But 'how it works' is not 'proof it works.' The next four slides show the measured evidence: capability health, trust metrics, efficiency, and cost — every number grounded in a real file, not a marketing claim."
---
## Slide 9 — Telemetry Architecture
This slide shows how Nova instruments itself — the CloudEvents envelope, the cold store, and the PowerBI export.
```mermaid
graph TB
A[Platform components] --> B[CloudEvents 1.0 envelope]
B --> C[metrics/events.jsonl]
B --> D[metrics/decision_ledger.db]
B --> E[metrics/runs/]
C --> F[Collector]
D --> F
E --> F
F --> G[metrics/nova_metrics.db]
G --> H[metrics/powerbi/]
H --> I[PowerBI dashboards]
```
- Platform components → CloudEvents 1.0 envelope → `metrics/events.jsonl` + `metrics/runs/` + `metrics/decision_ledger.db` → collector → `metrics/nova_metrics.db` (SQLite cold store) → `metrics/powerbi/` (CSV/JSON views) → PowerBI
- D-120 (Nova-native), D-125 (hybrid events/files), D-126 (cold-only)
- <span class="badge planned">Planned</span>: Hot-path (live ops dashboard) — D-126
> **Benefit:** you now know that every metric in this deck is traceable to a real emitted event — the architecture IS the trust substrate. When a CFO asks 'where does this number come from?', the answer is a file path, not a Slack thread.
> **Speaker notes:** The architecture is deliberately minimal (Nova-native, no Kafka/Prometheus/ClickHouse). The hot path is deferred (D-126) — the cold store is sufficient for batch/historical analysis. The key point: every number in the Proof act is traceable to a file path. This is the "no fabrication" constraint made architectural.
> **Transition:** "The architecture is sound — here is the measured proof."
---
## Slide 10 — Capability Health + Confidence Distribution
This slide shows the grounded proof: capability health and confidence distribution from real runs.
**Capability Health:** 18 Verified + 4 Skipped (post-D-096 teardown) from `.ciagent/REGRESSION_REPORT.json`
| Status | Count |
|--------|-------|
| Verified | 18 |
| Skipped | 4 |
| Broken | 0 |
| Decayed | 0 |
- The 4 Skipped are live-AWS capabilities (CAP-013..016) — honestly skipped because resources are torn down (D-096), not a failure
- Confidence distribution: from `metrics/nova_metrics.db` `fact_confidence` — score histogram, band breakdown (pass/halt)
> **Benefit:** you now know the platform is verified — 18 capabilities pass, 4 are honestly skipped, 0 broken. The honesty model (Skipped ≠ failure) is what makes the Verified count credible.
> **Speaker notes:** The 18V+4S number is the single most important proof point. It says "the platform works, and we're honest about what we can't test." The 4 Skipped are live-AWS capabilities — they're skipped because the live AWS resources are torn down (D-096), not because they're broken. When live AWS is re-provisioned, they reactivate.
> **Transition:** "Capability health is necessary — here is the trust substrate that makes autonomy defensible."
---
## Slide 11 — Decision Ledger + Attestation Coverage
This slide shows the trust metrics — Decision Ledger coverage and attestation coverage, both 100%.
- **Decision Ledger Coverage:** 100% of platform runs emit `ai.decision.made` with outcome backfill (source: `metrics/decision_ledger.db`)
- **Attestation Coverage:** 100% of prod/dr promotions attested by a human (source: `hitl_gates.py` + outbox `approver_*` attributes)
- **AI Decision Accuracy:** decisions not followed by apply.failed/incident within 5min
- The trust-snapshot report (`metrics/TRUST_SNAPSHOT.md`) with chain-integrity verdict
- <span class="badge planned">Planned</span>: Tamper-Evident Ledger Checkpoints (D-083)
> **Benefit:** you now know the trust is provable — not a marketing claim, a queryable record. The Decision Ledger is the moat; features can be copied, an immutable decision history cannot.
> **Speaker notes:** The trust metrics are the "provably trustworthy" proof. Decision Ledger Coverage = 100% means no AI decision is ever lost. Attestation Coverage = 100% means no prod/dr promotion lands without a human sign-off. The chain-integrity verdict (from the trust snapshot) proves the ledger hasn't been tampered with.
> **Transition:** "Trust is provable — here is the operational efficiency that makes the ROI real."
---
## Slide 12 — Zero-Touch Efficiency
This slide shows the zero-touch efficiency metrics — touchless resolution, human escalation, and MTTR.
- **Touchless Resolution Rate:** runs without operational HITL block ÷ total (attestation gates excluded)
- **Human Escalation Frequency:** operational HITL blocks only (confidence-driven; attestation sign-offs excluded)
- **MTTR (platform-run):** apply.failed → successful retry (D-131)
**Post-Pilot caveat:** these three metrics are computed on N internal runs today; the production-denominator activates when a pilot estate runs (see NORTH_STAR Post-Pilot Targets section).
> **Benefit:** you now know the zero-touch efficiency is measurable — the pipeline works today on internal runs, and the denominator expands to production estates when a pilot activates.
> **Speaker notes:** The Post-Pilot caveat is the honesty model. The pipeline is grounded (it works); the denominator is zero (0 consumers). This is not a fabricated "99% touchless" claim — it's "the measurement works, and the numbers fill when a pilot runs."
> **Transition:** "Efficiency is half the ROI story — here is the cost side."
---
## Slide 13 — Cost & ROI
This slide shows the cost estimates and the ROI formula — with honest caveats about the current denominator.
- **Cost Estimates via Infracost:** pre-apply, grounded (reads plan JSON, offline)
- **ROI formula (shown inline):** `Platform ROI = (FTE hours saved × blended rate + cloud savings + avoided downtime) ÷ platform op cost`
- **N=0 caveat:** "These derived metrics are computed on N internal runs today; the production-denominator activates post-pilot. The formula is grounded; the production numbers are not yet."
- **FTE Hours Saved** (derived), **Platform ROI** (derived formula)
- <span class="badge planned">Planned</span>: Live CUR Reconciliation (D-096), Drift Auto-Reversal (D-096)
> **Benefit:** you now know the ROI formula — and you know it's computed on internal runs today, not fabricated production numbers. The formula is ready; the production denominator activates with a pilot.
> **Speaker notes:** The ROI formula is shown inline — not hidden in a footnote. The N=0 caveat is stated explicitly. This is the "no fabrication" constraint in action: we show the formula, we show the caveat, we don't pretend the production numbers exist.
> **Transition:** "The proof is grounded — here is what is honestly deferred."
---
## Slide 14 — What's Deferred — and Why
This slide pairs each deferred metric with its blocking decision — honesty about what isn't measured yet.
**To be clear:** these deferrals are *measurement infrastructure*, not whether the platform runs without humans. The platform IS autonomous in operations. What's deferred is the *evidence pipeline* for certain metrics — not the autonomy itself.
| # | Deferred Metric | Blocking Decision |
|---|----------------|-------------------|
| 1 | Live Infrastructure Health | D-096 |
| 2 | Live Outbox Write Rate | D-096 |
| 3 | Tamper-Evident Ledger Checkpoints | D-083 |
| 4 | Onboarding Funnel (granted) | D-113/D-114/D-119 |
| 5 | Drift Auto-Reversal | D-096 + no scheduler |
| 6 | Live CUR Reconciliation | D-096 |
| 7 | SLA / Unplanned Downtime | D-096 |
| 8 | Predictive vs Reactive | future emitter |
From `docs/METRICS_DEFERRED_ROADMAP.md`.
> **Benefit:** you now know the boundaries — what Nova measures today, and exactly what blocks the rest. The autonomy is real; the measurement gaps are documented.
> **Speaker notes:** The preempt is critical: these deferrals are measurement infrastructure, not autonomy. The platform runs without humans in operations. What's deferred is the evidence pipeline for live-infra health, drift detection, predictive remediation — not the autonomy itself. Showing this slide to leadership demonstrates honesty, not weakness.
> **Transition:** "The proof is honest — here is the roadmap from here to the 1218 month targets."
---
## Slide 15 — Roadmap to the North Star
This slide shows the path from v1.17's grounded metrics to the 1218 month targets — the unblock path for each deferred metric.
- Each deferred metric → blocking decision → unblock requirement → candidate milestone
- The hot-path activation section (post-D-096, Nova-native only, D-120)
- Re-evaluation triggers: D-096 lift, D-083 lift, onboarding-grant lift
From `docs/METRICS_DEFERRED_ROADMAP.md`.
> **Benefit:** you now know the path — every deferred metric has an unblock requirement and a candidate milestone. Nothing is hand-waved; everything has a plan.
> **Speaker notes:** The roadmap is the bridge from "honestly deferred" to "here's how we get there." Each deferred metric has a specific unblock requirement and a candidate future milestone. The re-evaluation triggers ensure the metrics layer evolves when the blocking decisions lift.
> **Transition:** "The roadmap is clear — here is the recap and the ask."
---
## Slide 16 — Recap + Ask (the "what I told you" deck-level closing)
This slide recaps the 5 acts and states the ask.
**Recap:**
- **Problem:** the operator is the bottleneck; autonomy in operations, human at stage gates
- **Vision:** invisible operations with provable trust (NORTH_STAR)
- **How:** pipeline + Decision Ledger + 8-concern attestation matrix
- **Proof:** 18V+4S, 100% ledger coverage, 100% attestation, grounded ROI formula
- **Roadmap:** deferred metrics have unblock paths
**The ask:** "The ask is a business decision: approve a pilot estate to activate the production-denominator metrics (Touchless Resolution, Human Escalation, AI Decision Accuracy), and approve the tamper-evident ledger build-out (D-083 lift) to move from local hash-chain to S3 Object Lock + JWS. These two decisions move Nova from 'pipeline-ready' to 'production-proven.'"
> **Benefit:** you leave with a clear business decision to make — approve a pilot + the ledger build-out — and the confidence that every claim in this deck is grounded, derived, or honestly deferred.
> **Speaker notes:** The ask is a business decision, not insider language. "Approve a pilot estate" is something a C-suite can decide. "Approve the ledger build-out" is a budget decision. The recap reinforces the 5-act arc — the audience leaves with the structure, not a pile of facts.
---
## Appendix Slide A1 — Metrics Glossary
This appendix defines every KPI in one line with its grounding badge.
| KPI | Definition | Status |
|-----|-----------|--------|
| Touchless Resolution Rate | runs without operational HITL block ÷ total | partial (Post-Pilot) |
| Human Escalation Frequency | operational HITL blocks ÷ total | partial (Post-Pilot) |
| AI Decision Accuracy | decisions not followed by failure within 5min | partial (Post-Pilot) |
| MTTR (p95) | apply.failed → successful retry | grounded |
| Confidence-Gate Halt Rate | runs with band=block ÷ total | grounded |
| Provisioning Lead Time | run.completed run.started | grounded |
| Deployment Frequency | count(run.completed) per day | grounded |
| Cost Savings (Infracost) | sum(delta_usd where delta < 0) | partial (CUR deferred) |
| FTE Hours Saved | run count × manual baseline × rate | derived (N=0 caveat) |
| Platform ROI | (labor + cloud + avoided downtime) ÷ op cost | derived (N=0 caveat) |
| Decision Ledger Coverage | decisions with outcome ÷ total | grounded |
| Attestation Coverage | prod/dr attested ÷ total prod/dr | grounded |
| Policy Compliance Rate | 1 failed_assets ÷ total | grounded |
> **Benefit:** you now have a reference for every metric mentioned in the deck.
---
## Appendix Slide A2 — Operating Model & Cost
This appendix shows the real cost figures + the zero-cost steady state.
- **Cost figures** from `COST.md`: $0.001883 over 8 days, ~$0.007/month, S3-dominated, zero BAU compute
- **Zero-cost steady state:** all resources torn down post-v1.11 (D-096); the platform runs offline
- References the pre-mortem (`PRE_MORTEM.md`: v1.10 decay root cause + four forward failure modes + structural mitigations)
> **Benefit:** you now know the operating cost is negligible — and the structural mitigation that prevents decay.
---
> **End of deck.** 16 main slides + 2 appendix slides = 18 total.
> Both old decks (`how-the-platform-works` + `the-developer-experience`) are retired (D-130).
+97
View File
@@ -0,0 +1,97 @@
# RACI — Who Owns What
> This page is the citizen-developer-facing copy.
Nova's delivery lifecycle has four roles. This page clarifies who owns
what — so the citizen developer knows what they bring, what the platform
provides, what quality engineering guards, and what is co-owned with SRE.
## The Four Roles
### Citizen Developer (CD)
That's you — the consumer (technical developer L3A or non-technical L3B).
You are **Responsible** for all **Functional Requirements (FRs)** and
**User Acceptance Testing (UAT)**. You produce the FRs + UAT via your AI
coding agent, an upstream agentic SDLC platform, or any upstream
development platform. **The source does not matter** — all are subject to
the same compliance standards (the submission-readiness gate, the
contract schema, the policy envelope, the immutable audit stream). Nova
validates the submission, not the author.
### Platform (Nova)
Nova is **Responsible** for all **Non-Functional Requirements (NFRs)**,
**Infrastructure** (cloud resource lifecycle, state, IAM), and
**Production deployments to cloud** (the apply path, the pipeline, the
release mechanics).
### Quality Engineering (QE)
Quality Engineering is **Responsible** for the platform-side quality
checks: policy enforcement, confidence scoring, schema validation, and
the functional/contract/non-functional evidence that feeds attestation.
QE owns the **quality** of what the platform produces — the gate
evidence, not the gate decision.
### SRE — co-owned with you
Production readiness is **co-owned**. SRE owns operational readiness:
the operational attestation (incident response, capacity, resilience,
DR). The platform performs the QA + SRE attestations agentically (it
runs the confidence signal, the policy checks, the
separation-of-duties). The citizen developer **oversees and triggers**
the actual release — the human attestation at the stage gate is your
authorization. The platform runs the checks; you authorize the promotion.
This is the "autonomy in operations, human at stage gates" model.
## The Matrix
| Work Category | Citizen Developer | Platform | Quality Engineering | SRE |
|---|---|---|---|---|
| **Functional Requirements (FRs)** | **R/A** | C | I | I |
| **User Acceptance Testing (UAT)** | **R/A** | C | I | I |
| **Non-Functional Requirements (NFRs)** | I | **R/A** | C | C |
| **Infrastructure (cloud, state, IAM)** | I | **R/A** | I | C |
| **QA (policy, confidence, schema checks)** | C | R | **R/A** | I |
| **Production deployment to cloud** | I | **R/A** | C | C |
| **Quality attestation (QA sign-off)** | **A** | R | **R** | I |
| **Production readiness (SRE sign-off)** | **A** | R | C | **R** |
**Key:** **R** = Responsible (does the work) · **A** = Accountable (owns
the outcome, sign-off) · **C** = Consulted · **I** = Informed.
## What This Means in Practice
**You (Citizen Developer) bring:**
- Your application code + a contract that declares intent.
- Your FRs (what the application does).
- Your UAT (you accept the deployment when it meets your FRs).
**Nova (Platform) provides:**
- The NFRs (security, observability, compliance — baked into the
pipeline, not your concern).
- The infrastructure (cloud resources, state management, IAM scoping).
- The production deployment (the apply path, the pipeline, the release).
**Quality Engineering guards:**
- The policy enforcement, confidence scoring, schema validation.
- The quality attestation evidence that feeds the stage gates.
**You co-own production readiness with SRE:**
- Nova + SRE run the attestations (QA quality sign-off, SRE operational
readiness).
- You authorize the promotion at the stage gate. No promotion happens
without your recorded attestation.
## Compliance Standards Apply Equally
Your FRs + UAT may come from any source — an AI coding agent, an
agentic SDLC platform, or a traditional IDE. Nova does not
differentiate. All submissions pass through the same gate
(`schemas/submission-readiness.schema.json`): tags, environment
metadata, policy preconditions, profile markers. The compliance
standards are the same regardless of how the code was authored. This is
by design: the audit trail is the same, the policy envelope is the
same, the evidence stream is the same. The source does not matter; the
submission does.
+72
View File
@@ -0,0 +1,72 @@
# Scope — Nova is Downstream of PDLC
> This page is the citizen-developer-facing copy.
## The Boundary
The **Product Development Lifecycle (PDLC)** is **upstream** of Nova. The
PDLC includes:
- Product backlog / roadmap planning
- Code authorship (via AI coding agent, IDE, or agentic SDLC platform)
- Sprint planning / issue tracking
- Application business logic
- IDE workflows / developer experience
Nova never penetrates the PDLC. Nova's domain is **infrastructure +
delivery only**. Nova integrates with externally owned PDLC, SDLC,
Agentic, and Citizen Developer platforms with no regard for the source
of the intent: Nova provides a set of skills and MCP endpoints that help
the developer or AI agent make their application production-grade, and
all intents to deploy to production go through the same rigorous
controls, quality gates, attestation, and evidence stream.
## What Nova Does
Nova governs the downstream half:
- **Contract ingestion** — the validated entry point
- **Submission-readiness gate** — what is acceptable to start
(`schemas/submission-readiness.schema.json`)
- **Policy enforcement** — the confidence signal, Checkov, tagging
- **Cloud resource lifecycle** — Terraform plan/apply, state, IAM
- **Environment progression** — dev (autonomous) → qa (QA attestation) →
prod (SRE attestation) → dr (SRE attestation)
- **Immutable audit + attestation** — the Decision Ledger, the evidence
stream, the HITL gates
## The Integration Point
Integration between the PDLC and Nova is **only** through the validated,
published contract boundary:
```
PDLC (upstream) Nova (downstream)
───────────────── ─────────────────
product backlog contract ingestion
code authorship (AI agent / IDE / SDLC) → submission-readiness gate
sprint planning → policy enforcement
application business logic → cloud resource lifecycle
→ environment progression (dev→qa→prod→dr)
→ immutable audit + attestation
```
The citizen developer's AI coding agent, an upstream agentic SDLC
platform, or any upstream development platform may all produce
submissions. **The source does not matter** — all are subject to the
same compliance standards. Nova validates the submission, not the
author.
## What Nova is Not
- Not an upstream development platform (no product backlogs, IDE, code
authorship).
- Not a general-purpose AI agent platform (autonomy is narrow, bounded
by policy envelopes).
- Not a legacy infrastructure bridge (no VMs/bare metal/OS).
- Not a permissive delivery highway (no escape hatches past confidence
or HITL).
- Not a mutable audit log (VCS history ≠ regulatory evidence).
These anti-goals (from `docs/vision.md` §7 and Core Tenet #2) are
promoted here from buried tenets to an unmissable scope statement.
+88
View File
@@ -0,0 +1,88 @@
# Skills — Production-Grade Guidance for the Citizen Developer
> **Source of truth (v1.18, REQ-221, REQ-222).** The Nova skill catalog
> extends the BA.A 5-skill catalog (web API, worker, scheduled job, static
> asset, basic observability bootstrap) with Atelier-derived production-
> grade engineering principles. Each skill is a markdown file under
> `skills/` keyed to an Atelier domain path.
## How the Citizen Developer's AI Agent Consumes Skills
1. **Before completing a task**, read the relevant skill file(s) that
match the task's domain.
2. **Run `review/agent-checklist.md`** (from Atelier) before finishing —
the checklist items are the gate between "the code is written" and
"the task is done."
3. **Use the Atelier MCP server** (`mcp/atelier/server.py`, P5) for
agentic validation — the `atelier.validate_against_principles` tool
catches correctness/clarity/simplicity/observability gaps that
deterministic scanners (Wiz, Checkmarx, Mend) cannot.
## The 9 Skills
| Skill | Atelier Source | Core Principles | BA.A Mapping |
|---|---|---|---|
| [`api.md`](../skills/api.md) | `domains/api/` | C1, C2, C6 | web API |
| [`security.md`](../skills/security.md) | `domains/security/` | C1 | cross-cutting (all 5) |
| [`data.md`](../skills/data.md) | `domains/data/` | C1, C4, C6 | web API, worker, scheduled job |
| [`testing.md`](../skills/testing.md) | `domains/testing/` | C1, C5 | UAT (citizen-dev RACI) |
| [`observability.md`](../skills/observability.md) | `domains/observability/` | C7 | basic observability bootstrap |
| [`errors.md`](../skills/errors.md) | `domains/errors/` | C1, C7 | web API, worker, scheduled job |
| [`devops.md`](../skills/devops.md) | `domains/devops/` | C5, C7, C8 | scheduled job, worker |
| [`infrastructure-as-code.md`](../skills/infrastructure-as-code.md) | `domains/infrastructure-as-code/` | C1, C5, C8 | static asset |
| [`compliance.md`](../skills/compliance.md) | `domains/compliance/` | C1, C5 | cross-cutting (all 5) |
## Atelier Provenance
The skills are derived from [Atelier](https://example.com/atelier)
— a first-principles docs-as-code engineering framework with 8 core
principles (C1C8) and 19 domains, each with 10 derived P-rules. The
skills distill the citizen-developer-relevant subset of each domain's
first-principles, link to the agent-checklist triggers, and map to the
existing BA.A catalog.
Atelier is vendored under `mcp/atelier/vendor/` (pinned tag) for
audit reproducibility — an agentic validation result is replayable
against the exact principles that produced it.
## The 8 Core Principles (from Atelier)
| # | Principle | One-line |
|---|---|---|
| C1 | Correctness | The system does what it is supposed to do, and nothing else. |
| C2 | Clarity | The intent of the code is obvious to its reader. |
| C3 | Simplicity | The solution is as simple as possible, and no simpler. |
| C4 | Locality | Decisions and their consequences live near each other. |
| C5 | Reversibility | Every decision can be undone, and the cost of undoing is known. |
| C6 | Composability | Parts combine into wholes, and the parts are reusable. |
| C7 | Observability | The system's behavior is visible to those who must understand it. |
| C8 | Economy | The system uses no more resources than the task requires. |
Precedence: C1 > C2 > C3 > C4 > C5 > C6 > C7 > C8. Correctness is never
sacrificed.
## Reference-Only Domains (cited inside skills, not elevated to skill files)
These 4 Atelier domains are relevant to a citizen developer but are cited
inside the 9 skills above rather than getting their own skill file:
- **Performance** (`domains/performance/`) — cited in `observability.md` +
`devops.md` (bounded operations, timeouts, N+1)
- **Documentation** (`domains/documentation/`) — the runbook requirement
(W3.E prod mandatory) is the documentation skill in practice
- **Concurrency** (`domains/concurrency/`) — cited in `errors.md` +
`devops.md` (bounded queues, cancellation, timeout)
- **AI/ML** (`domains/ai-ml/`) — scope: engineering discipline (data
versioning, evaluation, serving, drift), not algorithm design
## Excluded Domains (not relevant to Nova citizen developer)
6 Atelier domains are excluded from the Nova skill catalog (not relevant
to a citizen developer building on Nova's infrastructure platform):
- UI/UX — Nova has no frontend (frontend-engineer deactivated, PERSONAS.md)
- Kubernetes — Nova is AWS-only this milestone (NORTH_STAR Non-Goal #7)
- GitOps + Operators — future roadmap (no GitOps reconciler today)
- Edge — not in scope (Nova is cloud, not edge)
- Messaging — not in scope (Nova deploys infra, not message brokers)
- i18n — application-level concern, not infrastructure
+150
View File
@@ -0,0 +1,150 @@
# Submission Readiness — What is Acceptable to Start
> **Source of truth:** `schemas/submission-readiness.schema.json` (v1.18,
> REQ-217). The validator is `core/submission_readiness.py` (REQ-218),
> invoked as `python3 -m core.lambda.contract_ingestor --check-readiness
> <submission.json>` (D-133).
Nova's submission-readiness gate defines what is **acceptable to start**.
It is a superset gate *above* contract-schema validity: the contract schema
(`schemas/contract.schema.json`) defines the **shape** (id / name /
environment / infrastructure); the readiness schema defines the **gate**
(tags, per-env mandatory metadata, policy preconditions, profile markers,
appSource). Both must pass before ingestion proceeds.
## How It Works
```
citizen developer submits
contract.schema.json validation (shape) ← the existing check
submission-readiness.schema.json (gate) ← the new check
├── contractId present (non-empty)
├── environment valid (dev/qa/prod/dr)
├── tags: all 5 Nova tags present (D-054)
├── policyPreconditions declared
├── profile: developer or agentic
│ └── if agentic: naturalLanguageIntent + confidenceAtSubmission + agentTrace
├── appSource: repo + ref (for runtime fetch)
└── per-env mandatory (W3.E):
dev → stack + environment
qa → + validation.e2eSuite + validation.loadTest
prod → + runbook + dashboard + oncall
dr → + drDrillRef
ready → proceed to contract ingestion
not ready → reject with citizen-developer-facing error (reason code)
```
## Reason Codes
When a submission is not ready, the validator returns one or more reason
codes. These are citizen-developer-facing — no stack traces.
| Code | Meaning |
|---|---|
| `MISSING_TAGS:<tag1>,<tag2>` | One or more required Nova tags are absent |
| `ENV_MISSING_MANDATORY:<env>:<field>` | A per-env mandatory field (W3.E) is missing |
| `AGENTIC_MISSING_INTENT:<marker>` | profile=agentic but a required marker is absent |
| `MISSING_APP_SOURCE` | appSource (repo + ref) is missing |
| `POLICY_PRECONDITION_MISSING` | No policy preconditions declared |
| `CONTRACT_SCHEMA_INVALID:<detail>` | The contract shape failed contract.schema.json |
| `READINESS_SCHEMA_INVALID:<detail>` | The submission failed the readiness schema |
## Good Example
```json
{
"contractId": "uuid-1234",
"id": "webapi",
"name": "Customer Web API",
"environment": "dev",
"tags": {
"nova:owner": "consumer-repo",
"nova:contract": "uuid-1234",
"nova:environment": "dev",
"nova:cost-center": "nova-default",
"nova:ref": "CHG0678912"
},
"policyPreconditions": {
"public-ingress": false,
"encryption_enabled": true,
"deletion_protection": true
},
"profile": "developer",
"appSource": {
"repo": "consumer/web-api",
"ref": "main"
},
"infrastructure": {
"static-assets": {
"inputs": {
"bucket_name": "webapi-assets"
}
}
}
}
```
Result: **READY** — passes the shape + the gate.
## Rejected Examples
### Missing Tags
```json
{
"contractId": "uuid-1234",
"environment": "dev",
"tags": {
"nova:owner": "consumer-repo"
},
"policyPreconditions": {"public-ingress": false},
"profile": "developer",
"appSource": {"repo": "consumer/repo", "ref": "main"}
}
```
Result: `NOT READY — MISSING_TAGS:nova:contract,nova:environment,nova:cost-center,nova:ref`
### Agentic Missing Intent
```json
{
"contractId": "uuid-1234",
"environment": "qa",
"tags": { "nova:owner": "x", "nova:contract": "x", "nova:environment": "qa", "nova:cost-center": "x", "nova:ref": "x" },
"policyPreconditions": {"public-ingress": false},
"profile": "agentic",
"appSource": {"repo": "x", "ref": "x"},
"validation": {"e2eSuite": true, "loadTest": true}
}
```
Result: `NOT READY — AGENTIC_MISSING_INTENT:naturalLanguageIntent; AGENTIC_MISSING_INTENT:confidenceAtSubmission; AGENTIC_MISSING_INTENT:agentTrace`
### Env Missing Mandatory (prod without runbook)
```json
{
"contractId": "uuid-1234",
"environment": "prod",
"tags": { "nova:owner": "x", "nova:contract": "x", "nova:environment": "prod", "nova:cost-center": "x", "nova:ref": "x" },
"policyPreconditions": {"public-ingress": false},
"profile": "developer",
"appSource": {"repo": "x", "ref": "x"}
}
```
Result: `NOT READY — ENV_MISSING_MANDATORY:prod:runbook; ENV_MISSING_MANDATORY:prod:dashboard; ENV_MISSING_MANDATORY:prod:oncall`
## Compliance-Standard Equivalence
The submission-readiness gate applies **equally** to all upstream sources.
Whether the citizen developer's submission originated from an AI coding
agent, an agentic SDLC platform, or a traditional development platform —
the same tags, the same env mandatory, the same policy preconditions, the
same profile markers are required. The source does not matter; the
submission does. This is the RACI compliance-standard equivalence note
(`docs/raci.md`) made machine-checkable.
+1
View File
@@ -0,0 +1 @@
# mcp/atelier — Nova Atelier MCP server package (v1.18)
+95
View File
@@ -0,0 +1,95 @@
# Nova Atelier MCP Server
> exposes Atelier engineering principles to the citizen developer's AI
> agent. Plugin-registry architecture; stdio transport;
> vendored Atelier for audit reproducibility.
## What This Is
The server exposes 4 tools that let a citizen developer's AI coding agent
look up production-grade engineering principles and validate code against
them — agentic validation that goes **beyond deterministic scanners**
(Wiz, Checkmarx, Mend) by catching correctness, clarity, simplicity, and
observability gaps.
## Tools
| Tool | Description |
|---|---|
| `atelier.lookup_principle(domain, principle_id)` | Look up a principle by domain + P-rule ID (e.g., `security`, `P4`). Returns the principle text + the core C-rule it derives from. |
| `atelier.list_domains()` | List the 19 Atelier domains with P-rule counts + Nova-relevance. |
| `atelier.matrix_lookup(domain)` | Look up the domain→core principle mapping for a given domain. |
| `atelier.validate_against_principles(snippet, domains?)` | Validate a code/diff snippet against the Atelier agent-checklist. Returns pass/fail per check item with the principle citation. |
## Architecture — Plugin Registry
```
mcp/atelier/
├── server.py # entrypoint: loads plugins, starts server
├── plugins/
│ ├── __init__.py
│ ├── principles.py # lookup_principle, list_domains, matrix_lookup
│ └── validation.py # validate_against_principles
├── vendor/ # pinned Atelier snapshot
│ ├── VERSION.md # pinned tag + upgrade instructions
│ ├── core/first-principles.md
│ ├── domains/security/first-principles.md
│ ├── review/agent-checklist.md
│ └── matrix/principles-matrix.md
└── README.md # this file
```
Each plugin module exposes `register(mcp) -> None` and calls `@mcp.tool()`
for its tools. `server.py` scans `plugins/` and calls `register` on each.
**Future capabilities drop in as a new plugin file — no `server.py` edits.**
## Running
### With the MCP Python SDK installed
```bash
pip install "mcp[cli]"
python3 -m mcp.atelier.server
```
The server runs over stdio. An MCP client (e.g., the citizen developer's
AI coding agent) spawns it as a subprocess and calls tools via JSON-RPC.
### Without the SDK (fallback / test mode)
The server degrades to a plain-Python tool registry. Tools are callable
directly — this is how tests run without the SDK installed:
```python
from mcp.atelier.server import NovaAtelierServer
s = NovaAtelierServer()
s.load_plugins()
result = s.call_tool("atelier_lookup_principle", {"domain": "security", "principle_id": "P4"})
```
## Vendoring
Atelier is vendored under `vendor/` at a pinned tag (`v0.3.6`, see
`vendor/VERSION.md`). An agentic validation result is only reproducible if
the principles that produced it are pinned. Live-fetch breaks replayability
(Atelier `main` drifts). To upgrade:
```bash
bash scripts/update_atelier_vendor.sh <new-tag>
```
## Extensibility
To add a new tool (e.g., a cost-estimation tool, a policy-as-code
evaluator): create `plugins/<name>.py`, expose `register(mcp)`, and call
`@mcp.tool()` on your function. The server picks it up automatically. No
`server.py` edit. This is the extensibility insurance for future
capabilities.
## Transport
- **Now:** stdio (local agent consumption — the citizen developer's AI
agent spawns the server as a subprocess).
- **Future:** Streamable HTTP (the MCP SDK supports it on the same
`MCPServer` object; adding it is a transport-only change in `server.py`,
not a rewrite).
+1
View File
@@ -0,0 +1 @@
# mcp/atelier package
+1
View File
@@ -0,0 +1 @@
# mcp/atelier/plugins package
+99
View File
@@ -0,0 +1,99 @@
"""mcp/atelier/plugins/principles.py — principle lookup, domain listing, matrix lookup.
Implements 3 MCP tools (REQ-223):
- atelier.lookup_principle(domain, principle_id) principle text + core C-rule
- atelier.list_domains() 19 domains with P-rule counts + Nova-relevance
- atelier.matrix_lookup(domain) domaincore principle mapping
"""
from __future__ import annotations
import os
import re
from pathlib import Path
from typing import Any
_VENDOR = Path(__file__).resolve().parent.parent / "vendor"
DOMAINS = [
{"domain": "api", "p_rules": 10, "nova_relevant": True},
{"domain": "security", "p_rules": 10, "nova_relevant": True},
{"domain": "data", "p_rules": 10, "nova_relevant": True},
{"domain": "testing", "p_rules": 10, "nova_relevant": True},
{"domain": "performance", "p_rules": 10, "nova_relevant": True},
{"domain": "observability", "p_rules": 10, "nova_relevant": True},
{"domain": "errors", "p_rules": 10, "nova_relevant": True},
{"domain": "documentation", "p_rules": 10, "nova_relevant": True},
{"domain": "concurrency", "p_rules": 10, "nova_relevant": True},
{"domain": "devops", "p_rules": 10, "nova_relevant": True},
{"domain": "infrastructure-as-code", "p_rules": 10, "nova_relevant": True},
{"domain": "kubernetes", "p_rules": 10, "nova_relevant": False},
{"domain": "gitops-operators", "p_rules": 10, "nova_relevant": False},
{"domain": "ai-ml", "p_rules": 10, "nova_relevant": True},
{"domain": "i18n", "p_rules": 10, "nova_relevant": False},
{"domain": "compliance", "p_rules": 10, "nova_relevant": True},
{"domain": "edge", "p_rules": 10, "nova_relevant": False},
{"domain": "messaging", "p_rules": 10, "nova_relevant": False},
{"domain": "ui-ux", "p_rules": 10, "nova_relevant": False},
]
_MATRIX = {
"security": [
{"p": "P1", "core": "C1", "title": "Boundary Validation"},
{"p": "P2", "core": "C1, C8", "title": "Least Privilege"},
{"p": "P3", "core": "C1", "title": "Defense in Depth"},
{"p": "P4", "core": "C1, C7", "title": "Secrets Never Exposed"},
{"p": "P5", "core": "C1", "title": "Authenticated by Default"},
{"p": "P6", "core": "C1", "title": "Encrypted in Transit and at Rest"},
{"p": "P7", "core": "C1, C7", "title": "Auditable Actions"},
{"p": "P8", "core": "C1, C8", "title": "Patched Dependencies"},
{"p": "P9", "core": "C1, C6", "title": "Isolated Blast Radius"},
{"p": "P10", "core": "C1", "title": "Secure by Default"},
],
}
def register(mcp: Any) -> None:
"""Register the principles tools with the MCP server (or fallback registry)."""
@mcp.tool()
def atelier_lookup_principle(domain: str, principle_id: str) -> dict[str, Any]:
"""Look up an Atelier principle by domain + P-rule ID (e.g., 'security', 'P4').
Returns the principle title, text, and the core C-rule(s) it derives from.
"""
fp = _VENDOR / "domains" / domain / "first-principles.md"
if not fp.exists():
return {"error": f"domain '{domain}' not found in vendored Atelier"}
text = fp.read_text()
# Parse the P-rule section
pattern = rf"## ({principle_id}\s*—\s*.+?)\n(.+?)(?=\n## |\Z)"
match = re.search(pattern, text, re.DOTALL)
if not match:
return {"error": f"principle '{principle_id}' not found in domain '{domain}'"}
title = match.group(1).strip()
body = match.group(2).strip()
# Find core C-rule from matrix
matrix_entry = next(
(e for e in _MATRIX.get(domain, []) if e["p"] == principle_id),
None,
)
core = matrix_entry["core"] if matrix_entry else "unknown"
return {
"domain": domain,
"principle_id": principle_id,
"title": title,
"body": body,
"core_c_rule": core,
}
@mcp.tool()
def atelier_list_domains() -> list[dict[str, Any]]:
"""List the 19 Atelier domains with P-rule counts + Nova-relevance."""
return DOMAINS
@mcp.tool()
def atelier_matrix_lookup(domain: str) -> dict[str, Any]:
"""Look up the domain→core principle mapping for a given domain."""
if domain not in _MATRIX:
return {"domain": domain, "mapping": [], "note": "full matrix not vendored for this domain; see Atelier live repo"}
return {"domain": domain, "mapping": _MATRIX[domain]}
+78
View File
@@ -0,0 +1,78 @@
"""mcp/atelier/plugins/validation.py — agentic validation against Atelier principles.
Implements 1 MCP tool (REQ-223):
- atelier.validate_against_principles(snippet, domains) pass/fail per
checklist item with the principle citation. This is the agentic
validation BEYOND deterministic scanners (Wiz/Checkmarx/Mend) it
catches correctness/clarity/simplicity/observability gaps that
deterministic tools cannot.
"""
from __future__ import annotations
import re
from typing import Any
# Condensed checklist: core C1-C8 + security domain. Each item is a
# (check_id, description, heuristic_pattern, principle_citation).
_CHECKLIST = [
# C1 Correctness
{"id": "C1.1", "desc": "Does the code do what the task asked, completely?", "heuristic": r"TODO|FIXME|pass\s*$", "principle": "C1 Correctness", "neg": True},
{"id": "C1.2", "desc": "Does it handle failure cases? (errors, timeouts)", "heuristic": r"except\s*:?\s*pass", "principle": "C1 Correctness", "neg": True},
{"id": "C1.3", "desc": "Is there a test that would fail if the code were wrong?", "heuristic": r"def test_|describe\(", "principle": "C1 Correctness", "neg": False, "optional": True},
# C2 Clarity
{"id": "C2.1", "desc": "Are names intent-revealing? (no 'data', 'temp', 'x')", "heuristic": r"\b(data|temp|x|foo|bar|doStuff)\b", "principle": "C2 Clarity", "neg": True},
# C3 Simplicity
{"id": "C3.1", "desc": "Is there dead code? (unreachable branches)", "heuristic": r"return\s+\w+\s*$.*return", "principle": "C3 Simplicity", "neg": True, "multiline": True},
# C7 Observability
{"id": "C7.1", "desc": "Are there logs for significant events?", "heuristic": r"log(ger|ging)?|print\(|console\.", "principle": "C7 Observability", "neg": False, "optional": True},
{"id": "C7.2", "desc": "Are there secrets in logs?", "heuristic": r"password|secret|token|api_key", "principle": "C7 Observability + Security P4", "neg": True},
# Security
{"id": "SEC.1", "desc": "No secrets in code/logs/URLs", "heuristic": r"(password|secret|token|api_key)\s*=\s*['\"]", "principle": "Security P4 Secrets Never Exposed", "neg": True},
{"id": "SEC.2", "desc": "Input validated at the boundary", "heuristic": r"validate|schema|assert", "principle": "Security P1 Boundary Validation", "neg": False, "optional": True},
{"id": "SEC.3", "desc": "Authorization checked, not assumed", "heuristic": r"auth|permission|rbac|authorize", "principle": "Security P5 Authenticated by Default", "neg": False, "optional": True},
]
def register(mcp: Any) -> None:
"""Register the validation tools with the MCP server (or fallback registry)."""
@mcp.tool()
def atelier_validate_against_principles(snippet: str, domains: list[str] | None = None) -> dict[str, Any]:
"""Validate a code/diff snippet against Atelier principles.
Runs the agent-checklist items against the snippet and returns
pass/fail per item with the principle citation. This is the
agentic validation BEYOND deterministic scanners (Wiz/Checkmarx/
Mend) it catches correctness/clarity/simplicity/observability
gaps that deterministic tools cannot.
Args:
snippet: The code or diff text to validate.
domains: Optional list of domains to include (default: core + security).
"""
results: list[dict[str, Any]] = []
for check in _CHECKLIST:
pattern = check["heuristic"]
flags = re.DOTALL if check.get("multiline") else 0
found = bool(re.search(pattern, snippet, flags))
# neg=True means finding the pattern is a FAIL; neg=False means finding is a PASS
if check.get("neg"):
status = "FAIL" if found else "PASS"
else:
if check.get("optional"):
status = "PASS" if found else "WARN"
else:
status = "PASS" if found else "WARN"
results.append({
"check_id": check["id"],
"description": check["desc"],
"status": status,
"principle": check["principle"],
})
all_pass = all(r["status"] == "PASS" for r in results)
return {
"overall": "PASS" if all_pass else "FAIL",
"results": results,
"domains_checked": domains or ["core", "security"],
"note": "Agentic validation beyond Wiz/Checkmarx/Mend — catches correctness, clarity, simplicity, observability gaps.",
}
+142
View File
@@ -0,0 +1,142 @@
"""mcp/atelier/server.py — Nova Atelier MCP server (REQ-223, D-135, D-137, D-140).
Plugin-registry architecture (D-140): plugins/<name>.py modules each expose
``register(mcp) -> None`` and call ``@mcp.tool()`` for their tools. This file
scans ``plugins/`` and calls ``register`` on each. Future capabilities drop
in as new plugin files no server.py edits.
Transport: stdio (D-135). The MCP Python SDK v2 (``modelcontextprotocol/
python-sdk``, D-137) is the target. If the SDK is not installed, the server
degrades to a plain-Python tool registry that can be tested directly the
tools are callable without MCP. This makes the server testable in CI
without the SDK installed.
Usage (with SDK):
python3 -m mcp.atelier.server
Usage (without SDK, for testing):
from mcp.atelier.server import NovaAtelierServer
s = NovaAtelierServer()
s.load_plugins()
result = s.call_tool("atelier.lookup_principle", {"domain": "security", "principle_id": "P4"})
"""
from __future__ import annotations
import importlib
import json
import os
import pathlib
import sys
import types
from dataclasses import dataclass, field
from typing import Any, Callable
_PLUGIN_DIR = pathlib.Path(__file__).parent / "plugins"
_VENDOR_DIR = pathlib.Path(__file__).parent / "vendor"
class _ToolRegistry:
"""A minimal tool registry that mimics the MCP ``@mcp.tool()`` decorator.
When the MCP SDK is available, ``NovaAtelierServer`` wraps a real
``MCPServer`` and the decorator registers tools with the SDK. When the
SDK is absent, this registry is the fallback tools are callable via
``call_tool()`` for testing.
"""
def __init__(self) -> None:
self._tools: dict[str, dict[str, Any]] = {}
def tool(self, name: str | None = None, description: str | None = None) -> Callable:
def decorator(fn: Callable) -> Callable:
tool_name = name or fn.__name__
self._tools[tool_name] = {
"fn": fn,
"description": description or fn.__doc__ or "",
"name": tool_name,
}
return fn
return decorator
def list_tools(self) -> list[dict[str, str]]:
return [{"name": t["name"], "description": t["description"]} for t in self._tools.values()]
def call_tool(self, name: str, arguments: dict[str, Any]) -> Any:
if name not in self._tools:
raise KeyError(f"Unknown tool: {name}")
return self._tools[name]["fn"](**arguments)
class NovaAtelierServer:
"""The Nova Atelier MCP server.
Wraps an MCP SDK ``MCPServer`` if available; otherwise uses the
``_ToolRegistry`` fallback. Plugins are loaded from ``plugins/``.
"""
def __init__(self) -> None:
self.registry = _ToolRegistry()
self._mcp = None
try:
from mcp.server import MCPServer # type: ignore[import-not-found]
self._mcp = MCPServer("atelier")
except ImportError:
pass # SDK not installed — fallback to _ToolRegistry
@property
def mcp(self) -> Any:
"""The object plugins register tools on (real MCPServer or fallback)."""
return self._mcp if self._mcp is not None else self.registry
def load_plugins(self) -> list[str]:
"""Scan plugins/ and call ``register(mcp)`` on each. Returns loaded names."""
loaded: list[str] = []
for p in sorted(_PLUGIN_DIR.glob("*.py")):
if p.stem == "__init__":
continue
mod_name = f"mcp.atelier.plugins.{p.stem}"
mod = importlib.import_module(mod_name)
if hasattr(mod, "register"):
mod.register(self.mcp if self._mcp else self.registry)
loaded.append(p.stem)
return loaded
def list_tools(self) -> list[dict[str, str]]:
if self._mcp is not None:
return [{"name": t.name, "description": t.description} for t in self._mcp._tools.values()] # type: ignore[attr-defined]
return self.registry.list_tools()
def call_tool(self, name: str, arguments: dict[str, Any]) -> Any:
if self._mcp is not None:
raise RuntimeError("MCP SDK call_tool not supported in fallback mode — use the MCP client")
return self.registry.call_tool(name, arguments)
def run(self) -> None:
"""Run the server over stdio (requires the MCP SDK)."""
if self._mcp is None:
raise RuntimeError("MCP SDK not installed — cannot run server. Install: pip install mcp")
self._mcp.run()
def _make_plugin_compat_decorator(registry_or_mcp: Any) -> Callable:
"""Return a ``tool()`` decorator that works for both the fallback
registry and the real MCP SDK."""
if hasattr(registry_or_mcp, "tool"):
return registry_or_mcp.tool
# Fallback: wrap registry.tool() as a decorator factory
return registry_or_mcp.tool
def main() -> None:
server = NovaAtelierServer()
loaded = server.load_plugins()
print(f"Atelier MCP server — {len(loaded)} plugins loaded: {', '.join(loaded)}", file=sys.stderr)
if server._mcp is None:
print("MCP SDK not installed — server is in fallback (test) mode.", file=sys.stderr)
print("Tools: " + ", ".join(t["name"] for t in server.list_tools()), file=sys.stderr)
else:
server.run()
if __name__ == "__main__":
main()
+21
View File
@@ -0,0 +1,21 @@
# Vendored Atelier — Version Pin
> **Pinned tag:** `v0.3.6` (the v0.4 milestone release, 2026-08-05)
> **Commit:** `666b137dbb3c00e81f8740d18b639bc67587d29f`
> **P-rule count:** 190 (19 domains × 10 P-rules)
> **Vendor date:** 2026-08-06
> **Vendor reason:** audit reproducibility (D-136) — an agentic validation
> result is only replayable if the principles that produced it are pinned.
## Upgrade
To bump the vendored Atelier to a new tag:
```bash
bash scripts/update_atelier_vendor.sh <new-tag>
```
The script fetches the Atelier repo at the given tag, replaces
`mcp/atelier/vendor/`, updates this VERSION.md, and commits the change.
Upgrades are **intentional** — never automatic. Atelier `main` is a
moving target; pinning is required for audit reproducibility.
+28
View File
@@ -0,0 +1,28 @@
# Core First Principles
The 8 universal axioms. Every domain principle derives from one or more
of these. Precedence: C1 > C2 > C3 > C4 > C5 > C6 > C7 > C8.
## C1 — Correctness
The system does what it is supposed to do, and nothing else.
## C2 — Clarity
The intent of the code is obvious to its reader.
## C3 — Simplicity
The solution is as simple as possible, and no simpler.
## C4 — Locality
Decisions and their consequences live near each other.
## C5 — Reversibility
Every decision can be undone, and the cost of undoing is known.
## C6 — Composability
Parts combine into wholes, and the parts are reusable.
## C7 — Observability
The system's behavior is visible to those who must understand it.
## C8 — Economy
The system uses no more resources than the task requires.
+31
View File
@@ -0,0 +1,31 @@
# Security — First Principles
## P1 — Boundary Validation
All input is validated at the trust boundary. (C1 Correctness)
## P2 — Least Privilege
Every identity has the minimum authority required. (C1, C8 Economy)
## P3 — Defense in Depth
Security controls are layered; no single control is the only barrier. (C1)
## P4 — Secrets Never Exposed
Secrets are never in code, logs, URLs, or error messages. (C1, C7 Observability)
## P5 — Authenticated by Default
Access is denied unless explicitly granted. (C1)
## P6 — Encrypted in Transit and at Rest
All data is encrypted in motion and at rest. (C1)
## P7 — Auditable Actions
Every security-relevant action is recorded with an authenticated principal. (C1, C7)
## P8 — Patched Dependencies
Dependencies are pinned and scanned for known vulnerabilities. (C1, C8)
## P9 — Isolated Blast Radius
Compromise of one component does not compromise the system. (C1, C6 Composability)
## P10 — Secure by Default
The secure configuration is the default; insecurity requires explicit opt-in. (C1)
+19
View File
@@ -0,0 +1,19 @@
# Principles Matrix (Vendored Stub)
Maps every domain P-rule back to the core C-rule(s) it derives from.
Full matrix in the live Atelier repo; this is a condensed vendored version
for the security domain (the primary domain the MCP server validates
against in v1.18).
| Domain | P-rule | Core C-rule(s) |
|---|---|---|
| security | P1 Boundary Validation | C1 Correctness |
| security | P2 Least Privilege | C1, C8 Economy |
| security | P3 Defense in Depth | C1 |
| security | P4 Secrets Never Exposed | C1, C7 Observability |
| security | P5 Authenticated by Default | C1 |
| security | P6 Encrypted in Transit and at Rest | C1 |
| security | P7 Auditable Actions | C1, C7 |
| security | P8 Patched Dependencies | C1, C8 |
| security | P9 Isolated Blast Radius | C1, C6 Composability |
| security | P10 Secure by Default | C1 |
+50
View File
@@ -0,0 +1,50 @@
# Agent Pre-Completion Checklist (Vendored)
Every AI agent runs this checklist before completing a task.
## Core Principles Checklist (C1C8)
### C1 Correctness
- Does the code do what the task asked, completely?
- Does it handle the specified edge cases? (nulls, empties, max, min)
- Does it handle the failure cases? (errors, timeouts, invalid input)
- Is there a test that would fail if the code were wrong?
### C2 Clarity
- Can a stranger read this and understand it without asking you?
- Are names intent-revealing? (No `data`, `temp`, `x`, `doStuff`)
- Do comments explain *why*, not *what*?
### C3 Simplicity
- Is this the simplest solution that is complete?
- Is there dead code? (Unreachable branches, unused variables)
- Is there premature abstraction? (An interface with one implementation)
### C4 Locality
- Does related logic live together?
- Are side effects near their causes?
### C5 Reversibility
- Is this change undoable? (migration has a `down`, deploy has a rollback)
- Did I avoid irreversible actions without explicit confirmation?
### C6 Composability
- Does this component/function do one thing?
- Is the boundary (props/args/return) explicit and typed?
### C7 Observability
- Are there logs for significant events?
- Do errors carry enough context to debug? (request ID, user, action)
- Are there no secrets in logs?
### C8 Economy
- Is memory bounded? (No unbounded growth, no loading everything)
- Is time bounded? (No N+1, no blocking without timeout)
## Domain-Specific (Security)
- No secrets in code, logs, URLs, or error messages
- Input is validated at the boundary
- Output is encoded for its context
- Crypto uses vetted libraries (no MD5/SHA1 for security)
- Authorization is checked, not assumed
-1
View File
@@ -1,6 +1,5 @@
# Nova Metrics Directory
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (D-128)
This directory holds Nova's telemetry/observability artifacts. The
metrics layer is **Nova-native** (D-120): JSONL event log + SQLite cold
-2
View File
@@ -1,7 +1,5 @@
# Nova PowerBI Dashboard — Import Guide
> v1.17 — Strategic Direction, Leadership Metrics & Unified Story (REQ-208)
> Generated: 2026-08-04
This guide documents how to import Nova's metrics views into PowerBI
via the folder connector, and suggests a starter visual model.
+34 -7
View File
@@ -48,6 +48,11 @@
"description": "Target group target type (ip or instance).",
"required": false,
"default": "ip"
},
"enabled": {
"type": "boolean",
"default": true,
"description": "Feature flag: enable/disable this module. Set to false to skip resource creation."
}
},
"outputs": {
@@ -85,20 +90,42 @@
{
"type": "aws:elbv2:loadbalancer",
"description": "Application load balancer in the VPC subnets.",
"inputs": ["name", "subnets", "security_group", "load_balancer_type"],
"outputs": ["lb_arn"]
"inputs": [
"name",
"subnets",
"security_group",
"load_balancer_type"
],
"outputs": [
"lb_arn"
]
},
{
"type": "aws:elbv2:targetgroup",
"description": "Target group for the ECS service tasks.",
"inputs": ["name", "port", "protocol", "vpc_id", "target_type"],
"outputs": ["target_group_arn"]
"inputs": [
"name",
"port",
"protocol",
"vpc_id",
"target_type"
],
"outputs": [
"target_group_arn"
]
},
{
"type": "aws:elbv2:listener",
"description": "Listener forwarding the LB port to the target group.",
"inputs": ["lb_arn", "port", "protocol", "target_group_arn"],
"outputs": ["listener_arn"]
"inputs": [
"lb_arn",
"port",
"protocol",
"target_group_arn"
],
"outputs": [
"listener_arn"
]
}
]
}
}
+5 -2
View File
@@ -1,4 +1,5 @@
resource "aws_lb" "this" {
count = var.enabled ? 1 : 0
name = var.name
load_balancer_type = var.load_balancer_type
subnets = local.subnet_list
@@ -6,6 +7,7 @@ resource "aws_lb" "this" {
}
resource "aws_lb_target_group" "this" {
count = var.enabled ? 1 : 0
name_prefix = "${var.name}-"
port = var.port
protocol = var.protocol
@@ -18,13 +20,14 @@ resource "aws_lb_target_group" "this" {
}
resource "aws_lb_listener" "this" {
load_balancer_arn = aws_lb.this.id
count = var.enabled ? 1 : 0
load_balancer_arn = aws_lb.this[0].id
port = var.port
protocol = var.protocol
default_action {
type = "forward"
target_group_arn = aws_lb_target_group.this.arn
target_group_arn = aws_lb_target_group.this[0].arn
}
depends_on = [aws_lb_target_group.this]
+3 -3
View File
@@ -1,14 +1,14 @@
output "lb_arn" {
value = aws_lb.this.id
value = aws_lb.this[0].id
description = "The load balancer ARN."
}
output "listener_arn" {
value = aws_lb_listener.this.arn
value = aws_lb_listener.this[0].arn
description = "The listener ARN."
}
output "target_group_arn" {
value = aws_lb_target_group.this.arn
value = aws_lb_target_group.this[0].arn
description = "The target group ARN."
}
+6
View File
@@ -50,3 +50,9 @@ variable "vpc_id" {
description = "VPC ID for the target group (ref to vpc or platform VPC)."
default = null
}
variable "enabled" {
type = bool
description = "Feature flag: enable/disable this module. Set to false to skip resource creation."
default = true
}
+31 -6
View File
@@ -43,6 +43,11 @@
"type": "string",
"description": "AWS region (CloudFront is global but the provider region is used for the OAC).",
"required": true
},
"enabled": {
"type": "boolean",
"default": true,
"description": "Feature flag: enable/disable this module. Set to false to skip resource creation."
}
},
"outputs": {
@@ -75,17 +80,37 @@
{
"type": "aws:cloudfront:distribution",
"description": "CloudFront distribution with S3 origin via OAC.",
"inputs": ["bucket_regional_domain_name", "price_class", "viewer_protocol_policy", "default_ttl", "max_ttl", "waf_web_acl_arn", "oac_id"],
"outputs": ["distribution_arn", "distribution_domain_name"]
"inputs": [
"bucket_regional_domain_name",
"price_class",
"viewer_protocol_policy",
"default_ttl",
"max_ttl",
"waf_web_acl_arn",
"oac_id"
],
"outputs": [
"distribution_arn",
"distribution_domain_name"
]
},
{
"type": "aws:cloudfront:originaccesscontrol",
"description": "Origin Access Control for the S3 origin.",
"inputs": ["name", "origin_type", "signing_behavior"],
"outputs": ["oac_id"]
"inputs": [
"name",
"origin_type",
"signing_behavior"
],
"outputs": [
"oac_id"
]
}
],
"intra_refs": [
{"from": "aws:cloudfront:distribution.oac_id", "to": "aws:cloudfront:originaccesscontrol.oac_id"}
{
"from": "aws:cloudfront:distribution.oac_id",
"to": "aws:cloudfront:originaccesscontrol.oac_id"
}
]
}
}
+3 -1
View File
@@ -1,4 +1,5 @@
resource "aws_cloudfront_origin_access_control" "this" {
count = var.enabled ? 1 : 0
name = local.oac_name
origin_access_control_origin_type = local.oac_origin_type
signing_behavior = local.oac_signing_behavior
@@ -6,10 +7,11 @@ resource "aws_cloudfront_origin_access_control" "this" {
}
resource "aws_cloudfront_distribution" "this" {
count = var.enabled ? 1 : 0
origin {
origin_id = "s3-origin"
domain_name = var.bucket_regional_domain_name
origin_access_control_id = aws_cloudfront_origin_access_control.this.id
origin_access_control_id = aws_cloudfront_origin_access_control.this[0].id
s3_origin_config {
origin_access_identity = ""
}
+3 -3
View File
@@ -1,14 +1,14 @@
output "distribution_arn" {
value = aws_cloudfront_distribution.this.arn
value = aws_cloudfront_distribution.this[0].arn
description = "The CloudFront distribution ARN."
}
output "distribution_domain_name" {
value = aws_cloudfront_distribution.this.domain_name
value = aws_cloudfront_distribution.this[0].domain_name
description = "The CloudFront distribution domain name."
}
output "oac_id" {
value = aws_cloudfront_origin_access_control.this.id
value = aws_cloudfront_origin_access_control.this[0].id
description = "The Origin Access Control ID."
}
@@ -38,3 +38,9 @@ variable "region" {
description = "AWS region (CloudFront is global but the provider region is used for the OAC)."
default = null
}
variable "enabled" {
type = bool
description = "Feature flag: enable/disable this module. Set to false to skip resource creation."
default = true
}
+6 -1
View File
@@ -19,6 +19,11 @@
"type": "string",
"description": "ARN of the CMK for repository encryption; if absent, uses AWS-managed key.",
"required": false
},
"enabled": {
"type": "boolean",
"default": true,
"description": "Feature flag: enable/disable this module. Set to false to skip resource creation."
}
},
"outputs": {
@@ -48,4 +53,4 @@
"default": true
}
}
}
}
+1
View File
@@ -6,6 +6,7 @@ locals {
}
resource "aws_ecr_repository" "this" {
count = var.enabled ? 1 : 0
name = var.name
image_tag_mutability = "MUTABLE"
image_scanning_configuration {
+2 -2
View File
@@ -1,9 +1,9 @@
output "repository_url" {
value = aws_ecr_repository.this.repository_url
value = aws_ecr_repository.this[0].repository_url
description = "The ECR repository URL."
}
output "repository_arn" {
value = aws_ecr_repository.this.arn
value = aws_ecr_repository.this[0].arn
description = "The ECR repository ARN."
}
+7 -1
View File
@@ -13,4 +13,10 @@ variable "kms_key_arn" {
type = string
description = "ARN of the CMK for ECR encryption; if absent, uses managed key."
default = null
}
}
variable "enabled" {
type = bool
description = "Feature flag: enable/disable this module. Set to false to skip resource creation."
default = true
}
+6 -1
View File
@@ -19,6 +19,11 @@
"type": "string",
"description": "ARN of the CMK for CloudWatch log group encryption; if absent, uses managed key.",
"required": false
},
"enabled": {
"type": "boolean",
"default": true,
"description": "Feature flag: enable/disable this module. Set to false to skip resource creation."
}
},
"outputs": {
@@ -43,4 +48,4 @@
"default": true
}
}
}
}
+1
View File
@@ -1,3 +1,4 @@
resource "aws_ecs_cluster" "this" {
count = var.enabled ? 1 : 0
name = var.name
}
+2 -2
View File
@@ -1,9 +1,9 @@
output "cluster_arn" {
value = aws_ecs_cluster.this.arn
value = aws_ecs_cluster.this[0].arn
description = "The ECS cluster ARN."
}
output "cluster_id" {
value = aws_ecs_cluster.this.id
value = aws_ecs_cluster.this[0].id
description = "The ECS cluster ID."
}
@@ -15,3 +15,9 @@ variable "kms_key_arn" {
description = "ARN of the CMK for CloudWatch log group encryption; if absent, uses managed key."
default = null
}
variable "enabled" {
type = bool
description = "Feature flag: enable/disable this module. Set to false to skip resource creation."
default = true
}
+28 -5
View File
@@ -79,6 +79,11 @@
"description": "ECS task definition family name.",
"required": false,
"default": "app"
},
"enabled": {
"type": "boolean",
"default": true,
"description": "Feature flag: enable/disable this module. Set to false to skip resource creation."
}
},
"outputs": {
@@ -107,14 +112,32 @@
{
"type": "aws:ecs:task_definition",
"description": "Fargate task definition; the adapter jsonencodes image/port/env into container_definitions.",
"inputs": ["image", "port", "cpu", "memory", "env", "family"],
"outputs": ["task_def_arn"]
"inputs": [
"image",
"port",
"cpu",
"memory",
"env",
"family"
],
"outputs": [
"task_def_arn"
]
},
{
"type": "aws:ecs:service",
"description": "Fargate service running the task definition in the cluster + subnets.",
"inputs": ["cluster_arn", "subnets", "security_group", "lb_target_group_arn", "desired_count", "launch_type"],
"outputs": ["service_arn"]
"inputs": [
"cluster_arn",
"subnets",
"security_group",
"lb_target_group_arn",
"desired_count",
"launch_type"
],
"outputs": [
"service_arn"
]
}
]
}
}
+3 -1
View File
@@ -1,4 +1,5 @@
resource "aws_ecs_task_definition" "this" {
count = var.enabled ? 1 : 0
family = var.family
cpu = tostring(var.cpu)
memory = tostring(var.memory)
@@ -8,9 +9,10 @@ resource "aws_ecs_task_definition" "this" {
}
resource "aws_ecs_service" "this" {
count = var.enabled ? 1 : 0
name = "nova-microservice"
cluster = var.cluster_arn
task_definition = aws_ecs_task_definition.this.arn
task_definition = aws_ecs_task_definition.this[0].arn
desired_count = var.desired_count
launch_type = var.launch_type
+2 -2
View File
@@ -1,9 +1,9 @@
output "service_arn" {
value = aws_ecs_service.this.id
value = aws_ecs_service.this[0].id
description = "The ECS service ARN."
}
output "task_def_arn" {
value = aws_ecs_task_definition.this.arn
value = aws_ecs_task_definition.this[0].arn
description = "The ECS task definition ARN."
}
@@ -78,3 +78,9 @@ variable "family" {
description = "ECS task definition family name."
default = "app"
}
variable "enabled" {
type = bool
description = "Feature flag: enable/disable this module. Set to false to skip resource creation."
default = true
}
+6 -1
View File
@@ -24,6 +24,11 @@
"type": "string",
"description": "AWS region the role is created in.",
"required": true
},
"enabled": {
"type": "boolean",
"default": true,
"description": "Feature flag: enable/disable this module. Set to false to skip resource creation."
}
},
"outputs": {
@@ -48,4 +53,4 @@
"default": true
}
}
}
}
+3 -2
View File
@@ -1,11 +1,12 @@
resource "aws_iam_role" "this" {
count = var.enabled ? 1 : 0
name = var.role_name
assume_role_policy = local.assume_role_policy
}
resource "aws_iam_role_policy" "ecr_logs" {
count = local.inline_policy != null ? 1 : 0
count = (local.inline_policy != null && var.enabled) ? 1 : 0
name = local.inline_policy.name
role = aws_iam_role.this.id
role = aws_iam_role.this[0].id
policy = local.inline_policy.policy
}
+2 -2
View File
@@ -1,9 +1,9 @@
output "role_arn" {
value = aws_iam_role.this.arn
value = aws_iam_role.this[0].arn
description = "The IAM role ARN."
}
output "role_id" {
value = aws_iam_role.this.id
value = aws_iam_role.this[0].id
description = "The IAM role ID."
}
@@ -21,3 +21,9 @@ variable "region" {
description = "AWS region (provider-level; not a resource arg)."
default = null
}
variable "enabled" {
type = bool
description = "Feature flag: enable/disable this module. Set to false to skip resource creation."
default = true
}
+6 -1
View File
@@ -20,6 +20,11 @@
"description": "Number of days before the key is deleted after deletion is requested (default 30).",
"required": false,
"default": 30
},
"enabled": {
"type": "boolean",
"default": true,
"description": "Feature flag: enable/disable this module. Set to false to skip resource creation."
}
},
"outputs": {
@@ -49,4 +54,4 @@
"default": true
}
}
}
}
+3 -1
View File
@@ -1,10 +1,12 @@
resource "aws_kms_key" "this" {
count = var.enabled ? 1 : 0
description = var.description
enable_key_rotation = true
deletion_window_in_days = var.deletion_window_days
}
resource "aws_kms_alias" "this" {
count = var.enabled ? 1 : 0
name = local.alias_name
target_key_id = aws_kms_key.this.key_id
target_key_id = aws_kms_key.this[0].key_id
}
+2 -2
View File
@@ -1,9 +1,9 @@
output "kms_key_arn" {
value = aws_kms_key.this.arn
value = aws_kms_key.this[0].arn
description = "The KMS key ARN."
}
output "kms_key_id" {
value = aws_kms_key.this.key_id
value = aws_kms_key.this[0].key_id
description = "The KMS key ID."
}
+7 -1
View File
@@ -14,4 +14,10 @@ variable "deletion_window_days" {
type = number
description = "Deletion window in days (7-30)."
default = 30
}
}
variable "enabled" {
type = bool
description = "Feature flag: enable/disable this module. Set to false to skip resource creation."
default = true
}
+5
View File
@@ -79,6 +79,11 @@
"description": "Database admin password",
"required": false,
"default": "ACdlcI2026!"
},
"enabled": {
"type": "boolean",
"default": true,
"description": "Feature flag: enable/disable this module. Set to false to skip resource creation."
}
},
"outputs": {
+1
View File
@@ -5,6 +5,7 @@ resource "aws_db_subnet_group" "this" {
}
resource "aws_db_instance" "this" {
count = var.enabled ? 1 : 0
engine = var.engine
engine_version = var.engine_version
instance_class = var.instance_class
+2 -2
View File
@@ -1,9 +1,9 @@
output "db_endpoint" {
value = aws_db_instance.this.endpoint
value = aws_db_instance.this[0].endpoint
description = "The RDS instance endpoint."
}
output "db_arn" {
value = aws_db_instance.this.arn
value = aws_db_instance.this[0].arn
description = "The RDS instance ARN."
}

Some files were not shown because too many files have changed in this diff Show More