Files
acdl/.ciagent/RESEARCH.md
T
Jon Chery 0404988465 docs(P00): research findings — v1.22 deck layout root cause (REQ-254..262)
8 findings (all confidence >= 0.8):
- F1 (VERY HIGH): theme CSS has zero section padding (/* @theme */ is
  a comment, not the directive; no @import of Marp default).
- F2 (VERY HIGH): overflow:hidden silently clips dense content (8/19
  slides overflow).
- F3 (HIGH): image aspect-ratio catastrophe (platform-pipeline 13.52x,
  telemetry-live-ops 0.63x).
- F4 (HIGH): header+footer chrome on every slide (~70px lost).
- F5 (MEDIUM-HIGH): render_deck.sh omits --theme (unthemed output).
- F6 (HIGH): render_slides.sh missing -s 2 -b transparent (1x PNGs).
- F7 (LOW): P5 marp-cli version bump — NOT the cause (theme CSS byte-
  identical P3->P5).
- F8 (VERY HIGH): test coverage gaps — no layout/overflow/aspect-ratio
  tests; static-file-property tests only.

Persona roster (v1.22): lead-developer (theme CSS + deck markdown +
mermaid + .ciagent), backend-engineer (render scripts + tests).
frontend-engineer + data-engineer deactivated. D-148 (theme CSS is
lead-developer, not frontend), D-149 (no new personas).

---ci---
project: acdl
phase: 0
milestone: v1.22
status: research
---/ci---
2026-08-11 19:22:45 +00:00

148 KiB
Raw Blame History

Nova — v1.11 RESTART Research Findings

Phase: research (pre-Phase 56). Milestone: v1.11 (RESTART). Status: research. Researcher: ci-researcher. Autonomy: full (CLARIFY auto-resolved; all binding decisions D-097..D-107 are committed in the CLARIFY stage). Branch: milestone/v1.11-restart (branched off tag v1.10.2, per D-097). Sources: ACDL codebase (v1.10.2 tree) + git history (failed first attempt on phase/56-iam-re-bootstrap + phase/57-live-deploy-microservice) + the CLARIFY commit (80b7286).

This file overwrites the prior v1.1 research artifact. v1.11 is a fresh milestone; the v1.1 research (Gitea OIDC, Checkov, IR shape, outbox) is historical and preserved in git history. This file documents the technical findings that ground the v1.11 restart plan.


Background — why v1.11 is a restart

v1.11 is a restart, not a continuation. The first attempt (phase/56 + phase/57, abandoned per D-097) made five defects worse, not better. The restart branches off the clean v1.10.2 tag and corrects three structural defects that the CLARIFY stage locked as binding decisions:

  1. Stateless adapter (D-098, D-099, D-100). The current adapter is a 750-line monolith with 3 constant tables and 39 type-specific branches that duplicate what interface.json already declares and hardcode defaults that belong in the module. v1.11 makes it a ~80-line stateless assembler; each L1 ships a real terraform/ module dir that owns its resource shape, nested blocks, and defaults.
  2. Terraform owns lifecycle (D-101). The first attempt added a Python script (verify_deploy_microservice.py) that ran terraform init -reconfigure in a fresh temp dir each time, which contributed to the 4-VPC bug. v1.11 deletes that script; run_platform.sh gains --apply and --destroy modes; Python never runs terraform.
  3. Pipeline-driven testing (D-102, D-103, D-104). No per-module Python/pytest. A modules-lifecycle pipeline matrix-runs each L1 module's examples/{simple,complex}.yml contracts through apply→modify→destroy against live AWS. The "test" = the pipeline cell going green.

FINDING 1 — Adapter monolith audit

1.1 The three constant tables

adapters/terraform/adapter.py (750 lines on the v1.10.2 tree) is built around three constant tables:

Table Line What it encodes Entries
TYPE_MAP 26 Stack type (aws:<service>:<kind>) → Terraform resource type (aws_s3_bucket, aws_vpc, …). 19
INPUT_MAP 51 Stack input name → Terraform arg name, per stack type. Only non-identity mappings are listed; an input not present uses the stack name as the Terraform arg (identity). 19 (one per stack type)
OUTPUT_MAP 75 Stack output name → Terraform attribute name, per stack type. Only non-identity mappings. 19 (one per stack type)

Why they duplicate interface.json. Each L1 module already declares its inputs, outputs, and stack type in interface.json (engine-agnostic). The three tables are the engine binding — the Terraform-specific name mappings that interface.json deliberately omits (it is engine-agnostic per ARCHITECTURE.md §12). The duplication is therefore intentional in the original design: the adapter was meant to be a thin translator that holds the engine binding in three tables, and the L1 holds the engine-agnostic content.

The drift. What was not intended is that the tables grew into 39 type-specific branches (§1.2) that hardcode resource shapes, nested HCL blocks, and defaults (§1.3) — content that belongs in the module, not the adapter. The adapter stopped being a thin translator and became a per-resource-type code generator. D-098 corrects this: the engine binding moves into a per-module terraform/ subdir (the real Terraform module), and the adapter becomes a stateless assembler that emits module "x" { source = "..." ... } blocks. The three tables are deleted.

1.2 The 39 type-specific branches across 18 stack types

_emit_resource (line 156) is a generic loop that, for each input, looks up the Terraform arg in INPUT_MAP, renders the value, and appends arg = value. But 18 of the 19 stack types have a specialized branch inside _emit_resource that runs after the generic loop and emits nested HCL blocks, hardcoded defaults, or resource-specific wiring. The count of 39 branches is the sum of the per-type specializations (some types have 23 branches). The full inventory:

# Stack type Terraform type Specialized logic (what the branch does)
1 aws:s3:bucket aws_s3_bucket versioning {} block (default true); server_side_encryption_configuration {} block (SSE-KMS, CMK ref or managed-key fallback with stderr warning); kms_key_arn is not a bare arg — emitted as the SSE block.
2 aws:ec2:vpc aws_vpc tags { Name = ... } from the name input; hardcoded cidr_block = "10.0.0.0/16" default when the L2 doesn't supply a CIDR (line 288).
3 aws:ec2:subnet aws_subnet vpc_id = aws_vpc.vpc-vpc.id hardcoded ref when not in inputs; hardcoded cidr_block = "10.0.1.0/24" default (line 296); tags { Name = ... }.
4 aws:ec2:routetable aws_route_table vpc_id = aws_vpc.vpc-vpc.id hardcoded ref; route { cidr_block = "0.0.0.0/0" gateway_id = aws_internet_gateway.vpc-igw.id } hardcoded default route; tags { Name = "<name>-rt" }.
5 aws:ecs:cluster aws_ecs_cluster Hardcoded name = "acdl-microservice" default when not in inputs (line 300).
6 aws:ecs:task_definition aws_ecs_task_definition _container_definitions() helper: jsonencodes image/port/env into a container_definitions block; hardcoded family = "app" default (line 279).
7 aws:ecs:service aws_ecs_service network_configuration {} block (subnets + security_groups wrapped in list brackets); load_balancer {} block from lb_target_group_arn with hardcoded container_name = "app" + container_port = 8080; hardcoded desired_count = 1, launch_type = "FARGATE", task_definition = aws_ecs_task_definition.service-task-definition.arn, name = "acdl-microservice".
8 aws:iam:role aws_iam_role managed_policy_arns = [...] from comma-separated string; hardcoded ECS task execution assume_role_policy JSON when not supplied (line 326331); hardcoded name = "acdl-microservice-role" default.
9 aws:elbv2:loadbalancer aws_lb subnets/security_group wrapped in list brackets; hardcoded load_balancer_type = "application" default.
10 aws:elbv2:listener aws_lb_listener default_action { type = "forward" target_group_arn = aws_lb_target_group.alb-targetgroup.arn } hardcoded; load_balancer_arn = aws_lb.alb-loadbalancer.id hardcoded ref.
11 aws:elbv2:targetgroup aws_lb_target_group Hardcoded target_type = "ip", vpc_id = aws_vpc.vpc-vpc.id, protocol = "HTTP", port = 8080.
12 aws:ecr:repository aws_ecr_repository Hardcoded name = "acdl-microservice" default; encryption_configuration {} block (not a bare kms_key_arn arg).
13 aws:cloudfront:distribution aws_cloudfront_distribution origin {} block (origin_id, domain_name, origin_access_control_id, s3_origin_config {}); default_cache_behavior {} block (viewer_protocol_policy, target_origin_id, ttls, allowed/cached methods); enabled = true; price_class; restrictions { geo_restriction {} }; viewer_certificate { cloudfront_default_certificate = true }; web_acl_id from WAF ref. ~8 nested blocks.
14 aws:cloudfront:originaccesscontrol aws_cloudfront_origin_access_control name; hardcoded origin_access_control_origin_type = "s3", signing_behavior = "always", signing_protocol = "sigv4".
15 aws:wafv2:webacl aws_wafv2_web_acl name; hardcoded scope = "CLOUDFRONT"; default_action {} (allow/block from input, default allow); visibility_config {}; custom rule {} blocks as nested HCL (P1-4 fix) or default AWS-managed-rules block. ~5 nested blocks.
16 aws:rds:instance aws_db_instance NFR-derived backup_retention_period (default 7), deletion_protection (default true); storage_encrypted = true default; skip_final_snapshot = true (dev safety).
17 aws:kms:key aws_kms_key NFR-derived enable_key_rotation = true default.
18 aws:ecs:uptime-service aws_ecs_service Feature-flag gate (returns "" when disabled); container_definitions jsonencode for uptime-kuma; hardcoded subnets = ["subnet-uptime"], security_groups = ["sg-uptime"], assign_public_ip = true; hardcoded desired_count = 1, launch_type = "FARGATE".

Plus a global prevent_destroy lifecycle block emitted for every resource when nfrs.deletion_protection is true (line 576581), and the _emit_igw() helper that synthesizes an internet gateway + route table association from the VPC resource (line 585).

1.3 Hardcoded defaults that belong in the module

The defaults below are emitted by the adapter when the L2 composition does not supply the input. They are resource shape decisions — CIDR ranges, trust policies, network config — that belong in the module's locals.tf (D-100), not in the adapter. The adapter should pass only resolved contract inputs; if a default is wrong, fix the module, not the adapter.

Default Adapter line What it is Where it belongs
cidr_block = "10.0.0.0/16" 288 VPC CIDR default modules/l1/vpc/terraform/locals.tf
cidr_block = "10.0.1.0/24" 296 Subnet CIDR default modules/l1/vpc/terraform/locals.tf
ECS task execution assume_role_policy JSON 326331 Trust policy for the IAM role modules/l1/iam-role/terraform/main.tf (or locals.tf)
ECR/logs inline policy / encryption_configuration {} 304315, 380 ECR KMS encryption block modules/l1/ecr/terraform/main.tf
Fargate requires_compatibilities / launch_type = "FARGATE" 261263 ECS launch config modules/l1/ecs-service/terraform/locals.tf
assign_public_ip (uptime) 565 ECS network config modules/l1/uptime/terraform/main.tf
Listener/target ports (port = 8080, container_port = 8080) 201, 348 ALB + ECS container ports modules/l1/alb/terraform/locals.tf + modules/l1/ecs-service/terraform/locals.tf
Security group emission (security_groups = [...]) 255258, 564 ECS network config modules/l1/ecs-service/terraform/main.tf
name = "acdl-microservice" (cluster, ECR, service) 265, 300, 303 Resource name defaults modules/l1/*/terraform/locals.tf
family = "app" 279 Task definition family modules/l1/ecs-service/terraform/locals.tf
target_type = "ip", protocol = "HTTP" 345, 347 ALB target group defaults modules/l1/alb/terraform/locals.tf
load_balancer_type = "application" 342 ALB type default modules/l1/alb/terraform/locals.tf
desired_count = 1 260 ECS desired count modules/l1/ecs-service/terraform/locals.tf
WAF scope = "CLOUDFRONT", managed-rules default block 421, 472490 WAF defaults modules/l1/waf/terraform/main.tf
CloudFront signing_behavior = "always", signing_protocol = "sigv4", origin_type = "s3" 365367 OAC defaults modules/l1/cloudfront/terraform/main.tf
CloudFront viewer_certificate { cloudfront_default_certificate = true }, restrictions {} 403410 Distribution defaults modules/l1/cloudfront/terraform/main.tf
RDS backup_retention_period = 7, skip_final_snapshot = true 497, 506 RDS defaults modules/l1/rds/terraform/locals.tf
KMS enable_key_rotation = true 510 KMS rotation default modules/l1/kms-key/terraform/main.tf
prevent_destroy = true lifecycle (global) 576581 Deletion protection Each module's main.tf (or a shared lifecycle.tf)

1.4 Why this is a drift from the original vision

ARCHITECTURE.md §12.2 states: "The adapter is a thin layer; it does not own L1/L2 content — it only translates." STANDARDS.md §8 (line 448506) documents the intended design: "a thin translator with 3 tables + specialized branches." The drift was baked into the standards doc itself — §8.2 explicitly blesses "specialized _emit_resource branches" for "resources with nested HCL blocks" and §8.3 step 4 instructs module authors to "add a specialized branch in _emit_resource keyed on that stack type" when a new L1 needs nested blocks.

The result: every new L1 with a nested block (CloudFront, WAF, ECS, uptime) added 3080 lines of resource-shape code to the adapter. The adapter grew from a spike-era ~150 lines to 750 lines, with the resource shape (CIDR ranges, trust policies, container ports, managed-rule sets) encoded as Python string concatenation rather than Terraform HCL. D-098 corrects the drift: the standards doc §8 must be rewritten to document the new pattern (per-module terraform/ subdir + stateless assembler), and the "specialized branch" guidance is removed.

Confidence: 0.95. The audit is a direct line-by-line read of the v1.10.2 adapter.py; the drift is structural and unambiguous.


FINDING 2 — State-key root cause of the 4-VPC bug

2.1 The state key

adapter.py line 664 + 676:

stack_name = stack.get("name", "spike")
terraform_tf = (
    ...
    f'    key    = "spike/{stack_name}/terraform.tfstate"\n'
    ...
)

core/contract_resolver.py line 569:

"stack": {
    "name": contract["id"],
    ...
}

So stack_name = contract["id"] and the state key is spike/{contract.id}/terraform.tfstate.

2.2 The 5 microservice contracts

All five microservice contracts share id: msvc and differ only in environment:

Contract file id environment
contracts/microservice.yml msvc dev
contracts/microservice.dev.yml msvc dev
contracts/microservice.qa.yml msvc qa
contracts/microservice.prod.yml msvc prod
contracts/microservice.dr.yml msvc dr

The state key does not include the environment. So all four environment contracts (dev/qa/prod/dr) collide on the same state key: spike/msvc/terraform.tfstate.

2.3 The two root causes

Root cause 1 — the adapter emits per-contract state keys with no VPC sharing. The microservice composition (modules/l2/microservice/ composition.json) includes a vpc child (vpc@1.0.0). Every contract that resolves through this composition emits its own VPC resource. There is no platform VPC to share; each contract deploys its own VPC. D-105 corrects this: terraform/platform owns ONE VPC; the microservice composition drops its vpc child and references the platform VPC via a data source. The standalone vpc L1 module stays (consumers deploy their own VPCs). No per-contract VPC ever again.

Root cause 2 — the state key does not distinguish environments. Because the state key is spike/{contract.id}/terraform.tfstate and all four env contracts share id: msvc, every environment's terraform apply writes to the same remote state key. Combined with the first attempt's verify_deploy_microservice.py running terraform init -reconfigure in a fresh temp dir each time, each run created a fresh local state that diverged from the remote key. The first run (dev) created VPC #1 and pushed it to spike/msvc/terraform.tfstate. The second run (qa) ran -reconfigure in a fresh temp dir, pulled the remote state (which had dev's VPC), but because the local state was fresh and the composition emitted a new VPC resource address, terraform saw the VPC as "to add" again — creating VPC #2 and overwriting the remote state. Repeating for prod and dr created VPCs #3 and #4. Four VPCs, one state key, no environment discrimination.

D-106 corrects this: the composition must be deterministic — same contract → same resolved stack → same state key, every time. State keys become env-aware and stable across apply/modify/destroy: spike/{id}/{env}/terraform.tfstate. The environment is part of the key, so dev/qa/prod/dr never collide.

2.4 Why -reconfigure in a fresh temp dir made it worse

terraform init -reconfigure forces terraform to re-read the backend config and pull remote state into the local working directory. When the working directory is a fresh temp dir (as verify_deploy_microservice.py did), there is no local .terraform/ state cache — terraform must pull the remote state fresh. If the remote state key is shared across environments (root cause 2) and the composition emits a new VPC each time (root cause 1), the -reconfigure pull merges the prior environment's state with the new resource addresses, and the subsequent apply creates a new VPC because the resource address in the new composition run differs from the one in the remote state (the L2 namespacing or the fresh temp dir caused terraform to treat the VPC as a new resource). D-101 deletes verify_deploy_microservice.py entirely; run_platform.sh gains --apply and --destroy modes that run terraform in a stable working directory (not a fresh temp dir per run), and Python never runs terraform.

Confidence: 0.90. The state-key derivation is a direct code read (adapter.py:664,676 + contract_resolver.py:569). The 5 contracts are read verbatim. The 4-VPC mechanism is the only consistent explanation for the observed symptom (4 VPCs in the account after 4 env runs). The 0.10 residual is for the possibility that the resource-address divergence was caused by a separate composition-namespacing bug rather than the fresh temp dir alone — but either way, the two root causes (shared state key + per-contract VPC) are confirmed and D-105/D-106 correct both.


FINDING 3 — Per-module terraform module design

3.1 What the per-module terraform/ subdir should contain

D-098/D-099: each L1 module ships a real terraform/ module dir. The canonical layout for a multi-resource module:

modules/l1/<name>/
  interface.json          # engine-agnostic (unchanged)
  instance.json          # regression baseline (unchanged)
  README.md
  examples/
    simple.yml
    complex.yml
  terraform/              # NEW — the engine binding
    versions.tf           # required_version + required_providers
    variables.tf          # from interface.json inputs
    locals.tf             # default interpolation (heavy use, D-099)
    main.tf               # resource blocks (resource shape + nested blocks)
    outputs.tf            # from interface.json outputs

Trivial single-resource modules (e.g. s3) may inline locals in main.tf (D-099). Multi-resource modules (vpc, ecs-service, alb, microservice-shaped) get the full split.

3.2 The three reference modules (from interface.json)

s3 (modules/l1/s3/interface.json):

  • variables.tf: bucket_name (string, required), region (string, required), kms_key_arn (string, optional).
  • locals.tf: sse_algorithm = "aws:kms", versioning default true, managed-key fallback (alias/aws/s3 when kms_key_arn is null), the prevent_destroy lifecycle.
  • main.tf: resource "aws_s3_bucket" "this" { bucket = var.bucket_name ... } + versioning {} block + server_side_encryption_configuration {} block (CMK ref or managed fallback).
  • outputs.tf: bucket_arn (→ aws_s3_bucket.this.arn), bucket_name (→ aws_s3_bucket.this.id), bucket_regional_domain_name (→ aws_s3_bucket.this.bucket_regional_domain_name).
  • versions.tf: terraform { required_version = ">= 1.9, < 1.10" required_providers { aws = { source = "hashicorp/aws", version = "~> 5.0" } } }.

vpc (modules/l1/vpc/interface.json — multi-resource: vpc + subnet + routetable):

  • variables.tf: cidr (string, required), azs (string, required), name (string, required), region (string, required).
  • locals.tf: cidr_block = coalesce(var.cidr, "10.0.0.0/16"), subnet CIDR derivation (cidrsubnets(local.cidr_block, 8, 8, ...) per AZ), name tag interpolation, the IGW + route table association.
  • main.tf: aws_vpc, aws_subnet (count/for_each over azs split), aws_route_table, aws_internet_gateway, aws_route_table_association — all the resources that the adapter's _emit_igw() helper synthesized dynamically now live here as real HCL.
  • outputs.tf: vpc_id, subnet_ids (join the subnet ids).
  • versions.tf: same provider block.

ecs-service (modules/l1/ecs-service/interface.json — multi-resource: task_definition + service):

  • variables.tf: image, port, cpu (default 256), memory (default 512), env (optional), cluster_arn, subnets, security_group, lb_target_group_arn (optional), region, kms_key_arn (optional), desired_count (default 1), launch_type (default "FARGATE"), family (default "app").
  • locals.tf: container_definitions jsonencode (image/port/env/cpu/ memory), requires_compatibilities = ["FARGATE"] when launch_type is FARGATE, log group name + KMS ref, the prevent_destroy lifecycle.
  • main.tf: aws_ecs_task_definition (family, container_definitions, requires_compatibilities, execution_role_arn) + aws_ecs_service (name, cluster, task_definition, desired_count, launch_type, network_configuration {}, load_balancer {} block).
  • outputs.tf: service_arn, task_def_arn.
  • versions.tf: same provider block.

3.3 How the stateless adapter assembles them

The new adapter (D-098) is a ~80-line stateless assembler. It:

  1. Reads modules/registry.json → for each resource in the resolved stack instance, looks up the L1 module by module field (<name>@<semver>).
  2. Gets the terraform_dir from the registry entry (or derives it as modules/l1/<name>/terraform/).
  3. Emits a root main.tf with one module "x" { source = "<terraform_dir>" ... } block per resource, passing the resolved contract inputs as module arguments.
  4. Wires refs via module "x".<output> interpolations: a ref:<id>.<out> input value becomes module.<id>.<out> in the consuming module block.
  5. Emits the stack-level output {} blocks (passthrough from the producing module's outputs).
  6. Emits terraform.tf (backend config with the env-aware state key, D-106) + providers.tf (aws provider, region from the first resource).

The adapter holds no TYPE_MAP, INPUT_MAP, OUTPUT_MAP, and no type-specific branches. The engine binding (stack type → Terraform resource type, input → arg name, output → attribute name, nested blocks, defaults) lives entirely in the per-module terraform/ subdir. interface.json stays engine-agnostic.

Confidence: 0.90. The module layout is grounded in the existing interface.json files (read verbatim) and the Terraform module convention (versions/variables/locals/main/outputs split). The assembler design is D-098/D-099 (user-confirmed). The 0.10 residual is for the exact terraform_dir registry field shape (not yet implemented) and the ref-wiring syntax (module.<id>.<out> vs a locals alias).


FINDING 4 — Existing pipeline architecture

4.1 The central pipeline contract

pipelines/contract.yml is the declarative deployment pipeline spec (a contract, not an executable workflow). It declares 9 stages: validate-contractresolve-stackterraform-plancheckovconfidenceapply (dev only) → publish-outputsdeploy-uptimecomment-outputs. Each stage has name, command, required (bool), and optional description. The executable workflow (.github/workflows/deploy.yml + .gitea/workflows/deploy.yml, byte-identical) implements these stages by invoking scripts/run_platform.sh. Validated against schemas/deploy-pipeline.schema.json.

4.2 The plan-only pipelines (existing, run on every PR)

Two platform pipelines run on every PR to main (offline, free):

Pipeline File Matrix What it does
Primitives plan .github/workflows/primitives-plan.yml (+ .gitea/ byte-identical) s3, vpc, ecs-cluster, ecs-service, iam-role, alb, ecr, cloudfront, waf, rds (10 primitives) For each L1 primitive, runs bash scripts/run_primitive_plan.sh --check-only <primitive> — resolves the primitive's instance.json, runs the adapter, validates the emitted Terraform structure (offline, no AWS).
Patterns plan .github/workflows/patterns-plan.yml (+ .gitea/ byte-identical) static-assets, microservice (2 modules) For each L2 module, runs bash scripts/run_pattern_plan.sh --check-only <module> — resolves the sample contract, runs the adapter, validates the emitted Terraform (offline).

Both trigger on pull_request: branches: [main], run on ubuntu-latest, install jsonschema pyyaml boto3. The --check-only mode is offline (no AWS, no Checkov, no DynamoDB) — it resolves the contract/instance, runs the adapter, and validates the emitted Terraform file structure. This is what makes the pipelines free.

4.3 run_platform.sh — plan only, never apply/destroy

scripts/run_platform.sh (521 lines) has three modes today:

  • --check-only (offline, no AWS): contract → resolver → adapter → stream TF → validate → exit 0.
  • --plan-only (requires AWS): contract → resolver → adapter → terraform init -reconfigure -lock=falseterraform validateterraform plan -lock=false -out=tfplan → exit 0 (line 274297).
  • default (requires AWS + Checkov + DynamoDB): contract → resolver → adapter → terraform plan → Checkov → confidence → outbox.

Critically, line 287 runs terraform plan only. There is no terraform apply and no terraform destroy in run_platform.sh today. The apply stage in pipelines/contract.yml (line 5457) declares command: bash scripts/run_platform.sh --plan-only — a misnomer; it runs plan, not apply. The lifecycle modes (--apply, --destroy) must be added (D-101). Python never runs terraform; run_platform.sh is the only shell entry point.

4.4 run_primitive_plan.sh

scripts/run_primitive_plan.sh (65 lines) runs the platform pipeline for a single primitive. --check-only mode: resolves instance.json, runs the adapter, validates the emitted {main.tf,terraform.tf,providers.tf} exist and main.tf is non-empty. Default mode (requires AWS): terraform init -backend=falseterraform validateterraform plan. This is the per-primitive plan check that the primitives-plan pipeline matrix invokes.

4.5 The byte-identical Gitea+GitHub convention

pipelines/README.md:22 documents the convention: "Create byte-identical workflow YAMLs in .gitea/workflows/<name>.yml and .github/workflows/<name>.yml." Both workflows must implement the same stages, commands, triggers, and runner declared in the contract. tests/test_pipeline_contract.py validates that the Gitea and GitHub workflow YAMLs are byte-identical and conform to the schema. The only difference is the forge runtime (Gitea Actions vs GitHub Actions). The new modules-lifecycle pipeline (D-102) must follow this convention: byte-identical .gitea/workflows/modules-lifecycle.yml + .github/workflows/modules-lifecycle.yml.

Confidence: 0.95. All pipeline files are read verbatim from the v1.10.2 tree. The "plan only, never apply/destroy" finding is a direct read of run_platform.sh line 287 + the --plan-only exit at line 293.


FINDING 5 — PERSONAS.md update for v1.11

The existing PERSONAS.md (v1.9) has 6 active personas: lead-developer, backend-engineer, platform-engineer (custom), security-engineer (custom), lambda-engineer (custom, v1.9), frontend-engineer. v1.11 changes the roster:

  • Deactivate lambda-engineer — no per-module Python this milestone (D-102: testing is pipeline-driven, not pytest). The v1.9 Lambda (core/lambda/contract_ingestor.py) persists but is not touched in v1.11.
  • Deactivate cost-engineer — not in the v1.9 roster (the v1.9 data-engineer is already deactivated). v1.11 has no cost-engineer work; cost is documented in COST.md (REQ-119) by the lead-developer.
  • Keep backend-engineer — owns the adapter rewrite (stateless assembler) + core/contract_resolver.py (env-aware state keys, D-106).
  • Keep data-engineer (reactivated) — owns terraform/ (platform VPC, D-105) + the per-module terraform/ subdirs (the engine binding, D-098/D-099/D-100). This is the heaviest territory in v1.11: 12 L1 modules each get a real terraform/ module dir.
  • Keep general (the lead-developer + backend-engineer pipeline work) — owns pipelines/ + .gitea/workflows/ + .github/workflows/ (the modules-lifecycle pipeline, D-102) + scripts/run_platform.sh (--apply/--destroy modes, D-101).

Territory alignment (v1.11)

Persona Territory Domain
backend-engineer adapters/terraform/adapter.py (rewrite to stateless assembler), core/contract_resolver.py (env-aware state keys), schemas/stack.schema.json (if touched) backend
data-engineer terraform/ (platform VPC, D-105), modules/l1/*/terraform/ (per-module terraform subdirs — the engine binding), modules/l1/*/interface.json (defaults move from adapter to interface), modules/registry.json (terraform_dir field) data
general (lead-developer + backend-engineer) pipelines/modules-lifecycle.yml, .gitea/workflows/modules-lifecycle.yml + .github/workflows/modules-lifecycle.yml (byte-identical), scripts/run_platform.sh (--apply/--destroy), scripts/run_primitive_plan.sh (if extended), modules/STANDARDS.md §8 rewrite coordination + pipelines

Territory enforcement: warn

Co-authoring is expected on the adapter + run_platform.sh boundary (backend-engineer rewrites the adapter; general adds the lifecycle modes to run_platform.sh that invoke it). warn keeps it frictionless — cross-territory edits are logged in the commit message but do not fail the task.

Domain priority (v1.11)

data → backend → general

Rationale: the terraform foundation (per-module terraform/ subdirs + platform VPC) is the binding constraint — the stateless adapter cannot be written until the reference s3 module exists (D-107: P56a proves the design with s3 first). Backend (adapter/resolver) follows once the module shape is proven. General (pipelines/workflows) wires the lifecycle modes last, once the adapter + modules produce valid terraform.

The updated PERSONAS.md is written to /root/acdl/.ciagent/PERSONAS.md (see that file). YAML frontmatter with active, phase_specific, and reason fields per persona.

Confidence: 0.90. The persona changes are grounded in the CLARIFY decisions (D-098..D-107) and the v1.11 scope (no per-module Python → lambda-engineer deactivated; terraform module authoring is the heaviest work → data-engineer reactivated).


Assumptions logged

ID Assumption Confidence Rationale
A-1.1 The terraform_dir field will be added to modules/registry.json entries (or derived as modules/l1/<name>/terraform/) so the stateless adapter can locate each module's terraform subdir. 0.85 D-098 says the adapter reads registry.json → gets terraform_dir. The exact field name is not yet locked; the derivation path is the obvious fallback.
A-1.2 The ref-wiring syntax in the root main.tf will be module.<id>.<output> (standard Terraform module output interpolation), not a locals alias. 0.85 The existing _ref_expr already produces <tf_type>.<id>.<attr>; the module equivalent is module.<id>.<output>. Standard Terraform convention.
A-2.1 The 4-VPC bug's resource-address divergence was caused by the fresh temp dir + -reconfigure pull merging remote state with new composition runs, not a separate composition-namespacing bug. 0.80 The two confirmed root causes (shared state key + per-contract VPC) are sufficient to explain 4 VPCs. The exact terraform-state mechanics of the divergence are inferred, not observed in a debug log.
A-3.1 Trivial single-resource modules (s3) may inline locals in main.tf; multi-resource modules (vpc, ecs-service, alb) get the full 5-file split. 0.90 D-099 states this explicitly.
A-4.1 The modules-lifecycle pipeline will matrix-run each L1 module's examples/{simple,complex}.yml contracts (the modify variants), not new contract files. 0.90 D-103: "Uses the module's own existing example contracts as the modify variants. No extra contract files needed."
A-5.1 platform-engineer and security-engineer from the v1.9 roster are folded into data-engineer and backend-engineer for v1.11 (the v1.11 scope is terraform + adapter + pipelines, not security adapters or HITL gates). 0.75 The v1.11 scope (D-097..D-107) does not touch Wiz/Kyverno/Checkov/HITL. The persona roster is simplified to the three active domains.

Decisions surfaced (research → already bound in CLARIFY)

All v1.11 binding decisions (D-097..D-107) were committed in the CLARIFY stage (80b7286) before this research ran. This research grounds those decisions with codebase evidence; it does not surface new binding decisions. The decisions are summarized in §Background above and documented in full in the CLARIFY commit.


v1.12 Addendum — Presentation Refinement Research

Generated: 2026-07-29. Phase 66. Milestone v1.12. Mode: docs-only NFR milestone focused on the leadership decks. Surface: docs/presentations/ (PW + DX, all four layers) + one real adapter fix + two probe fixes required to make deck claims true.

Background — why v1.12 exists

v1.11 (P56aP65) landed the stateless adapter, pipeline-driven lifecycle testing, single platform VPC, COST.md, PRE_MORTEM.md, and a teardown to zero-cost. P65's plan (REQ-118) required the decks to be rewritten to "Verified live-aws via lifecycle pipeline; torn down to zero-cost." That rewrite did not fully land on the deck artifacts. This research is a drift audit: a systematic comparison of the deck artifacts against the v1.11-verified reality.

FINDING 1 — Drift audit (9 items)

Systematic comparison of docs/presentations/* against .ciagent/CAPABILITY_INVENTORY.md, .ciagent/COST.md, .ciagent/PRE_MORTEM.md, .ciagent/ROADMAP.md, and git log.

  1. Wrong verification status. Both rendered HTML decks still say "6 cloud capabilities are design-verified + locally emulated, deploy-unverified (IAM drift)" (PW "Testing vs. Planned" slide; DX slide A6). CAPABILITY_INVENTORY.md says 22/22 Verified and the IAM-drift framing was removed in P65. The decks contradict the inventory. Verified: grep -c "deploy-unverified\|IAM drift\|design-verified" docs/presentations/*.html → 3 hits per deck.
  2. Re-verification header stale. Both source .md headers say "Re-verification (2026-07-27)… v1.10 Phase 54… 16/16… 6 IAM-gated escalated." Should reflect v1.11: 22/22 Verified, torn down.
  3. Road to the North Star diagram stale. docs/presentations/assets/mmd/road-to-north-star.mmd shows v1.10 as "NEXT" with "HITL wiring / all-runner OIDC / regulatory ledger". v1.11 is complete; the diagram must advance.
  4. Rendered HTML not re-rendered. git log shows the HTML was last touched at 10b87a6 (P57), before v1.11. P65's "re-render HTML" task did not reach the rendered artifacts.
  5. Version refs stale. Decks reference @v1.10 in deploy.yml uses: snippets (Safe Promotion Path, Safe Decommission). Ship tag is now v1.11.0; will be v1.12.0 at Phase 70 complete.
  6. Cost story has no real numbers. COST.md exists ($0.001883 over 8 days, ~$0.007/mo, S3-dominated, zero BAU compute) but the decks' A6 "Operating Model & Cost" slide is generic prose with no figures.
  7. Pre-mortem unreferenced. P65 planned to add a pre-mortem reference; PRE_MORTEM.md exists (v1.10 decay root cause + four forward failure modes) but no deck slide references it.
  8. Two v1.11 stories absent. (a) Architectural simplicity: adapter 918→~80 lines, defaults centralized in per-module terraform/ dirs. (b) Verifiable deploys: a modules-lifecycle pipeline matrix-runs each module apply→modify→destroy against live AWS. Neither is in the decks.
  9. Duplicated story-beat lines. how-the-platform-works.md slides 310 each repeat their intro line twice (a copy-paste artifact).

FINDING 2 — Regression gate surfaces real decay (D-091)

The v1.12 regression gate run (Phase 66) re-ran the D-091 regression gate to back every deck claim. It found 3 Broken capabilities: {'Verified': 19, 'Decayed': 0, 'Broken': 3}.

CAP-013 — live-aws — REAL platform defect (Class A)

adapters/terraform/adapter.py:159-172 (the seen dedup loop) collapses the two ecs-service sub-resources (service-task-definition

  • service-service, both module ecs-service@1.0.0) into ONE module "service-task-definition" block. But the stack output service_arn (resolver from: "service-service") is emitted as value = module.service-service.service_arn — referencing a module call that was never emitted. terraform validate fails: "No module call name." The same defect silently breaks the alb L1 too. Static- assets (CAP-014) doesn't hit it because its L1s are single-resource. Classification A — real platform defect. The adapter produces invalid Terraform for any multi-resource L1 with stack-level outputs. Fix required before decks can claim 22/22 Verified.

CAP-017 — lifecycle-pipeline — regression-probe bug (Class B/C)

core/regression_verify.py:444 hardcodes required = ["versions.tf", "variables.tf", "locals.tf", "main.tf", "outputs.tf"]. The CAP-017 probe targets the rds L1 module, whose main.tf uses only var.* and aws_db_subnet_group.this — no local.* references, so locals.tf is legitimately absent. The probe is over- strict. The rds module is correctly structured; the capability works. Classification B/C — trivial probe fix. Drop locals.tf from the required list, or make it conditional on local. usage.

CAP-018 — lifecycle-pipeline — regression-probe bug (Class B/C)

core/local_emulators.py:273 defines LocalLambdaStub as a dataclass with one required field outbox: FlatFileOutbox. Every real caller passes it (core/local_emulators.py:464, the tests). The CAP-018 probe at core/regression_verify.py:486-491 is the only caller that instantiates it bare: LocalLambdaStub()TypeError. The probe was added in P63 and never aligned with the real signature. The capability is exercised green by CAP-011. Classification B/C — trivial probe fix. Pass an outbox to the constructor.

Implication for the decks

CAPABILITY_INVENTORY.md claims 22/22 Verified, but the regression gate (D-091 — the exact mechanism PRE_MORTEM.md FM-3 says backs every deck claim) shows CAP-013 is genuinely broken. The inventory overstates. v1.12 cannot ship decks claiming 22/22 until CAP-013 is fixed and the gate re-runs clean. This is the structural mitigation the pre-mortem requires (verified-only claims; decks unfrozen only after re-verification). The user decision: fix the defect inside v1.12 (Phase 67), then the decks can honestly claim 22/22.

FINDING 3 — Talking points structure gap

Both talking-points files have only 5 appendix sections (A1A5) while the Marp decks have 6 (A6 = "Operating Model & Cost"). The A6 content exists in the Marp deck and source markdown but was never distilled into the talking points. The re-distill step (Phase 69) must add the A6 section to both talking-points files.

FINDING 4 — Versioning facts

  • Current ship tag: v1.11.0 (v1.11 complete).
  • deploy.yml still references v1.9 in comments + ref: v1.9 — v1.11 apparently did not bump the deploy workflow uses: tag (the bump is a separate concern; decks use the current ship tag).
  • Decks should show @v1.11 in examples (current state); Phase 70 bumps to @v1.12 after the tag exists.

Assumptions logged

  • No automated ci-doc-verifier script exists in the repo. The plan's "ci-doc-verifier confirms" is satisfied by a manual grep-based verification recorded in the Phase 70 VERIFY step (consistent with how prior NFR-patch phases handled it). Confidence 0.90 — verified by ls scripts/ | grep doc and grep -rl deck tests/.
  • PPTX export requires Chromium + Marp CLI; the environment has it (/root/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome). PPTX is uploaded to the Gitea release, not committed. Confidence 0.85 — README documents the path; the chromium binary exists.

Decisions surfaced (research → bound in CLARIFY-equivalent)

  • D-108 — v1.12 includes one real adapter fix (CAP-013) and two probe fixes (CAP-017, CAP-018) as Phase 67 prerequisites, so the decks can honestly claim 22/22 Verified. The milestone is "presentation refinement" but the verified-only-claims pre-mortem mitigation makes the fixes mandatory. The user confirmed this scope (interactive decision, 2026-07-29).
  • D-109 — Decks use @v1.11 in examples during Phase 68 (current state), bumped to @v1.12 at Phase 70 complete after the tag exists. Avoids a dangling reference to a tag that doesn't exist yet.

v1.14 Research Addendum — NFR Refinement scope audit (2026-07-29)

Phase 0 RESEARCH for milestone v1.14 (NFR Refinement). A full codebase survey (8 categories, file:line evidence) was conducted to populate the 20-phase scope. This addendum records the findings; the phase list is in ROADMAP.md §v1.14; the requirements are in REQUIREMENTS.md §v1.14.

Survey method

Read-only survey of /root/acdl at v1.13.2 (HEAD 139224ff, 533 tests collected). 8 categories: stubs, P1/P2 backlog, security, docs drift, test gaps, terraform gaps, workflow gaps, config hygiene. All file:line references verified against the live codebase.

Finding 1 — Open P1/P2 backlog (REVIEW.md v1.11)

5 P1 + 4 P2 findings from the v1.11 multi-persona review remain open:

ID File:Line Status v1.14 phase
P1-1 adapter.py:159-170 (silent drop of unregistered-module resources) open P1
P1-2 static-assets/composition.json (unwired cloudfront inputs; WAF unconditional) open P2
P1-3 run_l2_lifecycle_*.sh (vestigial [ci-vpc-outputs.json] arg) open P3
P1-4 CAPABILITY_INVENTORY.md:9-16 (summary table stale) fixed (now 22/22)
P1-5 regression_verify.py:432-519 (CAP-017..022 offline proxy, no terraform validate) open P4
P2-1 alb/main.tf:9 (name_prefix="tg-ci-" discards var.name) open P6
P2-2 test_adapter.py (no dedup-merge or remote-state-key test) open P5
P2-3 waf/complex.yml + locals.tf (redundant upper() + uppercase example) open (post-hoc) folded into P2
P2-4 COST.md:106 (account ID published; accepted exposure) open (post-hoc) folded into P8 (centralize code-side)

Finding 2 — Security posture gaps

Swallowed errors (6 sites):

  • core/local_emulators.py:374except Exception: pass in _fake_urlopen; if patching fails, urlopen stays real → network egress. [SEC] → P7.
  • core/lambda/contract_ingestor.py:157 — GitHub search failure → existing = [] → duplicate issues. → P7.
  • terraform/bootstrap/create_state_backend.py:51 — over-broad except Exception: on head_bucket → spurious create_bucket on permissions/network errors. → P7.
  • core/output_publisher.py:100,168 — SSM/GitHub failure → silent None/False. → P7.
  • terraform/bootstrap/apply_iam_baseline.py:78 — over-broad on old-version delete. → P7.

Hardcoded account ID 581513795199 (15+ sites): adapter.py:125,140, apply_iam_baseline.py:33, create_state_backend.py:33,35, push_consumer_image.py:32, terraform state-bucket names, ECR image ref. → P8 (externalize to ACDL_AWS_ACCOUNT_ID / data.aws_caller_identity).

IAM policy wildcards (6 Resource: "*" statements): spike_runner_policy.json — cloudfront (line 117), wafv2 (129), kms (218), iam (236). KMS allows key creation/deletion on ANY key; IAM allows role creation on ANY role. → P9 (scope to acdl-* ARNs).

Contract-ingestor identity validation gap: contract_ingestor.py:221-245_validate_caller_identity validates consumerRepo format only; doesn't verify caller owns the repo (ABAC reliance). No contractId/environment/error validation. → P10.

Schema validation gaps: contract.schema.json + environment.schema.json — no additionalProperties: false (undocumented fields pass silently); no format validation for bucket/ARN/CIDR. → P11.

Credential hygiene: .gitignore covers .env*/*.tfstate* but no credential-pattern catch-all (*.pem/*.key/*.p12). → P12.

Audit ledger integrity (D-083): audit_ledger_design.md:47-70 — JWS + Object Lock + DLQ deferred. Per D-096, stays deferred; documented in P19. The hash-chain + DynamoDB outbox is the v1.14 audit record.

Finding 3 — Stubs / missing functionality

  • adapters/kyverno/kyverno_adapter.py:11,115-116--kube-version parsed then discarded (_ = kube_version). → P13 (implement or remove + document).
  • scripts/__pycache__/verify_deploy_microservice.cpython-312.pyc — orphan bytecode for a deleted source file. → P14.
  • core/regression_verify.py:237 — DynamoDB write deferred to Phase 54 (outbox hash-chain verified, no real DynamoDB write). Accepted deferral.
  • adapters/wiz/wiz_adapter.py — real GraphQL client (not a stub); degrades gracefully. OK.
  • core/separation_of_duties.pyroute_halt_artifact is real (SNS + outbox fallback). OK.
  • core/lambda/contract_ingestor.pyreport_error is real (GitHub issues via Secrets Manager). OK.

Finding 4 — Documentation drift

  • ARCHITECTURE.md — no v1.11/v1.12/v1.13/v1.14 addendum; line 500-506 still describes the old parameterized adapter (pre-stateless rewrite). → P19.
  • Stale @v1.6@v1.9 workflow refs in README.md:225, docs/consumer-guide.md (12 sites), docs/architecture.md:233, docs/pipeline/versioning.md:29, docs/pipeline/index.md:42. → P19.
  • modules/STANDARDS.md §8 references TYPE_MAP (deleted in v1.11); §9.4 requires 5-file split but §489-492 allows inlining — inconsistent. → P18.
  • COST.md window stops at v1.10; no v1.11v1.13 spend. → P19.
  • GRILL.md G-005/G-008 escalations — CAP-017..022 now Verified via lifecycle pipeline; COST.md now exists. → P19 (mark resolved).
  • IAM_POLICY.md — reflects v1.11 re-bootstrap but not v1.12/v1.13. → P19.
  • Decks reference "v1.12" verification status; not re-synced for v1.13.2. → P19.

Finding 5 — Test coverage gaps

  • 533 tests collected; 5 @pytest.mark.slow (deselected from fast suite). 7 scripts with no test: seed_uptime_monitors.py, push_consumer_image.py, sync_to_gl.sh, post_stage_comment.sh, rotate_spike_key.sh, create_state_backend.py, create_iam_user.py. → P15.
  • Adapter dedup-merge + ACDL_REMOTE_STATE_KEY override — no unit test (P2-2). → P5.

Finding 6 — Terraform gaps

  • 3 L1 modules lack locals.tf (ecr, ecs-cluster, rds). → P18.
  • static-assets/composition.json unwired inputs (P1-2). → P2.
  • terraform/platform/main.tf:255 — hardcoded CIDR; count=2 subnets not data-driven. → P20.
  • terraform/bootstrap/create_state_backend.py:51 — over-broad except (Finding 2). → P7.

Finding 7 — Workflow / pipeline gaps

  • 4 GitHub-only workflows (patterns-plan, platform-test, primitives-plan, release) — no Gitea mirror. → P16.
  • rotate_spike_key.sh (only set -u), sync_to_gl.sh (no set flags). → P16.
  • 3 shared workflows (ci, deploy, modules-lifecycle) byte-identical (verified). OK.
  • modules-lifecycle matrix covers all 12 L1 + 2 L2. OK.

Finding 8 — Config / project hygiene

  • config.json bash_allowlist has dead JS entries (npm/node/jest/eslint /tsc — no package.json). → P14/P17.
  • config.json branching_strategy: "phase" mismatched with flat-workflow practice. → P17.
  • config.json ollama-cloud.base_url: "" (empty; no glm model configured). → P17.
  • config.json frontend-engineer persona still in personas[] (PERSONAS.md:80 says inactive). → P17.
  • pyproject.toml version 1.3.0 (stale); coverage source acdl_platform (renamed to core in v1.6). → P14.

Persona assessment (v1.14)

The v1.14 milestone is NFR-only (bug fixes, security, tests, docs). The active persona roster from v1.11 (PERSONAS.md) carries forward unchanged:

  • lead-developer (active) — coordination; owns the wave ordering + cross-phase dependencies.
  • backend-engineer (active) — owns adapters/, core/ (adapter dedup, contract ingestor, regression gate, output publisher).
  • data-engineer (active) — owns terraform/, modules/ (ALB fix, static-assets wiring, platform VPC, IAM policy, STANDARDS).
  • frontend-engineer (inactive) — no frontend; decks are markdown (lead-developer territory). Stays deactivated per PERSONAS.md:80.

No custom personas needed for v1.14 (no new domains). Territory enforcement = warn (config.json:167). The v1.14 work is concentrated in adapters/, core/, terraform/, scripts/, tests/, docs/, .ciagent/ — all within existing persona territories.


v1.15 Research Addendum — Nova Rebrand scope audit (2026-07-30)

Survey method

A thorough, exhaustive codebase survey (via the explore subagent) plus targeted grep -rni counts. The survey covered 346 tracked files, reporting occurrence counts and the mechanical-vs-judgment split per category. The full survey is recorded in the planning conversation transcript; the binding conclusions are summarized here.

Finding 1 — Brand string surface area

  • 1,465 total ACDL/acdl occurrences across 205 files.
  • 17 occurrences of the full "Agentic Cloud Delivery Platform" phrase (all prominent titles/headers: README, docs/index, vision, pyproject, decks, .ciagent/*.md).
  • 0 existing references to "nova" (case-insensitive) — no collision risk.
  • User-facing (docs/, README, decks, contracts, schemas, module READMEs): high-priority for rebrand.
  • Internal (.ciagent/*.md, tests/, terraform/, scripts/, workflows): mechanical but voluminous.

Finding 2 — Code identifiers (judgment category)

  • Python package name: pyproject.toml name = "acdl" (no acdl/ package dir exists — source lives in core/, adapters/; the name is a metadata label). Mechanical rename.
  • Python file: adapters/terraform/policy/custom_rules/acdl_tagging.py (+ Checkov registration in schemas/tagging-standard.json line 5 + adapter config). Rename file + update registration.
  • Env var prefixes: 21 distinct ACDL_* prefixes (ACDL_LIFECYCLE_MODE 66×, ACDL_AWS_ACCESS_KEY_ID 40×, ACDL_AWS_SECRET_ACCESS_KEY 37×, ACDL_REMOTE_STATE_KEY 22×, ACDL_AWS_ACCOUNT_ID 21×, ACDL_TAG_NAMING 20×, ACDL_KMS_KEY_ID 16×, ACDL_BOOTSTRAP_AWS_* 15× each, ACDL_SOD_HALT_TOPIC_ARN 14×, ACDL_LOCAL_TIER 13×, etc.). No centralized env loader exists today (scattered os.environ.get). D-108: a new core/env.py get_env() helper centralizes the dual-read fallback.
  • Workflow name:: .github/workflows/release.yml line 11 name: acdl-release — mechanical.

Finding 3 — AWS resource names (high-risk migration)

Terraform creates real AWS resources with acdl- prefixes. Renaming forces destroy+recreate (downtime, data loss for DynamoDB/state bucket). D-102: full rename with migration (user-directed).

Resource Type Migration
acdl-contracts / acdl-change-requests DynamoDB scan+copy data, verify row counts
acdl/github-token Secrets Manager recreate secret, repoint Lambda
acdl-contract-ingestor (role/policy/Lambda) IAM+Lambda recreate role/Lambda, update trigger
acdl-sod-halt SNS recreate topic, repoint publisher
acdl-ecs-sg SG recreate (brief ECS disruption)
alias/acdl-platform KMS alias repoint alias (cheap)
acdl-microservice (cluster/ECR/service/task/role) ECS+ECR re-push images, recreate service
acdl-spike-runner (user/policy) IAM re-bootstrap with new key
acdl-tfstate-581513795199-us-east-1 S3 state bucket terraform init -migrate-state, back up state JSON
acdl-alb (name prefix) ALB recreate (brief downtime)

Finding 4 — Consumer/infra conventions (judgment category, D-104)

  • AWS tag keys acdl:owner|environment|contract|cost-center|ref (5 keys, ~109 tag assignments) — matched by ABAC session policies. Parallel-tag period (add nova:*, swap policy, remove acdl:*).
  • SSM path /acdl/{env}/{contractId}/{output} (67 refs) — deploy outputs stored here. Migration script copies params, readers updated, old deleted.
  • Consumer path .acdl/contract.yml (23 refs) — consumer repos depend on this. Renamed .nova/contract.yml + migration guide.

Finding 5 — Docs & decks (mechanical)

  • README.md (16), docs/index.md, docs/vision.md, docs/architecture.md, docs/consumer-guide.md (35), docs/modules/index.md (28), all .ciagent/.md, modules/STANDARDS.md, schemas/README.md, pipelines/README.md, adapters/README.md, terraform//README.md.
  • Deck markdown + mermaid .mmd sources (5 files) + rendered HTML. PNGs re-exported from edited .mmd sources.
  • S&P visual theme (sp-theme.json, deck CSS) is client branding — D-107: untouched. Only product-brand text (ACDL→Nova) changes.
  • Schema $id URLs (https://acdl.cloudinit.dev/schemas/...) → https://nova.cloudinit.dev/schemas/... (D-110: illustrative, no DNS resolution needed for validation).

Finding 6 — CI / pipeline / release

  • Workflow files mirrored in .gitea/workflows/ + .github/workflows/ (modules-lifecycle 31×, deploy 22×, ci 2×, release 3×).
  • release.yml release title ACDL vX.Y.ZNova vX.Y.Z (forward only; past releases keep names).
  • Git branches/tags use milestone/v* / phase/* / v*no brand name present, no change needed (D-112: flat-branch convention preserved).
  • config.json release.gitea.repo stays acdl (D-105: real repo name unchanged; doc URLs illustrative only).

Finding 7 — External / URLs

  • github.com/acdl/... (~20 refs in docs + module READMEs + reusable-workflow uses: refs) — D-105: illustrative, updated to nova for prose. Real GitHub org/repo rename is out of scope.
  • git.cloudinit.dev/continuous-intelligence/acdl* (incl. sister repos acdl-contracts, acdl-evidence) — updated in prose to nova*.
  • README has no badges (no shields.io, no img src).

Finding 8 — Nomenclature / tagline

  • "DevSecOps", "New Dawn", "enabler" appear nowhere in the repo today — clean insertion, no collisions to reconcile (D-106).
  • Current tagline ("North Star" / "consumers declare intent") is retained; Nova tagline added alongside.
  • "bottleneck" (6 occurrences in deck talking points) — compatible with the Nova "no bottleneck" messaging; left in place.

Persona assessment (v1.15)

No new personas needed for v1.15 — the rebrand touches existing territories (docs, code, terraform, CI, tests). The active roster: lead-developer (docs/decks/.ciagent meta + verification + migration runbooks), backend-engineer (core/env.py dual-read helper, contract resolver path, Lambda, output_publisher, regression_verify), data-engineer (terraform resource names/tagging, state bucket migration, DynamoDB data migration, ECR re-push, schemas/tagging-standard). The frontend-engineer remains deactivated (no UI; decks are markdown = lead-developer territory). A security-engineer persona is not activated — the ABAC session-policy + tag-key migration (REQ-162) is data-engineer territory (terraform IAM) with lead-developer review. Territory enforcement = warn (co-authoring expected at the core/env.py + terraform boundary, and the contract-resolver + deploy-workflow boundary).

Assumptions logged

  • A1 (confidence 0.9): No live AWS access is available during P0P4 execution (the acdl-spike-runner IAM user's creds are in .env.secrets but live apply/modify/destroy is gated by NOVA_LIFECYCLE_MODE defaulting to plan-only). The terraform changes are validated via terraform validate; live apply is exercised by the modules-lifecycle workflow when explicitly set to full. This matches the v1.11v1.14 established pattern.
  • A2 (confidence 0.85): .env.secrets contains live rotated AWS credentials keyed by ACDL_AWS_*. P2 renames the KEYS only (values stay). The runtime reads via the new core/env.py dual-read helper (NOVA_AWS_ACCESS_KEY_ID preferred, ACDL_AWS_ACCESS_KEY_ID fallback), so no re-rotation is needed until P5 removes the fallback.
  • A3 (confidence 0.8): The Gitea release API (POST .../releases) is reachable for v1.14.x tags (the v1.14 milestone shipped releases through v1.13.24 / release id 285). P0 ship targets v1.14.0.

v1.16 NFR Simplification — Research Addendum (2026-07-30)

Milestone: v1.16-Nova-Simplification (NFR). Research is codebase- grounded (not domain/ecosystem) — the two explore passes identified concrete, file:line-verified residual debt the v1.15 rebrand left, plus genuine simplification and the onboarding request-path scaffolding.

R1. v1.14 NFR categories already closed (do NOT re-propose)

v1.14 (REQ-135..154) swept: over-broad excepts (REQ-141), hardcoded account-ID externalization (REQ-142, ACDL_AWS_ACCOUNT_ID env + live fallback G-102), IAM Resource:"*" scoping to nova-* ARNs (REQ-143, G-104), contractId/env/error validation (REQ-144), additionalProperties:false schemas (REQ-145), .gitignore credential catch-all (REQ-146), Kyverno --kube-version removal + deferral doc (REQ-147, G-103), orphan bytecode/dead-config cleanup (REQ-148), 7 untested-script test coverage (REQ-149), set -euo pipefail parity + Gitea workflow parity (REQ-150), config/persona/backend hygiene (REQ-151), STANDARDS.md TYPE_MAP consistency (REQ-152), ARCHITECTURE/COST/GRILL doc sync (REQ-153), platform-VPC CIDR/subnet parameterization (REQ-154).

The v1.16 grill (G-101..G-106) and escalation E-001 are all CLOSED. v1.16 finds NEW residual signals (D-117).

R2. Fresh debt the v1.15 rebrand left (verified file:line)

High-severity correctness regressions (P1):

  • adapters/terraform/adapter.py:117 — emits state_bucket = f"acdl-tfstate-{account_id}-us-east-1". The live state bucket was renamed to nova-tfstate-* in v1.15 P4 (REQ-163), but the adapter's emitted terraform backend still references acdl-tfstate-*. In plan-only mode this is latent (no real init against the bucket), but a full-mode lifecycle run would point at a non-existent bucket.
  • adapters/kyverno/policies/require-resource-labels.yml:6,21,25,33,37 — enforces acdl:owner/acdl:environment labels. nova_tagging.py hard-fails on any acdl:* key post-P5 (REQ-164). The Kyverno policy contradicts the Nova tagging standard.

User-facing brand misses (P2):

  • core/environment_check.py:59,61 — onboarding message header/body say "ACDL Environment Onboarding" / "ACDL environments are platform- managed" (user-facing).
  • core/lambda/contract_ingestor.py:145,191 — GitHub issue alert title [ACDL-ALERT] + body "auto-created by the ACDL platform Lambda" (user-facing artifact).
  • scripts/post_stage_comment.sh:39,46 — PR comment header "ACDL Stage"
    • footer "ACDL deploy pipeline" (user-facing).
  • scripts/run_ci.sh:39 — CI banner "ACL CI Pipeline".
  • Module docstrings: core/contract_resolver.py:1,474, core/confidence_signal.py:1, adapters/terraform/adapter.py:1, adapters/kyverno/kyverno_adapter.py:1, adapters/wiz/wiz_adapter.py:1, adapters/README.md:1, adapters/kyverno/README.md:4,18.

Dead code + stale comments (P3):

  • scripts/run_platform.sh:153export ACDL_ENVIRONMENT_OVERRIDE="$ENVIRONMENT_OVERRIDE" # legacy fallback, removed in P5 — comment says "removed in P5" but the line is STILL present (dead code, P5 already shipped).
  • Stale dual-read comments across core/local_emulators.py:15-16,503,505, core/regression_verify.py:318-319,333, scripts/run_regression.sh, scripts/run_lifecycle_*.sh (reference the retired G-106 fallback).
  • acdl_* temp-dir prefixes: core/local_emulators.py:71,252, core/regression_verify.py:183,234, scripts/run_pattern_plan.sh:29, scripts/run_primitive_plan.sh:29, scripts/run_lifecycle_*.sh:36,41.

R3. Simplification opportunities (verified)

  • core/regression_verify.py:328-409_check_live_terraform_plan_ microservice + _check_live_terraform_plan_static_assets are ~95% identical (resolve → adapt → init → validate → plan). Extract _check_live_terraform_plan(contract, label) (~35 lines saved).
  • core/regression_verify.py:149-178_check_resolver_static_assets + _check_resolver_microservice identical except contract path. Extract _check_resolver(contract).
  • core/regression_verify.py:477-503 — duplicated lifecycle-contract- resolve block. Extract _assert_contracts_resolve(module_dir).
  • scripts/run_platform.sh:336-350 + :452-466 — duplicated HITL attestation block. Extract run_hitl_gate() shell fn (~14 lines).
  • core/contract_resolver.py:50-68 duplicates core/environment_check.py: 36-54 env loader verbatim. Import instead.
  • scripts/run_platform.sh:145-146 — hardcoded CONTRACT_ID UUID + WORK="/tmp/acdl_platform_run_v18" (v18 stale). Make config/env- derived.
  • .gitea/workflows/.github/workflows/ — 3 byte-identical pairs (ci.yml, deploy.yml, modules-lifecycle.yml, ~20 KB) maintained by hand + a test asserting identity. Generator (D-115) eliminates manual-sync risk.
  • core/contract_resolver.py:540is_l2 = "l2" in interface_path or "composition" in interface_path fragile string heuristic. Add kind to registry entries (P7).
  • scripts/run_platform.sh (610 lines) — decommission block (:180-237)
    • uptime block (:520-606) are self-contained. Extract to scripts/run_decommission.sh + scripts/run_uptime.sh (P9).

R4. Security gaps (verified, NEW — not v1.14 duplicates)

  • core/lambda/contract_ingestor.py:251-252if not caller_arn: pass silently skips identity validation when IAM identity absent; relies on ABAC layer only (no defense-in-depth). Fail closed instead (P10).
  • core/lambda/contract_ingestor.py:269valid_envs = {"dev","qa", "prod","dr"} hardcoded; the core/environments/ dir is the source of truth. Derive from the directory (P10).
  • core/lambda/contract_ingestor.pysubmit_contract checks the contract key exists but never validates the blob's size or schema. Unbounded payload → DynamoDB write amplification. Size cap + schema validation (P11).
  • scripts/migrate_ssm_paths.py:113except Exception: pass (claims ParameterNotFound but catches all). Last true broad-swallow. Narrow to ParameterNotFound (P4).
  • core/output_publisher.py:112,182except Exception in publish_to_ssm + post_github_comment swallow all (not narrowed by REQ-141 which targeted 6 other sites). Narrow to specific exceptions.

R5. Onboarding request-path scaffolding (already present)

The infrastructure for a zero-human request path already exists:

  • terraform/platform/main.tf:153-183 deploys contract_ingestor Lambda
    • Function URL (IAM auth).
  • core/lambda/contract_ingestor.py:325-362 dispatches submit_contract | report_error | validate_change_request. Adding onboard_consumer is a small extension (P18, D-119: writes a pending CMDB row, no provisioning).
  • terraform/platform/consumer_invoke_policy.json is the ABAC policy template (aws:PrincipalTag/nova:owner == ${consumerRepo} scoped lambda:InvokeFunctionUrl).
  • core/environments/*.json are static JSON templates with placeholder account_id: "000000000000" — auto-generation from a request is straightforward (P19).

Missing for "no humans": (a) onboard_consumer action + onboarding schema (P18), (b) core/onboarding.py to auto-generate <env>.json + emit a PR (P19), (c) cross-account deploy-role + ABAC tag Terraform, offline-proven (P20, D-114). Real AWS account/network/state creation stays a future feature (D-113).

R6. Developer experience gaps

  • scripts/run_platform.sh has no --help (:82 rejects --* flags). --deploy-uptime (:532) is undocumented in the header. --local is absent from the README (P15).
  • No .github/workflows/README.md cataloging the 7 workflows' inputs/ secrets/triggers (P16).
  • No single getting-started path; README "How to run" lists 3 manual bootstrap steps. The offline happy path (run_ci.sh + run_platform.sh --check-only/--local) is not surfaced first (P17).

Assumptions logged (v1.16)

  • A1 (0.9): No live AWS access during execution (consistent with v1.11v1.15). NOVA_LIFECYCLE_MODE defaults to plan-only; terraform changes validated via terraform validate. The state-bucket drift (P1) is latent in plan-only mode but must still be fixed for correctness.
  • A2 (0.85): The onboard_consumer Lambda action (P18) is offline- testable via moto / the local Lambda stub (D-092), consistent with the existing submit_contract/report_error test pattern.
  • A3 (0.8): The workflow generator (P8, D-115) must preserve the byte-identity property as a test assertion (generated outputs match committed files), not lose it — the dedup is mechanical, not a semantic change to the workflows.
  • A4 (0.85): The regression gate (D-091, D-118) at P9 and P21 confirms "simplify without regressions" — 22/22 capabilities must stay Verified. The gate is the credible control for the simplification wave.

v1.17 Research — Strategic Direction, Leadership Metrics & Unified Story

Phase: research (P0). Milestone: v1.17. Status: research. Researcher: ci-researcher + explore agent (signal inventory). Autonomy: full. Decisions D-120..D-132 locked in the planning conversation (PROJECT.md). NORTH_STAR.md drafted (pending GRILL).

1. Telemetry Signal Inventory (grounding audit)

Methodology: every claim below is grounded in a concrete file path + line number in /root/acdl. No speculation. The explore agent performed a full sweep of the repo. The finding: Nova has no metrics/telemetry/ dashboard aggregation layer today. What exists is a set of discrete, structured, file-based signal artifacts (JSON reports, JSONL logs, hash-chained outbox events, PR comments, Checkov JSON) plus unstructured stdout logs. A metrics milestone must aggregate these existing signals — it must not invent new ones without first adding emitters.

(a) Signals that EXIST TODAY and are STRUCTURED (groundable)

Signal File / Emitter Schema Persistent?
Regression report (22 caps, status, duration_ms, gate) .ciagent/REGRESSION_REPORT.jsoncore/regression_verify.py:643-667 regression_verify.py:82-91 Yes (committed file)
Regression report (markdown mirror) .ciagent/REGRESSION_REPORT.md same Yes
Checkpoint (milestone/phase/tag/regression summary) .ciagent/CHECKPOINT.json (CIAgent-managed) ad-hoc Yes
PolicyCheckResult list (per-rule pass/fail/severity/resourceRef) $WORK/pcr.jsonrun_platform.sh:395 + checkov_adapter.py:50-71 schemas/policy_check_result.schema.json No (ephemeral /tmp/)
Confidence signal (score, band, perInput, reasonCodes) $WORK/signal.jsonrun_platform.sh:412-426 + confidence_signal.py:60-65 confidence_signal.py:60-65 No (ephemeral)
Outbox event (hash-chained, CONFIDENCE_COMPUTED) $WORK/event.json + $WORK/outbox_item.jsonrun_platform.sh:444-459 + outbox_writer.py:44-56 audit_ledger_design.md:44-45,81-97 No (ephemeral; live DynamoDB torn down D-096)
Resolved Target Stack $WORK/stack.jsoncontract_resolver.py:581-603 schemas/stack.schema.json No (ephemeral)
Lambda return bodies (submit/report_error/validate_cr/onboard) core/lambda/contract_ingestor.py:171,265,284,392,446 ad-hoc JSON No (Lambda not live; local stub only)
DynamoDB CMDB rows (submitted/pending contracts) nova-contracts table ← contract_ingestor.py:160-170,433-445 ad-hoc No (table torn down D-096)
SSM parameters (deploy outputs) /nova/<env>/<contractId>/<name>output_publisher.py:123-156 ad-hoc No (live AWS, torn down)
PR stage comment (mode, runId) GitHub PR API ← post_stage_comment.sh:34-48 + deploy.yml:141 markdown table Yes (GitHub)
PR deploy-outputs comment GitHub PR API ← output_publisher.py:159-189 markdown table Yes (GitHub)
GitHub issue (deploy failure alert) GitHub API ← contract_ingestor.py:179-290 + deploy.yml:143-152 issue body Yes (GitHub)
Local E2E result (stack_name, tier, outbox_events, chain_verified, lambda_status) stdout JSON ← core/local_emulators.py:498-508,519 ad-hoc No (stdout)
HITL gate result core/hitl_gates.py:87,90 + run_platform.sh:179-185 stdout HITL PASS/BLOCK No (stdout)
Attestation matrix result core/attestation_matrix.py:184,187 stdout ATTESTATION PASS/BLOCK No (stdout)
Cost figures .ciagent/COST.md (manual Cost Explorer query) markdown table Yes (manual, not automated)

(b) Signals that EXIST but are UNSTRUCTURED (log-only)

Signal Source Format
CI pipeline result scripts/run_ci.sh:70-71 stdout banner === CI PIPELINE OK ===
Platform stage banners + summaries scripts/run_platform.sh:222,241,258,263,315,383,411,442,463,490,496 stdout === Step N: ... === + summary lines
Terraform init/validate/plan/apply/destroy logs $WORK/tf-*.logrun_platform.sh:320,324,328,352,375 raw terraform stdout (via tee)
Lifecycle test results scripts/run_lifecycle_test.sh etc. exit code only (no report file)
Decommission step counts scripts/run_decommission.sh:40,54 stdout decommission step N: M resources...
Uptime endpoint count scripts/run_uptime.sh:72,87 stdout uptime: N endpoint(s) to monitor
Onboarding prompt core/environment_check.py:57-81 stdout text block
Pytest results pyproject.toml:25 (-v --tb=short) stdout only (no junit/json)
sync_workflows result scripts/sync_workflows.py:56,53 stdout OK: 3 workflow pairs match / DRIFT: ...

(c) Proposed executive metrics with NO grounding today (DEFERRED)

Proposed metric Why no grounding Controlling decision
Live infrastructure health (ECS running count, ALB 5xx, RPS) Live AWS torn down; CAP-013..016 Skipped D-096
Live outbox write rate / ledger append latency DynamoDB outbox table absent D-096
Tamper-evident ledger checkpoint count / JWS signature rate S3 Object Lock + JWS + async worker deferred D-083
Onboarding funnel: requested → granted conversion Only "requested" (pending row) is emitted; no grant event D-113, D-114, D-119
Time-to-provision (onboarding SLA) Real AWS provisioning deferred D-113
Cross-account role grant count Offline-proven only, no live apply D-114
Drift detection (scheduled terraform plan -detailed-exitcode) Needs live AWS workspaces + a scheduler Nova doesn't have D-096 + no scheduler
GreenOps / carbon (WattTime/Electricity Maps API) No grounding; new external API future emitter
Predictive vs Reactive ratio Requires an ML anomaly-forecasting service future emitter
Multi-cloud normalization (Azure/GCP/K8s, FOCUS spec) Nova is AWS-only future
Red Team MTTR No red-team program exists future
Self-healing velocity Nova has no auto-remediator future emitter
SLA / unplanned downtime Needs live service uptime monitoring against SLOs D-096
Per-module lifecycle success rate over time No structured report file written; only exit code gap (no decision)
Test pass rate / test count time-series No junit/json reporter configured gap (add --junitxml to addopts)
Code coverage trend pytest-cov installed but not in addopts gap
Deploy frequency / lead time / MTTR (DORA) No deploy-event emitter; pipeline runs not counted gap
Policy pass rate time-series pcr.json emitted but ephemeral; not persisted gap (D-096 blocks live persistence)
Confidence score distribution over time signal.json emitted but ephemeral gap
Consumer adoption count / active consumers PROJECT.md:487 explicitly states "0 consumer adoption today" honest scope
Cost time-series (automated) COST.md is a one-shot manual query; no automated emitter gap

Bottom line: the single richest existing structured signal is .ciagent/REGRESSION_REPORT.json (22 capabilities × {status, tier, duration_ms, detail} + summary counts + boolean gate). The next richest is the per-run $WORK/*.json family (pcr.json, signal.json, event.json, stack.json) — but these are ephemeral and not persisted in CI. The lowest-friction grounding for a "no-humans" dashboard is therefore: (1) regression report → capability health, (2) PR comments + GitHub issues → deploy/failure activity, (3) add --junitxml to pytest → test trend, (4) persist $WORK/*.json → policy/confidence/outbox time-series, (5) extend outbox_writer → Decision Ledger, (6) add Infracost → pre-apply cost estimates.

2. Telemetry Reference Architecture (Nova-native adaptation)

The PO provided a full distributed-system telemetry reference architecture (CloudEvents 1.0 envelope, OpenTelemetry SDK, Kafka/NATS event bus, Prometheus hot path, ClickHouse warehouse, QLDB decision ledger, Infracost, drift detection, ML anomaly forecasting). Per D-120, we adopt the principles but implement with Nova-native minimal tech. The mapping:

Direction's principle Nova-native implementation (v1.17)
Events are the source of truth; dashboards are projections Hybrid (D-125): existing file signals stay as files; collector reads them and emits normalized CloudEvents into metrics/events.jsonl + SQLite. New emitters emit CloudEvents directly.
Every AI action is logged with confidence + alternatives Decision Ledger (D-121): outbox_writer.py extended → SQLite append-only hash-chain table. ai.decision.made modeled from confidence_signal (D-122): decision_id=run_id, chosen_action=band, confidence=score, alternatives=perInput, human_override=HITL block.
Hot/cold storage split Cold-only SQLite (D-126): metrics/nova_metrics.db. Hot path deferred (no live ops, D-096).
Read-only external integrators Infracost (pre-apply, offline, reads plan JSON). Cloud billing CUR deferred (D-096). Carbon APIs deferred (future).
CloudEvents 1.0 envelope Adopted. core/metrics/event_envelope.py defines the envelope + platform.* semantic conventions.
Decision Ledger = append-only with hash chain + outcome backfill SQLite append-only table with hash chain (D-121). Outcome backfilled from apply.completed via decision_id → request_id correlation. Honors D-083 (no S3 Object Lock/JWS).
Cost governance: mandatory tags + Infracost pre-apply Nova already enforces nova:* tags (nova_tagging.py, hard mode). Infracost added as plan post-processor (D-120). Post-apply CUR deferred (D-096).
Definition-of-success docs for every KPI Per-KPI docs in docs/metrics/ (D-127).
Replay-ability SQLite store + JSONL event log are replayable by design.

CloudEvents envelope (Nova-native)

{
  "specversion": "1.0",
  "id": "<uuid>",
  "source": "nova.platform",
  "type": "nova.run.completed",
  "time": "<ISO8601>",
  "subject": "<contractId>/<env>",
  "datacontenttype": "application/json",
  "platform": {
    "tenant_id": "acdl",
    "run_id": "run-<epoch>",
    "contract_id": "<uuid>",
    "environment": "dev|qa|prod|dr",
    "actor": {"type": "confidence-gate", "id": "confidence_signal"},
    "trace_id": "<run_id>"
  },
  "data": {
    "duration_ms": 4800,
    "stages": ["resolve", "adapt", "validate", "plan", "apply"],
    "exit_code": 0,
    "confidence": {"score": 0.94, "band": "pass", "perInput": {...}},
    "policy": {"passed": 12, "failed": 0, "skipped": 0},
    "hitl": {"gate": "dev", "result": "autonomous", "block": false},
    "cost_estimate_usd": -12.40,
    "decision_id": "run-<epoch>",
    "outcome": "succeeded"
  }
}

Core event types (Nova-native minimum viable set)

Event type Emitted by Purpose Grounding
nova.run.started run_platform.sh Measures demand; provisioning lead time start new emitter (P1)
nova.run.completed run_platform.sh Run count, stage durations, exit, MTTR new emitter (P1)
nova.run.failed run_platform.sh Failure count, MTTR numerator new emitter (P1)
nova.policy.evaluated checkov_adapter.py Policy pass rate, compliance KPIs grounded (pcr.json → P1 persists)
nova.confidence.computed confidence_signal.py Confidence distribution, decision accuracy grounded (signal.json → P1 persists)
nova.ai.decision.made outbox_writer.py (extended) Decision Ledger entry grounded (D-121, D-122)
nova.attestation.recorded hitl_gates.py Attestation Coverage, human-in-the-loop audit grounded (D-132)
nova.cost.estimated Infracost post-processor Pre-apply cost estimate new emitter (P1, Infracost)
nova.capability.verified regression_verify.py Capability health, regression gate grounded (REGRESSION_REPORT.json)
nova.test.completed pytest (junit XML) Test count, pass rate new (P1 adds --junitxml)

3. Metric-to-Signal Scorecard (the "no fabrication" contract)

Executive metric (NORTH_STAR target) Status Source / formula Decision
Touchless Resolution Rate ≥99% grounded (after P1) runs without operational HITL block ÷ total runs (attestation gates excluded) D-122, D-132
Human Escalation Frequency <0.1% grounded (after P1) operational HITL blocks ÷ total runs (attestation sign-offs excluded) D-122, D-132
MTTR (p95) <60s grounded (platform-run) apply.failed.time → successful retry.time D-131
Predictive vs Reactive ≥3:1 deferred requires ML forecasting (future emitter) future
AI Decision Accuracy ≥99.5% grounded (after decision ledger) decisions not followed by apply.failed/incident within 5min D-121, D-122
Drift Auto-Reversal ≥95% deferred requires drift detection (D-096 + scheduler) D-096
Cloud Spend Reduction ≥25% partial pre-apply estimate grounded (Infracost); actuals deferred (D-096 CUR) D-120
L1/L2 Ops Hours Avoided ≥70% derived formula: run count × manual baseline minutes × blended rate D-127
Platform ROI ≥250% derived formula: (labor savings + cloud savings + avoided downtime) ÷ platform op cost D-127
Decision Ledger Coverage 100% grounded (this milestone) outbox_writer.py → SQLite hash-chain D-121
Attestation Coverage 100% grounded hitl_gates.py + outbox approver_* attributes; prod/dr D-132
AI-Agent Intent Share ≥40% future no AI-agent consumers today; placeholder view future
Capability health (18V+4S) grounded REGRESSION_REPORT.json existing
Confidence score distribution grounded (after P1) signal.json → decision ledger D-121
Policy pass rate grounded (after P1) pcr.json → persisted D-120
Test count / pass rate grounded (after P1) pytest --junitxml D-120
Provisioning Lead Time grounded (after P1) run.started → run.completed D-120
Deployment Frequency grounded (after P1) count(run.completed) per day D-120
Deploy-failure alert count grounded GitHub issues via Lambda report_error (D-055) existing
Cost figures (actuals) manual one-shot COST.md (Cost Explorer query) existing
FTE Hours Saved / TRV derived formula over run count + COST.md D-127
Self-healing velocity deferred no auto-remediator future
SLA / unplanned downtime deferred needs live service uptime (D-096) D-096
GreenOps / carbon deferred WattTime/Electricity Maps API (future) future
Red Team MTTR deferred no red-team program future
Multi-cloud normalization deferred Nova is AWS-only future
Live CUR reconciliation deferred needs live AWS billing (D-096) D-096

4. Deferred-Decision Ledger (constraints on this milestone)

Decision Scope Grounding impact
D-096 Live AWS torn down post-v1.11 BLOCKS all live-AWS metrics (CAP-013..016 Skipped; live outbox; live state bucket; live CUR)
D-083 S3 Object Lock + JWS + async worker deferred BLOCKS tamper-evident ledger; v1.17 uses SQLite hash-chain instead
D-113/D-114/D-119 Onboarding = request-path only; no auto-grant BLOCKS onboarding funnel "granted" half
D-091/D-118 Regression gate (D-091) gates milestone completion ENABLES the strongest metric signal (REGRESSION_REPORT.json)
D-092 Local emulating adapters ENABLES offline E2E metrics (CAP-011/012)
D-055 report_error Lambda action creates GitHub issues ENABLES deploy-failure alert metric
D-050 Publish deploy outputs to SSM + GitHub PR comment ENABLES outputs-published metric
D-054/D-043/D-109 Nova tagging standard (hard mode) ENABLES tagging-compliance metric
D-084 8-concern attestation matrix ENABLES attestation metrics (operator-supplied evidence artifacts)
D-089 Signature verification skipped when signing key unset (dev/CI) Signature metrics are no-ops in dev

5. Deck-Storytelling Research (x3 arc + per-slide benefit)

The "tell them x3" structure

The PO's direction: "Tell them what you're going to tell them, then tell them, then tell them what you told them." Applied at two levels:

Deck level (the 5-act arc):

  1. Opening slide = "what I'm going to tell you" — the full arc preview: Problem → Vision → How → Proof → Roadmap.
  2. Body (acts 15) = "tell them" — each act delivers its content.
  3. Closing slide = "what I told you" — recap of the 5 acts + the ask.

Per slide:

  1. Slide opens with what it'll cover (1 line: "This slide shows X").
  2. Slide delivers the content (bullets, diagram, or table).
  3. Slide closes with an explicit "benefit of this stage" callout (1 line: "Benefit: you now know Y" or "Why this matters: Z").

Fluidity conventions

  • Transitions are written, not hand-waved. Each slide's opening line references the previous slide's close ("Having seen X, now consider Y").
  • No disjointed jumps. If a topic shift is needed, a bridge slide or a transition sentence carries the audience across.
  • The arc is visible. A small "act indicator" in the Marp footer (e.g., Act 3/5: How it works) keeps the audience oriented.

Existing deck inventory (to be retired)

Two decks exist today in docs/presentations/:

  • how-the-platform-works.md (32,916 bytes) → marp → html → talking-points
  • the-developer-experience.md (27,509 bytes) → marp → html → talking-points

Both follow a 4-step process (source .md → Marp → HTML → talking-points) documented in docs/presentations/README.md. Per D-130, both are merged into one unified narrative deck and retired.

Grounded metrics already cited in existing decks

  • "22/22 auto-verifiable capabilities Verified" — stale vs current REGRESSION_REPORT.json (18V+4S post-D-096). The unified deck must derive this from the report, not copy the stale claim.
  • Confidence thresholds: dev ≥0.50, qa ≥0.75, prod ≥0.90, dr ≥0.95 — grounded in core/confidence_signal.py:57 (THRESHOLDS).
  • RPO = 0 (evidence write synchronous) — grounded in core/audit_ledger_design.md:27,103.
  • Cost figures — how-the-platform-works.md:461; cites COST.md.
  • Confidence signal 6 inputs + weights — grounded in core/confidence_signal.py:40-47.
  • "~80-line stateless adapter" vs "918-line monolith" — grounded in ROADMAP/RESEARCH prose.

Planned deck structure (for PLAN to detail)

The unified deck "Nova — The No-Humans Infrastructure Platform":

Act Slides Content Proof source
1. Problem 23 The no-humans imperative; why operators are the bottleneck; the trust gap NORTH_STAR vision
2. Vision/Direction 23 Nova's vision; 4 strategic objectives; anti-goals; the attestation model (autonomy in operations, human at stage gates) NORTH_STAR
3. How it works 34 Contract → resolver → adapter → confidence → HITL gate; the Decision Ledger; the 8-concern attestation matrix code grounding
4. Proof (metrics) 34 Capability health (18V+4S); confidence distribution; policy pass rate; Decision Ledger coverage; Attestation Coverage; cost estimates; the grounded/derived/deferred honesty model metrics export
5. Roadmap/Ask 2 1218mo targets (committed); deferred metrics (honest); the ask NORTH_STAR targets

Total: ~1216 slides. Opening = arc preview; closing = recap + ask.

6. Assumptions logged (v1.17)

  • A1 (0.9): No live AWS access during execution (consistent with v1.11v1.16). All metrics that require live AWS ship as placeholder views. The Infracost integration runs offline (reads plan JSON).
  • A2 (0.85): The Decision Ledger SQLite hash-chain is sufficient for v1.17's audit needs. The full tamper-evident ledger (S3 Object Lock + JWS, D-083) is a future milestone. The hash-chain provides append-only + integrity verification locally.
  • A3 (0.8): The "AI decision" framing (D-122) is honest: Nova's "AI" is the confidence-gated policy engine (confidence_signal + HITL gate), not an LLM planner. The deck and METRICS.md must frame this accurately — overclaiming "AI" would violate the "no fabrication" constraint.
  • A4 (0.85): The unified deck's "Proof" section cites only grounded metrics with real numbers. Deferred metrics are shown as "Planned" with the <span class="badge planned">Planned</span> badge. No fabricated numbers in any slide.
  • A5 (0.8): --junitxml + --json-report added to pytest addopts does not break the existing test suite (the flags are additive; pytest continues to run normally). CAP-009 (offline pytest suite passes) must remain Verified after the change.
  • A6 (0.75): Infracost is available as a CLI tool that can be installed in the CI environment and run locally. It reads terraform plan -out=plan.tfplan + terraform show -json plan.tfplan to produce a cost estimate. No live AWS access required. If Infracost is not available, the cost.estimated event is omitted (degraded mode, not a failure).

v1.18 Research — Citizen Developer & Production-Grade Guidance

Phase 0 RESEARCH. Autonomy = full. Findings are evidence-grounded (fetched from live sources, not assumed). The Atelier repo, the MCP Python SDK v2 docs, the existing Nova schemas/ingestor, and the Marp CLI README were all fetched directly. Decisions are logged with confidence scores; low-confidence items are flagged.

1. Atelier Integration Reference

1.1 The 8 core principles (C1C8)

Source: core/first-principles.md (fetched 2026-08-06 from https://git.cloudinit.dev/coreci/atelier/raw/branch/main/core/first-principles.md). Precedence is a total order — a lower-numbered principle is never sacrificed for a higher-numbered one (C1 never sacrificed; C2 only for C1; C3 only for C1/C2; C4C8 tradeable among themselves but always below C1C3).

ID Principle One-line description
C1 Correctness The system does what it is supposed to do, and nothing else. Highest principle; never overridden. Security is a subset (exploitable code is incorrect). Includes temporal correctness (a late answer is wrong when the deadline mattered).
C2 Clarity The intent of the code is obvious to its reader. Optimize for the reader; names reveal intent; comments explain why not what. Unclear code is where bugs hide.
C3 Simplicity The solution is as simple as possible, and no simpler. Complexity is the enemy of correctness; every line is a liability. Not laziness — the result of removing everything unnecessary.
C4 Locality Decisions and their consequences live near each other. State, logic, side effects that depend on each other live near each other. A change needing many distant files is a locality violation.
C5 Reversibility Every decision can be undone, and the cost of undoing is known. Migrations/deploys/schema/API changes reversible by default. Versioning, feature flags, rollback paths are the mechanisms.
C6 Composability Parts combine into wholes, and the parts are reusable in new wholes. A part that does one thing well composes; the boundary is its contract. Composable parts are understandable in isolation.
C7 Observability The system's behavior is visible to the people who must understand it. Logs/metrics/traces are first-class, designed in. An observable system answers "what/why/what next" without reading source.
C8 Economy The system uses no more resources than the task requires (time, memory, attention, money, complexity). Most tradeable principle; unbounded growth in any resource is a defect.

The precedence string (from core/first-principles.md §3): C1 Correctness > C2 Clarity > C3 Simplicity > C4 Locality > C5 Reversibility > C6 Composability > C7 Observability > C8 Economy.

Conflict resolution (core/conflict-resolution.md, fetched): a deterministic 6-step procedure. The hierarchy is core/first-principles.md > domains/<x>/first-principles.md > domains/<x>/<topic>.md > languages/<lang>.md > examples/<x>.md. Same-level conflicts resolve by core derivation (via the matrix), then by specificity, then by filing an issue (a tie is a defect). A domain's "non-tradeable" declaration (e.g. Security: 8 of 10) promotes those rules to C1-equivalent — a binding escalation recorded in the matrix's derivation.

1.2 The 19 domains and Nova-citizen-dev relevance

Source: matrix/principles-matrix.md (fetched) + the releases page (v0.4 milestone = 19 domains, 190 P-rules, confirmed in the v0.3.6 release notes and the matrix Coverage Summary).

# Domain Atelier path P-rules Nova-citizen-dev relevant? Reason
1 UI/UX domains/uiux/ 10 NO — excluded v1.18 has no frontend (Out of Scope: "A Nova-built frontend / dashboard"). Decks are markdown, not a UI.
2 API Design domains/api/ 10 YES A citizen developer building a web API / worker / scheduled job touches API contracts. Maps to skills/api.md.
3 Security domains/security/ 10 YES Zero-trust, input validation, secret hygiene, fail-securely — universal for any production-grade service. Maps to skills/security.md.
4 Data domains/data/ 10 YES Schema-as-truth, migration safety, referential integrity — applies to any stateful service. Maps to skills/data.md.
5 Testing domains/testing/ 10 YES Tests-as-specification, determinism, edge-case coverage — required for a citizen developer's UAT. Maps to skills/testing.md.
6 Performance domains/performance/ 10 YES (reference, not a skill) Measure-first, bounded operations, no N+1, timeouts. NOT one of the 9 REQ-221 skills; Performance principles are cited inside the 9 skills + the index.
7 Observability domains/observability/ 10 YES Structured logs, correlation IDs, no secrets in logs — the "basic observability bootstrap" BA.A skill. Maps to skills/observability.md.
8 Errors domains/errors/ 10 YES Errors are data, fail loudly + specifically, preserve context — production-grade error handling. Maps to skills/errors.md.
9 Documentation domains/documentation/ 10 YES (reference, not a skill) Docs-as-code, audience awareness, examples mandatory. REQ-221 does NOT list skills/documentation.md; a self-referential "documentation skill" is redundant. Principles cited inside docs/skills.md index.
10 Concurrency domains/concurrency/ 10 YES (reference, not a skill) Immutability, bounded queues, timeouts — advanced for a citizen developer's first 5 skills. REQ-221 does NOT list skills/concurrency.md. Top rules cross-referenced inside skills/api.md + skills/errors.md.
11 DevOps domains/devops/ 10 YES Reproducibility, rollback-first, config-as-code — the citizen developer co-owns Release Management (RACI). Maps to skills/devops.md.
12 Infrastructure as Code domains/infrastructure-as-code/ 10 YES Declarative intent, idempotence, plan-before-apply, no secrets in HCL — directly relevant to the Nova contract→Terraform path. Maps to skills/infrastructure-as-code.md.
13 Kubernetes domains/kubernetes/ 10 NO — excluded Nova emits Terraform (ECS/Fargate per the architecture), not K8s manifests. Kyverno adapter is "ready but inactive" (D-053). Not citizen-dev-relevant.
14 GitOps + Operators domains/gitops-operators/ 10 NO — excluded Nova uses a push pipeline (contract → resolve → plan → apply), not a pull-based reconciler. Not citizen-dev-relevant.
15 AI/ML domains/ai-ml/ 10 YES (reference, not a skill) Reproducibility, data versioning, drift detection — relevant to Nova itself (Nova is an agentic platform), but a citizen developer on Nova is NOT building ML models; they consume Nova's agentic capability. REQ-221 does NOT list skills/ai-ml.md; the Atelier AI/ML domain is platform-team guidance, not citizen-dev guidance.
16 i18n domains/i18n/ 10 NO — excluded Not relevant to a citizen developer's first production-grade service on Nova.
17 Compliance domains/compliance/ 10 YES Audit logs append-only, policy-as-code, evidence-by-operation — directly relevant (Nova's compliance posture is a selling point). Maps to skills/compliance.md.
18 Edge domains/edge/ 10 NO — excluded Nova does not deploy edge/CDN for the citizen developer's first 5 skills; Route53/ACM/CloudFront are consumer-supplied extension points (D-049).
19 Messaging domains/messaging/ 10 NO — excluded The citizen developer's first 5 skills (web API / worker / scheduled job / static asset / observability bootstrap) do not require a broker; messaging is a future capability.

Relevant count: 13 of 19 are relevant to some Nova audience (YES or YES-reference). Of those, 9 become skills (per REQ-221, the planned count). The other 4 relevant domains (Performance, Documentation, Concurrency, AI/ML) are reference-only — their principles are cited inside skills or the docs/skills.md index, but they do NOT get their own skill file. This matches REQ-221's exact 9-skill list.

Excluded count: 6 of 19 (UI/UX, Kubernetes, GitOps, i18n, Edge, Messaging) are not relevant to a Nova citizen developer building a production-grade application — confirmed.

1.3 Atelier domain → Nova skill mapping (final 9-skill list)

REQ-221 names exactly 9 skills. The research confirms the planned 9 — no adjustment needed. The mapping (each skill cites its Atelier source path + distills the citizen-developer-relevant subset + links to agent-checklist triggers + maps to the BA.A 5-skill catalog):

Nova skill file Atelier domain path P-rules distilled BA.A catalog skill it extends
skills/api.md domains/api/ P1 Contract Fidelity, P2 Clarity, P5 Versioning, P6 Idempotency, P8 Security, P9 Error Transparency web API
skills/security.md domains/security/ P1 Zero Trust, P2 Least Privilege, P4 Input Validation, P6 Crypto Correctness, P8 Fail Securely, P9 Secret Hygiene all 5 (cross-cutting)
skills/data.md domains/data/ P1 Truth, P3 Invariants in Schema, P4 Migration Safety, P7 Type Fidelity, P9 Referential Integrity web API, worker, scheduled job
skills/testing.md domains/testing/ P1 Tests as Specification, P3 Determinism, P5 Coverage of Behavior, P9 Edge Case Coverage, P10 No Test Theater all 5 (UAT is a citizen-developer RACI responsibility)
skills/observability.md domains/observability/ P1 Structured by Default, P2 Correlation, P6 No Secrets in Obs, P7 Actionable Alerts basic observability bootstrap
skills/errors.md domains/errors/ P1 Errors are Data, P2 Fail Loudly, P3 Fail Specifically, P4 Preserve Context, P5 Recoverable When Possible web API, worker, scheduled job
skills/devops.md domains/devops/ P1 Reproducibility, P4 Rollback First, P5 Progressive Delivery, P6 Config as Code, P8 Security at Every Layer scheduled job, worker (deploy/release is co-owned Release Mgmt)
skills/infrastructure-as-code.md domains/infrastructure-as-code/ P1 Declarative Intent, P2 Idempotence, P4 Plan Before Apply, P5 Version Everything, P10 Secrets Never in Code static asset (the contract→Terraform path)
skills/compliance.md domains/compliance/ P1 Audit Logs Append-Only, P2 Every Significant Action Logged, P4 Policy is Code, P5 Policy is Evaluated as a Gate, P9 Secrets Redacted in Audit all 5 (cross-cutting; Nova's compliance posture)

Final recommendation: 9 skills, exactly as REQ-221 planned. Confidence 0.95 — the planned list maps cleanly to the relevant Atelier domains and to the BA.A 5-skill catalog; the 4 "reference-only" domains (Performance, Documentation, Concurrency, AI/ML) are correctly not elevated to skills (a citizen developer's first production-grade service does not need a standalone Concurrency or AI/ML skill; Performance and Documentation principles are cited inside the 9 skills + the index).

1.4 Agent-checklist → MCP atelier.validate_against_principles checks

Source: review/agent-checklist.md (fetched). The checklist has a Core (C1C8) section (8 subsections, ~30 boolean items) plus domain-specific trigger sections (one per domain; Nova-relevant ones: API, Security, Data, Testing, Performance, Observability, Errors, Concurrency, DevOps, IaC, Compliance).

The MCP atelier.validate_against_principles tool (REQ-223, in plugins/validation.py) runs the relevant checklist items against a code/diff snippet. The tool input model:

class ValidateInput(BaseModel):
    snippet: str               # the code/diff to validate
    language: str               # e.g. "python", "terraform", "yaml"
    domains: list[str]          # e.g. ["security", "api"] — which domain triggers to run
    run_core: bool = True       # always run C1C8 unless explicitly skipped

The structured output model (Pydantic, returned as structured_content):

class Violation(BaseModel):
    principle: str              # e.g. "C1", "security/P9"
    checklist_item: str         # the verbatim checklist question
    severity: str               # "C1" (blocking) | "non-tradeable" | "tradeable"
    evidence: str               # the snippet substring + why it fails
    fix_hint: str               # the principle's remediation guidance

class ValidateResult(BaseModel):
    snippet_id: str             # hash of the snippet for replay
    passed: bool
    violations: list[Violation]
    domains_checked: list[str]
    core_checked: bool

Checklist → check mapping (the validation plugin encodes each checklist item as a boolean predicate over the snippet + language):

Checklist section MCP check behavior Nova-relevant?
C1 Correctness (4 items) Run all 4; any fail → severity: "C1" (blocking). YES — always run (core)
C2 Clarity (4 items) Heuristic checks: name smell (data/temp/x/doStuff), comment-why ratio. YES — always run
C3 Simplicity (4 items) Dead-code heuristic, function-length, premature-abstraction. YES — always run
C4 Locality (3 items) Cross-file-change heuristic (for diffs); within-file coupling. YES — always run
C5 Reversibility (3 items) Migration-has-down, deploy-has-rollback presence checks. YES — always run
C6 Composability (3 items) Single-responsibility heuristic, boundary-typed check. YES — always run
C7 Observability (4 items) Log-presence, error-context, metric, no-secrets-in-logs (hard check). YES — always run
C8 Economy (3 items) Unbounded-growth, no-timeout, resource-leak heuristics. YES — always run
If API 5 items: nouns-plural-lowercase, status codes, structured errors, schema validation, auth-required. YES — when domains includes "api"
If Security 5 items: no-secrets-in-code/logs/URLs, input-validation, output-encoding, vetted-crypto, authz-checked. All 5 are non-tradeable (Security domain §3). YES — when "security"
If Data 5 items: schema-reflects-domain, constraints-in-schema, migration-up-down, domain-types, no-SELECT-star. YES — when "data"
If Testing 4 items: independence, determinism, edge-cases, failure-specificity. YES — when "testing"
If Performance 4 items: no-unbounded, no-N+1, timeouts, cache-invalidation. YES — when "performance"
If Observability 4 items: structured-logs, correlation-id, no-high-cardinality, alerts-have-runbooks. YES — when "observability"
If Errors 4 items: not-swallowed, specific, context-preserved, recovery-attempted. YES — when "errors"
If Concurrency 5 items: shared-state-minimized, minimal-locks, bounded-queues, timeouts, cancellation. YES — when "concurrency"
If DevOps 4 items: pipeline-is-process, rollback-known, config-in-code, env-parity. YES — when "devops"
If IaC 8 items: declarative, pinned-providers, remote-locked-state, plan-before-apply, no-secrets-in-HCL, versioned-modules, drift-is-incident, least-priv-providers. YES — when "infrastructure-as-code"
If Compliance 10 items: append-only-audit, a-priori-action-set, retention-as-policy, policy-as-code, policy-as-gate, continuous-evidence, attributable-identity, subject-access, redacted-secrets, observable-posture. YES — when "compliance"

The validation plugin reads the vendored review/agent-checklist.md (frozen at the pinned tag — §1.6) so the checks are replayable against the exact checklist version that produced a result. The plugin maps each checklist line to a predicate function keyed by (language, principle) so a "no secrets in code" check runs differently for Python (ast scan for string-constant assignment) vs Terraform (HCL scan for hardcoded provider keys) vs YAML (scan for api_key: literals).

1.5 Principle-lookup query model

atelier.lookup_principle(domain: str, principle_id: str) (REQ-223, in plugins/principles.py) resolves a principle reference to its full text + core derivation + checklist items. Resolution model:

Input:

class LookupInput(BaseModel):
    domain: str         # "security" | "api" | "data" | ... | "core"
    principle_id: str   # "P4" | "C1" (core) | "P9"

Resolution path (the lookup algorithm):

  1. If domain == "core": load vendor/core/first-principles.md, parse the ### C<n>. <Name> section for principle_id (e.g. C1 → the "C1. Correctness" section). Return the full principle text.
  2. Else: load vendor/domains/<domain>/first-principles.md, parse the ### P<n>. <Name> section for principle_id (e.g. security/P4 → the "P4. Input Validation" section).
  3. Cross-reference the matrix: load vendor/matrix/principles-matrix.md, find the row for <Domain> P<n>, extract the Core column (e.g. Security P4 → C1). This is the core derivation.
  4. Cross-reference the checklist: load vendor/review/agent-checklist.md, find the If <Domain> section, extract the checklist items tagged with P<n> (the IaC section tags items with (P1), (P10) etc.; the Security section items map to P9, P4, P5, P6, P1/P10 by content).
  5. Check non-tradeable status: load vendor/domains/<domain>/first-principles.md §3 (Conflict Resolution); if the principle is listed as "never sacrificed", mark non_tradeable: true (escalates it to C1-equivalent per core/conflict-resolution.md §6).

Return (structured output):

class PrincipleLookup(BaseModel):
    domain: str                   # "security"
    principle_id: str             # "P4"
    name: str                     # "Input Validation"
    text: str                     # full principle body
    core_derivation: list[str]   # ["C1"]  (from the matrix)
    non_tradeable: bool           # True for security P1-P8, P9; False for P10
    checklist_items: list[str]    # the verbatim checklist questions for this P-rule
    source_path: str              # "domains/security/first-principles.md" (relative to vendor/)

Example resolution — atelier.lookup_principle("security", "P4"):

  • name: "Input Validation"
  • text: "All input is untrusted until proven otherwise. Validation happens at the boundary, against a schema, with explicit failure modes."
  • core_derivation: ["C1"] (matrix row: Security P4 → C1)
  • non_tradeable: true (Security §3 lists P4 as "never sacrificed")
  • checklist_items: ["Input is validated at the boundary", "Output is encoded for its context"] (from review/agent-checklist.md If Security)
  • source_path: "domains/security/first-principles.md"

The two companion tools:

  • atelier.list_domains() → returns the 19 domain names + their P-rule counts + relevance flag (the plugin hardcodes the Nova-relevance table from §1.2 so the citizen developer's agent can filter to the 13 relevant / 9 skill-bearing domains).
  • atelier.matrix_lookup(domain: str) → returns the full domain→core mapping for one domain (all 10 P-rules → their core C-rule(s)), used by validate_against_principles to set severity and by conflict resolution when two findings collide.

Recommendation: vendor tag v0.3.6 (the v0.4 milestone release).

Evidence (from https://git.cloudinit.dev/coreci/atelier/releases, fetched 2026-08-06):

  • The latest release is v0.3.6, dated 2026-08-05 16:22:58 +00:00, tagged v0.3.6 (commit 66b4767d25), marked Stable, with the title "v0.3.6 — v0.4 milestone: Edge + Messaging + Language-Derived Docs".
  • It is the v0.4 milestone release (the release notes state: "v0.4 — Edge + Messaging + Language-Derived Docs (Milestone Release). Tag: v0.3.6 (NFR milestone — final patch IS the deliverable; no separate minor tag per branch-strategy.md)").
  • The matrix is at its complete state: 19 domains, 190 P-rules (the Coverage Summary in matrix/principles-matrix.md confirms this exactly; the v0.3.6 release notes confirm "170 → 190 P-rules across 19 domains"). All 190 P-rules trace to ≥1 core C-rule (no orphans — verified in the release audit).
  • -11 commits to main since this release — there is post-release activity on main, which is exactly why pinning matters: vendoring main HEAD would be a moving target. v0.3.6 is the frozen, audited, reproducible snapshot. This satisfies D-136 (vendor for audit reproducibility) — an agentic validation result must be replayable against the exact principles that produced it.

Vendoring mechanics (for REQ-224):

  • mcp/atelier/vendor/ = a clean copy of the Atelier repo at tag v0.3.6 (the core/, domains/, matrix/, review/ directories — the docs the MCP tools read; examples/ and languages/ are optional but cheap to include for completeness).
  • mcp/atelier/vendor/VERSION.md records: tag v0.3.6, commit 66b4767d25, date 2026-08-05, milestone "v0.4 Edge + Messaging + Language-Derived Docs", P-rule count 190, domain count 19.
  • scripts/update_atelier_vendor.sh = a helper that takes a tag arg, fetches the tarball from https://git.cloudinit.dev/coreci/atelier/archive/<tag>.tar.gz, extracts the doc directories into mcp/atelier/vendor/, and updates VERSION.md. Intentional upgrades only (re-run + re-audit).

Confidence: 0.95. The only risk is that a v0.5 milestone lands before P5 ships — but the pinning model (VERSION.md + update script) makes a future upgrade a deliberate, audited action, not a silent drift.


2. MCP Python SDK v2 Reference

Source: https://py.sdk.modelcontextprotocol.io/ (the official Python SDK docs, fetched 2026-08-06) + the Tools page (.../servers/tools/) + the Structured Output page (.../servers/structured-output/). The docs document v2, the current stable release line (Python 3.10+).

2.1 Confirmed API patterns

  1. Server creation + import path. The v2 high-level server class is MCPServer (NOT FastMCP — that was v1; v2 renamed/restructured):

    from mcp.server import MCPServer
    mcp = MCPServer("atelier")   # one arg = server name
    

    This is the exact pattern shown in the docs' landing-page example and the Tools-page example. There is no FastMCP import in v2.

  2. @mcp.tool() decorator — inputSchema from type hints. Confirmed verbatim from the docs: "No JSON Schema. a: int, b: int is the schema." The SDK reads three things from the function:

    • name = the function name (search_books)
    • description = the docstring (the model sees this)
    • arguments = the type hints (query: str, limit: int) The SDK generates the JSON Schema and sends it during tools/list. Type hints are the contract — if a client sends "limit": "ten", the SDK rejects it before the function runs. Optional args = default values (limit: int = 10 → leaves required, gains default: 10). Richer constraints via Annotated[int, Field(ge=1, le=50, description="...")]. Enums via Literal["a", "b"]. Pydantic BaseModel parameter = structured "body" (nested as $defs).
  3. Multiple tools / dynamic registration (plugin-registry). The @mcp.tool() decorator is called on the mcp object. A plugin receives mcp and calls @mcp.tool() on it — this is plain Python decorator application, no registration magic. The plugin-registry pattern (D-140):

    # plugins/principles.py
    from mcp.server import MCPServer
    def register(mcp: MCPServer) -> None:
        @mcp.tool()
        def atelier_lookup_principle(domain: str, principle_id: str) -> PrincipleLookup:
            """Look up an Atelier principle by domain + ID."""
            ...
    

    server.py scans plugins/, imports each module, calls register(mcp). Each plugin's @mcp.tool() calls register the tool on the shared mcp object. This is the confirmed dynamic- registration pattern — no add_tool() API is needed; the decorator does it.

  4. stdio transport. The landing-page example shows uv run mcp dev server.py (Inspector). For stdio transport (D-135: stdio now), the server runs over stdio via the SDK's run entry point. The v2 server object supports stdio as the default transport. The exact run call is mcp.run() (the SDK handles the transport based on how the process is launched — stdio when invoked by an MCP host over stdio). The README's "no protocol handling" promise means mcp.run() is the only call needed. (HTTP transport is on the same server object — Out of Scope for v1.18, future milestone; the server object is transport-agnostic so adding HTTP later is a transport-only change, confirming D-135.)

  5. outputSchema / structured output. Confirmed: the return type annotation IS the output schema. From the Structured Output page: "the return type annotation is the output schema. It's published in tools/list as output_schema." A Pydantic BaseModel return type produces an unwrapped object schema (no result wrapper); a TypedDict or dataclass works identically. The result carries both content (text, for the model) and structured_content (data, for the application). Validation is enforced: whatever the function returns is validated against the schema before it leaves the server — a mismatch is a tool error (not a corrupt result). This is exactly what atelier.validate_against_principles needs: a ValidateResult(BaseModel) return type gives the host a structured violations list while giving the model a JSON-text rendering of the same object. structured_output=False opts out (text-only); we do NOT opt out for the validation tool.

  6. listChanged capability / dynamic tool registration. The v2 docs (Tools page + landing page) describe tool registration as declarative (@mcp.tool() at import time). The docs do NOT document a runtime listChanged notification API on the high-level MCPServer. For Nova's use case (plugins loaded once at server startup, not added/removed at runtime), this is fine — all 4 tools are registered before mcp.run(). A future milestone that adds tools at runtime would need the low-level Server (advanced/low-level-server/) for explicit notification control. Conclusion: no listChanged needed for v1.18; the plugin-registry loads at startup, before the stdio loop. Confidence 0.85 (the docs are silent on a high-level listChanged; the low-level server has it, but we use the high-level server).

2.2 Skeleton for mcp/atelier/server.py (P5 basis)

This is the 15-line pattern Nova's server should follow (the basis for P5 implementation):

import importlib, pathlib
from mcp.server import MCPServer

mcp = MCPServer("atelier")  # server name; stdio transport is the default

# Plugin-registry: scan plugins/, import each, call register(mcp).
for p in sorted(pathlib.Path(__file__).parent.glob("plugins/*.py")):
    if p.stem != "__init__": importlib.import_module(f".plugins.{p.stem}", __package__).register(mcp)

@mcp.tool()
def atelier_list_domains() -> list[dict]:
    """List the 19 Atelier domains with P-rule counts + Nova-relevance."""
    return [{"domain": "security", "p_rules": 10, "nova_relevant": True}, ...]

if __name__ == "__main__":
    mcp.run()  # stdio transport (D-135); HTTP-ready on the same object (future)

Notes on the skeleton:

  • MCPServer("atelier") — one import, one constructor arg (the name).
  • The plugin loop uses importlib + a register(mcp) convention (D-140). Each plugin's register body contains @mcp.tool() calls that register that plugin's tools on the shared mcp object. sorted() makes plugin load order deterministic (audit reproducibility — a plugin load order that changes between runs would break replay).
  • The sample tool shows the pattern: @mcp.tool(), type hints ARE the input schema, docstring IS the description, return type IS the output schema. The real atelier_list_domains returns a list[DomainInfo] (a list[BaseModel] → wrapped in {"result": [...]}, per the Structured Output docs).
  • mcp.run() — the single entry point; stdio is the default. No transport boilerplate. Adding HTTP later = a transport argument or a different run call on the same object (D-135, Out of Scope for v1.18).
  • The vendored Atelier snapshot (mcp/atelier/vendor/) is read by the plugin tool functions (not shown in the skeleton); the plugins load the markdown files lazily on first tool call and cache the parsed structure in module-level dicts (C8 Economy — don't re-parse the matrix on every lookup).

3. Submission-Readiness Gap Analysis

3.1 contract.schema.json defines SHAPE, not the readiness gate

Confirmed by reading /root/acdl/schemas/contract.schema.json (51 lines). The schema defines the contract shape only:

  • required: ["id", "name", "environment", "infrastructure"]
  • id: pattern ^[a-z][a-z0-9-]{2,5}$ (36 char acronym)
  • name: minLength 3
  • environment: enum ["dev", "qa", "prod", "dr"]
  • infrastructure: map keyed by module name, each entry has version (optional semver) + inputs (required, additionalProperties allowed)
  • additionalProperties: false (top-level + per-module)

What it does NOT define (the gap):

  • No tags field (the 5 required Nova tags per D-054)
  • No per-env mandatory metadata (the W3.E table: dev=stack+environment; qa+=e2eSuite+loadTest; prod+=runbook+dashboard+oncall; dr+=drDrillRef)
  • No policyPreconditions field (declared policy expectations)
  • No profile field (developer | agentic; agentic requires naturalLanguageIntent, confidenceAtSubmission, agentTrace)
  • No appSource field (repo + ref pointer for runtime fetch)
  • No contractId field at the top level (the ingestor payload has contractId in the Lambda envelope, but the contract blob itself does not — the readiness schema promotes it to a required field per REQ-217)

The schema's own description confirms this is the shape: "A consumer contract declares intent: which infrastructure to deploy, in which environment, with which inputs." It is the intent shape, not the ready-to-start gate.

3.2 The readiness schema is a SUPERSET gate ABOVE contract-schema validity

Confirmed by PROJECT.md (lines 635643, the v1.18 scope statement) and REQ-217. The relationship:

contract.schema.json  (SHAPE — id/name/environment/infrastructure)
        ▲
        │  references but does NOT redefine contract fields
        │
submission-readiness.schema.json  (GATE — superset above shape validity)
        = contract-shape-valid  (delegate to contract.schema.json)
        + tags                  (5 required Nova tags, D-054)
        + per-env mandatory     (W3.E table)
        + policyPreconditions   (declared policy expectations)
        + profile               (developer | agentic + agentic markers)
        + appSource             (repo + ref pointer)
        + contractId            (non-empty, promoted to required)

PROJECT.md hard constraint (line 674676): "The submission-readiness schema is a superset gate above contract.schema.json, NOT a duplicate — it references but does not redefine contract fields."

This means submission-readiness.schema.json uses $ref to contract.schema.json for the contract shape (or validates the contract blob against it as a first step), then adds the gate fields alongside it. The validator (REQ-218) calls contract.schema.json validation first (the existing _validate_contract_schema in the ingestor), then the readiness checks. This is a two-layer gate, not a merged schema.

3.3 Fields the new schemas/submission-readiness.schema.json must add

Per REQ-217 + W3.E (PROJECT.md line 888) + D-054 (tagging standard):

Field Type Required Source / rule
contractId string (non-empty) YES REQ-217. Promoted from the Lambda envelope to a contract-level required field.
environment enum dev/qa/prod/dr YES Already in contract.schema.json; the readiness schema references it (does not redefine) and uses it to select the per-env mandatory set.
tags object YES D-054 / schemas/tagging-standard.json. Required keys: nova:owner, nova:contract, nova:environment, nova:cost-center (nova:ref optional). The readiness schema references tagging-standard.json's required_tags shape.
policyPreconditions object (map of string→boolean/string) YES REQ-217. Declared policy expectations the platform will enforce (e.g. {"public-ingress": false}).
profile enum developer | agentic YES REQ-217 / W3.E.
profile == agentic → requires: naturalLanguageIntent (string), confidenceAtSubmission (number 01), agentTrace (object/string) per W3.E conditional REQ-22 / W3.E. These are "optional everywhere" per W3.E (a developer profile omits them) but required when profile is agentic.
appSource object {repo: string, ref: string} YES REQ-217. Repo + ref pointer for runtime fetch.
Per-env mandatory (W3.E):
dev stack + environment YES W3.E. (These are the base contract fields; the readiness schema enforces their presence for dev.)
qa adds validation.e2eSuite + validation.loadTest YES for qa W3.E.
prod adds runbook + dashboard + oncall YES for prod W3.E.
dr adds drDrillRef YES for dr W3.E.
inputs map object optional everywhere W3.E ("inputs map is always optional").

The per-env mandatory table is a conditional allOf in JSON Schema draft 2020-12: an if/then keyed on environment that requires the env-specific fields. The reason code ENV_MISSING_MANDATORY:<env>:<field> (REQ-218) maps directly to this conditional check.

3.4 How contract_ingestor.py currently works (P3 wiring point)

Read /root/acdl/core/lambda/contract_ingestor.py (502 lines). The current entry point + dispatch:

  • Entry point: lambda_handler(event, context) (line 460). Parses event["body"] (JSON string) → payload. Reads action (default "submit_contract").
  • Identity validation: _validate_caller_identity(event, payload) (line 293) — checks IAM caller ARN, consumerRepo format, contractId format (regex ^[a-zA-Z0-9][a-zA-Z0-9_-]{0,63}$), environment enum (from core/environments/*.json, P10/REQ-174), error-length cap. Fails closed if no IAM identity (P10).
  • Action dispatch (line 475):
    • submit_contract_submit_contract(payload) (line 135): validates required fields (consumerRepo, contractId, contract, environment), size-caps the contract blob (256 KB, P11/REQ-175), calls _validate_contract_schema(contract) (line 57 — validates against schemas/contract.schema.json via jsonschema; no-op if schema/jsonschema unavailable; bypassed by NOVA_LAMBDA_LOCAL_BYPASS), writes to DynamoDB nova-contracts (PK consumerRepo, SK contractId#submittedAt).
    • report_error_report_error (D-055, GitHub/Gitea issue).
    • validate_change_request_validate_change_request (REQ-93).
    • onboard_consumer_onboard_consumer (P18/REQ-182, validates against schemas/onboarding.schema.json).
  • Error mapping: ValueError → 400 (or 401 for identity failures); other Exception → 500 (defensive top-level guard, pragma: no cover).

Where P3 adds --check-readiness (D-133):

The ingestor is a Lambda handler, not a CLI. D-133 says the validator is "invoked as contract_ingestor.py --check-readiness subcommand" — this is a local CLI mode for citizen-developer pre-flight validation, NOT a new Lambda action. The implementation pattern (confirmed by the existing code structure):

  1. Add a if __name__ == "__main__": block at the bottom of contract_ingestor.py that parses sys.argv (argparse or manual). The existing file has NO __main__ block (it's Lambda-only); P3 adds one.
  2. The --check-readiness subcommand loads a contract file (or reads stdin), validates it against schemas/submission-readiness.schema.json (REQ-217) via the new core/submission_readiness.py validator (REQ-218), and prints a structured ReadinessResult (pass/fail per check + reason codes).
  3. The validator (core/submission_readiness.py) calls _validate_contract_schema(contract) first (reusing the existing function — the shape gate), then runs the readiness checks (tags, per-env mandatory, policyPreconditions, profile:agentic markers, appSource).
  4. On fail → the CLI exits non-zero with a citizen-developer-facing error (not a stack trace) — REQ-218. On pass → proceeds to existing ingestion (in the Lambda path, the readiness check would be a pre-write gate; in the CLI path, it's a pre-flight check that returns 0).

Reason codes (REQ-218, the validator's return vocabulary): MISSING_TAGS, ENV_MISSING_MANDATORY:<env>:<field>, AGENTIC_MISSING_INTENT, MISSING_APP_SOURCE, POLICY_PRECONDITION_MISSING. Each maps to a failed check in the schema's conditional allOf. The validator returns a list of these (not a single error) so a citizen developer sees all gaps at once, not one-at-a-time (C2 Clarity — the reader understands the full scope of fixes needed).

3.5 Gitea release-asset API endpoint (for scripts/attach_release_asset.py)

Confirmed from the existing scripts/ship_phase.sh (line 38) which already uses the Gitea releases API, and from the Gitea API swagger (https://gitea.com/api/swagger, fetched — the OpenAPI/Swagger JSON is published there; the endpoint is standard Gitea).

Release creation (existing pattern, ship_phase.sh line 38):

POST https://git.cloudinit.dev/api/v1/repos/continuous-intelligence/acdl/releases
Authorization: token <NOVA_GITEA_TOKEN>
Content-Type: application/json
Body: {"tag_name": "...", "name": "...", "body": "..."}
Response: {"id": <release_id>, ...}

Release asset attachment (the new endpoint, for attach_release_asset.py):

POST https://git.cloudinit.dev/api/v1/repos/continuous-intelligence/acdl/releases/{release_id}/assets
Authorization: token <NOVA_GITEA_TOKEN>
Content-Type: multipart/form-data
Form fields:
  name       = <filename, e.g. "nova-no-humans-platform.pptx">
  attachment = <the file, multipart>
Response: {"id": <asset_id>, "name": "...", "size": ..., "download_count": 0, ...}

The Gitea API endpoint is POST /api/v1/repos/{owner}/{repo}/releases/{id}/assets with a multipart form containing name (the display filename) and attachment (the file binary). The {id} is the numeric release ID returned by the release-creation call (the d.get('id') in ship_phase.sh line 40). attach_release_asset.py (REQ-228) takes a release tag (or ID) + a file path, resolves the tag → release ID (GET /api/v1/repos/.../releases/tags/{tag} if only the tag is known), then POSTs the multipart form. The token comes from .env.secrets (NOVA_GITEA_TOKEN, same as ship_phase.sh line 35).

Implementation note: urllib (used throughout contract_ingestor.py and ship_phase.sh) does not natively produce multipart form bodies — attach_release_asset.py must either (a) construct the multipart boundary + body manually (the standard urllib pattern), or (b) use requests if available. The repo's convention is stdlib-only (urllib, no requests dependency in the ingestor), so the script should construct the multipart body manually (C3 Simplicity — no new dependency for one script; C8 Economy — stdlib is sufficient). A ~30-line multipart_encode(fields, files) helper is the standard stdlib pattern.


4. Marp PPTX Theme Fidelity

4.1 The PPTX export path and inline-CSS survival

Source: the Marp CLI README (https://github.com/marp-team/marp-cli, fetched) + the existing docs/presentations/README.md (lines 93105)

  • the v1.9.2 theme commit ae0cb58 (verified via git show).

Confirmed export command (from docs/presentations/README.md line 9699):

CHROME_PATH=/root/.cache/ms-playwright/chromium-1217/chrome-linux64/chrome \
  npx --yes @marp-team/marp-cli@latest --allow-local-files \
  docs/presentations/nova-no-humans-platform-marp.md \
  -o <output-path>.pptx

How PPTX export works (from the Marp CLI README, --pptx section): The default (non-editable) PPTX "consists of pre-rendered background images." Marp renders each slide in a headless browser (Chrome/Chromium via the --browser-path / CHROME_PATH env), captures the rendered slide as a high-resolution image (default scale factor 2x — the README states: "By default, Marp CLI will use 2 as the default scale factor in PPTX"), and embeds those images as full-slide background pictures in the PPTX. Presenter notes are supported; the PPTX opens in PowerPoint, Keynote, Google Slides, LibreOffice Impress.

Inline style: CSS survival — CONFIRMED YES. Because the slides are rasterized in a headless browser, the browser's rendering engine applies the inline style: CSS block (H1/H2 #D6002A, title-slide bg #1B1B1B with 8px #D6002A accent, body text #1B1B1B, blockquote border #D6002A, table headers #F0F0F0, font 'Akkurat Pro' + fallbacks) exactly as it does for HTML export. The CSS is baked into the pixels of each slide image. The PPTX is a sequence of images, not editable PPTX shapes — so there is no "CSS stripping" step. The S&P Global Energy theme survives PPTX export in the standard (non-editable) path.

The current unified deck (docs/presentations/nova-no-humans-platform-marp.md) already has the style: | block in its frontmatter (verified: line 8 style: |, line 2 marp: true, line 3 theme: default). So the S&P theme is already inline; PPTX export will honor it.

Caveat — --pptx-editable (NOT used): The experimental --pptx-editable flag generates editable PPTX (texts/shapes, not images), and the README warns: "If the theme and inline styles are providing complex styles into the slide, --pptx-editable may throw an error or output the incomplete result." Nova does NOT use --pptx-editable (the S&P theme is complex inline CSS); the standard image-based PPTX is the path. REQ-228 specifies --pptx --allow-local-files, not --pptx-editable.

4.2 Fallback (NOT needed, documented for completeness)

If PPTX export ever strips inline CSS (it does NOT in the standard path, per §4.1), the fallback is a Marp custom theme CSS file referenced via --theme <path>:

CHROME_PATH=... npx @marp-team/marp-cli@latest --allow-local-files \
  --theme docs/presentations/assets/sp-theme.css \
  docs/presentations/nova-no-humans-platform-marp.md \
  -o output.pptx

Marp CLI supports custom theme CSS files via --theme <path> (the README's "Use custom theme" section: "A custom theme created by user also can use easily by passing the path of CSS file"). The CSS file would be docs/presentations/assets/sp-theme.css containing the same rules currently in the inline style: block, prefixed with the @theme meta comment (Marpit convention: /* @theme sp-energy */). The deck's frontmatter theme: directive would then be set to the custom theme name instead of default.

Recommendation: do NOT use the fallback. The inline style: block survives the standard PPTX path (rasterized images). The fallback adds a file to maintain in sync with the inline block (a DRY violation — two sources of truth for the S&P colors). REQ-214 restores the S&P theme in the unified deck's inline style: block (the v1.9.2 pattern); the PPTX export uses the same deck file. The S&P colors survive PPTX export via the inline style: block. No --theme flag, no separate CSS file needed. Confidence 0.90 (the only residual risk is a Marp CLI version regression that changes the rasterization path — mitigated by @marp-team/marp-cli@latest pinning in the render script and the PPTX slide-count/media verification step already in docs/presentations/README.md lines 332340).

4.3 ship_phase.sh release pattern + attach_release_asset.py extension

Confirmed from scripts/ship_phase.sh (read in full, 46 lines):

  • Line 38: POST https://git.cloudinit.dev/api/v1/repos/continuous-intelligence/acdl/releases with Authorization: token <NOVA_GITEA_TOKEN> (read from .env.secrets, line 35) + JSON body {"tag_name", "name", "body"} (line 37). The response's id is the release ID (line 40: d.get('id')).
  • The script creates the tag, pushes, creates the release, prints release_id: <id> tag: <tag>.

How attach_release_asset.py extends it (REQ-228): attach_release_asset.py is a separate script (not a modification to ship_phase.sh) that runs after the release exists. It takes a release tag (or ID) + a file path, then:

  1. Resolve tag → release ID (if only the tag is known): GET /api/v1/repos/continuous-intelligence/acdl/releases/tags/{tag} → the release object's id.
  2. Upload the asset: POST /api/v1/repos/continuous-intelligence/acdl/releases/{id}/assets with multipart form (name = filename, attachment = file binary)
    • Authorization: token <NOVA_GITEA_TOKEN> (same .env.secrets source).
  3. Print asset_id: <id> release: <tag> file: <name> for the ship log.

The render+attach flow (REQ-228, triggered by any docs/presentations/*-marp.md or docs/presentations/assets/ change, D-142):

render_deck.sh  →  HTML (committed) + PPTX (committed, D-141)
                ↓
attach_release_asset.py  →  PPTX uploaded to the phase's Gitea release

PPTX is a first-class artifact (PROJECT.md line 682683): committed to git (history) + attached to the release (download) — both always, not optional. This is the D-141 decision (no LFS — the binary is committed directly).


5. Assumptions logged (v1.18)

  • A1 (0.92): The Atelier v0.3.6 tag is the correct pin. It is the latest release (2026-08-05), the v0.4 milestone release, and the complete matrix state (19 domains, 190 P-rules). -11 commits to main since this release confirms main is a moving target — pinning is required for audit reproducibility (D-136). Risk: a v0.5 lands before P5 ships — mitigated by VERSION.md + update script (deliberate upgrade, not silent drift).
  • A2 (0.88): The MCP Python SDK v2 high-level server class is MCPServer (import from mcp.server import MCPServer), NOT FastMCP. The docs (landing page + Tools page) use MCPServer consistently; FastMCP was the v1 name. D-137 (MCP Python SDK v2) resolves to this import. Risk: the v1→v2 rename — if a future SDK patch restores a FastMCP alias, both imports would work, but the v2 canonical name is MCPServer.
  • A3 (0.85): mcp.run() starts the stdio transport by default (no explicit transport argument needed for the stdio path). The docs show uv run mcp dev server.py (Inspector) and the "no protocol handling" promise implies mcp.run() is the single entry point. The exact run() signature for stdio vs HTTP is not spelled out on the landing page (it's in the "Running your server" section, not fetched in full); the D-135 decision (stdio now, HTTP-ready on the same object) is consistent with a single run() entry point. P5 implementation should verify the exact run call from the "Running your server" docs page.
  • A4 (0.90): The standard (non-editable) PPTX export bakes inline style: CSS into the rasterized slide images. The Marp README states PPTX "consists of pre-rendered background images" — the browser rendering applies the CSS before rasterization. The S&P theme survives PPTX export. The --pptx-editable path (NOT used) is the only path that could strip CSS, and Nova does not use it.
  • A5 (0.88): The Gitea release-asset endpoint is POST /api/v1/repos/{owner}/{repo}/releases/{id}/assets with multipart name + attachment. This is the standard Gitea API (the swagger at gitea.com/api/swagger publishes the OpenAPI spec); the existing ship_phase.sh uses the sibling .../releases endpoint, confirming the API root + auth pattern. The {id} is the numeric release ID (resolvable from the tag via GET .../releases/tags/{tag}).
  • A6 (0.85): The submission-readiness schema uses JSON Schema draft 2020-12 conditional allOf / if-then for the per-env mandatory table (W3.E). This is the standard pattern for "if environment=qa then require validation.e2eSuite + validation.loadTest." The jsonschema library (already a dependency, used in contract_ingestor.py) supports draft 2020-12 conditionals. The validator (core/submission_readiness.py) may implement the per-env check in Python (clearer reason codes) rather than relying solely on schema conditionals — the schema is the shape, the validator is the gate with the citizen-developer-facing reason codes (REQ-218).
  • A7 (0.80): The --check-readiness CLI mode is added as a if __name__ == "__main__": block in contract_ingestor.py (which currently has none — it's Lambda-only). D-133 says "invoked as contract_ingestor.py --check-readiness" — this is a local pre-flight CLI, not a new Lambda action. The validator lives in core/submission_readiness.py (REQ-218); the ingestor dispatches to it. This keeps the Lambda path unchanged (the readiness gate is a pre-write step in _submit_contract only if desired; the CLI path is the citizen-developer pre-flight). Risk: the exact wiring (does the Lambda also gate on readiness, or only the CLI?) is a P3 implementation decision — REQ-218 says "On pass → proceeds to existing contract ingestion," implying the gate is in the submission path, but the CLI mode is the pre-flight surface.
  • A8 (0.90): The 9-skill list in REQ-221 is final (no adjustment). The research confirms the 9 Atelier domains map cleanly to the BA.A 5-skill catalog; the 4 "reference-only" domains (Performance, Documentation, Concurrency, AI/ML) are correctly NOT elevated to skills. Adding a 10th skill would break REQ-221's exact list and the BA.A mapping.
  • A9 (0.88): The mcp-engineer persona is NOT needed — it folds into backend-engineer. The MCP plugin-registry (D-140) is a Python backend pattern (decorators, type hints, stdio, urllib). The SDK v2 API surface is small and FastAPI/Pydantic-style (already in backend-engineer's range). D-143 logged in PERSONAS.md records this.

v1.22 Research — Nova Deck Layout Fix (2026-08-11)

Investigation into the systemic layout/formatting problems in the Nova presentation deck reported as "completely out of whack" after the v1.21 P5 re-render. This research IS the investigation — the findings below are the empirical root-cause analysis that drives the v1.22 requirements (REQ-254..262).

Background — why v1.22 exists

The v1.21 milestone shipped a refined deck (renamed to "Autonomous Cloud Delivery Platform", 4-beat arc, 18 main + 1 appendix slides). The P5 phase re-rendered the HTML + PPTX and added two new mermaid diagrams (platform-pipeline.png, telemetry-live-ops.png). After P5, the user reported that the layout is "completely out of whack" and that "they all have layout issues." This research identifies the root cause and the fix scope.

FINDING 1 — Theme CSS has ZERO section padding (CONFIDENCE: VERY HIGH)

docs/presentations/assets/nova-sp-theme.css line 1 is /* @theme nova-sp */ — a comment, not the @theme directive that Marp uses to register a theme name. The theme does not @import Marp's default theme. Marp's built-in default theme applies section { padding: 56px 64px; }. Because this custom theme neither imports the default nor sets its own padding, the rendered <section> has zero padding.

Verification: grep for padding:56px / padding:64px / padding:96px in the rendered HTML returns zero matches. The only section rules in the rendered HTML are:

  • section{width:1280px;height:720px;box-sizing:border-box;overflow:hidden;position:relative;...} (marpit base — no padding)
  • section{font-family:...;font-size:22px;color:var(--sp-black);background:var(--sp-white)} (theme — no padding)

Effect: Content is jammed against the slide edges (left/top/right/ bottom all 0px), header/footer chrome overlaps content, and there is no breathing room. This alone makes every slide look "out of whack."

FINDING 2 — overflow:hidden silently clips dense content (CONFIDENCE: VERY HIGH)

The marpit base rule sets overflow:hidden on section. The theme adds no overflow override, no scaling, no shrink-to-fit. Any slide whose content exceeds 720px is clipped with no visual indication. Combined with zero padding, content-dense slides (tables, image+bullets) lose their bottom rows / benefit paragraphs.

Per-slide overflow risk table (available content height ≈ 720px header(~35px) footer(~35px) padding(0px) = ~650px):

# Slide Est. height Fits? Issue
3 Objectives + Anti-Goals ~780px NO Densest text slide; nested list
5 RACI (8-row × 5-col) ~700px NO Cell text wraps to 2 lines
6 Pipeline (image + 4 bullets) ~750px NO Image + bullets overflow
8 Attestation (10-row × 4-col) ~780px NO Description column wraps
9 Telemetry (image + 4 bullets) ~720px NO Image + bullets overflow
12 Deferred (8-row table) ~720px NO Blocking-work column wraps
15 Quarter-by-Quarter (5-col) ~720px NO Wide table, long text
A1 Glossary (13-row × 3-col) ~700px NO On title-class (dark bg)
1,4,10,11,13,16,17,18 various ~620px TIGHT Cramped with 0 padding
2,7,14 various ~520px YES Manageable density

8 of 19 slides overflow; 8 more are cramped.

FINDING 3 — Image aspect-ratio catastrophe on slides 6 & 9 (CONFIDENCE: HIGH)

The two new P5 PNGs have extreme, opposite aspect ratios:

  • platform-pipeline.png = 1568×116 (aspect 13.52, ultra-wide/short). The deck uses ![w:1000]. At width=1000px, height = 1000/13.52 = 74px. The max-height:320px rule never engages. The image renders as a 1000×74 thin strip — text in nodes is nearly unreadable, and the 10-node LR flowchart is squashed.

  • telemetry-live-ops.png = 1024×1628 (aspect 0.63, tall). The deck uses ![w:900]. At width=900px the natural height would be 1428px — but max-height:320px clamps it, so the image actually renders at ~201×320. The w:900 directive is completely overridden by max-height:320px. The image is tiny and the explicit width is ignored. The .mmd uses flowchart TB (top-bottom) — exactly the failure mode the README (line 168) warns against.

FINDING 4 — Header+footer chrome on every slide (CONFIDENCE: HIGH)

The frontmatter sets both header: and footer: to the full 51-char deck title "Nova — The Autonomous Cloud Delivery Platform" on every slide (including the title slide, which has data-header/data-footer attributes present but _paginate: false only suppresses the page number, not the header/footer). The theme gives header a border-bottom and footer a border-top, each consuming ~30-40px of vertical chrome. With zero section padding, the header text sits at the very top edge and the footer at the very bottom edge, visually colliding with slide content. This reduces the effective content area from 720px to roughly 640-650px on every slide.

FINDING 5 — render_deck.sh produces unthemed output (CONFIDENCE: MEDIUM-HIGH)

scripts/render_deck.sh (line 44) runs marp-cli without --theme, relying on the frontmatter theme: nova-sp. But nova-sp is not a built-in Marp theme — it's a custom CSS file. Marp resolves theme: frontmatter against its built-in theme registry (default, gaia, uncover) and registered custom themes via --theme-set. Without --theme <file> or --theme-set, Marp cannot resolve nova-sp and falls back to the default theme (or errors). The committed HTML was rendered by render_slides.sh (which correctly passes --theme), so the committed artifact is fine — but render_deck.sh is a stale, dangerous script that would produce an unthemed/default-themed deck if anyone ran it. The README (line 201) documents render_slides.sh as canonical; render_deck.sh is not mentioned in the build-commands section.

FINDING 6 — render_slides.sh missing 2x scale + transparent bg (CONFIDENCE: HIGH)

render_slides.sh mermaid invocation (lines 51-55) does NOT pass -s 2 (2x scale) or -b transparent, despite the README (line 193) documenting both as required. This is why platform-pipeline.png is only 1568px wide (1x) instead of 3136px (2x) — the rendered PNGs are lower resolution than the README specifies, contributing to illegibility when scaled.

FINDING 7 — P5 marp-cli version bump (CONFIDENCE: LOW — not the cause)

The P3→P5 HTML diff is 831 changed lines, but the theme CSS portion is byte-identical (verified: font-size:22px, max-height:320px, marpit-root-font-size:22px, --sp-red:#D6002A all match; sp-red appears exactly once in both). The large diff is:

  • (a) marp-cli boilerplate (bespoke-marp presenter/overview/transition CSS) changed due to a marp-cli version bump (neither script pins a version — both use @latest), and
  • (b) content changes: title "No-Humans Infrastructure Platform" → "Autonomous Cloud Delivery Platform", footer "Act %{page}/5 — v1.20" → deck title, slide count 20 → 19.

The version bump did not alter the slide layout engine or the theme rules. This is not the regression source. The layout problems are inherent to the theme CSS (zero padding, no overflow handling, blunt image rule) which has been unchanged. P5 made the content denser (new diagrams with extreme aspect ratios, longer deck-title header/footer) which made the pre-existing theme flaws more visible.

FINDING 8 — Test coverage gaps (CONFIDENCE: VERY HIGH)

tests/test_slides_pipeline.py (268 lines) checks static file properties only:

  • Theme CSS file exists and contains #D6002A / #1B1B1B
  • Frontmatter references nova-sp, not default
  • render_slides.sh exists, is executable, invokes mermaid-cli + marp-cli
  • Every .mmd has a .png
  • No maturity badges, no version in footer, slide count = 18+1
  • No D-###/REQ-###/internal .py paths in slides

What is NOT tested (the gaps that let layout regressions through):

  1. NO rendered-dimension / overflow test — no test renders the HTML and checks that each slide's content height ≤ 720px.
  2. NO theme-CSS structural test — no test asserts section has padding, that overflow is handled, or that img rules don't conflict with w:/h: directives.
  3. NO image aspect-ratio / legibility test — no test checks that PNG dimensions are reasonable for a 16:9 slide.
  4. NO render-script theme-flag test — no test asserts render_deck.sh passes --theme (it doesn't), so the broken script passes CI.
  5. NO rendered-HTML structural assertion — no test parses the committed HTML to verify the theme is actually embedded.
  6. NO mermaid render-scale test — no test verifies PNGs are 2x scale.

Conclusion: A layout regression — including the current zero-padding, image-clamping, and table-overflow problems — would pass every existing test. This is why the user's "completely out of whack" report was not caught.

Theme CSS gaps (summary)

  1. NO padding on section (lines 21-26 set font/color/bg only).
  2. NO overflow handling on section.
  3. NO @import of a base theme (line 1 is a comment, not @theme).
  4. NO rule for the h2 + bold-lead-paragraph pattern (default <p> margins waste ~44px each).
  5. Table cell padding 6px 10px too generous for 10-13 row tables.
  6. img { max-height: 320px } is a blunt instrument that breaks w: directives on tall images and does nothing for ultra-wide images.
  7. Header/footer have no padding/margin — collide with content at 0 section padding.
  8. section.title does not suppress header/footer.
  9. NO rule for ol (only ul/li styled).
  10. NO @media print overrides for PPTX export fidelity.

Assumptions logged (v1.22)

  • A1 (0.95): The theme CSS is the primary root cause. Adding section { padding: 48px 56px 40px; } alone would fix the "jammed against edges" look on all 19 slides. Confidence grounded in the grep verification (zero padding matches in rendered HTML).
  • A2 (0.9): The P5 re-render is NOT a regression — the theme CSS is byte-identical P3→P5. P5's denser content (new diagrams, longer header/footer) made pre-existing flaws visible. Grounded in the byte-level diff comparison.
  • A3 (0.9): render_deck.sh should be deleted, not fixed. The README already documents render_slides.sh as canonical; keeping a second broken script is a footgun. Grounded in the README build- commands section (line 201) which does not mention render_deck.sh.
  • A4 (0.85): Splitting slides 3 and 8 (18 → 20 main) is preferable to trimming content, because the content is leadership-relevant and should not be lost. The slide-count test + README convention are updated to match. Grounded in the overflow estimates (slides 3 and 8 are the densest at ~780px).
  • A5 (0.8): Pinning marp-cli/mermaid-cli versions is necessary to prevent uncontrolled boilerplate-CSS drift. The exact pinned versions will be determined during P2 execution by testing which version produces stable output in this environment.

Decisions surfaced (research → bound in CLARIFY)

All 5 CLARIFY decisions are grounded in these findings:

  • Comprehensive scope (FINDINGS 1-8 span 4 layers)
  • Full pipeline (new milestone, complete audit trail)
  • Re-layout to LR (FINDING 3 — TB produced 0.63 aspect)
  • Delete render_deck.sh (FINDING 5 — stale, unthemed)
  • Split slides 3+8 (FINDING 2 — densest overflow)