Files
acdl/.ciagent/RESEARCH.md
T
Jon Chery 7585c828f0
acdl-ci / Lint (push) Successful in 7s
acdl-ci / Test (push) Successful in 23s
acdl-ci / Platform check-only (offline) (push) Successful in 8s
docs(P48): vision gaps + badge system + substrate→engine + CR format + agentic tags
9 requirements implemented across presentation decks and project docs:

1. DX closing slide: added 'Infrastructure as a utility, not a craft' bullet
   to convey the full vision (infrastructure consumed, not maintained;
   platform compounds value over time).
2. PW Problem slide: 'moving a merged change' → 'promoting a change'.
3. PW Problem slide: added 'Red tape' and 'Scalability without increasing
   headcount' bullets (4 frictions, not 2).
4. PW Roadmap slide: redesigned with side-by-side HTML table layout
   (Testing | Planned), 16px font, no overflow.
5. PW deck: added new slide 'What This Platform Is — and Isn't' after North
   Star (sovereign boundary, infrastructure as utility, 4 anti-goals).
   PW deck now 16 slides (was 15).
6. Maturity nomenclature: 'Available today'/'shipped' → 'Testing' across
   both decks + source markdown. New .testing badge (blue/teal #DBEAFE).
   Roadmap title: 'Testing vs. Planned'. The platform has 0 consumer
   adoption — 'shipped' was inaccurate.
7. Global: 'substrate' → 'engine' across entire project (88 matches, 30+
   files including .ciagent/, docs/, modules/, adapters/, schemas/, code).
8. Presentation files only: 'forge' → 'VCS' / 'version control system'
   (6 occurrences in 4 files). 'forge' retained in all technical docs and
   code as the industry-standard term.
9. New .agentic badge (purple/violet #EDE9FE) appended to agentic features
   in both decks: confidence signal, autonomous dev, pattern recognition,
   dynamic module creation, citizen developer surface, auto-promotion.

Also: Change Request ID format changed from 'CR-2026-001' to 'CHG0678912'
across presentation files, consumer guide, and test fixtures.

HTML re-rendered. PPTX rendered for release upload.

---ci---
phase: 48
milestone: v1.9
status: complete
requirements:
  covered: []
  partial: []
---/ci---
2026-07-23 14:58:29 +00:00

100 KiB
Raw Blame History

ACDL — v1.1 Research Findings

Phase: research (pre-Phase 06). Milestone: v1.1. Status: complete. Researcher: ci-researcher. Autonomy: full (CLARIFY auto-resolved parameters). Sources: web (Gitea docs, act_runner/gitea-runner repo, gitea/runner-images, Checkov docs, Checkov GitHub README, go-gitea issue tracker) + ACDL codebase. This file is a research artifact; no implementation code here, only sketches / snippets / schema shapes. Decisions surfaced are listed in §"Decisions surfaced" and are proposals for PROJECT.md until lead-developer adopts them.


TARGET 1 — Gitea / act_runner OIDC support (HIGHEST PRIORITY)

Findings

Verdict: Gitea Actions does NOT support emitting an OIDC id-token as of Gitea 1.27.x / gitea-runner (formerly act_runner) v2.1.0 (July 2026). This single finding gates Phase 08 feasibility.

Evidence (all verified 2026-07-21):

  1. Gitea docs — "Compared to GitHub Actions" (https://docs.gitea.com/usage/actions/comparison), section "Different behavior → Job token permissions (permissions)":

    "GitHub-only scopes such as statuses, checks, deployments, id-token, security-events, and pages are not supported, while Gitea-specific scopes such as code, releases, wiki, and projects are available."

    id-token is explicitly listed as an unsupported GitHub-only scope. The permissions: id-token: write block that the GitHub Actions AWS OIDC pattern relies on is therefore a no-op in Gitea Actions.

  2. Gitea docs — "Actions job token permissions (GITEA_TOKEN)" (https://docs.gitea.com/usage/actions/token-permissions), "Compatibility notes": identical text — id-token not supported. The supported scopes are contents, code, releases, issues, pull-requests, actions, wiki, projects, packages.

  3. Open issue: "Native OIDC Token for workload identity federation" (go-gitea/gitea#33681, opened 2025-02-21, status Open as of 2026-07-21, label type/proposal). The issue author explicitly asks for permissions.id-token: writestyle issuance so Gitea CI can federate to GCP/AWS without long-lived keys. No milestone assigned. Confirms the gap is still open at the proposal stage.

  4. Open issue: "Gitea as an OIDC IdP for Actions" (go-gitea/gitea#26383, opened 2023-08-07, status Open). Long-standing request; the original motivation for the OIDC work.

  5. Draft PR: "Add Actions OIDC provider with workflow permission gating" (go-gitea/gitea#36988, opened 2026-03-25 by @lunny, status Draft as of 2026-07-21, label lgtm/need 2 — needs two maintainer approvals, has unresolved review comments from 2026-05-27 re: case-insensitive bearer header, .well-known path location, and run_id/job_id query-param design). The PR implements the ACTIONS_ID_TOKEN_REQUEST_URL / ACTIONS_ID_TOKEN_REQUEST_TOKEN env-var contract, id-token permission parsing, and a JWT issuer — but it is not merged and not in any released Gitea version. No 1.27.x or 1.28-dev changelog mentions it.

  6. Open issue: "ci: consider replacing AWS access-key secrets with OIDC" (go-gitea/gitea#37980, opened 2026-06-03, status Open) — Gitea's own CI is still using long-lived AWS access keys because OIDC is not available. This is the strongest possible signal: the Gitea project itself has not been able to dogfood OIDC.

  7. Runner rename: act_runner was renamed to gitea-runner in gitea/runner#850 (2026-04-30). Latest runner release: v2.1.0 (2026-07-16). The v0.2.x line is the legacy act_runner naming. No gitea-runner release notes mention OIDC token issuance.

Conclusion: the GitHub Actions pattern (permissions: id-token: writeACTIONS_ID_TOKEN_REQUEST_URL + ACTIONS_ID_TOKEN_REQUEST_TOKENaws sts assume-role-with-web-identity via aws-actions/configure-aws-credentials) is not portable to Gitea Actions today. There is no env-var, no permission flag, and no documented mechanism. The draft PR #36988, if/when merged, would close the gap — but it cannot be a v1.1 dependency (draft, unmerged, no target milestone).

Confidence

0.95 — multiple primary-source confirmations (official docs + open issues + draft PR state). The 0.05 residual is for the possibility that the Gitea instance at https://git.cloudinit.dev runs a custom build with #36988 cherry-picked; this is unlikely (the instance is documented as standard Gitea per the v1.0 research) and should be verified in Phase 08 by checking https://git.cloudinit.dev/api/v1/version and the runner version label.

Assumptions logged

  • A-1.1 (0.90): the Gitea instance at git.cloudinit.dev runs upstream Gitea 1.27.x with no OIDC patches. Verifiable in Phase 08 via GET /api/v1/version and the runner admin page.
  • A-1.2 (0.85): PR #36988 will not merge and release in time for the v1.1 spike (Phases 0810). Even if it merged tomorrow, it would land in 1.28-dev at the earliest; the spike cannot block on it.

Fallbacks evaluated (for assuming an AWS IAM role from a Gitea Actions

step without a long-lived key)

Option Mechanism Viability for the spike
(a) GitHub-style OIDC (id-token: write) ACTIONS_ID_TOKEN_REQUEST_URL → JWT → sts assume-role-with-web-identity Not available (see above).
(b) Self-hosted OIDC broker Stand up a tiny OIDC IdP (e.g. dex, oauth2-proxy, or a custom JWKS endpoint) that the Gitea job authenticates to with its GITEA_TOKEN and that issues a JWT minted with a platform signing key; AWS IAM trusts the broker's JWKS. Workable but heavy for a spike — requires a second always-on service, a signing-key rotation story, and IAM trust plumbing. Better suited to v1.2.
(c) aws sts assume-role-with-web-identity with a token from Gitea's own API Use the job's GITEA_TOKEN (a PAT-equivalent, short-lived for the job) as the WebIdentityToken to STS. Rejected: STS rejects non-OIDC tokens; GITEA_TOKEN is not a JWT, has no iss/sub/aud claims, and AWS IAM has no Gitea OIDC provider to trust. (This is exactly the gap #33681 describes for GCP.)
(d) Short-lived AWS creds via a scheduled credential mint A platform job (cron) mints aws sts get-session-token (or a role-session) and writes the temp creds as a Gitea Actions secret with a TTL ≤ 1h. The spike workflow reads the secret. Workable, but reintroduces a long-lived key upstream (the mint job needs one) and a secret in Gitea — a narrower version of the very thing §12.5 forbids. Acceptable as a documented spike-only waiver if (a) and (b) are both rejected for the spike scope.
(e) LocalStack as an AWS stand-in Replace real AWS with LocalStack for the spike; no IAM trust needed at all (LocalStack mocks STS). Workable for the mechanics of terraform plan but invalidates REQ-23 ("real AWS via OIDC") and the spike's whole purpose of proving real-AWS feasibility. Reject for the spike; keep as a unit-test engine only.
(f) Documented spike-only waiver: rotate a long-lived key per-run One IAM access key, stored as a Gitea Actions secret, used by the workflow, rotated (deactivated + new key) after each spike run by the same workflow. The cleanest available option that still touches real AWS. Still violates the letter of §12.5 ("long-lived credentials are forbidden") but satisfies the intent for a time-boxed spike: the key's useful lifetime equals one workflow run (minutes), not "long-lived." Requires an explicit, logged waiver.
(g) GitHub-hosted mirror pipeline Run the OIDC-requiring step on GitHub Actions (which supports id-token: write) against the same repo mirrored from Gitea. Rejected: introduces a second forge, violates the "Forge: Gitea" constraint, and defeats the spike's purpose of proving the platform works on Gitea.

Recommendation (concrete)

Adopt option (f) as a documented spike-only waiver (propose as D-039), and commit to option (b) for v1.2. The spike achieves real terraform plan against real AWS without a persistently long-lived key: the key is minted by Phase 08, used by the Phase 09/10 workflow, and rotated immediately after each run. D-034 (the one-shot bootstrap waiver) already permits a single bootstrapping aws iam call; D-039 extends that with a per-run rotation discipline so the spike never leaves a usable key behind.

Exact Gitea Actions workflow snippet for the spike (Phase 09/10), assuming option (f):

# .gitea/workflows/spike-plan.yml (research sketch — not implementation)
name: acdl-spike-plan
on:
  workflow_dispatch:
    inputs:
      contract-ref:
        description: "Ref carrying the contract"
        required: false
        default: main
        type: string

jobs:
  plan:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: "Install terraform + checkov"
        run: |
          # HashiCorp apt repo (see TARGET 2)
          wget -O- https://apt.releases.hashicorp.com/gpg | sudo gpg --dearmor -o /usr/share/keyrings/hashicorp-archive-keyring.gpg
          echo "deb [signed-by=/usr/share/keyrings/hashicorp-archive-keyring.gpg] https://apt.releases.hashicorp.com $(lsb_release -cs) main" | sudo tee /etc/apt/sources.list.d/hashicorp.list
          sudo apt-get update && sudo apt-get install -y terraform jq python3-pip
          pip3 install --break-system-packages checkov

      - name: "Assume role via short-lived key (spike waiver D-039)"
        env:
          AWS_ACCESS_KEY_ID:     ${{ secrets.SPIKE_AWS_ACCESS_KEY_ID }}
          AWS_SECRET_ACCESS_KEY: ${{ secrets.SPIKE_AWS_SECRET_ACCESS_KEY }}
          AWS_SESSION_TOKEN:     ${{ secrets.SPIKE_AWS_SESSION_TOKEN }}   # set if the secret is a session token
          AWS_REGION:            us-east-1
        run: |
          # Mint a *fresh* short-lived session for this run (≤ 1h TTL).
          # The SPIKE_AWS_* secret is the one-key bootstrap (D-034); the
          # session creds below are what terraform sees.
          STS_JSON=$(aws sts get-session-token --duration-seconds 3600)
          aws sts assume-role --role-arn arn:aws:iam::${{ secrets.SPIKE_AWS_ACCOUNT_ID }}:role/acdl-act-runner-role \
            --role-session-name acdl-spike-${{ gitea.run_id }} \
            --duration-seconds 3600 > /tmp/role.json
          # Export the assumed-role creds into the subsequent step env.
          {
            echo "AWS_ACCESS_KEY_ID=$(jq -r .Credentials.AccessKeyId /tmp/role.json)"
            echo "AWS_SECRET_ACCESS_KEY=$(jq -r .Credentials.SecretAccessKey /tmp/role.json)"
            echo "AWS_SESSION_TOKEN=$(jq -r .Credentials.SessionToken /tmp/role.json)"
          } >> "$GITHUB_ENV"
          aws sts get-caller-identity

      - name: "terraform init + plan"
        run: |
          cd terraform
          terraform init -input=false
          terraform plan -input=false -out=tfplan.binary
          terraform show -json tfplan.binary > tfplan.json

      - name: "Checkov → PolicyCheckResult"
        run: |
          checkov -f terraform/tfplan.json --framework terraform_plan \
            --check CKV_AWS_18,CKV_AWS_19,CKV_AWS_20,CKV_AWS_21,CKV_AWS_41,CKV_AWS_45,CKV_AWS_46,CKV_AWS_1,CKV_AWS_24,CKV_AWS_25,CKV_AWS_33,CKV_AWS_7 \
            --output json --output-file-path checkov.json
          python3 adapters/terraform/policy/checkov_adapter.py checkov.json > policy_results.json
          # ... contract→IR → plan → confidence → outbox (Phase 10)

      - name: "Rotate the bootstrap key (D-039 discipline)"
        if: always()
        env:
          AWS_ACCESS_KEY_ID:     ${{ secrets.SPIKE_AWS_ACCESS_KEY_ID }}
          AWS_SECRET_ACCESS_KEY: ${{ secrets.SPIKE_AWS_SECRET_ACCESS_KEY }}
          AWS_REGION:            us-east-1
        run: |
          # Deactivate + delete the key this run just used, then mint a new
          # one and write it back to the Gitea secret via the API.
          # (Sketch — Phase 08 will implement scripts/rotate_spike_key.sh.)
          bash scripts/rotate_spike_key.sh

Note on the env-var name: Gitea Actions exposes the runner context as gitea.* (e.g. gitea.run_id) and the GitHub-compatible github.* aliases work too (confirmed in the Gitea FAQ). $GITHUB_ENV is the canonical step- env file in both. The snippet uses both forms deliberately to stay GitHub-compatible.

What this means for §12.5 and REQ-23

  • §12.5 ("long-lived credentials are forbidden") is a target architecture commitment, not a spike constraint. The architecture's own Gitea API surface table already hedged: id-token: write / OIDC was "To be confirmed in RESEARCH." This research confirms it is not available. The locked target stands for v1.2+; the spike uses the documented waiver.
  • REQ-23 ("AWS OIDC trust is configured … the temporary long-lived key … is rotated immediately after (waiver D-034)") is re-interpreted by this research: the "OIDC trust" cannot be configured in this environment yet. REQ-23's intent (real AWS, no persistent long-lived key) is met by D-039's per-run rotation. REQ-23 should be amended in Phase 07 to read "AWS trust is configured for the spike via a per-run-rotated key (waiver D-039); OIDC federation is the v1.2 target, blocked on go-gitea/gitea#36988."

TARGET 2 — Terraform + Checkov availability on the act_runner image

Findings

  • The default Gitea runner image is gitea/runner-images:ubuntu-latest, which is built from catthehocker/ubuntu:act-24.04 (https://gitea.com/gitea/runner-images). This mirrors the GitHub Actions ubuntu-latest image (catthehocker/ubuntu).
  • catthehocker/ubuntu:act-24.04 is a community-maintained clone of the GitHub-hosted Ubuntu runner. Neither terraform nor checkov is pre-installed on the GitHub-hosted runners, and therefore neither is on the Gitea runner image. Confirmed by the runner-images README ("Images are built from catthehocker/ubuntu:* or node:*. Additional packages will be installed if they are needed by runner") — i.e. only what gitea-runner itself needs is added.
  • terraform install: use the official HashiCorp apt repository (the recommended path; not tfenvtfenv adds a shell-init step that is unnecessary for a single-version spike).
  • checkov install: pip3 install checkov. Checkov requires Python ≥ 3.9 and ≤ 3.12 (per the GitHub README "Requirements"; Python 3.13 is not in the tested matrix — the README says "3.9 - 3.13" but the badge matrix historically stops at 3.12; the act-24.04 image ships Python 3.12, so this is fine). On Debian 12+ / Ubuntu 24.04, pip refuses to install into the system environment — use --break-system-packages (acceptable inside an ephemeral job container) or a venv. The spike uses --break-system-packages for simplicity (the container is thrown away after the job).

Confidence

0.85 — the runner-image composition is documented; the exact packages on catthehocker/ubuntu:act-24.04 are not enumerated in the Gitea docs (the image is third-party), but the GitHub-hosted-runner analogue is well known to exclude terraform. The 0.15 residual is for the possibility that a future gitea/runner-images:ubuntu-latest-full adds terraform.

Assumptions logged

  • A-2.1 (0.90): the act_runner at git.cloudinit.dev uses runs-on: ubuntu-latestdocker://docker.gitea.com/runner-images:ubuntu-latest (the documented default). Confirmable from the runner admin page in Phase 08.
  • A-2.2 (0.90): no runs-on: label on the instance maps to an image that pre-installs terraform/checkov.

Recommendation

  • runs-on: ubuntu-latest (the default).
  • Install step as sketched in the TARGET 1 workflow snippet. Pin terraform to a known version (e.g. terraform=1.9.* — the spike does not need the latest; pinning avoids a surprise major bump mid-spike).
  • Pin checkov to a known version (e.g. checkov>=3.2,<4) to keep the rule-ID set stable across spike runs.
  • Checkov rule IDs for the L2 checks (see TARGET 4 for the full mapping): the 4 L2 checks + tag/naming map to existing Checkov rules (CKV_AWS_18, CKV_AWS_19, CKV_AWS_20, CKV_AWS_21, CKV_AWS_41, CKV_AWS_45, CKV_AWS_46, CKV_AWS_1, CKV_AWS_24, CKV_AWS_25, CKV_AWS_33, CKV_AWS_7). No custom Checkov checks are needed for the spike. Tag/naming convention is not a built-in Checkov check; the spike implements it as a tiny custom YAML policy (see TARGET 4) or defers tag/naming to the confidence signal's NFR input for the spike and adds the custom Checkov rule in v1.2.

TARGET 3 — Target Stack IR prior art + v1 shape

Findings (prior art survey)

  • Pulumi resource model (https://www.pulumi.com/docs/concepts/): a program declares Resource objects with inputs (typed props), outputs (resolved after create), and explicit dependsOn / parent-child links. Pulumi's resource is not a tree — it's a DAG — and the parent relationship is for composition (e.g. a ComponentResource wrapping child resources), not for hard parent-single-child. This is more general than the ACDL IR needs in v1 (single-parent, max-depth-5 tree).

  • Terraform CDK (cdktf): compiles TypeScript/Python/etc. to Terraform HCL. The intermediate is a TerraformAsset/TerraformElement graph that is nearly 1:1 with HCL. Confirms the "nearly isomorphic to Terraform in v1" claim in architecture.md §12.1 — CDK's IR is HCL-shaped.

  • Crossplane Compositions: a CompositeResourceDefinition (XRD) declares a schema; a Composition templates patches from the composite to managed resources. Patching is by path, not by typed contract. Crossplane's model is runtime (the controller reconciles), whereas ACDL's IR is build-time (the adapter compiles to a plan). Crossplane validates that this separation (IR ≠ runtime) is a workable design.

  • ACDL v1.0 demo (modules/l1/l1-s3/manifest.yaml, modules/l2/l2-invoice-service/manifest.yaml): the demo's shape is:

    # L1 manifest (modules/l1/l1-s3/manifest.yaml)
    name: l1-s3
    kind: l1
    description: Object store primitive
    inputs:
      bucket_name: { description: ..., type: string }
      region:      { description: ..., type: string }
      retention_days: { description: ..., type: string }
    
    # L2 manifest (modules/l2/l2-invoice-service/manifest.yaml)
    name: l2-invoice-service
    kind: l2
    description: ...
    l1s:
      - name: l1-eks-fargate
        inputs: { cluster_name: invoice-cluster, region: us-east-1, cpu_arch: arm64 }
      - name: l1-s3
        inputs: { bucket_name: acdl-invoice-archive, region: us-east-1, retention_days: "365" }
    

    The demo's L2 l1s: list is a flat composition (no nesting, no relationships). The v1 IR upgrades this to a tree with explicit relationships (single parent per child, shared keyword for multi-rel) and typed outputs.

Confidence

0.80 — the prior-art survey is grounded; the v1 IR shape is a recommendation (not externally validated), confidence will rise once Phase 09 implements the adapter and the round-trip to Terraform is verified.

Assumptions logged

  • A-3.1 (0.85): the IR's "nearly isomorphic to Terraform in v1" claim (architecture.md §12.1) is the right v1 boundary — build a thin IR, defer engine-specific expressiveness to v2.
  • A-3.2 (0.80): single-parent-per-child is sufficient for v1 (no L1 needs two parents in the spike). The "shared keyword for multi-relationship" (architecture.md §12.1) is a v2 concern; the v1 schema reserves the field but the spike does not exercise it.

Recommendation — v1 IR JSON shape (sketch for schemas/ir.schema.json)

The IR is a tree of resources with typed inputs, outputs, nfrs, and a single parent reference. Composition metadata (max depth, registry version) lives at the root. Policy hooks are not in the IR — they attach at the L2-composition / pipeline stage, not in the resource definition (this keeps the IR purely descriptive, matching the architecture's "policy hooks are the points in the composition where policy checks attach" — the hooks are on the composition, not the resource).

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://acdl.cloudinit.dev/schemas/ir.schema.json",
  "title": "ACDL Target Stack IR",
  "type": "object",
  "required": ["version", "stack", "resources"],
  "properties": {
    "version": { "type": "string", "pattern": "^\\d+\\.\\d+\\.\\d+$" },
    "stack": {
      "type": "object",
      "required": ["name", "kind", "depth"],
      "properties": {
        "name": { "type": "string", "pattern": "^l[12]-[a-z][a-z0-9-]*$" },
        "kind": { "enum": ["l1", "l2"] },
        "depth": { "type": "integer", "minimum": 1, "maximum": 5 }
      }
    },
    "resources": {
      "type": "array",
      "items": { "$ref": "#/$defs/resource" },
      "minItems": 1
    },
    "relationships": {
      "type": "array",
      "items": { "$ref": "#/$defs/relationship" }
    }
  },
  "$defs": {
    "resource": {
      "type": "object",
      "required": ["id", "type", "module", "inputs"],
      "properties": {
        "id": { "type": "string", "pattern": "^[a-z][a-z0-9-]*$" },
        "type": { "type": "string", "description": "IR resource type, e.g. aws:s3:bucket" },
        "module": { "type": "string", "pattern": "^l1-[a-z][a-z0-9-]*@\\d+\\.\\d+\\.\\d+$",
          "description": "L1 registry reference (name@semver, per W3.D)" },
        "parent": { "type": "string", "description": "id of the single parent resource; absent for the root" },
        "inputs": {
          "type": "object",
          "additionalProperties": { "type": ["string", "number", "boolean", "object", "array", "null"] },
          "description": "Typed input contract; validated against the L1 module's declared inputs"
        },
        "outputs": {
          "type": "object",
          "additionalProperties": { "$ref": "#/$defs/outputSpec" },
          "description": "Typed output contract; the adapter translates these to Terraform outputs"
        },
        "nfrs": {
          "type": "object",
          "description": "Non-functional requirements (latency, throughput, rpo, rto, etc.); opaque to the adapter, consumed by the confidence signal's NFR input",
          "additionalProperties": true
        }
      }
    },
    "outputSpec": {
      "type": "object",
      "required": ["type"],
      "properties": {
        "type": { "type": "string", "description": "IR type, e.g. string, arn, ref:<resourceId>.<outputName>" },
        "description": { "type": "string" }
      }
    },
    "relationship": {
      "type": "object",
      "required": ["from", "to", "kind"],
      "properties": {
        "from": { "type": "string", "description": "resource id" },
        "to":   { "type": "string", "description": "resource id" },
        "kind": { "enum": ["parent", "depends_on", "uses_output"],
          "description": "parent = single-parent composition; depends_on = ordering only; uses_output = output reference. v1 uses parent + uses_output only." },
        "shared_keyword": { "type": "string",
          "description": "For multi-relationship (v2); reserved, unused in v1." }
      }
    }
  }
}

Why this shape round-trips to Terraform cleanly (v1):

  • resource.modulemodule "<id>" { source = "..."; version = "..."; <inputs> }.
  • resource.inputs → Terraform variable {} block in the L1 module + argument values in the L2 root module's module block.
  • resource.outputs → Terraform output {} block in the L1 module + output {} (passthrough) in the L2 root module.
  • relationship.kind = uses_output with to = "X.out" → Terraform interpolation module.X.<output>.
  • relationship.kind = parent → the child resource is inside the parent L1's module block (no Terraform construct; it's a composition hint the adapter uses to order module blocks). For the spike (l2-static-assetsl1-s3 only, depth 1) there is exactly one resource and zero relationships — the IR still validates, and the adapter produces a single module "s3" { ... } block.

TARGET 4 — PolicyCheckResult + Checkov adapter

Findings (Checkov JSON output shape)

Checkov's JSON output (per the GitHub README and CLI reference) is a JSON object keyed by framework, each containing results with passed_checks, failed_checks, and skipped_checks arrays. Each check record has at minimum: check_id, check_name, check_result ({ "result": "PASSED" | "FAILED" }, plus evaluations), file_path, file_abs_path, repo_file_path, resource, resource_address, code_block, severity (when available; requires Prisma Cloud API for full severity metadata, but Checkov emits severity for many built-in checks), guideline, bc_category_id. For terraform_plan framework, the file_path is the plan JSON and resource is the Terraform address (e.g. aws_s3_bucket.customer).

The adapter must:

  1. Run Checkov with --framework terraform_plan --output json --output-file-path checkov.json --soft-fail (so Checkov never exits non-zero; the confidence signal decides the gate, not Checkov's exit code — matching the v1.0 demo's confidence_signal.py discipline of always-exit-0).
  2. Read the JSON, iterate the failed_checks + passed_checks + (if present) skipped_checks, and emit one PolicyCheckResult per check.

Checkov rule → ACDL L2-check + severity mapping

ACDL L2 check Checkov rule ID(s) Checkov default severity PolicyCheckResult severity
secrets-in-plaintext CKV_AWS_41 (provider creds), CKV_AWS_45 (lambda env), CKV_AWS_46 (EC2 userdata) HIGH high
public ingress CKV_AWS_20 (S3 public read ACL), CKV_AWS_57 (S3 public write ACL), CKV_AWS_24 (SG 0.0.0.0/0 → 22), CKV_AWS_25 (SG 0.0.0.0/0 → 3389) HIGH (S3 ACL), MEDIUM (SG) high (S3 ACL), medium (SG) — the S3-public-ACL is the spike's exercised case
IAM wildcard CKV_AWS_1 (admin *:* policy document), CKV_AWS_40 (policy attached to user) HIGH (CKV_AWS_1), MEDIUM (CKV_AWS_40) high (CKV_AWS_1), medium (CKV_AWS_40)
KMS key reference CKV_AWS_7 (KMS rotation), CKV_AWS_33 (KMS wildcard principal) MEDIUM medium
tag compliance (no built-in Checkov rule for tag presence; CKV_AWS_51 is ECR immutable tags, not general tag compliance) n/a low (spike-only: emit a SKIPPED PolicyCheckResult with ruleId ACDL_TAG_NAMING and a "deferred to v1.2" message)
naming convention (no built-in) n/a low (same — ACDL_TAG_NAMING skipped in spike)

Severity mapping (Checkov → PolicyCheckResult): Checkov severities are CRITICAL, HIGH, MEDIUM, LOW, INFO (some require the Prisma Cloud API). Map 1:1 to the PolicyCheckResult severity enum (lowercase). Where Checkov does not emit a severity (no API key in the spike — we run without --bc-api-key), fall back to a default severity table baked into the adapter (above) — this is the spike's "no Prisma Cloud" path and matches the v1.0 demo's policy_checker.py deterministic discipline.

Critical-override semantics (architecture.md §8): if any PolicyCheckResult has severity: critical AND result: fail, the confidence signal hard-overrides the score to a mandatory block regardless of other inputs. None of the L2-check rules above are critical by default; the spike does not exercise the critical path (the v1.0 demo's Act 4 used a high-equivalent public-ingress). The architecture reserves critical for future rules (e.g. a custom Checkov rule that flags a resource that would expose customer data).

Confidence

0.85 — the Checkov JSON shape is documented and stable; the rule IDs are verified against the Checkov policy index. The 0.15 residual is for Checkov minor-version field-name drift (e.g. resource_address vs resource); the adapter must defensively read both.

Assumptions logged

  • A-4.1 (0.85): the spike runs Checkov without a Prisma Cloud API key, so severities come from the adapter's baked-in default table, not Checkov's severity field. If Checkov does emit a severity, the adapter prefers it.
  • A-4.2 (0.80): tag/naming is deferred for the spike (a SKIPPED PolicyCheckResult is sufficient to satisfy the "all six inputs present" gate per §8 — the policy input is present, it just says "skipped"). A custom Checkov YAML rule for tag presence lands in v1.2.

Recommendation — Checkov adapter sketch (Python)

# adapters/terraform/policy/checkov_adapter.py  (RESEARCH SKETCH — not implementation)
"""Translate Checkov JSON output to ACDL PolicyCheckResult records.

Reads Checkov's JSON output (one framework key, e.g. "terraform_plan"),
emits a list of PolicyCheckResult dicts conforming to
schemas/policy_check_result.schema.json.
"""
import json
import sys
import datetime

# Checkov rule ID → (ACDL L2 check name, default severity when Checkov omits one)
RULE_MAP = {
    "CKV_AWS_41": ("secrets-in-plaintext", "high"),
    "CKV_AWS_45": ("secrets-in-plaintext", "high"),
    "CKV_AWS_46": ("secrets-in-plaintext", "high"),
    "CKV_AWS_20": ("public-ingress", "high"),
    "CKV_AWS_57": ("public-ingress", "high"),
    "CKV_AWS_24": ("public-ingress", "medium"),
    "CKV_AWS_25": ("public-ingress", "medium"),
    "CKV_AWS_1":  ("iam-wildcard", "high"),
    "CKV_AWS_40": ("iam-wildcard", "medium"),
    "CKV_AWS_7":  ("kms-key-reference", "medium"),
    "CKV_AWS_33": ("kms-key-reference", "medium"),
}

def _iso8601_now():
    return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")

def _to_pcr(checkov_record, contract_id, result_str):
    rule_id = checkov_record["check_id"]
    acdl_check, default_sev = RULE_MAP.get(rule_id, (rule_id, "info"))
    severity = checkov_record.get("severity", default_sev).lower()
    return {
        "contractId": contract_id,
        "evaluatedAt": _iso8601_now(),
        "engine": "checkov",
        "ruleId": rule_id,
        "severity": severity,
        "result": {"PASSED": "pass", "FAILED": "fail",
                   "SKIPPED": "skipped"}.get(result_str, "error"),
        "message": checkov_record.get("check_name", ""),
        "evidence": {
            "file_path": checkov_record.get("file_path"),
            "resource": checkov_record.get("resource"),
            "resource_address": checkov_record.get("resource_address"),
            "code_block": checkov_record.get("code_block"),
        },
        "resourceRef": checkov_record.get("resource_address") or checkov_record.get("resource", ""),
    }

def adapt(checkov_json_path, contract_id):
    with open(checkov_json_path) as fh:
        data = json.load(fh)
    # Checkov JSON: { "<framework>": { "results": { "passed_checks":[], "failed_checks":[], "skipped_checks":[] } } }
    out = []
    for _framework, body in data.items():
        results = body.get("results", body)  # tolerate both shapes
        for rec in results.get("passed_checks", []):
            out.append(_to_pcr(rec, contract_id, "PASSED"))
        for rec in results.get("failed_checks", []):
            out.append(_to_pcr(rec, contract_id, "FAILED"))
        for rec in results.get("skipped_checks", []):
            out.append(_to_pcr(rec, contract_id, "SKIPPED"))
    return out

if __name__ == "__main__":
    checkov_path, contract_id = sys.argv[1], sys.argv[2]
    print(json.dumps(adapt(checkov_path, contract_id), indent=2))

Note on the spike's tag/naming handling: the spike emits a single PolicyCheckResult with ruleId: "ACDL_TAG_NAMING", result: "skipped", severity: "info", message: "tag/naming check deferred to v1.2" from the adapter after the Checkov pass — so the confidence signal sees the "all six inputs present" condition (the policy input is a non-empty list) without the spike needing a custom Checkov rule.


TARGET 5 — DynamoDB outbox pattern (RPO=0)

Findings

  • RPO=0 mechanics: the contract-submission handler writes to the DynamoDB outbox synchronously in the same transaction (or same step) that acks the submission. The ack does not return until the outbox write is durable (DynamoDB strong-consistent write). This is the standard "transactional outbox" pattern (https://microservices.io/patterns/data/transactional-outbox.html).
  • DynamoDB mode: PAY_PER_REQUEST (on-demand) for the spike — no capacity planning, no minimum cost, scales to zero. Provisioned is for steady-state high-throughput (v1.2+ when the pipeline has real load).
  • Worker design for the spike: there is no separate async worker in the spike. The pipeline step that computes the confidence signal writes the evidence event to the outbox synchronously in the same step. The "async worker + DLQ" is a v1.2 concern; the spike's RTO is "the step re-runs" (Gitea Actions re-runs the workflow on failure). This is acceptable because the spike is a single contract, single run, dev-only.
  • Outbox table schema (spike minimum):
    • PK: contractId (UUID)
    • SK: eventType#eventTs (e.g. POLICY_CHECKED#2026-07-21T12:00:00Z) — sorts events per contract chronologically.
    • Attributes: payload (the JWS-signed event body, see TARGET 7), prev_event_hash (the chain link), approver_qa (GitHub username of the QA approver; empty in the dev-only spike), approver_prod (SRE; empty in spike), environment, stack, score, band.
    • TTL: expire_at (1 year per §8 "1-year storage"; set to now + 365d).
  • Approver identities: stored on the first event for a contract (the submission event) as approver_qa / approver_prod, updated on the promotion events. The separation-of-duties check (§10.3) reads approver_qa at the prod-promotion step and compares to the new approver_prod. In the dev-only spike, both are empty — the check is a no-op stub that returns "distinct" (the spike does not exercise identity distinctness; it only proves the storage path works).

Confidence

0.85 — the transactional-outbox pattern is well-established; the spike minimum (synchronous write, no async worker) is a deliberate scope cut, not a design risk.

Assumptions logged

  • A-5.1 (0.90): DynamoDB on-demand is the right spike mode (zero cost when idle, no capacity planning).
  • A-5.2 (0.85): the spike's "RTO = re-run the workflow" is acceptable because the spike is a single dev-only submission; the v1.2 outbox worker + DLQ is the production RTO design.

Recommendation

  • DynamoDB table acdl-outbox, on-demand, single-region (us-east-1).
  • PK contractId, SK eventType#eventTs.
  • The pipeline step writes the evidence event synchronously via boto3 put_item (strong-consistent by default for DynamoDB).
  • The v1.0 demo's evidence_writer.py hash chain (see TARGET 7) is lifted for the spike's prev_event_hash field; the JWS signature is deferred to v1.2 (TARGET 7).
  • No separate worker / DLQ / EventBridge / Lambda for the spike. Phase 07's platform/audit_ledger_design.md documents the v1.2 async-worker
    • DLQ design; the spike implements only the synchronous write path.

TARGET 6 — Six-input confidence signal

Findings

The architecture (§8) locks "six canonical inputs" but does not enumerate them. Cross-referencing the vision's "Safety is Computed" tenet ("aggregating policy conformance, validation evidence, and historical behavior") and the HITL matrix in §10.4 (which enumerates: functional correctness, performance baseline, security posture, contract NFRs, operational readiness, incident response, capacity/cost, resilience), the 6 inputs must cover what the platform can compute autonomously in dev (dev has no HITL matrix — §5's "dev = Full autonomy, all six inputs present"). The HITL matrix's 8 concerns are qa/prod/dr concerns; the confidence signal's 6 inputs are the platform-computable subset that exists in every environment (including dev).

# Input Weight What it is Dev source
1 policy check results 0.30 List of PolicyCheckResult records (§12.6); severity-weighted penalty Checkov adapter (TARGET 4)
2 validation evidence 0.25 Schema-validity + IR-resolution-success + terraform-plan-success (the pipeline's own build/test gates) pipeline steps (schema validate → IR resolve → terraform validateterraform plan)
3 freshness 0.10 Age of the contract's declared validation evidence (e2eSuite, loadTest) relative to submission; in dev, this is the age of the L1/L2 module versions vs. the registry L1 registry publication timestamps
4 source / attestation 0.15 Identity of the submitter + the contract's source provenance (git ref, commit SHA, signed-by). In dev (autonomous), this is "any valid submitter" — the gate is presence, not identity. Gitea gitea.actor + commit SHA
5 historical behavior 0.10 Platform's observed history for this contract / stack / submitter: prior rollback count, prior policy-fail count. In the spike (first submission), this is a neutral 0.5 (no history). DynamoDB outbox (prior events for this contractId / stack)
6 NFR conformance 0.10 The contract's declared NFRs (latency, throughput, error rate) vs. the platform's measured baseline for this stack. In the spike, l2-static-assets declares no NFRs, so this input is "present + neutral 0.5" (the gate is presence, not conformance). contract nfrs block (optional) + platform baseline (none in spike)

Weights sum to 1.0. The base score (before severity penalties) is the weighted sum of each input's per-input score (each in [0,1]). The severity→penalty mapping (locked, §8) is then applied as a deduction from the weighted sum: any critical finding hard-overrides to 0 (block); each high finding deducts 0.2, medium 0.05, low 0.01, info 0.0. The final score is clamped to [0,1].

Per-env thresholds (locked, §8): dev ≥ 0.50, qa ≥ 0.75, prod ≥ 0.90, dr ≥ 0.95. band ∈ {pass, warn, block} where warn is the band between "pass" and "block" (e.g. for dev: ≥ 0.50 = pass, 0.400.50 = warn, < 0.40 = block — the warn band is a v1.2 signal for the HITL reviewer; in dev (autonomous) warn is treated as block since there is no reviewer).

Confidence

0.75 — the 6 inputs are a recommendation; the architecture does not enumerate them, so this research is choosing. The weights are a starting point (BA.B "thresholds frozen for v1, tuning begins v1.2" applies to the thresholds; the weights should be frozen alongside for v1 and tuned together in v1.2).

Assumptions logged

  • A-6.1 (0.80): "all six inputs present" (§5, dev gate) means each input produces a non-null per-input score; a missing input halts with an explicit reason (§8 "halt with explicit reason on missing input").
  • A-6.2 (0.75): the spike's inputs 3/5/6 are "present + neutral 0.5" because the spike is the first submission with no history and no declared NFRs. This is documented as the "cold start" baseline.
  • A-6.3 (0.85): warn band is needed in v1.2 for qa/prod/dr HITL review; in dev (autonomous, no reviewer) warn is treated as block.

Recommendation — Python module sketch

# platform/confidence_signal.py  (RESEARCH SKETCH — not implementation)
"""Six-input weighted-sum confidence signal.

Inputs (weights sum to 1.0):
  1. policy_results  (0.30) — list[PolicyCheckResult]
  2. validation     (0.25) — {schema: bool, ir_resolved: bool, tf_validated: bool, tf_planned: bool}
  3. freshness      (0.10) — {age_days: float, max_age_days: float}
  4. source         (0.15) — {submitter: str, commit_sha: str, signed: bool}
  5. history        (0.10) — {prior_rollbacks: int, prior_policy_fails: int}
  6. nfrs           (0.10) — {declared: list[str], conformance: float|None}

Severity → penalty (locked, §8):
  critical → hard override (score = 0, block)
  high     → -0.20
  medium   → -0.05
  low      → -0.01
  info     →  0.00
"""
from dataclasses import dataclass, field
from typing import Literal

WEIGHTS = {"policy":0.30,"validation":0.25,"freshness":0.10,"source":0.15,"history":0.10,"nfrs":0.10}
PENALTY = {"critical":None,"high":0.20,"medium":0.05,"low":0.01,"info":0.0}  # None = hard override
THRESHOLDS = {"dev":0.50,"qa":0.75,"prod":0.90,"dr":0.95}

@dataclass
class Signal:
    score: float
    band: Literal["pass","warn","block"]
    perInput: dict
    reasonCodes: list[str] = field(default_factory=list)

def _per_input_score(name, raw):
    """Normalize a raw input to [0,1]. Spike cold-start: unknown inputs → 0.5."""
    if raw is None:
        return 0.5, f"INPUT_MISSING:{name}"  # halt later if *required*
    # ... per-input scoring rules (see full spec in platform/confidence_signal.py)
    return 0.5, ""

def compute(contract_id, environment, inputs):
    if set(inputs.keys()) != set(WEIGHTS.keys()):
        missing = set(WEIGHTS) - set(inputs.keys())
        return Signal(0.0, "block", {},
                      [f"INPUT_MISSING:{m}" for m in missing])

    per_input = {}
    reasons = []
    base = 0.0
    for name, raw in inputs.items():
        score, reason = _per_input_score(name, raw)
        if reason:
            reasons.append(reason)
        per_input[name] = score
        base += WEIGHTS[name] * score

    # Severity penalties (from the policy_results input)
    penalty = 0.0
    for pcr in inputs.get("policy_results", []):
        if pcr["result"] != "fail":
            continue
        sev = pcr["severity"]
        if PENALTY[sev] is None:  # critical → hard override
            return Signal(0.0, "block", per_input,
                          reasons + [f"CRITICAL_OVERRIDE:{pcr['ruleId']}"])
        penalty += PENALTY[sev]

    score = max(0.0, min(1.0, base - penalty))
    threshold = THRESHOLDS[environment]
    band = "pass" if score >= threshold else ("block" if score < threshold - 0.10 else "warn")
    # dev (autonomous) treats warn as block (no reviewer)
    if environment == "dev" and band == "warn":
        band = "block"
    return Signal(score, band, per_input, reasons)

TARGET 7 — Tiered audit ledger (S3 Object Lock + JWS + chain)

Findings

  • S3 Object Lock modes:
    • Compliance mode: once written, no one (including the root account) can delete or overwrite the object until the retention expires. This is the WORM guarantee regulators want.
    • Governance mode: the root account can delete (with s3:BypassGovernanceRetention permission); privileged users can override. Useful for internal policy, not for regulatory evidence.
    • The architecture locks compliance mode (§9), 7-year retention. For the spike, Object Lock is deferred (see recommendation below) — the spike writes to the DynamoDB outbox + the existing acdl-evidence audit repo (the v1.0 demo's path), and S3 Object Lock is the v1.2 cold-tier build-out.
  • JWS (RFC 7515) detached signature: the event payload is canonical-JSON-serialized, hashed (SHA-256), and signed with a private key; the signature is stored separately (detached) alongside the payload. Verification re-canonicalizes the payload and checks the signature. The signing key question: per-contract (one key per contractId) vs platform (one key for the whole platform).
    • Per-contract: stronger isolation (a key compromise affects one contract), but requires key management per contract (expensive at scale, and the contract's first event has no key yet — a chicken-and- egg).
    • Platform: one signing key (or a small rotation set) for the whole platform. Simpler, matches the "platform is the only writer to the outbox" (§10.3 step 4) design. A compromise of the platform key compromises all evidence, but so would a compromise of the platform itself (the platform is the writer). Recommend platform-level key, stored in AWS KMS, rotated quarterly.
  • prev_event_hash chain: the v1.0 demo's evidence_writer.py (read at /root/acdl/scripts/evidence_writer.py) already implements this exactly:
    • Build event dict with hash = "" (empty string).
    • canonical = json.dumps(event, sort_keys=True, separators=(",", ":")).
    • hash = sha256(canonical.encode("utf-8")).hexdigest().
    • Set event["hash"] = hash.
    • The next event's prev_hash = the previous event's hash.
    • Auto-genesis: if the log is empty, insert a genesis event (seq=0, prev_hash="GENESIS").
    • Lift this verbatim for the spike's outbox events. The JWS signature wraps the same hash (or the canonical JSON) — the chain and the signature are orthogonal: the chain gives ordering/tamper- evidence within the log; JWS gives authenticity per event.
  • Daily checkpoints (§9): a daily job reads the last event hash and writes a "checkpoint" event to the ledger (and optionally to a public notarization service). For the spike, no daily checkpoint (the spike runs in minutes, not days).
  • GitHub audit repo (acdl-evidence) is the hot query index, not part of the chain (§9). The v1.0 demo already commits audit.json to acdl-evidence via finalize_evidence.py. The spike continues this: the outbox is the source of truth; acdl-evidence is a queryable mirror (the timeline UI at evidence-ui/index.html reads it).

Spike minimum (recommendation)

  • Defer S3 Object Lock + JWS to v1.2. The spike's evidence event is: the v1.0 demo's hash-chained shape, written to the DynamoDB outbox (TARGET 5) and mirrored to acdl-evidence (unchanged from v1.0).
  • The spike's audit_ledger_design.md (Phase 07) documents the v1.2 S3-Object-Lock-compliance + JWS + KMS-key + daily-checkpoint design in full, but the spike implements only the hash-chain + outbox write.
  • This is acceptable because the spike's goal (REQ-28) is to prove the IR commitments hold, not to prove the audit ledger's regulatory posture. The audit ledger's regulatory posture is a design artifact (Phase 07) + a v1.2 build.

Confidence

0.85 — the v1.0 demo's hash chain is read and confirmed; the deferral of Object Lock + JWS is a scope decision, not a design risk.

Assumptions logged

  • A-7.1 (0.85): the spike can defer S3 Object Lock + JWS to v1.2 without invalidating REQ-20 ("tiered audit ledger design authored") — REQ-20 is design, not implementation, and Phase 07 authors the design.
  • A-7.2 (0.80): platform-level KMS signing key is the right v1.2 choice (vs per-contract); the spike does not exercise this.

Recommendation — spike evidence event shape (lifted from v1.0 + outbox)

{
  "seq": 1,
  "ts": "2026-07-21T12:00:00Z",
  "stage": "dev",
  "event": "contract applied: l2-static-assets (confidence 0.82, band pass)",
  "prev_hash": "<sha256 of the genesis event, or GENESIS>",
  "hash": "<sha256 of the canonical JSON of this event with hash=''>",
  "contractId": "uuid",
  "environment": "dev",
  "stack": "l2-static-assets",
  "score": 0.82,
  "band": "pass"
}

The v1.2 addition is a jws field (detached signature over the canonical JSON) and a move of the storage from the outbox-only path to the outbox→S3-Object-Lock path.


TARGET 8 — HITL matrix + separation of duties

Findings

  • GitHub Environments + required reviewers: GitHub-native; Gitea has no Environments API (confirmed in v1.0 research, ARCHITECTURE.md "Gitea API surface" table; re-confirmed in the Gitea docs "Compared to GitHub Actions" → "jobs.<job_id>.environment ... It's ignored by Gitea Actions now"). The v1.0 demo's workaround (D-013: workflow_dispatch approval inputs) is the only available Gitea-native gate.
  • CODEOWNERS: Gitea supports CODEOWNERS files (for PR review routing); this is the routing layer (§10.2). It does not enforce identity distinctness (§10.3 — that's the DynamoDB outbox check).
  • The spike is dev-only (REQ-27 contract has environment: dev), so HITL is not exercised in the spike. Phase 07 authors the design; Phase 10's verify_phase10.sh does not assert any HITL behavior.
  • Identity-distinctness check sequence (§10.3):
    1. dev→qa promotion: read QA approver GitHub identity from the GitHub Deployment approval event → write to outbox keyed by contractId.
    2. qa→prod: read stored QA approver from outbox + new SRE approver from the approval event.
    3. If qaApprover == prodApprover: block, emit SEPARATION_OF_DUTIES_VIOLATION, route halt artifact to SRE on-call.
    4. The check is in the central pipeline (the platform is the only outbox writer).
    • Gitea adaptation: there is no "GitHub Deployment approval event" in Gitea. The v1.0 demo modeled this as a workflow_dispatch input (approve_qa: true / approve_prod: true). The approver identity in Gitea is gitea.actor of the dispatch event. The design doc must specify: "the approver identity is gitea.actor of the workflow_dispatch run that sets approve_qa=true (resp. approve_prod=true)."
  • Timeout (§10.5): 1 business day = warn + escalate; 2 business days = auto-freeze + re-submit (linked via supersedes). The spike does not implement the timer (no HITL in spike); the design doc specifies it as a Gitea scheduled workflow (on: schedule) that scans the outbox for PENDING_ATTESTATION events older than 1/2 business days and emits the warning/freeze events.
  • Rejection (§10.6): returns contract to HELD state; new submission linked via supersedes. The supersedes field is a contract-schema field (TARGET 9) pointing at the prior contractId.

Confidence

0.85 — the Gitea adaptation (using gitea.actor of the dispatch event) is the only viable path given the no-Environments-API constraint; it is documented in the v1.0 research and re-confirmed.

Assumptions logged

  • A-8.1 (0.90): the spike does not exercise HITL (dev-only); Phase 07 authors the design; v1.2 wires it.
  • A-8.2 (0.85): gitea.actor of the workflow_dispatch run is the approver identity of record. This is Gitea's only available approval-identity signal.
  • A-8.3 (0.80): the 1d/2d timeout is a v1.2 scheduled workflow (no spike implementation).

Recommendation — design doc sketch (platform/hitl_matrix_design.md)

The Phase 07 design doc should contain:

  1. The full 8-concern attestation matrix (lifted from docs/architecture.md §10.4, formatted as a Markdown table).
  2. The Gitea-specific pre-execution gate model:
    • qa gate: workflow_dispatch with approve_qa: true; the dispatch run's gitea.actor is the QA approver.
    • prod gate: workflow_dispatch with approve_prod: true; the dispatch run's gitea.actor is the SRE approver.
    • dr gate: workflow_dispatch with approve_dr: true; same.
  3. The separation-of-duties check (platform/separation_of_duties.py): reads approver_qa from the outbox for the contractId, compares to the new gitea.actor of the prod-dispatch run; blocks on equality; emits SEPARATION_OF_DUTIES_VIOLATION.
  4. The timeout design: a on: schedule workflow (runs hourly) that scans the outbox for PENDING_ATTESTATION events with ts older than 1/2 business days and emits the warn/freeze events.
  5. The supersedes linking on rejection (a contract-schema field).
# platform/separation_of_duties.py  (RESEARCH SKETCH — not implementation)
"""Check that qaApprover != prodApprover for a contract.

Reads the outbox for the contractId; returns (ok, reason).
Spike: always returns (True, "dev-only") because the spike is dev-only.
"""
def check(outbox_client, contract_id, current_prod_approver):
    item = outbox_client.get(contract_id)
    if item is None:
        return True, "no prior approver (first promotion)"  # dev→qa has no SoD check
    qa_approver = item.get("approver_qa")
    if not qa_approver:
        return True, "no QA approver recorded (dev-only spike)"
    if qa_approver == current_prod_approver:
        return False, f"SEPARATION_OF_DUTIES_VIOLATION: qaApprover==prodApprover=={qa_approver}"
    return True, "distinct"

TARGET 9 — Contract schema (JSON Schema draft 2020-12)

Findings

  • Per-env mandatory/optional (W3.E, resolved in PROJECT.md):
    • dev requires: stack, environment
    • qa adds: validation.e2eSuite, validation.loadTest
    • prod adds: runbook, dashboard, oncall
    • dr adds: drDrillRef
    • inputs always optional
    • profile: agentic fields optional everywhere
  • profile: agentic marker unlocks L3B fields: naturalLanguageIntent, confidenceAtSubmission, agentTrace (architecture.md §5).
  • Central repo + generated clients: the spike uses a local schemas/contract.schema.json; the central repo + generated client libraries are v1.2 (architecture.md §7).
  • Fail-fast with reason codes: schema validation failure produces a reason code from a published vocabulary (e.g. SCHEMA_MISSING_REQUIRED:stack, SCHEMA_TYPE_MISMATCH:inputs.bucket_name, SCHEMA_UNKNOWN_PROFILE). The vocabulary is a JSON list in schemas/reason_codes.json (Phase 07 authoring).
  • supersedes field (from TARGET 8): optional, points at the prior contractId on re-submission after rejection.

Confidence

0.85 — the per-env mandatory table is locked (W3.E); the schema shape is a direct formalization.

Assumptions logged

  • A-9.1 (0.90): the spike's contracts/spike.yaml has only stack, environment, and inputs — the minimal dev contract.
  • A-9.2 (0.85): inputs is a free-form object with string values (matching the v1.0 demo's contract.yaml shape) for v1; typed inputs per-L1 is a v1.2 enhancement (the IR's resource.inputs is typed, but the contract's inputs is the L2-level param map, free-form in v1).

Recommendation — schemas/contract.schema.json sketch

{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://acdl.cloudinit.dev/schemas/contract.schema.json",
  "title": "ACDL Contract",
  "type": "object",
  "required": ["stack", "environment"],
  "properties": {
    "stack": { "type": "string", "pattern": "^l2-[a-z][a-z0-9-]*$" },
    "environment": { "enum": ["dev", "qa", "prod", "dr"] },
    "inputs": {
      "type": "object",
      "additionalProperties": { "type": ["string","number","boolean"] },
      "description": "L2-level parameter map; free-form in v1, typed in v1.2"
    },
    "validation": {
      "type": "object",
      "properties": {
        "e2eSuite":  { "type": "string", "description": "ref to the e2e suite" },
        "loadTest":  { "type": "string" }
      }
    },
    "runbook":    { "type": "string" },
    "dashboard": { "type": "string" },
    "oncall":    { "type": "string" },
    "drDrillRef":{ "type": "string" },
    "profile":   { "enum": ["developer", "agentic"], "default": "developer" },
    "naturalLanguageIntent":  { "type": "string" },
    "confidenceAtSubmission": { "type": "number", "minimum": 0, "maximum": 1 },
    "agentTrace":             { "type": "string" },
    "supersedes": { "type": "string", "format": "uuid",
      "description": "prior contractId this re-submission replaces (after rejection)" }
  },
  "allOf": [
    { "if": { "properties": { "environment": { "const": "qa" } } },
      "then": { "required": ["validation"] } },
    { "if": { "properties": { "environment": { "const": "prod" } } },
      "then": { "required": ["runbook", "dashboard", "oncall"] } },
    { "if": { "properties": { "environment": { "const": "dr" } } },
      "then": { "required": ["drDrillRef"] } },
    { "if": { "properties": { "profile": { "const": "agentic" } } },
      "then": { "required": ["naturalLanguageIntent"] } }
  ]
}

Spike contract (contracts/spike.yaml) validates against this:

stack: l2-static-assets
environment: dev
inputs:
  bucket_name: acdl-spike-bucket
  region: us-east-1

dev requires only stack + environment; inputs optional; no profile (defaults to developer). Passes.


TARGET 10 — Archive strategy (v1.0 demo → demo/)

Findings (static analysis of the demo's path references)

The demo's scripts resolve paths relative to REPO_ROOT (the parent of scripts/). Moving the demo to demo/ means REPO_ROOT becomes the demo/ directory, and the relative paths modules/..., scripts/..., evidence-ui/... must still resolve inside demo/.

Path references found (static analysis only — no execution):

  1. scripts/run_demo.sh (read at /root/acdl/scripts/run_demo.sh):

    • SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
    • REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
    • ev() calls python3 "$SCRIPT_DIR/evidence_writer.py" — resolves to demo/scripts/evidence_writer.py ✓ (inside demo/).
    • python3 "$SCRIPT_DIR/policy_checker.py"
    • python3 "$SCRIPT_DIR/confidence_signal.py"
    • python3 "$SCRIPT_DIR/l3b_agent_stub.py"
    • python3 "$SCRIPT_DIR/finalize_evidence.py" ✓ (the upload path; skipped under --no-upload).
    • ( cd "$REPO_ROOT" && bash "$SCRIPT_DIR/mock_executor.sh" ... )REPO_ROOT = demo/, mock_executor.sh resolves L2 manifests relative to its cwd.
    • rm -f "$REPO_ROOT/state.json" — cleans up demo/state.json ✓.
    • --audit "$REPO_ROOT/evidence-ui/index.html" — resolves to demo/evidence-ui/index.html ✓.
  2. scripts/mock_executor.sh (read at /root/acdl/scripts/mock_executor.sh):

    • L2_MANIFEST="modules/l2/${STACK}/manifest.yaml" — resolved relative to cwd (REPO_ROOT = demo/), so this becomes demo/modules/l2/<stack>/manifest.yaml ✓.
    • L1_SCRIPT="modules/l1/${L1_NAME}/mock_apply.sh" — becomes demo/modules/l1/<l1>/mock_apply.sh ✓.
  3. .gitea/workflows/pipeline.yml (read at /root/acdl/.gitea/workflows/pipeline.yml): uses python3 scripts/policy_checker.py, scripts/confidence_signal.py, scripts/evidence_writer.py, scripts/finalize_evidence.py, scripts/mock_executor.sh — all relative to the checkout root. After the move, the checkout root for the demo workflow is demo/ (if the workflow is moved to demo/.gitea/workflows/), OR the workflow stays at .gitea/workflows/ and references demo/scripts/... (if the workflow is kept at the repo root for the demo to be runnable from the repo root). Decision: move the demo workflows to demo/.gitea/workflows/ so the demo is fully self-contained under demo/. The v1.0 demo's Gitea Actions runs are historical (tag v1.1.0); future demo re-runs are local (run_demo.sh --no-upload), not on Gitea Actions.

  4. scripts/verify_phase01..05.sh: these verify the demo's phases. After the move, they run from demo/scripts/ and reference modules/, scripts/, evidence-ui/ relative to demo/. The verify scripts use the same REPO_ROOT-relative pattern. They will continue to work from demo/ because all paths are relative to the script's parent.

  5. evidence-ui/index.html: fetches audit.json from a relative URL (the v1.0 demo's acdl-evidence raw URL). No path fixup needed; the UI is static.

Path fixups needed

None. All demo scripts resolve paths via SCRIPT_DIR/REPO_ROOT relative-to-script, which auto-adjusts when the demo moves to demo/. The only thing to verify (in Phase 06's verify_phase06.sh) is that demo/scripts/run_demo.sh --no-upload exits 0 from demo/ — but per the task instructions, this is static analysis only; Phase 06 executes the regression check.

Confidence

0.90 — the static analysis is thorough; the 0.10 residual is for an undetected absolute path in a script I did not read (e.g. scripts/gitea_setup.sh, scripts/finalize_evidence.py).

Assumptions logged

  • A-10.1 (0.90): the demo's scripts use only SCRIPT_DIR/REPO_ROOT relative paths; no absolute paths. (Verifyable by a grep -n '/' pass in Phase 06.)
  • A-10.2 (0.85): moving .gitea/workflows/pipeline.yml to demo/.gitea/workflows/pipeline.yml does not break any historical Gitea Actions run (those are pinned to tag v1.1.0 and immutable).

Recommendation — file-move list

From (repo root) To
modules/ demo/modules/
scripts/ demo/scripts/
evidence-ui/ demo/evidence-ui/
contracts/ demo/contracts/
.gitea/workflows/pipeline.yml demo/.gitea/workflows/pipeline.yml
.gitea/workflows/.gitkeep demo/.gitea/workflows/.gitkeep
ACDL_DEMO.md demo/ACDL_DEMO.md
contracts-repo/ demo/contracts-repo/ (if it is demo-only; verify in Phase 06)
runner-data/ demo/runner-data/ (if it is demo-only; verify in Phase 06)

New top-level dirs (scaffolded empty in Phase 06): platform/, schemas/, adapters/, terraform/, modules-ir/.

README rewrite: reflect the real platform (vision + architecture links, new layout); the demo README moves to demo/README.md (or demo/ACDL_DEMO.md).

Path fixups in demo/scripts/*.sh and demo/.gitea/workflows/*.yml: none (all paths are relative-to-script). The only fixup may be in demo/.gitea/workflows/pipeline.yml's ref: milestone/v1.0-initial (the demo branch) — if the demo is re-run on Gitea Actions (it won't be; the demo is local-only post-archive), but this is moot for the archive.


Risks (highest-risk items that could block Phase 0810, ranked)

  1. R-1 (HIGHEST): Gitea OIDC gap → spike must use a key-rotation waiver.

    • Risk: Phase 08 cannot configure real AWS OIDC trust (TARGET 1). The spike falls back to D-039 (per-run-rotated long-lived key).
    • Mitigation: adopt D-039 (this research's recommendation); Phase 08 implements scripts/rotate_spike_key.sh; the v1.2 path tracks go-gitea/gitea#36988.
    • Impact if unmitigated: Phase 08's success criterion ("a workflow step assumes the role via OIDC with no long-lived credential") is unachievable in this environment. The waiver is the only path.
  2. R-2: Checkov terraform_plan framework edge cases on the spike's minimal S3 plan.

    • Risk: Checkov's terraform_plan scanner ignores a few checks (CKV_AWS_217, 233, 237 — per the plan-scanning doc) that rely on lifecycle blocks not present in plan JSON. The spike's S3-only plan may surface fewer checks than expected, making the "all six inputs present" gate artificially pass.
    • Mitigation: the adapter (TARGET 4) emits the ACDL_TAG_NAMING skipped record to guarantee the policy input is non-empty; Phase 10's verify_phase10.sh asserts the confidence signal's perInput.policy is present and non-null.
    • Impact if unmitigated: the spike could pass without exercising the policy path meaningfully.
  3. R-3: The 6 confidence-signal inputs are a recommendation, not locked.

    • Risk: the architecture (§8) does not enumerate the 6 inputs; this research chose them (policy, validation, freshness, source, history, nfrs). If lead-developer or security-engineer disagrees, Phase 07 re-opens the design.
    • Mitigation: surface as D-040 (this research's recommended 6 + weights); Phase 07 adopts or amends.
    • Impact if unmitigated: Phase 07 scope creep; the spike's confidence signal (Phase 10) blocks on the enumeration.
  4. R-4: gitea-runner rename (act_runnergitea-runner).

    • Risk: the v1.0 docs and the architecture refer to act_runner; the runner was renamed in 2026-04 (gitea/runner#850). The binary is now gitea-runner, the image gitea/runner. Documentation drift.
    • Mitigation: Phase 07 updates ARCHITECTURE.md to use gitea-runner (the current name) with a note that v1.0 used act_runner.
    • Impact if unmitigated: confusion in Phase 08 operator docs; low impact.
  5. R-5: Demo archive may have an undetected absolute path.

    • Risk: TARGET 10's static analysis may have missed an absolute path in a script I did not read (gitea_setup.sh, finalize_evidence.py, l3b_agent_stub.py).
    • Mitigation: Phase 06's verify_phase06.sh runs demo/scripts/run_demo.sh --no-upload and asserts exit 0; a failure is caught there.
    • Impact if unmitigated: Phase 06 regression failure; fixable in-phase.
  6. R-6: S3 Object Lock + JWS deferred to v1.2 may be challenged.

    • Risk: REQ-20 says "tiered audit ledger design is authored" — this is satisfied by Phase 07 design doc. But if a reviewer reads REQ-20 as "implemented," the spike falls short.
    • Mitigation: PROJECT.md's REQ-20 wording is "design authored"; Phase 07 produces the design; v1.2 implements. Re-confirm in Phase 07's success criteria.
    • Impact if unmitigated: milestone audit debate; resolvable by pointing at the design doc.

Decisions surfaced (proposals for PROJECT.md adoption by lead-developer)

ID Decision Rationale Confidence Alternatives
D-039 Spike-only waiver: per-run-rotated long-lived AWS key (D-034's bootstrap key, rotated after each spike run by scripts/rotate_spike_key.sh). OIDC federation deferred to v1.2, blocked on go-gitea/gitea#36988. Gitea Actions does not support id-token: write / OIDC token issuance (TARGET 1). The waiver satisfies §12.5's intent (no persistent long-lived key) for the spike; v1.2 implements real OIDC when the Gitea PR merges. 0.95 (b) self-hosted OIDC broker [heavy for a spike]; (d) scheduled credential mint [reintroduces upstream long-lived key]; (e) LocalStack [invalidates REQ-23].
D-040 The 6 confidence-signal inputs are: policy (0.30), validation (0.25), freshness (0.10), source (0.15), history (0.10), nfrs (0.10). Weights frozen for v1, tuned in v1.2 alongside thresholds (BA.B). The architecture (§8) locks "six canonical inputs" but does not enumerate them. This research's enumeration covers the platform-computable subset present in every environment (including dev). 0.75 Different input set (e.g. split "validation" into schema/IR/plan); different weights.
D-041 Spike audit ledger = v1.0 hash chain + DynamoDB outbox + acdl-evidence mirror. S3 Object Lock (compliance mode, 7-yr) + JWS (platform-level KMS signing key, rotated quarterly) + daily checkpoints are v1.2 build-out, authored as design in Phase 07 (platform/audit_ledger_design.md). REQ-20 is "design authored," not "implemented." The spike proves the outbox write path; the regulatory ledger is a v1.2 build. 0.85 Implement Object Lock in the spike (scope creep; the spike's goal is the IR commitments, not the ledger).
D-042 HITL approver identity in Gitea = gitea.actor of the workflow_dispatch run that sets approve_qa=true / approve_prod=true / approve_dr=true. Separation-of-duties reads approver_qa from the DynamoDB outbox and compares to the prod-dispatch gitea.actor. Gitea has no Environments API and ignores environment: blocks (v1.0 D-013; re-confirmed). gitea.actor is the only available approval-identity signal. 0.85 Self-hosted approval portal (over-engineering for v1).
D-043 Tag/naming compliance is deferred for the spike: the Checkov adapter emits a single SKIPPED PolicyCheckResult (ruleId: ACDL_TAG_NAMING, severity: info) so the confidence signal's policy input is non-empty. A custom Checkov YAML rule for tag presence lands in v1.2. Checkov has no built-in tag-presence check for general AWS resources. Writing a custom Checkov rule in the spike is scope creep; the spike's goal is the IR + adapter path. 0.80 Write the custom Checkov rule in the spike (adds a Phase 09/10 dependency on Checkov's custom-rule API).
D-044 DynamoDB outbox mode = PAY_PER_REQUEST (on-demand); PK contractId, SK eventType#eventTs; TTL expire_at = now + 365d. No separate async worker / DLQ in the spike (RTO = workflow re-run); the v1.2 outbox worker + DLQ is a Phase 07 design artifact. On-demand is the zero-cost-at-idle spike mode; the spike is a single dev-only submission. 0.85 Provisioned capacity (overkill for the spike); separate Lambda worker (v1.2 scope).
D-045 Runner image tooling: runs-on: ubuntu-latest (default gitea/runner-images:ubuntu-latest); install terraform via HashiCorp apt repo (pin 1.9.*), checkov via pip (pin >=3.2,<4), use --break-system-packages for the pip install inside the ephemeral container. Neither tool is pre-installed on the default runner image (TARGET 2). Pinning avoids mid-spike version drift. 0.85 tfenv (extra shell-init step, unnecessary for single-version spike); gitea/runner-images:ubuntu-latest-full (amd64-only, large, not needed).
D-046 act_runnergitea-runner rename: Phase 07 updates ARCHITECTURE.md and PROJECT.md to use the current name gitea-runner (formerly act_runner, renamed 2026-04 in gitea/runner#850). The binary is gitea-runner; the image is gitea/runner. Naming drift between v1.0 docs and the current runner. 0.90 Keep act_runner (stale).

v1.2 Research Addendum (Phase 11, 2026-07-21)

Phase: research (Phase 11). Milestone: v1.2. Status: active. Researcher: ci-researcher (inline, docs phase). Autonomy: full. Sources: GitHub API (go-gitea/gitea#36988), ACDL codebase audit (terraform/bootstrap/, scripts/, adapters/terraform/, modules-ir/registry.json, .ciagent/VERIFY.md, .ciagent/PERSONAS.md). Scope: re-eval OIDC blocker, NFR audit of the v1.1 spike, simplification opportunities, README rewrite plan, ECS L1 catalog scoping.

TARGET 9 — go-gitea/gitea#36988 re-check (v1.2)

Verdict (conf 0.95): still open, not merged.

Re-checked 2026-07-21 via api.github.com/repos/go-gitea/gitea/pulls/36988:

  • state: open
  • merged: false
  • merged_at: null
  • updated_at: 2026-05-27T16:26:24Z
  • title: "Add Actions OIDC provider with workflow permission gating"

No movement since the v1.1 research (2026-07-21 v1.1 research also found it open). Real OIDC federation remains impossible for Gitea Actions. D-047 adopts: extend the D-039 per-run-rotated-key waiver for v1.2; real OIDC is deferred to v1.3+. The waiver continues to satisfy §12.5's intent: no persistently long-lived key (scripts/rotate_spike_key.sh rotates after each run; Phase 12 tightens IAM scoping + rotation hygiene).

TARGET 10 — NFR audit of the v1.1 spike

Audited the v1.1 spike's operational code for NFR gaps.

terraform/bootstrap/spike_runner_policy.json — least-privilege PASS already. Explicit Allow list (S3 state bucket R/W, DynamoDB outbox R/W, sts:GetCallerIdentity) + DenyEverythingElse on * with NotResource. No wildcards in the Allow statements. v1.2 gap: the policy only covers S3 + DynamoDB + STS — Phase 15's terraform apply to ECS needs ECS + ECR + ELB + IAM (plan + apply) permissions added. Phase 12 scopes the policy expansion; Phase 15 applies it.

terraform/bootstrap/create_state_backend.py — idempotent PASS already. head_bucket → skip-create if exists; describe_table → skip-create if exists; put_bucket_versioning is idempotent. No v1.2 change needed.

terraform/bootstrap/create_iam_user.py — idempotent PASS already. get_user → skip-create if exists; put_user_policy overwrites (idempotent); list_access_keys → skip-create if an active key exists. No v1.2 change needed.

scripts/run_spike_plan.sh + scripts/run_spike_e2e.sh — two scripts, overlapping setup (env loading, cd terraform/spike, terraform init). run_spike_e2e.sh is the superset (full pipeline); run_spike_plan.sh is the plan-only subset. v1.2 simplification (Phase 12): consolidate into one scripts/run_platform.sh with a --plan-only flag (default: full e2e). Removes ~30 lines of duplication.

scripts/rotate_spike_key.sh — idempotent PASS (always ends with exactly 1 active key). Uses the bootstrap root key to rotate; documented that D-034 closure (root key deactivation) is a manual user step. No v1.2 change needed (the root key is now deactivated per D-034 closure; rotation uses the spike key itself or a separate rotation credential — flagged as a v1.2 operational note in Phase 12).

Error handling: run_spike_e2e.sh uses set -u + a fail() helper — good. run_spike_plan.sh uses set -u + inline exits — adequate. The consolidated run_platform.sh should use set -euo pipefail + fail() for uniform strictness.

P1-1 redaction target (Phase 12 — DONE): the v1.1 .ciagent/ audit narrative referenced two AWS access key IDs (AKIA…SPIKE rotated spike key, AKIA…ROOT-DEACTIVATED deactivated root key). Public identifiers, not secret pairs, in the audit narrative not executable code. Phase 12 redacted them to AKIA…SPIKE / AKIA…ROOT-DEACTIVATED across .ciagent/RESEARCH.md, PROJECT.md, REVIEW.md, AUDIT.md.

P1-B stale paths: .ciagent/PERSONAS.md line 47 still has platform/registry/** (the rest were fixed at ab69d10). Phase 12 fixes line 47 to acdl_platform/registry/** (or removes it — there is no acdl_platform/registry/ dir; the registry is modules-ir/registry.json).

TARGET 11 — Simplification opportunities

  1. Script consolidation (above): run_spike_*.shrun_platform.sh.
  2. terraform/spike/.terraform/ artifacts: gitignored already (.gitignore covers .terraform/, .terraform.lock.hcl, tfplan, *.tfstate*). No change.
  3. acdl_platform/__pycache__/: gitignored already. No change.
  4. Dead code: none found — the spike is tight. The run_spike_plan.sh script is the only redundancy (subsumed by run_platform.sh --plan-only).
  5. demo/ archive: correctly separated; no v1.2 touch.

TARGET 12 — README rewrite plan

Current README.md (51 lines) is stale: "v1.1 (active)" framing, no "how to run the platform" section, no v1.2 objective. Phase 11 rewrites it to reflect:

  • v1.1 complete (tag v1.2.0); v1.0 demo archived under demo/.
  • The actual spike flow: contract → IR → terraform plan → Checkov → confidence signal → outbox.
  • How to run: scripts/run_platform.sh (after Phase 12; for now scripts/run_spike_e2e.sh).
  • Real repo layout table (the existing one is accurate; refresh the "Populated" column).
  • v1.2 objective (platform hardening + ECS microservice).

TARGET 13 — ECS L1 catalog scoping (for Phase 13)

Six L1s needed for an ECS Fargate microservice. Each maps to one or more AWS Terraform resources; the adapter TYPE_MAP (currently {"aws:s3:bucket": "aws_s3_bucket"}) needs expansion:

L1 IR type(s) Terraform resource(s) Key inputs
l1-vpc aws:ec2:vpc, aws:ec2:subnet, aws:ec2:routetable aws_vpc, aws_subnet, aws_route_table + associations cidr, azs
l1-ecs-cluster aws:ecs:cluster aws_ecs_cluster name
l1-ecs-service aws:ecs:service, aws:ecs:task_definition aws_ecs_service, aws_ecs_task_definition image, port, cpu, memory, env
l1-iam-role aws:iam:role, aws:iam:rolepolicyattachment aws_iam_role, aws_iam_role_policy_attachment task + exec role
l1-alb aws:elbv2:loadbalancer, aws:elbv2:listener, aws:elbv2:targetgroup aws_lb, aws_lb_listener, aws_lb_target_group port, protocol
l1-ecr aws:ecr:repository aws_ecr_repository name

The IR schema (schemas/ir.schema.json) is engine-agnostic and already supports arbitrary resource types — no schema change needed, only new interface.json files + TYPE_MAP entries. The l2-microservice thin-composition references all six (depth ≤ 5).

Decisions surfaced (v1.2)

ID Decision Rationale Confidence Alternatives
D-047 Extend D-039 per-run-rotated-key waiver for v1.2. Real OIDC deferred to v1.3+. go-gitea/gitea#36988 still open (TARGET 9). The waiver satisfies §12.5's intent for v1.2; Phase 12 tightens IAM + rotation hygiene. 0.95 (a) wait for #36988 (blocks v1.2 indefinitely); (b) self-hosted OIDC broker (heavy); (c) KMS-backed ephemeral creds (scope creep for v1.2).
D-048 Consolidate run_spike_plan.sh + run_spike_e2e.sh → one scripts/run_platform.sh with --plan-only flag (default: full e2e). Two scripts with overlapping setup (~30 lines duplicated). One script with a flag is simpler and matches the "streamline" scope axis. 0.90 Keep both (redundant); delete run_spike_plan.sh only (loses the plan-only convenience).
D-049 v1.2 L1 catalog = 6 L1s (l1-vpc, l1-ecs-cluster, l1-ecs-service, l1-iam-role, l1-alb, l1-ecr). The adapter TYPE_MAP expands to 9 IR types (3 new for VPC, 3 for the rest). Minimal set to deploy an ECS Fargate service end-to-end. VPC is split into vpc/subnet/routetable because the IR models one resource per interface.json entry, but the L1 groups them. 0.85 Fewer L1s (e.g. fold VPC into the ECS service — violates L1 single-purpose); more L1s (e.g. separate l1-securitygroup — scope creep for v1.2).

v1.8 Research Addendum

Phase: research (pre-Phase 28). Milestone: v1.8. Status: active. Researcher: ci-researcher. Autonomy: full. Sources: web (uptime-kuma GitHub, Terraform docs, AWS KMS docs, AWS ECS Fargate docs, GitHub Actions docs) + ACDL codebase analysis.

RESEARCH TARGET 1 — uptime-kuma deployment on ECS Fargate

Verdict: ECS Fargate is the most cost-effective cloud-native option for deploying uptime-kuma, consistent with the existing platform primitives (ecs-cluster, ecs-service, alb).

Findings (verified 2026-07-22):

  1. uptime-kuma Docker image: louislam/uptime-kuma:1 (v1) or louislam/uptime-kuma:2 (v2, latest stable 2.4.0 as of 2026-05-31). The container listens on port 3001. Data is stored in /app/data (SQLite + uploaded files). NFS is not supported for the data volume; EFS is the AWS-native equivalent and works with ECS Fargate.

  2. Monitoring capabilities: HTTP(s), TCP, HTTP(s) Keyword, HTTP(s) JSON Query, WebSocket, Ping, DNS Record, Push, Steam Game Server, Docker Containers. 20-second intervals minimum. Certificate info. Proxy support. 2FA support.

  3. Notification services (90+): Telegram, Discord, Gotify, Slack, Pushover, Email (SMTP), Microsoft Teams (via webhook), and many others. For the ACDL primitive, we expose: Teams webhook, email (SMTP), SMS (via SNS or an external gateway), and GitHub issues (via the GitHub API).

  4. ECS Fargate deployment shape:

    • Task definition: 1 container (louislam/uptime-kuma:1), port 3001, CPU 256 (.25 vCPU), Memory 512 (.5 GB) — minimal cost (~$5/mo at us-east-1 on-demand pricing for .25 vCPU + .5 GB running 24/7).
    • EFS volume for /app/data (persistent storage across task restarts; Fargate + EFS is the standard pattern for stateful containers).
    • ALB + listener for a stable public URL (the uptime dashboard).
    • CloudWatch log group (encrypted with the per-stack CMK).
  5. Endpoint seeding: uptime-kuma has a REST API (socket.io-based). The platform can seed monitors by either:

    • (a) Passing UPTIMA_KUMA__monitors env var (JSON array) consumed by a startup script — but uptime-kuma does not natively read env for monitor config.
    • (b) A post-deploy seeding script that calls the uptime-kuma API (POST /api/monitor) to create monitors from the monitored_endpoints input. This is the cleaner approach — the platform runs a Python script after the ECS service is up that creates monitors via the API.
    • Recommendation: (b) — a scripts/seed_uptime_monitors.py that reads the monitored_endpoints from the stack outputs + calls the uptime-kuma API. This is testable offline (mocked API) and decouples container startup from monitor configuration.
  6. Separate terraform state: The uptime stack uses a separate S3 key prefix (uptime/{consumerRepo}/{contractId}/) so it is independent of the consumer stack's state. The uptime stack has its own VPC + ALB + ECS cluster (or shares the consumer's — design decision: separate to avoid state coupling, per the requirement "separate terraform run, with a separate state").

  7. Feature flag: The feature_flag_enabled input (set from the consumer contract inputs.uptime_enabled, default true) controls whether the deploy-uptime pipeline stage runs. When false, the stage is skipped entirely (no resources emitted, no API calls).

RESEARCH TARGET 2 — Terraform prevent_destroy lifecycle

Verdict: lifecycle { prevent_destroy = true } is the correct Terraform mechanism for deletion protection. It prevents terraform destroy from destroying the resource without first setting prevent_destroy = false.

Findings (verified 2026-07-22):

  1. prevent_destroy is a meta-argument inside a lifecycle {} block within a resource. When set to true, any Terraform plan that would destroy the resource will fail with an error. To destroy, the user must first set prevent_destroy = false and apply, then destroy.

  2. This is exactly the 2-step decommission pattern the user requested: Step 1: set deletion_protection = false (which the adapter translates to prevent_destroy = false) + apply. Step 2: set all counts to 0 + apply (which destroys the resources now that prevent_destroy is false).

  3. Adapter emission: The adapter should emit lifecycle { prevent_destroy = true } inside each resource block when the deletion_protection NFR is true. When false, omit the lifecycle block (or set prevent_destroy = false). This is a per-resource meta-argument, not a provider-level setting.

  4. RDS special case: RDS already has a deletion_protection argument on aws_db_instance (not a lifecycle meta-arg). The adapter should emit BOTH: the deletion_protection argument (for the RDS API-level protection) AND lifecycle { prevent_destroy = true } (for the Terraform-level protection). This is defense-in-depth.

RESEARCH TARGET 3 — AWS KMS key rotation

Verdict: enable_key_rotation = true on aws_kms_key enables automatic annual rotation (AWS rotates the key material annually). For 90-day rotation, a custom key rotation policy is needed (AWS managed rotation is annual only; 90-day requires a manual rotation schedule or a custom multi-region key + rotation Lambda).

Findings (verified 2026-07-22):

  1. aws_kms_key with enable_key_rotation = true enables AWS's automatic key material rotation. AWS rotates the backing key material annually (365 days). This is the simplest option and is the AWS best practice for most use cases.

  2. 90-day rotation: AWS does not support custom rotation periods for managed keys. To achieve 90-day rotation:

    • (a) Use aws_kms_key with enable_key_rotation = true (annual AWS-managed rotation) + a CloudWatch Events rule that triggers a Lambda every 90 days to create a new key + update the alias. This is complex and overkill for v1.8.
    • (b) Accept annual AWS-managed rotation as the default and document that 90-day rotation requires a custom rotation pipeline (roadmap item). The enable_key_rotation = true is the v1.8 implementation; the 90-day requirement is a roadmap enhancement.

    Recommendation: (b) — enable_key_rotation = true (AWS-managed annual rotation) as the v1.8 implementation. The 90-day requirement is documented as a roadmap item (custom rotation Lambda). The NFR enable_rotation (default true) controls the enable_key_rotation argument. This is pragmatic; annual rotation is AWS's best practice and 90-day is a future enhancement.

  3. Per-stack CMK pattern: Each L2 deployment creates its own aws_kms_key + aws_kms_alias (alias/acdl--). The key is tagged with acdl:owner + acdl:environment. All primitives in the stack reference this key via kms_key_arn. No shared keys across stacks.

  4. Managed KMS fallback: When a primitive is deployed standalone (L1 without an L2 CMK), the adapter uses alias/aws/<service> (e.g. alias/aws/s3, alias/aws/rds). This is the AWS-managed key for that service. The adapter emits a stderr warning when falling back. The kms_key_arn input is optional; the encryption_enabled NFR defaults to true.

RESEARCH TARGET 4 — Forge-agnostic API URLs (P1-9)

Verdict: GitHub and Gitea have compatible issue APIs but different search endpoints. A GITHUB_API_BASE env var + _forge_type() helper branches the search URL.

Findings (verified 2026-07-22):

  1. GitHub API: https://api.github.com/search/issues?q=... for search; https://api.github.com/repos/{owner}/{repo}/issues for create; https://api.github.com/repos/{owner}/{repo}/issues/{n}/comments for comments.

  2. Gitea API: https://git.cloudinit.dev/api/v1/repos/{owner}/{repo}/issues?... for search (no /search/issues endpoint — issues are listed via the repo issues endpoint with query params); https://git.cloudinit.dev/api/v1/repos/{owner}/{repo}/issues for create; https://git.cloudinit.dev/api/v1/repos/{owner}/{repo}/issues/{n}/comments for comments.

  3. Detection: If GITHUB_API_BASE contains /api/v1, it's Gitea; otherwise it's GitHub. The _forge_type() helper returns "gitea" or "github" based on this. The search URL is branched accordingly; the create + comment URLs are the same pattern ({base}/repos/{owner}/{repo}/issues).

  4. Auth: Both use Authorization: token <token> header. GitHub also accepts Authorization: Bearer <token>; Gitea uses token. The existing token header works for both.

RESEARCH TARGET 5 — DynamoDB as CMDB for change requests

Verdict: A DynamoDB acdl-change-requests table is consistent with the existing platform Lambda + DynamoDB pattern (D-051). The validate_change_request Lambda action queries the table + asserts status=approved.

Findings (verified 2026-07-22):

  1. Table schema: PK changeRequestId (string), SK submittedAt (string). Attributes: consumerRepo, contractId, status (enum: requested|approved|rejected|executed), requestedBy, approvedBy, submittedAt, executedAt.

  2. Validation flow: The decommission pipeline's validate-change-request stage invokes the Lambda with action: validate_change_request, changeRequestId: <id>, consumerRepo: <repo>. The Lambda queries the table; if the item exists + status == "approved" + consumerRepo matches, returns 200 with the CR details. Otherwise returns 403.

  3. Terraform: Add the table to terraform/platform/main.tf with SSE via the platform CMK + point-in-time recovery (matching the acdl-contracts table pattern from D-051).

RESEARCH TARGET 6 — Module engineering standards (scan of current modules)

Verdict: The current modules follow a consistent pattern that can be codified into standards. Key patterns identified:

  1. L1 required files: interface.json, instance.json, README.md, examples/simple.yaml, examples/complex.yaml. Multi-resource L1s add resources[] + intra_refs[] to interface.json.

  2. L2 required files: composition.json, README.md, examples/simple.yaml, examples/complex.yaml. No instance.json.

  3. Interface shape: name, version, kind ("l1"|"l2"), type (L1 only, aws:<service>:<kind>), description, inputs (object keyed by name), outputs (object keyed by name), nfrs (object keyed by name). Multi-resource L1s add resources[] (array of {type, description, inputs[], outputs[]}) + intra_refs[] (array of {from, to}).

  4. Input shape: {type, description, required, [default], [enum]}. Output shape: {type, description}. NFR shape: {type, description, default}.

  5. NFR conventions (v1.8 additions): Every L1 MUST have deletion_protection (boolean, default true) + encryption_enabled (boolean, default true) NFRs. L2 modules MUST expose features.deletion_protection (default true) + features.uptime_enabled (default true).

  6. Registry: Every module MUST be registered in modules/registry.json at its semver. Entry: {"interface": "<path>", "published_at": "<iso>", "deprecated": false}.

  7. Adapter extension: 3-table pattern (TYPE_MAP + INPUT_MAP + OUTPUT_MAP) + specialized _emit_resource branches for complex resources (nested blocks like origin {}, rules {}, default_cache_behavior {}).

  8. README structure: # <name> — <description>, ## Resources, ## Inputs, ## Outputs, ## NFRs, ## Usage, ## Compliance extension points, ## Examples, ## Versioning.

  9. Catalog index gap: modules/README.md Primitives table is missing rds (flagged during scan). Must be fixed in Phase 35.

Decisions surfaced (v1.8)

ID Decision Rationale Confidence Alternatives
D-073 uptime-kuma v1 (louislam/uptime-kuma:1) as the default container image. v1 is stable + widely deployed. v2 (2.4.0) is newer but has breaking changes. v1 is the safer default; consumers can override via container_image input. 0.85 v2 (breaking changes risk); pin to a specific v1 tag (maintenance burden).
D-074 Monitor seeding via post-deploy API script (scripts/seed_uptime_monitors.py), not env vars. uptime-kuma does not natively read env for monitor config. A post-deploy script calling the API is cleaner + testable offline. 0.90 Env var config (not supported by uptime-kuma); manual config (defeats automation).
D-075 KMS rotation = enable_key_rotation = true (AWS-managed annual). 90-day rotation is a roadmap item (custom rotation Lambda). AWS does not support custom rotation periods for managed keys. Annual is the AWS best practice. 90-day requires a custom Lambda + CloudWatch Events rule — overkill for v1.8. 0.80 Custom rotation Lambda (complex, overkill); no rotation (violates requirement).
D-076 uptime stack = separate VPC + ALB + ECS cluster (not shared with consumer stack). Requirement says "separate terraform run, with a separate state". Sharing the consumer's VPC/ALB would couple the states. Separate infra is cleaner + isolates the uptime stack's lifecycle. 0.85 Share consumer's VPC/ALB (state coupling); use App Runner (new service type).
D-077 EFS volume for uptime-kuma /app/data (persistent storage across task restarts). Fargate + EFS is the standard pattern for stateful containers. NFS is not supported by uptime-kuma, but EFS is NFS-compatible + works with Fargate. 0.90 S3-backed (uptime-kuma doesn't support S3); no persistent storage (data lost on restart).

v1.9 Research Addendum (Phase 0, 2026-07-23)

Milestone v1.9. Researcher: lead-developer. Autonomy: full. The v1.9 scope is well-grounded in the existing codebase; the research is a focused addendum covering the four new implementation domains (interpolation, per-env workflow inputs, Wiz GraphQL, attestation matrix freshness validation) + the design-doc drift audit.

RA-1 — Contract interpolation prior art + syntax choice (D-081)

Finding: Variable expansion in declarative manifests is a solved pattern. Terraform uses ${var.x} / ${local.x}; Helm uses {{ .Values.x }}; GitHub Actions uses ${{ }}; CloudFormation uses !Ref / !Sub. The contract schema is YAML validated by jsonschema — the schema does not inspect string contents, so any token syntax is schema-safe.

Choice: ${env.<field>} + ${contract.<field>} (D-081). Rationale:

  • Shell-style ${...} is the most familiar to the platform's audience (DevOps engineers comfortable with Terraform/HCL).
  • Dotted paths (${env.state_backend.bucket}) mirror Python attribute access and the existing wire["from"] syntax (contract.inputs.x, <childId>.outputs.y).
  • No conflict with YAML (${} inside a YAML string is a literal until the resolver expands it) or with jsonschema (string content is not schema-constrained).
  • Jinja {{ }} was considered (supports future filters) but rejected — the contract is a data file, not a template; filters would invite logic-in-config anti-patterns.

Implementation shape: a single _expand_vars(value, context) recursive walker in core/contract_resolver.py. Context = {"env": <loaded env json>, "contract": <contract dict>}. Unknown token → ValueError with the token text (fail loud, no silent passthrough — consistent with the P1-3 SSM fail-loud precedent).

Confidence: 0.92. Risk: none — the expansion is post-schema-validation and pre-IR-resolution, so it cannot break the schema or the adapter.

RA-2 — GitHub Actions workflow_call environment input + per-env jobs (D-082)

Finding: GitHub Actions workflow_call inputs support type: string with no enum constraint at the workflow-call layer (enum constraints exist only for choice-typed workflow_dispatch inputs). The deploy workflow already uses workflow_call with contract + mode + changeRequestId string inputs. Adding an environment string input (default empty, validated by run_platform.sh) is a one-line addition.

Per-env job pattern: the consumer repo's caller workflow (.github/workflows/deploy-<env>.yml) does:

jobs:
  deploy-qa:
    uses: acdl/.github/workflows/deploy.yml@v1.9
    with:
      environment: qa
      contract: .acdl/static-assets.qa.yaml

One caller workflow per environment = one CI job per environment. The environment: field in the contract is not edited for promotion; promotion = running the qa caller. The hybrid model (D-082) also lets a single contract be promoted via the environment input alone.

Gitea caveat: Gitea Actions supports workflow_call (reuses the GitHub Actions workflow YAML). The environment input works identically. Gitea has no Environments API (D-013/D-042) — the HITL gate is the workflow_dispatch approval-input fallback (already documented in hitl_matrix_design.md). For workflow_call (reusable), the caller workflow's workflow_dispatch trigger carries the approval input.

Confidence: 0.90. Risk: the Gitea workflow_call + approval-input combination needs the caller to be workflow_dispatch-triggered (not workflow_call-triggered) for the gate to fire — documented in Phase 41.

RA-3 — Wiz GraphQL API shape (D-0xx, REQ-110)

Finding: Wiz exposes a GraphQL API at <WIZ_API_URL>/graphql. Auth = Authorization: Bearer <WIZ_API_TOKEN>. The primary query for issues:

query IssuesQuery($filterBy: IssueFilter) {
  issues(filterBy: $filterBy) {
    nodes { id severity title entity { name type } control { name }
            createdAt }
    pageInfo { hasNextPage endCursor }
  }
}

Wiz severity enum: CRITICAL | HIGH | MEDIUM | LOW | INFORMATIONAL. Mapping to PolicyCheckResult:

  • engine: "wiz"
  • ruleId: <control.name> (or WIZ_<issue.id> fallback)
  • severity: <wiz severity lowercased>
  • status: FAIL (Wiz issues are findings; pass = no issues returned)
  • message: <title>
  • resource: <entity.name>

Graceful degrade: when WIZ_API_TOKEN or WIZ_API_URL unset → emit the existing single SKIPPED WIZ_NOT_CONFIGURED record (no network call). Offline tests use a recorded JSON fixture (no live Wiz tenant).

Confidence: 0.80. Risk: Wiz API version drift — the query shape is stable as of Wiz API v2 (2026), but the fixture is the test's source of truth, not the live API.

RA-4 — Attestation matrix freshness validation (D-084, REQ-109)

Finding: The 8 concerns in hitl_matrix_design.md §10.4 have declared freshness windows (24h, 7d, 30d, 90d, 180d). Operator-supplied evidence (load test, DR drill, FinOps forecast, runbook) is uploaded as a signed blob. The matrix validates:

  1. Presence — the evidence artifact exists for the target env.
  2. Freshnessartifact.timestamp is within the declared window.
  3. Schema — the artifact matches a per-concern JSON schema (e.g. load-test artifact has p99_latency, throughput, pass_rate).
  4. Signature (when ACDL_ATTESTATION_SIGNING_KEY_ID set) — JWS detached signature verification against a platform KMS key. When unset (dev/CI), signature verification is skipped (offline-testable).

Offline-testable concerns (run for real, no operator input):

  • Contract NFRs (the platform's own contract validator).
  • Schema validity (jsonschema).
  • Policy pass (Checkov/Wiz/Kyverno PolicyCheckResult records).

Operator-supplied concerns (require uploaded artifact):

  • Functional correctness (e2e suite report).
  • Performance baseline (k6/Gatling report).
  • Security posture (Trivy/Snyk scan + Security signature).
  • Operational readiness (runbook/dashboard/oncall/alerts).
  • Incident response (Sev-1 drill record).
  • Capacity/cost (FinOps forecast).
  • Resilience (DR drill, chaos report, backup verification).
  • dr-region deploy (dr drill report).

Confidence: 0.88. Risk: the signature verification path is only exercised when a signing key is configured (dev/CI skips it); production deployment must set ACDL_ATTESTATION_SIGNING_KEY_ID.

RA-5 — Design doc drift audit (REQ-100, REQ-101)

core/hitl_matrix_design.md drift:

  • Status block says "v1.2 wires the gates" — stale (v1.9 wires them).
  • "Spike scope note" says "the spike is dev-only; HITL is not exercised" — stale (v1.9 exercises qa/prod/dr).
  • §10.4 matrix is presented as design-only — v1.9 implements the offline-testable subset (D-084).
  • D-042 approver-identity mechanics are still accurate (Gitea has no Environments API; gitea.actor is the approver of record).

core/audit_ledger_design.md drift:

  • "Spike scope (D-041)" says "Phases 08-10 implement" — stale (the outbox is shipped + production since v1.8).
  • "v1.2 build-out" (S3 Object Lock + JWS + worker + DLQ + checkpoints) never shipped; v1.9 defers it explicitly (D-083).
  • The outbox item shape is still accurate; the approver_qa/ approver_prod attributes are populated by v1.9's hitl_gates.attest.

Confidence: 0.95. Risk: none — doc-only.

RA-6 — P1-1 adapter defaults audit (D-085, REQ-102)

Hardcoded defaults found in adapters/terraform/adapter.py:

  • desired_count = 1 (ECS service, 2 occurrences: line 238, 481).
  • launch_type = "FARGATE" (ECS service, line 239, 482).
  • family = "app" (task def, line 254 — reads inputs.get("family", "app") so partially parameterized; the "app" default should move to the interface).
  • target_type = "ip" (ALB target group, line 274).
  • load_balancer_type = "application" (ALB, line 272).
  • Name = "acdl-microservice-rt" (route table, line 283) + Name = ... tags on VPC/IGW (lines 515, 542 name = "app").

Fix: add desired_count, launch_type, family, target_type, load_balancer_type, name (VPC/IGW/RT) to the corresponding L1 interface.json inputs with defaults. The adapter reads inputs.get("<name>", <default>) — but the resolver should populate the default from the interface so the adapter reads inputs["<name>"] with a fallback only for safety. Tests assert an override emits the overridden value.

Confidence: 0.90. Risk: low — the v1.1 S3 regression test must still pass (S3 has none of these inputs).

Decisions surfaced (v1.9)

ID Decision Rationale Confidence Alternatives
D-087 Interpolation expansion is recursive over dicts + lists + strings (not just top-level inputs). A nested input like env: { DATABASE_URL: "acdl-${env.environment}-db" } should expand too. 0.90 Top-level only (misses nested maps).
D-088 The environment workflow_call input overrides the contract's environment field before schema validation, so the schema sees the overridden value. Interpolation context depends on the resolved environment; override must happen pre-validation so ${env.environment} is consistent. 0.92 Override post-validation (inconsistent interpolation context).
D-089 Attestation artifact signature verification is skipped when ACDL_ATTESTATION_SIGNING_KEY_ID is unset (dev/CI); required for prod/dr. Offline tests cannot sign with a real KMS key. The skip is explicit + logged. 0.85 Always require signature (breaks offline tests).

End of RESEARCH.md v1.9 addendum.