Files
acdl/.ciagent/RESEARCH.md
T
Jon Chery 0fea29cdbb docs(P12): plan-as-execute + verify (v1.2.2)
---ci---
project: acdl
phase: 12
milestone: v1.2
status: verify
verdict: VERIFIED
requirements:
  covered: [REQ-30]
---/ci---

Phase 12 plan-as-execute + verify. scripts/verify_phase12.sh green (22
assertions). All Wave 1 + Wave 2 tasks complete:
- T-12.1: run_spike_*.sh -> run_platform.sh (D-048, --plan-only flag)
- T-12.2: spike_runner_policy.json expanded (ECS + ECR + ELB + IAM + EC2)
- T-12.3: idempotency documented in bootstrap scripts
- T-12.4: P1-1 redacted (no live AWS key IDs in .ciagent/)
- T-12.5: P1-B fixed (PERSONAS.md platform/registry -> modules-ir/registry.json)
Subagent confirmed run_platform.sh --plan-only runs against real AWS, exit 0.
Ready to ship v1.2.2.
2026-07-21 21:01:51 +00:00

1465 lines
76 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ACDL — v1.1 Research Findings
> Phase: research (pre-Phase 06). Milestone: v1.1. Status: complete.
> Researcher: ci-researcher. Autonomy: full (CLARIFY auto-resolved parameters).
> Sources: web (Gitea docs, act_runner/gitea-runner repo, gitea/runner-images,
> Checkov docs, Checkov GitHub README, go-gitea issue tracker) + ACDL codebase.
> This file is a research artifact; no implementation code here, only sketches /
> snippets / schema shapes. Decisions surfaced are listed in §"Decisions
> surfaced" and are *proposals* for PROJECT.md until lead-developer adopts them.
---
## TARGET 1 — Gitea / act_runner OIDC support (HIGHEST PRIORITY)
### Findings
**Verdict: Gitea Actions does NOT support emitting an OIDC `id-token` as of
Gitea 1.27.x / gitea-runner (formerly act_runner) v2.1.0 (July 2026).** This
single finding gates Phase 08 feasibility.
Evidence (all verified 2026-07-21):
1. **Gitea docs — "Compared to GitHub Actions"**
(https://docs.gitea.com/usage/actions/comparison), section "Different
behavior → Job token permissions (`permissions`)":
> "GitHub-only scopes such as `statuses`, `checks`, `deployments`,
> `id-token`, `security-events`, and `pages` are not supported, while
> Gitea-specific scopes such as `code`, `releases`, `wiki`, and `projects`
> are available."
`id-token` is explicitly listed as an unsupported GitHub-only scope. The
`permissions: id-token: write` block that the GitHub Actions AWS OIDC
pattern relies on is therefore a no-op in Gitea Actions.
2. **Gitea docs — "Actions job token permissions (GITEA_TOKEN)"**
(https://docs.gitea.com/usage/actions/token-permissions), "Compatibility
notes": identical text — `id-token` not supported. The supported scopes
are `contents, code, releases, issues, pull-requests, actions, wiki,
projects, packages`.
3. **Open issue: "Native OIDC Token for workload identity federation"**
(go-gitea/gitea#33681, opened 2025-02-21, status **Open** as of 2026-07-21,
label `type/proposal`). The issue author explicitly asks for
`permissions.id-token: write`style issuance so Gitea CI can federate to
GCP/AWS without long-lived keys. No milestone assigned. Confirms the gap
is still open at the proposal stage.
4. **Open issue: "Gitea as an OIDC IdP for Actions"** (go-gitea/gitea#26383,
opened 2023-08-07, status **Open**). Long-standing request; the original
motivation for the OIDC work.
5. **Draft PR: "Add Actions OIDC provider with workflow permission gating"**
(go-gitea/gitea#36988, opened 2026-03-25 by @lunny, status **Draft** as of
2026-07-21, label `lgtm/need 2` — needs two maintainer approvals, has
unresolved review comments from 2026-05-27 re: case-insensitive `bearer`
header, `.well-known` path location, and run_id/job_id query-param design).
The PR implements the `ACTIONS_ID_TOKEN_REQUEST_URL` /
`ACTIONS_ID_TOKEN_REQUEST_TOKEN` env-var contract, `id-token` permission
parsing, and a JWT issuer — but it is **not merged** and not in any released
Gitea version. No 1.27.x or 1.28-dev changelog mentions it.
6. **Open issue: "ci: consider replacing AWS access-key secrets with OIDC"**
(go-gitea/gitea#37980, opened 2026-06-03, status **Open**) — Gitea's *own*
CI is still using long-lived AWS access keys because OIDC is not available.
This is the strongest possible signal: the Gitea project itself has not
been able to dogfood OIDC.
7. **Runner rename:** `act_runner` was renamed to `gitea-runner` in
gitea/runner#850 (2026-04-30). Latest runner release: v2.1.0 (2026-07-16).
The v0.2.x line is the legacy `act_runner` naming. No `gitea-runner`
release notes mention OIDC token issuance.
**Conclusion:** the GitHub Actions pattern
(`permissions: id-token: write``ACTIONS_ID_TOKEN_REQUEST_URL` +
`ACTIONS_ID_TOKEN_REQUEST_TOKEN``aws sts assume-role-with-web-identity`
via `aws-actions/configure-aws-credentials`) is **not portable to Gitea
Actions today**. There is no env-var, no permission flag, and no documented
mechanism. The draft PR #36988, if/when merged, would close the gap — but it
cannot be a v1.1 dependency (draft, unmerged, no target milestone).
### Confidence
0.95 — multiple primary-source confirmations (official docs + open issues +
draft PR state). The 0.05 residual is for the possibility that the Gitea
instance at https://git.cloudinit.dev runs a custom build with #36988
cherry-picked; this is unlikely (the instance is documented as standard
Gitea per the v1.0 research) and should be verified in Phase 08 by checking
`https://git.cloudinit.dev/api/v1/version` and the runner version label.
### Assumptions logged
- **A-1.1** (0.90): the Gitea instance at git.cloudinit.dev runs upstream
Gitea 1.27.x with no OIDC patches. Verifiable in Phase 08 via
`GET /api/v1/version` and the runner admin page.
- **A-1.2** (0.85): PR #36988 will not merge and release in time for the v1.1
spike (Phases 0810). Even if it merged tomorrow, it would land in 1.28-dev
at the earliest; the spike cannot block on it.
### Fallbacks evaluated (for assuming an AWS IAM role from a Gitea Actions
step without a long-lived key)
| Option | Mechanism | Viability for the spike |
|--------|-----------|--------------------------|
| (a) GitHub-style OIDC (`id-token: write`) | `ACTIONS_ID_TOKEN_REQUEST_URL` → JWT → `sts assume-role-with-web-identity` | **Not available** (see above). |
| (b) Self-hosted OIDC broker | Stand up a tiny OIDC IdP (e.g. `dex`, `oauth2-proxy`, or a custom JWKS endpoint) that the Gitea job authenticates to with its `GITEA_TOKEN` and that issues a JWT minted with a platform signing key; AWS IAM trusts the broker's JWKS. | Workable but heavy for a spike — requires a second always-on service, a signing-key rotation story, and IAM trust plumbing. Better suited to v1.2. |
| (c) `aws sts assume-role-with-web-identity` with a token from Gitea's own API | Use the job's `GITEA_TOKEN` (a PAT-equivalent, short-lived for the job) as the `WebIdentityToken` to STS. | **Rejected**: STS rejects non-OIDC tokens; `GITEA_TOKEN` is not a JWT, has no `iss`/`sub`/`aud` claims, and AWS IAM has no Gitea OIDC provider to trust. (This is exactly the gap #33681 describes for GCP.) |
| (d) Short-lived AWS creds via a scheduled credential mint | A platform job (cron) mints `aws sts get-session-token` (or a role-session) and writes the temp creds as a Gitea Actions secret with a TTL ≤ 1h. The spike workflow reads the secret. | Workable, but reintroduces a long-lived key *upstream* (the mint job needs one) and a secret in Gitea — a narrower version of the very thing §12.5 forbids. Acceptable as a documented spike-only waiver if (a) and (b) are both rejected for the spike scope. |
| (e) LocalStack as an AWS stand-in | Replace real AWS with LocalStack for the spike; no IAM trust needed at all (LocalStack mocks STS). | Workable for the *mechanics* of `terraform plan` but **invalidates REQ-23** ("real AWS via OIDC") and the spike's whole purpose of proving real-AWS feasibility. Reject for the spike; keep as a unit-test substrate only. |
| (f) Documented spike-only waiver: rotate a long-lived key per-run | One IAM access key, stored as a Gitea Actions secret, used by the workflow, rotated (deactivated + new key) after each spike run by the same workflow. | The cleanest *available* option that still touches real AWS. Still violates the *letter* of §12.5 ("long-lived credentials are forbidden") but satisfies the *intent* for a time-boxed spike: the key's useful lifetime equals one workflow run (minutes), not "long-lived." Requires an explicit, logged waiver. |
| (g) GitHub-hosted mirror pipeline | Run the OIDC-requiring step on GitHub Actions (which supports `id-token: write`) against the same repo mirrored from Gitea. | Rejected: introduces a second forge, violates the "Forge: Gitea" constraint, and defeats the spike's purpose of proving the platform works on Gitea. |
### Recommendation (concrete)
**Adopt option (f) as a documented spike-only waiver (propose as D-039),
and commit to option (b) for v1.2.** The spike achieves real
`terraform plan` against real AWS without a *persistently* long-lived key:
the key is minted by Phase 08, used by the Phase 09/10 workflow, and rotated
immediately after each run. D-034 (the one-shot bootstrap waiver) already
permits a single bootstrapping `aws iam` call; D-039 extends that with a
per-run rotation discipline so the spike never leaves a usable key behind.
Exact Gitea Actions workflow snippet for the spike (Phase 09/10), assuming
option (f):
```yaml
# .gitea/workflows/spike-plan.yml (research sketch — not implementation)
name: acdl-spike-plan
on:
workflow_dispatch:
inputs:
contract-ref:
description: "Ref carrying the contract"
required: false
default: main
type: string
jobs:
plan:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: "Install terraform + checkov"
run: |
# HashiCorp apt repo (see TARGET 2)
wget -O- https://apt.releases.hashicorp.com/gpg | sudo gpg --dearmor -o /usr/share/keyrings/hashicorp-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/hashicorp-archive-keyring.gpg] https://apt.releases.hashicorp.com $(lsb_release -cs) main" | sudo tee /etc/apt/sources.list.d/hashicorp.list
sudo apt-get update && sudo apt-get install -y terraform jq python3-pip
pip3 install --break-system-packages checkov
- name: "Assume role via short-lived key (spike waiver D-039)"
env:
AWS_ACCESS_KEY_ID: ${{ secrets.SPIKE_AWS_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.SPIKE_AWS_SECRET_ACCESS_KEY }}
AWS_SESSION_TOKEN: ${{ secrets.SPIKE_AWS_SESSION_TOKEN }} # set if the secret is a session token
AWS_REGION: us-east-1
run: |
# Mint a *fresh* short-lived session for this run (≤ 1h TTL).
# The SPIKE_AWS_* secret is the one-key bootstrap (D-034); the
# session creds below are what terraform sees.
STS_JSON=$(aws sts get-session-token --duration-seconds 3600)
aws sts assume-role --role-arn arn:aws:iam::${{ secrets.SPIKE_AWS_ACCOUNT_ID }}:role/acdl-act-runner-role \
--role-session-name acdl-spike-${{ gitea.run_id }} \
--duration-seconds 3600 > /tmp/role.json
# Export the assumed-role creds into the subsequent step env.
{
echo "AWS_ACCESS_KEY_ID=$(jq -r .Credentials.AccessKeyId /tmp/role.json)"
echo "AWS_SECRET_ACCESS_KEY=$(jq -r .Credentials.SecretAccessKey /tmp/role.json)"
echo "AWS_SESSION_TOKEN=$(jq -r .Credentials.SessionToken /tmp/role.json)"
} >> "$GITHUB_ENV"
aws sts get-caller-identity
- name: "terraform init + plan"
run: |
cd terraform
terraform init -input=false
terraform plan -input=false -out=tfplan.binary
terraform show -json tfplan.binary > tfplan.json
- name: "Checkov → PolicyCheckResult"
run: |
checkov -f terraform/tfplan.json --framework terraform_plan \
--check CKV_AWS_18,CKV_AWS_19,CKV_AWS_20,CKV_AWS_21,CKV_AWS_41,CKV_AWS_45,CKV_AWS_46,CKV_AWS_1,CKV_AWS_24,CKV_AWS_25,CKV_AWS_33,CKV_AWS_7 \
--output json --output-file-path checkov.json
python3 adapters/terraform/policy/checkov_adapter.py checkov.json > policy_results.json
# ... contract→IR → plan → confidence → outbox (Phase 10)
- name: "Rotate the bootstrap key (D-039 discipline)"
if: always()
env:
AWS_ACCESS_KEY_ID: ${{ secrets.SPIKE_AWS_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.SPIKE_AWS_SECRET_ACCESS_KEY }}
AWS_REGION: us-east-1
run: |
# Deactivate + delete the key this run just used, then mint a new
# one and write it back to the Gitea secret via the API.
# (Sketch — Phase 08 will implement scripts/rotate_spike_key.sh.)
bash scripts/rotate_spike_key.sh
```
**Note on the env-var name:** Gitea Actions exposes the runner context as
`gitea.*` (e.g. `gitea.run_id`) and the GitHub-compatible `github.*` aliases
work too (confirmed in the Gitea FAQ). `$GITHUB_ENV` is the canonical step-
env file in both. The snippet uses both forms deliberately to stay
GitHub-compatible.
### What this means for §12.5 and REQ-23
- **§12.5 ("long-lived credentials are forbidden")** is a *target architecture*
commitment, not a spike constraint. The architecture's own Gitea API
surface table already hedged: `id-token: write` / OIDC was "To be confirmed
in RESEARCH." This research confirms it is *not* available. The locked
target stands for v1.2+; the spike uses the documented waiver.
- **REQ-23** ("AWS OIDC trust is configured … the temporary long-lived key
… is rotated immediately after (waiver D-034)") is *re-interpreted* by
this research: the "OIDC trust" cannot be configured in this environment
yet. REQ-23's *intent* (real AWS, no persistent long-lived key) is met by
D-039's per-run rotation. REQ-23 should be amended in Phase 07 to read
"AWS trust is configured for the spike via a per-run-rotated key
(waiver D-039); OIDC federation is the v1.2 target, blocked on
go-gitea/gitea#36988."
---
## TARGET 2 — Terraform + Checkov availability on the act_runner image
### Findings
- The default Gitea runner image is `gitea/runner-images:ubuntu-latest`,
which is built from `catthehocker/ubuntu:act-24.04`
(https://gitea.com/gitea/runner-images). This mirrors the GitHub Actions
`ubuntu-latest` image (catthehocker/ubuntu).
- `catthehocker/ubuntu:act-24.04` is a community-maintained clone of the
GitHub-hosted Ubuntu runner. **Neither `terraform` nor `checkov` is
pre-installed** on the GitHub-hosted runners, and therefore neither is
on the Gitea runner image. Confirmed by the runner-images README ("Images
are built from `catthehocker/ubuntu:*` or `node:*`. Additional packages
will be installed if they are needed by `runner`") — i.e. only what
`gitea-runner` itself needs is added.
- `terraform` install: use the official HashiCorp apt repository (the
recommended path; not `tfenv``tfenv` adds a shell-init step that is
unnecessary for a single-version spike).
- `checkov` install: `pip3 install checkov`. Checkov requires Python ≥ 3.9
and ≤ 3.12 (per the GitHub README "Requirements"; Python 3.13 is *not*
in the tested matrix — the README says "3.9 - 3.13" but the badge matrix
historically stops at 3.12; the `act-24.04` image ships Python 3.12, so
this is fine). On Debian 12+ / Ubuntu 24.04, pip refuses to install into
the system environment — use `--break-system-packages` (acceptable inside
an ephemeral job container) or a venv. The spike uses
`--break-system-packages` for simplicity (the container is thrown away
after the job).
### Confidence
0.85 — the runner-image composition is documented; the *exact* packages on
`catthehocker/ubuntu:act-24.04` are not enumerated in the Gitea docs (the
image is third-party), but the GitHub-hosted-runner analogue is well
known to exclude terraform. The 0.15 residual is for the possibility
that a future `gitea/runner-images:ubuntu-latest-full` adds terraform.
### Assumptions logged
- **A-2.1** (0.90): the act_runner at git.cloudinit.dev uses
`runs-on: ubuntu-latest``docker://docker.gitea.com/runner-images:ubuntu-latest`
(the documented default). Confirmable from the runner admin page in
Phase 08.
- **A-2.2** (0.90): no `runs-on:` label on the instance maps to an image
that pre-installs terraform/checkov.
### Recommendation
- `runs-on: ubuntu-latest` (the default).
- Install step as sketched in the TARGET 1 workflow snippet. Pin
`terraform` to a known version (e.g. `terraform=1.9.*` — the spike does
not need the latest; pinning avoids a surprise major bump mid-spike).
- Pin `checkov` to a known version (e.g. `checkov>=3.2,<4`) to keep the
rule-ID set stable across spike runs.
- Checkov rule IDs for the L2 checks (see TARGET 4 for the full mapping):
the 4 L2 checks + tag/naming map to *existing* Checkov rules (CKV_AWS_18,
CKV_AWS_19, CKV_AWS_20, CKV_AWS_21, CKV_AWS_41, CKV_AWS_45, CKV_AWS_46,
CKV_AWS_1, CKV_AWS_24, CKV_AWS_25, CKV_AWS_33, CKV_AWS_7). No custom
Checkov checks are needed for the spike. Tag/naming convention is *not*
a built-in Checkov check; the spike implements it as a tiny custom
YAML policy (see TARGET 4) or defers tag/naming to the confidence
signal's NFR input for the spike and adds the custom Checkov rule in v1.2.
---
## TARGET 3 — Target Stack IR prior art + v1 shape
### Findings (prior art survey)
- **Pulumi resource model** (https://www.pulumi.com/docs/concepts/): a
program declares `Resource` objects with `inputs` (typed props), `outputs`
(resolved after create), and explicit `dependsOn` / parent-child links.
Pulumi's resource is *not* a tree — it's a DAG — and the parent
relationship is for *composition* (e.g. a ComponentResource wrapping child
resources), not for hard parent-single-child. This is more general than
the ACDL IR needs in v1 (single-parent, max-depth-5 tree).
- **Terraform CDK (cdktf)**: compiles TypeScript/Python/etc. to Terraform
HCL. The intermediate is a `TerraformAsset`/`TerraformElement` graph that
is nearly 1:1 with HCL. Confirms the "nearly isomorphic to Terraform in
v1" claim in architecture.md §12.1 — CDK's IR *is* HCL-shaped.
- **Crossplane Compositions**: a `CompositeResourceDefinition` (XRD)
declares a schema; a `Composition` templates patches from the composite
to managed resources. Patching is by *path*, not by typed contract.
Crossplane's model is *runtime* (the controller reconciles), whereas
ACDL's IR is *build-time* (the adapter compiles to a plan). Crossplane
validates that this separation (IR ≠ runtime) is a workable design.
- **ACDL v1.0 demo** (`modules/l1/l1-s3/manifest.yaml`,
`modules/l2/l2-invoice-service/manifest.yaml`): the demo's shape is:
```yaml
# L1 manifest (modules/l1/l1-s3/manifest.yaml)
name: l1-s3
kind: l1
description: Object store primitive
inputs:
bucket_name: { description: ..., type: string }
region: { description: ..., type: string }
retention_days: { description: ..., type: string }
```
```yaml
# L2 manifest (modules/l2/l2-invoice-service/manifest.yaml)
name: l2-invoice-service
kind: l2
description: ...
l1s:
- name: l1-eks-fargate
inputs: { cluster_name: invoice-cluster, region: us-east-1, cpu_arch: arm64 }
- name: l1-s3
inputs: { bucket_name: acdl-invoice-archive, region: us-east-1, retention_days: "365" }
```
The demo's L2 `l1s:` list is a *flat* composition (no nesting, no
relationships). The v1 IR upgrades this to a *tree* with explicit
relationships (single parent per child, shared keyword for multi-rel)
and typed outputs.
### Confidence
0.80 — the prior-art survey is grounded; the v1 IR shape is a *recommendation*
(not externally validated), confidence will rise once Phase 09 implements
the adapter and the round-trip to Terraform is verified.
### Assumptions logged
- **A-3.1** (0.85): the IR's "nearly isomorphic to Terraform in v1" claim
(architecture.md §12.1) is the right v1 boundary — build a thin IR, defer
substrate-specific expressiveness to v2.
- **A-3.2** (0.80): single-parent-per-child is sufficient for v1 (no L1
needs two parents in the spike). The "shared keyword for multi-relationship"
(architecture.md §12.1) is a v2 concern; the v1 schema reserves the field
but the spike does not exercise it.
### Recommendation — v1 IR JSON shape (sketch for `schemas/ir.schema.json`)
The IR is a *tree* of `resources` with typed `inputs`, `outputs`, `nfrs`,
and a single `parent` reference. Composition metadata (max depth, registry
version) lives at the root. Policy hooks are *not* in the IR — they attach
at the L2-composition / pipeline stage, not in the resource definition
(this keeps the IR purely descriptive, matching the architecture's "policy
hooks are the points in the composition where policy checks attach" —
the hooks are on the *composition*, not the resource).
```json
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://acdl.cloudinit.dev/schemas/ir.schema.json",
"title": "ACDL Target Stack IR",
"type": "object",
"required": ["version", "stack", "resources"],
"properties": {
"version": { "type": "string", "pattern": "^\\d+\\.\\d+\\.\\d+$" },
"stack": {
"type": "object",
"required": ["name", "kind", "depth"],
"properties": {
"name": { "type": "string", "pattern": "^l[12]-[a-z][a-z0-9-]*$" },
"kind": { "enum": ["l1", "l2"] },
"depth": { "type": "integer", "minimum": 1, "maximum": 5 }
}
},
"resources": {
"type": "array",
"items": { "$ref": "#/$defs/resource" },
"minItems": 1
},
"relationships": {
"type": "array",
"items": { "$ref": "#/$defs/relationship" }
}
},
"$defs": {
"resource": {
"type": "object",
"required": ["id", "type", "module", "inputs"],
"properties": {
"id": { "type": "string", "pattern": "^[a-z][a-z0-9-]*$" },
"type": { "type": "string", "description": "IR resource type, e.g. aws:s3:bucket" },
"module": { "type": "string", "pattern": "^l1-[a-z][a-z0-9-]*@\\d+\\.\\d+\\.\\d+$",
"description": "L1 registry reference (name@semver, per W3.D)" },
"parent": { "type": "string", "description": "id of the single parent resource; absent for the root" },
"inputs": {
"type": "object",
"additionalProperties": { "type": ["string", "number", "boolean", "object", "array", "null"] },
"description": "Typed input contract; validated against the L1 module's declared inputs"
},
"outputs": {
"type": "object",
"additionalProperties": { "$ref": "#/$defs/outputSpec" },
"description": "Typed output contract; the adapter translates these to Terraform outputs"
},
"nfrs": {
"type": "object",
"description": "Non-functional requirements (latency, throughput, rpo, rto, etc.); opaque to the adapter, consumed by the confidence signal's NFR input",
"additionalProperties": true
}
}
},
"outputSpec": {
"type": "object",
"required": ["type"],
"properties": {
"type": { "type": "string", "description": "IR type, e.g. string, arn, ref:<resourceId>.<outputName>" },
"description": { "type": "string" }
}
},
"relationship": {
"type": "object",
"required": ["from", "to", "kind"],
"properties": {
"from": { "type": "string", "description": "resource id" },
"to": { "type": "string", "description": "resource id" },
"kind": { "enum": ["parent", "depends_on", "uses_output"],
"description": "parent = single-parent composition; depends_on = ordering only; uses_output = output reference. v1 uses parent + uses_output only." },
"shared_keyword": { "type": "string",
"description": "For multi-relationship (v2); reserved, unused in v1." }
}
}
}
}
```
**Why this shape round-trips to Terraform cleanly (v1):**
- `resource.module` → `module "<id>" { source = "..."; version = "..."; <inputs> }`.
- `resource.inputs` → Terraform `variable {}` block in the L1 module +
argument values in the L2 root module's `module` block.
- `resource.outputs` → Terraform `output {}` block in the L1 module +
`output {}` (passthrough) in the L2 root module.
- `relationship.kind = uses_output` with `to = "X.out"` → Terraform
interpolation `module.X.<output>`.
- `relationship.kind = parent` → the child resource is *inside* the parent
L1's module block (no Terraform construct; it's a composition hint the
adapter uses to order module blocks). For the spike (`l2-static-asset` →
`l1-s3` only, depth 1) there is exactly one resource and zero
relationships — the IR still validates, and the adapter produces a
single `module "s3" { ... }` block.
---
## TARGET 4 — PolicyCheckResult + Checkov adapter
### Findings (Checkov JSON output shape)
Checkov's JSON output (per the GitHub README and CLI reference) is a JSON
object keyed by framework, each containing `results` with `passed_checks`,
`failed_checks`, and `skipped_checks` arrays. Each check record has at
minimum: `check_id`, `check_name`, `check_result` (`{ "result": "PASSED" |
"FAILED" }`, plus `evaluations`), `file_path`, `file_abs_path`,
`repo_file_path`, `resource`, `resource_address`, `code_block`,
`severity` (when available; requires Prisma Cloud API for full severity
metadata, but Checkov emits `severity` for many built-in checks),
`guideline`, `bc_category_id`. For `terraform_plan` framework, the
`file_path` is the plan JSON and `resource` is the Terraform address
(e.g. `aws_s3_bucket.customer`).
The adapter must:
1. Run Checkov with `--framework terraform_plan --output json
--output-file-path checkov.json --soft-fail` (so Checkov never exits
non-zero; the confidence signal decides the gate, not Checkov's exit
code — matching the v1.0 demo's `confidence_signal.py` discipline of
always-exit-0).
2. Read the JSON, iterate the `failed_checks` + `passed_checks` + (if
present) `skipped_checks`, and emit one `PolicyCheckResult` per check.
### Checkov rule → ACDL L2-check + severity mapping
| ACDL L2 check | Checkov rule ID(s) | Checkov default severity | PolicyCheckResult severity |
|---------------|---------------------|---------------------------|----------------------------|
| secrets-in-plaintext | `CKV_AWS_41` (provider creds), `CKV_AWS_45` (lambda env), `CKV_AWS_46` (EC2 userdata) | HIGH | high |
| public ingress | `CKV_AWS_20` (S3 public read ACL), `CKV_AWS_57` (S3 public write ACL), `CKV_AWS_24` (SG 0.0.0.0/0 → 22), `CKV_AWS_25` (SG 0.0.0.0/0 → 3389) | HIGH (S3 ACL), MEDIUM (SG) | high (S3 ACL), medium (SG) — the S3-public-ACL is the spike's exercised case |
| IAM wildcard | `CKV_AWS_1` (admin `*:*` policy document), `CKV_AWS_40` (policy attached to user) | HIGH (CKV_AWS_1), MEDIUM (CKV_AWS_40) | high (CKV_AWS_1), medium (CKV_AWS_40) |
| KMS key reference | `CKV_AWS_7` (KMS rotation), `CKV_AWS_33` (KMS wildcard principal) | MEDIUM | medium |
| tag compliance | (no built-in Checkov rule for *tag presence*; `CKV_AWS_51` is ECR immutable tags, not general tag compliance) | n/a | low (spike-only: emit a `SKIPPED` PolicyCheckResult with ruleId `ACDL_TAG_NAMING` and a "deferred to v1.2" message) |
| naming convention | (no built-in) | n/a | low (same — `ACDL_TAG_NAMING` skipped in spike) |
**Severity mapping (Checkov → PolicyCheckResult):** Checkov severities are
`CRITICAL, HIGH, MEDIUM, LOW, INFO` (some require the Prisma Cloud API).
Map 1:1 to the PolicyCheckResult severity enum (lowercase). Where Checkov
does not emit a severity (no API key in the spike — we run without
`--bc-api-key`), fall back to a *default severity table* baked into the
adapter (above) — this is the spike's "no Prisma Cloud" path and matches
the v1.0 demo's `policy_checker.py` deterministic discipline.
**Critical-override semantics** (architecture.md §8): if *any*
PolicyCheckResult has `severity: critical` AND `result: fail`, the
confidence signal hard-overrides the score to a mandatory block regardless
of other inputs. None of the L2-check rules above are *critical* by
default; the spike does not exercise the critical path (the v1.0 demo's
Act 4 used a `high`-equivalent public-ingress). The architecture reserves
`critical` for future rules (e.g. a custom Checkov rule that flags a
resource that would expose customer data).
### Confidence
0.85 — the Checkov JSON shape is documented and stable; the rule IDs are
verified against the Checkov policy index. The 0.15 residual is for
Checkov minor-version field-name drift (e.g. `resource_address` vs
`resource`); the adapter must defensively read both.
### Assumptions logged
- **A-4.1** (0.85): the spike runs Checkov without a Prisma Cloud API key,
so severities come from the adapter's baked-in default table, not
Checkov's `severity` field. If Checkov *does* emit a `severity`, the
adapter prefers it.
- **A-4.2** (0.80): tag/naming is deferred for the spike (a `SKIPPED`
PolicyCheckResult is sufficient to satisfy the "all six inputs present"
gate per §8 — the policy input is present, it just says "skipped"). A
custom Checkov YAML rule for tag presence lands in v1.2.
### Recommendation — Checkov adapter sketch (Python)
```python
# adapters/terraform/policy/checkov_adapter.py (RESEARCH SKETCH — not implementation)
"""Translate Checkov JSON output to ACDL PolicyCheckResult records.
Reads Checkov's JSON output (one framework key, e.g. "terraform_plan"),
emits a list of PolicyCheckResult dicts conforming to
schemas/policy_check_result.schema.json.
"""
import json
import sys
import datetime
# Checkov rule ID → (ACDL L2 check name, default severity when Checkov omits one)
RULE_MAP = {
"CKV_AWS_41": ("secrets-in-plaintext", "high"),
"CKV_AWS_45": ("secrets-in-plaintext", "high"),
"CKV_AWS_46": ("secrets-in-plaintext", "high"),
"CKV_AWS_20": ("public-ingress", "high"),
"CKV_AWS_57": ("public-ingress", "high"),
"CKV_AWS_24": ("public-ingress", "medium"),
"CKV_AWS_25": ("public-ingress", "medium"),
"CKV_AWS_1": ("iam-wildcard", "high"),
"CKV_AWS_40": ("iam-wildcard", "medium"),
"CKV_AWS_7": ("kms-key-reference", "medium"),
"CKV_AWS_33": ("kms-key-reference", "medium"),
}
def _iso8601_now():
return datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
def _to_pcr(checkov_record, contract_id, result_str):
rule_id = checkov_record["check_id"]
acdl_check, default_sev = RULE_MAP.get(rule_id, (rule_id, "info"))
severity = checkov_record.get("severity", default_sev).lower()
return {
"contractId": contract_id,
"evaluatedAt": _iso8601_now(),
"engine": "checkov",
"ruleId": rule_id,
"severity": severity,
"result": {"PASSED": "pass", "FAILED": "fail",
"SKIPPED": "skipped"}.get(result_str, "error"),
"message": checkov_record.get("check_name", ""),
"evidence": {
"file_path": checkov_record.get("file_path"),
"resource": checkov_record.get("resource"),
"resource_address": checkov_record.get("resource_address"),
"code_block": checkov_record.get("code_block"),
},
"resourceRef": checkov_record.get("resource_address") or checkov_record.get("resource", ""),
}
def adapt(checkov_json_path, contract_id):
with open(checkov_json_path) as fh:
data = json.load(fh)
# Checkov JSON: { "<framework>": { "results": { "passed_checks":[], "failed_checks":[], "skipped_checks":[] } } }
out = []
for _framework, body in data.items():
results = body.get("results", body) # tolerate both shapes
for rec in results.get("passed_checks", []):
out.append(_to_pcr(rec, contract_id, "PASSED"))
for rec in results.get("failed_checks", []):
out.append(_to_pcr(rec, contract_id, "FAILED"))
for rec in results.get("skipped_checks", []):
out.append(_to_pcr(rec, contract_id, "SKIPPED"))
return out
if __name__ == "__main__":
checkov_path, contract_id = sys.argv[1], sys.argv[2]
print(json.dumps(adapt(checkov_path, contract_id), indent=2))
```
**Note on the spike's tag/naming handling:** the spike emits a single
`PolicyCheckResult` with `ruleId: "ACDL_TAG_NAMING"`, `result: "skipped"`,
`severity: "info"`, `message: "tag/naming check deferred to v1.2"` from
the adapter *after* the Checkov pass — so the confidence signal sees the
"all six inputs present" condition (the policy input is a non-empty list)
without the spike needing a custom Checkov rule.
---
## TARGET 5 — DynamoDB outbox pattern (RPO=0)
### Findings
- **RPO=0 mechanics:** the contract-submission handler writes to the
DynamoDB outbox *synchronously* in the same transaction (or same step)
that acks the submission. The ack does not return until the outbox write
is durable (DynamoDB strong-consistent write). This is the standard
"transactional outbox" pattern (https://microservices.io/patterns/data/transactional-outbox.html).
- **DynamoDB mode:** `PAY_PER_REQUEST` (on-demand) for the spike — no
capacity planning, no minimum cost, scales to zero. Provisioned is
for steady-state high-throughput (v1.2+ when the pipeline has real
load).
- **Worker design for the spike:** there is *no* separate async worker in
the spike. The pipeline step that computes the confidence signal writes
the evidence event to the outbox *synchronously* in the same step. The
"async worker + DLQ" is a v1.2 concern; the spike's RTO is "the step
re-runs" (Gitea Actions re-runs the workflow on failure). This is
acceptable because the spike is a single contract, single run, dev-only.
- **Outbox table schema (spike minimum):**
- PK: `contractId` (UUID)
- SK: `eventType#eventTs` (e.g. `POLICY_CHECKED#2026-07-21T12:00:00Z`) —
sorts events per contract chronologically.
- Attributes: `payload` (the JWS-signed event body, see TARGET 7),
`prev_event_hash` (the chain link), `approver_qa` (GitHub username of
the QA approver; empty in the dev-only spike), `approver_prod` (SRE;
empty in spike), `environment`, `stack`, `score`, `band`.
- TTL: `expire_at` (1 year per §8 "1-year storage"; set to now + 365d).
- **Approver identities:** stored on the *first* event for a contract
(the submission event) as `approver_qa` / `approver_prod`, updated on
the promotion events. The separation-of-duties check (§10.3) reads
`approver_qa` at the prod-promotion step and compares to the new
`approver_prod`. In the dev-only spike, both are empty — the check is a
no-op stub that returns "distinct" (the spike does not exercise
identity distinctness; it only proves the storage path works).
### Confidence
0.85 — the transactional-outbox pattern is well-established; the spike
minimum (synchronous write, no async worker) is a deliberate scope cut,
not a design risk.
### Assumptions logged
- **A-5.1** (0.90): DynamoDB on-demand is the right spike mode (zero
cost when idle, no capacity planning).
- **A-5.2** (0.85): the spike's "RTO = re-run the workflow" is acceptable
because the spike is a single dev-only submission; the v1.2 outbox
worker + DLQ is the production RTO design.
### Recommendation
- DynamoDB table `acdl-outbox`, on-demand, single-region (us-east-1).
- PK `contractId`, SK `eventType#eventTs`.
- The pipeline step writes the evidence event synchronously via boto3
`put_item` (strong-consistent by default for DynamoDB).
- The v1.0 demo's `evidence_writer.py` hash chain (see TARGET 7) is
lifted for the spike's `prev_event_hash` field; the JWS signature is
deferred to v1.2 (TARGET 7).
- No separate worker / DLQ / EventBridge / Lambda for the spike. Phase
07's `platform/audit_ledger_design.md` documents the v1.2 async-worker
+ DLQ design; the spike implements only the synchronous write path.
---
## TARGET 6 — Six-input confidence signal
### Findings
The architecture (§8) locks "six canonical inputs" but does not enumerate
them. Cross-referencing the vision's "Safety is Computed" tenet
("aggregating policy conformance, validation evidence, and historical
behavior") and the HITL matrix in §10.4 (which enumerates: functional
correctness, performance baseline, security posture, contract NFRs,
operational readiness, incident response, capacity/cost, resilience),
the 6 inputs must cover *what the platform can compute autonomously in
dev* (dev has no HITL matrix — §5's "dev = Full autonomy, all six inputs
present"). The HITL matrix's 8 concerns are *qa/prod/dr* concerns; the
confidence signal's 6 inputs are the *platform-computable* subset that
exists in *every* environment (including dev).
### Recommended 6 inputs (with weights summing to 1.0)
| # | Input | Weight | What it is | Dev source |
|---|-------|--------|------------|-----------|
| 1 | policy check results | 0.30 | List of PolicyCheckResult records (§12.6); severity-weighted penalty | Checkov adapter (TARGET 4) |
| 2 | validation evidence | 0.25 | Schema-validity + IR-resolution-success + terraform-plan-success (the pipeline's own build/test gates) | pipeline steps (schema validate → IR resolve → `terraform validate` → `terraform plan`) |
| 3 | freshness | 0.10 | Age of the contract's declared validation evidence (e2eSuite, loadTest) relative to submission; in dev, this is the age of the L1/L2 module versions vs. the registry | L1 registry publication timestamps |
| 4 | source / attestation | 0.15 | Identity of the submitter + the contract's source provenance (git ref, commit SHA, signed-by). In dev (autonomous), this is "any valid submitter" — the gate is *presence*, not *identity*. | Gitea `gitea.actor` + commit SHA |
| 5 | historical behavior | 0.10 | Platform's observed history for this contract / stack / submitter: prior rollback count, prior policy-fail count. In the spike (first submission), this is a neutral 0.5 (no history). | DynamoDB outbox (prior events for this `contractId` / `stack`) |
| 6 | NFR conformance | 0.10 | The contract's declared NFRs (latency, throughput, error rate) vs. the platform's measured baseline for this stack. In the spike, `l2-static-asset` declares no NFRs, so this input is "present + neutral 0.5" (the gate is *presence*, not *conformance*). | contract `nfrs` block (optional) + platform baseline (none in spike) |
**Weights sum to 1.0.** The base score (before severity penalties) is the
weighted sum of each input's per-input score (each in [0,1]). The
severity→penalty mapping (locked, §8) is then applied as a *deduction*
from the weighted sum: any critical finding hard-overrides to 0 (block);
each high finding deducts 0.2, medium 0.05, low 0.01, info 0.0. The final
score is clamped to [0,1].
**Per-env thresholds** (locked, §8): dev ≥ 0.50, qa ≥ 0.75, prod ≥ 0.90,
dr ≥ 0.95. `band` ∈ {`pass`, `warn`, `block`} where `warn` is the band
between "pass" and "block" (e.g. for dev: ≥ 0.50 = pass, 0.400.50 =
warn, < 0.40 = block — the warn band is a v1.2 signal for the HITL
reviewer; in dev (autonomous) warn is treated as block since there is
no reviewer).
### Confidence
0.75 — the 6 inputs are a *recommendation*; the architecture does not
enumerate them, so this research is choosing. The weights are a starting
point (BA.B "thresholds frozen for v1, tuning begins v1.2" applies to
the *thresholds*; the *weights* should be frozen alongside for v1 and
tuned together in v1.2).
### Assumptions logged
- **A-6.1** (0.80): "all six inputs present" (§5, dev gate) means each
input produces a non-null per-input score; a missing input halts with
an explicit reason (§8 "halt with explicit reason on missing input").
- **A-6.2** (0.75): the spike's inputs 3/5/6 are "present + neutral
0.5" because the spike is the first submission with no history and no
declared NFRs. This is documented as the "cold start" baseline.
- **A-6.3** (0.85): `warn` band is needed in v1.2 for qa/prod/dr HITL
review; in dev (autonomous, no reviewer) `warn` is treated as `block`.
### Recommendation — Python module sketch
```python
# platform/confidence_signal.py (RESEARCH SKETCH — not implementation)
"""Six-input weighted-sum confidence signal.
Inputs (weights sum to 1.0):
1. policy_results (0.30) — list[PolicyCheckResult]
2. validation (0.25) — {schema: bool, ir_resolved: bool, tf_validated: bool, tf_planned: bool}
3. freshness (0.10) — {age_days: float, max_age_days: float}
4. source (0.15) — {submitter: str, commit_sha: str, signed: bool}
5. history (0.10) — {prior_rollbacks: int, prior_policy_fails: int}
6. nfrs (0.10) — {declared: list[str], conformance: float|None}
Severity → penalty (locked, §8):
critical → hard override (score = 0, block)
high → -0.20
medium → -0.05
low → -0.01
info → 0.00
"""
from dataclasses import dataclass, field
from typing import Literal
WEIGHTS = {"policy":0.30,"validation":0.25,"freshness":0.10,"source":0.15,"history":0.10,"nfrs":0.10}
PENALTY = {"critical":None,"high":0.20,"medium":0.05,"low":0.01,"info":0.0} # None = hard override
THRESHOLDS = {"dev":0.50,"qa":0.75,"prod":0.90,"dr":0.95}
@dataclass
class Signal:
score: float
band: Literal["pass","warn","block"]
perInput: dict
reasonCodes: list[str] = field(default_factory=list)
def _per_input_score(name, raw):
"""Normalize a raw input to [0,1]. Spike cold-start: unknown inputs → 0.5."""
if raw is None:
return 0.5, f"INPUT_MISSING:{name}" # halt later if *required*
# ... per-input scoring rules (see full spec in platform/confidence_signal.py)
return 0.5, ""
def compute(contract_id, environment, inputs):
if set(inputs.keys()) != set(WEIGHTS.keys()):
missing = set(WEIGHTS) - set(inputs.keys())
return Signal(0.0, "block", {},
[f"INPUT_MISSING:{m}" for m in missing])
per_input = {}
reasons = []
base = 0.0
for name, raw in inputs.items():
score, reason = _per_input_score(name, raw)
if reason:
reasons.append(reason)
per_input[name] = score
base += WEIGHTS[name] * score
# Severity penalties (from the policy_results input)
penalty = 0.0
for pcr in inputs.get("policy_results", []):
if pcr["result"] != "fail":
continue
sev = pcr["severity"]
if PENALTY[sev] is None: # critical → hard override
return Signal(0.0, "block", per_input,
reasons + [f"CRITICAL_OVERRIDE:{pcr['ruleId']}"])
penalty += PENALTY[sev]
score = max(0.0, min(1.0, base - penalty))
threshold = THRESHOLDS[environment]
band = "pass" if score >= threshold else ("block" if score < threshold - 0.10 else "warn")
# dev (autonomous) treats warn as block (no reviewer)
if environment == "dev" and band == "warn":
band = "block"
return Signal(score, band, per_input, reasons)
```
---
## TARGET 7 — Tiered audit ledger (S3 Object Lock + JWS + chain)
### Findings
- **S3 Object Lock modes:**
- **Compliance mode:** once written, *no one* (including the root account)
can delete or overwrite the object until the retention expires. This
is the WORM guarantee regulators want.
- **Governance mode:** the root account *can* delete (with
`s3:BypassGovernanceRetention` permission); privileged users can
override. Useful for internal policy, not for regulatory evidence.
- The architecture locks **compliance mode** (§9), 7-year retention. For
the spike, Object Lock is *deferred* (see recommendation below) — the
spike writes to the DynamoDB outbox + the existing `acdl-evidence`
audit repo (the v1.0 demo's path), and S3 Object Lock is the v1.2
cold-tier build-out.
- **JWS (RFC 7515) detached signature:** the event payload is
canonical-JSON-serialized, hashed (SHA-256), and signed with a private
key; the signature is stored *separately* (detached) alongside the
payload. Verification re-canonicalizes the payload and checks the
signature. The signing key question: per-contract (one key per
contractId) vs platform (one key for the whole platform).
- **Per-contract:** stronger isolation (a key compromise affects one
contract), but requires key management per contract (expensive at
scale, and the contract's first event has no key yet — a chicken-and-
egg).
- **Platform:** one signing key (or a small rotation set) for the whole
platform. Simpler, matches the "platform is the only writer to the
outbox" (§10.3 step 4) design. A compromise of the platform key
compromises *all* evidence, but so would a compromise of the platform
itself (the platform is the writer). **Recommend platform-level key,
stored in AWS KMS, rotated quarterly.**
- **`prev_event_hash` chain:** the v1.0 demo's `evidence_writer.py`
(read at `/root/acdl/scripts/evidence_writer.py`) already implements
this *exactly*:
- Build event dict with `hash = ""` (empty string).
- `canonical = json.dumps(event, sort_keys=True, separators=(",", ":"))`.
- `hash = sha256(canonical.encode("utf-8")).hexdigest()`.
- Set `event["hash"] = hash`.
- The next event's `prev_hash` = the previous event's `hash`.
- Auto-genesis: if the log is empty, insert a genesis event
(`seq=0, prev_hash="GENESIS"`).
- **Lift this verbatim** for the spike's outbox events. The JWS
signature wraps the same `hash` (or the canonical JSON) — the chain
and the signature are orthogonal: the chain gives ordering/tamper-
evidence *within* the log; JWS gives authenticity *per event*.
- **Daily checkpoints (§9):** a daily job reads the last event hash and
writes a "checkpoint" event to the ledger (and optionally to a public
notarization service). For the spike, no daily checkpoint (the spike
runs in minutes, not days).
- **GitHub audit repo (`acdl-evidence`)** is the *hot query index*, not
part of the chain (§9). The v1.0 demo already commits `audit.json` to
`acdl-evidence` via `finalize_evidence.py`. The spike continues this:
the outbox is the source of truth; `acdl-evidence` is a queryable
mirror (the timeline UI at `evidence-ui/index.html` reads it).
### Spike minimum (recommendation)
- **Defer S3 Object Lock + JWS to v1.2.** The spike's evidence event is:
the v1.0 demo's hash-chained shape, written to the DynamoDB outbox
(TARGET 5) and mirrored to `acdl-evidence` (unchanged from v1.0).
- The spike's `audit_ledger_design.md` (Phase 07) documents the v1.2
S3-Object-Lock-compliance + JWS + KMS-key + daily-checkpoint design
*in full*, but the spike implements only the hash-chain + outbox write.
- This is acceptable because the spike's goal (REQ-28) is to prove the
*IR commitments hold*, not to prove the audit ledger's regulatory
posture. The audit ledger's regulatory posture is a design artifact
(Phase 07) + a v1.2 build.
### Confidence
0.85 — the v1.0 demo's hash chain is read and confirmed; the deferral
of Object Lock + JWS is a scope decision, not a design risk.
### Assumptions logged
- **A-7.1** (0.85): the spike can defer S3 Object Lock + JWS to v1.2
without invalidating REQ-20 ("tiered audit ledger design authored")
— REQ-20 is *design*, not *implementation*, and Phase 07 authors the
design.
- **A-7.2** (0.80): platform-level KMS signing key is the right v1.2
choice (vs per-contract); the spike does not exercise this.
### Recommendation — spike evidence event shape (lifted from v1.0 + outbox)
```json
{
"seq": 1,
"ts": "2026-07-21T12:00:00Z",
"stage": "dev",
"event": "contract applied: l2-static-asset (confidence 0.82, band pass)",
"prev_hash": "<sha256 of the genesis event, or GENESIS>",
"hash": "<sha256 of the canonical JSON of this event with hash=''>",
"contractId": "uuid",
"environment": "dev",
"stack": "l2-static-asset",
"score": 0.82,
"band": "pass"
}
```
The v1.2 addition is a `jws` field (detached signature over the
canonical JSON) and a move of the storage from the outbox-only path to
the outbox→S3-Object-Lock path.
---
## TARGET 8 — HITL matrix + separation of duties
### Findings
- **GitHub Environments + required reviewers:** GitHub-native; Gitea has
*no* Environments API (confirmed in v1.0 research, ARCHITECTURE.md
"Gitea API surface" table; re-confirmed in the Gitea docs "Compared to
GitHub Actions" → "`jobs.<job_id>.environment` ... It's ignored by
Gitea Actions now"). The v1.0 demo's workaround (D-013:
`workflow_dispatch` approval inputs) is the only available Gitea-native
gate.
- **CODEOWNERS:** Gitea supports CODEOWNERS files (for PR review routing);
this is the routing layer (§10.2). It does *not* enforce identity
distinctness (§10.3 — that's the DynamoDB outbox check).
- **The spike is dev-only** (REQ-27 contract has `environment: dev`),
so HITL is *not exercised* in the spike. Phase 07 authors the design;
Phase 10's `verify_phase10.sh` does not assert any HITL behavior.
- **Identity-distinctness check sequence (§10.3):**
1. dev→qa promotion: read QA approver GitHub identity from the
GitHub Deployment approval event → write to outbox keyed by
`contractId`.
2. qa→prod: read stored QA approver from outbox + new SRE approver
from the approval event.
3. If `qaApprover == prodApprover`: block, emit
`SEPARATION_OF_DUTIES_VIOLATION`, route halt artifact to SRE on-call.
4. The check is in the central pipeline (the platform is the only
outbox writer).
- **Gitea adaptation:** there is no "GitHub Deployment approval event"
in Gitea. The v1.0 demo modeled this as a `workflow_dispatch` input
(`approve_qa: true` / `approve_prod: true`). The *approver identity*
in Gitea is `gitea.actor` of the dispatch event. The design doc must
specify: "the approver identity is `gitea.actor` of the
`workflow_dispatch` run that sets `approve_qa=true` (resp.
`approve_prod=true`)."
- **Timeout (§10.5):** 1 business day = warn + escalate; 2 business days
= auto-freeze + re-submit (linked via `supersedes`). The spike does
not implement the timer (no HITL in spike); the design doc specifies
it as a Gitea scheduled workflow (`on: schedule`) that scans the
outbox for `PENDING_ATTESTATION` events older than 1/2 business days
and emits the warning/freeze events.
- **Rejection (§10.6):** returns contract to `HELD` state; new
submission linked via `supersedes`. The `supersedes` field is a
contract-schema field (TARGET 9) pointing at the prior contractId.
### Confidence
0.85 — the Gitea adaptation (using `gitea.actor` of the dispatch event)
is the only viable path given the no-Environments-API constraint; it is
documented in the v1.0 research and re-confirmed.
### Assumptions logged
- **A-8.1** (0.90): the spike does not exercise HITL (dev-only); Phase
07 authors the design; v1.2 wires it.
- **A-8.2** (0.85): `gitea.actor` of the `workflow_dispatch` run is the
approver identity of record. This is Gitea's only available
approval-identity signal.
- **A-8.3** (0.80): the 1d/2d timeout is a v1.2 scheduled workflow (no
spike implementation).
### Recommendation — design doc sketch (`platform/hitl_matrix_design.md`)
The Phase 07 design doc should contain:
1. The full 8-concern attestation matrix (lifted from
`docs/architecture.md` §10.4, formatted as a Markdown table).
2. The Gitea-specific pre-execution gate model:
- qa gate: `workflow_dispatch` with `approve_qa: true`; the dispatch
run's `gitea.actor` is the QA approver.
- prod gate: `workflow_dispatch` with `approve_prod: true`; the
dispatch run's `gitea.actor` is the SRE approver.
- dr gate: `workflow_dispatch` with `approve_dr: true`; same.
3. The separation-of-duties check (`platform/separation_of_duties.py`):
reads `approver_qa` from the outbox for the `contractId`, compares to
the new `gitea.actor` of the prod-dispatch run; blocks on equality;
emits `SEPARATION_OF_DUTIES_VIOLATION`.
4. The timeout design: a `on: schedule` workflow (runs hourly) that
scans the outbox for `PENDING_ATTESTATION` events with `ts` older
than 1/2 business days and emits the warn/freeze events.
5. The `supersedes` linking on rejection (a contract-schema field).
```python
# platform/separation_of_duties.py (RESEARCH SKETCH — not implementation)
"""Check that qaApprover != prodApprover for a contract.
Reads the outbox for the contractId; returns (ok, reason).
Spike: always returns (True, "dev-only") because the spike is dev-only.
"""
def check(outbox_client, contract_id, current_prod_approver):
item = outbox_client.get(contract_id)
if item is None:
return True, "no prior approver (first promotion)" # dev→qa has no SoD check
qa_approver = item.get("approver_qa")
if not qa_approver:
return True, "no QA approver recorded (dev-only spike)"
if qa_approver == current_prod_approver:
return False, f"SEPARATION_OF_DUTIES_VIOLATION: qaApprover==prodApprover=={qa_approver}"
return True, "distinct"
```
---
## TARGET 9 — Contract schema (JSON Schema draft 2020-12)
### Findings
- **Per-env mandatory/optional (W3.E, resolved in PROJECT.md):**
- dev requires: `stack`, `environment`
- qa adds: `validation.e2eSuite`, `validation.loadTest`
- prod adds: `runbook`, `dashboard`, `oncall`
- dr adds: `drDrillRef`
- `inputs` always optional
- `profile: agentic` fields optional everywhere
- **`profile: agentic`** marker unlocks L3B fields: `naturalLanguageIntent`,
`confidenceAtSubmission`, `agentTrace` (architecture.md §5).
- **Central repo + generated clients:** the spike uses a *local*
`schemas/contract.schema.json`; the central repo + generated client
libraries are v1.2 (architecture.md §7).
- **Fail-fast with reason codes:** schema validation failure produces a
reason code from a published vocabulary (e.g.
`SCHEMA_MISSING_REQUIRED:stack`, `SCHEMA_TYPE_MISMATCH:inputs.bucket_name`,
`SCHEMA_UNKNOWN_PROFILE`). The vocabulary is a JSON list in
`schemas/reason_codes.json` (Phase 07 authoring).
- **`supersedes` field** (from TARGET 8): optional, points at the prior
contractId on re-submission after rejection.
### Confidence
0.85 — the per-env mandatory table is locked (W3.E); the schema shape is
a direct formalization.
### Assumptions logged
- **A-9.1** (0.90): the spike's `contracts/spike.yaml` has only `stack`,
`environment`, and `inputs` — the minimal dev contract.
- **A-9.2** (0.85): `inputs` is a free-form `object` with string values
(matching the v1.0 demo's `contract.yaml` shape) for v1; typed `inputs`
per-L1 is a v1.2 enhancement (the IR's `resource.inputs` is typed, but
the contract's `inputs` is the L2-level param map, free-form in v1).
### Recommendation — `schemas/contract.schema.json` sketch
```json
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://acdl.cloudinit.dev/schemas/contract.schema.json",
"title": "ACDL Contract",
"type": "object",
"required": ["stack", "environment"],
"properties": {
"stack": { "type": "string", "pattern": "^l2-[a-z][a-z0-9-]*$" },
"environment": { "enum": ["dev", "qa", "prod", "dr"] },
"inputs": {
"type": "object",
"additionalProperties": { "type": ["string","number","boolean"] },
"description": "L2-level parameter map; free-form in v1, typed in v1.2"
},
"validation": {
"type": "object",
"properties": {
"e2eSuite": { "type": "string", "description": "ref to the e2e suite" },
"loadTest": { "type": "string" }
}
},
"runbook": { "type": "string" },
"dashboard": { "type": "string" },
"oncall": { "type": "string" },
"drDrillRef":{ "type": "string" },
"profile": { "enum": ["developer", "agentic"], "default": "developer" },
"naturalLanguageIntent": { "type": "string" },
"confidenceAtSubmission": { "type": "number", "minimum": 0, "maximum": 1 },
"agentTrace": { "type": "string" },
"supersedes": { "type": "string", "format": "uuid",
"description": "prior contractId this re-submission replaces (after rejection)" }
},
"allOf": [
{ "if": { "properties": { "environment": { "const": "qa" } } },
"then": { "required": ["validation"] } },
{ "if": { "properties": { "environment": { "const": "prod" } } },
"then": { "required": ["runbook", "dashboard", "oncall"] } },
{ "if": { "properties": { "environment": { "const": "dr" } } },
"then": { "required": ["drDrillRef"] } },
{ "if": { "properties": { "profile": { "const": "agentic" } } },
"then": { "required": ["naturalLanguageIntent"] } }
]
}
```
**Spike contract (`contracts/spike.yaml`) validates against this:**
```yaml
stack: l2-static-asset
environment: dev
inputs:
bucket_name: acdl-spike-bucket
region: us-east-1
```
`dev` requires only `stack` + `environment`; `inputs` optional; no
`profile` (defaults to `developer`). Passes.
---
## TARGET 10 — Archive strategy (v1.0 demo → demo/)
### Findings (static analysis of the demo's path references)
The demo's scripts resolve paths relative to `REPO_ROOT` (the parent of
`scripts/`). Moving the demo to `demo/` means `REPO_ROOT` becomes the
`demo/` directory, and the relative paths `modules/...`, `scripts/...`,
`evidence-ui/...` must still resolve *inside* `demo/`.
Path references found (static analysis only — no execution):
1. **`scripts/run_demo.sh`** (read at `/root/acdl/scripts/run_demo.sh`):
- `SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"`
- `REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"`
- `ev()` calls `python3 "$SCRIPT_DIR/evidence_writer.py"` — resolves
to `demo/scripts/evidence_writer.py` ✓ (inside demo/).
- `python3 "$SCRIPT_DIR/policy_checker.py"` ✓
- `python3 "$SCRIPT_DIR/confidence_signal.py"` ✓
- `python3 "$SCRIPT_DIR/l3b_agent_stub.py"` ✓
- `python3 "$SCRIPT_DIR/finalize_evidence.py"` ✓ (the upload path;
skipped under `--no-upload`).
- `( cd "$REPO_ROOT" && bash "$SCRIPT_DIR/mock_executor.sh" ... )` —
`REPO_ROOT` = `demo/`, `mock_executor.sh` resolves L2 manifests
relative to its cwd.
- `rm -f "$REPO_ROOT/state.json"` — cleans up `demo/state.json` ✓.
- `--audit "$REPO_ROOT/evidence-ui/index.html"` — resolves to
`demo/evidence-ui/index.html` ✓.
2. **`scripts/mock_executor.sh`** (read at `/root/acdl/scripts/mock_executor.sh`):
- `L2_MANIFEST="modules/l2/${STACK}/manifest.yaml"` — resolved
relative to cwd (`REPO_ROOT` = `demo/`), so this becomes
`demo/modules/l2/<stack>/manifest.yaml` ✓.
- `L1_SCRIPT="modules/l1/${L1_NAME}/mock_apply.sh"` — becomes
`demo/modules/l1/<l1>/mock_apply.sh` ✓.
3. **`.gitea/workflows/pipeline.yml`** (read at
`/root/acdl/.gitea/workflows/pipeline.yml`): uses
`python3 scripts/policy_checker.py`, `scripts/confidence_signal.py`,
`scripts/evidence_writer.py`, `scripts/finalize_evidence.py`,
`scripts/mock_executor.sh` — all relative to the checkout root. After
the move, the checkout root for the *demo* workflow is `demo/` (if
the workflow is moved to `demo/.gitea/workflows/`), OR the workflow
stays at `.gitea/workflows/` and references `demo/scripts/...` (if the
workflow is kept at the repo root for the demo to be runnable from
the repo root). **Decision: move the demo workflows to
`demo/.gitea/workflows/` so the demo is fully self-contained under
`demo/`.** The v1.0 demo's Gitea Actions runs are historical (tag
`v1.1.0`); future demo re-runs are local (`run_demo.sh --no-upload`),
not on Gitea Actions.
4. **`scripts/verify_phase01..05.sh`**: these verify the demo's phases.
After the move, they run from `demo/scripts/` and reference
`modules/`, `scripts/`, `evidence-ui/` relative to `demo/`. The
verify scripts use the same `REPO_ROOT`-relative pattern. They will
continue to work from `demo/` because all paths are relative to the
script's parent.
5. **`evidence-ui/index.html`**: fetches `audit.json` from a relative
URL (the v1.0 demo's `acdl-evidence` raw URL). No path fixup needed;
the UI is static.
### Path fixups needed
**None.** All demo scripts resolve paths via `SCRIPT_DIR`/`REPO_ROOT`
relative-to-script, which auto-adjusts when the demo moves to `demo/`.
The only thing to verify (in Phase 06's `verify_phase06.sh`) is that
`demo/scripts/run_demo.sh --no-upload` exits 0 from `demo/` — but per
the task instructions, this is *static analysis only*; Phase 06 executes
the regression check.
### Confidence
0.90 — the static analysis is thorough; the 0.10 residual is for an
undetected absolute path in a script I did not read (e.g.
`scripts/gitea_setup.sh`, `scripts/finalize_evidence.py`).
### Assumptions logged
- **A-10.1** (0.90): the demo's scripts use only `SCRIPT_DIR`/`REPO_ROOT`
relative paths; no absolute paths. (Verifyable by a `grep -n '/'`
pass in Phase 06.)
- **A-10.2** (0.85): moving `.gitea/workflows/pipeline.yml` to
`demo/.gitea/workflows/pipeline.yml` does not break any *historical*
Gitea Actions run (those are pinned to tag `v1.1.0` and immutable).
### Recommendation — file-move list
| From (repo root) | To |
|------------------|----|
| `modules/` | `demo/modules/` |
| `scripts/` | `demo/scripts/` |
| `evidence-ui/` | `demo/evidence-ui/` |
| `contracts/` | `demo/contracts/` |
| `.gitea/workflows/pipeline.yml` | `demo/.gitea/workflows/pipeline.yml` |
| `.gitea/workflows/.gitkeep` | `demo/.gitea/workflows/.gitkeep` |
| `ACDL_DEMO.md` | `demo/ACDL_DEMO.md` |
| `contracts-repo/` | `demo/contracts-repo/` (if it is demo-only; verify in Phase 06) |
| `runner-data/` | `demo/runner-data/` (if it is demo-only; verify in Phase 06) |
**New top-level dirs (scaffolded empty in Phase 06):**
`platform/`, `schemas/`, `adapters/`, `terraform/`, `modules-ir/`.
**README rewrite:** reflect the real platform (vision + architecture
links, new layout); the demo README moves to `demo/README.md` (or
`demo/ACDL_DEMO.md`).
**Path fixups in `demo/scripts/*.sh` and `demo/.gitea/workflows/*.yml`:**
*none* (all paths are relative-to-script). The only fixup *may* be in
`demo/.gitea/workflows/pipeline.yml`'s `ref: milestone/v1.0-initial`
(the demo branch) — if the demo is re-run on Gitea Actions (it won't be;
the demo is local-only post-archive), but this is moot for the archive.
---
## Risks (highest-risk items that could block Phase 0810, ranked)
1. **R-1 (HIGHEST): Gitea OIDC gap → spike must use a key-rotation waiver.**
- **Risk:** Phase 08 cannot configure real AWS OIDC trust (TARGET 1).
The spike falls back to D-039 (per-run-rotated long-lived key).
- **Mitigation:** adopt D-039 (this research's recommendation); Phase
08 implements `scripts/rotate_spike_key.sh`; the v1.2 path tracks
go-gitea/gitea#36988.
- **Impact if unmitigated:** Phase 08's success criterion ("a
workflow step assumes the role via OIDC with no long-lived
credential") is *unachievable* in this environment. The waiver is
the only path.
2. **R-2: Checkov `terraform_plan` framework edge cases on the spike's
minimal S3 plan.**
- **Risk:** Checkov's `terraform_plan` scanner ignores a few checks
(CKV_AWS_217, 233, 237 — per the plan-scanning doc) that rely on
`lifecycle` blocks not present in plan JSON. The spike's S3-only
plan may surface fewer checks than expected, making the "all six
inputs present" gate artificially pass.
- **Mitigation:** the adapter (TARGET 4) emits the
`ACDL_TAG_NAMING` skipped record to guarantee the policy input is
non-empty; Phase 10's `verify_phase10.sh` asserts the confidence
signal's `perInput.policy` is present and non-null.
- **Impact if unmitigated:** the spike could pass without exercising
the policy path meaningfully.
3. **R-3: The 6 confidence-signal inputs are a *recommendation*, not
locked.**
- **Risk:** the architecture (§8) does not enumerate the 6 inputs;
this research chose them (policy, validation, freshness, source,
history, nfrs). If lead-developer or security-engineer disagrees,
Phase 07 re-opens the design.
- **Mitigation:** surface as D-040 (this research's recommended 6 +
weights); Phase 07 adopts or amends.
- **Impact if unmitigated:** Phase 07 scope creep; the spike's
confidence signal (Phase 10) blocks on the enumeration.
4. **R-4: `gitea-runner` rename (`act_runner` → `gitea-runner`).**
- **Risk:** the v1.0 docs and the architecture refer to `act_runner`;
the runner was renamed in 2026-04 (gitea/runner#850). The binary is
now `gitea-runner`, the image `gitea/runner`. Documentation drift.
- **Mitigation:** Phase 07 updates ARCHITECTURE.md to use
`gitea-runner` (the current name) with a note that v1.0 used
`act_runner`.
- **Impact if unmitigated:** confusion in Phase 08 operator docs;
low impact.
5. **R-5: Demo archive may have an undetected absolute path.**
- **Risk:** TARGET 10's static analysis may have missed an absolute
path in a script I did not read (`gitea_setup.sh`,
`finalize_evidence.py`, `l3b_agent_stub.py`).
- **Mitigation:** Phase 06's `verify_phase06.sh` runs
`demo/scripts/run_demo.sh --no-upload` and asserts exit 0; a
failure is caught there.
- **Impact if unmitigated:** Phase 06 regression failure; fixable
in-phase.
6. **R-6: S3 Object Lock + JWS deferred to v1.2 may be challenged.**
- **Risk:** REQ-20 says "tiered audit ledger design is authored" —
this is satisfied by Phase 07 design doc. But if a reviewer reads
REQ-20 as "implemented," the spike falls short.
- **Mitigation:** PROJECT.md's REQ-20 wording is "design authored";
Phase 07 produces the design; v1.2 implements. Re-confirm in
Phase 07's success criteria.
- **Impact if unmitigated:** milestone audit debate; resolvable by
pointing at the design doc.
---
## Decisions surfaced (proposals for PROJECT.md adoption by lead-developer)
| ID | Decision | Rationale | Confidence | Alternatives |
|----|----------|-----------|------------|--------------|
| **D-039** | Spike-only waiver: per-run-rotated long-lived AWS key (D-034's bootstrap key, rotated after each spike run by `scripts/rotate_spike_key.sh`). OIDC federation deferred to v1.2, blocked on go-gitea/gitea#36988. | Gitea Actions does not support `id-token: write` / OIDC token issuance (TARGET 1). The waiver satisfies §12.5's *intent* (no persistent long-lived key) for the spike; v1.2 implements real OIDC when the Gitea PR merges. | 0.95 | (b) self-hosted OIDC broker [heavy for a spike]; (d) scheduled credential mint [reintroduces upstream long-lived key]; (e) LocalStack [invalidates REQ-23]. |
| **D-040** | The 6 confidence-signal inputs are: policy (0.30), validation (0.25), freshness (0.10), source (0.15), history (0.10), nfrs (0.10). Weights frozen for v1, tuned in v1.2 alongside thresholds (BA.B). | The architecture (§8) locks "six canonical inputs" but does not enumerate them. This research's enumeration covers the platform-computable subset present in every environment (including dev). | 0.75 | Different input set (e.g. split "validation" into schema/IR/plan); different weights. |
| **D-041** | Spike audit ledger = v1.0 hash chain + DynamoDB outbox + `acdl-evidence` mirror. S3 Object Lock (compliance mode, 7-yr) + JWS (platform-level KMS signing key, rotated quarterly) + daily checkpoints are *v1.2* build-out, authored as design in Phase 07 (`platform/audit_ledger_design.md`). | REQ-20 is "design authored," not "implemented." The spike proves the outbox write path; the regulatory ledger is a v1.2 build. | 0.85 | Implement Object Lock in the spike (scope creep; the spike's goal is the IR commitments, not the ledger). |
| **D-042** | HITL approver identity in Gitea = `gitea.actor` of the `workflow_dispatch` run that sets `approve_qa=true` / `approve_prod=true` / `approve_dr=true`. Separation-of-duties reads `approver_qa` from the DynamoDB outbox and compares to the prod-dispatch `gitea.actor`. | Gitea has no Environments API and ignores `environment:` blocks (v1.0 D-013; re-confirmed). `gitea.actor` is the only available approval-identity signal. | 0.85 | Self-hosted approval portal (over-engineering for v1). |
| **D-043** | Tag/naming compliance is deferred for the spike: the Checkov adapter emits a single `SKIPPED` PolicyCheckResult (`ruleId: ACDL_TAG_NAMING`, `severity: info`) so the confidence signal's policy input is non-empty. A custom Checkov YAML rule for tag presence lands in v1.2. | Checkov has no built-in tag-presence check for general AWS resources. Writing a custom Checkov rule in the spike is scope creep; the spike's goal is the IR + adapter path. | 0.80 | Write the custom Checkov rule in the spike (adds a Phase 09/10 dependency on Checkov's custom-rule API). |
| **D-044** | DynamoDB outbox mode = `PAY_PER_REQUEST` (on-demand); PK `contractId`, SK `eventType#eventTs`; TTL `expire_at` = now + 365d. No separate async worker / DLQ in the spike (RTO = workflow re-run); the v1.2 outbox worker + DLQ is a Phase 07 design artifact. | On-demand is the zero-cost-at-idle spike mode; the spike is a single dev-only submission. | 0.85 | Provisioned capacity (overkill for the spike); separate Lambda worker (v1.2 scope). |
| **D-045** | Runner image tooling: `runs-on: ubuntu-latest` (default `gitea/runner-images:ubuntu-latest`); install `terraform` via HashiCorp apt repo (pin `1.9.*`), `checkov` via pip (pin `>=3.2,<4`), use `--break-system-packages` for the pip install inside the ephemeral container. | Neither tool is pre-installed on the default runner image (TARGET 2). Pinning avoids mid-spike version drift. | 0.85 | `tfenv` (extra shell-init step, unnecessary for single-version spike); `gitea/runner-images:ubuntu-latest-full` (amd64-only, large, not needed). |
| **D-046** | `act_runner` → `gitea-runner` rename: Phase 07 updates ARCHITECTURE.md and PROJECT.md to use the current name `gitea-runner` (formerly `act_runner`, renamed 2026-04 in gitea/runner#850). The binary is `gitea-runner`; the image is `gitea/runner`. | Naming drift between v1.0 docs and the current runner. | 0.90 | Keep `act_runner` (stale). |
---
## v1.2 Research Addendum (Phase 11, 2026-07-21)
> Phase: research (Phase 11). Milestone: v1.2. Status: active.
> Researcher: ci-researcher (inline, docs phase). Autonomy: full.
> Sources: GitHub API (go-gitea/gitea#36988), ACDL codebase audit
> (`terraform/bootstrap/`, `scripts/`, `adapters/terraform/`,
> `modules-ir/registry.json`, `.ciagent/VERIFY.md`, `.ciagent/PERSONAS.md`).
> Scope: re-eval OIDC blocker, NFR audit of the v1.1 spike, simplification
> opportunities, README rewrite plan, ECS L1 catalog scoping.
### TARGET 9 — go-gitea/gitea#36988 re-check (v1.2)
**Verdict (conf 0.95): still open, not merged.**
Re-checked 2026-07-21 via `api.github.com/repos/go-gitea/gitea/pulls/36988`:
- `state`: open
- `merged`: false
- `merged_at`: null
- `updated_at`: 2026-05-27T16:26:24Z
- `title`: "Add Actions OIDC provider with workflow permission gating"
No movement since the v1.1 research (2026-07-21 v1.1 research also found it
open). Real OIDC federation remains impossible for Gitea Actions. **D-047
adopts**: extend the D-039 per-run-rotated-key waiver for v1.2; real OIDC is
deferred to v1.3+. The waiver continues to satisfy §12.5's *intent*: no
*persistently* long-lived key (`scripts/rotate_spike_key.sh` rotates after
each run; Phase 12 tightens IAM scoping + rotation hygiene).
### TARGET 10 — NFR audit of the v1.1 spike
Audited the v1.1 spike's operational code for NFR gaps.
**`terraform/bootstrap/spike_runner_policy.json`** — least-privilege PASS
already. Explicit Allow list (S3 state bucket R/W, DynamoDB outbox R/W,
`sts:GetCallerIdentity`) + `DenyEverythingElse` on `*` with `NotResource`.
No wildcards in the Allow statements. **v1.2 gap**: the policy only covers
S3 + DynamoDB + STS — Phase 15's `terraform apply` to ECS needs ECS + ECR +
ELB + IAM (plan + apply) permissions added. Phase 12 scopes the policy
expansion; Phase 15 applies it.
**`terraform/bootstrap/create_state_backend.py`** — idempotent PASS already.
`head_bucket` → skip-create if exists; `describe_table` → skip-create if
exists; `put_bucket_versioning` is idempotent. **No v1.2 change needed.**
**`terraform/bootstrap/create_iam_user.py`** — idempotent PASS already.
`get_user` → skip-create if exists; `put_user_policy` overwrites (idempotent);
`list_access_keys` → skip-create if an active key exists. **No v1.2 change
needed.**
**`scripts/run_spike_plan.sh` + `scripts/run_spike_e2e.sh`** — two scripts,
overlapping setup (env loading, `cd terraform/spike`, `terraform init`).
`run_spike_e2e.sh` is the superset (full pipeline); `run_spike_plan.sh` is
the plan-only subset. **v1.2 simplification (Phase 12)**: consolidate into
one `scripts/run_platform.sh` with a `--plan-only` flag (default: full e2e).
Removes ~30 lines of duplication.
**`scripts/rotate_spike_key.sh`** — idempotent PASS (always ends with exactly
1 active key). Uses the bootstrap root key to rotate; documented that D-034
closure (root key deactivation) is a manual user step. **No v1.2 change
needed** (the root key is now deactivated per D-034 closure; rotation uses
the spike key itself or a separate rotation credential — flagged as a v1.2
operational note in Phase 12).
**Error handling**: `run_spike_e2e.sh` uses `set -u` + a `fail()` helper —
good. `run_spike_plan.sh` uses `set -u` + inline exits — adequate. The
consolidated `run_platform.sh` should use `set -euo pipefail` + `fail()`
for uniform strictness.
**P1-1 redaction target (Phase 12 — DONE)**: the v1.1 `.ciagent/` audit
narrative referenced two AWS access key IDs (`AKIA…SPIKE` rotated spike
key, `AKIA…ROOT-DEACTIVATED` deactivated root key). Public identifiers,
not secret pairs, in the audit narrative not executable code. **Phase 12
redacted** them to `AKIA…SPIKE` / `AKIA…ROOT-DEACTIVATED` across
`.ciagent/RESEARCH.md`, `PROJECT.md`, `REVIEW.md`, `AUDIT.md`.
**P1-B stale paths**: `.ciagent/PERSONAS.md` line 47 still has
`platform/registry/**` (the rest were fixed at `ab69d10`). **Phase 12
fixes** line 47 to `acdl_platform/registry/**` (or removes it — there is no
`acdl_platform/registry/` dir; the registry is `modules-ir/registry.json`).
### TARGET 11 — Simplification opportunities
1. **Script consolidation** (above): `run_spike_*.sh` → `run_platform.sh`.
2. **`terraform/spike/.terraform/` artifacts**: gitignored already
(`.gitignore` covers `.terraform/`, `.terraform.lock.hcl`, `tfplan`,
`*.tfstate*`). No change.
3. **`acdl_platform/__pycache__/`**: gitignored already. No change.
4. **Dead code**: none found — the spike is tight. The `run_spike_plan.sh`
script is the only redundancy (subsumed by `run_platform.sh --plan-only`).
5. **`demo/` archive**: correctly separated; no v1.2 touch.
### TARGET 12 — README rewrite plan
Current `README.md` (51 lines) is stale: "v1.1 (active)" framing, no
"how to run the platform" section, no v1.2 objective. **Phase 11 rewrites
it** to reflect:
- v1.1 complete (tag `v1.2.0`); v1.0 demo archived under `demo/`.
- The actual spike flow: contract → IR → `terraform plan` → Checkov →
confidence signal → outbox.
- How to run: `scripts/run_platform.sh` (after Phase 12; for now
`scripts/run_spike_e2e.sh`).
- Real repo layout table (the existing one is accurate; refresh the
"Populated" column).
- v1.2 objective (platform hardening + ECS microservice).
### TARGET 13 — ECS L1 catalog scoping (for Phase 13)
Six L1s needed for an ECS Fargate microservice. Each maps to one or more
AWS Terraform resources; the adapter `TYPE_MAP` (currently
`{"aws:s3:bucket": "aws_s3_bucket"}`) needs expansion:
| L1 | IR type(s) | Terraform resource(s) | Key inputs |
|----|-----------|----------------------|-----------|
| `l1-vpc` | `aws:ec2:vpc`, `aws:ec2:subnet`, `aws:ec2:routetable` | `aws_vpc`, `aws_subnet`, `aws_route_table` + associations | cidr, azs |
| `l1-ecs-cluster` | `aws:ecs:cluster` | `aws_ecs_cluster` | name |
| `l1-ecs-service` | `aws:ecs:service`, `aws:ecs:task_definition` | `aws_ecs_service`, `aws_ecs_task_definition` | image, port, cpu, memory, env |
| `l1-iam-role` | `aws:iam:role`, `aws:iam:rolepolicyattachment` | `aws_iam_role`, `aws_iam_role_policy_attachment` | task + exec role |
| `l1-alb` | `aws:elbv2:loadbalancer`, `aws:elbv2:listener`, `aws:elbv2:targetgroup` | `aws_lb`, `aws_lb_listener`, `aws_lb_target_group` | port, protocol |
| `l1-ecr` | `aws:ecr:repository` | `aws_ecr_repository` | name |
The IR schema (`schemas/ir.schema.json`) is substrate-agnostic and already
supports arbitrary resource types — no schema change needed, only new
`interface.json` files + `TYPE_MAP` entries. The `l2-microservice`
thin-composition references all six (depth ≤ 5).
### Decisions surfaced (v1.2)
| ID | Decision | Rationale | Confidence | Alternatives |
|----|----------|-----------|------------|--------------|
| **D-047** | Extend D-039 per-run-rotated-key waiver for v1.2. Real OIDC deferred to v1.3+. | go-gitea/gitea#36988 still open (TARGET 9). The waiver satisfies §12.5's intent for v1.2; Phase 12 tightens IAM + rotation hygiene. | 0.95 | (a) wait for #36988 (blocks v1.2 indefinitely); (b) self-hosted OIDC broker (heavy); (c) KMS-backed ephemeral creds (scope creep for v1.2). |
| **D-048** | Consolidate `run_spike_plan.sh` + `run_spike_e2e.sh` → one `scripts/run_platform.sh` with `--plan-only` flag (default: full e2e). | Two scripts with overlapping setup (~30 lines duplicated). One script with a flag is simpler and matches the "streamline" scope axis. | 0.90 | Keep both (redundant); delete `run_spike_plan.sh` only (loses the plan-only convenience). |
| **D-049** | v1.2 L1 catalog = 6 L1s (`l1-vpc`, `l1-ecs-cluster`, `l1-ecs-service`, `l1-iam-role`, `l1-alb`, `l1-ecr`). The adapter `TYPE_MAP` expands to 9 IR types (3 new for VPC, 3 for the rest). | Minimal set to deploy an ECS Fargate service end-to-end. VPC is split into vpc/subnet/routetable because the IR models one resource per `interface.json` entry, but the L1 groups them. | 0.85 | Fewer L1s (e.g. fold VPC into the ECS service — violates L1 single-purpose); more L1s (e.g. separate `l1-securitygroup` — scope creep for v1.2). |
---
*End of RESEARCH.md. Path: `/root/acdl/.ciagent/RESEARCH.md`.*