Files
acdl/.ciagent/GRILL.md
T
Jon Chery cf8aa53c8d docs(P00): grill — v1.26 adversarial review (9 challenges, PROCEED 0.84, 2 revisions)
---ci---
project: acdl
phase: 0
milestone: v1.26
status: grill
verdict: PROCEED
confidence: 0.84
revisions: [G-Q4 REQ-322 to P2 W0, G-Q6 enforcement-deferred note, G-Q9 key-split future item]
---/ci---
2026-08-12 21:20:11 +00:00

225 lines
12 KiB
Markdown

# GRILL — v1.26 Live Pilot Estate Activation
> Adversarial review of the v1.26 SPECIFY + CLARIFY + RESEARCH + IDEATE +
> PLAN. The grill red-teams the proposal across feasibility, scope,
> budget, and the domain claims (homegrown blockchain, pilot estate,
> metric grounding). Each challenge gets a binding verdict
> (PROCEED / REVISE / ESCALATE). Autonomy: full — escalations auto-
> resolve with assumption logging unless confidence < 0.60.
## Verdict: PROCEED (0.84) — 0 escalations, 2 revisions
The milestone is feasible, scoped, and the domain claims hold. Two
plan revisions are binding (G-Q4, G-Q8) and are already captured in
PLAN.md. No work is blocked.
---
## Challenges
### G-Q1 — Is a homegrown PoA blockchain viable for a pilot, or is it reckless?
**Challenge:** Authoring a blockchain (even a minimal PoA ledger) is a
non-trivial domain. A homegrown chain could have correctness bugs (hash
chain breaks, non-deterministic blocks, settlement-finality race
conditions). Why not use a proven chain (Ethereum L2, Solana, Hyperledger
Fabric)?
**Verdict:** PROCEED (confidence 0.88). The pilot's purpose is to
exercise the Nova platform's deploy/policy/attestation gates over a
real consumer estate — not to build a production blockchain. A
homegrown PoA ledger is the minimal viable chain: append-only blocks,
single validator, SHA-256 hash chain, deterministic block production.
This is ~200 lines of Python (block + ledger + validator). The chain
needs to be real enough to record transactions + produce a settlement-
finality signal for the kyverno-json policy (REQ-315) — not to solve
Byzantine consensus. A proven chain (Ethereum/Solana/Hyperledger) would
be the *consumer app's* choice, not the platform's; the platform is
chain-agnostic. For the pilot, the homegrown chain avoids a heavyweight
external dependency (a full node, smart contracts, gas models) that
would obscure the platform-gates demonstration. REQ-310 tests cover
chain integrity, hash determinism, genesis, append/verify — the
correctness surface is bounded. Multi-validator BFT is a future
milestone (D-201). No revision needed.
### G-Q2 — Does "all types of securities" scope-explode the milestone?
**Challenge:** The user said "offering all types of securities." Equities
(D-200, pilot scope) is one type. Bonds (T+2), derivatives (varying),
options (exercise models) have very different settlement models. Does
the equities-only deferral betray the user's intent?
**Verdict:** PROCEED (confidence 0.85). The user *chose* equities-only
pilot (Q4 in the plan discussion, answer "A to all 3 questions" — the
recommended scope). "All types of securities" is the *product vision*;
v1.26 is the *pilot* (equities first). The roadmap documents the
deferral. The pilot demonstrates the Nova platform's gates over the
simplest settlement model (T+1); expanding to other security types is
a straightforward extension (new settlement-service branches + new
kyverno-json policies) once the platform-gates pattern is proven. No
revision needed — the scope decision is the user's, not the grill's.
### G-Q3 — Does the consumer-repo-as-2nd-project break single-project tooling?
**Challenge:** CIAgent has been single-project since v1.0. v1.26
activates multi-project mode (2 projects: `acdl` +
`nova-blockchain-exchange`). Does this break assumptions in the
CIAgent tooling (branch naming, `.ciagent/` paths, commit `---ci---`
blocks)?
**Verdict:** PROCEED (confidence 0.90). `run.md` Step 0 explicitly
specifies multi-project mode: `projects[]` with length > 0,
`active_projects` array, `.ciagent/<slug>/` subdirectory paths, branch
prefixes `<slug>/`. The `---ci---` block gains a `project: <slug>`
field (already in the v1.26 commits). The consumer's project files
live in `.ciagent/nova-blockchain-exchange/`. The platform's existing
flat `.ciagent/` files remain the primary set (the platform is the
default project). Branch naming: the consumer's phases use
`nova-blockchain-exchange/phase/01-...`; the platform's phases use
`acdl/phase/03-...` (or flat `phase/03-...` for platform-level work).
No tooling change needed — the multi-project spec is already in
`run.md`. D-206 records this. No revision needed.
### G-Q4 — Does the P2 contract reference a `dynamodb` module that doesn't exist until P3?
**Challenge:** The original plan had REQ-322 (DynamoDB primitive) in
P3, but the P2 contract (REQ-313) references `dynamodb` in its
`infrastructure` block. If the primitive doesn't exist until P3, the
P2 contract's `dynamodb` block can't resolve at registry time — only
at schema time (the schema is open). Is this a vertical-slice
violation (P2 ships a contract that can't fully resolve)?
**Verdict:** REVISE (confidence 0.92). This is a real vertical-slice
violation. PLAN.md already revised: REQ-322 moves to P2 W0 (before the
contract). The revised mapping (PLAN.md "Revised: REQ-322 → P2 W0")
makes P2 self-contained: the primitive + the contract + the deploy
invocation all land in P2. This is a binding revision — the original
P3 placement is superseded. ROADMAP.md is already updated (REQ-322 in
P2). No further revision needed — the plan self-corrected.
### G-Q5 — Does live-AWS pilot break the MTTR < 60s target?
**Challenge:** NORTH_STAR.md MTTR target: < 60s p95. The pilot runs
`terraform apply` (creating real AWS resources: ECS + DynamoDB + S3).
Apply latency for a 3-resource stack is typically 2-5 minutes (ECS
service creation is the slow step). Does this break the MTTR target?
**Verdict:** PROCEED (confidence 0.86). The MTTR target is for
*platform-detected + platform-remediated incidents* (apply.failed →
successful retry), not for first-time apply latency. The pilot's
first apply is a deployment, not an incident-remediation. The MTTR
metric measures the retry path: if the apply fails (e.g. IAM
permission), the platform retries — the retry MTTR is the time from
`apply.failed` to `apply.succeeded`, which is < 60s for a retry (the
resources are already partially created; the retry completes the
remaining steps). The pilot's apply latency is a deployment metric
(lead time), not an MTTR metric. RESEARCH §1.2 (v1.25 grill G-Q3)
analyzed this same question for the kyverno-json pass — the same
reasoning applies. No revision needed.
### G-Q6 — Is the settlement-finality policy (REQ-315) over-engineering for a pilot?
**Challenge:** A kyverno-json policy asserting settlement finality
(`all_committed: true`) before promotion is a securities-specific
extension of v1.25's policy engine. Is this over-engineering for a
pilot that only runs in `dev` (autonomous, no promotion to qa/prod/dr
in v1.26 per D-208)?
**Verdict:** PROCEED (confidence 0.80). The policy is *authored* in
v1.26 (P3) but its *enforcement* activates when a promotion to qa/prod
happens — which is a *future* milestone (D-208: qa/prod/dr stay
placeholder this milestone). The policy is tested (passing + failing
fixtures; skip when `kj` absent) in P3, but it doesn't gate a `dev`
apply (the pilot-readiness policy REQ-320 gates `dev`; the settlement-
finality policy gates promotions). Authoring + testing the policy in
v1.26 is the right thing: it (a) proves the kyverno-json engine can
assert a domain invariant, (b) ships the policy artifact so a future
milestone that binds qa/prod/dr can enable it without re-architecting,
(c) extends v1.25's moat (the policy engine is swappable + extensible
to new domains). The cost is ~1 policy file + 1 test file. No revision
needed — but the POLICY IS NOT ENFORCED in v1.26 (it's authored +
tested, enforcement is future). PLAN.md should note this. **Minor
revision: PLAN.md P3 W4 Task 4.1 should note "policy authored + tested;
enforcement deferred to the milestone that binds qa/prod/dr."** Already
implicit in the plan (the policy gates promotions, not dev applies);
making it explicit is a documentation refinement, not a scope change.
### G-Q7 — Is D-083 deferral defensible for a pilot with real money-like flows?
**Challenge:** The pilot is a stock exchange — securities trading. D-083
(S3 Object Lock / JWS tamper-evident ledger) is deferred (D-204). The
SQLite hash-chain + DynamoDB outbox is the audit record. Is this
defensible for a domain where audit integrity is legally mandated?
**Verdict:** PROCEED (confidence 0.82). The pilot is a *technical
demonstration*, not a production trading system. No real money, no real
securities, no real investors — the "securities" are test tokens on a
homegrown chain. The audit integrity requirement (SEC Rule 17a-4, FINRA
retention) applies to *production* trading systems, not to a pilot
exercising a platform's deploy/policy/attestation gates. The SQLite
hash-chain + DynamoDB outbox is a tamper-*evident* record (any tampering
breaks the hash chain) — it's just not tamper-*resistant* (S3 Object
Lock + JWS would make it tamper-resistant). For a pilot, tamper-evident
suffices. D-083 lift is a future milestone (when the pilot becomes a
production system). D-204 records this. No revision needed.
### G-Q8 — Does the outcome-backfill emitter (REQ-317) touch the PCR schema?
**Challenge:** REQ-317 wires `apply.completed`/`apply.failed`
`fact_decision.outcome`. The v1.25 hard constraint says "DO NOT change
`schemas/policy_check_result.schema.json`." Does the backfill touch the
PCR schema?
**Verdict:** PROCEED (confidence 0.95). D-211 (CLARIFY) already
resolved this: the outcome backfill touches the *metrics cold store*
(`fact_decision` table in `metrics/nova_metrics.db`), not the PCR
schema. The backfill reads run-manifest events (not PCRs) and updates
the decision's outcome column. The PCR schema is unchanged. This
respects the v1.25 hard constraint. No revision needed.
### G-Q9 — Does the `NOVA_AWS_*` root-equivalent key create a security risk?
**Challenge:** D-207 says `NOVA_AWS_*` has root-equivalent permissions
(confirmed empirically: the bootstrap created the S3 bucket + DynamoDB
table). Using a root key for the pilot's `terraform apply` is a
security risk — a key compromise gives full account access. Should the
pilot use a least-privilege key?
**Verdict:** PROCEED (confidence 0.78). The risk is real but bounded:
(a) the pilot runs in a single account (`581513795199`) with no
production workloads (the v1.11 teardown left it empty; the pilot is
the only workload), (b) the key is in `.env.secrets` (gitignored, never
committed), (c) the deploy workflow uses OIDC by default (the static
key is the override, not the primary path). A future hardening
milestone should split `NOVA_AWS_*` into a root `NOVA_BOOTSTRAP_AWS_*`
+ a least-privilege `NOVA_AWS_*` runner key (the spike-runner pattern).
For v1.26, the single key suffices (pilot scope). D-207 records this.
**Minor revision: PLAN.md should note the key-split as a future
hardening item.** Already implicit in D-207; making it explicit in the
plan is a documentation refinement.
---
## Summary
9 challenges; 0 escalations; 2 binding revisions (G-Q4, G-Q6/G-Q9
minor). Overall verdict: PROCEED (confidence 0.84).
**Binding revisions:**
- **G-Q4:** REQ-322 moves to P2 W0 (already revised in PLAN.md + ROADMAP.md).
- **G-Q6:** PLAN.md P3 W4 Task 4.1 should note the settlement-finality
policy is authored + tested in v1.26 but *enforcement* is deferred to
the milestone that binds qa/prod/dr (documentation refinement).
- **G-Q9:** PLAN.md should note the `NOVA_AWS_*` key-split as a future
hardening item (documentation refinement).
**No work is blocked.** The milestone is feasible, scoped, and the
domain claims hold. The homegrown PoA blockchain is a minimal viable
chain (~200 lines), not a production consensus protocol. The equities-
only scope is the user's choice. The multi-project mode is specified in
`run.md`. The P2→P3 dependency is resolved (REQ-322 → P2 W0). The
MTTR target is for incident-remediation, not first-time apply. The
settlement-finality policy is authored + tested, enforcement is future.
D-083 deferral is defensible for a technical pilot. The PCR schema is
unchanged. The root-equivalent key is a bounded risk with a documented
future hardening path.