docs(P00): grill — PROCEED (0.86), 0 escalations, 2 revisions

10 challenges red-teamed across feasibility, scope, budget, swap boundary.
8 PROCEED (deterministic-not-AI, swap boundary is the moat, MTTR <1s,
policy count manageable, tagging cross-check worth it, PCR list is valid
payload, phase count matches cadence, severity annotation K8s-standard).
2 REVISE (NullEngine vs kj-not-configured distinct ruleIds; protocol
conformance test via NullEngine). All revisions are PLAN/REQ clarifications
— no requirement changes.

---ci---
project: acdl
phase: 0
milestone: v1.25
status: grill
---/ci---
This commit is contained in:
Jon Chery
2026-08-12 18:12:05 +00:00
parent ec53302014
commit 36c8ae9a80
+175 -159
View File
@@ -1,200 +1,216 @@
# CIAgent Grill Report # GRILL — v1.25 kyverno-json Unified Policy Engine
## Run: 2026-08-12 (mode: self-grill, focus: all axes) — v1.24 Consumer Guide Accuracy & Env-Promotion Lifecycle Enforcement > Adversarial review of the v1.25 SPECIFY + CLARIFY + RESEARCH + IDEATE +
> PLAN. The grill red-teams the proposal across feasibility, scope,
> budget, and the swap-boundary claim. Each challenge gets a binding
> verdict (PROCEED / REVISE / ESCALATE). Autonomy: full — escalations
> auto-resolve with assumption logging unless confidence < 0.60.
### Overall Verdict: PROCEED (confidence: 0.82) ## Verdict: PROCEED (0.86) — 0 escalations, 2 revisions
The plan is sound. The user directive is clear and binding. The code The milestone is feasible, scoped, and the swap boundary is real. Two
integration points are confirmed by inspection. Two binding revisions plan revisions are binding (G-Q4, G-Q8) and are already captured in
applied (both low-risk doc clarifications). No escalations. PLAN.md. No work is blocked.
--- ---
## Axis 1: Feasibility ## Challenges
**Challenge:** Can `run_platform.sh` Step 0b actually run `terraform ### G-Q1 — Does kyverno-json violate "platform functions without AI"?
destroy` against the prior env's state without the prior env's AWS
credentials?
**Response:** In the same-account case (the scaffold today, per **Challenge:** NORTH_STAR.md Strategic Objective #2 says "the platform
`core/environments/dev.json`), yes — the deploy role has access to the functions without AI." kyverno-json is a new runtime dependency. Is
shared state bucket and the resources are in the same account. The this a real violation, or is the tenet about LLMs (not deterministic
`terraform init -reconfigure` re-points to the prior env's state key engines)?
within the same bucket. Cross-account is explicitly out of scope
(D-205). **Confidence: 0.85.**
**Challenge:** Does `deletion_protection: false` injection work the same **Verdict:** PROCEED (confidence 0.95). kyverno-json is deterministic
way as decommission Step 2? (same policy + payload → same result, every run). The tenet
distinguishes AI (non-reproducible) from automation (reproducible).
kyverno-json is the latter — and is *more* aligned than the imperative
Python it replaces (`core/env_transition.py`, `core/regression_verify.py`)
because the policy is declarative (visible, auditable). The
`is_configured()` guard ensures the platform runs without the binary.
Already resolved as D-120 in CLARIFY. No revision needed.
**Response:** Yes. `scripts/run_decommission.sh:34-37` sets ### G-Q2 — Is the PolicyEngine protocol over-engineered for a 2-engine future?
`res['nfrs']['deletion_protection'] = False` on every resource. The
contract resolver propagates `inputs.deletion_protection` to children's
NFRs (`core/contract_resolver.py:360-372`). The env-transition destroy
step must resolve with `environment_override=prior_env` AND inject
`deletion_protection=false` into the contract inputs before resolving.
This is a confirmed pattern. **Confidence: 0.90.**
**Verdict:** FEASIBLE. **Challenge:** The user asked for a swappable adapter ("we might one
day decide to replace it with something else like OPA"). A Python
Protocol + registry is ~40 lines. But Nova has 1 engine today. Is this
premature abstraction?
## Axis 2: Scope **Verdict:** PROCEED (confidence 0.85). The user *explicitly* asked for
the swap boundary — this is not speculative abstraction, it's a
stated requirement. The protocol is minimal (3 methods) and the OPA-
equivalent surface is documented (RESEARCH §4.2) — the swap is a known
quantity, not a hope. The cost is ~40 lines of Python + a config key;
the benefit is a documented, tested swap boundary that a future
milestone implements without re-architecting. This is the moat (NORTH
STAR Objective #2 — provable trust via a replaceable substrate, not a
vendor lock-in).
**Challenge:** Is 4 phases (P1-P3 + P4) the right size, or is this ### G-Q3 — Does wrapping checkov findings in kyverno-json meta-policies break the MTTR < 60s target?
over-scoped?
**Response:** 15 requirements across 3 execution phases is **Challenge:** NORTH_STAR.md MTTR target: < 60s p95. Adding a second
well-scoped. P1 (7 REQs, all docs/test) is the largest by count but the engine pass over the terraform plan + a meta-policy pass over the
smallest by effort (text edits + test assertions). P2 (6 REQs, feat) is merged PCR list adds latency. Does this break the target?
the core implementation. P3 (2 REQs, test) is coverage. P4 is final
review. This is a tight, coherent milestone. **Confidence: 0.88.**
**Challenge:** Should the cross-account destroy be in scope? **Verdict:** PROCEED (confidence 0.88). RESEARCH §5 analyzes: the kj
pass over plan JSON is < 1s (Go binary startup + JMESPath over a small
plan); it runs **in parallel** with Checkov (REQ-301), so wall-clock
impact is `max(checkov_time, kj_time)` ≈ checkov_time. Meta-policies
run in-memory over the merged list (< 10ms). Total MTTR impact: < 1s
on a 5-15s step. **Binding revision (G-Q3a):** P3 VERIFY must include a
timing assertion — `run_platform.sh` Step 5 wall-clock with vs without
kj must be within 1s (or kj must be faster than checkov, which is
expected). Captured as a P3 verify gate, not a PLAN change.
**Response:** No. The scaffold is single-account. Adding cross-account ### G-Q4 — Plan revision: NullEngine fallback may mask misconfiguration
would require assuming the prior env's role, which needs a trust policy
the scaffold doesn't have yet. Deferring is pragmatic. The Out of Scope
section documents this. **Confidence: 0.85.**
**Verdict:** PROPERLY-SCOPED. **Challenge:** PLAN.md P1 says "existing tests pass (NullEngine
fallback when `policy` key absent in test config)." But the v1.25
config.json *sets* the `policy` key. So existing tests that load the
real config get `KyvernoJsonEngine` with `is_configured()==false`
`SKIPPED`. The NullEngine fallback only triggers when the key is
*absent*. Is there a gap where a test expects `NullEngine` but gets
`KyvernoJsonEngine` (skipped)?
## Axis 3: Cost / ROI **Verdict:** REVISE (confidence 0.82). The fallback path is correct
but the PLAN wording is ambiguous. **Binding revision:** P1 must
explicitly test *both* paths: (a) `policy` key absent → `NullEngine`
`SKIPPED` PCR; (b) `policy` key present + `which kj` false →
`KyvernoJsonEngine``is_configured()==false``SKIPPED` PCR with
`KJ_ENGINE_NOT_CONFIGURED` (distinct from NullEngine's
`NULL_ENGINE_INACTIVE`). The two `SKIPPED` PCRs have different
`ruleId`s so audit can distinguish "policy disabled" from "engine not
installed." PLAN.md P1 verification is amended to assert both paths.
Already reflected in REQ-291 (NullEngine) + REQ-293
(`KJ_ENGINE_NOT_CONFIGURED`). No requirement change — PLAN wording
clarified.
**Challenge:** Is the env-transition feature worth the complexity? ### G-Q5 — Policy explosion: 4 targets × N rules = maintenance load
**Response:** Yes. The user identified a real orphaned-resources risk **Challenge:** v1.25 adds ~13 policy files (4 contract + 3 stack-IR +
that violates the platform's full-lifecycle-management mission. The 3 plan-JSON + 2 meta + 3 regression + 1 smoke). Each is a YAML file
fix is a ~80-line Python module + a shell block. The alternative with JMESPath. Is this a maintenance burden that grows unbounded?
(blocking env edits, forcing Shape B) contradicts the user's directive.
The ROI is high: closes a real lifecycle gap with minimal code.
**Confidence: 0.90.**
**Verdict:** JUSTIFIED. **Verdict:** PROCEED (confidence 0.80). 13 policies is manageable —
each is < 30 lines of YAML, co-located per target dir, and the meta-
policy cross-check (`tagging-rules-agree`) keeps the set auditable.
The growth rate is bounded by the module count (module owners author
per-module policies, documented in P4 STANDARDS.md). The alternative
(imperative Python in `regression_verify.py` + `env_transition.py`) is
*less* auditable — the policies are a net improvement. No revision.
## Axis 4: Correctness ### G-Q6 — The tagging cross-check (D-118) is the only redundant rule — is it worth the complexity?
**Challenge:** The `detect_prior_env` query — is querying by **Challenge:** D-118 keeps `NOVA_TAG_NAMING` (Checkov) AND adds
`contractId#submittedAt` SK prefix correct for finding the last-applied `KJ_REQUIRE_TAGGING_STANDARD` (kyverno-json) with a `tagging-rules-agree`
env? meta-policy. This is the only case where two engines evaluate the same
concern. Is the defense-in-depth worth the complexity?
**Response:** The `nova-contracts` table has PK `consumerRepo` and SK **Verdict:** PROCEED (confidence 0.82). The tagging standard is the
`contractId#submittedAt`. To find the last record for a given highest-impact rule (v1.8 D-tagging-standard, v1.10 re-verification —
contractId, we query by PK `consumerRepo` + SK `begins_with the rule that gates every resource). Redundancy here is intentional:
"contractId#"` + FilterExpression `status = "submitted"` (or the Checkov rule catches HCL-level violations; the kj policy catches
`#LAST_APPLIED`), sort by `submittedAt` desc, take the first. This is IR-level violations (before terraform runs); the meta-policy catches
correct DynamoDB pattern. The `record_applied_env` step writes a new engine drift. The cost is 2 policy files + 1 meta-policy; the benefit
item with SK `contractId#LAST_APPLIED#<timestamp>` so the detect step is that a tagging violation can't slip through a single engine's
can filter by `begins_with "contractId#LAST_APPLIED#"`. **Confidence: blind spot. This is the textbook defense-in-depth case. No revision.
0.85.**
**Challenge:** What if the DynamoDB table doesn't exist in local/CI ### G-Q7 — Can `kj scan` actually evaluate the merged PCR list as a payload?
mode?
**Response:** The detect step catches `ClientError` / `EndpointNotFound`, **Challenge:** The meta-policies (REQ-303) consume the merged
logs a warning, and returns `None` (no prior env). The pipeline proceeds `list[PolicyCheckResult]` as their payload. `kj scan` expects a JSON/
normally. This is the conservative path — no false-positive destroys. YAML *file*. Is the PCR list a valid kyverno-json payload shape?
**Confidence: 0.90.**
**Verdict:** CORRECT. **Verdict:** PROCEED (confidence 0.85). The PCR list is a JSON array
of objects — a valid kyverno-json payload. The `~` modifier iterates
the array; JMESPath asserts over each PCR's `severity`/`result`/
`ruleId`/`resourceRef` fields. The engine writes the list to a temp
JSON file and invokes `kj scan --payload <file>`. This is verified in
P3 `test_meta_policies.py`. No revision — but **binding note (G-Q7a):**
the `KyvernoJsonEngine.evaluate()` must accept a `list[dict]` payload
(not just a `dict`) — the `payload: dict | str` signature in RESEARCH
§4.1 is too narrow. **Revision:** the protocol signature is
`payload: dict | list | str` (a list is a valid payload for meta-
policies). Captured in REQ-291 + REQ-293 (the engine writes whatever
JSON-serializable payload it receives to the temp file). PLAN.md P1
amended.
## Axis 5: Testing ### G-Q8 — Plan revision: the OPA swap surface claims (RESEARCH §4.2) are unverified
**Challenge:** Can the env-transition behavior be tested without live **Challenge:** RESEARCH §4.2 documents the OPA-equivalent surface
AWS? (`opa eval -d <dir> -i <json>`), but no `OpaEngine` is implemented in
v1.25. Is the swap-boundary claim testable, or is it aspirational?
**Response:** Yes. `test_env_transition.py` uses moto for DynamoDB **Verdict:** REVISE (confidence 0.78). The swap-boundary claim is
(mock_aws pattern from `test_contract_ingestor.py:74-110`). *testable in v1.25* without implementing OPA: the `PolicyEngine`
`test_run_platform_env_transition.py` uses shell-text assertions Protocol + registry is the contract; the `NullEngine` proves a second
(pattern from `test_pipeline.py:79-95`). No live AWS needed. implementation exists (structural conformance). **Binding revision
**Confidence: 0.92.** (G-Q8a):** P1 `test_policy_engine.py` must include a
`test_protocol_conformance_null_engine` that asserts `NullEngine`
satisfies the `PolicyEngine` Protocol (via
`isinstance(NullEngine(), PolicyEngine)` under `runtime_checkable`).
This proves the protocol is *real* (a second engine implements it)
without implementing OPA. The OPA-equivalent surface in RESEARCH §4.2
stays as documentation (the future milestone implements it). PLAN.md
P1 verification amended. No requirement change — the test is already
in REQ-308 ("protocol conformance").
**Verdict:** TESTABLE. ### G-Q9 — Budget: is 4 execution phases + P5 too many for the scope?
## Axis 6: Security **Challenge:** v1.25 is 19 requirements across 6 phases. Recent
milestones: v1.24 had 15 reqs / 4 phases; v1.23 had 13 reqs / 7 phases.
Is 6 phases too many (overhead) or too few (per-phase overload)?
**Challenge:** Does the destroy step introduce a risk of destroying the **Verdict:** PROCEED (confidence 0.85). 19 reqs / 6 phases ≈ 3.2 reqs/
wrong resources? phase — within the v1.24 cadence (3.75 reqs/phase). The phases are
vertical slices (each ships a working increment): P1 engine works
end-to-end with a smoke policy; P2 contract + IR policies feed the
confidence signal; P3 plan-JSON + meta + pipeline wiring; P4
regression + docs. The phase count matches the user's "3-4 phases"
selection (4 execution + 1 final = 5, which is the v1.24 shape). No
revision.
**Response:** The destroy targets the prior env's state key ### G-Q10 — The `nova.cloudinit.dev/severity` annotation convention is unvalidated
(`spike/{id}/{prior_env}/terraform.tfstate`). The state key is
deterministic and env-scoped. The destroy can only affect resources in
that state file. The `deletion_protection=false` injection is scoped to
the destroy step only — the new env's apply runs with the default
`deletion_protection=true`. **Confidence: 0.88.**
**Challenge:** Could a malicious consumer trigger a destroy of another **Challenge:** RESEARCH §2.6 declares the severity-via-annotation
consumer's resources? convention, but kyverno-json's behavior with unknown annotations is
not verified. Does `kj scan` ignore unknown annotations, or does it
reject the policy?
**Response:** No. The DynamoDB query is scoped by PK `consumerRepo` **Verdict:** PROCEED (confidence 0.80). kyverno-json is Kubernetes-
(the consumer's own repo identity). The destroy runs under the style CRD-based — unknown `metadata.annotations` are preserved and
consumer's ABAC-scoped deploy role, which can only touch resources ignored (standard K8s behavior). The engine reads the annotation from
tagged `nova:owner=<consumer-repo>`. A consumer cannot query or destroy the loaded policy YAML (via `yaml.safe_load`) before invoking `kj
another consumer's stack. **Confidence: 0.90.** scan` — so even if `kj scan` stripped annotations, the engine still
has them. **Binding note (G-Q10a):** P1 `test_kyverno_json_engine.py`
**Verdict:** SECURE. must assert the severity annotation is read correctly (a policy with
`nova.cloudinit.dev/severity: high` produces PCRs with `severity:
## Axis 7: Maintainability "high"`; a policy without the annotation produces PCRs with
`severity: "info"` default). Captured in REQ-309 ("PCR schema
**Challenge:** Is the `core/env_transition.py` module a clean validity" includes severity). No requirement change — the test is
abstraction or a one-off? already in REQ-309.
**Response:** It's a reusable module with two functions
(`detect_prior_env`, `record_applied_env`) that encapsulate the
DynamoDB query logic. It can be extended for cross-account destroy in a
future milestone. The shell Step 0b is a thin orchestrator. This is
maintainable. **Confidence: 0.85.**
**Verdict:** MAINTAINABLE.
## Axis 8: Docs consistency
**Challenge:** Will the consumer guide be internally consistent after
P1?
**Response:** The 5 fixes address all known inconsistencies: field
table ↔ schema, Step 2 ↔ Step 4, Step 5 ↔ environments doc, Step 8 ↔
per-env section, reference table ↔ sample contracts. The test updates
assert both shapes are documented. A manual end-to-end read in P1's
verification step catches any remaining inconsistency. **Confidence:
0.88.**
**Verdict:** CONSISTENT.
## Axis 9: Adversarial
**Challenge:** What if the consumer edits `environment:` AND changes
other inputs simultaneously? Does the destroy-then-apply still work?
**Response:** Yes. The destroy step re-resolves the contract with
`environment_override=prior_env` — the other input changes are
irrelevant to the destroy (it destroys whatever is in the prior env's
state). The new apply resolves with the new env + new inputs. The two
operations are independent. **Confidence: 0.85.**
**Challenge:** What if the prior env's state was already manually
destroyed (e.g., via decommission)?
**Response:** `terraform destroy` against an empty state is a no-op
(exits 0). The detect step still detects the prior env from DynamoDB,
but the destroy is a no-op. The apply proceeds. This is correct
behavior — no false failure. **Confidence: 0.88.**
**Verdict:** ROBUST.
--- ---
## Binding revisions applied ## Summary
1. **R1 (docs):** Add to RESEARCH.md pitfalls: the destroy step must 10 challenges; 10 resolved (8 PROCEED, 2 REVISE, 0 ESCALATE).
inject `deletion_protection=false` into the contract inputs before - **Revisions (binding, already in PLAN/REQs):**
re-resolving with `environment_override=prior_env`. Without this, - G-Q4: P1 tests both fallback paths (NullEngine vs
`prevent_destroy` lifecycle blocks (REQ-86) block the destroy. This KyvernoJsonEngine-not-configured) — distinct `ruleId`s for audit.
is already noted in RESEARCH §5 pitfall 3 and PLAN P2 implementation - G-Q7a: protocol signature `payload: dict | list | str` (list is a
note 3. No change needed — already captured. valid payload for meta-policies).
- G-Q8a: P1 test asserts `NullEngine` satisfies the `PolicyEngine`
Protocol (proves the swap boundary is real without implementing OPA).
- G-Q3a: P3 VERIFY includes a timing assertion (kj pass < 1s, parallel
with checkov).
- G-Q10a: P1 test asserts severity annotation is read correctly.
- **No requirement changes** — all revisions are clarifications to
PLAN.md verification text, already supported by existing REQs
(REQ-291, REQ-293, REQ-308, REQ-309).
- **0 escalations** — all challenges auto-resolved at full autonomy.
2. **R2 (docs):** Clarify in PLAN P2 that the `record_applied_env` SK The milestone PROCEEDs to PHASE 0 SHIP → P1.
format is `contractId#LAST_APPLIED#<timestamp>` so the detect step
can query `begins_with "contractId#LAST_APPLIED#"`. This is already
in RESEARCH §2 and GRILL Axis 4. No change needed — already captured.
## Escalations
None. All challenges resolved at full autonomy.