36c8ae9a80
10 challenges red-teamed across feasibility, scope, budget, swap boundary. 8 PROCEED (deterministic-not-AI, swap boundary is the moat, MTTR <1s, policy count manageable, tagging cross-check worth it, PCR list is valid payload, phase count matches cadence, severity annotation K8s-standard). 2 REVISE (NullEngine vs kj-not-configured distinct ruleIds; protocol conformance test via NullEngine). All revisions are PLAN/REQ clarifications — no requirement changes. ---ci--- project: acdl phase: 0 milestone: v1.25 status: grill ---/ci---
216 lines
11 KiB
Markdown
216 lines
11 KiB
Markdown
# GRILL — v1.25 kyverno-json Unified Policy Engine
|
||
|
||
> Adversarial review of the v1.25 SPECIFY + CLARIFY + RESEARCH + IDEATE +
|
||
> PLAN. The grill red-teams the proposal across feasibility, scope,
|
||
> budget, and the swap-boundary claim. Each challenge gets a binding
|
||
> verdict (PROCEED / REVISE / ESCALATE). Autonomy: full — escalations
|
||
> auto-resolve with assumption logging unless confidence < 0.60.
|
||
|
||
## Verdict: PROCEED (0.86) — 0 escalations, 2 revisions
|
||
|
||
The milestone is feasible, scoped, and the swap boundary is real. Two
|
||
plan revisions are binding (G-Q4, G-Q8) and are already captured in
|
||
PLAN.md. No work is blocked.
|
||
|
||
---
|
||
|
||
## Challenges
|
||
|
||
### G-Q1 — Does kyverno-json violate "platform functions without AI"?
|
||
|
||
**Challenge:** NORTH_STAR.md Strategic Objective #2 says "the platform
|
||
functions without AI." kyverno-json is a new runtime dependency. Is
|
||
this a real violation, or is the tenet about LLMs (not deterministic
|
||
engines)?
|
||
|
||
**Verdict:** PROCEED (confidence 0.95). kyverno-json is deterministic
|
||
(same policy + payload → same result, every run). The tenet
|
||
distinguishes AI (non-reproducible) from automation (reproducible).
|
||
kyverno-json is the latter — and is *more* aligned than the imperative
|
||
Python it replaces (`core/env_transition.py`, `core/regression_verify.py`)
|
||
because the policy is declarative (visible, auditable). The
|
||
`is_configured()` guard ensures the platform runs without the binary.
|
||
Already resolved as D-120 in CLARIFY. No revision needed.
|
||
|
||
### G-Q2 — Is the PolicyEngine protocol over-engineered for a 2-engine future?
|
||
|
||
**Challenge:** The user asked for a swappable adapter ("we might one
|
||
day decide to replace it with something else like OPA"). A Python
|
||
Protocol + registry is ~40 lines. But Nova has 1 engine today. Is this
|
||
premature abstraction?
|
||
|
||
**Verdict:** PROCEED (confidence 0.85). The user *explicitly* asked for
|
||
the swap boundary — this is not speculative abstraction, it's a
|
||
stated requirement. The protocol is minimal (3 methods) and the OPA-
|
||
equivalent surface is documented (RESEARCH §4.2) — the swap is a known
|
||
quantity, not a hope. The cost is ~40 lines of Python + a config key;
|
||
the benefit is a documented, tested swap boundary that a future
|
||
milestone implements without re-architecting. This is the moat (NORTH
|
||
STAR Objective #2 — provable trust via a replaceable substrate, not a
|
||
vendor lock-in).
|
||
|
||
### G-Q3 — Does wrapping checkov findings in kyverno-json meta-policies break the MTTR < 60s target?
|
||
|
||
**Challenge:** NORTH_STAR.md MTTR target: < 60s p95. Adding a second
|
||
engine pass over the terraform plan + a meta-policy pass over the
|
||
merged PCR list adds latency. Does this break the target?
|
||
|
||
**Verdict:** PROCEED (confidence 0.88). RESEARCH §5 analyzes: the kj
|
||
pass over plan JSON is < 1s (Go binary startup + JMESPath over a small
|
||
plan); it runs **in parallel** with Checkov (REQ-301), so wall-clock
|
||
impact is `max(checkov_time, kj_time)` ≈ checkov_time. Meta-policies
|
||
run in-memory over the merged list (< 10ms). Total MTTR impact: < 1s
|
||
on a 5-15s step. **Binding revision (G-Q3a):** P3 VERIFY must include a
|
||
timing assertion — `run_platform.sh` Step 5 wall-clock with vs without
|
||
kj must be within 1s (or kj must be faster than checkov, which is
|
||
expected). Captured as a P3 verify gate, not a PLAN change.
|
||
|
||
### G-Q4 — Plan revision: NullEngine fallback may mask misconfiguration
|
||
|
||
**Challenge:** PLAN.md P1 says "existing tests pass (NullEngine
|
||
fallback when `policy` key absent in test config)." But the v1.25
|
||
config.json *sets* the `policy` key. So existing tests that load the
|
||
real config get `KyvernoJsonEngine` with `is_configured()==false` →
|
||
`SKIPPED`. The NullEngine fallback only triggers when the key is
|
||
*absent*. Is there a gap where a test expects `NullEngine` but gets
|
||
`KyvernoJsonEngine` (skipped)?
|
||
|
||
**Verdict:** REVISE (confidence 0.82). The fallback path is correct
|
||
but the PLAN wording is ambiguous. **Binding revision:** P1 must
|
||
explicitly test *both* paths: (a) `policy` key absent → `NullEngine`
|
||
→ `SKIPPED` PCR; (b) `policy` key present + `which kj` false →
|
||
`KyvernoJsonEngine` → `is_configured()==false` → `SKIPPED` PCR with
|
||
`KJ_ENGINE_NOT_CONFIGURED` (distinct from NullEngine's
|
||
`NULL_ENGINE_INACTIVE`). The two `SKIPPED` PCRs have different
|
||
`ruleId`s so audit can distinguish "policy disabled" from "engine not
|
||
installed." PLAN.md P1 verification is amended to assert both paths.
|
||
Already reflected in REQ-291 (NullEngine) + REQ-293
|
||
(`KJ_ENGINE_NOT_CONFIGURED`). No requirement change — PLAN wording
|
||
clarified.
|
||
|
||
### G-Q5 — Policy explosion: 4 targets × N rules = maintenance load
|
||
|
||
**Challenge:** v1.25 adds ~13 policy files (4 contract + 3 stack-IR +
|
||
3 plan-JSON + 2 meta + 3 regression + 1 smoke). Each is a YAML file
|
||
with JMESPath. Is this a maintenance burden that grows unbounded?
|
||
|
||
**Verdict:** PROCEED (confidence 0.80). 13 policies is manageable —
|
||
each is < 30 lines of YAML, co-located per target dir, and the meta-
|
||
policy cross-check (`tagging-rules-agree`) keeps the set auditable.
|
||
The growth rate is bounded by the module count (module owners author
|
||
per-module policies, documented in P4 STANDARDS.md). The alternative
|
||
(imperative Python in `regression_verify.py` + `env_transition.py`) is
|
||
*less* auditable — the policies are a net improvement. No revision.
|
||
|
||
### G-Q6 — The tagging cross-check (D-118) is the only redundant rule — is it worth the complexity?
|
||
|
||
**Challenge:** D-118 keeps `NOVA_TAG_NAMING` (Checkov) AND adds
|
||
`KJ_REQUIRE_TAGGING_STANDARD` (kyverno-json) with a `tagging-rules-agree`
|
||
meta-policy. This is the only case where two engines evaluate the same
|
||
concern. Is the defense-in-depth worth the complexity?
|
||
|
||
**Verdict:** PROCEED (confidence 0.82). The tagging standard is the
|
||
highest-impact rule (v1.8 D-tagging-standard, v1.10 re-verification —
|
||
the rule that gates every resource). Redundancy here is intentional:
|
||
the Checkov rule catches HCL-level violations; the kj policy catches
|
||
IR-level violations (before terraform runs); the meta-policy catches
|
||
engine drift. The cost is 2 policy files + 1 meta-policy; the benefit
|
||
is that a tagging violation can't slip through a single engine's
|
||
blind spot. This is the textbook defense-in-depth case. No revision.
|
||
|
||
### G-Q7 — Can `kj scan` actually evaluate the merged PCR list as a payload?
|
||
|
||
**Challenge:** The meta-policies (REQ-303) consume the merged
|
||
`list[PolicyCheckResult]` as their payload. `kj scan` expects a JSON/
|
||
YAML *file*. Is the PCR list a valid kyverno-json payload shape?
|
||
|
||
**Verdict:** PROCEED (confidence 0.85). The PCR list is a JSON array
|
||
of objects — a valid kyverno-json payload. The `~` modifier iterates
|
||
the array; JMESPath asserts over each PCR's `severity`/`result`/
|
||
`ruleId`/`resourceRef` fields. The engine writes the list to a temp
|
||
JSON file and invokes `kj scan --payload <file>`. This is verified in
|
||
P3 `test_meta_policies.py`. No revision — but **binding note (G-Q7a):**
|
||
the `KyvernoJsonEngine.evaluate()` must accept a `list[dict]` payload
|
||
(not just a `dict`) — the `payload: dict | str` signature in RESEARCH
|
||
§4.1 is too narrow. **Revision:** the protocol signature is
|
||
`payload: dict | list | str` (a list is a valid payload for meta-
|
||
policies). Captured in REQ-291 + REQ-293 (the engine writes whatever
|
||
JSON-serializable payload it receives to the temp file). PLAN.md P1
|
||
amended.
|
||
|
||
### G-Q8 — Plan revision: the OPA swap surface claims (RESEARCH §4.2) are unverified
|
||
|
||
**Challenge:** RESEARCH §4.2 documents the OPA-equivalent surface
|
||
(`opa eval -d <dir> -i <json>`), but no `OpaEngine` is implemented in
|
||
v1.25. Is the swap-boundary claim testable, or is it aspirational?
|
||
|
||
**Verdict:** REVISE (confidence 0.78). The swap-boundary claim is
|
||
*testable in v1.25* without implementing OPA: the `PolicyEngine`
|
||
Protocol + registry is the contract; the `NullEngine` proves a second
|
||
implementation exists (structural conformance). **Binding revision
|
||
(G-Q8a):** P1 `test_policy_engine.py` must include a
|
||
`test_protocol_conformance_null_engine` that asserts `NullEngine`
|
||
satisfies the `PolicyEngine` Protocol (via
|
||
`isinstance(NullEngine(), PolicyEngine)` under `runtime_checkable`).
|
||
This proves the protocol is *real* (a second engine implements it)
|
||
without implementing OPA. The OPA-equivalent surface in RESEARCH §4.2
|
||
stays as documentation (the future milestone implements it). PLAN.md
|
||
P1 verification amended. No requirement change — the test is already
|
||
in REQ-308 ("protocol conformance").
|
||
|
||
### G-Q9 — Budget: is 4 execution phases + P5 too many for the scope?
|
||
|
||
**Challenge:** v1.25 is 19 requirements across 6 phases. Recent
|
||
milestones: v1.24 had 15 reqs / 4 phases; v1.23 had 13 reqs / 7 phases.
|
||
Is 6 phases too many (overhead) or too few (per-phase overload)?
|
||
|
||
**Verdict:** PROCEED (confidence 0.85). 19 reqs / 6 phases ≈ 3.2 reqs/
|
||
phase — within the v1.24 cadence (3.75 reqs/phase). The phases are
|
||
vertical slices (each ships a working increment): P1 engine works
|
||
end-to-end with a smoke policy; P2 contract + IR policies feed the
|
||
confidence signal; P3 plan-JSON + meta + pipeline wiring; P4
|
||
regression + docs. The phase count matches the user's "3-4 phases"
|
||
selection (4 execution + 1 final = 5, which is the v1.24 shape). No
|
||
revision.
|
||
|
||
### G-Q10 — The `nova.cloudinit.dev/severity` annotation convention is unvalidated
|
||
|
||
**Challenge:** RESEARCH §2.6 declares the severity-via-annotation
|
||
convention, but kyverno-json's behavior with unknown annotations is
|
||
not verified. Does `kj scan` ignore unknown annotations, or does it
|
||
reject the policy?
|
||
|
||
**Verdict:** PROCEED (confidence 0.80). kyverno-json is Kubernetes-
|
||
style CRD-based — unknown `metadata.annotations` are preserved and
|
||
ignored (standard K8s behavior). The engine reads the annotation from
|
||
the loaded policy YAML (via `yaml.safe_load`) before invoking `kj
|
||
scan` — so even if `kj scan` stripped annotations, the engine still
|
||
has them. **Binding note (G-Q10a):** P1 `test_kyverno_json_engine.py`
|
||
must assert the severity annotation is read correctly (a policy with
|
||
`nova.cloudinit.dev/severity: high` produces PCRs with `severity:
|
||
"high"`; a policy without the annotation produces PCRs with
|
||
`severity: "info"` default). Captured in REQ-309 ("PCR schema
|
||
validity" includes severity). No requirement change — the test is
|
||
already in REQ-309.
|
||
|
||
---
|
||
|
||
## Summary
|
||
|
||
10 challenges; 10 resolved (8 PROCEED, 2 REVISE, 0 ESCALATE).
|
||
- **Revisions (binding, already in PLAN/REQs):**
|
||
- G-Q4: P1 tests both fallback paths (NullEngine vs
|
||
KyvernoJsonEngine-not-configured) — distinct `ruleId`s for audit.
|
||
- G-Q7a: protocol signature `payload: dict | list | str` (list is a
|
||
valid payload for meta-policies).
|
||
- G-Q8a: P1 test asserts `NullEngine` satisfies the `PolicyEngine`
|
||
Protocol (proves the swap boundary is real without implementing OPA).
|
||
- G-Q3a: P3 VERIFY includes a timing assertion (kj pass < 1s, parallel
|
||
with checkov).
|
||
- G-Q10a: P1 test asserts severity annotation is read correctly.
|
||
- **No requirement changes** — all revisions are clarifications to
|
||
PLAN.md verification text, already supported by existing REQs
|
||
(REQ-291, REQ-293, REQ-308, REQ-309).
|
||
- **0 escalations** — all challenges auto-resolved at full autonomy.
|
||
|
||
The milestone PROCEEDs to PHASE 0 SHIP → P1. |