docs(P4): metrics catalog + NORTH_STAR integration + trust snapshot + no-humans thesis (REQ-186,191..195,204,210..213)

P4 (Wave 3, docs) — REQ-186, 191, 192, 193, 194, 195, 204, 210, 211, 212, 213

New docs:
- docs/METRICS.md — canonical KPI catalog (grounded/derived/deferred)
- docs/metrics/*.md — 13 per-KPI definition-of-success docs (D-127)
- docs/METRICS_DEFERRED_ROADMAP.md — 8 deferred metrics + hot-path plan + re-eval triggers (REQ-210)
- docs/NO_HUMANS_THESIS.md — thesis defensibility brief (REQ-213)

New tools:
- core/metrics/trust_snapshot.py — 5 trust metrics + chain-integrity verdict + snapshot hash (REQ-211)
- scripts/check_north_star_diff.sh — CI check for NORTH_STAR strategic section changes (REQ-204)

Modified:
- .ciagent/config.json — strategic_direction_file: .ciagent/NORTH_STAR.md (REQ-186)

---ci---
project: acdl
phase: 4
milestone: v1.17
status: execute
---/ci---
This commit is contained in:
Jon Chery
2026-08-04 20:05:06 +00:00
parent 942185c85b
commit b054849a99
20 changed files with 793 additions and 1 deletions
+21
View File
@@ -0,0 +1,21 @@
# AI Decision Accuracy — Definition of Success
> KPI: AI Decision Accuracy
> Target: ≥ 99.5% (no rollback, no follow-up incident within 5 min of action)
**What this number means:** the percentage of AI decisions (confidence-
gated policy engine outcomes) that were NOT followed by an apply failure
or incident within 5 minutes. A high-confidence decision that later
caused an incident does NOT count as accurate.
**How it's computed:** `count(decisions WHERE outcome = 'succeeded' AND
no incident within 5min)` ÷ `total decisions`. Correlation via
`decision_id``run_id` → subsequent `apply.failed` or `incident.detected`
events.
**What "good" looks like:** ≥ 99.5% means fewer than 1 in 200 decisions
cause a secondary failure. The 0.5% allowance is for novel edge cases.
**D-122 honesty:** Nova's "AI" is the confidence-gated policy engine
(confidence_signal + HITL gate), not an LLM planner. The Decision Ledger
captures this real decision path — not a fabricated "AI agent."
+18
View File
@@ -0,0 +1,18 @@
# Attestation Coverage — Definition of Success
> KPI: Attestation Coverage
> Target: 100% of prod/dr promotions attested by a human
**What this number means:** every production and disaster-recovery
promotion has a recorded human attestation (approver identity, 8-concern
matrix result, separation-of-duties check on prod). This is the
"autonomy in operations, human in accountability" proof.
**How it's computed:** `count(prod/dr promotions with attestation.recorded
event) ÷ count(total prod/dr promotions)`. Sourced from the Decision
Ledger (`attestation.recorded` events) + `hitl_gates.py` + outbox
`approver_*` attributes.
**What "good" looks like:** 100% means no prod/dr promotion ever lands
without a human sign-off on record. The absence of an operator is never
the absence of a record (NORTH_STAR Anti-Goal #3).
+16
View File
@@ -0,0 +1,16 @@
# Confidence-Gate Halt Rate — Definition of Success
> KPI: Confidence-Gate Halt Rate
> Target: not a committed target (operational signal)
**What this number means:** how often the confidence gate itself halted
a run (band = block), independent of HITL blocks. The gate is the AI's
self-halt; HITL is the human gate. This distinguishes the AI's
self-regulation from human escalation.
**How it's computed:** `count(runs WHERE confidence_band = 'block')` ÷
`total runs`.
**What "good" looks like:** a low but non-zero rate means the gate is
working (catching genuinely uncertain runs) without being overly
conservative (blocking everything).
+18
View File
@@ -0,0 +1,18 @@
# Cost Savings via Infracost Estimates — Definition of Success
> KPI: Cost Savings via Infracost Estimates
> Target: ≥ 25% on pilot estates (partial)
**What this number means:** the pre-apply cost estimate from Infracost
shows the delta between the planned infrastructure and the current
state. Negative deltas = savings.
**How it's computed:** `sum(fact_cost_estimate.delta_usd WHERE delta < 0)`
per period.
**What's grounded:** the pre-apply estimate (Infracost reads plan JSON,
offline).
**What's deferred:** actual-spend reconciliation from AWS CUR (D-096 —
needs live AWS billing). The placeholder view
`placeholder_live_cur_reconciliation.csv` has the schema ready.
+16
View File
@@ -0,0 +1,16 @@
# Decision Ledger Coverage — Definition of Success
> KPI: Decision Ledger Coverage
> Target: 100% of AI actions with backfilled outcome
**What this number means:** every AI decision (confidence-gated policy
engine outcome) is captured in the Decision Ledger with its outcome
backfilled from the subsequent apply.completed/failed event.
**How it's computed:** `count(decision_ledger rows WHERE outcome ≠
'pending') ÷ count(decision_ledger rows)`. Sourced from
`metrics/decision_ledger.db`.
**What "good" looks like:** 100% means no AI decision is ever lost or
left without an outcome. The ledger is the trust substrate (NORTH_STAR
Objective #2).
+13
View File
@@ -0,0 +1,13 @@
# Deployment Frequency — Definition of Success
> KPI: Deployment Frequency
> Target: not a committed target (operational signal)
**What this number means:** the rate of infrastructure state updates
deployed safely per day. A DORA-adjacent metric for infrastructure.
**How it's computed:** `count(run.completed WHERE exit_code = 0)` per
day.
**What "good" looks like:** multiple deploys per day (vs. weekly/monthly
for human ops teams).
+17
View File
@@ -0,0 +1,17 @@
# FTE Hours Saved (Toil Reallocation Value) — Definition of Success
> KPI: FTE Hours Saved
> Target: ≥ 70% of pre-Nova FTE allocation (derived)
**What this number means:** the engineering hours saved by automated
operations, valued at the blended engineering rate. This is what those
hours were spent on instead (the "toil reallocation" — capital freed
up from ops to feature development).
**How it's computed:** `run count × manual baseline minutes per run ÷ 60
× blended hourly rate`. The manual baseline is the estimated time a
human team would take for the same operation (e.g., 30 min/ticket).
**Honesty caveat:** computed on N internal runs today; the production-
denominator activates post-pilot. The formula is grounded; the
production numbers are not yet.
@@ -0,0 +1,18 @@
# Human Escalation Frequency — Definition of Success
> KPI: Human Escalation Frequency
> Target: < 0.1% of platform actions (Post-Pilot)
**What this number means:** how often the AI platform was forced to fall
back or escalate to a human operator due to low confidence. This is the
inverse of Touchless Resolution Rate, scoped to operational escalations
only.
**How it's computed:** `count(runs WHERE hitl_block = 1 AND reason =
'confidence')` ÷ `total runs`. Attestation sign-offs are excluded.
**What "good" looks like:** < 0.1% means fewer than 1 in 1000 runs
require human intervention. Near-zero is the goal.
**What would be "gamer metrics":** counting attestation sign-offs as
escalations (they're not — they're designed controls).
+19
View File
@@ -0,0 +1,19 @@
# MTTR (Platform-Run) — Definition of Success
> KPI: MTTR (p95)
> Target: < 60 seconds
**What this number means:** the time from a platform-run failure
(apply.failed) to a successful retry. This is platform-run MTTR, not
infra-incident MTTR (which requires an incident detection system that
Nova doesn't have yet — deferred).
**How it's computed:** p95 of `successful_retry.time failed_run.time`
across all runs that failed then succeeded.
**What "good" looks like:** < 60 seconds means the platform recovers
from a failed run in under a minute, 95% of the time.
**What's deferred:** infra-incident MTTR (anomaly detected → healed)
requires an incident detection/remediation system (self-healing
velocity). That's a future emitter.
+15
View File
@@ -0,0 +1,15 @@
# Platform ROI — Definition of Success
> KPI: Platform ROI
> Target: ≥ 250% measured annually (derived)
**What this number means:** the total financial value delivered (labor
savings + cloud cost optimization + avoided downtime losses) vs. the
platform's operational/licensing cost.
**Formula:** `(FTE hours saved × blended rate + cloud savings + avoided
downtime) ÷ platform op cost`.
**Honesty caveat:** computed on N internal runs today; the production-
denominator activates post-pilot. The formula is grounded; the
production numbers are not yet.
+14
View File
@@ -0,0 +1,14 @@
# Zero-Trust Policy Compliance Rate — Definition of Success
> KPI: Zero-Trust Policy Compliance Rate
> Target: not a committed target (operational signal)
**What this number means:** the percentage of infrastructure assets
continuously verified as compliant with security baselines and policies.
**How it's computed:** `1 count(assets WHERE last_scan.status ≠ pass)
÷ count(assets)`. Sourced from `fact_policy_check` (Checkov results).
**What "good" looks like:** 100% means every resource passed every
policy check. The Nova tagging standard (nova_tagging.py, hard mode) is
the primary check.
+13
View File
@@ -0,0 +1,13 @@
# Provisioning Lead Time — Definition of Success
> KPI: Provisioning Lead Time
> Target: not a committed target (operational signal)
**What this number means:** the time from intent received (run.started)
to apply completed (run.completed). Measures how fast Nova provisions
compliant environments.
**How it's computed:** `run.completed_at run.started_at` per run.
**What "good" looks like:** minutes, not days. The reduction from days
(human ops) to minutes (autonomous) is the velocity proof.
+23
View File
@@ -0,0 +1,23 @@
# Touchless Resolution Rate — Definition of Success
> KPI: Touchless Resolution Rate
> Target: ≥ 99% across production estates (Post-Pilot)
**What this number means:** the percentage of platform runs that complete
end-to-end without an operational HITL block. An operational HITL block
is a confidence-driven escalation (the AI's confidence was too low to
proceed). Attestation gates (qa/prod/dr sign-offs) are NOT counted as
escalations — they are designed controls, not autonomy failures.
**How it's computed:** `runs WHERE hitl_block = 0 AND environment = 'dev'`
÷ `total runs` (dev environment only, where attestation gates don't apply).
For production estates: `runs WHERE hitl_block = 0` ÷ `total runs`
excluding attestation-gate sign-offs.
**What "good" looks like:** ≥ 99% means fewer than 1 in 100 runs require
human intervention due to low confidence. The 1% allowance is for
genuine edge cases (novel failure modes, blast-radius exceedances).
**What would be "gamer metrics":** counting attestation gates as
"touchless" (they're not — they're human by design) or counting only
dev runs (cherry-picking the easiest environment).