b054849a99
P4 (Wave 3, docs) — REQ-186, 191, 192, 193, 194, 195, 204, 210, 211, 212, 213 New docs: - docs/METRICS.md — canonical KPI catalog (grounded/derived/deferred) - docs/metrics/*.md — 13 per-KPI definition-of-success docs (D-127) - docs/METRICS_DEFERRED_ROADMAP.md — 8 deferred metrics + hot-path plan + re-eval triggers (REQ-210) - docs/NO_HUMANS_THESIS.md — thesis defensibility brief (REQ-213) New tools: - core/metrics/trust_snapshot.py — 5 trust metrics + chain-integrity verdict + snapshot hash (REQ-211) - scripts/check_north_star_diff.sh — CI check for NORTH_STAR strategic section changes (REQ-204) Modified: - .ciagent/config.json — strategic_direction_file: .ciagent/NORTH_STAR.md (REQ-186) ---ci--- project: acdl phase: 4 milestone: v1.17 status: execute ---/ci---
18 lines
747 B
Markdown
18 lines
747 B
Markdown
# Human Escalation Frequency — Definition of Success
|
|
|
|
> KPI: Human Escalation Frequency
|
|
> Target: < 0.1% of platform actions (Post-Pilot)
|
|
|
|
**What this number means:** how often the AI platform was forced to fall
|
|
back or escalate to a human operator due to low confidence. This is the
|
|
inverse of Touchless Resolution Rate, scoped to operational escalations
|
|
only.
|
|
|
|
**How it's computed:** `count(runs WHERE hitl_block = 1 AND reason =
|
|
'confidence')` ÷ `total runs`. Attestation sign-offs are excluded.
|
|
|
|
**What "good" looks like:** < 0.1% means fewer than 1 in 1000 runs
|
|
require human intervention. Near-zero is the goal.
|
|
|
|
**What would be "gamer metrics":** counting attestation sign-offs as
|
|
escalations (they're not — they're designed controls). |