Files
acdl/.ciagent/NORTH_STAR.md
T
Jon Chery 97560e3c88 docs(grill): v1.17 binding decisions applied — 12 fixes (NORTH_STAR reclassification + deck plan rewrites)
NORTH_STAR-CHANGE: reclassify 3 targets to Post-Pilot section (E-003);
move AI-Agent Intent Share to Future Horizons (E-004).

GRILL binding decisions applied:
- G-Q4: D-122 honesty sentence on slide 7 (Nova's AI = confidence-gated policy engine, not LLM)
- G-Q5/G-Q15: derived metrics annotated with N=0 caveat; ROI formula shown inline on slide 13
- G-Q6: NORTH_STAR targets reclassified (3 to Post-Pilot, honesty note added)
- G-Q8: stake line (18V+0 consumers) added to slide 1
- G-Q9: 4 filler benefit closes rewritten (slides 1, 4, 13, 16)
- G-Q10: slide 12 split into Zero-Touch Efficiency + Cost & ROI (deck now 18 slides)
- G-Q11: preempt added to slide 14 (deferrals are measurement infra, not autonomy)
- G-Q13: Act 3->4 transition rewritten ('how it works' is not 'proof it works')
- G-Q14: slide 9 benefit reframed from data plumbing to trust substrate
- G-Q16: slide 16 ask reframed as business decision (approve pilot + ledger build-out)

---ci---
project: acdl
phase: 0
milestone: v1.17
status: grill
---/ci---
2026-08-04 19:44:28 +00:00

11 KiB
Raw Blame History

NORTH_STAR — Nova

Status: Draft (pending interactive GRILL → final) Milestone: v1.17 — Strategic Direction, Leadership Metrics & Unified Story Owner: Product Owner Purpose: Durable strategic intent. Read by CIAgent in every future /ci-run so the platform's direction survives across milestones. This is NOT a status document (that's PROJECT.md) and NOT an engineering architecture (that's the telemetry reference in RESEARCH.md/ ARCHITECTURE.md). It is the PO's committed direction: what we're building toward, what we refuse to build, and how we'll know we won.


Vision

Infrastructure operations become invisible. Every environment provisioned, every incident healed, every risk remediated — by an autonomous system whose trustworthiness is provable, not promised. Human attestation remains required at stage gates — QA signs off for production, SRE greenlights based on operational readiness — but the operator is never in the loop of normal operations.

Nova is the autonomous infrastructure layer that lets product teams ship without engaging an operator, and lets executives trust the AI not because it never fails but because every decision is captured, scored, and accountable.


Strategic Objectives (4)

1. Demonstrate production-grade zero-touch operations. Nova must run real customer estates with no human in the loop of normal operations — autonomy as the default, not the demo. Stage-gate attestation (QA for production, SRE for operational readiness) remains human by design; operational escalations (AI confidence too low to proceed) are the failure mode we drive toward zero. Everything else collapses if autonomy isn't real.

2. Establish provable trust in AI decisions. Build the audit substrate — Decision Ledger, confidence scoring, circuit breakers, blast-radius controls — that turns "autonomous" from a marketing claim into a defensible one. Trust is the moat. Features can be copied; an immutable, queryable decision history cannot.

3. Deliver compounding, quantifiable ROI for customers. Each quarter on Nova must reduce cloud spend, free engineering hours, and avoid downtime measurably. If the CFO can't point to a number that improves quarter-over-quarter, Nova fails its commercial test, regardless of how clever the AI is.

4. Become the default substrate for agentic infrastructure consumption. AI agents are already becoming the largest consumers of cloud infrastructure. Nova must be the platform through which those agents declare, deploy, and verify infrastructure — not a vendor scrambling into that market two quarters late.


Anti-Goals (5 — what Nova is fundamentally NOT)

  1. Not a Terraform, Kubernetes, or hyperscaler competitor. We orchestrate them. Replacing them is the most expensive possible distraction from the value we create.
  2. Not a general-purpose AI agent platform. We are purpose-built for infrastructure operations. Breadth here produces shallow tools; depth here wins the category.
  3. Not a system that removes humans from accountability. Only from operations. Every AI decision lands in an immutable ledger. Every stage-gate promotion (qa/prod/dr) requires a human attestation recorded with approver identity, separation-of-duties check, and the 8-concern evidence matrix. The absence of an operator is never the absence of a record.
  4. Not for legacy, untagged, or freeform infrastructure. Nova requires Terraform-managed, policy-aligned, fully-tagged inputs. We optimize for the disciplined 95%, not the chaotic 5%.
  5. Not sold to operators. Nova is sold to leadership on outcomes — cost, velocity, risk. Selling to operators inverts the incentive and breaks the autonomy thesis.

Non-Goals (v1.17 milestone scope — deferred work, not permanent boundaries)

Anti-Goals are what Nova fundamentally is not. Non-Goals are what we will not do this milestone — deferred work, not permanent boundaries. Each Non-Goal cites the controlling decision ID.

  1. Live AWS re-provisioning (deferred — D-096). Metrics that require live infrastructure ship as placeholder PowerBI views with documented schemas.
  2. Onboarding auto-grant (deferred — D-113/D-114/D-119). Only the request-path metric is grounded; the requested→granted funnel is a placeholder.
  3. ML anomaly-forecasting / predictive remediation (no emitter today). The Predictive-vs-Reactive metric ships as a placeholder.
  4. Drift detection scheduled job (deferred — D-096 + no scheduler). Drift metrics ship as placeholders.
  5. Live cost CUR reconciliation (deferred — D-096). Pre-apply Infracost estimates are grounded; actual-spend reconciliation is a placeholder.
  6. S3 Object Lock / JWS tamper-evident ledger (deferred — D-083). The Decision Ledger uses a local SQLite hash-chain this milestone; the Object-Lock/JWS build-out is a future milestone.
  7. Multi-cloud support (Azure/GCP/K8s). Nova is AWS-only this milestone.

1218 Month Targets

Targets are committed, not aspirational. Each is a number a board member can repeat back to us. The grounding column records whether the metric is measurable this milestone, and if not, what blocks it.

Honesty note (GRILL G-Q6 binding): Nova has 0 consumer adoption today (PROJECT.md:495). Three targets (Touchless Resolution, Human Escalation, AI Decision Accuracy) are scoped "across production estates" — the measurement pipeline is grounded this milestone, but the denominator is zero until a pilot estate activates. These targets are reclassified as Post-Pilot (the pipeline works; the numbers fill when consumers exist). This is the same honesty model as Cloud Spend Reduction (partial: pipeline grounded, actuals deferred).

Current-milestone targets (grounded or derived this milestone)

Domain Target Grounding (v1.17) Note
MTTR (p95) < 60 seconds grounded (platform-run MTTR) apply.failed → successful retry; infra-incident MTTR deferred (no incident detection)
Cloud Spend Reduction ≥ 25% on pilot estates vs. 12-month pre-Nova baseline partial pre-apply estimate grounded (Infracost); actual-spend deferred (D-096 CUR)
L1 / L2 Ops Hours Avoided ≥ 70% of pre-Nova FTE allocation derived formula over run count × manual baseline (computed on N internal runs; production-denominator activates post-pilot)
Platform ROI ≥ 250% measured annually derived formula (labor savings + cloud savings + avoided downtime) ÷ platform op cost (computed on N internal runs; production-denominator activates post-pilot)
Decision Ledger Coverage 100% of AI actions with backfilled outcome grounded (this milestone builds it) outbox_writer.py → SQLite hash-chain
Attestation Coverage 100% of prod/dr promotions attested by a human grounded hitl_gates.py + outbox approver_* attributes; separation-of-duties on prod

Post-Pilot targets (pipeline grounded this milestone; denominator activates when a pilot estate runs)

Domain Target Grounding (v1.17) Note
Touchless Resolution Rate ≥ 99% across production estates partial (pipeline grounded; denominator = 0 today) runs completing without operational HITL block ÷ total runs (attestation gates excluded); activates post-pilot
Human Escalation Frequency < 0.1% of platform actions partial (pipeline grounded; denominator = 0 today) operational HITL blocks only (confidence-driven); attestation sign-offs excluded; activates post-pilot
AI Decision Accuracy ≥ 99.5% (no rollback, no follow-up incident within 5 min of action) partial (pipeline grounded; denominator = 0 today) decisions not followed by apply.failed/incident within 5min; activates post-pilot

Deferred targets (measurement requires future systems)

Domain Target Grounding (v1.17) Note
Predictive vs. Reactive Ratio ≥ 3 : 1 (prevention dominates reaction) deferred requires ML forecasting service (future emitter)
Drift Auto-Reversal Rate ≥ 95% within one detection cycle deferred requires drift detection (D-096 + scheduler)

Committed targets whose measurement is deferred remain committed — the target is the destination; the metric is the odometer, and some odometers aren't built yet. Each deferred metric ships as a placeholder PowerBI view + a definition-of-success doc recording the dependency. Post-Pilot targets are committed targets whose measurement pipeline is grounded this milestone; the numbers activate when a pilot estate runs.

Future Horizons (strategic direction, not committed targets)

Domain Aspiration Note
AI-Agent Intent Share ≥ 40% of total intent volume originated by non-human consumers Strategic Objective #4 direction. No backing requirement, no placeholder view, no emitter today. Moves to a committed target when agentic consumption is real.

Success Criteria (v1.17 — what constitutes success for THIS milestone)

Distinct from the 1218mo targets: those are the destination. These are the milestone's exit criteria.

v1.17 is a success if:

  1. Decision Ledger emits ai.decision.made for 100% of platform runs with outcome backfill, AND attestation.recorded events for 100% of qa/prod/dr promotions (event completeness — all 3 gates captured; grounded in outbox_writer.py → SQLite hash-chain; honors D-083). The Attestation Coverage metric (target 100%) measures prod/dr promotions specifically — see REQ-194.
  2. docs/METRICS.md catalogs every executive KPI with a grounded / derived / deferred status, a source file or decision ID, and a per-KPI definition-of-success doc in docs/metrics/.
  3. The PowerBI export produces all fact/dimension views + 8 empty placeholder views for deferred metrics (with documented schemas ready to fill when their blocking decisions lift).
  4. The unified narrative deck ships with the x3 arc (Problem→Vision→How→Proof→Roadmap) at deck + slide level, per-slide benefit callouts, and fluid transitions; both old decks retired.
  5. NORTH_STAR.md is wired into CIAgent context-loading so every future /ci-run reads it.
  6. CAP-023 (metrics collector) + CAP-024 (deck structure) pass in the regression gate.

What "won" looks like

By month 18, Nova is the layer enterprise leadership points to when they say "we don't have an infrastructure ops team anymore, and the audit trail is stronger than it ever was" — and it is the default substrate their AI engineering teams reach for first when an agent needs to deploy.


Relationship to v1.17 engineering

  • Pillar A (this file): strategic direction — durable, PO-authored.
  • Pillar B (engineering): the telemetry reference architecture (adapted from the PO's technical-direction input) lives in RESEARCH.md/ARCHITECTURE.md. It is the how; this file is the why.
  • Pillar C (story): the unified narrative deck proves Pillars A+B to leadership. The deck's Proof section cites grounded metrics; its Roadmap section cites deferred targets honestly.