25ddc894c2
9 requirements complete (REQ-254..262): - P1: theme-css — section padding + overflow + image rules + title chrome + spacing tightening (REQ-254,255,256) - P2: render-scripts — delete render_deck.sh, pin CLI versions, 2x scale + transparent bg (REQ-257,258) - P3: mermaid-relayout — telemetry TB + platform-pipeline 4-node TB, re-rendered 2x transparent (REQ-259,260) - P4: deck-content — split slides 3+8 (18->20 main), trim 8 overflowing slides, remove redundant header (REQ-261) - P5: render-and-test — re-render HTML+PPTX, add 9 layout/aspect- ratio/theme-structural tests (REQ-262) - P6: final review + audit + ship (this commit) Final review fixes: source .md + talking-points re-synced to 20-slide structure; ![h:480 class:tall] directives applied; README stale references updated; CSS trailing newline added. Root cause: nova-sp-theme.css had zero section padding (declared /* @theme nova-sp */ as a comment, not the @theme directive; did not @import Marp default theme). Combined with overflow:hidden, blunt img max-height:320px, header+footer chrome on every slide, and two P5 diagrams with extreme aspect ratios (13.52x and 0.63x), 8 of 19 slides overflowed. NOT a P5 regression — theme CSS byte-identical P3->P5; P5 denser content made pre-existing flaws visible. Tags on v1.21.x line (v1.21.0 P0 -> v1.21.6 P6 final = milestone release). 32 slide tests pass (23 original + 9 new). 94 key-file tests pass. Pipeline check exit 0. ---ci--- project: acdl phase: 6 milestone: v1.22 status: complete phase_role: final requirements: covered: [REQ-254,REQ-255,REQ-256,REQ-257,REQ-258,REQ-259,REQ-260,REQ-261,REQ-262] partial: [] ---/ci---
747 lines
35 KiB
Markdown
747 lines
35 KiB
Markdown
# Nova — The Autonomous Cloud Delivery Platform
|
||
|
||
> **Source of truth** (Step 1 of the 4-step deck process).
|
||
> Unified narrative deck. 4-beat arc: Problem → Solution → Proof →
|
||
> Roadmap + Ask. x3 structure at deck level (opening = the problem + the
|
||
> arc, body = tell them, closing = recap + ask) AND per slide (opens with
|
||
> what it covers, delivers, closes with a benefit callout written for a
|
||
> tech-leadership audience).
|
||
>
|
||
> **Honesty model:** every metric cited is grounded (cites a source),
|
||
> derived (documented formula), or deferred (cites the blocking work).
|
||
> No fabricated numbers. Internal provenance (decision IDs, requirement
|
||
> IDs, internal file paths) is kept out of the audience-facing slides —
|
||
> those live in the appendix and the `.ciagent/` files only.
|
||
>
|
||
> v1.21 — Deck Refinement & Pipeline Hardening
|
||
|
||
---
|
||
|
||
## Slide 1 — The Problem
|
||
|
||
**Product teams now own their cloud infrastructure — but ownership without
|
||
discipline is destroying value.**
|
||
|
||
The broad shift to "you build it, you run it" put Terraform into the hands
|
||
of product teams. The intention was right: teams that own their stack ship
|
||
faster. The reality is that infrastructure-as-code is a different craft
|
||
from software development, and the engineering standards that teams apply
|
||
to application code are rarely applied to the infrastructure that carries
|
||
it.
|
||
|
||
- **No lifecycle planning.** Resources are authored for creation, not for
|
||
patching, decommissioning, or rollback. When a change is needed, the
|
||
change is destructive — because no one planned the lifecycle.
|
||
- **Proactive scanning is not part of authoring.** In a year where
|
||
AI-frontier models discover and exploit zero-day vulnerabilities at a
|
||
rapid pace, teams cannot keep up by reacting. Infrastructure modules
|
||
must be scanned as code and at runtime, post-deployment — and remediated
|
||
at the pace the threat moves, not the pace a sprint allows.
|
||
- **Bandwidth gaps in infrastructure operations.** An unusual amount of
|
||
time is spent on remediation, the push for innovation does not pause,
|
||
and the result is that operational work is chronically under-resourced.
|
||
Gaps open. Detections are missed. Incidents grow.
|
||
- **Tribal knowledge and the rockstar-operator problem.** Operations
|
||
depend on a handful of administrators who hold the infrastructure in
|
||
their heads. When they leave, the knowledge leaves with them. The
|
||
platform should encode the discipline, not the person.
|
||
|
||
Every hour a developer spends writing, deploying, fixing, or remediating
|
||
infrastructure is an hour not spent releasing features to production and
|
||
generating value.
|
||
|
||
> **Benefit:** the rest of this deck shows the answer — an autonomous
|
||
> cloud delivery platform that encodes infrastructure discipline as
|
||
> policy, scans proactively, remediates rapidly, and makes operations
|
||
> visible to leadership rather than hidden in tribal knowledge.
|
||
|
||
> **Speaker notes:** Do not frame this as "humans are the problem." The
|
||
> problem is that ownership was granted without the discipline, tooling,
|
||
> and lifecycle planning that infrastructure requires. The operator is
|
||
> not the bottleneck because operators exist — the bottleneck is that
|
||
> operations depend on a few individuals instead of an encoded system.
|
||
|
||
> **Transition:** "Here is the destination Nova is building toward."
|
||
|
||
---
|
||
|
||
## Slide 2 — Nova's Vision
|
||
|
||
**Infrastructure operations become visible. Every environment provisioned,
|
||
every incident healed, every risk remediated — by an autonomous system
|
||
whose trustworthiness is provable, not promised. Human attestation remains
|
||
required at stage gates; the operator is never in the loop of normal
|
||
operations.**
|
||
|
||
- **Visibility is the recurring theme.** Security posture, remediation
|
||
velocity, reliability, and lead time are surfaced as queryable signals —
|
||
not hidden in a person's head or a Slack thread.
|
||
- **Provable, not promised.** Trust is established by deterministic
|
||
scripts that calculate a score and gate the action. The platform
|
||
functions without AI. "AI decisions" are really automated decisions.
|
||
- **Autonomy in operations, human at stage gates.** QA signs off for
|
||
production; SRE greenlights based on operational readiness. The
|
||
absence of an operator in the loop is never the absence of a record.
|
||
|
||
> **Benefit:** the destination is autonomous operations with provable
|
||
> trust — security, remediation velocity, reliability, and lead time made
|
||
> visible to leadership, not promised to them.
|
||
|
||
> **Speaker notes:** "Visible" is the operative word. The vision is not
|
||
> just that operations run without an operator — it is that operations
|
||
> become observable, queryable, and accountable. That is what makes the
|
||
> trust defensible.
|
||
|
||
> **Transition:** "The vision is ambitious — here are the strategic
|
||
> objectives that make it concrete, and the anti-goals that keep it
|
||
> focused."
|
||
|
||
---
|
||
|
||
## Slide 3 — Strategic Objectives
|
||
|
||
**4 Strategic Objectives:**
|
||
1. **Demonstrate production-grade zero-touch operations** — autonomy as
|
||
the default, not the demo. Stage-gate attestation (QA, SRE) remains
|
||
human by design.
|
||
2. **Establish provable trust in automated decisions** — deterministic
|
||
scripts calculate a score; a band outcome gates the action. The
|
||
platform functions without AI. The Decision Ledger, confidence
|
||
scoring, circuit breakers, and blast-radius controls make
|
||
"autonomous" a defensible claim, not a marketing one.
|
||
3. **Deliver compounding, quantifiable ROI** — measured on four CTO-grade
|
||
metrics, all flowing into PowerBI:
|
||
- **Lead Time** (PR → Production) — downward trend.
|
||
- **Infrastructure Vulnerability Count** — downward trend
|
||
(proactive scanning keeps up with the AI-era 0-day pace).
|
||
- **MTTR** — for platform-detected and platform-remediated incidents.
|
||
- **Cloud Spend Reduction** — on pilot estates vs. the pre-Nova
|
||
baseline.
|
||
4. **Integrate with externally owned development platforms — regardless
|
||
of source.** Nova integrates with externally owned PDLC, SDLC,
|
||
Agentic, and Citizen Developer platforms. Nova provides skills and
|
||
MCP endpoints that help the developer or AI agent make their
|
||
application production-grade. Regardless of the source, all intents
|
||
to deploy to production go through the same rigorous controls,
|
||
quality gates, attestation, and evidence stream.
|
||
|
||
> **Benefit:** the scope is explicit — Nova governs infrastructure and
|
||
> delivery, integrates with any upstream source through one validated
|
||
> contract, and measures success on four metrics a CTO can repeat back.
|
||
|
||
> **Speaker notes:** Objective #2 is the one to land carefully: trust is
|
||
> established by deterministic scoring, not by an LLM. The platform
|
||
> functions without AI.
|
||
|
||
> **Transition:** "The objectives are concrete — here is what Nova is
|
||
> NOT, to keep it focused."
|
||
|
||
---
|
||
|
||
## Slide 4 — Anti-Goals (What Nova Is NOT)
|
||
|
||
1. Not a general-purpose AI agent platform.
|
||
2. Not a system that removes humans from accountability — only from
|
||
normal operations.
|
||
3. Not an upstream development platform (no product backlogs, IDE, code
|
||
authorship).
|
||
4. Not a replacement for the Product Development Lifecycle (PDLC).
|
||
|
||
> **Benefit:** the boundaries are explicit — Nova is purpose-built for
|
||
> infrastructure operations and delivery, not a general-purpose AI agent
|
||
> or an upstream development platform.
|
||
|
||
> **Speaker notes:** Anti-goals #3 and #4 protect the scope boundary —
|
||
> Nova will not become an IDE or a product-planning tool.
|
||
|
||
> **Transition:** "The scope boundary is explicit — here is exactly
|
||
> where Nova sits relative to the product development lifecycle."
|
||
|
||
---
|
||
|
||
## Slide 5 — Scope: Downstream of PDLC
|
||
|
||
**Nova governs infrastructure and delivery. The PDLC is upstream — Nova
|
||
never penetrates it. Integration is through one validated contract.**
|
||
|
||
- **The PDLC is upstream:** product backlog, code authorship (AI agent,
|
||
IDE, agentic SDLC), sprint planning, application business logic.
|
||
- **Nova is downstream:** contract ingestion → submission-readiness gate
|
||
→ policy enforcement → cloud resource lifecycle → environment
|
||
progression (dev → qa → prod → dr) → immutable audit + attestation.
|
||
- **The integration point is one contract.** The citizen developer's AI
|
||
coding agent, an upstream agentic SDLC platform, or any development
|
||
platform may all produce submissions — the source does not matter
|
||
because all are subject to the same compliance standards.
|
||
- **Nova validates the submission, not the author.** The audit trail is
|
||
the same; the policy envelope is the same; the evidence stream is the
|
||
same.
|
||
|
||
> **Benefit:** a clean scope boundary — Nova is purpose-built for
|
||
> infrastructure operations and integrates with any upstream source
|
||
> through one validated contract, so the platform team's surface area
|
||
> stays bounded.
|
||
|
||
> **Speaker notes:** This slide protects the scope. The moment Nova
|
||
> starts owning the PDLC, it loses focus. The contract boundary is what
|
||
> keeps Nova deep on infrastructure and delivery rather than shallow on
|
||
> everything.
|
||
|
||
> **Transition:** "With the scope clear, here is who owns what across the
|
||
> delivery lifecycle."
|
||
|
||
---
|
||
|
||
## Slide 6 — RACI: Who Owns What
|
||
|
||
**Four roles, one matrix — the citizen developer owns FRs + UAT, the
|
||
platform owns NFRs + infra, quality engineering owns the gate evidence,
|
||
and SRE owns operational readiness.**
|
||
|
||
| Work Category | Citizen Dev | Platform | Quality Eng | SRE |
|
||
|---|---|---|---|---|
|
||
| Functional Requirements | **R/A** | C | I | I |
|
||
| User Acceptance Testing | **R/A** | C | I | I |
|
||
| Non-Functional Requirements | I | **R/A** | C | C |
|
||
| Infrastructure (cloud, state, IAM) | I | **R/A** | I | C |
|
||
| QA (policy, confidence, schema) | C | R | **R/A** | I |
|
||
| Production deployment to cloud | I | **R/A** | C | C |
|
||
| Quality attestation (QA sign-off) | **A** | R | **R** | I |
|
||
| Production readiness (SRE sign-off) | **A** | R | C | **R** |
|
||
|
||
**R** = Responsible · **A** = Accountable (sign-off) · **C** = Consulted · **I** = Informed.
|
||
|
||
- **Compliance-standard equivalence:** FRs + UAT may come from any
|
||
upstream source (AI agent, agentic SDLC, dev platform) — all pass the
|
||
same submission-readiness gate.
|
||
- **Production readiness is co-owned:** the platform runs the
|
||
attestations agentically; the citizen developer authorizes the
|
||
promotion at the stage gate.
|
||
|
||
> **Benefit:** every party knows what they bring, what the platform
|
||
> provides, what quality engineering guards, and where SRE signs off —
|
||
> accountability is explicit, never diffuse.
|
||
|
||
> **Speaker notes:** Quality attestation is now owned by Quality
|
||
> Engineering (not the Platform), and Production readiness is owned by
|
||
> SRE. The Platform runs the checks agentically but is never the
|
||
> Accountable party for the gate — that separation keeps the platform
|
||
> honest.
|
||
|
||
> **Transition:** "With ownership clear, here is how the pipeline
|
||
> enforces it."
|
||
|
||
---
|
||
|
||
## Slide 7 — The Platform Pipeline
|
||
|
||
**How intent becomes verified infrastructure — with fail-fast policy
|
||
scanning before the plan and runtime scanning after it.**
|
||
|
||
```mermaid
|
||
graph LR
|
||
A[Contract] --> B[Resolver]
|
||
B --> C[Adapter]
|
||
C --> D["Checkov (static code)"]
|
||
D --> E[Terraform Plan]
|
||
E --> F["Wiz (on plan)"]
|
||
F --> G[Confidence Signal]
|
||
G --> H{Stage Gate}
|
||
H -->|dev: autonomous| I[Apply]
|
||
H -->|qa/prod/dr: attested| I
|
||
I --> J[Evidence + Ledger]
|
||
```
|
||
|
||
- **Contract → resolver → adapter → Checkov on static code (before the
|
||
plan) → terraform plan → Wiz on the plan → confidence signal → stage
|
||
gate → apply → evidence + ledger.**
|
||
- **Fail-fast, quick feedback.** Checkov runs on the authored Terraform
|
||
code before `terraform plan` so developers get immediate policy
|
||
feedback, not a delayed plan-stage failure.
|
||
- **Wiz on the plan when configured; Checkov as a drop-in otherwise.**
|
||
Wiz scans the terraform plan output. When Wiz credentials are not
|
||
available, Checkov runs against the plan as a drop-in replacement. Wiz
|
||
and Checkov are never both run on the plan.
|
||
- **Dev is autonomous** (no stage gate); **qa/prod/dr require human
|
||
attestation** (QA for quality, SRE for production readiness).
|
||
|
||
> **Benefit:** the pipeline gives developers fast, deterministic feedback
|
||
> on policy at authoring time and gives the platform a runtime scan on the
|
||
> resolved plan — two layers of scanning, zero operator involvement in
|
||
> normal operations.
|
||
|
||
> **Speaker notes:** The two-stage scan is the key design: static code
|
||
> scanning catches policy violations before the cost of a plan; runtime
|
||
> plan scanning catches what the static code cannot (resolved values,
|
||
cross-resource issues). The platform picks the runtime scanner based on
|
||
configuration — never both, to avoid duplicate noise.
|
||
|
||
> **Transition:** "The pipeline produces decisions — here is how every
|
||
> decision is captured and made accountable."
|
||
|
||
---
|
||
|
||
## Slide 8 — The Decision Ledger
|
||
|
||
**Every automated decision is captured, immutable, queryable — and
|
||
accountable.**
|
||
|
||
- **What is captured:** every action the platform takes — the chosen
|
||
action, the confidence score, the alternatives considered, whether a
|
||
human overrode it, and the outcome (backfilled once the apply
|
||
completes). Every stage-gate attestation (QA sign-off, SRE
|
||
production-readiness sign-off) is captured with approver identity and
|
||
the evidence that was presented.
|
||
- **"AI decisions" are really automated decisions.** The decisions are
|
||
made by deterministic scripts that calculate a score and a band; the
|
||
platform functions without AI. The ledger captures the real decision
|
||
path — not a fabricated "AI agent." When an LLM planner is added later,
|
||
it will emit richer alternatives without breaking the schema.
|
||
- **The value is accountability, not the storage engine.** The ledger is
|
||
an append-only, tamper-evident record. The point is not which database
|
||
it lives in — the point is that every decision is queryable for
|
||
auditing, traceable to an outcome, and impossible to rewrite after the
|
||
fact.
|
||
|
||
> **Benefit:** "autonomous" is defensible because every decision the
|
||
> platform makes is immutable, queryable, and accountable — and the
|
||
> audience knows exactly what "automated" means here: deterministic
|
||
> scoring, not a black-box LLM.
|
||
|
||
> **Speaker notes:** Do not dwell on the storage substrate. The audience
|
||
> cares that the ledger is append-only, queryable, and tied to outcomes —
|
||
> not that it is a hash-chain in a SQLite file. The D-122 honesty point
|
||
> is restated without the decision ID: the platform's decisions are
|
||
> deterministic; the ledger captures that real path.
|
||
|
||
> **Transition:** "Decisions are captured — here is how stage-gate
|
||
> attestation keeps humans in accountability."
|
||
|
||
---
|
||
|
||
## Slide 9 — Attestation Matrix: QA
|
||
|
||
**The designed controls that keep humans at stage gates — QA concerns,
|
||
freshness-validated.**
|
||
|
||
| Concern | Env | Freshness | Description |
|
||
|---------|-----|-----------|-------------|
|
||
| Functional correctness | qa | 24h | The application behaves as specified; evidence accepted from the consumer's UAT. |
|
||
| Performance baseline | qa | 7d | The deployment meets its performance envelope vs. the agreed baseline. |
|
||
| Security posture | qa | 24h | The deployment's security findings have been reviewed and accepted. |
|
||
|
||
- Each concern has a freshness window — evidence older than the window
|
||
does not satisfy the gate.
|
||
- Concerns that are offline-testable run for real; concerns that require
|
||
external evidence accept signed artifacts.
|
||
|
||
> **Benefit:** QA signs off on quality before any promotion — the gate
|
||
> is explicit, not implicit.
|
||
|
||
> **Speaker notes:** The matrix is not a rubber stamp. Each concern has a
|
||
> freshness window and a plain-language description of what is being
|
||
> attested. The "operator-supplied" label from the prior deck was
|
||
> dropped — every concern now has a plain-language description.
|
||
|
||
> **Transition:** "QA is half the matrix — here are the production and
|
||
> DR controls."
|
||
|
||
---
|
||
|
||
## Slide 10 — Attestation Matrix: Prod/DR
|
||
|
||
**Production and DR controls — operational readiness, resilience, and
|
||
disaster recovery.**
|
||
|
||
| Concern | Env | Freshness | Description |
|
||
|---------|-----|-----------|-------------|
|
||
| Operational readiness | prod | 30d | SRE confirms the deployment is operable: runbooks, dashboards, on-call coverage. |
|
||
| Incident response | prod | 90d | The on-call path has been exercised; the deployment has a working incident-response plan. |
|
||
| Capacity & cost | prod | 30d | Capacity headroom and monthly cost are within the agreed envelope. |
|
||
| Resilience: DR drill | prod | 180d | A DR drill has been run and the deployment recovered within the RTO. |
|
||
| Resilience: chaos | prod | 90d | A chaos exercise has been run and the deployment absorbed the failure. |
|
||
| Resilience: backup | prod | 30d | Backups are restorable and have been tested within the freshness window. |
|
||
| DR region deploy | dr | 180d | The DR region can be deployed and the deployment is reachable from it. |
|
||
|
||
- Each concern has a freshness window — evidence older than the window
|
||
does not satisfy the gate.
|
||
- **Separation-of-duties on prod:** the approver cannot be the same
|
||
person who built the deployment.
|
||
|
||
> **Benefit:** the gate model is explicit — autonomy in operations,
|
||
> human in accountability, by design. The matrix is what makes autonomous
|
||
> operations safe enough to trust in production.
|
||
|
||
> **Speaker notes:** The prod/DR rows are the operational-readiness and
|
||
> resilience gates — SRE signs off on operability, incident response,
|
||
> capacity, and the three resilience checks (DR drill, chaos, backup).
|
||
> Separation-of-duties on prod is the rule that keeps the gate honest:
|
||
> the approver cannot be the same person who built the deployment.
|
||
|
||
> **Transition:** "You've seen how Nova works — the pipeline, the ledger,
|
||
> the attestation gates. Here is how Nova instruments itself so that
|
||
> every claim in this deck is traceable to a real signal."
|
||
|
||
---
|
||
|
||
## Slide 11 — Telemetry & Live Ops
|
||
|
||
**Every metric in this deck is traceable to a real emitted signal — and
|
||
the live-ops dashboard makes operations visible in PowerBI.**
|
||
|
||
```mermaid
|
||
graph TB
|
||
A[Platform components] --> B[CloudEvents envelope]
|
||
B --> C[Event log]
|
||
B --> D[Decision ledger]
|
||
B --> E[Run records]
|
||
C --> F[Collector]
|
||
D --> F
|
||
E --> F
|
||
F --> G[Cold store]
|
||
G --> H[PowerBI views]
|
||
H --> I[Live ops dashboard]
|
||
```
|
||
|
||
- **Platform components emit a CloudEvents envelope** → event log,
|
||
decision ledger, and run records → collector → cold store → PowerBI
|
||
views → **live ops dashboard.**
|
||
- **The live ops dashboard (PowerBI)** surfaces the four CTO-grade
|
||
metrics — Lead Time, Infrastructure Vulnerability Count, MTTR, Cloud
|
||
Spend — alongside the trust metrics (Decision Ledger coverage,
|
||
Attestation coverage) and the efficiency metrics (touchless
|
||
resolution, escalation frequency).
|
||
- **The architecture is deliberately minimal.** Nova-native envelopes;
|
||
no Kafka, no Prometheus, no ClickHouse. The cold store is sufficient
|
||
for batch and historical analysis; the live-ops surface is built in
|
||
PowerBI on top of the exported views.
|
||
- **Every number in the Proof slides is traceable to a signal.** When a
|
||
CFO asks "where does this number come from?", the answer is a query
|
||
against the cold store, not a Slack thread.
|
||
|
||
> **Benefit:** the architecture is the trust substrate — leadership sees
|
||
> the same numbers the platform produces, in PowerBI, with full
|
||
> traceability to the emitted signal. Operations become visible.
|
||
|
||
> **Speaker notes:** The value is not the plumbing — it is that the
|
||
> platform's metrics surface in a tool leadership already uses (PowerBI),
|
||
> and every number is traceable. The live-ops dashboard is where the
|
||
> "infrastructure operations become visible" theme lands concretely.
|
||
|
||
> **Transition:** "The architecture is sound — here is the measured
|
||
> proof."
|
||
|
||
---
|
||
|
||
## Slide 12 — Decision Ledger + Attestation Coverage
|
||
|
||
**By design, no change reaches production without a ledger entry and a
|
||
human attestation — both queryable for auditing, with full
|
||
traceability.**
|
||
|
||
- **Decision Ledger coverage: 100%.** Every platform run emits a
|
||
decision record with outcome backfill. No automated decision is ever
|
||
lost.
|
||
- **Attestation coverage: 100%.** Every prod/dr promotion is attested by
|
||
a human — QA for quality, SRE for production readiness — recorded with
|
||
approver identity, separation-of-duties check, and the evidence matrix.
|
||
- **No change to production without both.** The ledger entry and the
|
||
human attestation are mandatory, not optional. This is enforced by the
|
||
pipeline, not by policy.
|
||
- **Easily queried for auditing.** The ledger and the attestation
|
||
records are queryable by run, by environment, by approver, and by
|
||
outcome — the audit trail is a query, not a forensic exercise.
|
||
- **Full traceability.** A production change is traceable from the
|
||
contract that declared intent, through the policy scan, the confidence
|
||
score, the attestation, to the applied outcome. Nothing is opaque.
|
||
|
||
> **Benefit:** trust is provable — not a marketing claim, a queryable
|
||
> record. An auditor can answer "who approved this, when, on what
|
||
> evidence?" in one query; a CTO can answer "how many of last quarter's
|
||
> prod changes were touchless?" in one query.
|
||
|
||
> **Speaker notes:** The mandatory-by-design point is the one to land.
|
||
> The ledger + attestation are not a best-effort feature; they are a
|
||
> gate. No change reaches production without both. That is what makes
|
||
> the 100% numbers credible — they are enforced, not aspirational.
|
||
|
||
> **Transition:** "Trust is provable — here is the cost side of the ROI."
|
||
|
||
---
|
||
|
||
## Slide 13 — Cost & ROI
|
||
|
||
**The ROI formula and the cost estimates — grounded, with the production
|
||
denominator honestly flagged.**
|
||
|
||
- **Cost estimates are pre-apply and offline.** The platform reads the
|
||
terraform plan and estimates cost before anything is applied — so a
|
||
regression in cost is caught before the spend happens, not after.
|
||
- **The ROI formula:**
|
||
`Platform ROI = (FTE hours saved × blended rate + cloud savings + avoided downtime) ÷ platform op cost`
|
||
- **The four CTO-grade metrics (from Slide 3) are the ROI proof:**
|
||
Lead Time (PR → Prod), Infrastructure Vulnerability Count (trend), MTTR,
|
||
Cloud Spend Reduction. All flow into PowerBI.
|
||
- **Honest caveat:** the derived metrics are computed on internal runs
|
||
today; the production-denominator activates when a pilot estate runs.
|
||
The formula is grounded; the production numbers are not yet.
|
||
|
||
> **Benefit:** the ROI is not a black box — the formula is shown, the
|
||
> four metrics are committed, and the production-denominator caveat is
|
||
> stated up front. The CFO can see exactly what is real today and what
|
||
> activates with a pilot.
|
||
|
||
> **Speaker notes:** The formula is shown inline, not hidden. The
|
||
> "no fabrication" constraint in action: show the formula, show the
|
||
> caveat, do not pretend the production numbers exist.
|
||
|
||
> **Transition:** "The proof is grounded — here is what is honestly
|
||
> deferred, and why."
|
||
|
||
---
|
||
|
||
## Slide 14 — What's Deferred — and Why
|
||
|
||
**Honesty about what is not measured yet — and the blocking work for
|
||
each.**
|
||
|
||
To be clear: these deferrals are measurement infrastructure, not the
|
||
autonomy itself. The platform runs without an operator in the loop of
|
||
normal operations. What is deferred is the evidence pipeline for certain
|
||
metrics — not the autonomy.
|
||
|
||
| # | Deferred metric | Blocking work |
|
||
|---|-----------------|---------------|
|
||
| 1 | Live infrastructure health | Live AWS re-provisioning (currently torn down to a zero-cost steady state) |
|
||
| 2 | Live outbox write rate | Live AWS re-provisioning |
|
||
| 3 | Tamper-evident ledger checkpoints | Audit-ledger build-out (S3 Object Lock + signed checkpoints) |
|
||
| 4 | Onboarding funnel (requested → granted) | Auto-grant implementation |
|
||
| 5 | Drift auto-reversal | Drift-detection scheduler (not yet built) |
|
||
| 6 | Live cost reconciliation | Live AWS re-provisioning + actual-spend feed |
|
||
| 7 | SLA / unplanned downtime | Live AWS re-provisioning |
|
||
| 8 | Predictive vs reactive ratio | ML anomaly-forecasting service (not yet built) |
|
||
|
||
> **Benefit:** the boundaries are explicit — what Nova measures today,
|
||
> and exactly what blocks the rest. The autonomy is real; the measurement
|
||
> gaps are documented with the work that unblocks each one.
|
||
|
||
> **Speaker notes:** The preempt is critical: these deferrals are
|
||
> measurement infrastructure, not autonomy. The platform runs without an
|
||
> operator in the loop. What is deferred is the evidence pipeline for
|
||
> live-infra health, drift, predictive remediation — not the autonomy
|
||
> itself.
|
||
|
||
> **Transition:** "The proof is honest — here is the roadmap from here to
|
||
> the targets."
|
||
|
||
---
|
||
|
||
## Slide 15 — Roadmap to the North Star
|
||
|
||
**The path from the grounded metrics to the 12–18 month targets — each
|
||
deferred metric has an unblock path and a candidate milestone.**
|
||
|
||
| Timeframe | Work | Unblocks |
|
||
|-----------|------|----------|
|
||
| Near-term | Live AWS re-provisioning | Live infra health, live outbox write rate, live cost reconciliation, SLA |
|
||
| Near-term | Auto-grant implementation | Onboarding funnel (requested → granted) |
|
||
| Mid-term | Drift-detection scheduler | Drift auto-reversal |
|
||
| Mid-term | Audit-ledger build-out (Object Lock + signed checkpoints) | Tamper-evident ledger checkpoints |
|
||
| Mid-term | Hot-path activation (live-ops dashboard goes from batch to near-real-time) | Live-ops dashboard freshness |
|
||
| Longer-term | ML anomaly-forecasting service | Predictive vs reactive ratio |
|
||
|
||
- Each deferred metric has a specific unblock requirement and a
|
||
candidate future milestone.
|
||
- Re-evaluation triggers: each blocking piece of work lifts on its own
|
||
schedule; the metrics layer evolves as each one lands.
|
||
|
||
> **Benefit:** every deferred metric has an unblock path — nothing is
|
||
> hand-waved; everything has a plan and a timeframe.
|
||
|
||
> **Speaker notes:** This is the bridge from "honestly deferred" to
|
||
> "here is how we get there." The roadmap uses timeframes, not status —
|
||
> most of it is not implemented yet, so a status column would be noise.
|
||
|
||
> **Transition:** "The unblock path is clear — here is the 12-month
|
||
> product arc."
|
||
|
||
---
|
||
|
||
## Slide 16 — 12-Month Product Roadmap
|
||
|
||
**The product arc from pilot activation to integration — four quarters,
|
||
four outcomes.**
|
||
|
||
| Quarter | Theme | Board-level outcome |
|
||
|---------|-------|---------------------|
|
||
| **Q1** | Pilot Activation | Nova runs a real customer estate end-to-end, autonomously, with a measurable zero-touch rate. |
|
||
| **Q2** | Provable Trust | Every automated decision lands in a tamper-evident ledger; the CFO sees real cloud-spend reconciliation. |
|
||
| **Q3** | Compounding ROI | Quarter-over-quarter cloud spend drops; drift is detected and reversed without a human. |
|
||
| **Q4** | Integration & Predictive | AI agents deploy through Nova by default; the ML anomaly-forecasting service goes live. |
|
||
|
||
Grounded in the four strategic objectives (autonomy, provable trust, ROI,
|
||
integration) and the deferred-metric unblock paths.
|
||
|
||
> **Benefit:** the 12-month product arc — each quarter activates a
|
||
> strategic objective and its corresponding board-level metric, from
|
||
> pilot activation through integration leadership.
|
||
|
||
> **Speaker notes:** The roadmap is organized by product outcome, not
|
||
> by technical milestone. Each quarter activates one strategic
|
||
> objective from the North Star.
|
||
|
||
> **Transition:** "Here is the quarter-by-quarter detail."
|
||
|
||
---
|
||
|
||
## Slide 17 — Quarter-by-Quarter Outcomes
|
||
|
||
| Quarter | Product theme | Key deliverable | Target metric | Grounding |
|
||
|---------|---------------|-----------------|---------------|-----------|
|
||
| **Q1** | Pilot Activation | Re-provision live AWS; activate first pilot estate; onboarding auto-grant | Touchless ≥ 99% · Escalation < 0.1% · Accuracy ≥ 99.5% | Objective #1 — autonomy as the default |
|
||
| **Q2** | Provable Trust | Tamper-evident ledger (Object Lock + signed checkpoints); daily checkpoints; live cost reconciliation | Decision Ledger Coverage 100% · Cost Savings ≥ 25% | Objective #2 — trust is the moat |
|
||
| **Q3** | Compounding ROI + Drift | Drift-detection scheduler; auto-reversal; pre-apply → actual-spend reconciliation on the pilot estate | Drift Auto-Reversal ≥ 95% · Spend Reduction ≥ 25% | Objective #3 — CFO-pointable numbers |
|
||
| **Q4** | Integration + Predictive | ML anomaly-forecasting; AI-agent intent surface; multi-cloud (Azure/GCP) preview | Predictive:Reactive ≥ 3:1 · AI-Agent Intent Share (first measurement) | Objective #4 — default substrate for agents |
|
||
|
||
**Month-18 destination:** *"Nova is the layer enterprise leadership
|
||
points to when they say 'we don't have an infrastructure ops team
|
||
anymore, and the audit trail is stronger than it ever was.'"*
|
||
|
||
> **Benefit:** each quarter has a concrete deliverable, a target metric
|
||
> grounded in a strategic objective, and a path from "honestly deferred"
|
||
> to "shipped and measured."
|
||
|
||
> **Speaker notes:** Q1–Q3 are committed (grounded pipeline + known
|
||
> unblock paths). Q4 targets are committed-deliverable,
|
||
> aspirational-metric — the ML service ships, the intent-share number is
|
||
> a first measurement (we do not control adoption rate).
|
||
|
||
> **Transition:** "Production-grade guidance is how Nova helps the
|
||
> citizen developer's AI agent meet the bar — here is the first half."
|
||
|
||
---
|
||
|
||
## Slide 18 — Production-Grade Guidance via Atelier (1/2)
|
||
|
||
**Nova instructs the citizen developer's AI agent on production-grade
|
||
engineering — a set of skills and an MCP server.**
|
||
|
||
- **Skills** — markdown files keyed to production-grade engineering
|
||
domains (API, security, data, testing, observability, errors, DevOps,
|
||
infrastructure-as-code, compliance). The skills extend the baseline
|
||
catalog with Nova-specific production-grade principles.
|
||
- **MCP server** — a plugin-registry, stdio server exposing four tools:
|
||
`lookup_principle`, `list_domains`, `matrix_lookup`, and
|
||
`validate_against_principles`. The developer's AI agent (or any
|
||
agentic SDLC platform) calls these tools to look up the principles
|
||
that apply to its submission.
|
||
- **The integration point is the same regardless of source.** Whether
|
||
the submission comes from an AI coding agent, an agentic SDLC
|
||
platform, or a traditional IDE, the same skills and MCP server apply.
|
||
This is how Nova makes the citizen developer production-grade without
|
||
owning the PDLC.
|
||
|
||
> **Benefit:** the citizen developer's AI agent is not unguided — Nova
|
||
> provides production-grade engineering principles as skills and as an
|
||
> MCP surface, so submissions arrive at the contract boundary already
|
||
> aligned with the platform's standards.
|
||
|
||
> **Speaker notes:** This is the first half of the Atelier story — the
|
||
> surface (skills + MCP). The next slide is what the surface catches
|
||
> that deterministic scanners cannot.
|
||
|
||
> **Transition:** "Here is what that guidance catches that deterministic
|
||
> scanners cannot."
|
||
|
||
---
|
||
|
||
## Slide 19 — Production-Grade Guidance via Atelier (2/2)
|
||
|
||
**Agentic validation catches engineering-discipline gaps that deterministic
|
||
scanners miss — and the validation is reproducible.**
|
||
|
||
- **Beyond deterministic scanners.** Wiz, Checkmarx, and Mend check
|
||
policy and secrets — they do not check engineering discipline. The
|
||
Atelier MCP server catches correctness, clarity, and observability gaps
|
||
that deterministic tools cannot: "is this service observable?",
|
||
"is this error path handled?", "is this API contract clear?"
|
||
- **Agentic validation, not a second policy engine.** The MCP server
|
||
gives the AI agent the principles to validate against; the agent does
|
||
the validation. This is agentic validation — the agent reasons about
|
||
the submission against the principles, not a second static scan.
|
||
- **Vendored for audit reproducibility.** Atelier is vendored at a
|
||
pinned tag. A validation result is replayable against the exact
|
||
principles that produced it — so an audit can reproduce a validation
|
||
months later, not just trust a log line.
|
||
|
||
> **Benefit:** the citizen developer's submission is checked for
|
||
> engineering discipline, not just policy compliance — and the check is
|
||
> reproducible for audit. That is what makes the submission
|
||
> production-grade, regardless of which upstream platform produced it.
|
||
|
||
> **Speaker notes:** The value is the gap deterministic scanners leave:
|
||
engineering discipline. Policy scanners catch "is this S3 bucket
|
||
public?"; the MCP server catches "is this service observable if that
|
||
bucket fails?". The vendoring point is audit reproducibility — the
|
||
validation is not a black box.
|
||
|
||
> **Transition:** "You've seen the problem, the solution, and the proof.
|
||
> Here is the recap and the ask."
|
||
|
||
---
|
||
|
||
## Slide 20 — Recap + Ask
|
||
|
||
**The 4-beat recap + the business decision.**
|
||
|
||
**Recap:**
|
||
- **Problem:** product teams own infrastructure without the discipline
|
||
and lifecycle planning it requires; bandwidth gaps and tribal
|
||
knowledge leave operations exposed.
|
||
- **Solution:** autonomous cloud delivery — operations become visible,
|
||
trust is provable (deterministic scoring), humans at stage gates.
|
||
- **Proof:** 100% ledger coverage, 100% attestation coverage, grounded
|
||
ROI formula, four CTO-grade metrics flowing into PowerBI.
|
||
- **Roadmap:** deferred metrics have unblock paths; the 12-month product
|
||
arc activates one strategic objective per quarter.
|
||
|
||
**The ask:** "Approve a pilot estate to activate the production-denominator
|
||
metrics (Lead Time, Vulnerability Count, MTTR, Cloud Spend), and approve
|
||
the tamper-evident ledger build-out to move from the local hash-chain to
|
||
S3 Object Lock + signed checkpoints. These two decisions move Nova from
|
||
'pipeline-ready' to 'production-proven.'"
|
||
|
||
> **Benefit:** a clear business decision — approve a pilot and the ledger
|
||
> build-out — with the confidence that every claim in this deck is
|
||
> grounded, derived, or honestly deferred.
|
||
|
||
> **Speaker notes:** The ask is a business decision, not insider
|
||
> language. "Approve a pilot estate" is a C-suite decision. "Approve the
|
||
> ledger build-out" is a budget decision. The recap reinforces the 4-beat
|
||
> arc — the audience leaves with the structure, not a pile of facts.
|
||
|
||
---
|
||
|
||
## Appendix A1 — Metrics Glossary
|
||
|
||
| KPI | Definition | Status |
|
||
|-----|-----------|--------|
|
||
| Touchless Resolution Rate | runs without operational stage-gate block ÷ total | partial (Post-Pilot) |
|
||
| Human Escalation Frequency | operational stage-gate blocks ÷ total | partial (Post-Pilot) |
|
||
| Automated Decision Accuracy | decisions not followed by failure within 5min | partial (Post-Pilot) |
|
||
| MTTR (p95) | apply.failed → successful retry | grounded |
|
||
| Confidence-Gate Halt Rate | runs with band=block ÷ total | grounded |
|
||
| Provisioning Lead Time | run.completed − run.started | grounded |
|
||
| Deployment Frequency | count(run.completed) per day | grounded |
|
||
| Cost Savings (pre-apply) | sum(delta_usd where delta < 0) | partial (live reconciliation deferred) |
|
||
| FTE Hours Saved | run count × manual baseline × rate | derived (N=0 caveat) |
|
||
| Platform ROI | (labor + cloud + avoided downtime) ÷ op cost | derived (N=0 caveat) |
|
||
| Decision Ledger Coverage | decisions with outcome ÷ total | grounded |
|
||
| Attestation Coverage | prod/dr attested ÷ total prod/dr | grounded |
|
||
| Policy Compliance Rate | 1 − failed_assets ÷ total | grounded |
|
||
|
||
> **Benefit:** a reference for every metric mentioned in the deck.
|
||
|
||
---
|
||
|
||
> **End of deck.** 20 main slides + 1 appendix slide = 21 total. |