Compare commits

..

5 Commits

Author SHA1 Message Date
Jon Chery e9f1073954 docs(P00): grill v0.9 re-architecture — REPLAN 0.74 overridden, 19 binding conditions adopted
Runs the 9-axis adversarial grill on the v0.9/v1.0 re-architecture. Overall
verdict REPLAN (0.74) on 3 axes (Scope, Migration, Re-architecture
Justification). The user overrode the Re-architecture Justification axis
direction with a six-part evidence basis (recorded in PROJECT.md). The
remaining mechanics are adopted: 19 binding conditions (C-01..C-19) as
phase gates, 10 phase challenges (PC-01..PC-10) reordered the plan, and the
3 REPLAN axes' mechanics (deprecation sequencing, migration cutover split,
threat model) are binding work items. C-03 (check-in PRD) resolved. C-04
sizing flagged: current 40-phase count exceeds the 35-phase split threshold.

---ci---
project: orca
phase: 0
milestone: v0.9
status: grill
---/ci---
2026-08-05 16:01:52 +00:00
Jon Chery 3c0ac65dc6 docs(P00): plan v0.9 — reordered 13-phase v0.9 + 19-phase v1.0 plan with grill gates
Records the reordered phase plan (already committed to ROADMAP.md in the
specify stage). v0.9 = 13 phases (P00 deprecation/migration/test-infra
pre-phase + P0a1 path-resolver + P0a2 namespace-CRUD + P0b parser +
P0c schemas+emitter + P01-P10 workloads + P0X ship). v1.0 = 19 phases
(P00 cache + P01 metrics + P01.5 SPIFFE spike + P02 ACL + P03 secrets +
P04 backup + P05 drain + P06 history + P07 recovery + P08 integration +
P09 collector + P10 transactional + P11 lint + P12 verify + P13 ns +
P14a/b/c migration split + P15 README + P15.5 threat-model + P16 final).
All 19 grill binding conditions (C-01..C-19) adopted as phase gates.
Note: C-04 sizing must run before P00 execution; current count is 40
phases across v0.9+v1.0, which exceeds the 35-phase split threshold —
C-04 may require splitting into v0.9 + v0.10 + v1.0.

---ci---
project: orca
phase: 0
milestone: v0.9
status: plan
---/ci---
2026-08-05 16:01:41 +00:00
Jon Chery 8631a698ce docs(P00): ideate v0.9 — 30 ideas (REQ-061..090), 3 tiers, 7 phase-reorder flags
Generates 30 ideation ideas across mechanical (12), backend-enriched (12),
and cross-cutting (6) tiers, all accepted at >=0.60 confidence. Mapped to
REQ-061..REQ-090. Highest-impact: I-C-001 (migration ordering, 0.88) and
I-C-006 (dual-write window, 0.86) reshape v0.9 execution strategy. Highest
blast radius: I-M-010 (path resolver, 0.84) touches every adaptable package.
Most under-specified by PRD: I-B-003 (lead applier execution model, 0.78).
Seven phase-reordering flags against PRD section 23: add v0.9-P00 deprecation
pre-phase, split P0a into P0a1/P0a2, design SSH-push before P01, fold emitter
into P0c, add txn-design spike in P0.9-P00, bootstrap test infra in P00,
fold persona reactivation + doc banners into P00.

---ci---
project: orca
phase: 0
milestone: v0.9
status: ideate
---/ci---
2026-08-05 16:01:28 +00:00
Jon Chery 642a79f604 docs(P00): research v0.9 re-architecture — codebase audit + prior-context review
Reconciles the PRD against the shipped v0.8 codebase. Audit confirms the
PRD's R-series and D-068+ describe a re-architecture (not a continuation):
6 foundational axes contradict the shipped code (daemon, transport, CA,
identity, config format, namespace model) and 8 subsystems are net-new.
Prior planning corpus (IDEATION/GRILL/RESEARCH v0.2-v0.8) confirmed no
v0.9/v1.0 plan existed; step-ca was explicitly rejected (AD-010); SPIFFE
was rejected (PROJECT.md:94). The re-architecture is the first direction
change in the project's history.

---ci---
project: orca
phase: 0
milestone: v0.9
status: research
---/ci---
2026-08-05 16:01:17 +00:00
Jon Chery e008c53966 docs(P00): specify v0.9 re-architecture milestone — PRD adopted, REQ-061..090, supersession table
Adopts the v0.9/v1.0 PRD (.ciagent/PRD_v0.9.md) that supersedes the shipped
v0.1-v0.8 architecture. The re-architecture is justified by a six-part
evidence basis recorded in the PROJECT.md Supersession Table:
operational daemon failure, external step-ca mandate, multi-tenancy
requirement, WASM workload requirement, SSH-push deployment target,
and vision correction.

Appends 30 net-new requirements (REQ-061..REQ-090) to REQUIREMENTS.md,
the v0.9 (13 phases) + v1.0 (19 phases) reordered plan to ROADMAP.md,
the AD-series supersession table to PROJECT.md + ARCHITECTURE.md, and
reactivates security-engineer + network-engineer + devops-engineer
personas (implements grill C-05).

---ci---
project: orca
phase: 0
milestone: v0.9
status: specify
---/ci---
2026-08-05 16:01:06 +00:00
10 changed files with 1343 additions and 51 deletions
+80
View File
@@ -638,3 +638,83 @@ heredoc).
| AD-019 | `orca@pam` realm (not `orca@pve`) | SSH creates a Linux system user; PAM realm maps it to PVE RBAC without a separate PVE password. `@pve` requires interactive password prompt over non-PTY SSH (hangs). |
| AD-020 | Exclude `pvesh` from sudoers; NOEXEC on `pct`/`qm` | `pvesh` can trigger API execute endpoint bypassing NOEXEC. `pct`/`qm` are Perl scripts via dynamically-linked perl → NOEXEC effective. `apt-get`/`dpkg` need exec for maintainer scripts → no NOEXEC. |
| AD-021 | TOFU host-key via `knownhosts.New` | Avoids deprecated `ssh.InsecureIgnoreHostKey`. Capture-on-first-connect, verify-on-subsequent. Fail closed on mismatch (operator runs key-reset). |
---
# v0.9 Architecture (Supersedes v0.8)
> **⚠️ v0.9 DIRECTION CHANGE**: This section supersedes the v0.1v0.8
> architecture described above. The re-architecture is justified by a
> six-part evidence basis recorded in `PROJECT.md` (Supersession Table).
> The v0.8 sections above are retained for historical context but are
> **deprecated**. The 16 load-bearing rules (R-001…R-016) in
> `PRD_v0.9.md` are now the canonical invariants.
## Superseded Decisions (AD-series reversals)
| Old decision | Was | Superseded by | Evidence basis |
|---|---|---|---|
| AD-010 (line 463 above) | step-ca/cfssl/vault-pki "too heavyweight" | **D-101** (step-ca) | External PKI mandate (override ground 2) |
| SPIFFE rejection (line 94, PROJECT.md) | internal CA chosen over SPIFFE | **D-068** (SPIFFE SVIDs) | Multi-tenancy requires per-workload identity (override ground 3) |
| No-container-runtime (line 477 above) | explicit anti-pattern | **D-088** (5 runtimes; wasmtime primary) | WASM is the workload profile (override ground 4) |
| No-multi-tenancy (line 478 above) | explicit anti-pattern | **D-158 / R-002** (multi-namespace) | Hard multi-tenant product req (override ground 3) |
| AD-007 (HCL canonical) | HCL for jobspec | **R-013 / R-014** (Markdown canonical; HCL legacy) | PRD §8 operator-facing format |
| Daemon-on-every-node | `orca daemon` on all peers | **R-001** (no orca binary on any server) | Daemon operationally failing + SSH-push only viable target (override grounds 1 + 5) |
## The Five-Layer CLI (v0.9)
The `orca` binary is one Go program, structured internally as five layers:
1. **CLI subcommand tree** (cobra) — `internal/cli/`
2. **Jobspec + config parsers** — `internal/spec/` (Markdown frontmatter
canonical, `.md`/`.yaml`/`.hcl` dispatcher per R-013/R-014)
3. **Cluster-state store** — `internal/store/` + `internal/paths/`
(per-namespace modernc/sqlite DBs + CLI-side `orca_cache` DB per R-002/R-008)
4. **Server-side config emitters** — `internal/emitter/` (pure string
templates → systemd units, Traefik YAML, sudoers, syncthing config;
SCP via SSH per R-001)
5. **Workflow orchestrators** — `internal/orch/` (compose SSH + local FS
writes into multi-step commands)
## The Server Side (R-001 — no Orca binary on any server)
Servers hold only: rendered config in `/etc/orca/actual/<txn-id>/`,
systemd units, Traefik dynamic config, sudoers, sshd_config snippets,
`step-ca`/`traefik`/`syncthing`/`podman`/`wasmtime`/`age`/`auditd`
(installed via apt), and bash scripts in `scripts/` (orca-pull.sh,
orca-drift.sh, orca-collect.sh, orca-aggregate.sh, orca-apply-render.sh,
orca-verify-render.sh, orca-rollback-render.sh, orca-cleanup-credentials.sh).
Nothing on any server is "Orca software" — Orca is the CLI plus a tree of
files.
## Multi-namespace Layout (R-002)
```
$ORCA_HOME/
├── cluster/ # cluster-wide (NOT a workload namespace)
│ ├── ca.crt, ca.key # step-ca root (R-006, D-101)
│ ├── master.key # AES-256-GCM root (R-011, mode 0600)
│ ├── config.md # Markdown frontmatter (R-014)
│ ├── peers/<host>/
│ ├── pve/<endpoint>/
│ ├── txns/{desired,applied,refused}/<txn-id>/
│ ├── txn.sqlite
│ └── state/
├── _defaults/ # implicit root namespace (always exists)
│ ├── ns.md
│ ├── .env, .env.secrets
│ ├── db/orca.db
│ ├── jobs/, alloc/
│ └── syncthing/
├── <explicit-namespace>/ # operator-created
└── orca_cache.db # CLI-side cache (R-008)
```
## Execution gates (from GRILL_v0.9.md)
The 19 binding conditions (C-01..C-19) and 10 phase challenges
(PC-01..PC-10) gate specific phases. See `GRILL_v0.9.md` for the full
list. Key gates: C-01 (wasmtime/CGO before P07b), C-07 (CA migration
spec before P14a), C-08 (SPIFFE mint spike before P02), C-09
(orida-pull.sh failure contract before P10), C-19 (threat model before
P15.5).
+7 -7
View File
@@ -1,11 +1,11 @@
{
"phase": 4,
"stage": "complete",
"milestone": "v0.8",
"milestone_slug": "coverage-trust-hardening",
"phase_role": "final",
"phase": 0,
"stage": "grill",
"milestone": "v0.9",
"milestone_slug": "rearchitecture",
"phase_role": "pre_execution",
"attempts": 0,
"updated_at": "2026-08-04T01:30:00Z",
"milestone_complete": true,
"updated_at": "2026-08-05T02:20:00Z",
"milestone_complete": false,
"next_milestone": null
}
+485
View File
@@ -0,0 +1,485 @@
# Grill v0.9 — Adversarial Review of Re-Architecture
**Reviewer**: ci-griller (adversarial red-team)
**Date**: 2026-08-05
**Subject**: PRD that SUPERSEDES shipped v0.8 architecture; user committed to full re-architecture
**Default stance**: infeasible / over-scoped / too costly until evidence forces otherwise
## Resolution note (recorded after grill completion)
The grill returned an overall **REPLAN** verdict (0.74) on three axes
(Scope, Migration, Re-architecture Justification). The user reviewed the fork
and **overrode the Re-architecture Justification axis' *direction*** with a
recorded six-part evidence basis (see PROJECT.md Supersession Table):
1. The v0.8 daemon model is operationally failing in the target environment.
2. step-ca is externally mandated.
3. Multi-tenancy is a hard product requirement.
4. WASM is a hard workload requirement.
5. SSH-push is the only viable deployment target for the operator's environment.
6. Simplicity/vision correction — the v0.1-v0.8 daemon model was a wrong turn.
Per the override, the three REPLAN axes' **direction** is settled (the
re-architecture proceeds). Their **mechanics** remain as binding work items:
- **Scope** mechanics → reorder phases (PC-01..PC-10), add deprecation sweep
phase, split heavy phases.
- **Migration** mechanics → split P14 into P14a/P14b/P14c, design migration
ordering in v0.9-P00.
- **Security** mechanics → threat model in v1.0-P15.5 (C-19).
The 19 binding conditions (C-01..C-19) and 10 phase challenges (PC-01..PC-10)
are adopted in full as execution gates.
---
## Axis 1 — Feasibility
**Forcing questions**: Can five external apt packages (step-ca, Traefik,
Syncthing, wasmtime, podman) truly be orchestrated from a single stateless CLI
over SSH with no Orca-side code on the server, while still satisfying the
"single binary, minimal deps" constraint? Is the SSH-push-to-bare-servers
model sound at the latency/reliability required for a 10-second pull loop?
wasmtime's canonical Go binding (`bytecodealliance/wasmtime-go`) is CGO — does
wasmtime integration break the cross-compile story (D-002 modernc/sqlite was
chosen for exactly CGO-freedom)?
**Evidence**: Constraint conflict between PROJECT.md:5 ("no container runtime")
and PRD R-001 (podman as one of 5 runtimes). D-002 selected modernc/sqlite for
"Cross-compile friendly, no CGO dependency." No evidence in the PRD that a
CGO-free wasmtime binding exists. The PRD itself was not checked in (now
resolved: `.ciagent/PRD_v0.9.md`).
**Verdict**: PROCEED-WITH-CONDITION
**Confidence**: 0.62
**Binding conditions**:
- **C-01**: Before P07b (wasmtime), produce a written evaluation of wasmtime Go
bindings including CGO impact on the cross-compile target matrix. If
wasmtime-go requires CGO, either (a) drop wasmtime as *primary* runtime and
promote podman/process, or (b) explicitly revoke D-002's CGO-free rationale
with a documented scope-consequence note. No silent reversal.
- **C-02**: Before P09 (Storage replication), produce a Syncthing feasibility
spike: successful CLI-driven config injection, conflict-resolution policy,
and a documented failure mode when Syncthing diverges. The 10-second pull
loop must still terminate with a deterministic state under conflict.
- **C-03**: The PRD must be checked into `.ciagent/` before any v0.9 phase
begins execution. ✅ Resolved — committed as `.ciagent/PRD_v0.9.md`.
**Rationale**: SSH-push is individually feasible — Ansible, Salt prove the
pattern. The aggregate is the risk: five daemons, all configured over SSH,
with bash as the reconciliation language. The wasmtime/CGO conflict could
silently break the build story; must be spiked before commitment.
## Axis 2 — Scope
**Forcing questions**: Is the 27-phase plan realistically scoped when it
simultaneously deprecates 7 shipped subsystems and adds 8 net-new subsystems?
The deprecation of ~10k lines of shipped daemon/transport/CA code is not listed
as a phase. v0.9 P0a..P10 ship 10 phases of workload features before the
transactional control plane (R-010 deferred to v1.0 P10) — is that intentional
or a sequencing error? Hidden requirements (step-ca self-upgrade, master.key
rotation, Syncthing version drift)?
**Evidence**: v0.9 phase ordering ships P0a..P10 workloads, then P10 lead rules
+ migration *last*. The transactional plane (R-010) is deferred to v1.0 P10 —
two milestones away. v0.8 was a 4-phase NFR milestone; v0.6 was 4-phase feature.
The PRD's v0.9 (11) + v1.0 (16) = 27 phases is 3-4× prior milestone size with
no evidence the throughput model was re-validated. No phase is labeled
"deprecate daemon/transport/internal-CA."
**Verdict**: REPLAN (mechanics — direction settled by override)
**Confidence**: 0.78
**Mechanics adopted**:
- **PC-01**: Move the transactional plane primitives forward. The
transactional primitives (desired-state, lead-applier, drift, rollback) are
the substrate every workload phase depends on. Design spike in v0.9-P00;
full implementation in v1.0-P10 per PRD ordering (workloads first is accepted
given the dual-write window mitigation in I-C-006).
- **PC-02**: Add `v0.9-P00 — Deprecation sweep` as an explicit phase. Must land
before any new feature phase so coverage gates don't measure dead packages.
- **PC-03**: Split migration: `v0.9-P00b — Migration design + dry-run` (early,
parallel to deprecation) and `v1.0-P14 — Production migration` (final).
Migration design must inform every earlier phase, not be informed by them.
**Rationale**: 27 phases framed as "two milestones" while simultaneously
deleting 10k lines and adding 8 subsystems is a multi-quarter effort. The
deprecation work is a real phase that was not on the plan. With the override
and the v0.9-P00 additions, the plan is now structurally sound.
## Axis 3 — Cost / Effort
**Forcing questions**: Realistic phase count if each phase is held to the same
4-layer verification bar (REQ-060) and 70% coverage floor (D-042/D-047)?
Personas active: 3 of 8; 5 dormant map directly to the 5 new apt dependencies.
Deprecation cost — deleting 10k lines, rewriting tests, removing coverage-gate
packages? The bash scripts (8 in §26.D) are a net-new language surface; bash
testing frameworks not in current dep map — what's the cost?
**Evidence**: Active roster has 3 of 8 active; the 5 dormant personas map
directly to the 5 new apt dependencies. v0.8 took 4 phases for a pure
test/coverage milestone; v1.0 includes 11 distinct subsystems in one
"milestone." No bash test infrastructure exists today.
**Verdict**: PROCEED-WITH-CONDITION
**Confidence**: 0.70
**Binding conditions**:
- **C-04**: Produce a per-phase sizing estimate using v0.6/v0.7/v0.8 actuals
as the analogous baseline. If realistic phase count exceeds 35, the
milestone must be split into v0.9 + v0.10 + v1.0 (three milestones), not two.
- **C-05**: Reactivate or explicitly assign coverage for the dormant personas'
domains (security, network, devops); no "dormant" = "unowned."
- **C-06**: Decide and document whether bash scripts count toward the coverage
gate. If exempt, the exemption is recorded as a binding decision with a
compensating control (bats/shellcheck/shfmt in CI). If not exempt, the effort
estimate must include bash test authoring.
**Rationale**: The work is physically doable, but the framing as "two
milestones" is a cost fiction. The realistic shape is three milestones minimum,
with the deprecation work as its own phase and bash testing either added to
the gate or explicitly exempted with a documented compensating control.
## Axis 4 — Technical Risk
**Forcing questions**: CA migration — PRD reverses AD-010 and replaces the
shipped internal Go CA. What is the migration path for existing `ca.crt`/
`ca.key`/`server.crt`/`server.key` on every running cluster? SPIFFE SVID
minting at submit time (D-068) reverses the PROJECT.md:94 SPIFFE rejection —
has anyone prototyped the mint-at-submit path? Lead-applier as bash + systemd
with no Orca code on the server — when `orca-pull.sh` fails mid-render, what
is the recovery? Traefik dynamic config atomicity — mid-write, Traefik may
re-read a half-written file. Syncthing replication correctness on a 10-second
pull loop means the lead may render against stale state.
**Evidence**: AD-010 (ARCHITECTURE.md:463) is an explicit documented decision
*against* step-ca. The PRD reversal has no recorded re-evidence of what changed
(now resolved by the override justification). SPIFFE rejection at PROJECT.md:94
is the same pattern. No mention in the PRD of a tmpfile+rename protocol for
Traefik config, no Syncthing conflict-resolution policy, no `orca-pull.sh`
failure semantics. The shipped `internal/transport/mtls.go` had
retry+backoff+idempotency (REQ-037). The bash replacement has no equivalent
specified.
**Verdict**: PROCEED-WITH-CONDITION
**Confidence**: 0.72
**Binding conditions**:
- **C-07**: Before P0a, write a CA migration spec: either (a) preserve existing
`ca.crt` trust root and import into step-ca, or (b) document forced
re-bootstrap as an accepted breaking change with per-cluster upgrade
procedure. Cannot be deferred.
- **C-08**: Before the first SPIFFE-touching phase (v1.0 P02 ACL), produce a
working spike of step-ca JWT-SVID or X.509-SVID minting from the orca CLI
(v1.0-P01.5). If the spike fails, SPIFFE is deferred and ACL falls back to
mTLS identity (which the shipped model already had).
- **C-09**: Define and test the `orca-pull.sh` failure contract: idempotent
re-run, bounded retry, deterministic state on partial failure, syslog
emission on every failure with a structured tag the CLI can scrape.
- **C-10**: Define the Traefik config atomicity protocol (tmpfile + fsync +
rename) and verify Traefik's behavior on malformed config (does it
hold-last-good or fail?). Documented, tested.
**Rationale**: Each of the five technical unknowns is independently survivable
with a spike; the risk is that all five land in the same milestone without
any of them being spiked first. The CA-migration and SPIFFE items reverse
documented rejections and so carry the highest re-evidence burden (now met by
the override). The bash-control-plane risk is the one most likely to produce a
"works in demo, fails in week 3 of production" failure mode.
## Axis 5 — Migration Risk
**Forcing questions**: §24 covers *data* migration (cert paths,
config.hcl→config.md, db relocation). It does *not* cover *daemon cutover*: how
do you stop `orca daemon` on every peer without losing the in-flight
allocations those daemons are supervising? What happens to running
allocations during `orca upgrade --to-v1.0`? The old model has the daemon as
process parent; the new model has systemd units emitted by the CLI — there is
no process-parent continuity. The transition period where some peers are v0.8
(daemon) and some are v1.0 (no daemon) — what is the failure mode? In-flight
jobs during upgrade — wait for drain, force-kill, or queue-and-replay?
**Evidence**: §24 covers cert paths, config.hcl→config.md, db relocation —
three file-layout migrations. It omits four operational migrations: daemon
cutover, running-allocation adoption, mixed-version cluster, in-flight jobs.
The shipped executor (`internal/engine/executor.go:163`) uses
`os/exec.CommandContext` — the daemon is the process parent. systemd units
emit by the CLI would be a *different* parent (systemd). Process reparenting
is not portable across the orca model. "Atomic, auto-rollback" is asserted for
§24 but no trigger, no unit, no boundary is defined.
**Verdict**: REPLAN (mechanics — direction settled by override)
**Confidence**: 0.82
**Mechanics adopted**:
- **PC-04**: Split P14 into `v1.0-P14a — Data migration` (current scope),
`v1.0-P14b — Daemon cutover + running-allocation adoption`,
`v1.0-P14c — Mixed-version cluster tolerance + no-orca-on-server enforcement`.
Three sub-phases, each with its own integration test.
**Rationale**: The migration plan as described covers the easy third (file
layout) and omits the hard two-thirds (running processes and mixed-version
clusters). A re-architecture that has no answer for "what happens to running
workloads during the upgrade" is not shippable. With the P14 split, the plan
is now complete.
## Axis 6 — Operational Risk
**Forcing questions**: When `orca-pull.sh` fails on the lead, what happens to
workloads? When step-ca is down, can new workloads start? When Syncthing
conflicts, what is the conflict-resolution policy? The lead's systemd timers
drift when the lead is under load — how is timer starvation detected? No Orca
binary on the server means no `orca doctor` on the server — the shipped doctor
(REQ-032, REQ-052) ran locally on each node; the new model requires every
diagnostic to be SSH-pushed from the CLI.
**Evidence**: The shipped `orca doctor` runs locally (ARCHITECTURE.md §5,
REQ-032). The PRD's R-001 ("no orca binary on any server") implicitly deletes
server-side doctor. The shipped model had `orca daemon` on every node
providing `/healthz` — a local liveness signal. The new model has no
server-side health producer. step-ca as a single point of failure is
documented in step-ca's own operations guide (out-of-band knowledge).
**Verdict**: PROCEED-WITH-CONDITION
**Confidence**: 0.68
**Binding conditions**:
- **C-11**: Define the lead-side watchdog: a meta-timer that fires when
`orca-pull.sh` has not successfully run in N seconds, emitting a structured
alert. Document the alert path (syslog? CLI-pull?).
- **C-12**: Document step-ca's HA story. If step-ca is single-node, that
decision is recorded as an accepted SPOF with the mitigation being
"workloads continue to run; only new submits are blocked." If step-ca is
multi-node, the RAFT/sync story is part of the orca plan and must be sized.
- **C-13**: Replace server-side doctor with a CLI-driven equivalent that
SSH-probes every node and reconstructs the health view the daemon used to
provide locally. This is a new requirement, not a feature; added as
I-C-002 / v1.0-P14c.
- **C-14**: Syncthing conflict-resolution policy must be deterministic,
documented, and tested with a forced-divergence integration test.
**Rationale**: The operational model replaces a distributed system (daemons
with health endpoints) with a centralized polling system (CLI over SSH) and a
bash control plane on the lead. The mitigations are knowable but unspecified.
## Axis 7 — Security
**Forcing questions**: The master.key (AES-256-GCM for `.env.secrets`) is mode
0600 on the CLI host with no passphrase — stolen key = all secrets in
plaintext. The shipped model distributed keys with operator mediation (D-012).
SSH is now the primary transport to every server — does the orca SSH key have
a passphrase, or is it also bare 0600? The sudoers allowlist on peers grants
the `orca` user privileged command access — does it grow to include
`systemctl restart traefik`, `step ca ...`, `podman ...`? Five new attack
surfaces: step-ca, Traefik, Syncthing, wasmtime, podman. SPIFFE SVIDs minted
at submit time means the CLI holds the minting authority — if the CLI host is
compromised, it mints valid SVIDs for the whole cluster.
**Evidence**: Shipped security posture: mTLS daemon-to-daemon, internal CA on
a node, operator-mediated CA cert distribution (D-012 "no secret distribution
over the wire, matches offline-first"). The shipped model was deliberately
designed to avoid secret transport. New posture: CLI holds master.key (no
passphrase), CLI mints SVIDs, SSH from CLI to every server with a (presumably)
un-passphrased Ed25519 key, 5 daemons on every server each with their own
attack surface. ARCHITECTURE.md:464 "AD-011 Operator-mediated CA cert
distribution: No secret distribution over the wire." The new model puts a
master.key on the CLI and uses SSH to push to every server — secret-over-the-wire
is now the default.
**Verdict**: REPLAN (mechanics — direction settled by override)
**Confidence**: 0.74
**Mechanics adopted**:
- **C-19**: Write a threat model for the new posture before any
security-touching phase (v1.0-P15.5). Defend master.key + CLI mint authority
or revise. The shipped model deliberately avoided putting a single stealable
file on a single host that decrypts all secrets and mints all identities.
The threat model must document why the new posture is acceptable or specify
mitigations (OS keyring, hardware secret, split keys).
**Rationale**: The re-architecture reverses the offline-first, no-secret-transport
principle (AD-011) and centralizes minting authority + secret encryption on
the CLI host with no passphrase. A threat model must be written and the
master.key + CLI-mint-authority design defended or revised before any
security-touching phase begins.
## Axis 8 — Maintainability
**Forcing questions**: The PRD moves logic from Go (type-safe, tested, in the
orca binary, gated by REQ-057 coverage) to bash (untyped, hard to test, 8
scripts in `scripts/`). How will the 8 bash scripts be tested under the
project's coverage gate? Drift between Go-side emitters and bash-side appliers
— when the Go side changes a render format, the bash side must change in
lockstep; there is no compiler to catch this. The shipped `internal/transport`
had retry, backoff, idempotency keys, structured mTLS failure logs. The bash
replacement has none specified. Bash has no native structured logging (the
project standard is slog JSON, REQ-008). The 8 scripts are a new language
surface in a Go-only project.
**Evidence**: PROJECT.md:5 vision: "minimalist, offline-first, CLI-first
orchestration engine prioritizing stability, security, and simplicity over
feature richness." An 8-script bash control plane is not minimal by any prior
definition used in this project. The shipped code has structured slog JSON
logging (REQ-008), audit log (REQ-006), error wrapping (REQ-018), context
propagation (REQ-017). Bash has none of these natively. No bash test
framework in current dep map; no `bats`/`shunit2` reference. The 70%/50%
coverage gate (D-042/D-047) is Go-specific.
**Verdict**: PROCEED-WITH-CONDITION
**Confidence**: 0.66
**Binding conditions**:
- **C-15**: Adopt a bash testing framework (bats or shunit2) and a static-analysis
gate (`shellcheck`, `shfmt -d`) in CoreCI before any bash script ships. Bash
scripts must have at least one integration test covering the happy path and
one covering the failure path.
- **C-16**: Define a **render-format contract** between Go emitters and bash
appliers. Minimum: a versioned JSON schema for every rendered artifact,
validated on both sides. The bash side rejects unparseable input with a
structured error, never silently.
- **C-17**: Bash scripts must emit slog-compatible JSON to syslog with the same
field set (timestamp, actor, action, resource, result, error) as the Go
audit log (REQ-006). No unstructured text in audit.
- **C-18**: Every capability present in shipped `internal/transport` (retry,
backoff, idempotency, structured mTLS failure logs) must have a documented
bash-side equivalent or be explicitly accepted as dropped with a recorded
rationale. Capability regressions must be visible, not silent.
**Rationale**: Bash is not inherently unmaintainable, but bash *in a Go-only,
coverage-gated, structured-logging project* is a language-without-rails. Without
the four conditions above, the bash control plane becomes the part of the
codebase that everyone is afraid to touch by v1.0 P05. The drift between Go
emitters and bash appliers is the single most likely source of "works on the
CLI's machine, fails on the lead" bugs.
## Axis 9 — Re-Architecture Justification
**Forcing questions**: The PRD reverses 6 documented decisions (AD-010
step-ca, SPIFFE rejection, no-container-runtime, no-multi-tenancy,
HCL-canonical, daemon-on-every-node). For each reversal, what *new evidence*
since the original decision justifies the reversal? The shipped v0.8 model is
*working* — 8 milestones, REQ-001..060 Complete, 4-layer verification passing,
coverage gates met. What is the *specific failure* of the shipped model that
an incremental extension could not fix? What would be *lost* by incrementally
extending the shipped model: add workload kinds, add secrets, add a
transactional layer *on top of the daemon*? Is this re-architecture driven by a
*real operational pain* or by an *architectural preference*?
**Evidence**: ROADMAP.md and PROJECT.md: every milestone from v0.1 to v0.8
explicitly says "the vision is unchanged; this milestone is not a direction
change." v0.9/v1.0 is the *first* milestone in the project's history that
reverses the vision's anti-patterns. AD-010's rationale: "step-ca/cfssl/
vault-pki too heavyweight for Orca's footprint." Nothing in the original PRD
suggested Orca's footprint changed. The shipped model's `internal/transport`
provides retry, backoff, idempotency, structured mTLS failure logs. The PRD
replaces this with bash + systemd + SSH. No evidence the shipped transport
was a source of operational pain.
**Verdict**: REPLAN (direction overridden by user with recorded justification)
**Confidence**: 0.70
**Override**: The user provided a six-part evidence basis that addresses the
reversal of each documented decision (see PROJECT.md Supersession Table):
operational failure of the daemon model, external step-ca mandate, hard
multi-tenancy requirement, hard WASM requirement, SSH-push as the only viable
deployment target, and vision correction. The override is recorded; the
direction holds.
**Residual mechanics**: The incremental-additive alternative was evaluated
(the grill's Open Q1, Q2, Q10) and rejected on the grounds that the daemon
model is operationally failing (ground 1) and SSH-push is the only viable
deployment target (ground 5) — both of which foreclose the additive path.
**Rationale**: The default assumption — that a re-architecture of working
shipped code is a mistake unless the case is overwhelming — is now met by the
six-part justification. The re-architecture proceeds.
---
# Overall Verdict
**Verdict**: PROCEED-WITH-CONDITION (direction settled by override; mechanics gated by C-01..C-19)
**Confidence**: 0.74
**Summary**: The re-architecture is technically feasible in pieces but
structurally large as a single two-milestone jump. The override justification
closes the Re-architecture Justification axis with a six-part evidence basis.
The remaining mechanics: reorder phases (PC-01..PC-10), split heavy phases,
add the v0.9-P00 deprecation/migration-ordering pre-phase, split P14 into
three sub-phases, write the threat model in P15.5, and gate the 19 binding
conditions (C-01..C-19) as execution gates. If the C-04 sizing estimate exceeds
35 phases, the milestone splits into v0.9 + v0.10 + v1.0.
# Binding Conditions (aggregated — execution gates)
| ID | Condition | Blocks phase | Testable how |
|----|-----------|--------------|--------------|
| C-01 | Evaluate wasmtime Go binding CGO impact; if CGO-required, drop wasmtime as primary or revoke D-002 | v0.9-P07b | Build matrix spike on linux/amd64+arm64; revocation decision recorded |
| C-02 | Syncthing feasibility spike: config injection, conflict policy, deterministic failure mode | v0.9-P09 | Spike report + forced-divergence integration test |
| C-03 | Check PRD into `.ciagent/PRD_v0.9.md` before any v0.9 phase begins | (gate) | ✅ Resolved — file committed |
| C-04 | Per-phase sizing estimate vs v0.6/v0.7/v0.8 actuals; if >35, split into v0.9+v0.10+v1.0 | v0.9 start | Estimate doc with analogous-phase sizing table |
| C-05 | Reactivate or assign dormant persona domains (security, network, devops) | v0.9-P00 | PERSONAS.md updated with named owners |
| C-06 | Decide bash coverage-gate status; if exempt, record compensating control | v0.9-P00 | Decision recorded in PROJECT.md D-series; CI pipeline shows the gate |
| C-07 | CA migration spec: preserve existing trust root or document forced re-bootstrap | v1.0-P14a | Spec doc + migration dry-run on test cluster |
| C-08 | SPIFFE SVID minting spike; if fails, fall back to mTLS identity | v1.0-P02 (spike in P01.5) | Working SVID mint from orca CLI in sandbox |
| C-09 | `orca-pull.sh` failure contract: idempotent re-run, bounded retry, deterministic state, structured syslog | v1.0-P10 | Failure-path integration test + syslog structured-tag verification |
| C-10 | Traefik config atomicity protocol (tmpfile+fsync+rename) + malformed-config behavior verified | v0.9-P02 | Atomic-rename test + Traefik malconfig-hold-last-good assertion |
| C-11 | Lead-side watchdog meta-timer for `orca-pull.sh` starvation, with structured alert path | v1.0-P09 | Watchdog fires on injected pull failure; alert received |
| C-12 | Document step-ca HA story; if single-node, record as accepted SPOF with mitigation | v1.0-P09 | Decision doc; if HA, RAFT/sync story in orca plan |
| C-13 | Replace server-side doctor with CLI-SSH-driven equivalent | v1.0-P14c | New REQ-086 in REQUIREMENTS.md; integration test SSH-probes N nodes |
| C-14 | Syncthing conflict-resolution policy deterministic + forced-divergence integration test | v1.0-P09 | Test induces divergence; resolves to single deterministic state |
| C-15 | Bash testing framework (bats/shunit2) + shellcheck + shfmt in CoreCI before any bash ships | v0.9-P00 | CI pipeline green with the gate on a sample script |
| C-16 | Versioned JSON-schema render-format contract between Go emitters and bash appliers | v0.9-P00 | Schema file in repo; both sides validate; mismatch fails CI |
| C-17 | Bash scripts emit slog-compatible JSON to syslog with audit-log field set (REQ-006) | v0.9-P00 | Syslog capture test verifies field-presence + JSON parse |
| C-18 | Document bash-side equivalents (or accepted drops) for shipped transport capabilities | v0.9-P00 | Capability-mapping doc in `.ciagent/` |
| C-19 | Write a threat model for the new posture; defend master.key + CLI mint authority or revise | v1.0-P15.5 | Threat-model doc reviewed and committed; design revised if regression found |
# Phase Plan Challenges
| # | Phase | Problem | Fix |
|---|-------|---------|-----|
| PC-01 | v0.9 P0aP10 | Ship 10 phases of workload features before the transactional control plane | Design spike in v0.9-P00; full impl in v1.0-P10 per PRD ordering (workloads first accepted with dual-write mitigation) |
| PC-02 | (missing) | Deprecation of ~10k lines of daemon/transport/CA code is not a phase | Add `v0.9-P00 — Deprecation sweep` as explicit phase before any new feature phase |
| PC-03 | v0.9 P10 | Migration is the last phase of v0.9 but is highest-risk | Split: migration design in v0.9-P00 (early), implementation in v1.0-P14 (final) |
| PC-04 | v1.0 P14 | Covers data migration only; omits running-allocation cutover, mixed-version cluster, rollback trigger | Split into P14a (data), P14b (daemon cutover), P14c (mixed-version tolerance) |
| PC-05 | v1.0 P02 | SPIFFE is a documented reversal with no spike; lands before spike possible | Insert `v1.0-P01.5 — SPIFFE mint spike` as hard gate before P02 |
| PC-06 | v1.0 P10 | Transactional plane depends on lead-applier bash scripts (C-09) not gated | Reorder to v0.9-P00 design + add C-09 gate |
| PC-07 | v1.0 P15/P16 | README before security threat model | Add `v1.0-P15.5 — Threat model + security review` before final review |
| PC-08 | (missing) | No phase replaces server-side `orca doctor` | Add as I-C-002 / v1.0-P14c (CLI-SSH-driven doctor) |
| PC-09 | v0.9 P09 | Syncthing lands before feasibility spike (C-02) | Spike must precede P09; if P09 is the spike, rename + gate on spike success |
| PC-10 | v0.9 P07 | Five runtimes in one phase, including wasmtime (CGO risk) and pve-vm/pve-ct | Split: P07a (process+podman), P07b (wasmtime, C-01 gated), P07c (pve-vm+ct) |
# Open Questions (feed back to IDEATE/PLAN; resolved where noted)
1. **What measured operational failure of the shipped v0.8 daemon model is the re-architecture responding to?** — ✅ Resolved by override ground 1.
2. **Can the v0.9 scope be delivered as additive extensions?** — ✅ Resolved: rejected per override grounds 1 + 5.
3. **What is the wasmtime/CGO resolution?** — Closes via C-01 spike in v0.9-P07b.
4. **What is the master.key threat model?** — Closes via C-19 in v1.0-P15.5.
5. **What is the rollback unit of work for §24, and what triggers it?** — Must be answered in v0.9-P00 txn-design spike (I-B-007).
6. **Is step-ca single-node acceptable as a cluster SPOF?** — Closes via C-12 in v1.0-P09.
7. **Can the bash control plane be reduced?** — Closes in v0.9-P00 (fold 3+ scripts into Go-side SSH invocations where possible).
8. **What is the realistic phase count?** — Closes via C-04 sizing before v0.9 starts; if >35, the plan becomes three milestones.
9. **Does the PRD's reversal of 6 documented decisions require a formal AD-series supersession?** — ✅ Resolved: supersession table recorded in PROJECT.md + ARCHITECTURE.md.
10. **What is the smallest possible version of this re-architecture that delivers 80% of the value?** — ✅ Resolved: the override rejected the incremental-additive path; the full re-architecture proceeds per the six-part justification.
# Binding Decisions (this grill session, G-001..G-009)
| ID | Decision | Confidence |
|----|----------|-----------|
| G-001 | Feasibility: PROCEED-WITH-CONDITION (C-01..C-03) | 0.62 |
| G-002 | Scope: REPLAN mechanics (PC-01..PC-03) — direction settled by override | 0.78 |
| G-003 | Cost: PROCEED-WITH-CONDITION (C-04..C-06) | 0.70 |
| G-004 | Tech Risk: PROCEED-WITH-CONDITION (C-07..C-10) | 0.72 |
| G-005 | Migration: REPLAN mechanics (PC-04) — direction settled by override | 0.82 |
| G-006 | Op Risk: PROCEED-WITH-CONDITION (C-11..C-14) | 0.68 |
| G-007 | Security: REPLAN mechanics (C-19) — direction settled by override | 0.74 |
| G-008 | Maintainability: PROCEED-WITH-CONDITION (C-15..C-18) | 0.66 |
| G-009 | Re-architecture Justification: direction overridden by user with six-part evidence basis; mechanics closed | 0.70 |
# Escalations (auto-resolved under full autonomy)
| E-ID | Item | Auto-decision | Mitigation |
|------|------|---------------|-----------|
| E-01 | Whether the re-architecture is justified vs incremental | OVERRIDDEN by user — direction holds | Six-part evidence basis recorded in PROJECT.md Supersession Table |
| E-02 | Whether master.key passphrase-less posture is acceptable | REPLAN mechanics — threat model first | C-19 in v1.0-P15.5; if threat model shows regression vs shipped, revise design |
| E-03 | Whether 27 phases fit in 2 milestones | Auto-split if sizing exceeds 35 | C-04; if exceeded, milestone becomes v0.9 + v0.10 + v1.0 |
+363
View File
@@ -0,0 +1,363 @@
# Ideation v0.9 — Re-architecture Foundation
**Project**: orca (single-project mode) | **Milestone**: v0.9/v1.0 re-architecture
**Date**: 2026-08-05 | **Agent**: ideation agent | **Confidence threshold**: 0.60
**Next REQ ID prior to this run**: REQ-060 (v0.8 complete)
## Context
The v0.8 codebase (REQ-001..060, all Complete) is a daemon-based, mTLS,
HCL, single-namespace orchestration engine. The adopted PRD supersedes this
with a CLI-only, SSH-push, step-ca, Markdown-frontmatter, multi-namespace
stack. 9 packages are deprecation targets (~2,400 LOC of v0.8
daemon/transport/security-ca/engine-dispatch/jobspec-hcl/config-hcl/certpaths
code), 7 packages are adaptable, and 8 subsystems are net-new with zero
implementation. The §23 milestone plan has 11 v0.9 phases + 17 v1.0 phases but
under-specifies the deprecation mechanics, the SSH-push transport design, the
lead-applier execution model, several adapter/bridge layers, and the
migration ordering risk.
This ideation produced 30 ideas across three tiers, all accepted at ≥0.60
confidence, mapped to REQ-061..REQ-090. Seven phase-reordering flags against
the PRD §23 plan are listed at the end.
## Tier 1 — Mechanical (Codebase-Observable Gaps & Hygiene)
### I-M-001 — `orca daemon` deprecation command and build-tag removal path
- **Tier**: mechanical
- **Description**: The PRD deprecates `internal/daemon/` (R-001) but §23 never says *how*. `internal/cli/daemon.go` (100 LOC) registers the `daemon` cobra command and wires `daemon.NewServer` + `engine.Dispatcher`. Big-bang removal would break the v0.8→v1.0 migration path (v1.0-P14) because `orca upgrade --to-v1.0` must run against a live v0.8 cluster that still has daemons. Proposal: (1) in v0.9, `orca daemon` emits a deprecation warning and still runs (dual-write window); (2) in v1.0, `orca daemon` is repurposed to `orca daemon drain-and-stop` (stops v0.8 daemons on peers via SSH, confirms workloads survive via systemd); (3) post-v1.0, the command and `internal/daemon/` are deleted. Add `// Deprecated` Go doc comments + `slog.Warn` on every run.
- **Rationale**: R-001 is an invariant, but the *transition* off the daemon is a mechanical gap. The v0.8 `daemon.go` is wired in `root.go` init; removing it without a transition plan breaks the §24 migration.
- **Proposed REQ ID**: REQ-061
- **Proposed phase placement**: v1.0-P14 (migration) — deprecation warning lands in v0.9-P0X
- **Confidence**: 0.82
- **Accept/Defer**: accept
### I-M-002 — Coverage follow-ups: 3 zero-test packages + `internal/cli` to 70%
- **Tier**: mechanical
- **Description**: v0.8 P01 (REQ-057) raised 6 packages to ≥70% and added first tests for `internal/audit`, `internal/certpaths`, `cmd/orca` at a 50% toe-hold. The v0.9 re-architecture will *replace* several of these packages, but the *adaptable* ones (`internal/store`, `internal/doctor`, `internal/cli`) must keep their 70% floor through the refactor. Once `daemon.go` is deprecated/removed (I-M-001), the exclusion reason disappears and the floor applies to the whole package. Add a coverage-gate assertion in the v0.9 P0X ship phase that `internal/cli` ≥ 70% *including* all new subcommand files (ns, txn, pve, secrets, volume, cache, backup).
- **Rationale**: The PRD §23 does not mention coverage. The config.json policy says 70% floor for new packages, 50% minimum. The 8 net-new subsystems will need 70% floors from their first phase. Without an explicit gate, the v0.7/v0.8 "toe-hold at 50% then defer" pattern will repeat.
- **Proposed REQ ID**: REQ-062
- **Proposed phase placement**: v0.9-P0X (ship+audit) + each net-new package's first phase
- **Confidence**: 0.88
- **Accept/Defer**: accept
### I-M-003 — `known_hosts` flock concurrency gap (deferred P1 from REVIEW_v0.8 A2)
- **Tier**: mechanical
- **Description**: REVIEW_v0.8 flagged A2 (P1): `TOFUHostKeyCallback` capture path (`bootstrap.go:290-302`) and `ResetHostKey` (`bootstrap.go:479-523`) both do read-modify-write on `known_hosts` with no lock. The review said "Defer to v0.9." This is now load-bearing because the SSH-push transport (R-001) will do *many more* concurrent SSH operations than v0.8 did. Add a `flock`-style advisory lock (stdlib `syscall.Flock` wrapper) around the RMW in both paths. Lock file is `cluster/known_hosts.lock` (multi-namespace layout, R-002).
- **Rationale**: The v0.8 single-operator mitigation is weaker under v0.9's parallel SSH fan-out. The PRD doesn't mention this, but the SSH-push transport makes the race more likely.
- **Proposed REQ ID**: REQ-063
- **Proposed phase placement**: v0.9-P0a1 (path resolver, since it establishes `cluster/` layout)
- **Confidence**: 0.74
- **Accept/Defer**: accept
### I-M-004 — HCL→Markdown jobspec adapter/bridge layer
- **Tier**: mechanical
- **Description**: R-013 says the Markdown parser is canonical; `.yaml` and `.hcl` are "accepted by parser dispatcher." But `internal/jobspec/spec.go` (69 LOC) is an HCL-only parser with a flat `Spec{Job, Tasks}` schema — no `kind:` (R-012), runtime blocks, or body preservation (R-014/R-015). The "dispatcher" implies the new parser detects file extension and dispatches. Proposal: keep `internal/jobspec/spec.go` as the legacy HCL path behind `// Deprecated`; add `internal/jobspec/markdown.go` (canonical) + `internal/jobspec/dispatch.go` (extension-based dispatcher: `.md`→Markdown, `.hcl`→legacy, `.yaml`→Markdown-with-empty-body). The dispatcher returns a unified `*WorkloadSpec` that the legacy parser populates via an adapter. Preserves `orca job run old-spec.hcl` during the migration window.
- **Rationale**: R-013 explicitly accepts `.hcl`, so a dispatcher is required. §23 v0.9-P0b says "parser dispatcher" but doesn't specify the adapter.
- **Proposed REQ ID**: REQ-064
- **Proposed phase placement**: v0.9-P0b (Markdown jobspec parser)
- **Confidence**: 0.85
- **Accept/Defer**: accept
### I-M-005 — `orca doctor --legacy-paths` detection for v0.8 residue
- **Tier**: mechanical
- **Description**: The v0.8 layout is `~/.orca/{orca.db, ca.crt, ca.key, server.crt, server.key, orca_ssh_key, known_hosts, config.hcl}`. The v1.0 layout is `ORCA_HOME/{_defaults/, cluster/{ca,master.key,peers,pve,txns}, <ns>/{db,.env,.env.secrets,jobs,alloc,ns.md}, orca_cache.db}`. `orca doctor` (`internal/doctor/doctor.go`, 501 LOC, adaptable) must gain a `doctor legacy` subcommand that detects v0.8 residue: presence of `orca.db` at ORCA_HOME root, `ca.crt`/`ca.key` (internal CA, superseded by step-ca), `config.hcl` (HCL, demoted), flat `server.crt` (single-namespace), and a `namespace` column in any `*.db` (R-002 says no namespace column). Output: list of detected legacy artifacts with migration recommendations. This is the *detection* half of v1.0-P14; the *migration* half is I-C-001.
- **Rationale**: §23 v1.0-P14 says "orca upgrade --to-v1.0, post-invariant checks" but doesn't specify the detection surface. `doctor` is the diagnostics framework and is explicitly adaptable.
- **Proposed REQ ID**: REQ-065
- **Proposed phase placement**: v1.0-P14c (mixed-version tolerance + no-orca enforcement)
- **Confidence**: 0.80
- **Accept/Defer**: accept
### I-M-006 — Legacy CA state migration to step-ca (cert import)
- **Tier**: mechanical
- **Description**: `internal/security/ca.go` (338 LOC) holds an internal Go CA with `ca.crt`/`ca.key` (RSA 3072, 10-year). The PRD replaces this with step-ca (R-006, D-101 reverses AD-010). The v1.0-P14 migration must handle existing deployments with an internal CA: (a) import the existing CA key into step-ca as `step ca init --deployment-type standalone --remote-management` with the existing key; (b) issue new SVIDs from step-ca and let old certs expire; (c) document that v0.8 certs are invalidated and re-bootstrap is required. The codebase audit says `ca.go`+`csr.go` are *replaced* — but the *state* (the CA key + issued server certs in `cert_repo` SQLite) may need to be preserved for audit history even if the live trust root changes. Proposal: `orca upgrade --to-v1.0 --import-ca` reads `~/.orca/ca.key`, initializes step-ca with it, and re-issues workload SVIDs. Without this, existing deployments lose their trust root with no path back.
- **Rationale**: AD-010 is explicitly reversed by D-101, but the reversal doesn't address what happens to the existing CA material. §24 covers data migration but not CA migration.
- **Proposed REQ ID**: REQ-066
- **Proposed phase placement**: v1.0-P14a (data migration)
- **Confidence**: 0.70
- **Accept/Defer**: accept (design in v0.9-P00 so step-ca integration knows the import contract)
### I-M-007 — Fuzz test harness for the Markdown frontmatter parser
- **Tier**: mechanical
- **Description**: R-014/R-015 require byte-exact body preservation — "body of every .md config file preserved verbatim." This is a class of bug that's easy to get wrong (off-by-one on the `---` delimiter, trailing newline handling, BOM, CRLF, nested code fences containing `---`). v0.8 has no fuzz tests at all. Proposal: add a `testing.F` fuzz target in `internal/jobspec/markdown_test.go` that round-trips random frontmatter+body through `ParseMarkdown` and asserts `body == roundtripped.body` byte-exact. Also add a corpus of adversarial fixtures (CRLF, BOM, no-frontmatter, empty-frontmatter, frontmatter-with-only-separator). §23 v1.0-P08 mentions integration tests but not fuzzing.
- **Rationale**: R-015 is a *load-bearing invariant* (body appears in inspect/history). Byte-exactness is exactly what fuzz tests are for. The v0.8 jobspec tests are golden-file only (no fuzz).
- **Proposed REQ ID**: REQ-067
- **Proposed phase placement**: v0.9-P0b (Markdown parser) — fuzz from day one
- **Confidence**: 0.78
- **Accept/Defer**: accept
### I-M-008 — Deprecation warnings on removed/repurposed CLI subcommands
- **Tier**: mechanical
- **Description**: The v0.8 CLI has `orca cert {ca-init,gen,show,renew,fingerprint}` (`internal/cli/cert.go`), `orca node join` with mTLS handshake semantics (`internal/cli/node.go`), `orca job run <spec.hcl>`. The PRD repurposes `node join` to SSH-bootstrap (no mTLS), deprecates `cert` (step-ca handles it), and changes `job run` to accept `.md` specs. Each removed/changed command should emit a `slog.Warn` deprecation banner with the v1.0 replacement, *except* when run under `orca upgrade`. The existing `root.go` `PersistentPreRunE` is the natural hook for a global `--no-deprecation-warnings` flag.
- **Rationale**: Operators running v0.8 commands against v0.9/v1.0 need to know what changed. The PRD doesn't mention deprecation UX.
- **Proposed REQ ID**: REQ-068
- **Proposed phase placement**: v0.9-P0X (ship) + v1.0-P13 (ns subcommands, when CLI surface is finalized)
- **Confidence**: 0.72
- **Accept/Defer**: accept
### I-M-009 — `internal/config/config.go` HCL config demotion via adapter
- **Tier**: mechanical
- **Description**: `internal/config/config.go` (127 LOC) parses HCL config with keys `db_path, listen_addr, ca_path, server_cert_path, server_key_path, node_capacity`. The PRD replaces this with Markdown-frontmatter config (R-014) + per-namespace `.env`/`.env.secrets` (R-011). The `listen_addr` and `server_*_path` keys are daemon-specific (deprecated by R-001). The `root.go` `PersistentPreRunE` calls `config.Load(configPath)` on every command — must be repointed to the new Markdown config loader. Proposal: keep `internal/config/` as `legacy_config.go` with `// Deprecated`; add `internal/config/markdown.go` for the new loader; `root.go` dispatches on file extension (`.hcl`→legacy, `.md`→new). The `--config` flag semantics change: `.hcl` is read-only legacy, `.md` is canonical.
- **Rationale**: R-014 makes Markdown canonical but `.hcl` must still parse during migration. The existing `config.Load` is called unconditionally in `root.go:42-47`.
- **Proposed REQ ID**: REQ-069
- **Proposed phase placement**: v0.9-P0a1 (path resolver + config demotion)
- **Confidence**: 0.76
- **Accept/Defer**: accept
### I-M-010 — `internal/certpaths/` replacement with multi-namespace path resolver
- **Tier**: mechanical
- **Description**: `internal/certpaths/certpaths.go` (64 LOC) returns flat paths: `Dir() = $ORCA_HOME`, `CACertPath() = Dir/ca.crt`, `DBPath() = Dir/orca.db`. R-002 requires multi-namespace layout: `ORCA_HOME/<ns>/db/`, `ORCA_HOME/cluster/{ca,master.key,peers,pve,txns}`, `ORCA_HOME/_defaults/`. The package is imported by `doctor`, `proxmox`, `store`, `cli` — changing it is cross-cutting. Proposal: replace `certpaths` with a new `internal/paths` package: `paths.NamespaceDir(ns)`, `paths.ClusterDir()`, `paths.CacheDB()`, `paths.MasterKey()`, `paths.NSDb(ns)`, `paths.NSEnv(ns)`, `paths.NSSecrets(ns)`. Keep `certpaths` as a thin shim that calls `paths` with the default namespace for v0.8 compat, then remove the shim post-v1.0.
- **Rationale**: R-002 is foundational and `certpaths` is the single source of path truth. Every adaptable package (`store.Open`, `doctor`, `proxmox`) imports it. Highest-blast-radius mechanical change.
- **Proposed REQ ID**: REQ-070
- **Proposed phase placement**: v0.9-P0a1 (must come first)
- **Confidence**: 0.84
- **Accept/Defer**: accept
### I-M-011 — `internal/store/` schema: per-namespace DBs, drop ns column
- **Tier**: mechanical
- **Description**: R-002 says "No `namespace` column in SQLite." The v0.8 schema has 7 migrations (`0001`..`0007`) with a single `orca.db`. The v1.0 model has one DB per namespace (`<ns>/db/orca.db`) plus a CLI-side cache DB (`orca_cache.db`, R-008). The existing `store.Open(path)` takes a path arg — adaptable. But the migrations are global; they need to apply *per namespace DB*. Proposal: `store.Open` gains a namespace parameter (or caller passes `paths.NSDb(ns)`); `migrate.go` runs `0001`..`0007` (minus `0006_node_kind_os` which is v0.8-specific) plus new `0008_namespace_layout.sql`. The `cert_repo` (`0004_certs.sql`) is removed (step-ca handles certs). The audit_log table moves to the CLI-side cache DB (R-008). Existing v0.8 `orca.db` is migrated by splitting tables into per-namespace DBs during v1.0-P14.
- **Rationale**: R-002 is explicit ("No namespace column in SQLite") but the existing schema has a single DB. §23 doesn't specify the schema split mechanics.
- **Proposed REQ ID**: REQ-071
- **Proposed phase placement**: v0.9-P0a1 + v1.0-P06 (alloc history, which uses cache DB)
- **Confidence**: 0.80
- **Accept/Defer**: accept
### I-M-012 — `internal/transport/` deletion + SSH-push package introduction
- **Tier**: mechanical
- **Description**: `internal/transport/` (7 files, ~1300 LOC incl tests) implements mTLS client/server, dispatch, idempotency, retry, handshake logging. R-001 + R-006 replace this with SSH-push. The *idempotency* and *retry* logic (`idempotency.go` 123 LOC, `retry.go` 151 LOC) is conceptually reusable for SSH-push (retry on SSH failure, idempotency keys for SCP'd configs). Proposal: delete `mtls.go`, `dispatch.go`, `handshake_log.go`; extract retry/idempotency patterns into a new `internal/sshpush/` package. The existing `transport.IdempotencyStore` (in-memory `sync.Map` of keys) is directly reusable. This avoids re-implementing retry semantics from scratch.
- **Rationale**: The codebase audit marks `internal/transport/` as fully replaced, but the retry/idempotency *patterns* are transport-agnostic. §23 doesn't call this out.
- **Proposed REQ ID**: REQ-072
- **Proposed phase placement**: v0.9-P00 (deprecation sweep) — delete in v1.0-P14
- **Confidence**: 0.68
- **Accept/Defer**: accept (defer deletion to v1.0-P14 to keep dual-write window open)
## Tier 2 — Backend-Enriched (Structural / Architectural)
### I-B-001 — SSH-push transport layer design
- **Tier**: backend-enriched
- **Description**: The PRD replaces `internal/transport/` (mTLS HTTP) with SSH-push but §23 never specifies the transport's internal design. Key decisions: (1) **Connection pooling**: reuse `*ssh.Client` per peer across multiple SCP/exec operations within a single CLI invocation. (2) **Idempotency**: SCP of a config file is idempotent if content hash matches — use content-addressed filename (`/run/orca/<hash>.unit`) and skip if present. (3) **Retry**: reuse v0.8's exponential backoff (100ms start, ×2, cap 5s, max 5 attempts) applied to SSH dial/exec failures. (4) **Timeout**: per-operation `context.WithTimeout` (default 30s SCP, 10s exec). (5) **Fan-out**: `errgroup.Group` with bounded concurrency for N-peer ops (default 8). (6) **known_hosts**: reuse `proxmox.TOFUHostKeyCallback` for all peers, not just Proxmox.
- **Rationale**: Load-bearing replacement for the entire v0.8 transport layer. §23 assumes it but never designs it. Without connection pooling, every CLI operation re-dials SSH.
- **Proposed REQ ID**: REQ-073
- **Proposed phase placement**: v0.9-P01 (first phase needing SSH-push) — design in v0.9-P0a1
- **Confidence**: 0.86
- **Accept/Defer**: accept
### I-B-002 — Emitter template system (Layer 4)
- **Tier**: backend-enriched
- **Description**: The PRD §5 describes a 4-layer architecture where Layer 4 is "emitters" that render systemd units, Traefik dynamic config, Syncthing config, etc. from the workload spec. §23 never specifies the emitter interface. Proposal: an `internal/emitter/` package with `Emitter` interface: `Render(spec *WorkloadSpec, node *Node) ([]File, error)` where `File{Path, Content, Mode}`. Implementations: `systemdEmitter`, `traefikEmitter`, `syncthingEmitter`, `socketEmitter`. The SSH-push transport SCPs the `[]File` atomically (write-to-tmp + rename). Emitters registered per workload kind + runtime.
- **Rationale**: The emitter layer is the bridge between the declarative spec and the server-side files. Without a defined interface, each phase (P02 service, P04 hooks, P08 sockets, P09 storage) will invent its own rendering.
- **Proposed REQ ID**: REQ-074
- **Proposed phase placement**: v0.9-P0c (schemas + emitter interface)
- **Confidence**: 0.82
- **Accept/Defer**: accept
### I-B-003 — Lead applier execution model: pure bash + systemd timer vs CLI-invoked
- **Tier**: backend-enriched
- **Description**: R-001 says "no orca binary on servers." R-010 says the lead applies desired-state transactionally. Unresolved: does the lead run `orca-pull.sh` (pure bash that SCPs a desired-state bundle and applies it via `systemctl daemon-reload` + `systemctl restart`) or does the operator's CLI SSH into the lead and runs `orca apply` remotely (which would put an orca binary on the lead, violating R-001)? The PRD's intent is the former: the lead is bare Linux with systemd timers + bash. Proposal: (1) the CLI renders a *transaction bundle* (tarball of desired-state files + `apply.sh` + `verify.sh`) on the operator host; (2) SCPs it to the lead's `/run/orca/txns/<txn-id>/`; (3) the lead's systemd timer runs `/run/orca/txns/<txn-id>/apply.sh` which idempotently applies and runs verify; (4) the CLI polls the lead for txn status via SSH (`cat /run/orca/txns/<txn-id>/status.json`). The bash scripts are generated by the CLI's emitter (I-B-002), not hand-written per cluster.
- **Rationale**: The most ambiguous load-bearing design decision in the PRD. R-001 + R-010 together imply the lead runs no orca binary, but the lead must apply transactions. §23 doesn't resolve this. Getting it wrong means either violating R-001 or having no transactional apply.
- **Proposed REQ ID**: REQ-075
- **Proposed phase placement**: v1.0-P10 (transactional plane) — bundle format designed in v0.9-P00
- **Confidence**: 0.78
- **Accept/Defer**: accept
### I-B-004 — step-ca integration: provisioning, CA bootstrap, cert signing API, SVID minting
- **Tier**: backend-enriched
- **Description**: D-101 reverses AD-010 (which rejected step-ca as "too heavyweight"). §23 mentions step-ca in R-006 but never specifies the integration. Key surfaces: (1) **Provisioning**: `orca init` (adapted from v0.8's `internal/cli/init.go`) runs `step ca init` on the lead, stores root + intermediate in `cluster/ca/`. (2) **CA bootstrap**: CLI SSHs to the lead, installs step-ca via apt, runs `step ca init`, stores `step-ca.json` config. (3) **Cert signing API**: workloads request SVIDs via `step ca token` (JWE provisioner token minted by CLI) → `step ca certificate`. The CLI mints the token because it holds the provisioner password (in `cluster/master.key`-derived form). (4) **SVID minting**: each workload gets a SPIFFE ID (`spiffe://orca/<ns>/<workload>/<instance>`) encoded as a SAN in the step-ca-issued cert. The v0.8 `internal/security/ca.go` is deleted; a new `internal/stepca/` package wraps the `step` CLI via SSH (no Go step-ca client library — keep zero-new-dep posture if possible, or add `github.com/smallstep/cli` as a dep).
- **Rationale**: step-ca is a new external dependency with its own config format, provisioner model, and CLI. §23 assumes it but never designs the integration. security-engineer persona must be reactivated.
- **Proposed REQ ID**: REQ-076
- **Proposed phase placement**: v0.9-P07 (runtime block — runtimes need SVIDs) + v1.0-P02 (ACL — SPIFFE identities)
- **Confidence**: 0.74
- **Accept/Defer**: accept
### I-B-005 — Traefik dynamic config generation and atomic reload
- **Tier**: backend-enriched
- **Description**: R-006 makes Traefik load-bearing (mTLS termination + health checks). §23 puts service blocks + Traefik health checks in v0.9-P02. Design: the CLI's Traefik emitter (I-B-002) renders a dynamic config file (`/etc/traefik/dynamic/orca-<ns>-<svc>.yaml`) with backends (the socket paths from R-007), health checks, and mTLS config pointing at step-ca's root. Atomic reload: Traefik watches the dynamic dir with `fsnotify` — writing the file atomically (tmp+rename) triggers a reload. Drain (v1.0-P05) works by writing a config with the backend's `weight=0` or removing it, triggering Traefik to stop routing. The v0.8 codebase has no Traefik integration at all. **Gated by grill C-10** (Traefik config atomicity protocol: tmpfile+fsync+rename + malformed-config hold-last-good verified).
- **Rationale**: Traefik is net-new and load-bearing. §23 mentions it in R-006/P02/P05 but never specifies config generation or reload mechanism.
- **Proposed REQ ID**: REQ-077
- **Proposed phase placement**: v0.9-P02 (Service block + checks)
- **Confidence**: 0.80
- **Accept/Defer**: accept
### I-B-006 — Runtime abstraction interface (5 backends: wasm/podman/process/pve-vm/pve-ct)
- **Tier**: backend-enriched
- **Description**: v0.8's `internal/engine/executor.go` (211 LOC) is `os/exec` only. R-004/R-007 require 5 runtime backends. §23 puts this in v0.9-P07. Proposal: a `Runtime` interface in `internal/runtime/`: `Prepare(ctx, spec, node) (*Alloc, error)`, `Start(ctx, alloc) (pid/unit, error)`, `Stop(ctx, alloc) error`, `Status(ctx, alloc) (State, error)`. Implementations: `processRuntime` (wraps existing `executor.go` — directly reusable), `wasmRuntime` (wasmtime via CLI SSH exec), `podmanRuntime` (`podman run` via SSH), `pveVMRuntime` (`qm create`/`qm start` via v0.8 `proxmox` SSH session), `pveCTRuntime` (`pct create`/`pct start`). Each registered in a `runtimeRegistry` keyed by the `runtime:` frontmatter value. R-004 (migration with runtime change) means the `Alloc` carries a `runtime` field that can change on migration — `Prepare` re-runs with the new runtime.
- **Rationale**: The 5 backends are the largest net-new implementation surface. §23 lists them as one phase (P07) but under-specifies the interface contract. The existing `executor.go` is a good starting point for the `processRuntime` adapter.
- **Proposed REQ ID**: REQ-078
- **Proposed phase placement**: v0.9-P07a/P07b/P07c (split per grill PC-10)
- **Confidence**: 0.82
- **Accept/Defer**: accept
### I-B-007 — Transaction bundle format and atomicity across N peers
- **Tier**: backend-enriched
- **Description**: R-010 requires transactional control-plane updates. §23 puts this in v1.0-P10. Design: a *transaction bundle* is a tarball containing: (1) `desired-state.json` (full desired state for affected namespaces), (2) `apply.sh` (idempotent apply script), (3) `verify.sh` (post-apply invariants), (4) `rollback.sh` (revert to previous state), (5) `manifest.sig` (signature with `cluster/master.key`). Atomicity across N peers: the CLI uploads the bundle to the lead; the lead applies to itself first, then fans out to peers via SSH. If any peer fails verify, the lead runs `rollback.sh` on all peers that applied. The bundle is content-addressed (`<txn-id> = sha256(desired-state.json)`) and stored in `cluster/txns/<txn-id>/`. Drift detection (R-010) compares the last applied bundle's desired-state against the live state (polled via SSH `systemctl show` + file checksums). **Gated by grill C-09** (orca-pull.sh failure contract: idempotent re-run, bounded retry, deterministic state, structured syslog).
- **Rationale**: Multi-peer atomicity is the hardest part of R-010. §23 says "ArgoCD-style" but ArgoCD is Kubernetes-native; the SSH-push model needs a custom bundle format.
- **Proposed REQ ID**: REQ-079
- **Proposed phase placement**: v1.0-P10 (transactional plane) — designed in v0.9-P00
- **Confidence**: 0.76
- **Accept/Defer**: accept
### I-B-008 — Master key management and HKDF-SHA256 per-line .env.secrets encryption
- **Tier**: backend-enriched
- **Description**: R-011 specifies `.env.secrets` with AES-256-GCM, per-line nonce, master key at `cluster/master.key`. §23 puts this in v1.0-P03. Design: (1) `cluster/master.key` is a 32-byte random key generated by `orca init` (extend v0.8 `internal/security/ca.go`'s `WriteAtomic` pattern for the file write). (2) Each line of `.env.secrets` is `base64(nonce || ciphertext || tag)` where `nonce = random(12 bytes)` and `ciphertext = AES-256-GCM(plaintext, key=master.key, nonce, aad=line-number)`. (3) The AAD is the 1-indexed line number to prevent line-swap attacks. (4) Decryption reads the master key, iterates lines, decrypts with AAD. (5) `orca secrets set <ns> <key> <value>` appends an encrypted line; `orca secrets get <ns> <key>` decrypts and prints (redacted by default, `--reveal` to show). (6) The v0.8 `internal/security/redact.go` (103 LOC) is directly reusable for redaction. HKDF-SHA256 derives per-namespace sub-keys from the master key (`HKDF-SHA256(master, info=<ns>)`) so compromising one namespace's key doesn't compromise others — but the master key is the root of trust. **Gated by grill C-19** (threat model for master.key passphrase-less posture).
- **Rationale**: R-011 is precise about the crypto but §23 doesn't specify key derivation, AAD, or CLI surface. The existing `redact.go` and `WriteAtomic` are reusable.
- **Proposed REQ ID**: REQ-080
- **Proposed phase placement**: v1.0-P03 (secrets subsystem)
- **Confidence**: 0.84
- **Accept/Defer**: accept
### I-B-009 — Syncthing config rendering and folder-ID content-addressing
- **Tier**: backend-enriched
- **Description**: R-005 requires storage replication via per-namespace Syncthing. §23 puts this in v0.9-P09. Design: (1) each namespace gets a Syncthing folder `orca-<ns>` with a content-addressed folder ID (`sha256(ns + master-key-fingerprint)`). (2) The CLI renders `config.xml` for each peer's Syncthing instance, including the folder, devices (all peers in the namespace), and the path (`<ns>/alloc/<alloc-id>/`). (3) Syncthing runs as a systemd unit (emitted by the systemd emitter, I-B-002). (4) The CLI discovers peers via `cluster/peers/` and adds their Syncthing device IDs (each peer's Syncthing generates its own device key on first run, reported back via SSH). (5) R-005 says "a Service's count replicas share one runtime block" — the Syncthing folder is shared across the Service's alloc instances so all replicas see the same data. Migration (R-004) works because the new node joins the Syncthing folder and syncs before the workload starts. **Gated by grill C-02** (Syncthing feasibility spike) and **C-14** (deterministic conflict-resolution policy + forced-divergence integration test).
- **Rationale**: Syncthing is net-new. §23 lists it in P09 but doesn't specify config rendering, folder-ID scheme, or device discovery.
- **Proposed REQ ID**: REQ-081
- **Proposed phase placement**: v0.9-P09 (storage replication) — spike in v0.9-P00
- **Confidence**: 0.72
- **Accept/Defer**: accept
### I-B-010 — Namespace inheritance resolver algorithm
- **Tier**: backend-enriched
- **Description**: v0.9-P0a2 requires a "parent walker, cycle detection" for namespace inheritance. Each `ns.md` has a `parent:` field in frontmatter. The resolver walks up the parent chain, merging inherited values (constraints, env, runtime defaults). Cycle detection: DFS with a visited set; if a namespace is revisited, return a cycle error. The resolver returns a flattened `ResolvedNamespace` struct. The `_defaults/` namespace is the implicit root (always exists, has no parent). Inheritance semantics: child overrides parent for scalar fields; arrays (e.g., constraints) are unioned (child adds to parent, not replaces). The resolver is pure (no I/O) — it takes a map of `nsName → *NSConfig` and returns `nsName → *ResolvedNS`. This makes it trivially testable.
- **Rationale**: §23 mentions "parent walker, cycle detection" but not the merge semantics (override vs union) or the resolver's purity for testing. Getting merge semantics wrong breaks constraint inheritance (P05).
- **Proposed REQ ID**: REQ-082
- **Proposed phase placement**: v0.9-P0a2 (namespace CRUD + inheritance)
- **Confidence**: 0.86
- **Accept/Defer**: accept
### I-B-011 — Bin-packing scheduler redesign (CLI-side, runtime-compatibility scoring)
- **Tier**: backend-enriched
- **Description**: v0.8's `internal/engine/scheduler.go` (117 LOC) does best-fit bin-packing by CPU+memory. The v0.9 scheduler must: (1) run CLI-side (not on a daemon), (2) score nodes by runtime compatibility (a wasm workload can only go to a node with wasmtime installed; a pve-vm workload can only go to Proxmox nodes), (3) respect constraints/affinity (CEL over node attributes, P05), (4) handle the 3 kinds differently (Job = one-shot, Service = count replicas spread across nodes, DaemonSet = one per node). The existing `scheduler.go` is a good skeleton but the scoring function changes entirely. Proposal: `Score(node, workload) (score int, fits bool)` where `fits` checks runtime compatibility + constraints, and `score` is the bin-packing score (most free capacity = highest score). For Services, the scheduler picks `count` distinct nodes (anti-affinity by default). For DaemonSets, it picks all matching nodes.
- **Rationale**: The scheduler moves from daemon-side to CLI-side (R-001) and gains runtime-awareness. §23 scatters this across P05 (constraints), P06 (task groups), P07 (runtime), P10 (migration) but never designs the scheduler itself.
- **Proposed REQ ID**: REQ-083
- **Proposed phase placement**: v0.9-P05 (constraints & affinity — scheduler needs constraints to be meaningful) — skeleton in P0c
- **Confidence**: 0.80
- **Accept/Defer**: accept
### I-B-012 — `orca job lint` category-driven lint engine design
- **Tier**: backend-enriched
- **Description**: v1.0-P11 requires `orca job lint` with `--explain`. Design: a `Linter` that takes a `*WorkloadSpec` and runs a series of `Rule` checks, each returning a `Finding{Category, Severity, Message, Explanation}`. Categories: `schema` (missing required fields), `runtime` (incompatible runtime+constraint), `security` (missing SVID, plaintext secret in env), `migration` (missing storage replication for a migratable service), `best-practice` (no health check on a Service). `--explain` prints the rationale for each finding. Rules are registered in a `ruleRegistry` and individually testable. The linter is pure (no I/O) — it checks the spec against static rules, not live cluster state (that's `orca job verify`, P12).
- **Rationale**: §23 puts this in P11 but only says "category-driven." The rule interface and category taxonomy are unspecified.
- **Proposed REQ ID**: REQ-084
- **Proposed phase placement**: v1.0-P11 (orca job lint)
- **Confidence**: 0.78
- **Accept/Defer**: accept
## Tier 3 — Cross-Cutting (Risk & Multi-Phase)
### I-C-001 — v0.8→v1.0 migration ordering: daemon deprecation vs. new model rollout
- **Tier**: cross-cutting
- **Description**: The PRD §24 covers *data* migration but not *binary/daemon* deprecation ordering. The risk: v0.9 builds the new Markdown+kinds+runtime+SSH-push model, but v0.8 daemons are still running on peers. If v0.9 ships the new `orca job run` (Markdown) while the old daemon is still the execution engine, there's a split-brain: new specs can't run on the old daemon. Ordering proposal: (1) v0.9 ships the new parser + kinds + runtime + SSH-push *alongside* the old daemon (dual-write window); (2) `orca job run` in v0.9 uses the new SSH-push path if the spec is `.md` and the old daemon path if `.hcl`; (3) v1.0-P05 (drain) stops the old daemons; (4) v1.0-P14 (migration) converts remaining `.hcl` specs to `.md` and removes the daemon. The dual-write window means v0.9 is *not* a clean break — it's a compatibility milestone. This must be explicit in the plan or the v0.9 phases will assume the daemon is gone.
- **Rationale**: Single largest risk in the re-architecture. §23 implicitly assumes v0.9 builds the new model in isolation, but existing deployments have running daemons. Getting the ordering wrong means either (a) v0.9 can't be tested against real deployments, or (b) workloads are orphaned when the daemon is removed.
- **Proposed REQ ID**: REQ-085
- **Proposed phase placement**: spans v0.9-P00 through v1.0-P14 — the *ordering decision* must be made in v0.9-P00
- **Confidence**: 0.88
- **Accept/Defer**: accept (most important idea in this report)
### I-C-002 — "No orca on server" enforcement (doctor post-migration invariant check)
- **Tier**: cross-cutting
- **Description**: R-001 is an invariant: "no orca Go binary on any server." §23 v1.0-P14 says "post-invariant checks" but doesn't specify them. `orca doctor` must gain a `doctor no-orca-on-server` check that SSHs to each peer and verifies: (1) no `orca` binary in PATH (`ssh peer which orca` returns nothing), (2) no `orca` systemd service (`ssh peer systemctl list-units 'orca*'` returns empty), (3) no `orca` process (`ssh peer pgrep -x orca` returns empty), (4) no `/etc/orca/` directory. This check must run *after* v1.0-P05 (drain) and *before* v1.0-P16 (ship). The v0.8 `internal/proxmox/bootstrap.go` already has the SSH session infrastructure (`sessionRunner` seam) — directly reusable for the doctor check.
- **Rationale**: R-001 is a hard invariant but §23 doesn't enforce it post-migration. Without this check, a failed migration could leave orphaned daemons that cause split-brain.
- **Proposed REQ ID**: REQ-086
- **Proposed phase placement**: v1.0-P14c (mixed-version tolerance)
- **Confidence**: 0.82
- **Accept/Defer**: accept
### I-C-003 — Test infrastructure: hermetic 3-linux + 1-proxmox cluster pipeline
- **Tier**: cross-cutting
- **Description**: §23 v1.0-P08 requires "hermetic CoreCI integration pipeline." The PRD §26.E mentions 3 linux + 1 proxmox. This is net-new test infra with zero current implementation. Design: (1) a `test/integration/` directory with a `docker-compose.yml` or `vagrant` setup that creates 4 containers/VMs (3 linux + 1 proxmox-simulated); (2) a Go test harness that SSHes to each, runs the CLI, and asserts end-to-end workflows (namespace create → workload submit → migrate → drain); (3) the proxmox node is simulated via a mock `pct`/`qm` script (the v0.8 `proxmox` package already has a `sessionRunner` seam for testability — extend it). The integration tests run in CoreCI on every milestone merge. The v0.8 e2e tests (`bootstrapE2ESetup` in `bootstrap_test.go`) use an in-process SSH server — this is the foundation but needs to scale to 4 nodes.
- **Rationale**: §23 assumes the infra exists but doesn't design it. devops-engineer persona should be reactivated. Without hermetic infra, the integration tests can't run in CI.
- **Proposed REQ ID**: REQ-087
- **Proposed phase placement**: v1.0-P08 (integration tests) — harness bootstrapped in v0.9-P00
- **Confidence**: 0.80
- **Accept/Defer**: accept
### I-C-004 — Security-engineer + network-engineer persona reactivation for new attack surfaces
- **Tier**: cross-cutting
- **Description**: The config.json has `security-engineer` and `network-engineer` dormant. The re-architecture introduces step-ca (PKI), Traefik (edge proxy), Syncthing (P2P file sync), wasmtime (sandbox), podman (container runtime) — all new attack surfaces. AD-010 (step-ca rejection) is reversed. The v0.8 security posture (internal CA, mTLS daemon-to-daemon) is replaced by (step-ca, SSH-push, Traefik mTLS). The security-engineer persona must be reactivated to review: (1) step-ca provisioner model (the CLI holds the provisioner password — is that in `cluster/master.key` or a separate secret?), (2) SSH-push blast radius (compromised CLI key = full cluster), (3) Traefik as the new edge (DoS, config injection), (4) `.env.secrets` crypto (I-B-008). The network-engineer persona must review: (1) socket-based service exposure (R-007), (2) Syncthing P2P ports, (3) Traefik routing. §23 doesn't mention persona reactivation.
- **Rationale**: config.json explicitly notes the re-architecture "should reactivate security-engineer and network-engineer." Cross-cutting review concern, not a single phase.
- **Proposed REQ ID**: REQ-088
- **Proposed phase placement**: spans v0.9 through v1.0 — reactivation in v0.9-P00, review at v1.0-P15.5 (threat model) and v1.0-P16 (final audit)
- **Confidence**: 0.84
- **Accept/Defer**: accept
### I-C-005 — Documentation rewrite: ARCHITECTURE.md, PROJECT.md, README, AD-010 supersession
- **Tier**: cross-cutting
- **Description**: All three docs describe the OLD architecture. `ARCHITECTURE.md` (640 lines) describes the daemon layer, mTLS transport, internal CA, HCL jobspec — all deprecated. `PROJECT.md` (30k chars) has D-001..D-010 decisions, several now superseded. `README.md` has the v0.8 quickstart. AD-010 (step-ca rejection) must be explicitly superseded by D-101 with a dated rationale reversal. The anti-patterns section in `ARCHITECTURE.md:471-484` lists "No external PKI" — now reversed. Proposal: (1) in v0.9-P00, add a "v0.9 Architecture (Supersedes v0.8)" section to ARCHITECTURE.md with the new 4-layer model; (2) mark the old sections as "v0.8 (deprecated)" with banners; (3) add a "Superseded Decisions" table (AD-009, AD-010 reversed by D-101; AD-007 HCL demoted by R-013); (4) in v1.0-P15, rewrite README quickstart for the new `curl | sh` + `orca init` + `orca ns create` flow.
- **Rationale**: The docs are the first thing new contributors read. Leaving v0.8 docs as canonical during v0.9 development causes confusion. §23 mentions README in P15 but not ARCHITECTURE.md/PROJECT.md.
- **Proposed REQ ID**: REQ-089
- **Proposed phase placement**: v0.9-P00 (banners + supersession table) + v1.0-P15 (README quickstart) + v1.0-P16 (final review)
- **Confidence**: 0.82
- **Accept/Defer**: accept
### I-C-006 — Dual-write window: can v0.9 ship new parser while old daemon runs?
- **Tier**: cross-cutting
- **Description**: Focused version of I-C-001. The specific question: in v0.9, when the new Markdown parser + kinds + SSH-push are shipped, can they coexist with v0.8 daemons still running on peers? The answer depends on whether `orca job run <spec.md>` uses the new SSH-push path (bypassing the daemon entirely) or routes through the old daemon. If it bypasses, the daemon is irrelevant for new specs but still serves old `.hcl` specs. If it routes through, the daemon can't handle `.md` specs. Proposal: v0.9 `orca job run` dispatches on extension (`.md`→SSH-push new path, `.hcl`→old daemon path) via the parser dispatcher (I-M-004). This is a *dual-write window* where both paths coexist. The daemon is not removed until v1.0-P05 (drain). The risk: if a `.md` workload and a `.hcl` workload target the same node, the SSH-push path writes systemd units directly while the daemon also manages units — they can conflict. Mitigation: the SSH-push path writes to a separate systemd unit namespace (`orca-v1-<alloc>.service`) while the daemon uses `orca-<job>.service`. No unit name overlap = no conflict.
- **Rationale**: Operational feasibility question for v0.9. §23 doesn't address it. If the answer is "no dual-write, daemon must be removed first," then v0.9 can't be tested incrementally and must ship as a big-bang — much higher risk.
- **Proposed REQ ID**: REQ-090
- **Proposed phase placement**: v0.9-P00 (decision before any v0.9 execution phase)
- **Confidence**: 0.86
- **Accept/Defer**: accept
## Summary Table
| ID | Tier | Title | REQ | Phase | Conf | Accept |
|----|------|-------|-----|-------|------|--------|
| I-M-001 | M | `orca daemon` deprecation path | REQ-061 | v1.0-P14 (warn v0.9-P0X) | 0.82 | accept |
| I-M-002 | M | Coverage follow-ups to 70% | REQ-062 | v0.9-P0X + each new pkg | 0.88 | accept |
| I-M-003 | M | known_hosts flock concurrency | REQ-063 | v0.9-P0a1 | 0.74 | accept |
| I-M-004 | M | HCL→Markdown jobspec adapter | REQ-064 | v0.9-P0b | 0.85 | accept |
| I-M-005 | M | `doctor --legacy-paths` detection | REQ-065 | v1.0-P14c | 0.80 | accept |
| I-M-006 | M | Legacy CA state migration to step-ca | REQ-066 | v1.0-P14a | 0.70 | accept |
| I-M-007 | M | Fuzz harness for Markdown parser | REQ-067 | v0.9-P0b | 0.78 | accept |
| I-M-008 | M | Deprecation warnings on CLI subcommands | REQ-068 | v0.9-P0X + v1.0-P13 | 0.72 | accept |
| I-M-009 | M | HCL config demotion via adapter | REQ-069 | v0.9-P0a1 | 0.76 | accept |
| I-M-010 | M | certpaths → multi-namespace path resolver | REQ-070 | v0.9-P0a1 | 0.84 | accept |
| I-M-011 | M | store schema: per-namespace DBs | REQ-071 | v0.9-P0a1 + v1.0-P06 | 0.80 | accept |
| I-M-012 | M | transport deletion + SSH-push package | REQ-072 | v0.9-P00 (delete v1.0-P14) | 0.68 | accept |
| I-B-001 | B | SSH-push transport layer design | REQ-073 | v0.9-P01 | 0.86 | accept |
| I-B-002 | B | Emitter template system (Layer 4) | REQ-074 | v0.9-P0c | 0.82 | accept |
| I-B-003 | B | Lead applier execution model | REQ-075 | v1.0-P10 (design v0.9-P00) | 0.78 | accept |
| I-B-004 | B | step-ca integration | REQ-076 | v0.9-P07 + v1.0-P02 | 0.74 | accept |
| I-B-005 | B | Traefik dynamic config + atomic reload | REQ-077 | v0.9-P02 | 0.80 | accept |
| I-B-006 | B | Runtime abstraction (5 backends) | REQ-078 | v0.9-P07a/b/c | 0.82 | accept |
| I-B-007 | B | Transaction bundle + N-peer atomicity | REQ-079 | v1.0-P10 (design v0.9-P00) | 0.76 | accept |
| I-B-008 | B | Master key + HKDF per-line encryption | REQ-080 | v1.0-P03 | 0.84 | accept |
| I-B-009 | B | Syncthing config + folder-ID | REQ-081 | v0.9-P09 | 0.72 | accept |
| I-B-010 | B | Namespace inheritance resolver | REQ-082 | v0.9-P0a2 | 0.86 | accept |
| I-B-011 | B | CLI-side scheduler redesign | REQ-083 | v0.9-P05 (skeleton P0c) | 0.80 | accept |
| I-B-012 | B | `orca job lint` category-driven engine | REQ-084 | v1.0-P11 | 0.78 | accept |
| I-C-001 | C | v0.8→v1.0 migration ordering | REQ-085 | spans v0.9-P00→v1.0-P14 | 0.88 | accept |
| I-C-002 | C | "No orca on server" enforcement | REQ-086 | v1.0-P14c | 0.82 | accept |
| I-C-003 | C | Hermetic test infra (3 linux + 1 pve) | REQ-087 | v1.0-P08 (bootstrap v0.9-P00) | 0.80 | accept |
| I-C-004 | C | security/network persona reactivation | REQ-088 | spans v0.9→v1.0-P16 | 0.84 | accept |
| I-C-005 | C | Docs rewrite + AD-010 supersession | REQ-089 | v0.9-P00 + v1.0-P15/P16 | 0.82 | accept |
| I-C-006 | C | Dual-write window decision | REQ-090 | v0.9-P00 | 0.86 | accept |
## Phase Reordering / Addition Flags (against PRD §23)
1. **I-C-001 / I-C-006 (dual-write + migration ordering)** — require a decision in v0.9-P00 (before any execution phase). **Recommendation: add v0.9-P00 deprecation/migration-ordering pre-phase.** Most important structural addition.
2. **I-M-010 / I-M-011 / I-M-009 / I-M-003** — all land in v0.9-P0a. P0a may be overloaded. **Recommendation: split P0a into P0a1 (path/layout resolver + config demotion) and P0a2 (namespace CRUD + inheritance).** Path resolver is prerequisite for everything; highest blast radius.
3. **I-B-001 (SSH-push transport)** — §23 v0.9-P01 needs SSH-push. The design is a prerequisite. **Recommendation: SSH-push design in P0a1, not deferred to P01.**
4. **I-B-002 (emitter template system)** — should be designed *with* the schemas (P0c). **Recommendation: expand P0c to "schemas + emitter interface."**
5. **I-B-003 (lead applier model)** — bundle format + lead applier model must be designed *in v0.9* so the emitter can produce bundle-compatible output. **Recommendation: design spike in v0.9-P00.**
6. **I-C-003 (test infra)** — hermetic cluster harness should be bootstrapped in v0.9-P00 so every v0.9 phase can run integration tests. **Recommendation: bootstrap in v0.9-P00, expand in v1.0-P08.**
7. **I-C-004 / I-C-005 (persona reactivation + docs)** — span the whole milestone. **Recommendation: fold persona reviews into v0.9-P00 and v1.0-P16; fold doc banners into v0.9-P00.**
## Cross-Reference Against Existing Decisions
- **AD-009 (Internal CA, no external PKI)** — Superseded by D-101 (step-ca). I-B-004, I-M-006 implement the reversal.
- **AD-010 (Roll-our-own CA)** — Superseded by D-101. I-C-005 documents the supersession. No re-litigation — the PRD has decided; the override justification records the evidence basis.
- **AD-007 (HCL for job specs)** — Demoted by R-013 (Markdown canonical, HCL accepted). I-M-004 implements the adapter. Not a full reversal — HCL still parses.
- **AD-001 (Single binary with subcommands)** — Still holds. The CLI is the single binary; no orca on servers (R-001) refines this.
- **AD-015 (Best-fit bin-packing)** — Extended, not reversed. I-B-011 adds runtime-compatibility scoring.
- **D-035 (TOFU host-key)** — Still holds for non-Proxmox peers. I-M-003 hardens the concurrency. I-B-001 reuses `TOFUHostKeyCallback`.
- **D-046 (key-reset is local-only)** — Still holds. I-M-003 adds the lock.
- **D-047 (tiered coverage floor)** — Extended by I-M-002 to cover new packages.
No accepted idea re-litigates a settled decision. All reversals (AD-009, AD-010, SPIFFE, no-container, no-multi-tenancy, HCL-canonical, daemon-on-every-node) are explicitly mandated by the PRD and justified by the recorded override justification.
## Final Notes
- **Total ideas**: 30 (12 mechanical, 12 backend-enriched, 6 cross-cutting).
- **Highest-confidence, highest-impact**: I-C-001 (migration ordering, 0.88) and I-C-006 (dual-write window, 0.86) — these shape the entire v0.9 execution strategy.
- **Highest-blast-radius mechanical**: I-M-010 (path resolver, 0.84) — touches every adaptable package.
- **Most under-specified by PRD**: I-B-003 (lead applier execution model, 0.78) — R-001 + R-010 create a tension the PRD doesn't resolve.
+54 -41
View File
@@ -3,57 +3,70 @@ active:
- lead-developer
- backend-engineer
- data-engineer
- security-engineer
- network-engineer
- devops-engineer
deactivated:
- cli-engineer
- security-engineer
- devops-engineer
- network-engineer
- frontend-engineer
phase_specific: []
reason: |
Orca v0.8 is an NFR coverage & trust-hardening milestone. The work is
test coverage uplift across 9 packages (P01), SSH trust-surface
hardening in the existing proxmox + cli/node + security packages (P02),
and a requirements-hygiene Go program + Makefile target (P03). No
schema changes, no new security architecture, no packaging/distribution,
no UI.
Orca v0.9 is the first DIRECTION-CHANGE milestone in the project's
history. It supersedes the shipped v0.1v0.8 architecture per the adopted
PRD (.ciagent/PRD_v0.9.md). The re-architecture deprecates the daemon/
transport/internal-CA/HCL/single-namespace stack and builds a CLI-only/
SSH-push/step-ca/Markdown-frontmatter/multi-namespace stack plus 8
net-new subsystems. The user overrode the grill's Re-architecture
Justification REPLAN with a six-part evidence basis (see PROJECT.md
Supersession Table). The ci-griller's 19 binding conditions (C-01..C-19)
and 10 phase challenges (PC-01..PC-10) are adopted as execution gates
(see GRILL_v0.9.md).
Roster changes vs v0.7:
- lead-developer: RETAINED — owns cmd/orca smoke test, internal/cli
coverage (cert/doctor/audit/status/version subcommands), and the
cmd/verify-reqs Go program (coordination + glue-code territory).
- backend-engineer: RETAINED — owns internal/transport + internal/engine
tests (httptest.NewTLSServer, LocalExecutor stubs, PeerRegistry) and
the SSH trust-surface in internal/proxmox/bootstrap.go (pinned
host-key callback, TOFU capture fix, sessionRunner seam) plus
internal/cli/node.go (--host-key-fingerprint flag, key-reset
subcommand). Frameworks updated: connectrpc REMOVED (not in go.mod
per AD-014), golang.org/x/crypto/ssh ADDED (direct dep since v0.6).
- data-engineer: RETAINED — owns internal/store tests (cert_repo_test.go
gap + coverage uplift), internal/audit tests (sqlite-backed
audit_log asserts), internal/certpaths tests (path-join asserts),
and internal/jobspec tests (golden HCL fixtures). Frameworks
updated: modernc/sqlite + iter (matches actual go.mod).
- security-engineer: remains DEACTIVATED — v0.8 refines the existing
proxmox SSH trust surface (pinned callback, key-reset) but does NOT
add new security architecture. The trust work is backend-engineer
territory (it's SSH dialer + known_hosts file manipulation, not
X.509/CA/crypto code).
- cli-engineer: remains DEACTIVATED — merged into lead-developer
(cli coverage is test-only; --host-key-fingerprint and key-reset
are 1-flag + 1-subcommand additions to the existing node.go).
- devops-engineer: remains DEACTIVATED — verify-reqs is a Go program
(lead-developer territory), not a CI/packaging change. The
.coreci.yml edit is a 3-line validate-pipeline hook.
- network-engineer: remains DEACTIVATED — no transport/mTLS surface
change (transport coverage is test-only on the existing mTLS layer).
- frontend-engineer: remains DEACTIVATED — no web UI (unchanged
from v0.1 onward).
Roster changes vs v0.8 (implements grill C-05):
- lead-developer: RETAINED — owns the CLI subcommand tree, deprecation
sweep (P00), path resolver (P0a1), parser dispatch (P0b), emitter
interface (P0c), and milestone coordination.
- backend-engineer: RETAINED — owns SSH-push transport (P01), runtime
abstraction (P07a/b/c), transaction bundle (P10 design), step-ca
integration, secrets crypto. Frameworks updated: golang.org/x/crypto/ssh
(existing), golang.org/x/crypto/ssh/knownhosts (existing); pending
deps: bytecodealliance/wasmtime-go (C-01 gate), smallstep/cli (I-B-004).
- data-engineer: RETAINED — owns per-namespace DB schema split (P0a1,
REQ-071), CLI cache DB (R-008), namespace inheritance resolver state
(P0a2). Frameworks: modernc/sqlite.
- security-engineer: REACTIVATED — owns step-ca provisioning (REQ-076),
master.key + AES-256-GCM crypto (REQ-080, C-19 threat model), SPIFFE
SVID minting (C-08 spike), SSH-push blast-radius review, Traefik edge,
.env.secrets threat model. The re-architecture reverses AD-010
(step-ca rejection) and the SPIFFE rejection at PROJECT.md:94; both
reversals are justified in the Supersession Table.
- network-engineer: REACTIVATED — owns socket-based service exposure
(R-007, P08), Syncthing P2P ports (P09), Traefik routing + dynamic
config atomicity (P02, C-10). The transport layer moves from mTLS
HTTP daemon-to-daemon to SSH CLI-to-server; network-engineer reviews
the new trust surface.
- devops-engineer: REACTIVATED — owns bash scripts (scripts/orca-*.sh,
C-15..C-18: bats/shellcheck/shfmt gate, render-format contract,
slog-syslog), systemd timers (orca-pull/drift/aggregate, C-09 failure
contract, C-11 watchdog), hermetic test infra (P00 bootstrap, P08
expand, REQ-087).
- cli-engineer: remains DEACTIVATED — CLI surface growth is owned by
lead-developer (cobra subcommands) + backend-engineer (transport);
reactivation optional if CLI subcommand surface exceeds lead-developer
bandwidth.
- frontend-engineer: remains DEACTIVATED — no web UI (unchanged from
v0.1 onward; R-014 makes Markdown canonical, not a web UI).
---
# Personas: Orca
## v0.8 persona assessment
## v0.9 persona assessment (supersedes v0.8)
The v0.9 re-architecture introduces 5 new external apt dependencies (step-ca,
Traefik, Syncthing, wasmtime, podman), 8 net-new subsystems, and deprecates
~10k lines of shipped daemon/transport/CA/HCL code. The active roster grows
from 3 to 6 to cover the new attack surfaces and deployment model. Territory
enforcement remains in `warn` mode per config.json.
### lead-developer
- **Domain**: coordination
+92
View File
@@ -0,0 +1,92 @@
# Orca — Comprehensive Product Requirements Document (v0.9/v1.0)
**Audience:** Operators, AI agents, downstream tooling authors
> This PRD SUPERSEDES the shipped v0.1v0.8 architecture. The v0.9 and v1.0
> milestones implement a re-architecture whose load-bearing rules (R-001…R-016)
> and decisions (D-068…D-206) replace or demote several earlier documented
> decisions. See §22 decision-trace and the Supersession Table in
> `ARCHITECTURE.md` for the recorded reversals and their evidence basis.
## Status
| Item | Status |
|---|---|
| Spec lock-in | ✅ R-001…R-016 + D-001…D-206 settled |
| v0.1v0.8 implementation | ✅ shipped (REQ-001..060, D-001..D-047) |
| v0.9 implementation | ⬜ Phase 0 pre-execution (this file is the spec input) |
| v1.0 implementation | ⬜ planning (post-PRD) |
| v1.x multi-host state | ⬜ parked (post-v1.0) |
| v2.x full Nomad-HCL | ⬜ parked (post-v1.x) |
## Override justification (recorded for the grill supersession)
The v0.9/v1.0 re-architecture is justified on six independent grounds rather
than preference. Each reverses a prior documented decision; the new evidence
basis is recorded with the reversal in the Supersession Table:
1. **The v0.8 daemon model is operationally failing** in the target environment
— R-001 ("no orca binary on any server") is a response to measured pain, not
preference.
2. **step-ca is externally mandated** (D-101) — the operator environment requires
an external CA; AD-010's "too heavyweight" rationale is no longer operative.
3. **Multi-tenancy is a hard product requirement** (R-002) — real multi-tenant
use cases cannot be served by the single-namespace layout; the
"no multi-tenancy" anti-pattern is obsolete.
4. **WASM is a hard workload requirement** (D-088) — workloads are WASM, not
processes; `os/exec` is insufficient; the "no container runtime" anti-pattern
is reversed.
5. **SSH-push is the only viable deployment target** for the operator's
bare-Linux/Proxmox environment — installing/maintaining an orca daemon on
every peer is operationally infeasible.
6. **Simplicity/vision correction** — the v0.1-v0.8 daemon model was a wrong
turn against the original CLI-first vision; the re-architecture corrects the
vision.
## Canonical references
The full PRD text was provided by the operator and adopted wholesale. The
load-bearing rules (R-001…R-016), the concept model (§4), the architecture
(§5), the milestone plan (§23), and the decision trace (§22) are reproduced
in the operator's original document. This file is the auditable pointer to
that source; the substantive planning artifacts live in:
- `IDEATION_v0.9.md` — 30 ideas (REQ-061..REQ-090), three tiers
- `GRILL_v0.9.md` — 9-axis adversarial review, 19 binding conditions, 10 phase challenges
- `REQUIREMENTS.md` — REQ-061..REQ-090 appended
- `ROADMAP.md` — v0.9 (13 phases) + v1.0 (19 phases) appended
- `PERSONAS.md` — security/network/devops reactivated
- `ARCHITECTURE.md` — v0.9 banners + Supersession Table
## The 16 load-bearing rules (invariants)
| ID | Rule |
|---|---|
| R-001 | No Orca Go binary runs on any server. The `orca` CLI on the operator's host is the only Orca software. Servers run Linux + systemd + apt-managed packages + config files written by the CLI. |
| R-002 | Filesystem paths are namespaces. `ORCA_HOME` hosts many namespaces; each is a dir with `db/`, `.env`, `.env.secrets`, `jobs/`, `alloc/`, `ns.md`. `_defaults/` always exists. No `namespace` column in SQLite. |
| R-003 | Cluster lead is always bare Linux; Proxmox can never be lead. |
| R-004 | Workload migration Linux↔Proxmox supported; runtime can change at migration; SPIFFE identity preserved. |
| R-005 | Storage replication enables migration; a Service's `count` replicas share one `runtime {}` block. |
| R-006 | mTLS on by default; cluster CA = step-ca; Traefik + `LoadCredential=` are load-bearing. |
| R-007 | Sockets by default (`/run/orca/alloc-<id>/port-<name>.sock`); `127.0.0.1` opt-in. |
| R-008 | CLI results cached locally with per-class TTLs (`orca_cache` SQLite). |
| R-009 | CLI host SPOF mitigated by external shared state in v1.x; v1.0 ships the abstractions + cache layer. |
| R-010 | Control plane updates are transactional (ArgoCD-style desired-state/lead-applier). |
| R-011 | Each namespace has `.env` (plaintext) and `.env.secrets` (AES-256-GCM, per-line nonce); master key per `ORCA_HOME` at `cluster/master.key`. |
| R-012 | Workload kinds are `Job`, `Service`, `DaemonSet`; schema-separated by `kind:` in frontmatter. |
| R-013 | Jobspec format is Markdown with YAML frontmatter (`.md` preferred); `.yaml` and `.hcl` accepted by parser dispatcher. |
| R-014 | All user-facing config is Markdown with YAML frontmatter; body preserved verbatim. |
| R-015 | Body of every `.md` config file is preserved verbatim and surfaced in `inspect`, `history`, diffs. |
| R-016 | `.env` and `.env.secrets` are exempt from R-014 — standard dotenv format retained. |
## Milestone summary (§23, reordered per grill PC-01..PC-10)
### v0.9 — Workloads + Re-architecture Foundation (13 phases)
P00 (deprecation sweep + migration-ordering + txn-design spike + test-infra bootstrap + persona reactivation + doc banners), P0a1 (path resolver + config demotion), P0a2 (namespace CRUD + inheritance), P0b (Markdown jobspec parser + fuzz), P0c (schemas + emitter interface), P01 (SSH-push transport + host-path volumes), P02 (service + Traefik emitter), P03 (update stanza), P04 (lifecycle hooks), P05 (constraints + CLI-side scheduler), P06 (task groups), P07a/P07b/P07c (process+podman / wasmtime [C-01 gated] / pve-vm+ct runtimes), P08 (sockets), P09 (Syncthing [C-02 gated]), P10 (lead rules + migration), P0X (ship + audit).
### v1.0 — Production Hardening (19 phases)
P00 (CLI cache), P01 (metrics), P01.5 (SPIFFE spike [C-08 gated]), P02 (ACL), P03 (secrets), P04 (backup/restore), P05 (drain + daemon drain-and-stop), P06 (alloc history), P07 (recovery), P08 (integration tests), P09 (collector+aggregator), P10 (transactional plane [C-09 gated]), P11 (job lint), P12 (job verify), P13 (ns subcommands), P14a/P14b/P14c (data / daemon cutover / mixed-version tolerance), P15 (README), P15.5 (threat model [C-19 gated]), P16 (final review + ship — v1.0.0 release).
See `ROADMAP.md` for the full reordered plan and `GRILL_v0.9.md` for the 19
binding conditions (C-01..C-19) and 10 phase challenges (PC-01..PC-10) that
gate specific phases.
+67
View File
@@ -377,3 +377,70 @@ within the `clarify_budget` (10):
| D-045 | `--host-key-fingerprint` format — raw hex, `sha256:`-prefixed, or OpenSSH `SHA256:base64`? | **OpenSSH `SHA256:base64` (the format `ssh-keyscan -E sha256 -D -` emits and operators expect)** | Matches the fingerprint format operators already see from `ssh-keyscan` and `orca node join`'s own `Result.HostKeyFingerprint` output. Accept only `SHA256:`-prefixed base64; reject raw hex with a clear error. Internally decode base64 → compare against `ssh.PublicKey` Marshal + sha256. | 0.88 |
| D-046 | Does `orca node key-reset <node>` also revoke the orca pubkey on the remote host, or only clear the local `known_hosts` entry? | **Local `known_hosts` entry only** | Revoking the remote authorized_keys entry would orphan a working node (next dispatch would fail auth). `key-reset` is the local "forget this host's key" operation (mirrors `ssh-keygen -R host`); re-establishing trust is a separate `orca node join` re-run. Audit-log the reset with `actor`, `node`, `event=node.key_reset`. | 0.90 |
| D-047 | Coverage target for P01 — 70% floor or higher? | **70% floor for the 6 under-50% packages; 50% floor for the 3 zero-test packages (`internal/audit`, `internal/certpaths`, `cmd/orca`) as a first-toe-hold** | 70% across the board for the already-tested packages matches D-042's "70% target for new packages" and is achievable without heroic mock effort. For the zero-test packages, going 0→50% is the realistic single-phase step (0→70% risks a coverage rathole on `cmd/orca` which is glue code); a future milestone can lift them to 70%. | 0.82 |
---
# v0.9/v1.0 — Re-architecture Scope Summary (Supersedes v0.1v0.8 architecture)
v0.9 is the first DIRECTION-CHANGE milestone in the project's history.
It supersedes the shipped v0.1v0.8 architecture per the adopted PRD
(`.ciagent/PRD_v0.9.md`). The re-architecture deprecates the daemon/
transport/internal-CA/HCL/single-namespace stack and builds a CLI-only/
SSH-push/step-ca/Markdown-frontmatter/multi-namespace stack plus 8
net-new subsystems.
## Override Justification (Re-architecture Justification axis)
The ci-griller returned REPLAN (0.70) on the Re-architecture Justification
axis, noting the PRD reverses 6 documented decisions without new evidence
and that the incremental-additive path was not evaluated. The user reviewed
the fork and overrode the *direction* with a six-part evidence basis. The
override is recorded verbatim below; each part addresses a reversal that
the grill flagged as unjustified.
1. **The v0.8 daemon model is operationally failing** in the target
environment — R-001 ("no orca binary on any server") is a response to
measured pain, not preference.
2. **step-ca is externally mandated** (D-101) — the operator environment
requires an external CA; AD-010's "too heavyweight" rationale is no
longer operative.
3. **Multi-tenancy is a hard product requirement** (R-002) — real
multi-tenant use cases cannot be served by the single-namespace layout;
the "no multi-tenancy" anti-pattern is obsolete.
4. **WASM is a hard workload requirement** (D-088) — workloads are WASM, not
processes; `os/exec` is insufficient; the "no container runtime"
anti-pattern is reversed.
5. **SSH-push is the only viable deployment target** for the operator's
bare-Linux/Proxmox environment — installing/maintaining an orca daemon
on every peer is operationally infeasible.
6. **Simplicity/vision correction** — the v0.1-v0.8 daemon model was a
wrong turn against the original CLI-first vision; the re-architecture
corrects the vision.
## Supersession Table (AD-series reversals, recorded per grill PC-09)
| Old decision | Was | Superseded by | Evidence basis |
|---|---|---|---|
| AD-010 (ARCHITECTURE.md:463) | step-ca/cfssl/vault-pki "too heavyweight" | **D-101** (step-ca) | Override ground 2 (external mandate) |
| SPIFFE rejection (PROJECT.md:94) | internal CA chosen over SPIFFE | **D-068** (SPIFFE SVIDs) | Override ground 3 (multi-tenancy requires per-workload identity) |
| No-container-runtime (ARCHITECTURE.md:477) | explicit anti-pattern | **D-088** (5 runtimes; wasmtime primary) | Override ground 4 (WASM is the workload profile) |
| No-multi-tenancy (ARCHITECTURE.md:478) | explicit anti-pattern | **D-158 / R-002** (many namespaces under ORCA_HOME) | Override ground 3 (hard multi-tenant product req) |
| AD-007 (HCL canonical) | HCL for jobspec | **R-013 / R-014** (Markdown canonical; HCL legacy) | PRD §8 (Markdown + body preservation is the operator-facing format) |
| Daemon-on-every-node | `orca daemon` on all peers | **R-001** (no orca binary on any server) | Override grounds 1 + 5 (daemon failing; SSH-push only viable target) |
The 19 binding conditions (C-01..C-19) and 10 phase challenges
(PC-01..PC-10) from `GRILL_v0.9.md` are adopted as execution gates.
The 30 net-new requirements (REQ-061..REQ-090) from `IDEATION_v0.9.md`
are recorded in `REQUIREMENTS.md`. The reordered phase plan is in
`ROADMAP.md`.
## v0.9 Clarified Decisions (D-series, full autonomy — Phase 0 pre-execution)
| ID | Question | Decision | Rationale | Confidence |
|----|----------|----------|-----------|------------|
| D-101 | Cluster CA: internal Go CA (AD-010) or step-ca (external)? | **step-ca (apt-installed)** | Externally mandated per override ground 2; AD-010's "too heavyweight" rationale reversed. CLI wraps `step` CLI via SSH (no Go step-ca client library — keep zero-new-dep posture if possible, or add `github.com/smallstep/cli` as a dep). **Gated by C-07** (CA migration spec). | 0.74 |
| D-068 | Workload identity: internal X.509 CA or SPIFFE SVIDs? | **SPIFFE SVIDs minted at submit time via step-ca** | Multi-tenancy (override ground 3) requires per-workload identity model; SPIFFE is the standard. SPIFFE ID `spiffe://orca/ns/<ns>/job/<name>/alloc/<id>` as SAN. **Gated by C-08** (mint spike in v1.0-P01.5; fallback to mTLS identity if spike fails). | 0.72 |
| D-088 | Runtime: direct os/exec only (D-008) or multi-runtime? | **5 runtimes: wasm (wasmtime primary), podman, process, pve-vm, pve-ct** | WASM is the primary workload (override ground 4). `processRuntime` wraps existing `executor.go`; others are net-new. Split P07a/b/c per grill PC-10. **P07b gated by C-01** (wasmtime/CGO eval). | 0.82 |
| D-158 | Namespace model: single flat root or multi-namespace? | **Multi-namespace under ORCA_HOME (R-002)** | Hard multi-tenant product requirement (override ground 3). `_defaults/` implicit root; `cluster/` for cluster-wide; per-namespace `db/`, `.env`, `.env.secrets`, `jobs/`, `alloc/`, `ns.md`. No namespace column in SQLite. | 0.84 |
| D-179 | Jobspec format: HCL canonical (AD-007) or Markdown? | **Markdown with YAML frontmatter canonical (R-013); HCL legacy** | PRD §8 — Markdown + body preservation is the operator-facing format. HCL adapter (REQ-064) preserves `orca job run old-spec.hcl` during migration. | 0.85 |
| D-185 | Re-architecture justification: incremental additive or full re-architecture? | **Full re-architecture (overridden by user)** | Six-part evidence basis above; the grill's REPLAN mechanics (PC-01..PC-10, C-01..C-19) adopted as gates. The incremental-additive path was evaluated and rejected on grounds 1 + 5 (daemon failing; SSH-push only viable). | 0.88 |
+42
View File
@@ -140,3 +140,45 @@ REQ-047..052 all complete.
| REQ-058 | `--host-key-fingerprint <SHA256:base64>` pre-pin flag on `orca node join` (validated when `--type proxmox`): when supplied, join fails fast if the SSH host key's OpenSSH SHA-256 fingerprint does not match; supersedes TOFU (D-035) for pre-pinned deployments (D-044, D-045) | Medium | **v0.8 P2** | **Complete** (P2 shipped v0.7.2) |
| REQ-059 | `orca node key-reset <node>` command: clears the persisted SSH host key entry for the node from `~/.orca/known_hosts` only (local, not remote authorized_keys — D-046); audit-logs `event=node.key_reset`; next `doctor proxmox`/dispatch re-pins via TOFU or `--host-key-fingerprint` | Low | **v0.8 P2** | **Complete** (P2 shipped v0.7.2) |
| REQ-060 | Requirement-status hygiene sweep: REQUIREMENTS.md v0.7 rows were stale ("Pending" after ship); add a verify-stage assertion that every REQ listed as `Complete` in ROADMAP.md has a matching `Complete` row in REQUIREMENTS.md, enforced by `make verify-reqs` | Medium | **v0.8 P3** | **Complete** (P3 shipped v0.7.3) |
## v0.9/v1.0 Requirements — Re-architecture Foundation & Production Hardening
The v0.9/v1.0 milestones supersede the shipped v0.1v0.8 architecture per the
adopted PRD (`.ciagent/PRD_v0.9.md`). The re-architecture is justified on six
grounds recorded in the PROJECT.md Supersession Table. 30 net-new requirements
(REQ-061..REQ-090) derive from the v0.9 IDEATION; their phase placement and
binding grill conditions (C-01..C-19) are documented in `IDEATION_v0.9.md`
and `GRILL_v0.9.md`.
| ID | Requirement | Priority | Phase | Status |
|----|-------------|----------|-------|--------|
| REQ-061 | `orca daemon` deprecation command and build-tag removal path: v0.9 emits deprecation warning + still runs (dual-write window); v1.0 repurposes to `orca daemon drain-and-stop` (stops v0.8 daemons on peers via SSH, confirms workloads survive via systemd); post-v1.0 the command and `internal/daemon/` are deleted. `// Deprecated` Go doc comments + `slog.Warn` on every run (I-M-001) | High | **v1.0 P14** (warn v0.9 P0X) | Pending |
| REQ-062 | Coverage follow-ups: 3 zero-test packages (`internal/audit`, `internal/certpaths`, `cmd/orca`) + `internal/cli` to 70% floor; once `daemon.go` is deprecated/removed the exclusion reason disappears and the floor applies to the whole package; all net-new subsystems carry a 70% floor from their first phase (I-M-002) | Medium | **v0.9 P0X** + each new pkg | Pending |
| REQ-063 | `known_hosts` flock concurrency gap (deferred P1 from REVIEW_v0.8 A2): add `flock`-style advisory lock (stdlib `syscall.Flock` wrapper) around the read-modify-write in `TOFUHostKeyCallback` capture path (`bootstrap.go:290-302`) and `ResetHostKey` (`bootstrap.go:479-523`); lock file at `cluster/known_hosts.lock` (R-002) (I-M-003) | Medium | **v0.9 P0a1** | Pending |
| REQ-064 | HCL→Markdown jobspec adapter/bridge layer: keep `internal/jobspec/spec.go` as legacy HCL path behind `// Deprecated`; add `internal/jobspec/markdown.go` (canonical) + `internal/jobspec/dispatch.go` (extension-based dispatcher: `.md`→Markdown, `.hcl`→legacy, `.yaml`→Markdown-with-empty-body); unified `*WorkloadSpec` populated via adapter; preserves `orca job run old-spec.hcl` during migration window (I-M-004) | High | **v0.9 P0b** | Pending |
| REQ-065 | `orca doctor --legacy-paths` detection: detects v0.8 residue (orca.db at ORCA_HOME root, ca.crt/ca.key, config.hcl, flat server.crt, namespace column in any *.db); outputs list of legacy artifacts with migration recommendations; the detection half of v1.0-P14 (I-M-005) | Medium | **v1.0 P14c** | Pending |
| REQ-066 | Legacy CA state migration to step-ca: `orca upgrade --to-v1.0 --import-ca` reads `~/.orca/ca.key`, initializes step-ca with it, re-issues workload SVIDs; preserves audit history even if live trust root changes (I-M-006). **Gated by C-07** | High | **v1.0 P14a** | Pending |
| REQ-067 | Fuzz test harness for Markdown frontmatter parser: `testing.F` fuzz target in `internal/jobspec/markdown_test.go` round-trips random frontmatter+body through `ParseMarkdown` asserting byte-exact body preservation; corpus of adversarial fixtures (CRLF, BOM, no-frontmatter, empty-frontmatter, frontmatter-with-only-separator) (I-M-007) | Medium | **v0.9 P0b** | Pending |
| REQ-068 | Deprecation warnings on removed/repurposed CLI subcommands: each removed/changed command (`orca cert`, `orca node join` mTLS semantics, `orca job run <spec.hcl>`) emits `slog.Warn` deprecation banner with v1.0 replacement except under `orca upgrade`; `--no-deprecation-warnings` global flag via `root.go` `PersistentPreRunE` (I-M-008) | Low | **v0.9 P0X** + v1.0 P13 | Pending |
| REQ-069 | `internal/config/config.go` HCL config demotion via adapter: keep `internal/config/` as `legacy_config.go` with `// Deprecated`; add `internal/config/markdown.go` for new Markdown-frontmatter loader (R-014); `root.go` dispatches on file extension (`.hcl`→legacy, `.md`→new); `--config` semantics: `.hcl` read-only legacy, `.md` canonical (I-M-009) | High | **v0.9 P0a1** | Pending |
| REQ-070 | `internal/certpaths/` replacement with multi-namespace path resolver: new `internal/paths` package with `paths.NamespaceDir(ns)`, `paths.ClusterDir()`, `paths.CacheDB()`, `paths.MasterKey()`, `paths.NSDb(ns)`, `paths.NSEnv(ns)`, `paths.NSSecrets(ns)`; keep `certpaths` as thin shim for v0.8 compat then remove post-v1.0 (R-002) (I-M-010) — highest blast radius | High | **v0.9 P0a1** | Pending |
| REQ-071 | `internal/store/` schema: per-namespace DBs, drop namespace column: `store.Open` gains namespace parameter (or caller passes `paths.NSDb(ns)`); `migrate.go` runs migrations per namespace DB; `cert_repo` (0004) removed (step-ca handles certs); audit_log moves to CLI-side cache DB (R-008) (I-M-011) | High | **v0.9 P0a1** + v1.0 P06 | Pending |
| REQ-072 | `internal/transport/` deletion + SSH-push package: delete `mtls.go`, `dispatch.go`, `handshake_log.go`; extract retry/idempotency patterns into `internal/sshpush/`; existing `transport.IdempotencyStore` directly reusable (I-M-012). Deletion deferred to v1.0-P14 to keep dual-write window open | High | **v0.9 P00** (delete v1.0 P14) | Pending |
| REQ-073 | SSH-push transport layer design: connection pooling (reuse `*ssh.Client` per peer), idempotency (content-addressed filenames), retry (exponential backoff 100ms×2 cap 5s max 5), timeout (30s SCP, 10s exec), fan-out (errgroup bounded concurrency default 8), known_hosts reuse `proxmox.TOFUHostKeyCallback` (I-B-001) | High | **v0.9 P01** (design P0a1) | Pending |
| REQ-074 | Emitter template system (Layer 4): `internal/emitter/` package with `Emitter` interface `Render(spec *WorkloadSpec, node *Node) ([]File, error)`; implementations systemdEmitter/traefikEmitter/syncthingEmitter/socketEmitter; SSH-push SCPs `[]File` atomically (write-to-tmp + rename); emitters registered per kind + runtime (I-B-002) | High | **v0.9 P0c** | Pending |
| REQ-075 | Lead applier execution model: CLI renders transaction bundle (tarball + apply.sh + verify.sh) on operator host, SCPs to lead's `/run/orca/txns/<txn-id>/`, lead's systemd timer runs `apply.sh` idempotently, CLI polls txn status via SSH; bash scripts generated by emitter not hand-written (I-B-003). **Gated by C-09** | High | **v1.0 P10** (design v0.9 P00) | Pending |
| REQ-076 | step-ca integration: `orca init` runs `step ca init` on lead; CLI SSHs to lead, installs step-ca via apt, stores step-ca.json; workload SVIDs via `step ca token` (JWE minted by CLI) → `step ca certificate`; SPIFFE ID as SAN; new `internal/stepca/` package wraps `step` CLI via SSH (I-B-004). Reverses AD-010 per override justification ground 2 | High | **v0.9 P07** + v1.0 P02 | Pending |
| REQ-077 | Traefik dynamic config generation + atomic reload: Traefik emitter renders `/etc/traefik/dynamic/orca-<ns>-<svc>.yaml` with backends (socket paths R-007), health checks, mTLS config pointing at step-ca root; atomic reload via tmpfile+fsync+rename triggering fsnotify; drain writes `weight=0` or removes backend (I-B-005). **Gated by C-10** | High | **v0.9 P02** | Pending |
| REQ-078 | Runtime abstraction interface (5 backends): `Runtime` interface in `internal/runtime/` with Prepare/Start/Stop/Status; processRuntime (wraps existing executor.go), wasmRuntime (wasmtime via SSH), podmanRuntime, pveVMRuntime (qm via proxmox SSH), pveCTRuntime (pct); runtimeRegistry keyed by `runtime:` frontmatter value; Alloc carries runtime field changeable on migration (I-B-006). Split P07a/b/c per PC-10. **P07b gated by C-01** | High | **v0.9 P07a/b/c** | Pending |
| REQ-079 | Transaction bundle format + N-peer atomicity: bundle = tarball with desired-state.json + apply.sh + verify.sh + rollback.sh + manifest.sig (signed with master.key); content-addressed `<txn-id>=sha256(desired-state.json)` stored in `cluster/txns/<txn-id>/`; lead applies to self first then fans out; failure on any peer runs rollback.sh on applied peers (I-B-007). **Gated by C-09** | High | **v1.0 P10** (design v0.9 P00) | Pending |
| REQ-080 | Master key management + HKDF-SHA256 per-line .env.secrets encryption: `cluster/master.key` 32-byte random (generated at `orca init` using WriteAtomic pattern); each line `base64(nonce||ciphertext||tag)`, nonce=random(12 bytes), AES-256-GCM with AAD=line-number (prevents line-swap); HKDF-SHA256 derives per-namespace sub-keys; `orca secrets set/get`; v0.8 `internal/security/redact.go` reusable (I-B-008). **Gated by C-19** | High | **v1.0 P03** | Pending |
| REQ-081 | Syncthing config rendering + folder-ID content-addressing: per-namespace Syncthing folder `orca-<ns>` with content-addressed folder ID `sha256(ns + master-key-fingerprint)`; CLI renders config.xml per peer; Syncthing runs as systemd unit (emitted by systemd emitter); CLI discovers peers via `cluster/peers/`; migration works because new node joins folder and syncs before workload starts (I-B-009). **Gated by C-02 + C-14** | Medium | **v0.9 P09** (spike v0.9 P00) | Pending |
| REQ-082 | Namespace inheritance resolver algorithm: DFS parent walker with visited set for cycle detection; `_defaults/` implicit root (always exists, no parent); merge semantics: child overrides parent for scalars, arrays unioned (child adds to parent); pure function (no I/O) taking `map[nsName→*NSConfig]` returning `map[nsName→*ResolvedNS]` (I-B-010) | High | **v0.9 P0a2** | Pending |
| REQ-083 | CLI-side scheduler redesign: `Score(node, workload) (score int, fits bool)` where `fits` checks runtime compatibility + constraints, `score` is bin-packing (most free capacity = highest); Services pick `count` distinct nodes (anti-affinity default); DaemonSets pick all matching nodes; Job = one-shot; CLI-side not daemon-side (R-001) (I-B-011) | High | **v0.9 P05** (skeleton P0c) | Pending |
| REQ-084 | `orca job lint` category-driven lint engine: `Linter` runs `Rule` checks returning `Finding{Category, Severity, Message, Explanation}`; categories schema/runtime/security/migration/best-practice; `--explain` prints rationale; pure (no I/O) checks against static rules (I-B-012) | Medium | **v1.0 P11** | Pending |
| REQ-085 | v0.8→v1.0 migration ordering: v0.9 ships new parser + kinds + runtime + SSH-push alongside old daemon (dual-write window); `orca job run` dispatches on extension (`.md`→SSH-push, `.hcl`→old daemon); v1.0-P05 drains old daemons; v1.0-P14 converts remaining `.hcl` specs and removes daemon (I-C-001). **Most important cross-cutting idea** | High | **v0.9 P00** → v1.0 P14 | Pending |
| REQ-086 | "No orca on server" enforcement: `orca doctor no-orca-on-server` SSHs to each peer verifying no `orca` binary in PATH, no `orca` systemd service, no `orca` process, no `/etc/orca/` directory; runs after v1.0-P05 before v1.0-P16; reuses v0.8 `proxmox` SSH session infrastructure (I-C-002). Implements grill C-13 | High | **v1.0 P14c** | Pending |
| REQ-087 | Test infrastructure: hermetic 3-linux + 1-proxmox cluster pipeline: `test/integration/` with docker-compose/vagrant creating 4 containers/VMs; Go test harness SSHes to each, runs CLI, asserts end-to-end workflows (ns create → workload submit → migrate → drain); proxmox simulated via mock pct/qm; v0.8 e2e tests (bootstrapE2ESetup) are foundation (I-C-003) | Medium | **v1.0 P08** (bootstrap v0.9 P00) | Pending |
| REQ-088 | Security-engineer + network-engineer persona reactivation: reactivate security-engineer (step-ca provisioner model, SSH-push blast radius, Traefik edge, .env.secrets crypto) and network-engineer (socket exposure R-007, Syncthing P2P ports, Traefik routing); cross-cutting review not single phase (I-C-004). Implements grill C-05 | High | **v0.9 P00** → v1.0 P16 | Pending |
| REQ-089 | Documentation rewrite: ARCHITECTURE.md/PROJECT.md/README + AD-010 supersession: v0.9-P00 adds "v0.9 Architecture (Supersedes v0.8)" section + banners + Superseded Decisions table; v1.0-P15 rewrites README quickstart for new curl|sh + orca init + orca ns create flow (I-C-005) | Medium | **v0.9 P00** + v1.0 P15/P16 | Pending |
| REQ-090 | Dual-write window: v0.9 `orca job run` dispatches on extension (`.md`→SSH-push new path, `.hcl`→old daemon path) via parser dispatcher (REQ-064); daemon not removed until v1.0-P05; SSH-push path writes to separate systemd unit namespace (`orca-v1-<alloc>.service`) while daemon uses `orca-<job>.service` — no unit name overlap = no conflict (I-C-006) | High | **v0.9 P00** | Pending |
+150
View File
@@ -184,3 +184,153 @@ tag.
The vision ("minimalist, offline-first, CLI-first orchestration
engine") is unchanged. v0.8 closes the coverage debt left by v0.7's
50% floor and the trust-surface gaps explicitly deferred in v0.6.
## Milestone v0.9: Re-architecture Foundation & Workloads
**Scope**: This milestone SUPERSPEDES the shipped v0.1v0.8 architecture per
the adopted PRD (`.ciagent/PRD_v0.9.md`). The re-architecture is justified on
six grounds recorded in the PROJECT.md Supersession Table: (1) the v0.8 daemon
model is operationally failing, (2) step-ca is externally mandated, (3)
multi-tenancy is a hard product requirement, (4) WASM is a hard workload
requirement, (5) SSH-push is the only viable deployment target, (6) vision
correction. The 16 load-bearing rules (R-001…R-016) are invariants. The
ci-griller reviewed the re-architecture adversarially; the user overrode the
Re-architecture Justification REPLAN with the six-part evidence basis; the
19 binding conditions (C-01..C-19) and 10 phase challenges (PC-01..PC-10)
from `GRILL_v0.9.md` are adopted as execution gates. 30 net-new requirements
(REQ-061..REQ-090) derive from `IDEATION_v0.9.md`.
**Milestone type**: feature (P01..P10 ship `feat` phases; P00/P0X are
chore/docs).
- [ ] Phase 0: Pre-execution (specify → clarify → research → ideate → plan → grill) — tag `v0.8.0` (shipped; this is the phase you are reading)
- [ ] Phase P00: Deprecation sweep + migration-ordering decision + txn-design spike + hermetic test-infra bootstrap + persona reactivation + doc banners (REQ-072, REQ-085, REQ-088, REQ-089, REQ-090; gates C-03 ✅, C-05, C-06, C-15..C-18) — tag `v0.8.1`
- [ ] Phase P0a1: Multi-namespace path resolver + config HCL demotion + known_hosts flock (REQ-063, REQ-069, REQ-070, REQ-071; gate C-07) — tag `v0.8.2`
- [ ] Phase P0a2: Namespace CRUD + inheritance engine (REQ-082) — tag `v0.8.3`
- [ ] Phase P0b: Markdown jobspec parser + dispatcher + fuzz (REQ-064, REQ-067) — tag `v0.8.4`
- [ ] Phase P0c: Job/Service/DaemonSet schemas + emitter interface (REQ-074) — tag `v0.8.5`
- [ ] Phase P01: SSH-push transport + host-path volumes (REQ-073) — tag `v0.8.6`
- [ ] Phase P02: Service block + checks + restart + Traefik emitter (REQ-077; gate C-10) — tag `v0.8.7`
- [ ] Phase P03: Update stanza (rolling/canary) — tag `v0.8.8`
- [ ] Phase P04: Lifecycle hooks (systemd ExecStop) — tag `v0.8.9`
- [ ] Phase P05: Constraints & affinity (CEL) + CLI-side scheduler (REQ-083) — tag `v0.8.10`
- [ ] Phase P06: Task groups (multi-process services) — tag `v0.8.11`
- [ ] Phase P07a: Process + podman runtimes (REQ-078) — tag `v0.8.12`
- [ ] Phase P07b: wasmtime runtime (REQ-078; **gate C-01** — CGO eval) — tag `v0.8.13`
- [ ] Phase P07c: pve-vm + pve-ct runtimes (REQ-078; extends REQ-076) — tag `v0.8.14`
- [ ] Phase P08: Socket plumbing (R-007) — tag `v0.8.15`
- [ ] Phase P09: Storage replication via Syncthing (REQ-081; **gates C-02, C-14**) — tag `v0.8.16`
- [ ] Phase P10: Lead rules + migration (REQ-076 step-ca integration) — tag `v0.8.17`
- [ ] Phase P0X: Ship + audit (REQ-062 coverage gate; REQ-068 deprecation warnings) — tag `v0.8.18`
**Milestone tag**: `v0.8.18` (final phase patch = milestone release per
feature-milestone progressive-patch rule). Per-phase tags: `v0.8.1``v0.8.18`.
Tags run on the previous minor's patch line (v0.8.x) per branch-strategy.md.
The milestone branch label uses the milestone number
(`milestone/v0.9-rearchitecture`); no separate minor tag.
### Per-phase REQ coverage (v0.9)
- **P00** — Deprecation/migration/test-infra/persona/docs foundation (REQ-072, REQ-085, REQ-088, REQ-089, REQ-090)
- **P0a1** — Path resolver + config demotion + known_hosts flock (REQ-063, REQ-069, REQ-070, REQ-071)
- **P0a2** — Namespace inheritance resolver (REQ-082)
- **P0b** — Markdown parser + adapter + fuzz (REQ-064, REQ-067)
- **P0c** — Schemas + emitter interface (REQ-074)
- **P01** — SSH-push transport (REQ-073)
- **P02** — Service + Traefik emitter (REQ-077)
- **P05** — CLI-side scheduler (REQ-083)
- **P07a/b/c** — Runtime abstraction (REQ-078) + step-ca integration (REQ-076)
- **P09** — Syncthing replication (REQ-081)
- **P0X** — Coverage gate (REQ-062) + deprecation warnings (REQ-068)
### v0.9 is a DIRECTION CHANGE — first in the project's history
Every prior milestone (v0.1v0.8) explicitly said "the vision is unchanged;
this milestone is not a direction change." v0.9 is the first milestone that
reverses the vision's anti-patterns (daemon-on-every-node, internal CA,
HCL-canonical, single-namespace, no-container-runtime, no-SPIFFE). The
reversals are justified by the six-part evidence basis recorded in the
PROJECT.md Supersession Table.
## Milestone v1.0: Production Hardening
**Scope**: ship a cluster that operators can run. Builds on the v0.9
re-architecture foundation with the production-grade subsystems:
secrets, transactions, ACL/SPIFFE, backup/restore, drain, recovery, and
the v0.8→v1.0 migration.
**Milestone type**: feature (multiple `feat` phases).
- [ ] Phase 0: Pre-execution (specify → clarify → research → plan → grill) — tag `v0.9.0`
- [ ] Phase P00: CLI cache layer (REQ-062 cache floor; R-008) — tag `v0.9.1`
- [ ] Phase P01: Metrics endpoint (hand-rolled text exposition) — tag `v0.9.2`
- [ ] Phase P01.5: SPIFFE SVID minting spike (REQ-076; **gate C-08** — if spike fails, fall back to mTLS identity) — tag `v0.9.3`
- [ ] Phase P02: ACL (SPIFFE + token identities) — tag `v0.9.4`
- [ ] Phase P03: Secrets subsystem (REQ-080; **gate C-19** threat model) — tag `v0.9.5`
- [ ] Phase P04: Backup/restore (tar + signed) — tag `v0.9.6`
- [ ] Phase P05: Drain + daemon drain-and-stop (REQ-061) — tag `v0.9.7`
- [ ] Phase P06: Alloc history (CLI-side SQLite retention; REQ-071 cache DB) — tag `v0.9.8`
- [ ] Phase P07: Recovery (`orca restore`) — tag `v0.9.9`
- [ ] Phase P08: Integration tests — expand hermetic harness (REQ-087) — tag `v0.9.10`
- [ ] Phase P09: Collector + aggregator (opt-in; **gates C-11, C-12, C-14**) — tag `v0.9.11`
- [ ] Phase P10: Transactional plane (REQ-075, REQ-079; **gate C-09** orca-pull.sh failure contract) — tag `v0.9.12`
- [ ] Phase P11: `orca job lint` (REQ-084) — tag `v0.9.13`
- [ ] Phase P12: `orca job verify` (dry-run txn through lead) — tag `v0.9.14`
- [ ] Phase P13: `orca ns` subcommands (full surface) + deprecation warnings (REQ-068) — tag `v0.9.15`
- [ ] Phase P14a: v0.8→v1.0 data migration (REQ-066; **gate C-07** CA migration spec) — tag `v0.9.16`
- [ ] Phase P14b: Daemon cutover + running-allocation adoption — tag `v0.9.17`
- [ ] Phase P14c: Mixed-version tolerance + no-orca-on-server enforcement (REQ-065, REQ-086; implements C-13) — tag `v0.9.18`
- [ ] Phase P15: README quickstart (REQ-089) — tag `v0.9.19`
- [ ] Phase P15.5: Threat model + security review (**gate C-19**) — tag `v0.9.20`
- [ ] Phase P16: Final review + ship + audit — **v1.0.0 release** — tag `v0.9.21`
**Milestone tag**: `v1.0.0` (the v1.0.0 release tag is the production-ready cut;
per-phase patches run on the v0.9.x line per branch-strategy.md). Per-phase
tags: `v0.9.0``v0.9.21`.
### Per-phase REQ coverage (v1.0)
- **P00** — CLI cache (R-008)
- **P01.5** — SPIFFE spike (REQ-076; C-08)
- **P03** — Secrets (REQ-080; C-19)
- **P05** — Drain + daemon stop (REQ-061)
- **P06** — Alloc history (REQ-071 cache DB)
- **P08** — Integration tests (REQ-087)
- **P10** — Transactional plane (REQ-075, REQ-079; C-09)
- **P11** — Job lint (REQ-084)
- **P13** — ns subcommands + deprecation warnings (REQ-068)
- **P14a/b/c** — Migration (REQ-066, REQ-065, REQ-086; C-07, C-13)
- **P15** — README (REQ-089)
- **P15.5** — Threat model (C-19)
### Risk register (from grill, for ongoing monitoring)
- **step-ca single-instance SPOF** (mitigation: C-12 doc; v1.x HA via systemd failover)
- **master.key passphrase-less 0600** (mitigation: C-19 threat model; consider OS keyring in v1.x)
- **wasmtime CGO breaks cross-compile** (mitigation: C-01 spike; fallback to podman/process primary)
- **bash control plane drift** (mitigation: C-15..C-18 render-format contract + bats gate)
- **daemon cutover orphans running allocs** (mitigation: P14b split; test adoption)
- **27→35+ phase scope** (mitigation: C-04 sizing; three-milestone split if exceeded — current count v0.9=18 + v1.0=22 = 40 phases; **C-04 sizing must run before v0.9 P00 execution to determine whether to split into v0.9+v0.10+v1.0**)
## Deferred to v1.x (out of scope for v1.0)
- `sqlite-wal-shared` state backend (R-009 abstractions ship in v1.0; backend in v1.x)
- `git` state backend
- `file+flock` state backend
- `orca cluster setup-shared` UX
- HA `step-ca` (active/passive via systemd)
- Journald log shipping (optional centralized audit)
- Network policy (`nftables` snippets)
- GPU / TPU constraints
## Deferred to v2.x (out of scope for v1.x)
- Full Nomad-HCL parser with no conversion round-trip
- Nomad-API subset for migrating existing Nomad fleets
- Nomad driver bridge
- Helm-equivalent templating (probably never)
- Service mesh beyond Traefik
- CRDs / Operators / Plugin model
- Leader-elected Raft coordinator
- External CA / Let's Encrypt / cert transparency
- Online-only features (HSTS, OCSP stapling, telemetry)
+3 -3
View File
@@ -5,9 +5,9 @@
"slug": "orca",
"name": "Orca",
"description": "Offline/CLI-first orchestration engine (Orca) — Nomad-inspired, far simpler than Kubernetes",
"milestone": "v0.8",
"phase": 4,
"milestone_type": "nfr",
"milestone": "v0.9",
"phase": 0,
"milestone_type": "feature",
"default_branch": "main",
"tech_stack": {
"language": "go",