CLARIFY stage: 5 clarifications resolved per locked decisions + synthesis. C1: hybrid default for fresh init, migrate existing via orca upgrade C2: thin wrapper upgrade (v0.11); full rolling upgrade defers to v1.x C3: drain+reschedule migrate (v0.11); live-migrate defers to v1.x C4: remediation cooldown on success only; transient failures retry next tick C5: doctor mTLS = chain verification + live handshake probe ---ci--- project: orca phase: 0 milestone: v0.11 status: clarify ---/ci---
3.5 KiB
CLARIFY v0.11: Production Hardening
Status: resolved (full autonomy, 2026-08-07). All 5 clarifications resolved with the operator's locked decisions (Q1=A, Q2=C, Q3=A, Q4=A, Q5=A) and the research-ingestion synthesis. No open questions remain for Phase 0.
Resolved clarifications
C1 — Ingress default binding (resolved)
Question: Is 127.0.0.1:8443 + nft the shipped default, or is the
v0.8 behavior (:443 on Traefik) still the default and hybrid is opt-in?
Decision: R-017 makes the hybrid the default for fresh orca init
(new clusters). Existing v0.9/v0.10 clusters get an opt-in migration path
via orca upgrade (REQ-115), which handles the Traefik binding cutover
from :443 to 127.0.0.1:8443. This is a behavioral change for existing
operators but it ships in a controlled migration phase (P14a), not as a
surprise default flip.
Affected REQs: REQ-100 (Traefik binding), REQ-115 (orca upgrade).
Affected phase: P14a (migration), P15.5 (new default).
C2 — orca upgrade scope (resolved)
Question: Is orca upgrade --to-vX (a) a thin wrapper around
install.sh + orca restore (binary upgrade only), or (b) a full
cluster-rolling-upgrade orchestrator (drain → upgrade binary → restart →
next node)?
Decision: (a) thin wrapper for v0.11. Full cluster-rolling-upgrade (b) defers to v1.x. The thin wrapper handles the R-017 binding cutover (REQ-115) for existing clusters. A full rolling-upgrade orchestrator is a v1.x concern (it requires P05 drain + P09 syncthing + P14c mixed-version tolerance to be production-tested first).
Affected REQs: REQ-115. Affected phase: P14a.
C3 — orca job migrate semantics (resolved)
Question: Does orca job migrate --to <node> (a) drain+reschedule
(uses P05 drain + P06 alloc history), or (b) live-migrate with storage
replication (uses P09 syncthing, much harder)?
Decision: (a) drain+reschedule for v0.11. It composes existing P05/P06 work. Live-migrate with storage replication (b) is a v1.x concern (requires P09 syncthing replication to be production-tested + a storage-replication-aware scheduler).
Affected REQs: REQ-116. Affected phase: P05.
C4 — Remediation cooldown refinement (resolved, design refinement)
Question: Doc 5's D-232 cooldown (5-min) should not apply on transient remediation failures (SSH down, render tree missing) — only on successful remediation. Else a 30s network blip blocks re-remediation for 5 min.
Decision: Refine D-232: cooldown applies only on successful remediation; transient failures (SSH down, render tree missing, applier non-zero exit) retry on the next aggregator tick (10s) without entering cooldown. This is baked into REQ-108 and the D-232 rationale in PROJECT.md.
Affected REQs: REQ-108. Affected phase: P10.
C5 — orca doctor mTLS depth (resolved)
Question: Does orca doctor mTLS just verify the trust chain (CA →
server cert → workload SVIDs exist + not expired), or does it also do a
live mTLS handshake probe to each peer?
Decision: Both. Chain verification is cheap (local file reads + cert parsing); live probe reuses P01 (metrics endpoint) + P01.5 (SPIFFE spike) infrastructure. The doctor check reports both: chain integrity (static) + live handshake (dynamic). A failed live handshake with a valid chain indicates a network/config problem, not a cert problem.
Affected REQs: REQ-118. Affected phase: P15.5.
Open questions
None. All 5 clarifications resolved. Phase 0 proceeds to RESEARCH.