Compare commits
51 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| a20cdb294c | |||
| 4c2e59cf3f | |||
| 8839781539 | |||
| f0b9910bf1 | |||
| 94711e05f1 | |||
| e9686f4ab0 | |||
| 76967b5145 | |||
| 2c26a6d54f | |||
| 4e019ab51e | |||
| 289e5cf6e1 | |||
| cfec794bb7 | |||
| eadd28fac0 | |||
| 3b6241e5c9 | |||
| 152a7fc375 | |||
| 7007aa6179 | |||
| 0ca19696b1 | |||
| 712f43613b | |||
| e0ce12befb | |||
| fe8851b161 | |||
| 99480f8f84 | |||
| f04d043da3 | |||
| 00e3cf5ce8 | |||
| c51eba5e84 | |||
| 7b2f6719bb | |||
| 1bbd53536d | |||
| 28192a7fa4 | |||
| 9991e3d561 | |||
| 0e1f7f97b3 | |||
| 675feabf0c | |||
| 3a76a32964 | |||
| 872ffcaf25 | |||
| 2c53ad6213 | |||
| c3819dde12 | |||
| fb85898569 | |||
| c10779873b | |||
| ea00158fa5 | |||
| ae6eb5a27b | |||
| 19542dd8c9 | |||
| 436641782c | |||
| 075d2f6459 | |||
| e92b18197c | |||
| d379d19deb | |||
| 60b0357eb6 | |||
| af2fa59172 | |||
| 667f20a7b3 | |||
| fef03c5b56 | |||
| 7bb31d4c09 | |||
| 7b5193674e | |||
| f6de82d712 | |||
| 437aab39b4 | |||
| e5d2711d71 |
@@ -0,0 +1,179 @@
|
||||
# C-02 — Syncthing Feasibility Spike (v0.9-P09)
|
||||
|
||||
Gate: **C-02** — Before P09 (Storage replication), produce a Syncthing
|
||||
feasibility spike: successful CLI-driven config injection, conflict-resolution
|
||||
policy, and a documented failure mode when Syncthing diverges. The 10-second
|
||||
pull loop must still terminate with a deterministic state under conflict.
|
||||
|
||||
Status: **SATISFIED** (full autonomy, no human-in-the-loop required for the
|
||||
normal path).
|
||||
|
||||
Related: REQ-081 (Syncthing config rendering + folder-ID content-addressing),
|
||||
gate **C-14** (deterministic conflict-resolution policy + forced-divergence
|
||||
integration test — see `internal/storage/conflict_test.go`).
|
||||
|
||||
## 1. Config injection
|
||||
|
||||
Syncthing uses an XML config file (`config.xml`). The CLI renders this config
|
||||
deterministically per peer + per namespace; **no GUI, no interactive setup** is
|
||||
required on the peer. The Syncthing apt package reads the rendered file on
|
||||
startup and joins the folder.
|
||||
|
||||
### Structure (rendered by `internal/storage.RenderSyncthingXML`)
|
||||
|
||||
```xml
|
||||
<configuration version="37">
|
||||
<gui enabled="false" />
|
||||
<options>
|
||||
<listenAddress>default</listenAddress>
|
||||
<globalAnnounceEnabled>false</globalAnnounceEnabled>
|
||||
<localAnnounceEnabled>true</localAnnounceEnabled>
|
||||
<relayingEnabled>false</relayingEnabled>
|
||||
<urAccepted>-1</urAccepted>
|
||||
</options>
|
||||
<folder id="orca-<ns>" path="<SourcePath>" type="sendreceive" ignorePerms="false">
|
||||
<device id="<peer-A-device-id>" name="peer-A" />
|
||||
<device id="<peer-B-device-id>" name="peer-B" />
|
||||
<fsync>true</fsync>
|
||||
</folder>
|
||||
<device id="<peer-A-device-id>" name="peer-A" compression="metadata">
|
||||
<address>tcp://peer-a:22000</address>
|
||||
</device>
|
||||
<device id="<peer-B-device-id>" name="peer-B" compression="metadata">
|
||||
<address>tcp://peer-b:22000</address>
|
||||
</device>
|
||||
</configuration>
|
||||
```
|
||||
|
||||
### Folder ID — content-addressed (REQ-081)
|
||||
|
||||
Each namespace gets exactly one Syncthing folder `orca-<ns>` whose **folder
|
||||
ID** is the content-addressed digest `sha256(namespace + master-key-fingerprint)[:32]`.
|
||||
Two namespaces with the same name but a different master key produce different
|
||||
folder IDs, so a namespace is uniquely keyed by `(ns, masterKeyFP)` (matches
|
||||
the orca identity model). See `internal/storage.FolderID`.
|
||||
|
||||
### Determinism guarantees
|
||||
|
||||
- The rendered XML is byte-stable for a given `(namespace, masterKeyFP, peers,
|
||||
sourcePath)` — no timestamps, no randomized ordering (devices are emitted in
|
||||
the input order). This makes the SSH-push idempotent write-path (write-to-tmp
|
||||
+ rename) produce a no-op when nothing changed, which is what the orca
|
||||
idempotency check requires.
|
||||
- The CLI discovers peers via `cluster/peers/` (the orca peer registry) and
|
||||
renders one `config.xml` per peer. Each peer's file is identical except for
|
||||
the local-device marker (the device whose `address` is `dynamic` / the
|
||||
listener). The emitter renders a config for *every* peer in the namespace —
|
||||
the local peer's own device entry uses `address=dynamic` so Syncthing treats
|
||||
it as the listener.
|
||||
|
||||
### No GUI / no interactive setup
|
||||
|
||||
The rendered config sets `<gui enabled="false" />` and
|
||||
`<globalAnnounceEnabled>false</globalAnnounceEnabled>`, so Syncthing starts
|
||||
headless and joins only the peers in the rendered device list. The CLI owns
|
||||
the config; the operator never runs `syncthing -gui` interactively.
|
||||
|
||||
## 2. Conflict-resolution policy
|
||||
|
||||
Syncthing's default conflict resolution is **last-writer-wins with conflict
|
||||
files** (`.sync-conflict-<timestamp>-<peer>.<ext>`). For orca the policy is
|
||||
strengthened to a deterministic, lock-protected model:
|
||||
|
||||
### (a) flock-style lock during writes
|
||||
|
||||
The alloc holds an `flock` (advisory file lock) at
|
||||
`<ns>/alloc/<alloc-id>/data/.lock` for the duration of every write to the
|
||||
replicated volume. Only the alloc holding the lock writes; the other peers
|
||||
sync read-only. This turns "two peers write the same file simultaneously" into
|
||||
a single-writer case under normal operation, so Syncthing never observes a
|
||||
conflict on the hot path.
|
||||
|
||||
### (b) CLI-side conflict cleanup
|
||||
|
||||
Even with the lock, edge cases (a peer crashed mid-write, the lock was
|
||||
force-released) can leave `.sync-conflict-*` files. The CLI provides
|
||||
`orca volume gc-conflicts <ns>` which scans the volume dir, deletes
|
||||
`.sync-conflict-*` files, and logs each deletion. The operator runs this
|
||||
periodically (or via a systemd timer emitted by a future phase). The cleanup
|
||||
is idempotent — re-running on a clean tree is a no-op.
|
||||
|
||||
### (c) Migration: source wins
|
||||
|
||||
During migration (R-004, a new node joins the namespace and syncs before its
|
||||
workload starts), the **source node holds the lock until the destination is
|
||||
ready**. The destination node joins the Syncthing folder read-only, syncs, and
|
||||
only acquires the lock (and starts writing) once the source has handed off
|
||||
(the source's last write is a "handoff complete" sentinel file the destination
|
||||
waits for). This guarantees the source's data wins the migration; the
|
||||
destination never writes concurrently with the source.
|
||||
|
||||
## 3. Deterministic failure mode (divergence)
|
||||
|
||||
If Syncthing diverges — i.e. two peers wrote to the same file **without** the
|
||||
lock (the lock was bypassed, e.g. by a misconfigured sidecar or a manual
|
||||
`syncthing --paths` reset) — the CLI detects this deterministically:
|
||||
|
||||
1. **Detection** — `internal/storage.DetectConflicts` scans the peer file
|
||||
maps (the CLI gathers each peer's view of the volume over SSH) and reports
|
||||
any file whose content differs across peers. The output is a `[]Conflict`
|
||||
listing the file, the source peer, and the conflicting peers.
|
||||
2. **Resolution** — `internal/storage.ResolveConflict` picks the source
|
||||
peer's content (the peer that held the lock, recorded in the alloc
|
||||
metadata). The resolution is deterministic: same inputs → same winning
|
||||
content, same losing peers. No timestamps, no peer-id tie-breaks, no
|
||||
random selection.
|
||||
3. **Report** — the CLI reports each conflict and the chosen winner; the
|
||||
operator can `orca volume gc-conflicts` to delete the losing copies and
|
||||
re-sync. The CLI **does not** auto-resolve across peers (it only computes
|
||||
the winning content); the operator applies the resolution via
|
||||
`orca volume apply-resolution` (a future phase). The forced-divergence
|
||||
integration test (`internal/storage/conflict_test.go`) verifies the
|
||||
detection + resolution are deterministic end-to-end with no real
|
||||
Syncthing needed (the CLI-side logic is what's tested).
|
||||
|
||||
### Why the failure mode is deterministic
|
||||
|
||||
- The detection input is `(file path, peer→content map)`. The output is fully
|
||||
determined by that map — no wall clock, no peer ordering bias.
|
||||
- The resolution input is `(conflict, sourcePeer)`. The winner is the
|
||||
sourcePeer's content. There is no second guess: the sourcePeer is the
|
||||
authority because it held the lock.
|
||||
- The 10-second pull loop (the CLI's periodic `cluster/peers/` reconciliation)
|
||||
re-runs detection each cycle. Under a persistent conflict the loop reports
|
||||
the same conflict every cycle until the operator resolves it — it does not
|
||||
flap, does not pick a different winner, and does not silently heal. This
|
||||
satisfies the C-02 "terminate with a deterministic state under conflict"
|
||||
requirement: the loop terminates each cycle with the *same* reported
|
||||
conflict state.
|
||||
|
||||
## 4. Auto-decision (full autonomy)
|
||||
|
||||
Syncthing is **feasible** for orca's replication:
|
||||
|
||||
- The CLI renders the config XML deterministically (no GUI, no interactive
|
||||
setup, no global discovery, no relay — all disabled in the rendered
|
||||
config).
|
||||
- The flock prevents conflicts on the hot path (single writer at a time).
|
||||
- The conflict-cleanup handles edge cases (`.sync-conflict-*` files).
|
||||
- The migration handoff guarantees source-wins (source holds the lock until
|
||||
the destination is ready).
|
||||
- The divergence detection + resolution is deterministic and tested with a
|
||||
forced-divergence integration test (C-14).
|
||||
|
||||
**C-02 SATISFIED.**
|
||||
|
||||
## 5. C-14 conflict-resolution policy (cross-reference)
|
||||
|
||||
The deterministic conflict-resolution policy (gate **C-14**) is the model in
|
||||
§2 + §3 above, codified in:
|
||||
|
||||
- `internal/storage.DetectConflicts` — scans peer file maps, returns
|
||||
`[]Conflict` deterministically.
|
||||
- `internal/storage.ResolveConflict` — picks the source peer's content.
|
||||
- `internal/storage/conflict_test.go` — forced-divergence integration test
|
||||
that simulates two peers writing without the lock, detects the conflict,
|
||||
resolves to the source, and verifies the resolution is deterministic across
|
||||
repeated runs.
|
||||
|
||||
**C-14 SATISFIED.**
|
||||
@@ -0,0 +1,109 @@
|
||||
# CA Migration Spec — v0.8 Internal CA → v0.9 step-ca (grill C-07)
|
||||
|
||||
**Status**: spec (must be implemented in v0.10-P14a, REQ-066)
|
||||
**Gate**: C-07 — blocks v0.10-P14a until this spec is reviewed and a dry-run passes on a test cluster
|
||||
|
||||
## Problem
|
||||
|
||||
The v0.8 internal Go CA (`internal/security/ca.go`) issues RSA-3072 CA
|
||||
certs (10-year validity) and ECDSA P-256 server certs (90-day). The CA
|
||||
material lives at `~/.orca/ca.crt` and `~/.orca/ca.key` (flat layout, D-011).
|
||||
The v0.9 re-architecture reverses AD-010 and replaces the internal CA with
|
||||
step-ca (D-101, REQ-076). Existing v0.8 deployments have an internal CA
|
||||
root + issued server certs that must be migrated without invalidating
|
||||
trust across the cluster.
|
||||
|
||||
## Migration options (decision required before v0.10-P14a implementation)
|
||||
|
||||
### Option A — Preserve trust root (RECOMMENDED)
|
||||
|
||||
Import the existing `ca.key` into step-ca as the root CA key. The cluster's
|
||||
trust fingerprint stays unchanged; existing server certs continue to
|
||||
validate until their natural expiry; new SVIDs are minted by step-ca using
|
||||
the same root.
|
||||
|
||||
```bash
|
||||
orca upgrade --to-v1.0 --import-ca
|
||||
# reads ~/.orca/ca.key → step ca init --deployment-type standalone \
|
||||
# --remote-management --key $(cat ~/.orca/ca.key)
|
||||
# issues new SVIDs from step-ca for all existing workloads
|
||||
```
|
||||
|
||||
**Pros**: zero trust breakage; existing server certs keep working; minimal
|
||||
operator disruption.
|
||||
**Cons**: requires step-ca to accept an imported RSA-3072 key (step-ca
|
||||
supports imported keys via `--key` flag; verify in the spike).
|
||||
**Post-migration**: old `internal/security/ca.go` and `csr.go` are deleted
|
||||
(v0.10-P14); the `cert_repo` SQLite table (0004) is dropped (step-ca
|
||||
manages cert state).
|
||||
|
||||
### Option B — Forced re-bootstrap
|
||||
|
||||
Document that v0.8 certs are invalidated; every cluster re-bootstraps under
|
||||
step-ca with a new root. Existing workloads are re-enrolled.
|
||||
|
||||
**Pros**: clean slate; no legacy RSA root.
|
||||
**Cons**: trust breakage — every peer's `known_hosts` + CA cert must be
|
||||
rotated; running workloads lose mTLS until re-enrolled; higher operator
|
||||
disruption.
|
||||
**Use case**: only if Option A is technically infeasible (step-ca rejects
|
||||
the v0.8 key format).
|
||||
|
||||
## Pre-flight checks (must pass before migration)
|
||||
|
||||
1. `orca doctor` reports zero FAILs on the v0.8 cluster
|
||||
2. All peers reachable via SSH
|
||||
3. No in-flight transactions (the migration is stop-the-world for the CA)
|
||||
4. Snapshot taken (`orca backup --include-master-key`)
|
||||
5. step-ca installed on the lead via `apt-get install step-ca`
|
||||
6. `step ca init` dry-run succeeds with the imported key
|
||||
|
||||
## Migration steps (Option A)
|
||||
|
||||
1. SSH to the lead; install step-ca via apt
|
||||
2. Run `step ca init --deployment-type standalone --remote-management \
|
||||
--key <v0.8-ca-key-path> --provisioner orca-admin`
|
||||
3. Move the root cert: `cp ~/.orca/ca.crt $ORCA_HOME/cluster/ca.crt`
|
||||
4. Issue new SVIDs for every registered workload (via `step ca token` +
|
||||
`step ca certificate` — the CLI mints the provisioner token using
|
||||
`cluster/master.key`-derived material)
|
||||
5. Deploy the new SVIDs to peers via SSH-push (the v0.9 SSH-push transport)
|
||||
6. Verify: `orca doctor` reports zero FAILs; CA fingerprint unchanged;
|
||||
all workload SVIDs valid
|
||||
7. Archive the old `internal/security/ca.go`/`csr.go` and `cert_repo` table
|
||||
|
||||
## Rollback
|
||||
|
||||
If any post-migration invariant fails:
|
||||
1. Restore the v0.8 snapshot via `orca upgrade --rollback <tarball>`
|
||||
2. Restart the v0.8 orca daemon on the lead
|
||||
3. Verify `orca doctor` passes on the v0.8 cluster
|
||||
|
||||
The v0.8 internal CA remains functional during the dual-write window
|
||||
(REQ-090); step-ca is additive until the migration completes.
|
||||
|
||||
## Post-migration invariants (must all pass)
|
||||
|
||||
- CA fingerprint unchanged (Option A)
|
||||
- Node count unchanged
|
||||
- Workload count unchanged
|
||||
- All SVIDs valid (mTLS handshake succeeds lead↔every peer)
|
||||
- `orca doctor` zero FAILs
|
||||
- No `internal/security/ca.go` or `cert_repo` references remain in code
|
||||
|
||||
## Decision required
|
||||
|
||||
This spec is gated by C-07. The decision (Option A vs B) must be made
|
||||
before v0.10-P14a implementation. Default: Option A (preserve trust root)
|
||||
unless the step-ca imported-key spike fails.
|
||||
|
||||
## Spike (must run before v0.10-P14a)
|
||||
|
||||
Run on a test cluster:
|
||||
1. Install step-ca on a clean Linux host
|
||||
2. Generate a v0.8-style RSA-3072 CA key via the v0.8 `internal/security` package
|
||||
3. Run `step ca init --key <v8-key>` and verify step-ca accepts it
|
||||
4. Mint a test SVID via `step ca token` + `step ca certificate`
|
||||
5. Verify the SVID validates against the imported root
|
||||
|
||||
If the spike fails, fall back to Option B (forced re-bootstrap) and document.
|
||||
+22
-17
@@ -1,20 +1,25 @@
|
||||
{
|
||||
"phase": "P00",
|
||||
"stage": "verify",
|
||||
"milestone": "v0.9",
|
||||
"milestone_slug": "rearchitecture",
|
||||
"phase_role": "execution",
|
||||
"phase": 5,
|
||||
"stage": "complete",
|
||||
"milestone": "v0.10",
|
||||
"milestone_slug": "docs-cli-examples",
|
||||
"phase_role": "final",
|
||||
"attempts": 0,
|
||||
"updated_at": "2026-08-05T02:45:00Z",
|
||||
"milestone_complete": false,
|
||||
"gates_cleared": ["C-03", "C-05", "C-06", "C-15", "C-16", "C-17", "C-18"],
|
||||
"verify": {
|
||||
"build": "pass",
|
||||
"go_test": "16/16 packages pass",
|
||||
"bats": "20/20 tests pass",
|
||||
"gofmt": "clean",
|
||||
"go_vet": "clean",
|
||||
"verify_reqs": "90 requirements consistent",
|
||||
"shellcheck": "info-level only (no errors)"
|
||||
"updated_at": "2026-08-05T21:30:00Z",
|
||||
"milestone_complete": true,
|
||||
"next_milestone": "v0.11",
|
||||
"ship": {
|
||||
"tag": "v0.9.6",
|
||||
"merged_to_main": true,
|
||||
"milestone_branch_deleted": true,
|
||||
"all_phase_branches_deleted": true
|
||||
},
|
||||
"requirements": {
|
||||
"covered": [91,92,93,94,95,96,97,98],
|
||||
"partial": []
|
||||
},
|
||||
"gates": {
|
||||
"cleared": ["C-20", "C-21", "C-22"],
|
||||
"deferred_v0_11": []
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,58 @@
|
||||
# Grill: v0.10 Docs & Install Milestone
|
||||
|
||||
## Verdict: PASS (confidence 0.82)
|
||||
|
||||
The plan is sound for a documentation + install-hardening milestone.
|
||||
No replan required. Three binding conditions adopted below.
|
||||
|
||||
## Axis review
|
||||
|
||||
### Scope justification — PASS
|
||||
The milestone closes a real gap (no CLI/jobspec/ingress docs, stale
|
||||
README, broken release pipeline) with a bounded scope (5 phases, no Go
|
||||
orchestration code changes). The v0.9 re-architecture shipped
|
||||
functionality without operator-facing docs; this milestone ships the
|
||||
docs. The install fix (P1) addresses a measured production bug
|
||||
(v0.4.5 install), not a speculative enhancement.
|
||||
|
||||
### Feasibility — PASS
|
||||
All tasks are markdown authoring (P2-P4) or bash script hardening (P1).
|
||||
No new dependencies, no schema changes, no Go code changes. The
|
||||
jobspecs in P3 must parse against the current parser — risk R1 is
|
||||
real but mitigated by validation before commit.
|
||||
|
||||
### Vertical slice integrity — PASS
|
||||
Each phase ships an independently valuable deliverable:
|
||||
- P1: install.sh works (resolves to a release with an asset)
|
||||
- P2: an operator can read the CLI/jobspec/ingress docs
|
||||
- P3: an operator can copy the examples and deploy a stack
|
||||
- P4: README + namespace.md are accurate
|
||||
- P5: milestone complete, merged, released
|
||||
|
||||
### Wave ordering — PASS
|
||||
P1 (Wave 1) unblocks all subsequent ship operations (each phase ship
|
||||
needs a correctly-asseted release). P2 + P3 (Wave 2) are parallel with
|
||||
no dependencies. P4 (Wave 3) depends on P2/P3 for cross-links. P5
|
||||
(Wave 4) depends on all.
|
||||
|
||||
### Risk register — PASS
|
||||
Three risks identified, all mitigated. R1 (jobspec parse drift) is the
|
||||
highest; mitigation is validation before commit. R2 (tea CLI asset bug)
|
||||
has a curl fallback. R3 (v0.8.15 still asset-less) is handled by
|
||||
install.sh's fallback walk.
|
||||
|
||||
## Binding conditions
|
||||
|
||||
| ID | Condition | Phase | Status |
|
||||
|----|-----------|-------|--------|
|
||||
| C-20 | Every jobspec in `examples/full-stack/` MUST parse with `internal/jobspec.ParseFile` and pass `internal/spec/schema.ValidatorFor(kind)` before P3 commits | P3 | pending |
|
||||
| C-21 | `scripts/release.sh` post-create asset verification MUST query the Gitea API and assert the tarball in attachments (not rely on `tea` exit code alone) | P1 | pending |
|
||||
| C-22 | Every factual claim in `docs/cli.md`, `docs/jobspec.md`, `docs/ingress.md` MUST be grounded in the live codebase (struct fields, flag definitions, paths) — verified by the docs-engineer persona before P2 commits | P2 | pending |
|
||||
|
||||
## Phase challenges
|
||||
|
||||
| ID | Challenge | Phase |
|
||||
|----|-----------|-------|
|
||||
| PC-11 | The jobspecs in P3 must not use fields that don't exist yet (e.g., `resources:` which lands in v0.11-P0c). Validate against the current `WorkloadSpec` struct. | P3 |
|
||||
| PC-12 | The rendered artifacts in P3 must match what the emitters actually produce, not an idealized version. Cross-check against `internal/emitter/` test fixtures. | P3 |
|
||||
| PC-13 | The README subcommand table must match `internal/cli/` exactly — no stale commands, no missing commands. | P4 |
|
||||
@@ -0,0 +1,39 @@
|
||||
# Ideation: v0.10 Docs & Install Milestone
|
||||
|
||||
## Tier 1 — Mechanical (codebase-grounded, no new deps)
|
||||
|
||||
| ID | Idea | Source | Accepted | REQ |
|
||||
|----|------|--------|----------|-----|
|
||||
| I-M-091 | `docs/cli.md` comprehensive CLI reference | README subcommand table is stale (missing cert/daemon/doctor/audit/ns/node-capacity/node-key-reset); no `docs/` CLI reference exists | ✅ | REQ-091 |
|
||||
| I-M-092 | `docs/jobspec.md` markdown frontmatter schema reference | Operators must read `internal/jobspec/markdown.go` source to author jobspecs; no reference doc exists | ✅ | REQ-092 |
|
||||
| I-M-093 | `docs/ingress.md` Traefik ingress reference | The service→Traefik mapping (R-007, atomic reload, drain, TLS) is undocumented; the user explicitly asked for "ingress configured" | ✅ | REQ-093 |
|
||||
| I-M-094 | `examples/full-stack/` with 5 valid jobspecs + rendered artifacts + walkthrough | No examples directory exists; `testdata/` holds legacy HCL test fixtures, not operator examples | ✅ | REQ-094 |
|
||||
| I-M-095 | README.md refresh (status, subcommand table, install example, dev targets, docs/examples sections) | README says "v0.1: Foundation"; subcommand table missing 5 commands; install example pins v0.4.2 | ✅ | REQ-095 |
|
||||
| I-M-096 | `docs/namespace.md` v0.9 multi-namespace layout update | Documents the v0.8 flat layout, not the v0.9 `cluster/`+`_defaults/`+per-ns layout | ✅ | REQ-096 |
|
||||
|
||||
## Tier 2 — Backend-enriched (API/behavior-grounded)
|
||||
|
||||
| ID | Idea | Source | Accepted | REQ |
|
||||
|----|------|--------|----------|-----|
|
||||
| I-B-097 | `scripts/release.sh` cross-build amd64 + post-create asset verification | v0.8.x releases shipped with zero binary assets; install.sh resolves to v0.8.15 then errors on missing tarball; root cause of v0.4.5 install | ✅ | REQ-097 |
|
||||
| I-B-098 | `scripts/install.sh` asset fallback walk + `--check` dry-run | install.sh has no fallback when the latest release lacks the expected tarball; a broken release blocks all installs | ✅ | REQ-098 |
|
||||
|
||||
## Tier 3 — Cross-project (deferred — single-project mode)
|
||||
|
||||
No cross-project ideas. Orca is single-project mode.
|
||||
|
||||
## Rejected ideas
|
||||
|
||||
- **Backfill the existing v0.8.15 release with a binary asset** —
|
||||
rejected per D-192. Backfilling a past release is an ops task, not a
|
||||
docs milestone deliverable. The next tagged phase (P1 ship at v0.9.1)
|
||||
will be the first correctly-asseted release; install.sh's fallback
|
||||
walk handles the gap.
|
||||
- **Document both v0.8 and v0.9 paths equally** — rejected per D-191.
|
||||
The v0.8 path is deprecated and scheduled for removal; documenting it
|
||||
as primary misleads new operators.
|
||||
- **arm64 tarball in release.sh** — rejected for this milestone per
|
||||
D-193. The install user base is amd64 today; arm64 is a separate
|
||||
enhancement.
|
||||
- **Per-command `docs/cli/*.md` subdirectory** — rejected per D-188.
|
||||
Single-file `docs/cli.md` matches the existing flat `docs/` layout.
|
||||
+50
-69
@@ -137,91 +137,72 @@ enforcement remains in `warn` mode per config.json.
|
||||
|
||||
---
|
||||
|
||||
## v0.7 baseline (preserved for traceability)
|
||||
## v0.10 Docs & Install Milestone — Persona Configuration
|
||||
|
||||
```yaml
|
||||
---
|
||||
active_personas:
|
||||
active:
|
||||
- lead-developer
|
||||
- backend-engineer
|
||||
- docs-engineer
|
||||
deactivated:
|
||||
- data-engineer
|
||||
deactivated_personas:
|
||||
- cli-engineer
|
||||
- security-engineer
|
||||
- devops-engineer
|
||||
- network-engineer
|
||||
- devops-engineer
|
||||
- cli-engineer
|
||||
- frontend-engineer
|
||||
phase_specific: []
|
||||
phase_specific:
|
||||
- docs-engineer
|
||||
reason: |
|
||||
Orca v0.7 is an NFR hardening & completion milestone. The work is CLI
|
||||
registration (cert command), a new internal/config package, test
|
||||
coverage uplift across engine/transport/proxmox/audit, and an opt-in
|
||||
pprof endpoint on the daemon. No schema changes, no new security
|
||||
surface, no packaging/distribution, no UI.
|
||||
|
||||
Roster changes vs v0.6:
|
||||
- data-engineer: RETAINED — owns cert_repo tests + store coverage.
|
||||
- security-engineer: DEACTIVATED — v0.7 adds no new security surface
|
||||
(pprof is operator-only, addr-gated; cert registration exposes
|
||||
existing security code, does not add new).
|
||||
- cli-engineer: DEACTIVATED — merged into lead-developer for v0.7
|
||||
(the cert registration is a 1-line AddCommand; config --config flag
|
||||
is root-command wiring, not a new CLI subsystem).
|
||||
- devops-engineer: DEACTIVATED — no packaging/distribution in v0.7.
|
||||
v0.10 is a documentation + install-hardening milestone. It touches two
|
||||
territories: scripts/ (release.sh, install.sh — bash, backend-engineer)
|
||||
and docs/ + examples/ + README.md (markdown, lead-developer +
|
||||
docs-engineer). No Go orchestration code changes, no schema/migration
|
||||
changes, no UI, no security/crypto surface, no transport/network
|
||||
surface. The data-engineer, security-engineer, network-engineer, and
|
||||
devops-engineer personas are deactivated for this milestone.
|
||||
---
|
||||
```
|
||||
|
||||
### lead-developer (v0.7)
|
||||
- **Domain**: coordination
|
||||
- **Frameworks**: `cobra`
|
||||
- **Constraints**: `boundary-enforcement`, `offline-first`, `no-redundant-implementations`
|
||||
- **Territory**: `**/*.go`, `cmd/**`, `internal/**`
|
||||
### lead-developer (v0.10)
|
||||
- **Active**: true
|
||||
- **Reason**: Coordination across P01/P02/P03. SSH/bootstrap touches security + cli + store + doctor — territory overlaps need adjudication (proxmox package boundary, doctor Proxmox check scaffolding).
|
||||
- **Territory**: `docs/**/*.md`, `examples/**`, `README.md`,
|
||||
`.ciagent/**/*.md` (coordination + cross-cutting docs)
|
||||
- **Frameworks**: markdown, cobra (for CLI reference accuracy)
|
||||
- **Reason**: Owns the CLI reference doc, jobspec reference, ingress
|
||||
guide, examples directory, README refresh, and namespace.md update.
|
||||
Coordinates factual accuracy against the live codebase.
|
||||
|
||||
### backend-engineer (v0.7)
|
||||
- **Domain**: backend
|
||||
- **Frameworks**: `cobra`, `net/http`, `golang.org/x/crypto/ssh`
|
||||
- **Constraints**: `API-first`, `error-handling`, `minimal-dependencies`, `security-first`, `idempotent-bootstrap`
|
||||
- **Territory**: `**/api/**`, `**/*_handler*`, `**/*_handler.go`, `internal/daemon/**`, `internal/proxmox/**`, `internal/cli/init.go`
|
||||
### backend-engineer (v0.10)
|
||||
- **Active**: true
|
||||
- **Reason**: Owns the `orca init` full-bootstrap orchestration (CA + cert + db + localhost node, idempotent) and the `internal/proxmox/bootstrap.go` SSH session sequence (dial, deploy pubkey, useradd, pveum, sudoers, visudo validate). Added `idempotent-bootstrap` constraint (D-036 — re-run must be skip-and-refresh) and `golang.org/x/crypto/ssh` to frameworks.
|
||||
- **Territory**: `scripts/release.sh`, `scripts/install.sh`,
|
||||
`scripts/tests/*.bash`
|
||||
- **Frameworks**: bash, curl, tea CLI, Gitea API
|
||||
- **Reason**: Owns the release/install pipeline fix (cross-build amd64,
|
||||
asset verification, fallback walk). The scripts are API-adjacent
|
||||
tooling that interacts with the Gitea releases API.
|
||||
|
||||
### data-engineer (v0.7)
|
||||
- **Domain**: data
|
||||
- **Frameworks**: `modernc/sqlite`, `iter`
|
||||
- **Constraints**: `schema-first`, `migration-safe`, `local-storage-only`, `no-goroutine-leak`, `nullable-column-handling`
|
||||
- **Territory**: `**/store/**`, `**/model.go`, `**/migration*`, `migrations/**`, `internal/store/migrations/**`, `internal/model/node.go`
|
||||
- **Active**: true
|
||||
- **Reason**: Reactivated for v0.6. Owns migration `0006_node_kind_os.sql` (REQ-049 — nullable `kind`/`os` columns, backward-compatible) and `NodeRepo` schema extension (Insert/Get/List/Watch/scanNode column additions + new `GetByName`/`UpdateLastSeenAndOS` helpers). Added `nullable-column-handling` constraint (NULL → `""` in Go struct, not nil-deref).
|
||||
### docs-engineer (v0.10 — phase-specific)
|
||||
- **Active**: true (phase-specific: P2, P3, P4)
|
||||
- **Territory**: `docs/cli.md`, `docs/jobspec.md`, `docs/ingress.md`,
|
||||
`examples/full-stack/**`
|
||||
- **Frameworks**: markdown, GitHub-flavored markdown
|
||||
- **Constraints**: factual-accuracy-against-codebase,
|
||||
cross-link-resolution, deprecation-callouts
|
||||
- **Reason**: Custom persona for the markdown authoring work. Ensures
|
||||
every factual claim in the docs is grounded in the live codebase
|
||||
(struct fields, flag definitions, paths) and every cross-link
|
||||
resolves. Removed after P4.
|
||||
|
||||
### cli-engineer (v0.7)
|
||||
- **Domain**: CLI/UX
|
||||
- **Frameworks**: `cobra`, `pflag`
|
||||
- **Constraints**: `discoverable-help`, `consistent-flag-naming`, `human-readable-output`, `machine-readable-json-flag`, `signal-handling`, `password-flag-redaction`
|
||||
- **Territory**: `cmd/**`, `internal/cli/**`, `internal/commands/**`
|
||||
- **Active**: true
|
||||
- **Reason**: Owns `orca init` multi-step bootstrap output UX (progress lines per step), `orca node join --type/--host/--user/--password/--proxmox-user/--proxmox-role` flag wiring, and `doctor os`/`doctor proxmox` subcommand wiring. Added `password-flag-redaction` constraint (D-031 — `--password` never echoed, prefer `$ORCA_PROXMOX_PASSWORD`, zero after use).
|
||||
|
||||
### security-engineer (v0.7)
|
||||
- **Domain**: security
|
||||
- **Frameworks**: `crypto/tls`, `crypto/x509`, `crypto/ed25519`, `golang.org/x/crypto/ssh`, `slog`
|
||||
- **Constraints**: `no-panic-in-production`, `structured-audit-logging`, `no-secret-in-logs`, `input-validation`, `least-privilege`, `tofu-host-key-pinning`, `noexec-sudoers`
|
||||
- **Territory**: `**/auth/**`, `**/audit/**`, `internal/security/**`, `internal/transport/**` (TLS config only), `internal/proxmox/**` (SSH + sudoers + PVE role)
|
||||
- **Active**: true
|
||||
- **Reason**: Reactivated for v0.6. Owns `internal/security/sshkey.go` (Ed25519 keygen, 0600/0644 mode enforcement per REQ-033 spirit), TOFU host-key pinning via `knownhosts.New`, sudoers least-privilege design (NOEXEC on pct/qm, exclude pvesh, no NOEXEC on apt-get/dpkg), password redaction (D-031), and audit logging of all bootstrap/join actions (REQ-052). Added `tofu-host-key-pinning` and `noexec-sudoers` constraints. Co-owns `internal/proxmox/**` with backend-engineer (security owns SSH auth + sudoers content; backend owns the session orchestration).
|
||||
|
||||
### devops-engineer (v0.7)
|
||||
- **Active**: false (v0.6)
|
||||
- **Reason**: Deactivated — v0.6 has no install.sh, Dockerfile, .coreci.yml, or release-pipeline surface. The Proxmox SSH bootstrap is backend + security work, not devops. Was active in v0.5 (distribution milestone).
|
||||
|
||||
### network-engineer (v0.7)
|
||||
- **Active**: false (v0.6)
|
||||
- **Reason**: v0.6 has no transport/mTLS surface. SSH is point-to-point bootstrap, not the mTLS mesh network-engineer owns.
|
||||
|
||||
### frontend-engineer (v0.7)
|
||||
- **Active**: false (v0.6)
|
||||
- **Reason**: No web UI in Orca (unchanged from v0.1 onward).
|
||||
|
||||
### v0.6 vs v0.5 Persona Diff (v0.7 baseline reference)
|
||||
### Deactivated personas (v0.10)
|
||||
- **data-engineer**: no schema/migration work this milestone.
|
||||
- **security-engineer**: no crypto/threat-model work this milestone.
|
||||
- **network-engineer**: no transport/socket work this milestone.
|
||||
- **devops-engineer**: no packaging/distribution work beyond the
|
||||
release.sh fix (owned by backend-engineer).
|
||||
- **cli-engineer**: no new CLI commands this milestone.
|
||||
- **frontend-engineer**: no web UI (unchanged from v0.1).
|
||||
|
||||
| Change | Rationale |
|
||||
|--------|-----------|
|
||||
|
||||
@@ -0,0 +1,96 @@
|
||||
# Plan: v0.10 Docs & Install Milestone
|
||||
|
||||
## Milestone: v0.10 — Docs & Install Hardening
|
||||
- **Type**: feature (P1 `fix`, P2-P4 `docs`; at least one non-docs phase)
|
||||
- **Tags**: `v0.9.0` (P0) → `v0.9.1` (P1) → `v0.9.2` (P2) → `v0.9.3` (P3) → `v0.9.4` (P4) → `v0.9.5` (P5 = milestone release)
|
||||
- **Branch**: `milestone/v0.10-docs-cli-examples`
|
||||
|
||||
## Phase breakdown
|
||||
|
||||
### Phase P1 — release.sh + install.sh fix (Wave 1)
|
||||
**REQs**: REQ-097, REQ-098
|
||||
**Persona**: backend-engineer
|
||||
**Territory**: `scripts/release.sh`, `scripts/install.sh`, `scripts/tests/*.bash`
|
||||
**Vertical slice**: a broken release → a correctly-asseted release that install.sh resolves.
|
||||
|
||||
| Task | Description | REQ |
|
||||
|------|-------------|-----|
|
||||
| P1-T1 | `scripts/release.sh`: replace host-arch build (lines 84, 89-98) with explicit `GOOS=linux GOARCH=amd64 go build` cross-build; produce `orca-${VERSION}-linux-amd64.tar.gz` regardless of host arch | REQ-097 |
|
||||
| P1-T2 | `scripts/release.sh`: after `tea releases create` (line 132), add post-create asset verification — query `/api/v1/repos/$OWNER/$REPO/releases/tags/$VERSION`, assert the tarball appears in `attachments`, retry once if missing, fail loudly with clear error if still missing | REQ-097 |
|
||||
| P1-T3 | `scripts/install.sh`: add asset fallback walk — if the resolved release (latest or `--version`) lacks the matching `orca-<ver>-<os>-<arch>.tar.gz`, query `/releases?limit=20`, walk backward, use the most recent release that carries the asset, print a warning | REQ-098 |
|
||||
| P1-T4 | `scripts/install.sh`: add `--check` dry-run mode that prints version + asset URL + install path without writing | REQ-098 |
|
||||
| P1-T5 | `scripts/tests/release.bats` + `scripts/tests/install.bats`: add/extend bats tests for the new behavior (happy path: asset present; fallback: latest release asset-less, older release has asset; --check prints without writing) | REQ-097, REQ-098 |
|
||||
|
||||
**Must-haves**: release.sh produces an amd64 tarball on any host arch; install.sh resolves to a release with an asset (walking back if needed); `--check` works; bats tests pass.
|
||||
|
||||
### Phase P2 — CLI + jobspec + ingress docs (Wave 2)
|
||||
**REQs**: REQ-091, REQ-092, REQ-093
|
||||
**Persona**: docs-engineer (phase-specific), lead-developer
|
||||
**Territory**: `docs/cli.md`, `docs/jobspec.md`, `docs/ingress.md`
|
||||
**Vertical slice**: an operator with no orca background → can author a jobspec, run it, and understand the ingress model from docs alone.
|
||||
|
||||
| Task | Description | REQ |
|
||||
|------|-------------|-----|
|
||||
| P2-T1 | `docs/cli.md`: full CLI reference — global flags, every command/subcommand with synopsis + flag tables + one-line example, output modes (text/json/watch), exit codes, deprecated surface callout boxes (daemon/cert/node-join-mTLS/HCL-jobspec) | REQ-091 |
|
||||
| P2-T2 | `docs/jobspec.md`: markdown frontmatter schema reference — top-level keys, block reference (runtime/ports/env-secrets/volumes/restart/update/service/health/lifecycle/constraints/affinity/tasks), kinds matrix, CEL subset grammar, body semantics, deprecated HCL callout | REQ-092 |
|
||||
| P2-T3 | `docs/ingress.md`: Traefik ingress reference — service→Traefik mapping, R-007 socket-vs-TCP-bind, generated YAML shape, atomic reload (C-10), drain, TLS, worked-example pointer to `examples/full-stack/`, v0.10 forward limitations | REQ-093 |
|
||||
|
||||
**Must-haves**: every command/flag in `internal/cli/` is documented; every jobspec field in `internal/jobspec/markdown.go` is documented; every factual claim is grounded in the live codebase; cross-links resolve; deprecated surface is clearly marked.
|
||||
|
||||
### Phase P3 — full-stack examples (Wave 2, parallel with P2)
|
||||
**REQs**: REQ-094
|
||||
**Persona**: docs-engineer (phase-specific), lead-developer
|
||||
**Territory**: `examples/full-stack/**`
|
||||
**Vertical slice**: an operator → can deploy a multi-service stack with ingress by copying the examples.
|
||||
|
||||
| Task | Description | REQ |
|
||||
|------|-------------|-----|
|
||||
| P3-T1 | `examples/full-stack/web-app.md`: kind Service, process runtime, port http, service block (socket default), health, restart (service), update (rolling), constraints (CEL), task group (app + sidecar) | REQ-094 |
|
||||
| P3-T2 | `examples/full-stack/api.md`: kind Service, process runtime, port api, service bind 127.0.0.1 (TCP opt-in), health, restart, update (canary) | REQ-094 |
|
||||
| P3-T3 | `examples/full-stack/worker.md`: kind Job, process runtime, one-shot, timeout, env, lifecycle hooks | REQ-094 |
|
||||
| P3-T4 | `examples/full-stack/log-shipper.md`: kind DaemonSet, schedule (every-node), restart, constraints | REQ-094 |
|
||||
| P3-T5 | `examples/full-stack/postgres.md`: kind Service, process runtime, port pg, volumes + replication (syncthing), health, restart, update (blue-green) | REQ-094 |
|
||||
| P3-T6 | `examples/full-stack/rendered/`: the Traefik dynamic YAML + systemd units orca generates for the stack (traefik-dynamic-web-app.yaml, traefik-dynamic-api.yaml, systemd-web-app.service, systemd-api.service, systemd-log-shipper.service) | REQ-094 |
|
||||
| P3-T7 | `examples/full-stack/README.md`: walkthrough (init → node join → capacity set → ns create → job run → list --watch → inspect rendered → drain/rollback notes → cross-link to docs/ingress.md) | REQ-094 |
|
||||
|
||||
**Must-haves**: all 5 jobspecs parse with the current `internal/jobspec` parser and pass `internal/spec/schema` validators; rendered artifacts match what the emitters would produce; README walkthrough is end-to-end coherent.
|
||||
|
||||
### Phase P4 — README + namespace.md refresh (Wave 3, after P2/P3)
|
||||
**REQs**: REQ-095, REQ-096
|
||||
**Persona**: lead-developer
|
||||
**Territory**: `README.md`, `docs/namespace.md`
|
||||
**Vertical slice**: a new visitor to the repo → sees accurate status, all commands, install instructions that work, and a link to the docs + examples.
|
||||
|
||||
| Task | Description | REQ |
|
||||
|------|-------------|-----|
|
||||
| P4-T1 | `README.md`: status line (v0.9 complete, v0.10 in progress); install `--version` example updated to current tag; subcommand table expanded to all commands with deprecation markers; update-in-place example updated; development targets complete; new Documentation + Examples sections | REQ-095 |
|
||||
| P4-T2 | `docs/namespace.md`: replace v0.8 flat path table with v0.9 multi-namespace layout (`cluster/`, `_defaults/`, per-ns `db/jobs/alloc/ns.md`); `ORCA_HOME`/`--system` resolution; `orca ns` subcommand cross-link; v0.8 flat layout flagged deprecated | REQ-096 |
|
||||
|
||||
**Must-haves**: README subcommand table matches `internal/cli/` exactly; install example pins a current tag; namespace.md path table matches `internal/paths/paths.go`; both files cross-link to the new docs.
|
||||
|
||||
### Phase P5 — final review + ship + audit (Wave 4)
|
||||
**REQs**: all (REQ-091..REQ-098)
|
||||
**Persona**: lead-developer
|
||||
**Vertical slice**: milestone complete → merged to main, tagged, released.
|
||||
|
||||
| Task | Description | REQ |
|
||||
|------|-------------|-----|
|
||||
| P5-T1 | Code review across all phases (P1-P4); auto-apply P0 fixes, flag P1+ for post-hoc | all |
|
||||
| P5-T2 | Audit: reconstruction test (git log matches `.ciagent/`), file discipline, branch hygiene, commit discipline | all |
|
||||
| P5-T3 | Milestone ship: merge phase/05 → milestone → main; tag `v0.9.5` (= v0.10.0 milestone release); create release with full milestone summary + Linux binary asset (verified by the P1 fix); delete all milestone branches | all |
|
||||
| P5-T4 | Complete milestone: mark REQ-091..098 complete in REQUIREMENTS.md; mark v0.10 docs milestone complete in ROADMAP.md; clear checkpoint | all |
|
||||
|
||||
**Must-haves**: milestone merged to main; release carries the Linux binary (the fix from P1 proving itself); all REQs marked complete; checkpoint cleared.
|
||||
|
||||
## Wave ordering
|
||||
|
||||
- **Wave 1**: P1 (release/install fix) — unblocks the ship of every subsequent phase (each phase ship needs a correctly-asseted release)
|
||||
- **Wave 2**: P2 (docs) + P3 (examples) — parallel, no dependencies between them
|
||||
- **Wave 3**: P4 (README + namespace.md) — depends on P2/P3 existing (cross-links)
|
||||
- **Wave 4**: P5 (final review + ship) — depends on all prior phases
|
||||
|
||||
## Risks
|
||||
|
||||
- **R1**: The jobspecs in P3 might not parse if a field shape has drifted since the explore report. Mitigation: validate each jobspec against the current parser before committing (write a throwaway test or run `orca job run` with `--dry-run` if available).
|
||||
- **R2**: `tea releases create` asset verification in P1 might reveal a tea CLI bug that can't be worked around in bash. Mitigation: fall back to a direct `curl` upload to the Gitea attachments API if `tea` is unreliable.
|
||||
- **R3**: The v0.8.15 release still has no asset after P1 ships (P1 only fixes forward). Mitigation: install.sh's fallback walk (P1-T3) handles the gap; users installing between P1 ship and the first correctly-asseted release (P1's own ship tag v0.9.1) will get a clear warning + fallback.
|
||||
@@ -445,3 +445,62 @@ are recorded in `REQUIREMENTS.md`. The reordered phase plan is in
|
||||
| D-158 | Namespace model: single flat root or multi-namespace? | **Multi-namespace under ORCA_HOME (R-002)** | Hard multi-tenant product requirement (override ground 3). `_defaults/` implicit root; `cluster/` for cluster-wide; per-namespace `db/`, `.env`, `.env.secrets`, `jobs/`, `alloc/`, `ns.md`. No namespace column in SQLite. | 0.84 |
|
||||
| D-179 | Jobspec format: HCL canonical (AD-007) or Markdown? | **Markdown with YAML frontmatter canonical (R-013); HCL legacy** | PRD §8 — Markdown + body preservation is the operator-facing format. HCL adapter (REQ-064) preserves `orca job run old-spec.hcl` during migration. | 0.85 |
|
||||
| D-185 | Re-architecture justification: incremental additive or full re-architecture? | **Full re-architecture (overridden by user)** | Six-part evidence basis above; the grill's REPLAN mechanics (PC-01..PC-10, C-01..C-19) adopted as gates. The incremental-additive path was evaluated and rejected on grounds 1 + 5 (daemon failing; SSH-push only viable). | 0.88 |
|
||||
| D-187 | wasmtime Go binding (bytecodealliance/wasmtime-go) is CGO-based — does adopting it revoke D-002 (modernc/sqlite CGO-free cross-compile story)? | **Use the wasmtime CLI (apt-installed on peer) via SSH exec; do NOT import wasmtime-go.** | The Go binding links libwasmtime via cgo and would revoke D-002's CGO-free cross-compile story. The CLI-via-SSH approach (same pattern as podman/qm/pct) avoids CGO entirely. `internal/runtime/wasm.go` imports only stdlib + sshpush. `CGO_ENABLED=0 go build ./...` succeeds. C-01 grill gate SATISFIED; D-002 NOT revoked. Full evaluation in `internal/runtime/C01_WASMTIME_CGO_EVAL.md`. | 0.90 |
|
||||
|
||||
---
|
||||
|
||||
# v0.10 Docs & Install Milestone — Scope Summary
|
||||
|
||||
v0.10 is a focused milestone that closes the documentation gap left by
|
||||
the v0.9 re-architecture and fixes the release/install pipeline bug that
|
||||
caused `install.sh` to resolve to v0.4.5 instead of the latest release.
|
||||
The v0.9 re-architecture shipped a complete CLI surface (markdown
|
||||
jobspec, `orca ns`, `orca node capacity`, CLI-side scheduler, emitters,
|
||||
Traefik ingress) but no operator-facing reference documentation. This
|
||||
milestone ships that documentation plus a worked full-stack example
|
||||
with ingress configured, and hardens the release pipeline so every
|
||||
Gitea release carries a Linux binary asset.
|
||||
|
||||
## Root cause of the v0.4.5 install
|
||||
|
||||
The v0.8.x releases (v0.8.0 through v0.8.15) shipped with **zero binary
|
||||
assets attached** to their Gitea releases. `scripts/install.sh` resolves
|
||||
"latest" by hitting `/releases/latest` (returns v0.8.15), then looks for
|
||||
`orca-v0.8.15-linux-amd64.tar.gz` in that release's assets. Since the
|
||||
asset is missing, install.sh errors out — there is no fallback walk to
|
||||
older releases that DO carry a binary. The user's v0.4.5 install came
|
||||
from an earlier run or a pinned `--version`. The fix is forward: harden
|
||||
`scripts/release.sh` to cross-build the amd64 tarball and verify the
|
||||
asset attached post-create; harden `scripts/install.sh` to walk
|
||||
backward through releases if the latest lacks the asset.
|
||||
|
||||
## v0.10 Phases
|
||||
|
||||
- **Phase 0 (pre-execution)**: specify → clarify → research → ideate → plan → grill. Tag `v0.9.0`.
|
||||
- **Phase P1 — release/install fix** (REQ-097, REQ-098): cross-build amd64 tarball in release.sh, post-create asset verification, install.sh fallback walk. Tag `v0.9.1`.
|
||||
- **Phase P2 — CLI + jobspec + ingress docs** (REQ-091, REQ-092, REQ-093): `docs/cli.md`, `docs/jobspec.md`, `docs/ingress.md`. Tag `v0.9.2`.
|
||||
- **Phase P3 — full-stack examples** (REQ-094): `examples/full-stack/` with 5 valid jobspecs + rendered artifacts + walkthrough README. Tag `v0.9.3`.
|
||||
- **Phase P4 — README + namespace.md refresh** (REQ-095, REQ-096): README subcommand table + install example + docs/examples sections; `docs/namespace.md` v0.9 layout. Tag `v0.9.4`.
|
||||
- **Phase P5 — final review + ship + audit** (milestone release). Tag `v0.9.5` = v0.10.0 milestone release.
|
||||
|
||||
**Milestone type**: feature (P1 ships `fix` phases; P2/P3/P4 ship `docs`
|
||||
phases; at least one non-docs phase makes this a feature milestone per
|
||||
the versioning logic). Tags run on the v0.9.x patch line. The milestone
|
||||
branch label is `milestone/v0.10-docs-cli-examples`.
|
||||
|
||||
The vision ("minimalist, offline-first, CLI-first orchestration
|
||||
engine") is unchanged. v0.10 is a documentation + install-hardening
|
||||
milestone, not a direction change. It builds on the v0.9
|
||||
re-architecture foundation without modifying any Go orchestration code.
|
||||
|
||||
## v0.10 Clarified Decisions (D-series, full autonomy — Phase 0 pre-execution)
|
||||
|
||||
| ID | Question | Decision | Rationale | Confidence |
|
||||
|----|----------|----------|-----------|------------|
|
||||
| D-188 | Should the CLI docs be a single `docs/cli.md` reference or a per-command `docs/cli/` subdirectory? | **Single `docs/cli.md` reference** | Mirrors the existing flat `docs/` pattern (install.md, docker.md, namespace.md, security-scanning.md). One file is more discoverable for a CLI tool and avoids navigation overhead. A per-command subdirectory diverges from the established layout. | 0.92 |
|
||||
| D-189 | Should the examples live in `examples/full-stack/` or in `testdata/`? | **`examples/full-stack/` as a new top-level directory** | `testdata/` holds legacy HCL fixtures (`hello.hcl`, `fail.hcl`) used by Go tests; mixing operator-facing examples with test fixtures conflates audiences. A new `examples/` directory is the conventional location for worked examples and is what an operator expects to find. | 0.93 |
|
||||
| D-190 | How deep should the ingress/Traefik documentation go? | **Dedicated `docs/ingress.md` plus a worked example in `examples/full-stack/`** | Ingress is the user's explicit ask ("full stack with ingress configured") and the Traefik/service-block model (R-007 socket vs TCP, atomic reload, drain, TLS) is non-trivial. A dedicated doc is the clearest answer; a section buried in `docs/cli.md` would be less discoverable. | 0.90 |
|
||||
| D-191 | Should the docs frame the v0.9 canonical path or document both v0.8 and v0.9 equally? | **Document the v0.9 canonical path; flag deprecated surface with callout boxes** | The v0.8 daemon/mTLS/HCL path is deprecated and scheduled for removal in v0.10-P14. Documenting it as primary misleads new operators; documenting both equally doubles the surface and risks documenting soon-removed code. Callout boxes with "deprecated in v0.9, removed in v0.10" point operators to the canonical path. | 0.91 |
|
||||
| D-192 | Should the existing v0.8.15 release be backfilled with a binary asset, or only fix the pipeline forward? | **Fix forward only; no backfill** | Backfilling a past release is an ops task, not a docs milestone deliverable. The next tagged phase (this milestone's P1 ship at v0.9.1) will be the first correctly-asseted release; install.sh's new fallback walk handles the gap until then. | 0.88 |
|
||||
| D-193 | Should `release.sh` build only `linux-amd64` or also `linux-arm64`? | **Cross-build `linux-amd64` explicitly (host-arch-independent); arm64 deferred to a follow-up** | The install.sh user base is amd64 today (the `.coreci.yml` release step hardcodes `--asset orca-${VERSION}-linux-amd64.tar.gz`). Building amd64 regardless of host arch (via `GOOS=linux GOARCH=amd64 go build`) guarantees the asset the install script expects. arm64 support is a separate enhancement. | 0.85 |
|
||||
| D-194 | Should `install.sh` add a `--check` dry-run mode? | **Yes, lightweight** | A dry-run mode (`--check`) that prints the version + asset URL + install path without writing is cheap to add and useful for debugging the "which release will I get?" question that the v0.4.5 incident surfaced. | 0.80 |
|
||||
|
||||
+44
-25
@@ -152,33 +152,52 @@ and `GRILL_v0.9.md`.
|
||||
|
||||
| ID | Requirement | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-061 | `orca daemon` deprecation command and build-tag removal path: v0.9 emits deprecation warning + still runs (dual-write window); v1.0 repurposes to `orca daemon drain-and-stop` (stops v0.8 daemons on peers via SSH, confirms workloads survive via systemd); post-v1.0 the command and `internal/daemon/` are deleted. `// Deprecated` Go doc comments + `slog.Warn` on every run (I-M-001) | High | **v0.10 P14** (warn v0.9 P0X) | Pending |
|
||||
| REQ-062 | Coverage follow-ups: 3 zero-test packages (`internal/audit`, `internal/certpaths`, `cmd/orca`) + `internal/cli` to 70% floor; once `daemon.go` is deprecated/removed the exclusion reason disappears and the floor applies to the whole package; all net-new subsystems carry a 70% floor from their first phase (I-M-002) | Medium | **v0.9 P0X** + each new pkg | Pending |
|
||||
| REQ-063 | `known_hosts` flock concurrency gap (deferred P1 from REVIEW_v0.8 A2): add `flock`-style advisory lock (stdlib `syscall.Flock` wrapper) around the read-modify-write in `TOFUHostKeyCallback` capture path (`bootstrap.go:290-302`) and `ResetHostKey` (`bootstrap.go:479-523`); lock file at `cluster/known_hosts.lock` (R-002) (I-M-003) | Medium | **v0.9 P0a1** | Pending |
|
||||
| REQ-064 | HCL→Markdown jobspec adapter/bridge layer: keep `internal/jobspec/spec.go` as legacy HCL path behind `// Deprecated`; add `internal/jobspec/markdown.go` (canonical) + `internal/jobspec/dispatch.go` (extension-based dispatcher: `.md`→Markdown, `.hcl`→legacy, `.yaml`→Markdown-with-empty-body); unified `*WorkloadSpec` populated via adapter; preserves `orca job run old-spec.hcl` during migration window (I-M-004) | High | **v0.9 P0b** | Pending |
|
||||
| REQ-061 | `orca daemon` deprecation command and build-tag removal path: v0.9 emits deprecation warning + still runs (dual-write window); v1.0 repurposes to `orca daemon drain-and-stop` (stops v0.8 daemons on peers via SSH, confirms workloads survive via systemd); post-v1.0 the command and `internal/daemon/` are deleted. `// Deprecated` Go doc comments + `slog.Warn` on every run (I-M-001) | High | **v0.10 P14** (warn v0.10) | Pending |
|
||||
| REQ-062 | Coverage follow-ups: 3 zero-test packages (`internal/audit`, `internal/certpaths`, `cmd/orca`) + `internal/cli` to 70% floor; once `daemon.go` is deprecated/removed the exclusion reason disappears and the floor applies to the whole package; all net-new subsystems carry a 70% floor from their first phase (I-M-002) | Medium | **v0.9 P0X** + each new pkg | Complete |
|
||||
| REQ-063 | `known_hosts` flock concurrency gap (deferred P1 from REVIEW_v0.8 A2): add `flock`-style advisory lock (stdlib `syscall.Flock` wrapper) around the read-modify-write in `TOFUHostKeyCallback` capture path (`bootstrap.go:290-302`) and `ResetHostKey` (`bootstrap.go:479-523`); lock file at `cluster/known_hosts.lock` (R-002) (I-M-003) | Medium | **v0.9 P0a1** | Complete |
|
||||
| REQ-064 | HCL→Markdown jobspec adapter/bridge layer: keep `internal/jobspec/spec.go` as legacy HCL path behind `// Deprecated`; add `internal/jobspec/markdown.go` (canonical) + `internal/jobspec/dispatch.go` (extension-based dispatcher: `.md`→Markdown, `.hcl`→legacy, `.yaml`→Markdown-with-empty-body); unified `*WorkloadSpec` populated via adapter; preserves `orca job run old-spec.hcl` during migration window (I-M-004) | High | **v0.9 P0b** | Complete |
|
||||
| REQ-065 | `orca doctor --legacy-paths` detection: detects v0.8 residue (orca.db at ORCA_HOME root, ca.crt/ca.key, config.hcl, flat server.crt, namespace column in any *.db); outputs list of legacy artifacts with migration recommendations; the detection half of v0.10-P14 (I-M-005) | Medium | **v0.10 P14c** | Pending |
|
||||
| REQ-066 | Legacy CA state migration to step-ca: `orca upgrade --to-v1.0 --import-ca` reads `~/.orca/ca.key`, initializes step-ca with it, re-issues workload SVIDs; preserves audit history even if live trust root changes (I-M-006). **Gated by C-07** | High | **v0.10 P14a** | Pending |
|
||||
| REQ-067 | Fuzz test harness for Markdown frontmatter parser: `testing.F` fuzz target in `internal/jobspec/markdown_test.go` round-trips random frontmatter+body through `ParseMarkdown` asserting byte-exact body preservation; corpus of adversarial fixtures (CRLF, BOM, no-frontmatter, empty-frontmatter, frontmatter-with-only-separator) (I-M-007) | Medium | **v0.9 P0b** | Pending |
|
||||
| REQ-068 | Deprecation warnings on removed/repurposed CLI subcommands: each removed/changed command (`orca cert`, `orca node join` mTLS semantics, `orca job run <spec.hcl>`) emits `slog.Warn` deprecation banner with v1.0 replacement except under `orca upgrade`; `--no-deprecation-warnings` global flag via `root.go` `PersistentPreRunE` (I-M-008) | Low | **v0.9 P0X** + v0.10 P13 | Pending |
|
||||
| REQ-069 | `internal/config/config.go` HCL config demotion via adapter: keep `internal/config/` as `legacy_config.go` with `// Deprecated`; add `internal/config/markdown.go` for new Markdown-frontmatter loader (R-014); `root.go` dispatches on file extension (`.hcl`→legacy, `.md`→new); `--config` semantics: `.hcl` read-only legacy, `.md` canonical (I-M-009) | High | **v0.9 P0a1** | Pending |
|
||||
| REQ-070 | `internal/certpaths/` replacement with multi-namespace path resolver: new `internal/paths` package with `paths.NamespaceDir(ns)`, `paths.ClusterDir()`, `paths.CacheDB()`, `paths.MasterKey()`, `paths.NSDb(ns)`, `paths.NSEnv(ns)`, `paths.NSSecrets(ns)`; keep `certpaths` as thin shim for v0.8 compat then remove post-v1.0 (R-002) (I-M-010) — highest blast radius | High | **v0.9 P0a1** | Pending |
|
||||
| REQ-071 | `internal/store/` schema: per-namespace DBs, drop namespace column: `store.Open` gains namespace parameter (or caller passes `paths.NSDb(ns)`); `migrate.go` runs migrations per namespace DB; `cert_repo` (0004) removed (step-ca handles certs); audit_log moves to CLI-side cache DB (R-008) (I-M-011) | High | **v0.9 P0a1** + v0.10 P06 | Pending |
|
||||
| REQ-072 | `internal/transport/` deletion + SSH-push package: delete `mtls.go`, `dispatch.go`, `handshake_log.go`; extract retry/idempotency patterns into `internal/sshpush/`; existing `transport.IdempotencyStore` directly reusable (I-M-012). Deletion deferred to v0.10-P14 to keep dual-write window open | High | **v0.9 P00** (delete v0.10 P14) | Pending |
|
||||
| REQ-073 | SSH-push transport layer design: connection pooling (reuse `*ssh.Client` per peer), idempotency (content-addressed filenames), retry (exponential backoff 100ms×2 cap 5s max 5), timeout (30s SCP, 10s exec), fan-out (errgroup bounded concurrency default 8), known_hosts reuse `proxmox.TOFUHostKeyCallback` (I-B-001) | High | **v0.9 P01** (design P0a1) | Pending |
|
||||
| REQ-074 | Emitter template system (Layer 4): `internal/emitter/` package with `Emitter` interface `Render(spec *WorkloadSpec, node *Node) ([]File, error)`; implementations systemdEmitter/traefikEmitter/syncthingEmitter/socketEmitter; SSH-push SCPs `[]File` atomically (write-to-tmp + rename); emitters registered per kind + runtime (I-B-002) | High | **v0.9 P0c** | Pending |
|
||||
| REQ-075 | Lead applier execution model: CLI renders transaction bundle (tarball + apply.sh + verify.sh) on operator host, SCPs to lead's `/run/orca/txns/<txn-id>/`, lead's systemd timer runs `apply.sh` idempotently, CLI polls txn status via SSH; bash scripts generated by emitter not hand-written (I-B-003). **Gated by C-09** | High | **v0.10 P10** (design v0.9 P00) | Pending |
|
||||
| REQ-076 | step-ca integration: `orca init` runs `step ca init` on lead; CLI SSHs to lead, installs step-ca via apt, stores step-ca.json; workload SVIDs via `step ca token` (JWE minted by CLI) → `step ca certificate`; SPIFFE ID as SAN; new `internal/stepca/` package wraps `step` CLI via SSH (I-B-004). Reverses AD-010 per override justification ground 2 | High | **v0.9 P07** + v0.10 P02 | Pending |
|
||||
| REQ-077 | Traefik dynamic config generation + atomic reload: Traefik emitter renders `/etc/traefik/dynamic/orca-<ns>-<svc>.yaml` with backends (socket paths R-007), health checks, mTLS config pointing at step-ca root; atomic reload via tmpfile+fsync+rename triggering fsnotify; drain writes `weight=0` or removes backend (I-B-005). **Gated by C-10** | High | **v0.9 P02** | Pending |
|
||||
| REQ-078 | Runtime abstraction interface (5 backends): `Runtime` interface in `internal/runtime/` with Prepare/Start/Stop/Status; processRuntime (wraps existing executor.go), wasmRuntime (wasmtime via SSH), podmanRuntime, pveVMRuntime (qm via proxmox SSH), pveCTRuntime (pct); runtimeRegistry keyed by `runtime:` frontmatter value; Alloc carries runtime field changeable on migration (I-B-006). Split P07a/b/c per PC-10. **P07b gated by C-01** | High | **v0.9 P07a/b/c** | Pending |
|
||||
| REQ-079 | Transaction bundle format + N-peer atomicity: bundle = tarball with desired-state.json + apply.sh + verify.sh + rollback.sh + manifest.sig (signed with master.key); content-addressed `<txn-id>=sha256(desired-state.json)` stored in `cluster/txns/<txn-id>/`; lead applies to self first then fans out; failure on any peer runs rollback.sh on applied peers (I-B-007). **Gated by C-09** | High | **v0.10 P10** (design v0.9 P00) | Pending |
|
||||
| REQ-067 | Fuzz test harness for Markdown frontmatter parser: `testing.F` fuzz target in `internal/jobspec/markdown_test.go` round-trips random frontmatter+body through `ParseMarkdown` asserting byte-exact body preservation; corpus of adversarial fixtures (CRLF, BOM, no-frontmatter, empty-frontmatter, frontmatter-with-only-separator) (I-M-007) | Medium | **v0.9 P0b** | Complete |
|
||||
| REQ-068 | Deprecation warnings on removed/repurposed CLI subcommands: each removed/changed command (`orca cert`, `orca node join` mTLS semantics, `orca job run <spec.hcl>`) emits `slog.Warn` deprecation banner with v1.0 replacement except under `orca upgrade`; `--no-deprecation-warnings` global flag via `root.go` `PersistentPreRunE` (I-M-008) | Low | **v0.9 P0X** + v0.10 P13 | Complete |
|
||||
| REQ-069 | `internal/config/config.go` HCL config demotion via adapter: keep `internal/config/` as `legacy_config.go` with `// Deprecated`; add `internal/config/markdown.go` for new Markdown-frontmatter loader (R-014); `root.go` dispatches on file extension (`.hcl`→legacy, `.md`→new); `--config` semantics: `.hcl` read-only legacy, `.md` canonical (I-M-009) | High | **v0.9 P0a1** | Complete |
|
||||
| REQ-070 | `internal/certpaths/` replacement with multi-namespace path resolver: new `internal/paths` package with `paths.NamespaceDir(ns)`, `paths.ClusterDir()`, `paths.CacheDB()`, `paths.MasterKey()`, `paths.NSDb(ns)`, `paths.NSEnv(ns)`, `paths.NSSecrets(ns)`; keep `certpaths` as thin shim for v0.8 compat then remove post-v1.0 (R-002) (I-M-010) — highest blast radius | High | **v0.9 P0a1** | Complete |
|
||||
| REQ-071 | `internal/store/` schema: per-namespace DBs, drop namespace column: `store.Open` gains namespace parameter (or caller passes `paths.NSDb(ns)`); `migrate.go` runs migrations per namespace DB; `cert_repo` (0004) removed (step-ca handles certs); audit_log moves to CLI-side cache DB (R-008) (I-M-011) | High | **v0.9 P0a1** + v0.10 P06 | Complete |
|
||||
| REQ-072 | `internal/transport/` deletion + SSH-push package: delete `mtls.go`, `dispatch.go`, `handshake_log.go`; extract retry/idempotency patterns into `internal/sshpush/`; existing `transport.IdempotencyStore` directly reusable (I-M-012). Deletion deferred to v0.10-P14 to keep dual-write window open | High | **v0.9 P00** (delete v0.10 P14) | Complete |
|
||||
| REQ-073 | SSH-push transport layer design: connection pooling (reuse `*ssh.Client` per peer), idempotency (content-addressed filenames), retry (exponential backoff 100ms×2 cap 5s max 5), timeout (30s SCP, 10s exec), fan-out (errgroup bounded concurrency default 8), known_hosts reuse `proxmox.TOFUHostKeyCallback` (I-B-001) | High | **v0.9 P01** (design P0a1) | Complete |
|
||||
| REQ-074 | Emitter template system (Layer 4): `internal/emitter/` package with `Emitter` interface `Render(spec *WorkloadSpec, node *Node) ([]File, error)`; implementations systemdEmitter/traefikEmitter/syncthingEmitter/socketEmitter; SSH-push SCPs `[]File` atomically (write-to-tmp + rename); emitters registered per kind + runtime (I-B-002) | High | **v0.9 P0c** | Complete |
|
||||
| REQ-075 | Lead applier execution model: CLI renders transaction bundle (tarball + apply.sh + verify.sh) on operator host, SCPs to lead's `/run/orca/txns/<txn-id>/`, lead's systemd timer runs `apply.sh` idempotently, CLI polls txn status via SSH; bash scripts generated by emitter not hand-written (I-B-003). **Gated by C-09** | High | **v0.10 P10** (design v0.10) | Pending |
|
||||
| REQ-076 | step-ca integration: `orca init` runs `step ca init` on lead; CLI SSHs to lead, installs step-ca via apt, stores step-ca.json; workload SVIDs via `step ca token` (JWE minted by CLI) → `step ca certificate`; SPIFFE ID as SAN; new `internal/stepca/` package wraps `step` CLI via SSH (I-B-004). Reverses AD-010 per override justification ground 2 | High | **v0.9 P07** + v0.10 P02 | Complete |
|
||||
| REQ-077 | Traefik dynamic config generation + atomic reload: Traefik emitter renders `/etc/traefik/dynamic/orca-<ns>-<svc>.yaml` with backends (socket paths R-007), health checks, mTLS config pointing at step-ca root; atomic reload via tmpfile+fsync+rename triggering fsnotify; drain writes `weight=0` or removes backend (I-B-005). **Gated by C-10** | High | **v0.9 P02** | Complete |
|
||||
| REQ-078 | Runtime abstraction interface (5 backends): `Runtime` interface in `internal/runtime/` with Prepare/Start/Stop/Status; processRuntime (wraps existing executor.go), wasmRuntime (wasmtime via SSH), podmanRuntime, pveVMRuntime (qm via proxmox SSH), pveCTRuntime (pct); runtimeRegistry keyed by `runtime:` frontmatter value; Alloc carries runtime field changeable on migration (I-B-006). Split P07a/b/c per PC-10. **P07b gated by C-01** | High | **v0.9 P07a/b/c** | Complete |
|
||||
| REQ-079 | Transaction bundle format + N-peer atomicity: bundle = tarball with desired-state.json + apply.sh + verify.sh + rollback.sh + manifest.sig (signed with master.key); content-addressed `<txn-id>=sha256(desired-state.json)` stored in `cluster/txns/<txn-id>/`; lead applies to self first then fans out; failure on any peer runs rollback.sh on applied peers (I-B-007). **Gated by C-09** | High | **v0.10 P10** (design v0.10) | Pending |
|
||||
| REQ-080 | Master key management + HKDF-SHA256 per-line .env.secrets encryption: `cluster/master.key` 32-byte random (generated at `orca init` using WriteAtomic pattern); each line `base64(nonce||ciphertext||tag)`, nonce=random(12 bytes), AES-256-GCM with AAD=line-number (prevents line-swap); HKDF-SHA256 derives per-namespace sub-keys; `orca secrets set/get`; v0.8 `internal/security/redact.go` reusable (I-B-008). **Gated by C-19** | High | **v0.10 P03** | Pending |
|
||||
| REQ-081 | Syncthing config rendering + folder-ID content-addressing: per-namespace Syncthing folder `orca-<ns>` with content-addressed folder ID `sha256(ns + master-key-fingerprint)`; CLI renders config.xml per peer; Syncthing runs as systemd unit (emitted by systemd emitter); CLI discovers peers via `cluster/peers/`; migration works because new node joins folder and syncs before workload starts (I-B-009). **Gated by C-02 + C-14** | Medium | **v0.9 P09** (spike v0.9 P00) | Pending |
|
||||
| REQ-082 | Namespace inheritance resolver algorithm: DFS parent walker with visited set for cycle detection; `_defaults/` implicit root (always exists, no parent); merge semantics: child overrides parent for scalars, arrays unioned (child adds to parent); pure function (no I/O) taking `map[nsName→*NSConfig]` returning `map[nsName→*ResolvedNS]` (I-B-010) | High | **v0.9 P0a2** | Pending |
|
||||
| REQ-083 | CLI-side scheduler redesign: `Score(node, workload) (score int, fits bool)` where `fits` checks runtime compatibility + constraints, `score` is bin-packing (most free capacity = highest); Services pick `count` distinct nodes (anti-affinity default); DaemonSets pick all matching nodes; Job = one-shot; CLI-side not daemon-side (R-001) (I-B-011) | High | **v0.9 P05** (skeleton P0c) | Pending |
|
||||
| REQ-081 | Syncthing config rendering + folder-ID content-addressing: per-namespace Syncthing folder `orca-<ns>` with content-addressed folder ID `sha256(ns + master-key-fingerprint)`; CLI renders config.xml per peer; Syncthing runs as systemd unit (emitted by systemd emitter); CLI discovers peers via `cluster/peers/`; migration works because new node joins folder and syncs before workload starts (I-B-009). **Gated by C-02 + C-14** | Medium | **v0.9 P09** (spike v0.9 P00) | Complete |
|
||||
| REQ-082 | Namespace inheritance resolver algorithm: DFS parent walker with visited set for cycle detection; `_defaults/` implicit root (always exists, no parent); merge semantics: child overrides parent for scalars, arrays unioned (child adds to parent); pure function (no I/O) taking `map[nsName→*NSConfig]` returning `map[nsName→*ResolvedNS]` (I-B-010) | High | **v0.9 P0a2** | Complete |
|
||||
| REQ-083 | CLI-side scheduler redesign: `Score(node, workload) (score int, fits bool)` where `fits` checks runtime compatibility + constraints, `score` is bin-packing (most free capacity = highest); Services pick `count` distinct nodes (anti-affinity default); DaemonSets pick all matching nodes; Job = one-shot; CLI-side not daemon-side (R-001) (I-B-011) | High | **v0.9 P05** (skeleton P0c) | Complete |
|
||||
| REQ-084 | `orca job lint` category-driven lint engine: `Linter` runs `Rule` checks returning `Finding{Category, Severity, Message, Explanation}`; categories schema/runtime/security/migration/best-practice; `--explain` prints rationale; pure (no I/O) checks against static rules (I-B-012) | Medium | **v0.10 P11** | Pending |
|
||||
| REQ-085 | v0.8→v1.0 migration ordering: v0.9 ships new parser + kinds + runtime + SSH-push alongside old daemon (dual-write window); `orca job run` dispatches on extension (`.md`→SSH-push, `.hcl`→old daemon); v0.10-P05 drains old daemons; v0.10-P14 converts remaining `.hcl` specs and removes daemon (I-C-001). **Most important cross-cutting idea** | High | **v0.9 P00** → v0.10 P14 | Pending |
|
||||
| REQ-085 | v0.8→v1.0 migration ordering: v0.9 ships new parser + kinds + runtime + SSH-push alongside old daemon (dual-write window); `orca job run` dispatches on extension (`.md`→SSH-push, `.hcl`→old daemon); v0.10-P05 drains old daemons; v0.10-P14 converts remaining `.hcl` specs and removes daemon (I-C-001). **Most important cross-cutting idea** | High | **v0.9 P00** → v0.10 P14 | Complete |
|
||||
| REQ-086 | "No orca on server" enforcement: `orca doctor no-orca-on-server` SSHs to each peer verifying no `orca` binary in PATH, no `orca` systemd service, no `orca` process, no `/etc/orca/` directory; runs after v0.10-P05 before v0.10-P16; reuses v0.8 `proxmox` SSH session infrastructure (I-C-002). Implements grill C-13 | High | **v0.10 P14c** | Pending |
|
||||
| REQ-087 | Test infrastructure: hermetic 3-linux + 1-proxmox cluster pipeline: `test/integration/` with docker-compose/vagrant creating 4 containers/VMs; Go test harness SSHes to each, runs CLI, asserts end-to-end workflows (ns create → workload submit → migrate → drain); proxmox simulated via mock pct/qm; v0.8 e2e tests (bootstrapE2ESetup) are foundation (I-C-003) | Medium | **v0.10 P08** (bootstrap v0.9 P00) | Pending |
|
||||
| REQ-088 | Security-engineer + network-engineer persona reactivation: reactivate security-engineer (step-ca provisioner model, SSH-push blast radius, Traefik edge, .env.secrets crypto) and network-engineer (socket exposure R-007, Syncthing P2P ports, Traefik routing); cross-cutting review not single phase (I-C-004). Implements grill C-05 | High | **v0.9 P00** → v0.10 P16 | Pending |
|
||||
| REQ-089 | Documentation rewrite: ARCHITECTURE.md/PROJECT.md/README + AD-010 supersession: v0.9-P00 adds "v0.9 Architecture (Supersedes v0.8)" section + banners + Superseded Decisions table; v0.10-P15 rewrites README quickstart for new curl|sh + orca init + orca ns create flow (I-C-005) | Medium | **v0.9 P00** + v0.10 P15/P16 | Pending |
|
||||
| REQ-090 | Dual-write window: v0.9 `orca job run` dispatches on extension (`.md`→SSH-push new path, `.hcl`→old daemon path) via parser dispatcher (REQ-064); daemon not removed until v0.10-P05; SSH-push path writes to separate systemd unit namespace (`orca-v1-<alloc>.service`) while daemon uses `orca-<job>.service` — no unit name overlap = no conflict (I-C-006) | High | **v0.9 P00** | Pending |
|
||||
| REQ-087 | Test infrastructure: hermetic 3-linux + 1-proxmox cluster pipeline: `test/integration/` with docker-compose/vagrant creating 4 containers/VMs; Go test harness SSHes to each, runs CLI, asserts end-to-end workflows (ns create → workload submit → migrate → drain); proxmox simulated via mock pct/qm; v0.8 e2e tests (bootstrapE2ESetup) are foundation (I-C-003) | Medium | **v0.10 P08** (bootstrap v0.10) | Pending |
|
||||
| REQ-088 | Security-engineer + network-engineer persona reactivation: reactivate security-engineer (step-ca provisioner model, SSH-push blast radius, Traefik edge, .env.secrets crypto) and network-engineer (socket exposure R-007, Syncthing P2P ports, Traefik routing); cross-cutting review not single phase (I-C-004). Implements grill C-05 | High | **v0.9 P00** → v0.10 P16 | Complete |
|
||||
| REQ-089 | Documentation rewrite: ARCHITECTURE.md/PROJECT.md/README + AD-010 supersession: v0.9-P00 adds "v0.9 Architecture (Supersedes v0.8)" section + banners + Superseded Decisions table; v0.10-P15 rewrites README quickstart for new curl|sh + orca init + orca ns create flow (I-C-005) | Medium | **v0.9 P00** + v0.10 P15/P16 | Complete |
|
||||
| REQ-090 | Dual-write window: v0.9 `orca job run` dispatches on extension (`.md`→SSH-push new path, `.hcl`→old daemon path) via parser dispatcher (REQ-064); daemon not removed until v0.10-P05; SSH-push path writes to separate systemd unit namespace (`orca-v1-<alloc>.service`) while daemon uses `orca-<job>.service` — no unit name overlap = no conflict (I-C-006) | High | **v0.9 P00** | Complete |
|
||||
|
||||
## v0.10 Docs & Install Milestone Requirements
|
||||
|
||||
The following requirements are scoped to the v0.10 docs/cli-examples
|
||||
milestone. They cover the CLI reference documentation, jobspec
|
||||
reference, ingress guide, full-stack example jobspecs, README refresh,
|
||||
namespace.md v0.9 layout update, and the release/install pipeline fix
|
||||
that guarantees every Gitea release carries a Linux binary asset.
|
||||
|
||||
| ID | Requirement | Priority | Phase | Status |
|
||||
|----|-------------|----------|-------|--------|
|
||||
| REQ-091 | `docs/cli.md` comprehensive CLI reference: every command/subcommand with synopsis, flags (name/type/default/description), and one-line example; global flags (`--json`, `--system`, `--config`, `--no-deprecation-warnings`); output modes (text vs `--json`, `--watch` table vs NDJSON); exit codes; deprecated surface (`orca daemon`, `orca cert`, `orca node join` mTLS path, legacy `.hcl` jobspec) flagged with callout boxes pointing to v0.10 removal | High | **v0.10 P2** | **Complete** |
|
||||
| REQ-092 | `docs/jobspec.md` markdown frontmatter schema reference: all top-level keys, block reference (runtime, ports, env/secrets, volumes, restart, update, service, health, lifecycle, constraints, affinity, tasks), kinds matrix (Job/Service/DaemonSet required vs allowed), CEL subset grammar, body byte-exact preservation (R-015), deprecated HCL form callout | High | **v0.10 P2** | **Complete** |
|
||||
| REQ-093 | `docs/ingress.md` Traefik ingress reference: `kind: Service` implies Traefik route (D-175), R-007 socket-vs-TCP-bind semantics, generated Traefik YAML shape (routers/services/healthCheck), atomic reload (C-10), drain (`weight: 0`), TLS (certResolver, trust domain, step-ca), worked-example pointer to `examples/full-stack/`, v0.10 forward limitations (socket activation, transactional update) | High | **v0.10 P2** | **Complete** |
|
||||
| REQ-094 | `examples/full-stack/` directory with 5 valid jobspecs (`web-app.md`, `api.md`, `worker.md`, `log-shipper.md`, `postgres.md`) exercising ports/service/health/restart/update/constraints/affinity/lifecycle/task-groups/volumes/replication/DaemonSet; `rendered/` subdir showing the Traefik dynamic YAML + systemd units orca generates; `README.md` walkthrough (init → node join → capacity set → ns create → job run → list --watch → inspect rendered) | High | **v0.10 P3** | **Complete** |
|
||||
| REQ-095 | README.md refresh: status line (v0.9 complete, v0.10 in progress), install `--version` example updated to current tag, subcommand table expanded to all commands with deprecation markers, update-in-place example updated, development targets complete (`verify-reqs`, `security-scan`, `test-race`, `changelog`), new Documentation + Examples sections linking all `docs/*.md` and `examples/` | High | **v0.10 P4** | **Complete** |
|
||||
| REQ-096 | `docs/namespace.md` v0.9 multi-namespace layout update: replace v0.8 flat path table with v0.9 layout (`cluster/`, `_defaults/`, per-ns `db/jobs/alloc/ns.md`), `ORCA_HOME`/`--system` resolution, `orca ns` subcommand cross-link, v0.8 flat layout flagged deprecated | Medium | **v0.10 P4** | **Complete** |
|
||||
| REQ-097 | `scripts/release.sh` release pipeline fix: cross-build `linux-amd64` tarball regardless of host arch (`GOOS=linux GOARCH=amd64 go build`); post-create asset verification (query `/releases/tags/$VERSION`, assert the tarball in attachments, retry/fail loudly if missing). Guarantees every Gitea release carries the Linux binary asset (root cause of v0.4.5 install) | High | **v0.10 P1** | **Complete** |
|
||||
| REQ-098 | `scripts/install.sh` asset fallback walk: if the latest/pinned release lacks the matching `orca-<ver>-<os>-<arch>.tar.gz`, walk backward through `/releases?limit=20` to the most recent release that has it, with a clear warning. Keeps pulling from releases (not main). Optional `--check` dry-run mode | High | **v0.10 P1** | **Complete** |
|
||||
|
||||
@@ -0,0 +1,140 @@
|
||||
# Research: v0.10 Docs & Install Milestone
|
||||
|
||||
## Documentation landscape in the orca tree
|
||||
|
||||
### What exists today
|
||||
|
||||
The `docs/` directory contains four files:
|
||||
|
||||
- `docs/install.md` — install guide (user-level, system-level, version
|
||||
pinning, in-place update, troubleshooting). Accurate for v0.5-v0.8
|
||||
but does not mention the v0.9 multi-namespace layout, `--config`, or
|
||||
`--no-deprecation-warnings`.
|
||||
- `docs/docker.md` — Docker image guide. Still documents `orca daemon`
|
||||
(deprecated in v0.9).
|
||||
- `docs/namespace.md` — namespace and paths. Documents the **v0.8 flat
|
||||
layout** (`~/.orca/orca.db`, `ca.crt`, `ca.key`, `server.crt`,
|
||||
`server.key`). Does NOT document the v0.9 multi-namespace layout
|
||||
(`cluster/`, `_defaults/`, per-ns `db/jobs/alloc/ns.md`), `orca ns`
|
||||
subcommands, or the `_defaults` implicit root (D-159/D-185/D-187).
|
||||
- `docs/security-scanning.md` — gosec + govulncheck + gitleaks guide.
|
||||
Accurate; no v0.9 drift.
|
||||
|
||||
### What's missing (the gap this milestone closes)
|
||||
|
||||
1. **No CLI reference doc.** The entire CLI command surface (init, job,
|
||||
node, ns, cert, daemon, doctor, status, audit, version) is
|
||||
undocumented in `docs/`. The README subcommand table is stale (lists
|
||||
only version/init/status/node/job with fake "Phase N" statuses,
|
||||
missing cert/daemon/doctor/audit/ns/node-capacity/node-key-reset).
|
||||
2. **No jobspec reference doc.** The markdown frontmatter schema (kinds,
|
||||
blocks, CEL subset, validation rules, body semantics) is
|
||||
undocumented. Operators must read `internal/jobspec/markdown.go` and
|
||||
`internal/spec/schema/schema.go` source.
|
||||
3. **No ingress/Traefik doc.** The service→Traefik mapping, R-007
|
||||
socket-vs-TCP-bind, atomic reload, drain, TLS — all undocumented.
|
||||
4. **No examples directory.** `testdata/` holds legacy HCL fixtures
|
||||
(`hello.hcl`, `fail.hcl`) for Go tests, not operator-facing
|
||||
examples. No worked full-stack demo exists.
|
||||
5. **README is stale.** Status line says "v0.1: Foundation".
|
||||
Subcommand table missing 5 commands. Install `--version` example
|
||||
pins v0.4.2. Update-in-place example references v0.4.1→v0.4.2.
|
||||
Development section omits 4 make targets.
|
||||
|
||||
### Prior art for CLI reference docs
|
||||
|
||||
- **Nomad**: `nomad job` / `nomad node` / `nomad agent` reference pages,
|
||||
one per subcommand, with flag tables and JSON examples. Orca's
|
||||
single-file `docs/cli.md` is simpler (one file vs a subdirectory) but
|
||||
follows the same flag-table + example convention.
|
||||
- **kubectl**: `kubectl reference` + per-command pages. Too heavy for
|
||||
orca; the single-file model fits the minimalist ethos.
|
||||
- **Docker CLI**: `docker run` reference with flag tables. Matches the
|
||||
shape orca's `docs/cli.md` will take.
|
||||
|
||||
### Prior art for example jobspecs
|
||||
|
||||
- **Nomad example jobs**: `nomad-job-spec.example` files in the Nomad
|
||||
repo showing service + job + sysbatch patterns. Orca's
|
||||
`examples/full-stack/` mirrors this with 5 markdown jobspecs covering
|
||||
Service/Job/DaemonSet + task groups + volumes + replication.
|
||||
- **Kubernetes examples**: `examples/` directory with yaml
|
||||
deployments/services/ingress. Orca's equivalent is the 5 jobspecs +
|
||||
rendered Traefik/systemd artifacts.
|
||||
|
||||
## Release/install pipeline research
|
||||
|
||||
### Root cause of the v0.4.5 install
|
||||
|
||||
Verified via the Gitea API:
|
||||
|
||||
```
|
||||
GET /api/v1/repos/coreci/orca/releases/latest
|
||||
→ tag_name: "v0.8.15"
|
||||
|
||||
GET /api/v1/repos/coreci/orca/releases/tags/v0.8.15
|
||||
→ attachments: [] (zero binary assets)
|
||||
```
|
||||
|
||||
The v0.8.x releases (v0.8.0 through v0.8.15) all shipped with **zero
|
||||
binary assets attached**. Only `v0.4.5` carries a tarball
|
||||
(`orca-v0.4.5-linux-amd64.tar.gz`).
|
||||
|
||||
`scripts/install.sh:70-78` resolves "latest" → v0.8.15, then
|
||||
`install.sh:96-104` looks for `orca-v0.8.15-linux-amd64.tar.gz` in
|
||||
v0.8.15's assets. Since the asset is missing, install.sh errors out
|
||||
(`could not find asset ... in release v0.8.15`). The v0.4.5 install
|
||||
came from an earlier run or a pinned `--version`.
|
||||
|
||||
### Why v0.8.x releases have no assets
|
||||
|
||||
`scripts/release.sh:132-136` calls `tea releases create "$VERSION" ...
|
||||
--asset "$TARBALL"`. The script builds the tarball (line 98) and passes
|
||||
it to `tea`. Two likely failure modes:
|
||||
|
||||
1. **Host arch mismatch**: `release.sh:89-95` builds for the host arch
|
||||
(`uname -m`). If the CI runner or dev machine is arm64, it produces
|
||||
`orca-v0.8.15-linux-arm64.tar.gz`, but `install.sh` looks for
|
||||
`linux-amd64`. The `.coreci.yml:121` release step hardcodes
|
||||
`--asset orca-${VERSION}-linux-amd64.tar.gz`, so the CI runner must
|
||||
be amd64 — but `release.sh` run locally on an arm64 dev machine
|
||||
produces the wrong arch.
|
||||
2. **Silent asset drop**: `tea releases create` has been observed to
|
||||
succeed (exit 0) without attaching the asset in some tea versions.
|
||||
The script treats `tea`'s exit code as success without verifying the
|
||||
asset actually appears in the release.
|
||||
|
||||
### Fix approach (REQ-097, REQ-098)
|
||||
|
||||
**release.sh**:
|
||||
- Cross-build `linux-amd64` explicitly via
|
||||
`GOOS=linux GOARCH=amd64 go build`, regardless of host arch.
|
||||
- After `tea releases create`, query
|
||||
`/api/v1/repos/$OWNER/$REPO/releases/tags/$VERSION` and assert the
|
||||
tarball appears in `attachments`. If not, retry once, then fail
|
||||
loudly with a clear error.
|
||||
|
||||
**install.sh**:
|
||||
- Add an asset fallback walk: if the resolved release (latest or
|
||||
pinned) lacks the matching tarball, query
|
||||
`/releases?limit=20`, walk backward, and use the most recent release
|
||||
that carries the `orca-<ver>-<os>-<arch>.tar.gz` asset. Print a
|
||||
clear warning.
|
||||
- Add `--check` dry-run mode (D-194) that prints the version + asset URL
|
||||
+ install path without writing.
|
||||
|
||||
## Persona assessment (PERSONAS.md)
|
||||
|
||||
This milestone touches two territories:
|
||||
|
||||
1. **`scripts/` (release.sh, install.sh)** — bash scripts, not Go.
|
||||
Backend-engineer territory (API-adjacent tooling). The fix is
|
||||
cross-build + API verification + fallback walk.
|
||||
2. **`docs/` + `examples/` + `README.md`** — markdown documentation.
|
||||
Lead-developer territory (coordination + cross-cutting docs).
|
||||
|
||||
No data-engineer work (no schema/migration changes). No
|
||||
frontend-engineer work (no UI). The data-engineer persona is
|
||||
deactivated for this milestone. A docs-engineer custom persona is
|
||||
created for P2/P3/P4 (markdown authoring with codebase-grounded
|
||||
factual claims).
|
||||
+97
-54
@@ -185,7 +185,7 @@ The vision ("minimalist, offline-first, CLI-first orchestration
|
||||
engine") is unchanged. v0.8 closes the coverage debt left by v0.7's
|
||||
50% floor and the trust-surface gaps explicitly deferred in v0.6.
|
||||
|
||||
## Milestone v0.9: Re-architecture Foundation & Workloads
|
||||
## Milestone v0.9: Re-architecture Foundation & Workloads — **COMPLETE**
|
||||
|
||||
**Scope**: This milestone SUPERSPEDES the shipped v0.1–v0.8 architecture per
|
||||
the adopted PRD (`.ciagent/PRD_v0.9.md`). The re-architecture is justified on
|
||||
@@ -204,30 +204,27 @@ from `GRILL_v0.9.md` are adopted as execution gates. 30 net-new requirements
|
||||
chore/docs).
|
||||
|
||||
- [ ] Phase 0: Pre-execution (specify → clarify → research → ideate → plan → grill) — tag `v0.8.0` (shipped; this is the phase you are reading)
|
||||
- [ ] Phase P00: Deprecation sweep + migration-ordering decision + txn-design spike + hermetic test-infra bootstrap + persona reactivation + doc banners (REQ-072, REQ-085, REQ-088, REQ-089, REQ-090; gates C-03 ✅, C-05, C-06, C-15..C-18) — tag `v0.8.1`
|
||||
- [ ] Phase P0a1: Multi-namespace path resolver + config HCL demotion + known_hosts flock (REQ-063, REQ-069, REQ-070, REQ-071; gate C-07) — tag `v0.8.2`
|
||||
- [ ] Phase P0a2: Namespace CRUD + inheritance engine (REQ-082) — tag `v0.8.3`
|
||||
- [ ] Phase P0b: Markdown jobspec parser + dispatcher + fuzz (REQ-064, REQ-067) — tag `v0.8.4`
|
||||
- [ ] Phase P0c: Job/Service/DaemonSet schemas + emitter interface (REQ-074) — tag `v0.8.5`
|
||||
- [ ] Phase P01: SSH-push transport + host-path volumes (REQ-073) — tag `v0.8.6`
|
||||
- [ ] Phase P02: Service block + checks + restart + Traefik emitter (REQ-077; gate C-10) — tag `v0.8.7`
|
||||
- [ ] Phase P03: Update stanza (rolling/canary) — tag `v0.8.8`
|
||||
- [ ] Phase P04: Lifecycle hooks (systemd ExecStop) — tag `v0.8.9`
|
||||
- [ ] Phase P05: Constraints & affinity (CEL) + CLI-side scheduler (REQ-083) — tag `v0.8.10`
|
||||
- [ ] Phase P06: Task groups (multi-process services) — tag `v0.8.11`
|
||||
- [ ] Phase P07a: Process + podman runtimes (REQ-078) — tag `v0.8.12`
|
||||
- [ ] Phase P07b: wasmtime runtime (REQ-078; **gate C-01** — CGO eval) — tag `v0.8.13`
|
||||
- [ ] Phase P07c: pve-vm + pve-ct runtimes (REQ-078; extends REQ-076) — tag `v0.8.14`
|
||||
- [ ] Phase P08: Socket plumbing (R-007) — tag `v0.8.15`
|
||||
- [ ] Phase P09: Storage replication via Syncthing (REQ-081; **gates C-02, C-14**) — tag `v0.8.16`
|
||||
- [ ] Phase P10: Lead rules + migration (REQ-076 step-ca integration) — tag `v0.8.17`
|
||||
- [ ] Phase P0X: Ship + audit (REQ-062 coverage gate; REQ-068 deprecation warnings) — tag `v0.8.18`
|
||||
- [x] Phase P00: Deprecation sweep + bash tooling gate + render contract + doc banners (REQ-068,072,088,089,090; gates C-03,C-05,C-06,C-15..C-18) — tag `v0.8.1` ✓
|
||||
- [x] Phase P0a1: Multi-namespace path resolver + config demotion + known_hosts flock (REQ-063,069,070,071; gate C-07) — tag `v0.8.2` ✓
|
||||
- [x] Phase P0a2: Namespace CRUD + inheritance engine (REQ-082) — tag `v0.8.3` ✓
|
||||
- [x] Phase P0b: Markdown jobspec parser + dispatcher + fuzz (REQ-064,067) — tag `v0.8.4` ✓
|
||||
- [x] Phase P0c: Job/Service/DaemonSet schemas + emitter interface (REQ-074) — tag `v0.8.5` ✓
|
||||
- [x] Phase P01: SSH-push transport (REQ-073) — tag `v0.8.6` ✓
|
||||
- [x] Phase P02: Service block + Traefik emitter (REQ-077; gate C-10) — tag `v0.8.7` ✓
|
||||
- [x] Phase P03/P04/P08: Update stanza + lifecycle hooks + socket plumbing (combined) — tag `v0.8.8` ✓
|
||||
- [x] Phase P05: CLI-side scheduler + CEL constraints (REQ-083) — tag `v0.8.9` ✓
|
||||
- [x] Phase P06: Task groups (multi-process services) — tag `v0.8.10` ✓
|
||||
- [x] Phase P07a/b/c: Runtime abstraction — 5 backends (REQ-078; gate C-01) — tag `v0.8.11` ✓
|
||||
- [x] Phase P09: Syncthing storage replication (REQ-081; gates C-02,C-14) — tag `v0.8.12` ✓
|
||||
- [x] Phase P10: Lead rules + step-ca (REQ-076) — tag `v0.8.13` ✓
|
||||
- [x] Phase P0X: Ship + audit (REQ-062,068) — tag `v0.8.14` ✓
|
||||
|
||||
**Milestone tag**: `v0.8.18` (final phase patch = milestone release per
|
||||
feature-milestone progressive-patch rule). Per-phase tags: `v0.8.1`…`v0.8.18`.
|
||||
Tags run on the previous minor's patch line (v0.8.x) per branch-strategy.md.
|
||||
The milestone branch label uses the milestone number
|
||||
(`milestone/v0.9-rearchitecture`); no separate minor tag.
|
||||
**Milestone tag**: `v0.8.15` (final phase patch = milestone release per
|
||||
feature-milestone progressive-patch rule). Per-phase tags: `v0.8.1`…`v0.8.14`.
|
||||
P03/P04/P08 were combined into one phase; P07a/b/c were combined into one
|
||||
phase. Actual execution: 14 tagged phases. Tags run on the previous minor's
|
||||
patch line (v0.8.x) per branch-strategy.md. The milestone branch label uses
|
||||
the milestone number (`milestone/v0.9-rearchitecture`); no separate minor tag.
|
||||
|
||||
### Per-phase REQ coverage (v0.9)
|
||||
|
||||
@@ -252,7 +249,53 @@ HCL-canonical, single-namespace, no-container-runtime, no-SPIFFE). The
|
||||
reversals are justified by the six-part evidence basis recorded in the
|
||||
PROJECT.md Supersession Table.
|
||||
|
||||
## Milestone v0.10: Production Hardening
|
||||
## Milestone v0.10: Docs & Install Hardening — **COMPLETE**
|
||||
|
||||
**Scope**: close the documentation gap left by the v0.9 re-architecture
|
||||
and fix the release/install pipeline bug that caused `install.sh` to
|
||||
resolve to v0.4.5 instead of the latest release. The v0.9
|
||||
re-architecture shipped a complete CLI surface (markdown jobspec,
|
||||
`orca ns`, `orca node capacity`, CLI-side scheduler, emitters, Traefik
|
||||
ingress) but no operator-facing reference documentation. This milestone
|
||||
ships that documentation plus a worked full-stack example with ingress
|
||||
configured, and hardens the release pipeline so every Gitea release
|
||||
carries a Linux binary asset.
|
||||
|
||||
**Milestone type**: feature (P1 ships `fix` phases; P2/P3/P4 ship `docs`
|
||||
phases; at least one non-docs phase makes this a feature milestone per
|
||||
the versioning logic).
|
||||
|
||||
- [x] Phase 0: Pre-execution (specify → clarify → research → ideate → plan → grill) — tag `v0.9.0`
|
||||
- [x] Phase P1: release.sh + install.sh fix (REQ-097, REQ-098) — tag `v0.9.1`
|
||||
- [x] Phase P2: docs/cli.md + docs/jobspec.md + docs/ingress.md (REQ-091, REQ-092, REQ-093) — tag `v0.9.2`
|
||||
- [x] Phase P3: examples/full-stack/ (REQ-094) — tag `v0.9.3`
|
||||
- [x] Phase P4: README.md + docs/namespace.md refresh (REQ-095, REQ-096) — tag `v0.9.4`
|
||||
- [x] Phase P5: Final review + ship + audit (milestone release) — tag `v0.9.5` = v0.10.0 milestone release
|
||||
|
||||
**Milestone tag**: `v0.9.5` (final phase patch = milestone release per
|
||||
feature-milestone progressive-patch rule). Per-phase tags: `v0.9.0`…`v0.9.5`.
|
||||
Tags run on the previous minor's patch line (v0.9.x) per
|
||||
branch-strategy.md. The milestone branch label uses the milestone
|
||||
number (`milestone/v0.10-docs-cli-examples`); no separate minor tag.
|
||||
|
||||
### Per-phase REQ coverage (v0.10 docs milestone)
|
||||
|
||||
- **P1** — release.sh cross-build + asset verification (REQ-097); install.sh fallback walk (REQ-098)
|
||||
- **P2** — CLI reference (REQ-091); jobspec reference (REQ-092); ingress guide (REQ-093)
|
||||
- **P3** — full-stack examples (REQ-094)
|
||||
- **P4** — README refresh (REQ-095); namespace.md v0.9 layout (REQ-096)
|
||||
|
||||
### Root cause of the v0.4.5 install (documented in RESEARCH_v0.10.md)
|
||||
|
||||
The v0.8.x releases (v0.8.0–v0.8.15) shipped with zero binary assets
|
||||
attached to their Gitea releases. `install.sh` resolves "latest" →
|
||||
v0.8.15, looks for `orca-v0.8.15-linux-amd64.tar.gz`, finds nothing, and
|
||||
errors out. The v0.4.5 install came from an earlier run or a pinned
|
||||
`--version`. The fix is forward: release.sh cross-builds amd64 and
|
||||
verifies the asset post-create; install.sh walks backward through
|
||||
releases if the latest lacks the asset.
|
||||
|
||||
## Milestone v0.11: Production Hardening
|
||||
|
||||
**Scope**: ship a cluster that operators can run. Builds on the v0.9
|
||||
re-architecture foundation with the production-grade subsystems:
|
||||
@@ -261,36 +304,36 @@ the v0.8→v1.0 migration.
|
||||
|
||||
**Milestone type**: feature (multiple `feat` phases).
|
||||
|
||||
- [ ] Phase 0: Pre-execution (specify → clarify → research → plan → grill) — tag `v0.9.0`
|
||||
- [ ] Phase P00: CLI cache layer (REQ-062 cache floor; R-008) — tag `v0.9.1`
|
||||
- [ ] Phase P01: Metrics endpoint (hand-rolled text exposition) — tag `v0.9.2`
|
||||
- [ ] Phase P01.5: SPIFFE SVID minting spike (REQ-076; **gate C-08** — if spike fails, fall back to mTLS identity) — tag `v0.9.3`
|
||||
- [ ] Phase P02: ACL (SPIFFE + token identities) — tag `v0.9.4`
|
||||
- [ ] Phase P03: Secrets subsystem (REQ-080; **gate C-19** threat model) — tag `v0.9.5`
|
||||
- [ ] Phase P04: Backup/restore (tar + signed) — tag `v0.9.6`
|
||||
- [ ] Phase P05: Drain + daemon drain-and-stop (REQ-061) — tag `v0.9.7`
|
||||
- [ ] Phase P06: Alloc history (CLI-side SQLite retention; REQ-071 cache DB) — tag `v0.9.8`
|
||||
- [ ] Phase P07: Recovery (`orca restore`) — tag `v0.9.9`
|
||||
- [ ] Phase P08: Integration tests — expand hermetic harness (REQ-087) — tag `v0.9.10`
|
||||
- [ ] Phase P09: Collector + aggregator (opt-in; **gates C-11, C-12, C-14**) — tag `v0.9.11`
|
||||
- [ ] Phase P10: Transactional plane (REQ-075, REQ-079; **gate C-09** orca-pull.sh failure contract) — tag `v0.9.12`
|
||||
- [ ] Phase P11: `orca job lint` (REQ-084) — tag `v0.9.13`
|
||||
- [ ] Phase P12: `orca job verify` (dry-run txn through lead) — tag `v0.9.14`
|
||||
- [ ] Phase P13: `orca ns` subcommands (full surface) + deprecation warnings (REQ-068) — tag `v0.9.15`
|
||||
- [ ] Phase P14a: v0.8→v1.0 data migration (REQ-066; **gate C-07** CA migration spec) — tag `v0.9.16`
|
||||
- [ ] Phase P14b: Daemon cutover + running-allocation adoption — tag `v0.9.17`
|
||||
- [ ] Phase P14c: Mixed-version tolerance + no-orca-on-server enforcement (REQ-065, REQ-086; implements C-13) — tag `v0.9.18`
|
||||
- [ ] Phase P15: README quickstart (REQ-089) — tag `v0.9.19`
|
||||
- [ ] Phase P15.5: Threat model + security review (**gate C-19**) — tag `v0.9.20`
|
||||
- [ ] Phase P16: Final review + ship + audit — **v0.10.0 milestone release** — tag `v0.9.21` (v1.0.0 cut separately after UAT sign-off)
|
||||
- [ ] Phase 0: Pre-execution (specify → clarify → research → plan → grill) — tag `v0.10.0`
|
||||
- [ ] Phase P00: CLI cache layer (REQ-062 cache floor; R-008) — tag `v0.10.1`
|
||||
- [ ] Phase P01: Metrics endpoint (hand-rolled text exposition) — tag `v0.10.2`
|
||||
- [ ] Phase P01.5: SPIFFE SVID minting spike (REQ-076; **gate C-08** — if spike fails, fall back to mTLS identity) — tag `v0.10.3`
|
||||
- [ ] Phase P02: ACL (SPIFFE + token identities) — tag `v0.10.4`
|
||||
- [ ] Phase P03: Secrets subsystem (REQ-080; **gate C-19** threat model) — tag `v0.10.5`
|
||||
- [ ] Phase P04: Backup/restore (tar + signed) — tag `v0.10.6`
|
||||
- [ ] Phase P05: Drain + daemon drain-and-stop (REQ-061) — tag `v0.10.7`
|
||||
- [ ] Phase P06: Alloc history (CLI-side SQLite retention; REQ-071 cache DB) — tag `v0.10.8`
|
||||
- [ ] Phase P07: Recovery (`orca restore`) — tag `v0.10.9`
|
||||
- [ ] Phase P08: Integration tests — expand hermetic harness (REQ-087) — tag `v0.10.10`
|
||||
- [ ] Phase P09: Collector + aggregator (opt-in; **gates C-11, C-12, C-14**) — tag `v0.10.11`
|
||||
- [ ] Phase P10: Transactional plane (REQ-075, REQ-079; **gate C-09** orca-pull.sh failure contract) — tag `v0.10.12`
|
||||
- [ ] Phase P11: `orca job lint` (REQ-084) — tag `v0.10.13`
|
||||
- [ ] Phase P12: `orca job verify` (dry-run txn through lead) — tag `v0.10.14`
|
||||
- [ ] Phase P13: `orca ns` subcommands (full surface) + deprecation warnings (REQ-068) — tag `v0.10.15`
|
||||
- [ ] Phase P14a: v0.8→v1.0 data migration (REQ-066; **gate C-07** CA migration spec) — tag `v0.10.16`
|
||||
- [ ] Phase P14b: Daemon cutover + running-allocation adoption — tag `v0.10.17`
|
||||
- [ ] Phase P14c: Mixed-version tolerance + no-orca-on-server enforcement (REQ-065, REQ-086; implements C-13) — tag `v0.10.18`
|
||||
- [ ] Phase P15: README quickstart (REQ-089) — tag `v0.10.19`
|
||||
- [ ] Phase P15.5: Threat model + security review (**gate C-19**) — tag `v0.10.20`
|
||||
- [ ] Phase P16: Final review + ship + audit — **v0.11.0 milestone release** — tag `v0.10.21` (v1.0.0 cut separately after UAT sign-off)
|
||||
|
||||
**Milestone tag**: `v0.10.0` (the v0.10 milestone release tag; v1.0.0 is
|
||||
UAT-gated and cut separately after v0.10 completion per operator decision —
|
||||
**Milestone tag**: `v0.11.0` (the v0.11 milestone release tag; v1.0.0 is
|
||||
UAT-gated and cut separately after v0.11 completion per operator decision —
|
||||
the v1.0.0 tag marks production-ready sign-off, not a separate milestone).
|
||||
Per-phase patches run on the v0.9.x line per branch-strategy.md. Per-phase
|
||||
tags: `v0.9.0`…`v0.9.21`.
|
||||
Per-phase patches run on the v0.10.x line per branch-strategy.md. Per-phase
|
||||
tags: `v0.10.0`…`v0.10.21`.
|
||||
|
||||
### Per-phase REQ coverage (v0.10)
|
||||
### Per-phase REQ coverage (v0.11)
|
||||
|
||||
- **P00** — CLI cache (R-008)
|
||||
- **P01.5** — SPIFFE spike (REQ-076; C-08)
|
||||
@@ -312,9 +355,9 @@ tags: `v0.9.0`…`v0.9.21`.
|
||||
- **wasmtime CGO breaks cross-compile** (mitigation: C-01 spike; fallback to podman/process primary)
|
||||
- **bash control plane drift** (mitigation: C-15..C-18 render-format contract + bats gate)
|
||||
- **daemon cutover orphans running allocs** (mitigation: P14b split; test adoption)
|
||||
- **27→35+ phase scope** (mitigation: C-04 resolved — operator decision: keep 2 milestones v0.9 + v0.10, keep all phases, v1.0 is UAT-gated after v0.10; current count v0.9=18 + v0.10=22 = 40 phases, exceeds 35 soft limit but operator accepted)
|
||||
- **27→35+ phase scope** (mitigation: C-04 resolved — operator decision: keep 2 milestones v0.9 + v0.11, keep all phases, v1.0 is UAT-gated after v0.11; current count v0.9=18 + v0.11=22 = 40 phases, exceeds 35 soft limit but operator accepted)
|
||||
|
||||
## Deferred to v1.x (out of scope for v0.10)
|
||||
## Deferred to v1.x (out of scope for v0.11)
|
||||
|
||||
- `sqlite-wal-shared` state backend (R-009 abstractions ship in v1.0; backend in v1.x)
|
||||
- `git` state backend
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
"slug": "orca",
|
||||
"name": "Orca",
|
||||
"description": "Offline/CLI-first orchestration engine (Orca) — Nomad-inspired, far simpler than Kubernetes",
|
||||
"milestone": "v0.9",
|
||||
"milestone": "v0.10",
|
||||
"phase": 0,
|
||||
"milestone_type": "feature",
|
||||
"default_branch": "main",
|
||||
|
||||
@@ -4,7 +4,9 @@ Offline/CLI-first orchestration engine inspired by HashiCorp Nomad, far simpler
|
||||
|
||||
## Status
|
||||
|
||||
**v0.1: Foundation** — see [.ciagent/ROADMAP.md](.ciagent/ROADMAP.md) for the 6-phase plan.
|
||||
**v0.9: Re-architecture Foundation — COMPLETE** | **v0.10: Docs & Install Hardening — IN PROGRESS**
|
||||
|
||||
See [.ciagent/ROADMAP.md](.ciagent/ROADMAP.md) for the full roadmap.
|
||||
|
||||
## Pillars
|
||||
|
||||
@@ -28,7 +30,10 @@ curl -fsSL https://git.cloudinit.dev/coreci/orca/raw/branch/main/scripts/install
|
||||
curl -fsSL https://git.cloudinit.dev/coreci/orca/raw/branch/main/scripts/install.sh | sudo bash -s -- --system
|
||||
|
||||
# Pin a specific version
|
||||
curl -fsSL https://git.cloudinit.dev/coreci/orca/raw/branch/main/scripts/install.sh | bash -s -- --version v0.4.2
|
||||
curl -fsSL https://git.cloudinit.dev/coreci/orca/raw/branch/main/scripts/install.sh | bash -s -- --version v0.9.1
|
||||
|
||||
# Dry-run: check what would be installed without writing
|
||||
curl -fsSL https://git.cloudinit.dev/coreci/orca/raw/branch/main/scripts/install.sh | bash -s -- --check
|
||||
```
|
||||
|
||||
Then initialize local state and verify:
|
||||
@@ -54,27 +59,67 @@ config, database, and certificates in the namespace dir:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://git.cloudinit.dev/coreci/orca/raw/branch/main/scripts/install.sh | bash
|
||||
# → "updated orca from v0.4.1 to v0.4.2"
|
||||
# → "updated orca from v0.8.15 to v0.9.1"
|
||||
```
|
||||
|
||||
## Subcommands
|
||||
|
||||
| Command | Description | Status |
|
||||
|---------|-------------|--------|
|
||||
| `orca version` | Print version info | ✅ Phase 1 |
|
||||
| `orca init` | Initialize local orca state | ✅ Phase 1 (stub) |
|
||||
| `orca status` | Show orca daemon status | ✅ Phase 1 (stub) |
|
||||
| `orca node` | Node management (`join`, `leave`, `list`) | Phase 2 |
|
||||
| `orca job` | Job management (`run`, `list`, `stop`, `logs`) | Phase 3 |
|
||||
| Command | Description | Since |
|
||||
|---------|-------------|-------|
|
||||
| `orca init` | Initialize local orca state (full bootstrap) | v0.6 |
|
||||
| `orca version` | Print version info | v0.1 |
|
||||
| `orca status` | Show orca daemon status (**deprecated** v0.9) | v0.1 |
|
||||
| `orca job run` | Run a job from a spec file (`.md`/`.yaml`/`.hcl`) | v0.1 |
|
||||
| `orca job list` | List all jobs (`--watch` for streaming) | v0.1 |
|
||||
| `orca job stop` | Stop a running job | v0.1 |
|
||||
| `orca job logs` | Show task output for a job | v0.1 |
|
||||
| `orca node join` | Join a node (`--type proxmox` for SSH-push) | v0.2 |
|
||||
| `orca node leave` | Remove a node from the registry | v0.2 |
|
||||
| `orca node list` | List all nodes (`--watch` for streaming) | v0.2 |
|
||||
| `orca node key-reset` | Reset SSH known_hosts entry for a node | v0.8 |
|
||||
| `orca node capacity` | Manage node capacity (show/set/list) | v0.2 |
|
||||
| `orca ns list` | List all namespaces | v0.9 |
|
||||
| `orca ns create` | Create a namespace directory + ns.md | v0.9 |
|
||||
| `orca ns delete` | Remove an empty namespace | v0.9 |
|
||||
| `orca ns inspect` | Print effective chain, merged env, constraints | v0.9 |
|
||||
| `orca ns validate` | Run cycle + missing-parent + schema checks | v0.9 |
|
||||
| `orca doctor` | Run self-checks (cert/network/db/os/proxmox) | v0.2 |
|
||||
| `orca audit list` | View audit log entries | v0.1 |
|
||||
| `orca daemon` | Run the daemon (**deprecated** v0.9) | v0.1 |
|
||||
| `orca cert` | Manage certificates (**deprecated** v0.9) | v0.2 |
|
||||
|
||||
See [docs/cli.md](docs/cli.md) for the full CLI reference with all flags and examples.
|
||||
|
||||
## Documentation
|
||||
|
||||
| Document | Description |
|
||||
|----------|-------------|
|
||||
| [docs/cli.md](docs/cli.md) | CLI reference — every command, flag, and example |
|
||||
| [docs/jobspec.md](docs/jobspec.md) | Jobspec reference — markdown frontmatter schema |
|
||||
| [docs/ingress.md](docs/ingress.md) | Ingress guide — Traefik configuration |
|
||||
| [docs/namespace.md](docs/namespace.md) | Namespace and path layout |
|
||||
| [docs/install.md](docs/install.md) | Installation guide |
|
||||
| [docs/docker.md](docs/docker.md) | Docker image guide |
|
||||
| [docs/security-scanning.md](docs/security-scanning.md) | Security scanning tools |
|
||||
|
||||
## Examples
|
||||
|
||||
| Example | Description |
|
||||
|---------|-------------|
|
||||
| [examples/full-stack/](examples/full-stack/) | Full-stack deployment with ingress (5 services + rendered artifacts) |
|
||||
|
||||
## Development
|
||||
|
||||
```bash
|
||||
make build # Build binary to ./bin/orca
|
||||
make test # Run tests with race detection
|
||||
make lint # Run golangci-lint
|
||||
make fmt # Format code
|
||||
make release # Build + create Gitea release (Phase 6)
|
||||
make build # Build binary to ./bin/orca
|
||||
make test # Run tests
|
||||
make test-race # Run tests with race detection
|
||||
make lint # Run gofmt + go vet + shellcheck
|
||||
make fmt # Format code
|
||||
make security-scan # Run gosec + govulncheck + gitleaks
|
||||
make verify-reqs # Assert ROADMAP ↔ REQUIREMENTS consistency
|
||||
make changelog # Generate CHANGELOG.md from ---ci--- blocks
|
||||
make release # Build + create Gitea release (VERSION required)
|
||||
```
|
||||
|
||||
## Architecture
|
||||
@@ -83,4 +128,4 @@ See [.ciagent/ARCHITECTURE.md](.ciagent/ARCHITECTURE.md) for full architecture d
|
||||
|
||||
## License
|
||||
|
||||
MIT — see [LICENSE](LICENSE).
|
||||
MIT — see [LICENSE](LICENSE).
|
||||
+522
@@ -0,0 +1,522 @@
|
||||
# Orca CLI Reference
|
||||
|
||||
This document is the complete reference for the `orca` command-line
|
||||
interface. Every command, subcommand, and flag is documented here.
|
||||
|
||||
> **Canonical path (v0.9)**: The v0.9 re-architecture introduced the
|
||||
> SSH-push deployment model, markdown jobspec, multi-namespace layout,
|
||||
> and CLI-side scheduler. Commands marked **deprecated** below are from
|
||||
> the v0.8 daemon/mTLS model and will be removed in v0.11. Use the
|
||||
> v0.9 canonical path for all new work.
|
||||
|
||||
## Global flags
|
||||
|
||||
These flags are available on every `orca` command.
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--json` | bool | `false` | Output in JSON format (machine-readable) |
|
||||
| `--system` | bool | `false` | Use system-level namespace root (`/root/.orca`) instead of user-level (`~/.orca`). Errors if `ORCA_HOME` is already set to a conflicting value. |
|
||||
| `--config` | string | `""` | Path to config file (overrides `~/.orca/config.hcl`). Supports `.hcl` (legacy) and `.md` (v0.9 canonical) formats. |
|
||||
| `--no-deprecation-warnings` | bool | `false` | Suppress v0.9 deprecation warnings. Use during `orca upgrade` migrations. |
|
||||
|
||||
### Output modes
|
||||
|
||||
- **Text** (default): human-readable tables and messages.
|
||||
- **JSON** (`--json`): structured JSON output for machine consumption
|
||||
and AI agents.
|
||||
- **Watch** (`--watch` on list commands): table refresh (text default)
|
||||
or NDJSON streaming (`--json`), one line per event until Ctrl-C.
|
||||
|
||||
### Environment variables
|
||||
|
||||
| Variable | Description |
|
||||
|----------|-------------|
|
||||
| `ORCA_HOME` | Namespace root directory (default `~/.orca`). Overrides all on-disk paths. |
|
||||
| `ORCA_DB` | Fine-grained database path override. |
|
||||
| `ORCA_PROXMOX_PASSWORD` | SSH password for `orca node join --type proxmox` (never persisted). |
|
||||
| `ORCA_LISTEN_ADDR` | Daemon listen address (deprecated). |
|
||||
| `ORCA_CA_PATH` | CA certificate path override. |
|
||||
| `ORCA_SERVER_CERT_PATH` | Server certificate path override. |
|
||||
| `ORCA_SERVER_KEY_PATH` | Server key path override. |
|
||||
| `ORCA_NODE_CPU` | Node CPU capacity override (millicores). |
|
||||
| `ORCA_NODE_MEMORY_MB` | Node memory capacity override (MiB). |
|
||||
|
||||
### Exit codes
|
||||
|
||||
| Code | Meaning |
|
||||
|------|---------|
|
||||
| `0` | Success |
|
||||
| `1` | Error (printed to stderr) |
|
||||
|
||||
---
|
||||
|
||||
## `orca init`
|
||||
|
||||
Initialize local orca state with full bootstrap.
|
||||
|
||||
```
|
||||
orca init
|
||||
```
|
||||
|
||||
Performs a 6-step idempotent bootstrap:
|
||||
|
||||
1. Create the namespace directory (honors `$ORCA_HOME`; defaults to `~/.orca`)
|
||||
2. Open and migrate the SQLite database (migrations 0001–0006)
|
||||
3. Bootstrap the internal CA (`ca.crt` + `ca.key`) if not already present
|
||||
4. Generate the server cert (`server.crt` + `server.key`) if not already present
|
||||
5. Auto-detect the local OS via `/etc/os-release`
|
||||
6. Register a localhost node (kind=localhost, os=\<detected\>)
|
||||
|
||||
Re-running `orca init` is safe — it refreshes `last_seen` and `os` on
|
||||
the localhost node without regenerating certs or changing the node ID.
|
||||
|
||||
**Flags**: none.
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
orca init
|
||||
orca --system init # system-level bootstrap at /root/.orca
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `orca job`
|
||||
|
||||
Manage orca jobs — run, list, stop, and inspect.
|
||||
|
||||
### `orca job run`
|
||||
|
||||
Run a job from a spec file.
|
||||
|
||||
```
|
||||
orca job run <spec> [flags]
|
||||
```
|
||||
|
||||
Dispatches by file extension:
|
||||
- `.md` → Markdown frontmatter parser (v0.9 canonical)
|
||||
- `.yaml` / `.yml` → YAML frontmatter parser
|
||||
- `.hcl` → Legacy HCL adapter (deprecated, see callout below)
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--target` | string | `""` | Pin job to a specific node ID (overrides bin-packing scheduler) |
|
||||
| `--idempotency-key` | string | `""` | Idempotency key for cross-node dispatch dedupe |
|
||||
|
||||
**Examples**:
|
||||
```bash
|
||||
orca job run web-app.md
|
||||
orca job run api.yaml --target node-abc-123
|
||||
orca job run worker.md --idempotency-key deploy-2026-08-05
|
||||
```
|
||||
|
||||
> **Deprecated**: `orca job run <spec.hcl>` (legacy HCL jobspec) still
|
||||
> works via the adapter but emits a deprecation warning. Migrate `.hcl`
|
||||
> specs to `.md` (see [docs/jobspec.md](jobspec.md)). Removed in v0.11.
|
||||
|
||||
### `orca job list`
|
||||
|
||||
List all jobs.
|
||||
|
||||
```
|
||||
orca job list [flags]
|
||||
```
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--watch` | bool | `false` | Stream jobs until Ctrl-C (table refresh or `--json` per-event) |
|
||||
|
||||
**Output columns**: `ID NAME STATUS EXIT`
|
||||
|
||||
**Examples**:
|
||||
```bash
|
||||
orca job list
|
||||
orca job list --watch # table refresh
|
||||
orca job list --watch --json # NDJSON: {"event":"update","job":{...}}
|
||||
```
|
||||
|
||||
### `orca job stop`
|
||||
|
||||
Stop a running job (soft stop).
|
||||
|
||||
```
|
||||
orca job stop [job-id] [flags]
|
||||
```
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--id` | string | `""` | Job ID (alternative to positional argument) |
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
orca job stop abc-123-def
|
||||
orca job stop --id abc-123-def
|
||||
```
|
||||
|
||||
### `orca job logs`
|
||||
|
||||
Show task output for a job.
|
||||
|
||||
```
|
||||
orca job logs [job-id] [flags]
|
||||
```
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--id` | string | `""` | Job ID (alternative to positional argument) |
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
orca job logs abc-123-def
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `orca node`
|
||||
|
||||
Manage orca nodes — join, leave, or list nodes in the registry.
|
||||
|
||||
### `orca node join`
|
||||
|
||||
Join a node to the orca registry.
|
||||
|
||||
```
|
||||
orca node join [flags]
|
||||
```
|
||||
|
||||
Node types (via `--type`):
|
||||
- `localhost` (default): register a local or Linux node
|
||||
- `proxmox`: SSH-bootstrap a remote Proxmox VE 8/9 host (deploys orca
|
||||
pubkey, creates orca user + PVE role + sudoers allowlist; requires
|
||||
`--host` + `--password`)
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--name` | string | `""` | Node name (required for `--type localhost`) |
|
||||
| `--addr` | string | `""` | Node address (default `localhost:8443`) |
|
||||
| `--ca-fingerprint` | string | `""` | Pin CA cert SHA-256 (fails if on-disk CA doesn't match) |
|
||||
| `--type` | string | `"localhost"` | Node type: `localhost` or `proxmox` |
|
||||
| `--host` | string | `""` | Proxmox host address (IP/hostname; required for `--type proxmox`) |
|
||||
| `--ssh-user` | string | `"root"` | SSH username for proxmox bootstrap |
|
||||
| `--password` | string | `""` | SSH password for proxmox bootstrap (never persisted; prefer `$ORCA_PROXMOX_PASSWORD`) |
|
||||
| `--ssh-port` | int | `22` | SSH port for proxmox bootstrap |
|
||||
| `--proxmox-user` | string | `"orca"` | Linux system user to create on the proxmox host |
|
||||
| `--proxmox-role` | string | `"OrcaOperator"` | PVE custom role to create |
|
||||
| `--host-key-fingerprint` | string | `""` | SSH host key `SHA256:base64` fingerprint (pre-pin; supersedes TOFU for `--type proxmox`) |
|
||||
|
||||
**Examples**:
|
||||
```bash
|
||||
# Localhost (deprecated mTLS path)
|
||||
orca node join --name my-node
|
||||
|
||||
# Proxmox (v0.9 canonical SSH-push path)
|
||||
orca node join --type proxmox --host 192.168.1.100 --ssh-user root
|
||||
ORCA_PROXMOX_PASSWORD=secret orca node join --type proxmox --host 192.168.1.100
|
||||
|
||||
# Proxmox with pre-pinned host key
|
||||
orca node join --type proxmox --host 192.168.1.100 --host-key-fingerprint SHA256:abc123...
|
||||
```
|
||||
|
||||
> **Deprecated**: `orca node join` without `--type proxmox` (the
|
||||
> localhost mTLS join path) is deprecated in v0.9. The v0.9 canonical
|
||||
> path is SSH-push (`--type proxmox`) or local execution (no join
|
||||
> needed). Removed in v0.11.
|
||||
|
||||
### `orca node leave`
|
||||
|
||||
Remove a node from the orca registry.
|
||||
|
||||
```
|
||||
orca node leave [node-id] [flags]
|
||||
```
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--id` | string | `""` | Node ID (alternative to positional argument) |
|
||||
|
||||
### `orca node list`
|
||||
|
||||
List all nodes in the orca registry.
|
||||
|
||||
```
|
||||
orca node list [flags]
|
||||
```
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--watch` | bool | `false` | Stream nodes until Ctrl-C (table refresh or `--json` per-event) |
|
||||
|
||||
**Output columns**: `ID NAME ADDRESS STATE`
|
||||
|
||||
### `orca node key-reset`
|
||||
|
||||
Reset the SSH known_hosts entry for a node.
|
||||
|
||||
```
|
||||
orca node key-reset <node>
|
||||
```
|
||||
|
||||
Removes the pinned SSH host key for `<node>` from the local
|
||||
`known_hosts` file. The next connect re-pins the key via TOFU or
|
||||
`--host-key-fingerprint`. Local only — does not touch the remote
|
||||
host's `authorized_keys`.
|
||||
|
||||
`<node>` is the node name (for proxmox nodes, this is the host address).
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
orca node key-reset 192.168.1.100
|
||||
```
|
||||
|
||||
### `orca node capacity`
|
||||
|
||||
Manage node capacity declarations (bin-packing scheduler input).
|
||||
|
||||
```
|
||||
orca node capacity <subcommand>
|
||||
```
|
||||
|
||||
#### `orca node capacity show`
|
||||
|
||||
Show capacity for a node (defaults to `self`).
|
||||
|
||||
```
|
||||
orca node capacity show [node-id] [flags]
|
||||
```
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--node` | string | `""` | Node ID (defaults to `self`) |
|
||||
|
||||
**Output**: `Node:`, `CPU:` (millicores), `Memory:` (MiB), `Disk:` (MiB), `Updated:`
|
||||
|
||||
#### `orca node capacity set`
|
||||
|
||||
Declare capacity for a node.
|
||||
|
||||
```
|
||||
orca node capacity set [flags]
|
||||
```
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--cpu` | int64 | `0` | CPU capacity in millicores (1000 = 1 vCPU) |
|
||||
| `--memory` | int64 | `0` | Memory capacity in MiB |
|
||||
| `--disk` | int64 | `0` | Disk capacity in MiB |
|
||||
| `--node` | string | `""` | Node ID (defaults to `self`) |
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
orca node capacity set --cpu 4000 --memory 8192 --disk 100000
|
||||
orca node capacity set --cpu 2000 --memory 4096 --node web-1
|
||||
```
|
||||
|
||||
#### `orca node capacity list`
|
||||
|
||||
List all node capacity declarations.
|
||||
|
||||
```
|
||||
orca node capacity list
|
||||
```
|
||||
|
||||
**Output columns**: `NODE CPU(mc) MEM(MiB) DISK(MiB) UPDATED`
|
||||
|
||||
---
|
||||
|
||||
## `orca ns`
|
||||
|
||||
Manage orca namespaces under `ORCA_HOME` (R-002).
|
||||
|
||||
Each namespace is a directory with `ns.md`, `.env`, `.env.secrets`,
|
||||
`db/`, `jobs/`, `alloc/`. The implicit root namespace `_defaults`
|
||||
always exists; every namespace inherits from `_defaults` and cannot
|
||||
opt out.
|
||||
|
||||
### `orca ns list`
|
||||
|
||||
List all namespaces under `ORCA_HOME`.
|
||||
|
||||
```
|
||||
orca ns list
|
||||
```
|
||||
|
||||
**Output columns**: `NAME DEFAULT PATH` (`_defaults` marked `*`)
|
||||
|
||||
### `orca ns create`
|
||||
|
||||
Create a namespace directory + `ns.md`.
|
||||
|
||||
```
|
||||
orca ns create <name> [flags]
|
||||
```
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--parent` | string | `""` | Parent namespace (default `_defaults`; implicit root always appended last) |
|
||||
| `--inherits-env` | bool | `true` | Inherit env from parents |
|
||||
| `--inherits-secrets` | bool | `true` | Inherit secrets from parents |
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
orca ns create prod --parent _defaults
|
||||
orca ns create staging --parent prod
|
||||
```
|
||||
|
||||
### `orca ns delete`
|
||||
|
||||
Remove an empty namespace directory.
|
||||
|
||||
```
|
||||
orca ns delete <name>
|
||||
```
|
||||
|
||||
Refuses if `jobs/` or `alloc/` contain files. The implicit root
|
||||
`_defaults` cannot be deleted.
|
||||
|
||||
### `orca ns inspect`
|
||||
|
||||
Print the effective inheritance chain, merged env, and constraints.
|
||||
|
||||
```
|
||||
orca ns inspect <name>
|
||||
```
|
||||
|
||||
**Output**: `Namespace:`, `Chain:` (e.g., `prod -> _defaults`), `Env:`
|
||||
(sorted keys), `Constraints:` (unioned CEL expressions).
|
||||
|
||||
### `orca ns validate`
|
||||
|
||||
Run cycle + missing-parent + schema checks on a namespace.
|
||||
|
||||
```
|
||||
orca ns validate <name>
|
||||
```
|
||||
|
||||
Exits 0 if valid, 1 on error. Runs over ALL namespaces under
|
||||
`ORCA_HOME` (parsing + resolving validates cycles and missing parents
|
||||
across the set).
|
||||
|
||||
---
|
||||
|
||||
## `orca doctor`
|
||||
|
||||
Run self-checks on the orca installation.
|
||||
|
||||
```
|
||||
orca doctor [subcommand]
|
||||
```
|
||||
|
||||
Without a subcommand, runs all checks and prints a PASS/WARN/FAIL
|
||||
report per check.
|
||||
|
||||
### Subcommands
|
||||
|
||||
| Command | Description |
|
||||
|---------|-------------|
|
||||
| `orca doctor cert` | CA, server cert, expiry, fingerprint checks |
|
||||
| `orca doctor network` | Network reachability via mTLS `/healthz` probe |
|
||||
| `orca doctor db` | Database integrity (`PRAGMA integrity_check` + migration version) |
|
||||
| `orca doctor os` | OS detection self-check (verifies `/etc/os-release` matches stored node) |
|
||||
| `orca doctor proxmox` | Proxmox node reachability via SSH `pveversion`/`pvecmd status` probe |
|
||||
|
||||
**Example**:
|
||||
```bash
|
||||
orca doctor
|
||||
orca doctor cert
|
||||
orca doctor proxmox --json
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `orca audit`
|
||||
|
||||
View orca audit log (security-first observability).
|
||||
|
||||
### `orca audit list`
|
||||
|
||||
List recent audit log entries.
|
||||
|
||||
```
|
||||
orca audit list [flags]
|
||||
```
|
||||
|
||||
| Flag | Type | Default | Description |
|
||||
|------|------|---------|-------------|
|
||||
| `--limit` | int | `50` | Max entries to show |
|
||||
|
||||
**Output columns**: `TIMESTAMP ACTOR ACTION RESOURCE RESULT`
|
||||
|
||||
---
|
||||
|
||||
## `orca version`
|
||||
|
||||
Print version information.
|
||||
|
||||
```
|
||||
orca version
|
||||
```
|
||||
|
||||
**Output**:
|
||||
```
|
||||
orca version v0.9.1
|
||||
git commit: abc1234
|
||||
build time: 2026-08-05T20:30:00Z
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## `orca status`
|
||||
|
||||
Show orca daemon status.
|
||||
|
||||
```
|
||||
orca status
|
||||
```
|
||||
|
||||
> **Deprecated**: The daemon model is deprecated in v0.9 (replaced by
|
||||
> SSH-push, R-001). This command returns a stub status. Removed in
|
||||
> v0.11.
|
||||
|
||||
---
|
||||
|
||||
## Deprecated commands
|
||||
|
||||
The following commands are from the v0.8 daemon/mTLS model and are
|
||||
**deprecated in v0.9**. They still work during the dual-write window
|
||||
but emit `slog.Warn` deprecation warnings. They will be **removed in
|
||||
v0.11**.
|
||||
|
||||
> **`orca daemon`** — Run the orca daemon (HTTP API + health checks).
|
||||
> The v0.9 re-architecture replaces the daemon with SSH-push (R-001).
|
||||
> The daemon is repurposed to `drain-and-stop` in v0.11-P05 and deleted
|
||||
> in v0.11-P14. Flags: `--addr` (default `:8080`), `--pprof` (pprof
|
||||
> endpoint, default disabled).
|
||||
|
||||
> **`orca cert`** — Manage orca certificates (CA, server, rotation).
|
||||
> The v0.9 re-architecture replaces the internal CA with step-ca
|
||||
> (D-101). Subcommands: `ca-init`, `gen`, `show`, `renew`,
|
||||
> `fingerprint`. Removed in v0.11.
|
||||
|
||||
> **`orca node join` (mTLS path)** — The localhost mTLS join path
|
||||
> (without `--type proxmox`) is deprecated. The v0.9 canonical path is
|
||||
> SSH-push (`--type proxmox`) or local execution (no join needed).
|
||||
|
||||
> **`orca job run <spec.hcl>`** — Legacy HCL jobspec. Migrate to `.md`
|
||||
> (see [docs/jobspec.md](jobspec.md)). The HCL adapter preserves
|
||||
> `orca job run old-spec.hcl` during the migration window.
|
||||
|
||||
To suppress deprecation warnings during migration, use
|
||||
`--no-deprecation-warnings`:
|
||||
```bash
|
||||
orca --no-deprecation-warnings daemon
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## See also
|
||||
|
||||
- [docs/jobspec.md](jobspec.md) — Markdown frontmatter jobspec reference
|
||||
- [docs/ingress.md](ingress.md) — Traefik ingress configuration guide
|
||||
- [docs/namespace.md](namespace.md) — Namespace and path layout
|
||||
- [docs/install.md](install.md) — Installation guide
|
||||
- [examples/full-stack/](../examples/full-stack/) — Full-stack example with ingress
|
||||
+209
@@ -0,0 +1,209 @@
|
||||
# Orca Ingress & Traefik Guide
|
||||
|
||||
This document explains how Orca configures ingress via Traefik dynamic
|
||||
configuration. It covers the service→Traefik mapping, the R-007
|
||||
socket-vs-TCP-bind model, atomic reload, drain, TLS, and a worked
|
||||
example.
|
||||
|
||||
> **Canonical path (v0.9)**: Orca generates Traefik dynamic
|
||||
> configuration files via the `TraefikEmitter`. The `kind: Service`
|
||||
> workload implies a Traefik route. The emitter renders one YAML file
|
||||
> per Service; Traefik watches the dynamic config directory and reloads
|
||||
> atomically on change.
|
||||
|
||||
## The model
|
||||
|
||||
A `kind: Service` jobspec **implies** a Traefik route (D-175). `Job`
|
||||
and `DaemonSet` do **not** carry a Traefik route by default — a
|
||||
`service:` block on a `Job` is rejected by the validator.
|
||||
|
||||
When `orca job run` submits a `kind: Service` workload, the
|
||||
`TraefikEmitter` renders a Traefik dynamic config file at:
|
||||
|
||||
```
|
||||
/etc/traefik/dynamic/orca-<service-name>.yaml
|
||||
```
|
||||
|
||||
This file contains:
|
||||
- One **router** (`orca-<name>`) with a `PathPrefix` rule and TLS config.
|
||||
- One **service** (`orca-<name>`) as a `loadBalancer` with one **server**
|
||||
per port, pointing at the workload's Unix socket (or TCP port).
|
||||
- A **healthCheck** stanza when the `health:` block is present.
|
||||
|
||||
Traefik watches `/etc/traefik/dynamic/` via `fsnotify` and reloads
|
||||
whenever a file changes. Orca writes config atomically (write-tmp +
|
||||
rename) so Traefik sees a single `IN_MOVED_TO` event and never observes
|
||||
a half-written file.
|
||||
|
||||
## R-007: socket vs TCP bind
|
||||
|
||||
Orca workloads bind to a **Unix socket** by default, not a TCP port.
|
||||
This is the R-007 security model: loopback-only by default, no network
|
||||
exposure.
|
||||
|
||||
### Default: Unix socket
|
||||
|
||||
When `service.bind` is empty (default), the workload binds a Unix
|
||||
socket at:
|
||||
|
||||
```
|
||||
/run/orca/alloc-<alloc-id>/port-<port-name>.sock
|
||||
```
|
||||
|
||||
systemd creates `/run/orca/alloc-<alloc-id>/` via
|
||||
`RuntimeDirectory=orca/alloc-<alloc-id>` (mode 0750, owned by
|
||||
`orca:orca`). The Traefik backend server URL is:
|
||||
|
||||
```yaml
|
||||
servers:
|
||||
- url: "unix:///run/orca/alloc-<alloc-id>/port-<port-name>.sock"
|
||||
```
|
||||
|
||||
### TCP opt-in: `service.bind: 127.0.0.1`
|
||||
|
||||
When `service.bind: 127.0.0.1` is set, the workload binds a TCP port
|
||||
directly (loopback only). The emitter adds an `ExecStartPre` marker to
|
||||
the systemd unit so the bind mode is visible:
|
||||
|
||||
```ini
|
||||
ExecStartPre=/bin/echo orca: bind 127.0.0.1 port <name> (tcp, R-007 opt-in)
|
||||
```
|
||||
|
||||
`service.bind` must be a valid IP address. Empty (socket default) or
|
||||
`127.0.0.1` (TCP opt-in) are the documented values; any other valid IP
|
||||
is accepted but the bind happens in the process, not the emitter.
|
||||
|
||||
## Generated Traefik YAML
|
||||
|
||||
For a Service named `web` with port `http`:
|
||||
|
||||
```yaml
|
||||
http:
|
||||
routers:
|
||||
orca-web:
|
||||
rule: PathPrefix("/web")
|
||||
service: orca-web
|
||||
tls:
|
||||
certResolver: orca
|
||||
domains:
|
||||
- main: "cluster.orca.local"
|
||||
services:
|
||||
orca-web:
|
||||
loadBalancer:
|
||||
servers:
|
||||
- url: "unix:///run/orca/alloc-<alloc-id>/port-http.sock"
|
||||
healthCheck:
|
||||
path: /healthz
|
||||
interval: 5s
|
||||
timeout: 1s
|
||||
```
|
||||
|
||||
- One router per Service, named `orca-<service-name>`.
|
||||
- Router rule: `PathPrefix("/<service-name>")`.
|
||||
- TLS: `certResolver: orca`, trust domain `cluster.orca.local`
|
||||
(placeholder; step-ca provisioner overrides in v0.11).
|
||||
- One service per Service, named `orca-<service-name>`.
|
||||
- One server per port, URL is `unix://<socket-path>`.
|
||||
- `healthCheck` stanza present when `health:` block is set (required
|
||||
for Service). Path is `/healthz`; interval and timeout come from the
|
||||
`health:` block.
|
||||
|
||||
## Atomic reload (gate C-10)
|
||||
|
||||
Orca writes Traefik config atomically to avoid Traefik observing a
|
||||
half-written file:
|
||||
|
||||
1. Write to `<path>.tmp` via `WriteFileIdempotent` (write + fsync).
|
||||
2. `mv -f <path>.tmp <path>` (atomic POSIX rename).
|
||||
|
||||
Traefik's `fsnotify` watcher sees a single `IN_MOVED_TO` event and
|
||||
reloads. If the new config is malformed, Traefik logs an error and
|
||||
**holds last-good config** — the cluster keeps serving traffic on the
|
||||
previous config.
|
||||
|
||||
## Drain
|
||||
|
||||
`RenderDrain` produces the same Traefik YAML with `weight: 0` on every
|
||||
server in the load balancer:
|
||||
|
||||
```yaml
|
||||
servers:
|
||||
- url: "unix:///run/orca/alloc-<alloc-id>/port-http.sock"
|
||||
weight: 0
|
||||
```
|
||||
|
||||
Traefik stops sending traffic to the drained backend. The workload
|
||||
keeps running; drain is reversible (re-submit the normal config to
|
||||
restore traffic).
|
||||
|
||||
## TLS
|
||||
|
||||
- **certResolver**: `orca` (references the Traefik ACME/step-ca
|
||||
certificate resolver configured in Traefik's static config).
|
||||
- **Trust domain**: `cluster.orca.local` (placeholder in v0.9; step-ca
|
||||
provisioner in v0.11 overrides with the real cluster trust domain).
|
||||
- **SPIFFE SVIDs**: workload identity via SPIFFE SVIDs minted at submit
|
||||
time via step-ca (v0.11-P01.5, gate C-08). The SVID is a URI SAN in
|
||||
the workload's X.509 cert.
|
||||
|
||||
## Health checks
|
||||
|
||||
The `health:` block (required for `Service`) maps to the Traefik
|
||||
`healthCheck` stanza:
|
||||
|
||||
```yaml
|
||||
health:
|
||||
check_type: http
|
||||
interval: 5s
|
||||
timeout: 1s
|
||||
unhealthy_threshold: 2
|
||||
```
|
||||
|
||||
→
|
||||
|
||||
```yaml
|
||||
healthCheck:
|
||||
path: /healthz
|
||||
interval: 5s
|
||||
timeout: 1s
|
||||
```
|
||||
|
||||
Traefik polls each backend's `/healthz` at the configured interval. An
|
||||
unhealthy backend is removed from the load balancer pool until it
|
||||
passes the health check again.
|
||||
|
||||
## Worked example
|
||||
|
||||
See [examples/full-stack/](../examples/full-stack/) for a complete
|
||||
multi-service stack with ingress configured:
|
||||
- `web-app.md` — frontend Service (socket bind, PathPrefix route)
|
||||
- `api.md` — backend API Service (TCP opt-in, `127.0.0.1` bind)
|
||||
- `examples/full-stack/rendered/traefik-dynamic-web-app.yaml` — the
|
||||
Traefik config Orca generates
|
||||
|
||||
## v0.11 forward (limitations)
|
||||
|
||||
The following are not yet implemented in v0.9 and will land in v0.11:
|
||||
|
||||
- **`service.host` / `service.route_id`**: stored on the `ServiceBlock`
|
||||
but not yet consumed by the `TraefikEmitter`. The router rule is
|
||||
hardcoded `PathPrefix("/<name>")`. Custom host-based routing lands in
|
||||
v0.11.
|
||||
- **Socket activation**: real socket-activation (socket unit files, fd
|
||||
passing) lands in v0.11-P08. The current emitter renders the
|
||||
`RuntimeDirectory` + socket path comments but does not create socket
|
||||
units.
|
||||
- **Transactional update execution**: the `update:` block's rolling/
|
||||
canary/blue-green plan is computed by the emitter but not yet
|
||||
executed transactionally. Transactional execution lands in
|
||||
v0.11-P10.
|
||||
- **SPIFFE SVID minting**: workload identity via step-ca SVIDs lands in
|
||||
v0.11-P01.5 (gate C-08).
|
||||
- **Secrets in env**: `env: { KEY: { from: "secret:..." } }` resolution
|
||||
to `EnvironmentFile=`/`LoadCredential=` lands in v0.11-P03.
|
||||
|
||||
## See also
|
||||
|
||||
- [docs/cli.md](cli.md) — CLI reference
|
||||
- [docs/jobspec.md](jobspec.md) — Jobspec reference (`service:`, `health:`, `ports:` blocks)
|
||||
- [examples/full-stack/](../examples/full-stack/) — Full-stack example with ingress
|
||||
+442
@@ -0,0 +1,442 @@
|
||||
# Orca Jobspec Reference
|
||||
|
||||
This document is the complete reference for the Orca jobspec format —
|
||||
the Markdown-with-frontmatter specification that describes workloads.
|
||||
|
||||
> **Canonical format (v0.9)**: Orca uses Markdown with YAML frontmatter
|
||||
> as the canonical jobspec format (R-013/R-014). The legacy HCL format
|
||||
> is supported via an adapter during the migration window but is
|
||||
> deprecated (see [HCL jobspec](#deprecated-hcl-jobspec) below).
|
||||
|
||||
## File formats
|
||||
|
||||
The `orca job run` command dispatches by file extension:
|
||||
|
||||
| Extension | Parser | Body |
|
||||
|-----------|--------|------|
|
||||
| `.md` | `ParseMarkdown` (canonical) | Verbatim after closing `---` (R-015 byte-exact) |
|
||||
| `.yaml` / `.yml` | `parseYAMLFile` | Whole file as frontmatter; body empty |
|
||||
| `.hcl` | `ParseHCL` (legacy adapter) | Empty (deprecated) |
|
||||
|
||||
## Minimal example
|
||||
|
||||
```yaml
|
||||
---
|
||||
kind: Job
|
||||
name: my-job
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /bin/echo hello
|
||||
---
|
||||
# My Job
|
||||
|
||||
This body is preserved byte-exact and carried to the target node.
|
||||
```
|
||||
|
||||
## Top-level keys
|
||||
|
||||
| Key | Type | Default | Required | Notes |
|
||||
|-----|------|---------|----------|-------|
|
||||
| `orca-spec-version` | string | `""` | no | Free-form version tag (e.g. `"1"`) |
|
||||
| `kind` | enum | — | **yes** | One of `Job`, `Service`, `DaemonSet` |
|
||||
| `name` | string | — | **yes** | Workload name (trimmed, non-empty) |
|
||||
| `count` | int | `1` | no | Job: must be 1; Service: ≥1; DaemonSet: not allowed |
|
||||
| `runtime` | block | nil | see kinds | Runtime block (or per-task runtimes in a task group) |
|
||||
| `ports` | block list | nil | Service: **yes** | Array of port mappings |
|
||||
| `env` | block map | nil | no | Environment variables |
|
||||
| `secrets` | inline/block list | nil | no | Secret names (resolution in v0.11) |
|
||||
| `volumes` | block list | nil | no | Volume mounts |
|
||||
| `restart` | block | nil | Service/DaemonSet: **yes** | Restart policy |
|
||||
| `update` | block | nil | Service: **yes** | Update strategy |
|
||||
| `service` | block | nil | no | Traefik route definition (implied for Service; not allowed for Job/DaemonSet) |
|
||||
| `health` | block | nil | Service: **yes** | Health check |
|
||||
| `lifecycle` | block | nil | no | Pre-stop / post-start hooks |
|
||||
| `constraints` | list | nil | no | CEL expressions (node selection) |
|
||||
| `affinity` | block list | nil | no | Co-location / anti-affinity rules |
|
||||
| `tasks` | block list | nil | no | Task group (multi-process alloc) |
|
||||
| `timeout` | duration string | `""` | no | Job timeout |
|
||||
| `schedule` | block | nil | DaemonSet: **yes** | Schedule mode |
|
||||
|
||||
## Kinds
|
||||
|
||||
### `Job`
|
||||
|
||||
A one-shot batch task. Runs once and exits.
|
||||
|
||||
- `count` must be 1 (or unset). Use `Service` for replicas.
|
||||
- `service` block is **not allowed** (no Traefik route for Jobs).
|
||||
- `restart` optional (defaults to `never` / `on-failure`).
|
||||
- `timeout` optional.
|
||||
|
||||
**Example**:
|
||||
```yaml
|
||||
---
|
||||
kind: Job
|
||||
name: data-migration
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /usr/bin/python3 migrate.py
|
||||
timeout: 300s
|
||||
env:
|
||||
DB_URL: postgres://localhost/mydb
|
||||
---
|
||||
```
|
||||
|
||||
### `Service`
|
||||
|
||||
A long-running, load-balanced workload with a Traefik route.
|
||||
|
||||
- `count` ≥ 1 (number of replicas).
|
||||
- `ports` required (at least one).
|
||||
- `restart` required; `mode` one of `service`, `on-failure`, `never`.
|
||||
- `update` required; `strategy` one of `rolling`, `canary`, `blue-green`.
|
||||
- `runtime` required (or a task group with per-task runtimes).
|
||||
- `health` required (Traefik routing requires health checks).
|
||||
- `service` block optional (implied for Service; use for `bind` override).
|
||||
- `service.bind` if present must be a valid IP (`127.0.0.1` = TCP opt-in;
|
||||
default = Unix socket).
|
||||
|
||||
**Example**:
|
||||
```yaml
|
||||
---
|
||||
kind: Service
|
||||
name: web
|
||||
count: 3
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /usr/bin/httpd
|
||||
ports:
|
||||
- name: http
|
||||
port: 8080
|
||||
restart:
|
||||
mode: service
|
||||
attempts: 5
|
||||
delay: 2s
|
||||
update:
|
||||
strategy: rolling
|
||||
max_parallel: 1
|
||||
health:
|
||||
check_type: http
|
||||
interval: 5s
|
||||
timeout: 1s
|
||||
unhealthy_threshold: 2
|
||||
constraints:
|
||||
- node.role == "web"
|
||||
---
|
||||
```
|
||||
|
||||
### `DaemonSet`
|
||||
|
||||
A workload that runs on every matching node.
|
||||
|
||||
- `schedule` required; `mode` one of `every-node`, `matching`, `mandatory`.
|
||||
- `ports` **not allowed** (no Traefik route by default).
|
||||
- `count` **not allowed** (implicit = matching nodes).
|
||||
- `restart` required.
|
||||
|
||||
**Example**:
|
||||
```yaml
|
||||
---
|
||||
kind: DaemonSet
|
||||
name: log-shipper
|
||||
schedule:
|
||||
mode: every-node
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /usr/bin/fluent-bit
|
||||
restart:
|
||||
mode: service
|
||||
---
|
||||
```
|
||||
|
||||
## Block reference
|
||||
|
||||
### `runtime`
|
||||
|
||||
The runtime backend for the workload.
|
||||
|
||||
| Field | Key | Type | Default | Notes |
|
||||
|-------|-----|------|---------|-------|
|
||||
| `one_of` | `one_of` | string | — | Runtime type (see below) |
|
||||
| `image` | `image` | string | `""` | Container image (for `podman`) |
|
||||
| `command` | `command` | string | — | ExecStart command |
|
||||
|
||||
**Supported runtime types** (`one_of`):
|
||||
|
||||
| Type | Description | Requires |
|
||||
|------|-------------|----------|
|
||||
| `process` | Direct process execution via systemd (default) | systemd on target |
|
||||
| `wasm` / `wasmtime` | WASM via wasmtime CLI (apt-installed on peer, SSH exec) | wasmtime on target |
|
||||
| `podman` | Container via podman | podman on target |
|
||||
| `pve-vm` | Proxmox VM via `qm` | Proxmox node |
|
||||
| `pve-ct` | Proxmox container via `pct` | Proxmox node |
|
||||
| `proxmox` | Alias for Proxmox runtime | Proxmox node |
|
||||
|
||||
An empty/missing `Runtime` or `OneOf` is runtime-agnostic (always fits
|
||||
the runtime axis in the scheduler).
|
||||
|
||||
### `ports`
|
||||
|
||||
Array of port mappings. Required for `Service`.
|
||||
|
||||
| Field | Key | Type | Default | Notes |
|
||||
|-------|-----|------|---------|-------|
|
||||
| `name` | `name` | string | — | Port name (used in socket path) |
|
||||
| `port` | `port` | int | — | Container port |
|
||||
| `host_port` | `host_port` | int | `0` | Host port |
|
||||
| `protocol` | `protocol` | string | `""` | Protocol (e.g. `tcp`) |
|
||||
| `host_ip` | `host_ip` | string | `""` | Host IP |
|
||||
|
||||
**Example**:
|
||||
```yaml
|
||||
ports:
|
||||
- name: http
|
||||
port: 8080
|
||||
host_port: 80
|
||||
protocol: tcp
|
||||
- name: https
|
||||
port: 8443
|
||||
host_port: 443
|
||||
```
|
||||
|
||||
### `env`
|
||||
|
||||
Environment variables. Scalar values or secret references.
|
||||
|
||||
```yaml
|
||||
env:
|
||||
FOO: bar
|
||||
BAZ: "qux"
|
||||
SECRET_REF:
|
||||
from: "secret:db-password"
|
||||
INLINE: {from: "secret:token"}
|
||||
```
|
||||
|
||||
> Secret resolution (`from: "secret:..."`) lands in v0.11-P03. The
|
||||
> parser stores the reference; the emitter will emit
|
||||
> `EnvironmentFile=`/`LoadCredential=` in v0.11.
|
||||
|
||||
### `secrets`
|
||||
|
||||
List of secret names. Inline array or block list.
|
||||
|
||||
```yaml
|
||||
secrets: ["db-password", "api-token"]
|
||||
# or
|
||||
secrets:
|
||||
- db-password
|
||||
- api-token
|
||||
```
|
||||
|
||||
### `volumes`
|
||||
|
||||
Array of volume mounts.
|
||||
|
||||
| Field | Key | Type | Default | Notes |
|
||||
|-------|-----|------|---------|-------|
|
||||
| `name` | `name` | string | — | Volume name |
|
||||
| `type` | `type` | string | — | Volume type (e.g. `host`) |
|
||||
| `source` | `source` | string | — | Source path (or `replicate:<peer>,<peer>` for Syncthing) |
|
||||
| `target` | `target` | string | — | Mount target |
|
||||
| `read_only` | `read_only` | bool | `false` | Read-only mount (`true`/`yes`/`on`/`1`) |
|
||||
|
||||
**Example**:
|
||||
```yaml
|
||||
volumes:
|
||||
- name: data
|
||||
type: host
|
||||
source: /data
|
||||
target: /data
|
||||
read_only: true
|
||||
```
|
||||
|
||||
### `restart`
|
||||
|
||||
Restart policy.
|
||||
|
||||
| Field | Key | Type | Default | Notes |
|
||||
|-------|-----|------|---------|-------|
|
||||
| `mode` | `mode` | enum | — | `never`, `on-failure`, `service` |
|
||||
| `attempts` / `max_retries` | `attempts` or `max_retries` | int | `0` | Max retries (both keys accepted) |
|
||||
| `delay` | `delay` | duration string | `""` | Retry delay (e.g. `2s`) |
|
||||
|
||||
### `update`
|
||||
|
||||
Update strategy. Required for `Service`.
|
||||
|
||||
| Field | Key | Type | Default | Notes |
|
||||
|-------|-----|------|---------|-------|
|
||||
| `strategy` | `strategy` | enum | — | `rolling`, `canary`, `blue-green` |
|
||||
| `max_surge` | `max_surge` | int | `0` | Max surge |
|
||||
| `max_parallel` | `max_parallel` | int | `1` (clamped to `count`) | Max parallel updates |
|
||||
| `min_healthy_time` | `min_healthy_time` | duration | `""` | Min time healthy before next batch |
|
||||
| `healthy_deadline` | `healthy_deadline` | duration | `""` | Deadline for health |
|
||||
| `canary` | `canary` | int or `"<n>%"` | — | Canary size (int count or percentage) |
|
||||
| `auto_promote` | `auto_promote` | bool | `false` | Auto-promote canary (`true`/`yes`/`on`/`1`) |
|
||||
|
||||
**Strategies**:
|
||||
- **rolling**: batches of `max_parallel`, each batch waits for healthy.
|
||||
- **canary**: canary batch first, then `promote` (manual or `auto_promote`), then remaining in `max_parallel` batches.
|
||||
- **blue-green**: all new allocs start in parallel, wait healthy, then `cutover`.
|
||||
|
||||
> Transactional update execution lands in v0.11-P10. The current
|
||||
> emitter computes the plan; execution is a v0.11 deliverable.
|
||||
|
||||
### `service`
|
||||
|
||||
Traefik route definition. Implied for `Service`; not allowed for
|
||||
`Job`/`DaemonSet`. See [docs/ingress.md](ingress.md) for details.
|
||||
|
||||
| Field | Key | Type | Default | Notes |
|
||||
|-------|-----|------|---------|-------|
|
||||
| `name` | `name` | string | — | Service name |
|
||||
| `port` | `port` | int | — | Service port |
|
||||
| `bind` | `bind` | string (IP) | `""` | Bind mode: empty = Unix socket (default); `127.0.0.1` = TCP opt-in (R-007) |
|
||||
| `host` | `host` | string | `""` | Host (stored, not yet consumed by emitter) |
|
||||
| `route_id` | `route_id` | string | `""` | Route ID (stored, not yet consumed by emitter) |
|
||||
|
||||
### `health`
|
||||
|
||||
Health check. Required for `Service`.
|
||||
|
||||
| Field | Key | Type | Default | Notes |
|
||||
|-------|-----|------|---------|-------|
|
||||
| `check_type` | `check_type` | string | — | Check type (e.g. `http`) |
|
||||
| `interval` | `interval` | duration string | — | Check interval (e.g. `5s`) |
|
||||
| `timeout` | `timeout` | duration string | — | Check timeout |
|
||||
| `unhealthy_threshold` | `unhealthy_threshold` | int | `0` | Failures before unhealthy |
|
||||
|
||||
Maps to Traefik `healthCheck` stanza (`path: /healthz`).
|
||||
|
||||
### `lifecycle`
|
||||
|
||||
Lifecycle hooks. Maps to systemd `ExecStartPost` / `ExecStop`.
|
||||
|
||||
| Field | Key | Type | Default | systemd mapping |
|
||||
|-------|-----|------|---------|-----------------|
|
||||
| `post_start` | `post_start` | string list | nil | `ExecStartPost=` (runs after main starts) |
|
||||
| `pre_stop` | `pre_stop` | string list | nil | `ExecStop=` (runs before kill) |
|
||||
|
||||
**Example**:
|
||||
```yaml
|
||||
lifecycle:
|
||||
pre_stop:
|
||||
- /bin/sh -c 'sleep 5'
|
||||
- /usr/local/bin/drain.sh
|
||||
post_start:
|
||||
- /usr/local/bin/warm-cache.sh
|
||||
```
|
||||
|
||||
### `constraints`
|
||||
|
||||
CEL-subset expressions for node selection. Inline array or block list.
|
||||
|
||||
```yaml
|
||||
constraints:
|
||||
- node.role == "web"
|
||||
- region == "us"
|
||||
# or inline
|
||||
constraints: ['node.role == "web"', 'region == "us"']
|
||||
```
|
||||
|
||||
**CEL subset grammar** (hand-rolled, no CEL dependency):
|
||||
- Node attributes: `node.hostname`, `node.kind`, `node.cpus`,
|
||||
`node.memory`, `node.tags`, `node.runtimes`
|
||||
- Bare identifiers: equivalent to `node.<name>`
|
||||
- Literals: string (`"..."`), int
|
||||
- Comparisons: `==`, `!=`, `>=`, `<=`, `>`, `<`
|
||||
- Membership: `in`, `not in`
|
||||
- Boolean: `and`, `or`, `not`, parentheses
|
||||
- Anything outside the subset returns an error (node skipped, not
|
||||
silently mis-evaluated)
|
||||
|
||||
### `affinity`
|
||||
|
||||
Co-location / anti-affinity rules.
|
||||
|
||||
```yaml
|
||||
affinity:
|
||||
- target: zone == "a"
|
||||
weight: 80
|
||||
- target: web
|
||||
weight: -50 # anti-affinity (negative weight)
|
||||
```
|
||||
|
||||
- `target`: CEL expression or bare workload name (for name-based
|
||||
co-location).
|
||||
- `weight`: positive = co-locate, negative = anti-affinity.
|
||||
- Affinity is a **hint** (not a gate); evaluation failures are ignored.
|
||||
|
||||
### `tasks` (task group)
|
||||
|
||||
Multi-process alloc (P06). When `tasks` is non-empty, the alloc runs
|
||||
multiple processes, each as its own systemd unit, grouped under a
|
||||
systemd target.
|
||||
|
||||
```yaml
|
||||
tasks:
|
||||
- name: app
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /usr/bin/httpd -f
|
||||
env:
|
||||
LOG_LEVEL: debug
|
||||
- name: sidecar
|
||||
runtime:
|
||||
one_of: wasm
|
||||
command: /bin/wasm-runner sidecar.wasm
|
||||
```
|
||||
|
||||
- A task with no `runtime:` inherits the top-level `spec.Runtime`.
|
||||
- Each task can have its own `env:` overlay.
|
||||
- `command` falls back: `task.Command` → `task.Runtime.Command` →
|
||||
`spec.Runtime.Command`.
|
||||
- Task names must be unique within the group.
|
||||
|
||||
## Kinds matrix
|
||||
|
||||
| Feature | Job | Service | DaemonSet |
|
||||
|---------|-----|---------|-----------|
|
||||
| `count` | must be 1 | ≥ 1 | not allowed |
|
||||
| `ports` | optional | **required** | not allowed |
|
||||
| `service` block | not allowed | optional (implied) | not allowed |
|
||||
| `restart` | optional | **required** | **required** |
|
||||
| `update` | optional | **required** | optional |
|
||||
| `health` | optional | **required** | optional |
|
||||
| `runtime` | optional | **required** (or task group) | optional |
|
||||
| `schedule` | optional | optional | **required** |
|
||||
| `tasks` | optional | optional | optional |
|
||||
| Traefik route | no | yes (implied) | no (by default) |
|
||||
|
||||
## Body semantics
|
||||
|
||||
The body after the closing `---` is preserved **byte-exact** (R-015) —
|
||||
including trailing newlines, CRLF, BOM in body, and `---` inside code
|
||||
fences. The body is carried verbatim to the target node. It is not
|
||||
interpreted as commands/scripts by the parser today.
|
||||
|
||||
## Deprecated: HCL jobspec
|
||||
|
||||
The legacy HCL jobspec format is supported via an adapter during the
|
||||
migration window. It is deprecated in v0.9 and will be removed in
|
||||
v0.11.
|
||||
|
||||
```hcl
|
||||
job "hello-orca" {
|
||||
}
|
||||
|
||||
task "greet" {
|
||||
command = "/bin/echo"
|
||||
args = ["hello", "from", "orca"]
|
||||
}
|
||||
```
|
||||
|
||||
The adapter converts this to a `*WorkloadSpec{Kind: "Job", Name:
|
||||
"hello-orca", Count: 1, Runtime: {OneOf: "process", Command:
|
||||
"/bin/echo"}}`. Use `.md` for all new jobspecs.
|
||||
|
||||
## See also
|
||||
|
||||
- [docs/cli.md](cli.md) — CLI reference (including `orca job run`)
|
||||
- [docs/ingress.md](ingress.md) — Traefik ingress configuration
|
||||
- [examples/full-stack/](../examples/full-stack/) — Full-stack example jobspecs
|
||||
+158
-77
@@ -1,96 +1,177 @@
|
||||
# Namespace and Paths
|
||||
|
||||
Orca stores all on-disk state (SQLite database, CA certs, server certs,
|
||||
config) under a single **namespace root** directory. This document
|
||||
describes how that root is resolved and how to override it.
|
||||
Orca stores all on-disk state under a single **namespace root**
|
||||
directory. The v0.9 re-architecture introduced a multi-namespace
|
||||
layout (R-002) where each namespace is a self-contained directory tree
|
||||
with its own database, jobs, allocs, env, and secrets. A `cluster/`
|
||||
directory holds cluster-wide artifacts shared across namespaces.
|
||||
|
||||
## Default: User-Level (`~/.orca`)
|
||||
> **v0.9 layout (canonical)**: This document describes the v0.9
|
||||
> multi-namespace layout. The v0.8 flat layout (`orca.db`, `ca.crt`,
|
||||
> `server.crt` at the root) is deprecated and will be removed in
|
||||
> v0.11. See [v0.8 flat layout](#deprecated-v08-flat-layout) below.
|
||||
|
||||
By default, the namespace root is `~/.orca` (i.e., `$HOME/.orca`).
|
||||
All orca state lives under this directory:
|
||||
## Namespace root resolution
|
||||
|
||||
| Path | Contents |
|
||||
|------|----------|
|
||||
| `~/.orca/orca.db` | SQLite database (jobs, nodes, tasks, audit log, capacity) |
|
||||
| `~/.orca/ca.crt` | CA certificate (PEM, mode 0644) |
|
||||
| `~/.orca/ca.key` | CA private key (PEM, mode 0600) |
|
||||
| `~/.orca/server.crt` | Server certificate (PEM, mode 0644) |
|
||||
| `~/.orca/server.key` | Server private key (PEM, mode 0600) |
|
||||
|
||||
## Override: `ORCA_HOME` Environment Variable (REQ-041)
|
||||
|
||||
Set the `ORCA_HOME` environment variable to change the namespace root
|
||||
for **all** orca components (database, certs, init, daemon):
|
||||
|
||||
```bash
|
||||
export ORCA_HOME=/var/lib/orca
|
||||
orca init # creates /var/lib/orca/
|
||||
orca daemon # reads /var/lib/orca/orca.db
|
||||
orca cert ca-init # writes CA to /var/lib/orca/
|
||||
```
|
||||
|
||||
This is the single source of truth for the namespace root. Every
|
||||
component that reads or writes on-disk state resolves the root via
|
||||
`ORCA_HOME` (falling back to `~/.orca` when unset).
|
||||
|
||||
### Use cases
|
||||
|
||||
- **Testing**: point `ORCA_HOME` at a temp directory.
|
||||
- **Multi-instance**: run multiple orca daemons on the same host with
|
||||
different `ORCA_HOME` values.
|
||||
- **Custom layout**: store state on a mounted volume
|
||||
(`ORCA_HOME=/mnt/orca-data`).
|
||||
|
||||
## System-Level: `--system` Flag (REQ-042)
|
||||
|
||||
The `--system` persistent flag selects the system-level namespace root
|
||||
`/root/.orca`. This is intended for root-owned system deployments
|
||||
(where orca runs as a system service under root):
|
||||
|
||||
```bash
|
||||
sudo orca --system init # creates /root/.orca/
|
||||
sudo orca --system daemon # reads /root/.orca/orca.db
|
||||
sudo orca --system cert ca-init # writes CA to /root/.orca/
|
||||
```
|
||||
|
||||
The `--system` flag is equivalent to setting `ORCA_HOME=/root/.orca`,
|
||||
but it is a CLI convenience that does not require exporting an env var.
|
||||
If `ORCA_HOME` is already set to a different value, `--system` returns
|
||||
an error (to avoid silent namespace mismatches).
|
||||
|
||||
### Path layout
|
||||
|
||||
System-level uses the same directory shape as user-level, just under
|
||||
`/root/.orca` instead of `~/.orca`:
|
||||
|
||||
| Path | Contents |
|
||||
|------|----------|
|
||||
| `/root/.orca/orca.db` | SQLite database |
|
||||
| `/root/.orca/ca.crt` | CA certificate |
|
||||
| `/root/.orca/ca.key` | CA private key |
|
||||
| `/root/.orca/server.crt` | Server certificate |
|
||||
| `/root/.orca/server.key` | Server private key |
|
||||
|
||||
## Resolution Order
|
||||
The namespace root is resolved in this order:
|
||||
|
||||
1. If `--system` flag is passed → root is `/root/.orca` (errors if
|
||||
`ORCA_HOME` is set to a conflicting value).
|
||||
2. Else if `ORCA_HOME` is set → root is `$ORCA_HOME`.
|
||||
3. Else → root is `~/.orca` (`$HOME/.orca`).
|
||||
|
||||
## `ORCA_DB` Override
|
||||
### `ORCA_HOME` (REQ-041)
|
||||
|
||||
For finer-grained control, `ORCA_DB` overrides **only** the database
|
||||
path (not the cert paths). This is primarily a testing affordance. When
|
||||
`ORCA_DB` is set, certs still resolve under `ORCA_HOME` (or `~/.orca`).
|
||||
Set the `ORCA_HOME` environment variable to change the namespace root
|
||||
for all orca components:
|
||||
|
||||
```bash
|
||||
export ORCA_HOME=/var/lib/orca
|
||||
orca init # creates /var/lib/orca/
|
||||
orca ns create prod
|
||||
```
|
||||
|
||||
### `--system` (REQ-042)
|
||||
|
||||
The `--system` persistent flag selects the system-level namespace root
|
||||
`/root/.orca`:
|
||||
|
||||
```bash
|
||||
sudo orca --system init # creates /root/.orca/
|
||||
sudo orca --system ns list
|
||||
```
|
||||
|
||||
If `ORCA_HOME` is already set to a different value, `--system` returns
|
||||
an error (to avoid silent namespace mismatches).
|
||||
|
||||
## v0.9 multi-namespace layout (R-002)
|
||||
|
||||
```
|
||||
$ORCA_HOME/
|
||||
├── cluster/ # cluster-wide (NOT a workload namespace)
|
||||
│ ├── ca.crt, ca.key # step-ca root (R-006, D-101)
|
||||
│ ├── master.key # AES-256-GCM root (R-011, mode 0600)
|
||||
│ ├── config.md # Markdown frontmatter config (R-014)
|
||||
│ ├── known_hosts # SSH known_hosts (D-035)
|
||||
│ ├── orca_ssh_key # orca SSH private key (D-037)
|
||||
│ ├── orca_ssh_key.pub # orca SSH public key
|
||||
│ ├── peers/<host>/ # per-peer directory
|
||||
│ ├── txns/ # cluster transaction log (R-016)
|
||||
│ └── state/ # cluster state
|
||||
├── _defaults/ # implicit root namespace (always exists)
|
||||
│ ├── ns.md # namespace frontmatter (kind: Namespace)
|
||||
│ ├── .env # per-namespace env
|
||||
│ ├── .env.secrets # encrypted secrets
|
||||
│ ├── db/orca.db # per-namespace SQLite database
|
||||
│ ├── jobs/ # submitted jobspecs
|
||||
│ └── alloc/ # allocation state
|
||||
├── <explicit-namespace>/ # operator-created (e.g., prod, staging)
|
||||
│ ├── ns.md
|
||||
│ ├── .env, .env.secrets
|
||||
│ ├── db/orca.db
|
||||
│ ├── jobs/, alloc/
|
||||
│ └── syncthing/ # Syncthing config (if replicated volumes)
|
||||
└── orca_cache.db # CLI-side cache (R-008)
|
||||
```
|
||||
|
||||
### Key points
|
||||
|
||||
- **`_defaults/`** is the implicit root namespace (D-159). It always
|
||||
exists. Every namespace inherits from `_defaults` and cannot opt out
|
||||
(D-185, D-187).
|
||||
- **`cluster/`** is NOT a workload namespace — it holds cluster-wide
|
||||
artifacts (CA, master key, SSH keys, known_hosts, peers, txns).
|
||||
- **Per-namespace DBs**: each namespace has its own
|
||||
`db/orca.db` (R-002). No namespace column in SQLite.
|
||||
- **Namespace inheritance**: child namespaces inherit env and
|
||||
constraints from parents (via `ns.md` frontmatter `parents:` field).
|
||||
`_defaults` is always appended last in the inheritance chain.
|
||||
- **`orca ns` subcommands**: `list`, `create`, `delete`, `inspect`,
|
||||
`validate` — see [docs/cli.md](cli.md#orca-ns).
|
||||
|
||||
### Path reference (`internal/paths/`)
|
||||
|
||||
| Function | Path | Contents |
|
||||
|----------|------|----------|
|
||||
| `Root()` | `$ORCA_HOME` | Namespace root |
|
||||
| `ClusterDir()` | `Root()/cluster` | Cluster-wide artifacts |
|
||||
| `NamespaceDir(ns)` | `Root()/ns` | Per-namespace directory |
|
||||
| `NSDb(ns)` | `Root()/ns/db/orca.db` | Per-namespace SQLite DB |
|
||||
| `NSEnv(ns)` | `Root()/ns/.env` | Per-namespace env |
|
||||
| `NSSecrets(ns)` | `Root()/ns/.env.secrets` | Encrypted secrets |
|
||||
| `NSJobs(ns)` | `Root()/ns/jobs` | Jobs dir |
|
||||
| `NSAlloc(ns)` | `Root()/ns/alloc` | Alloc dir |
|
||||
| `NSMd(ns)` | `Root()/ns/ns.md` | Namespace frontmatter |
|
||||
| `DefaultNamespace()` | `_defaults` | Implicit root (D-159) |
|
||||
| `CACertPath()` | `ClusterDir()/ca.crt` | step-ca root (D-101) |
|
||||
| `MasterKeyPath()` | `ClusterDir()/master.key` | AES-256-GCM root key |
|
||||
| `KnownHostsPath()` | `ClusterDir()/known_hosts` | SSH known_hosts |
|
||||
| `SSHKeyPath()` | `ClusterDir()/orca_ssh_key` | orca SSH private key |
|
||||
| `ConfigPath()` | `ClusterDir()/config.md` | Markdown config (R-014) |
|
||||
| `CacheDB()` | `Root()/orca_cache.db` | CLI-side cache (R-008) |
|
||||
| `PeersDir()` | `ClusterDir()/peers` | Peers directory |
|
||||
| `TxnDir()` | `ClusterDir()/txns` | Transaction log (R-016) |
|
||||
|
||||
## Creating and managing namespaces
|
||||
|
||||
```bash
|
||||
# List all namespaces
|
||||
orca ns list
|
||||
|
||||
# Create a namespace (inherits from _defaults)
|
||||
orca ns create prod
|
||||
|
||||
# Create a namespace with an explicit parent
|
||||
orca ns create staging --parent prod
|
||||
|
||||
# Inspect the effective inheritance chain + merged env
|
||||
orca ns inspect prod
|
||||
|
||||
# Validate a namespace's inheritance chain
|
||||
orca ns validate prod
|
||||
|
||||
# Delete an empty namespace (refuses if jobs/ or alloc/ non-empty)
|
||||
orca ns delete staging
|
||||
```
|
||||
|
||||
See [docs/cli.md](cli.md#orca-ns) for the full `orca ns` reference.
|
||||
|
||||
## `ORCA_DB` override
|
||||
|
||||
For finer-grained control, `ORCA_DB` overrides only the database path
|
||||
(not the cert/namespace paths). This is primarily a testing affordance.
|
||||
|
||||
```bash
|
||||
export ORCA_DB=/tmp/test.db
|
||||
orca daemon # uses /tmp/test.db for the DB, ~/.orca/ for certs
|
||||
orca init # uses /tmp/test.db for the DB, ~/.orca/ for everything else
|
||||
```
|
||||
|
||||
## See Also
|
||||
## Deprecated: v0.8 flat layout
|
||||
|
||||
> **Deprecated in v0.9**: The v0.8 flat layout (`orca.db`, `ca.crt`,
|
||||
> `ca.key`, `server.crt`, `server.key` at the namespace root) is
|
||||
> superseded by the v0.9 multi-namespace layout (R-002). The v0.8
|
||||
> layout is supported during the dual-write window via
|
||||
> `internal/certpaths` (a thin shim) and will be removed in v0.11.
|
||||
|
||||
The v0.8 flat layout stored all state at the namespace root:
|
||||
|
||||
| Path | Contents |
|
||||
|------|----------|
|
||||
| `~/.orca/orca.db` | SQLite database |
|
||||
| `~/.orca/ca.crt` | CA certificate |
|
||||
| `~/.orca/ca.key` | CA private key |
|
||||
| `~/.orca/server.crt` | Server certificate |
|
||||
| `~/.orca/server.key` | Server private key |
|
||||
|
||||
The v0.9 re-architecture moved these to `cluster/` (CA, SSH keys) and
|
||||
per-namespace `db/` (SQLite) to support multi-tenancy (R-002). The
|
||||
`orca doctor --legacy-paths` command (v0.11-P14c) will detect v0.8
|
||||
residue and recommend migration.
|
||||
|
||||
## See also
|
||||
|
||||
- [Install Guide](install.md) — 1-liner install with `install.sh`.
|
||||
- [Docker Guide](docker.md) — running orca in a container (uses
|
||||
`ORCA_HOME=/var/lib/orca` inside the image).
|
||||
- [Docker Guide](docker.md) — running orca in a container.
|
||||
- [CLI Reference](cli.md) — `orca ns` subcommands.
|
||||
- [Jobspec Reference](jobspec.md) — markdown frontmatter schema.
|
||||
@@ -0,0 +1,191 @@
|
||||
# Full-Stack Example with Ingress
|
||||
|
||||
This directory contains a complete multi-service stack deployed with
|
||||
Orca, including Traefik ingress configuration. Each file is a valid
|
||||
Orca jobspec (`.md` frontmatter) that passes the v0.9 parser and schema
|
||||
validators.
|
||||
|
||||
> **Runnable out-of-the-box**: The `runtime.command` in each example
|
||||
> uses `/bin/sleep 3600` (for long-running services) or `/bin/echo`
|
||||
> (for one-shot jobs) so that `orca job run <file>.md` succeeds on any
|
||||
> Linux machine without installing any software. Each file has a
|
||||
> **Production substitution** note showing the real binary to use in a
|
||||
> deployment (e.g. `/usr/bin/httpd`,
|
||||
> `/usr/lib/postgresql/16/bin/postgres`).
|
||||
|
||||
## Stack overview
|
||||
|
||||
| File | Kind | Runtime | Ingress | Description |
|
||||
|------|------|---------|---------|-------------|
|
||||
| `web-app.md` | Service | process | Unix socket (default) | Frontend HTTP server, 3 replicas, rolling update |
|
||||
| `api.md` | Service | process | TCP `127.0.0.1:9090` (R-007 opt-in) | Backend API, 2 replicas, canary update |
|
||||
| `worker.md` | Job | process | none | One-shot batch worker with lifecycle hooks |
|
||||
| `log-shipper.md` | Service | process | Unix socket (metrics) | Log shipper on a dedicated node |
|
||||
| `postgres.md` | Service | process | Unix socket | Database with volume replication, blue-green update |
|
||||
|
||||
## Rendered artifacts
|
||||
|
||||
The `rendered/` directory shows what Orca generates on the target nodes
|
||||
when you submit these jobspecs:
|
||||
|
||||
| File | Description |
|
||||
|------|-------------|
|
||||
| `traefik-dynamic-web-app.yaml` | Traefik dynamic config for the web-app Service |
|
||||
| `traefik-dynamic-api.yaml` | Traefik dynamic config for the api Service (TCP bind) |
|
||||
| `systemd-web-app.service` | Systemd unit for the web-app alloc |
|
||||
| `systemd-api.service` | Systemd unit for the api alloc (with TCP bind marker) |
|
||||
| `systemd-log-shipper.service` | Systemd unit for the log-shipper alloc |
|
||||
|
||||
## Walkthrough
|
||||
|
||||
### Prerequisites
|
||||
|
||||
- Orca installed (`orca version` works)
|
||||
- 2+ Linux nodes reachable over SSH (for multi-node scheduling)
|
||||
- Traefik installed on the lead node (watches `/etc/traefik/dynamic/`)
|
||||
|
||||
### Step 1: Initialize the cluster
|
||||
|
||||
```bash
|
||||
# On the operator laptop
|
||||
orca init
|
||||
```
|
||||
|
||||
This creates `~/.orca/` (or `/root/.orca` with `--system`), bootstraps
|
||||
the CA, generates the server cert, auto-detects the OS, and registers
|
||||
a localhost node.
|
||||
|
||||
### Step 2: Join remote nodes
|
||||
|
||||
```bash
|
||||
# Join a Proxmox node (v0.9 canonical SSH-push path)
|
||||
orca node join --type proxmox --host 192.168.1.100 --ssh-user root
|
||||
|
||||
# Join a second node
|
||||
ORCA_PROXMOX_PASSWORD=secret orca node join --type proxmox --host 192.168.1.101
|
||||
```
|
||||
|
||||
### Step 3: Declare node capacity
|
||||
|
||||
The CLI-side scheduler uses capacity declarations for bin-packing:
|
||||
|
||||
```bash
|
||||
orca node capacity set --cpu 4000 --memory 8192 --disk 100000 --node 192.168.1.100
|
||||
orca node capacity set --cpu 4000 --memory 8192 --disk 100000 --node 192.168.1.101
|
||||
```
|
||||
|
||||
### Step 4: Create a namespace
|
||||
|
||||
```bash
|
||||
orca ns create prod --parent _defaults
|
||||
```
|
||||
|
||||
This creates `~/.orca/prod/` with `db/`, `jobs/`, `alloc/`, and `ns.md`.
|
||||
|
||||
### Step 5: Submit the stack
|
||||
|
||||
```bash
|
||||
orca job run web-app.md
|
||||
orca job run api.md
|
||||
orca job run worker.md
|
||||
orca job run log-shipper.md
|
||||
orca job run postgres.md
|
||||
```
|
||||
|
||||
Each `orca job run` parses the `.md` jobspec, validates it against the
|
||||
schema, schedules it via the CLI-side bin-packing scheduler, and
|
||||
generates the systemd + Traefik artifacts on the target node via
|
||||
SSH-push.
|
||||
|
||||
### Step 6: Observe placements
|
||||
|
||||
```bash
|
||||
orca job list --watch
|
||||
|
||||
# Output:
|
||||
# ID NAME STATUS EXIT
|
||||
# abc-123... web-app running 0
|
||||
# def-456... api running 0
|
||||
# ghi-789... worker complete 0
|
||||
# jkl-012... log-shipper running 0
|
||||
# mno-345... postgres running 0
|
||||
```
|
||||
|
||||
### Step 7: Inspect rendered artifacts
|
||||
|
||||
After submission, the target nodes have:
|
||||
|
||||
```
|
||||
/etc/systemd/system/orca-v1-web-app.service # systemd unit
|
||||
/etc/systemd/system/orca-v1-api.service # systemd unit (TCP bind)
|
||||
/etc/traefik/dynamic/orca-web-app.yaml # Traefik dynamic config
|
||||
/etc/traefik/dynamic/orca-api.yaml # Traefik dynamic config
|
||||
/run/orca/alloc-web-app-0/port-http.sock # Unix socket (R-007 default)
|
||||
```
|
||||
|
||||
See the `rendered/` directory in this example for the exact file
|
||||
contents.
|
||||
|
||||
### Step 8: Verify ingress
|
||||
|
||||
Traefik watches `/etc/traefik/dynamic/` and atomically reloads when a
|
||||
file changes (write-tmp + rename, gate C-10). The web-app is reachable
|
||||
at `https://<cluster-domain>/web-app` and the API at
|
||||
`https://<cluster-domain>/api`.
|
||||
|
||||
Health checks (`/healthz` on each backend) ensure Traefik only routes
|
||||
to healthy instances.
|
||||
|
||||
### Step 9: Drain and rollback
|
||||
|
||||
To drain a service (stop traffic, keep the workload running):
|
||||
|
||||
```bash
|
||||
# Orca writes a Traefik config with weight:0 on every backend
|
||||
# (RenderDrain). Traefik stops sending traffic.
|
||||
```
|
||||
|
||||
To roll back, re-submit the normal jobspec — Orca writes the
|
||||
non-drained Traefik config and Traefik resumes routing.
|
||||
|
||||
## Ingress model
|
||||
|
||||
See [docs/ingress.md](../../docs/ingress.md) for the full Traefik
|
||||
ingress reference. Key points:
|
||||
|
||||
- `kind: Service` **implies** a Traefik route (D-175).
|
||||
- Default bind is a **Unix socket** at
|
||||
`/run/orca/alloc-<id>/port-<name>.sock` (R-007).
|
||||
- `service.bind: 127.0.0.1` opts in to **TCP** (loopback only).
|
||||
- One Traefik dynamic file per Service at
|
||||
`/etc/traefik/dynamic/orca-<name>.yaml`.
|
||||
- Atomic reload via write-tmp + rename (gate C-10).
|
||||
- Drain sets `weight: 0` per backend.
|
||||
|
||||
## Validation
|
||||
|
||||
All jobspecs in this directory are validated by a Go test:
|
||||
|
||||
```bash
|
||||
go test ./examples/full-stack/ -v -run TestExamplesValidate
|
||||
```
|
||||
|
||||
This test parses each `.md` file with `jobspec.ParseFile` and validates
|
||||
it against `schema.ValidatorFor(kind)` — ensuring every field used in
|
||||
the examples exists in the current `WorkloadSpec` struct and passes the
|
||||
per-kind validators (gate C-20).
|
||||
|
||||
## v0.11 forward
|
||||
|
||||
The following are not yet implemented in v0.9 and will land in v0.11:
|
||||
|
||||
- **DaemonSet `schedule:` block**: the parser does not yet populate the
|
||||
`schedule:` frontmatter block (v0.9 parser gap). The `log-shipper`
|
||||
example uses `kind: Service` with `count: 1` and a `node.role`
|
||||
constraint as a workaround.
|
||||
- **Secret resolution**: `env: { KEY: { from: "secret:..." } }` is
|
||||
parsed but not resolved to `EnvironmentFile=`/`LoadCredential=` until
|
||||
v0.11-P03.
|
||||
- **Transactional update execution**: the `update:` block's plan is
|
||||
computed but not executed transactionally until v0.11-P10.
|
||||
- **Socket activation**: real socket unit files land in v0.11-P08.
|
||||
@@ -0,0 +1,46 @@
|
||||
---
|
||||
kind: Service
|
||||
name: api
|
||||
count: 2
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /bin/sleep 3600
|
||||
ports:
|
||||
- name: api
|
||||
port: 9090
|
||||
restart:
|
||||
mode: service
|
||||
attempts: 3
|
||||
delay: 5s
|
||||
update:
|
||||
strategy: canary
|
||||
canary: 1
|
||||
max_parallel: 1
|
||||
auto_promote: false
|
||||
min_healthy_time: 30s
|
||||
healthy_deadline: 5m
|
||||
service:
|
||||
name: api
|
||||
port: 9090
|
||||
bind: 127.0.0.1
|
||||
health:
|
||||
check_type: http
|
||||
interval: 10s
|
||||
timeout: 2s
|
||||
unhealthy_threshold: 3
|
||||
constraints:
|
||||
- node.role == "api"
|
||||
- node.cpus >= 2
|
||||
env:
|
||||
DB_HOST: postgres
|
||||
DB_PORT: "5432"
|
||||
LOG_LEVEL: info
|
||||
---
|
||||
# API Server
|
||||
|
||||
Backend API service binding to 127.0.0.1:9090 (TCP opt-in, R-007).
|
||||
Canary update strategy with manual promote. Two replicas with CPU
|
||||
constraint (>= 2 vCPUs) and API-role node selection.
|
||||
|
||||
> **Production substitution**: replace `runtime.command` with your
|
||||
> actual API binary, e.g. `/usr/bin/api-server --listen 127.0.0.1:9090`.
|
||||
@@ -0,0 +1,60 @@
|
||||
package fullstack_test
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
"git.cloudinit.dev/coreci/orca/internal/spec/schema"
|
||||
)
|
||||
|
||||
// TestExamplesValidate parses and validates every jobspec in
|
||||
// examples/full-stack/ against the current parser and schema validators
|
||||
// (gate C-20, REQ-094). This ensures the example jobspecs use only
|
||||
// fields that exist in the current WorkloadSpec struct and pass the
|
||||
// per-kind validators.
|
||||
func TestExamplesValidate(t *testing.T) {
|
||||
dir := filepath.Join("..", "..", "examples", "full-stack")
|
||||
entries, err := os.ReadDir(dir)
|
||||
if err != nil {
|
||||
t.Fatalf("read examples dir: %v", err)
|
||||
}
|
||||
for _, e := range entries {
|
||||
if e.IsDir() {
|
||||
continue
|
||||
}
|
||||
name := e.Name()
|
||||
// Skip README.md and other non-jobspec markdown files.
|
||||
if name == "README.md" {
|
||||
continue
|
||||
}
|
||||
ext := filepath.Ext(name)
|
||||
if ext != ".md" && ext != ".yaml" && ext != ".yml" {
|
||||
continue
|
||||
}
|
||||
t.Run(name, func(t *testing.T) {
|
||||
path := filepath.Join(dir, name)
|
||||
spec, err := jobspec.ParseFile(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseFile %s: %v", name, err)
|
||||
}
|
||||
if spec == nil {
|
||||
t.Fatalf("ParseFile %s: spec is nil", name)
|
||||
}
|
||||
if spec.Kind == "" {
|
||||
t.Fatalf("ParseFile %s: kind is empty", name)
|
||||
}
|
||||
if spec.Name == "" {
|
||||
t.Fatalf("ParseFile %s: name is empty", name)
|
||||
}
|
||||
validator, err := schema.ValidatorFor(spec.Kind)
|
||||
if err != nil {
|
||||
t.Fatalf("ValidatorFor %s (kind %s): %v", name, spec.Kind, err)
|
||||
}
|
||||
if err := validator.Validate(spec); err != nil {
|
||||
t.Fatalf("Validate %s: %v", name, err)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,43 @@
|
||||
---
|
||||
kind: Service
|
||||
name: log-shipper
|
||||
count: 1
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /bin/sleep 3600
|
||||
ports:
|
||||
- name: metrics
|
||||
port: 2024
|
||||
restart:
|
||||
mode: service
|
||||
attempts: 3
|
||||
delay: 10s
|
||||
update:
|
||||
strategy: rolling
|
||||
max_parallel: 1
|
||||
health:
|
||||
check_type: http
|
||||
interval: 30s
|
||||
timeout: 5s
|
||||
unhealthy_threshold: 3
|
||||
constraints:
|
||||
- node.role == "logs"
|
||||
env:
|
||||
LOG_LEVEL: warn
|
||||
OUTPUT: unix:///run/orca/alloc-log-collector/ingest.sock
|
||||
---
|
||||
# Log Shipper
|
||||
|
||||
Log shipper service (fluent-bit) running on a dedicated logs-role node.
|
||||
Exposes a metrics port for health checking. Ships logs to a central
|
||||
collector via Unix socket.
|
||||
|
||||
> **Production substitution**: replace `runtime.command` with your
|
||||
> actual log shipper binary, e.g.
|
||||
> `/usr/bin/fluent-bit -c /etc/orca/log-shipper/fluent-bit.conf`.
|
||||
|
||||
> **Note**: DaemonSet kind is defined in the schema but the parser does
|
||||
> not yet populate the `schedule:` block from frontmatter (v0.9 parser
|
||||
> gap). This example uses `kind: Service` with `count: 1` and a
|
||||
> `node.role == "logs"` constraint to achieve single-node placement
|
||||
> until the parser gains `schedule:` support (v0.11).
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
kind: Service
|
||||
name: postgres
|
||||
count: 1
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /bin/sleep 3600
|
||||
ports:
|
||||
- name: pg
|
||||
port: 5432
|
||||
restart:
|
||||
mode: service
|
||||
attempts: 5
|
||||
delay: 10s
|
||||
update:
|
||||
strategy: blue-green
|
||||
min_healthy_time: 60s
|
||||
healthy_deadline: 10m
|
||||
service:
|
||||
name: postgres
|
||||
port: 5432
|
||||
health:
|
||||
check_type: http
|
||||
interval: 15s
|
||||
timeout: 5s
|
||||
unhealthy_threshold: 3
|
||||
volumes:
|
||||
- name: data
|
||||
type: host
|
||||
source: replicate:peer-b,peer-c
|
||||
target: /var/lib/postgresql/data
|
||||
read_only: false
|
||||
constraints:
|
||||
- node.role == "db"
|
||||
- node.cpus >= 4
|
||||
- node.memory >= 8192
|
||||
env:
|
||||
POSTGRES_DB: appdb
|
||||
POSTGRES_USER: orca
|
||||
PGDATA: /var/lib/postgresql/data
|
||||
---
|
||||
# PostgreSQL
|
||||
|
||||
Database service with a single replica, blue-green update strategy,
|
||||
and volume replication via Syncthing (replicate:peer-b,peer-c). The
|
||||
data volume is replicated to two peers for fault tolerance. Health
|
||||
check on port 5432. Constraints require DB-role nodes with >= 4 vCPUs
|
||||
and >= 8 GiB memory.
|
||||
|
||||
> **Production substitution**: replace `runtime.command` with your
|
||||
> actual postgres binary, e.g.
|
||||
> `/usr/lib/postgresql/16/bin/postgres -D /var/lib/postgresql/data`.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Systemd unit for orca api service (alloc api-0)
|
||||
# Generated by SystemdEmitter (internal/emitter/systemd.go)
|
||||
# Path on target node: /etc/systemd/system/orca-v1-api.service
|
||||
# service.bind: 127.0.0.1 (TCP opt-in, R-007)
|
||||
[Service]
|
||||
ExecStart=/bin/sleep 3600
|
||||
RuntimeDirectory=orca/alloc-api-0
|
||||
# socket: /run/orca/alloc-api-0/port-api.sock
|
||||
ExecStartPre=/bin/echo orca: bind 127.0.0.1 port api (tcp, R-007 opt-in)
|
||||
@@ -0,0 +1,7 @@
|
||||
# Systemd unit for orca log-shipper service
|
||||
# Generated by SystemdEmitter (internal/emitter/systemd.go)
|
||||
# Path on target node: /etc/systemd/system/orca-v1-log-shipper.service
|
||||
[Service]
|
||||
ExecStart=/bin/sleep 3600
|
||||
RuntimeDirectory=orca/alloc-log-shipper-0
|
||||
# socket: /run/orca/alloc-log-shipper-0/port-metrics.sock
|
||||
@@ -0,0 +1,11 @@
|
||||
# Systemd unit for orca web-app service (alloc web-app-0)
|
||||
# Generated by SystemdEmitter (internal/emitter/systemd.go)
|
||||
# Path on target node: /etc/systemd/system/orca-v1-web-app.service
|
||||
# Unit name prefix orca-v1- (dual-write window, REQ-090)
|
||||
[Service]
|
||||
ExecStart=/bin/sleep 3600
|
||||
ExecStartPost=/bin/echo cache warmed
|
||||
ExecStop=/bin/sleep 5
|
||||
ExecStop=/bin/echo draining web-app
|
||||
RuntimeDirectory=orca/alloc-web-app-0
|
||||
# socket: /run/orca/alloc-web-app-0/port-http.sock
|
||||
@@ -0,0 +1,23 @@
|
||||
# Traefik dynamic config for orca api service
|
||||
# Generated by TraefikEmitter (internal/emitter/traefik.go)
|
||||
# Path on target node: /etc/traefik/dynamic/orca-api.yaml
|
||||
# service.bind: 127.0.0.1 (TCP opt-in, R-007)
|
||||
http:
|
||||
routers:
|
||||
orca-api:
|
||||
rule: PathPrefix("/api")
|
||||
service: orca-api
|
||||
tls:
|
||||
certResolver: orca
|
||||
domains:
|
||||
- main: "cluster.orca.local"
|
||||
services:
|
||||
orca-api:
|
||||
loadBalancer:
|
||||
servers:
|
||||
- url: "http://127.0.0.1:9090"
|
||||
- url: "http://127.0.0.1:9090"
|
||||
healthCheck:
|
||||
path: /healthz
|
||||
interval: 10s
|
||||
timeout: 2s
|
||||
@@ -0,0 +1,24 @@
|
||||
# Traefik dynamic config for orca web-app service
|
||||
# Generated by TraefikEmitter (internal/emitter/traefik.go)
|
||||
# Path on target node: /etc/traefik/dynamic/orca-web-app.yaml
|
||||
# Atomic reload: write to .tmp + mv (gate C-10)
|
||||
http:
|
||||
routers:
|
||||
orca-web-app:
|
||||
rule: PathPrefix("/web-app")
|
||||
service: orca-web-app
|
||||
tls:
|
||||
certResolver: orca
|
||||
domains:
|
||||
- main: "cluster.orca.local"
|
||||
services:
|
||||
orca-web-app:
|
||||
loadBalancer:
|
||||
servers:
|
||||
- url: "unix:///run/orca/alloc-web-app-0/port-http.sock"
|
||||
- url: "unix:///run/orca/alloc-web-app-1/port-http.sock"
|
||||
- url: "unix:///run/orca/alloc-web-app-2/port-http.sock"
|
||||
healthCheck:
|
||||
path: /healthz
|
||||
interval: 5s
|
||||
timeout: 1s
|
||||
@@ -0,0 +1,51 @@
|
||||
---
|
||||
kind: Service
|
||||
name: web-app
|
||||
count: 3
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /bin/sleep 3600
|
||||
ports:
|
||||
- name: http
|
||||
port: 8080
|
||||
restart:
|
||||
mode: service
|
||||
attempts: 5
|
||||
delay: 2s
|
||||
update:
|
||||
strategy: rolling
|
||||
max_parallel: 1
|
||||
min_healthy_time: 10s
|
||||
healthy_deadline: 2m
|
||||
service:
|
||||
name: web-app
|
||||
port: 8080
|
||||
health:
|
||||
check_type: http
|
||||
interval: 5s
|
||||
timeout: 1s
|
||||
unhealthy_threshold: 2
|
||||
constraints:
|
||||
- node.role == "web"
|
||||
affinity:
|
||||
- target: zone == "a"
|
||||
weight: 80
|
||||
lifecycle:
|
||||
post_start:
|
||||
- /bin/sh -c 'echo cache warmed'
|
||||
pre_stop:
|
||||
- /bin/sh -c 'sleep 5'
|
||||
- /bin/sh -c 'echo draining web-app'
|
||||
---
|
||||
# Web App
|
||||
|
||||
Frontend web application serving HTTP on port 8080 via Unix socket.
|
||||
Three replicas with rolling updates, anti-affinity for zone spreading,
|
||||
and lifecycle hooks for cache warm-up and graceful drain.
|
||||
|
||||
> **Production substitution**: this example uses `/bin/sh -c 'echo ...
|
||||
> sleep 3600'` so it runs out-of-the-box on any Linux machine. In a
|
||||
> real deployment, replace the `runtime.command` with your actual
|
||||
> binary, e.g. `/usr/bin/httpd -f /etc/orca/web-app/httpd.conf`, and
|
||||
> replace the lifecycle hooks with your real scripts
|
||||
> (`/usr/local/bin/warm-cache.sh`, `/usr/local/bin/drain.sh`).
|
||||
@@ -0,0 +1,27 @@
|
||||
---
|
||||
kind: Job
|
||||
name: worker
|
||||
runtime:
|
||||
one_of: process
|
||||
command: /bin/echo worker processing batch
|
||||
timeout: 300s
|
||||
env:
|
||||
QUEUE_URL: unix:///run/orca/alloc-worker/queue.sock
|
||||
BATCH_SIZE: "100"
|
||||
LOG_LEVEL: debug
|
||||
lifecycle:
|
||||
post_start:
|
||||
- /bin/sh -c 'echo worker registered'
|
||||
pre_stop:
|
||||
- /bin/sh -c 'echo draining worker queue'
|
||||
---
|
||||
# Worker
|
||||
|
||||
One-shot batch worker that processes items from a queue. Runs once,
|
||||
exits on completion or after 300s timeout. Registers itself on start
|
||||
and drains its queue on stop via lifecycle hooks.
|
||||
|
||||
> **Production substitution**: replace `runtime.command` with your
|
||||
> actual worker binary, e.g. `/usr/bin/python3 /opt/orca/jobs/worker.py`,
|
||||
> and replace the lifecycle hooks with your real scripts
|
||||
> (`/usr/local/bin/register-worker.sh`, `/usr/local/bin/drain-queue.sh`).
|
||||
@@ -1,64 +1,73 @@
|
||||
// Package certpaths centralizes the on-disk locations of the CA and
|
||||
// server cert/key files. The CLI layer, the security layer, and the
|
||||
// doctor layer all need to agree on these paths, so they're factored
|
||||
// into their own package to avoid import cycles (cli <-> doctor).
|
||||
// Package certpaths is the v0.8 path shim. It returns v0.8 flat-layout
|
||||
// paths for backward compatibility during the v0.9 dual-write window
|
||||
// (REQ-090). The v0.9 paths package (internal/paths) returns the new
|
||||
// multi-namespace layout (R-002).
|
||||
//
|
||||
// certpaths will be deleted after the v0.10-P14 migration. New code
|
||||
// should use internal/paths, NOT certpaths.
|
||||
//
|
||||
// Migration notes (per v0.10-P14):
|
||||
// - CA cert/key, server cert/key, SSH key/pub, known_hosts currently
|
||||
// live at the flat Root() location. The v0.9 internal/paths package
|
||||
// returns the new ClusterDir()/... locations; certpaths keeps the
|
||||
// v0.8 flat locations until the CA migration moves them.
|
||||
// - DBPath keeps returning Root()/orca.db (v0.8 location). The new
|
||||
// paths.NSDb("_defaults") returns Root()/_defaults/db/orca.db; the DB
|
||||
// moves in v0.10-P14.
|
||||
package certpaths
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/paths"
|
||||
)
|
||||
|
||||
const (
|
||||
defaultCADir = ".orca"
|
||||
caCertFilename = "ca.crt"
|
||||
caKeyFilename = "ca.key"
|
||||
)
|
||||
// Dir returns the v0.8 flat root directory. Delegates to paths.Root()
|
||||
// (which honors $ORCA_HOME, else ~/.orca). v0.8 callers expect the CA
|
||||
// and DB to live directly under this directory; that does not change
|
||||
// until the v0.10-P14 migration.
|
||||
func Dir() string { return paths.Root() }
|
||||
|
||||
// Dir returns the directory the local CA lives in. Honors $ORCA_HOME
|
||||
// for testability; otherwise defaults to ~/.orca.
|
||||
func Dir() string {
|
||||
if p := os.Getenv("ORCA_HOME"); p != "" {
|
||||
return p
|
||||
}
|
||||
home, _ := os.UserHomeDir()
|
||||
return filepath.Join(home, defaultCADir)
|
||||
}
|
||||
// CACertPath returns the v0.8 CA cert path: Dir()/ca.crt.
|
||||
// The v0.9 location is paths.CACertPath() = ClusterDir()/ca.crt; certpaths
|
||||
// keeps the v0.8 flat location until the CA migration in v0.10-P14.
|
||||
func CACertPath() string { return filepath.Join(paths.Root(), "ca.crt") }
|
||||
|
||||
// CACertPath returns the path to ca.crt.
|
||||
func CACertPath() string { return filepath.Join(Dir(), caCertFilename) }
|
||||
// CAKeyPath returns the v0.8 CA key path: Dir()/ca.key.
|
||||
// See CACertPath for migration notes.
|
||||
func CAKeyPath() string { return filepath.Join(paths.Root(), "ca.key") }
|
||||
|
||||
// CAKeyPath returns the path to ca.key.
|
||||
func CAKeyPath() string { return filepath.Join(Dir(), caKeyFilename) }
|
||||
// ServerCertPath returns the v0.8 server cert path: Dir()/server.crt.
|
||||
// See CACertPath for migration notes.
|
||||
func ServerCertPath() string { return filepath.Join(paths.Root(), "server.crt") }
|
||||
|
||||
// ServerCertPath returns the path to server.crt.
|
||||
func ServerCertPath() string { return filepath.Join(Dir(), "server.crt") }
|
||||
|
||||
// ServerKeyPath returns the path to server.key.
|
||||
func ServerKeyPath() string { return filepath.Join(Dir(), "server.key") }
|
||||
// ServerKeyPath returns the v0.8 server key path: Dir()/server.key.
|
||||
// See CACertPath for migration notes.
|
||||
func ServerKeyPath() string { return filepath.Join(paths.Root(), "server.key") }
|
||||
|
||||
// DBPath returns the path to the orca SQLite database. Honors $ORCA_DB
|
||||
// for testability and explicit override; otherwise defaults to
|
||||
// ~/.orca/orca.db under the same Dir() as the cert files.
|
||||
// for testability and explicit override; otherwise defaults to the v0.8
|
||||
// flat location Dir()/orca.db. The v0.9 location is
|
||||
// paths.NSDb(paths.DefaultNamespace()) = Root()/_defaults/db/orca.db;
|
||||
// certpaths keeps the v0.8 flat location until the DB move in v0.10-P14.
|
||||
func DBPath() string {
|
||||
if p := os.Getenv("ORCA_DB"); p != "" {
|
||||
return p
|
||||
}
|
||||
return filepath.Join(Dir(), "orca.db")
|
||||
return filepath.Join(paths.Root(), "orca.db")
|
||||
}
|
||||
|
||||
// SSHKeyPath returns the path to the orca SSH private key (Ed25519,
|
||||
// D-037). Used by `orca node join --type proxmox` to authenticate
|
||||
// to remote Proxmox hosts after the initial password-based bootstrap.
|
||||
// File mode 0600 (enforced by security.WriteKey).
|
||||
func SSHKeyPath() string { return filepath.Join(Dir(), "orca_ssh_key") }
|
||||
// SSHKeyPath returns the v0.8 SSH private key path: Dir()/orca_ssh_key.
|
||||
// The v0.9 location is paths.SSHKeyPath() = ClusterDir()/orca_ssh_key;
|
||||
// certpaths keeps the v0.8 flat location until the migration.
|
||||
func SSHKeyPath() string { return filepath.Join(paths.Root(), "orca_ssh_key") }
|
||||
|
||||
// SSHPubPath returns the path to the orca SSH public key (authorized_keys
|
||||
// format). Deployed to remote Proxmox hosts during `orca node join`.
|
||||
// File mode 0644 (enforced by security.WriteCert).
|
||||
func SSHPubPath() string { return filepath.Join(Dir(), "orca_ssh_key.pub") }
|
||||
// SSHPubPath returns the v0.8 SSH public key path: Dir()/orca_ssh_key.pub.
|
||||
// See SSHKeyPath for migration notes.
|
||||
func SSHPubPath() string { return filepath.Join(paths.Root(), "orca_ssh_key.pub") }
|
||||
|
||||
// KnownHostsPath returns the path to the SSH known_hosts file used for
|
||||
// TOFU host-key pinning (D-035). Captured on first connect, verified
|
||||
// on all subsequent connects via golang.org/x/crypto/ssh/knownhosts.
|
||||
func KnownHostsPath() string { return filepath.Join(Dir(), "known_hosts") }
|
||||
// KnownHostsPath returns the v0.8 known_hosts path: Dir()/known_hosts.
|
||||
// The v0.9 location is paths.KnownHostsPath() = ClusterDir()/known_hosts;
|
||||
// certpaths keeps the v0.8 flat location until the migration.
|
||||
func KnownHostsPath() string { return filepath.Join(paths.Root(), "known_hosts") }
|
||||
|
||||
@@ -3,15 +3,17 @@ package certpaths
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"runtime"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/paths"
|
||||
)
|
||||
|
||||
const defaultHomeSubdir = ".orca"
|
||||
|
||||
func TestPaths_HonorORCAHOME(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
// Ensure ORCA_DB doesn't leak from the environment / prior tests.
|
||||
t.Setenv("ORCA_DB", "")
|
||||
|
||||
cases := []struct {
|
||||
@@ -36,17 +38,26 @@ func TestPaths_HonorORCAHOME(t *testing.T) {
|
||||
})
|
||||
}
|
||||
|
||||
// DBPath defaults to $ORCA_HOME/orca.db.
|
||||
if got, want := DBPath(), filepath.Join(dir, "orca.db"); got != want {
|
||||
t.Errorf("DBPath = %q, want %q", got, want)
|
||||
}
|
||||
|
||||
// Dir() returns ORCA_HOME verbatim.
|
||||
if got, want := Dir(), dir; got != want {
|
||||
t.Errorf("Dir = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestShim_DelegatesDirToPaths(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
if got, want := Dir(), paths.Root(); got != want {
|
||||
t.Errorf("Dir() = %q, paths.Root() = %q (shim must delegate)", got, want)
|
||||
}
|
||||
if got, want := Dir(), dir; got != want {
|
||||
t.Errorf("Dir() = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDBPath_OrcaDBOverride(t *testing.T) {
|
||||
home := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", home)
|
||||
@@ -70,19 +81,14 @@ func TestDBPath_OrcaDBEmptyStringFallsBackToHome(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestDir_DefaultHomeFallback(t *testing.T) {
|
||||
// Unset ORCA_HOME so Dir() falls back to ~/.orca.
|
||||
// We can't reliably mutate the real HOME in a portable way, so just
|
||||
// assert that the returned path ends with the default subdir on the
|
||||
// current OS and is absolute.
|
||||
os.Unsetenv("ORCA_HOME")
|
||||
// Also clear ORCA_DB so DBPath's fallback to Dir() is exercised.
|
||||
os.Unsetenv("ORCA_DB")
|
||||
|
||||
home, err := os.UserHomeDir()
|
||||
if err != nil {
|
||||
t.Skipf("os.UserHomeDir: %v (cannot verify default fallback)", err)
|
||||
}
|
||||
want := filepath.Join(home, defaultCADir)
|
||||
want := filepath.Join(home, defaultHomeSubdir)
|
||||
if got := Dir(); got != want {
|
||||
t.Errorf("Dir() default = %q, want %q", got, want)
|
||||
}
|
||||
@@ -92,25 +98,22 @@ func TestDir_DefaultHomeFallback(t *testing.T) {
|
||||
}
|
||||
|
||||
func TestDir_ORCAHOMEEmptyFallsBack(t *testing.T) {
|
||||
// Empty string ORCA_HOME is treated as unset → ~/.orca fallback.
|
||||
t.Setenv("ORCA_HOME", "")
|
||||
home, err := os.UserHomeDir()
|
||||
if err != nil {
|
||||
t.Skipf("os.UserHomeDir: %v", err)
|
||||
}
|
||||
want := filepath.Join(home, defaultCADir)
|
||||
want := filepath.Join(home, defaultHomeSubdir)
|
||||
if got := Dir(); got != want {
|
||||
t.Errorf("Dir() with empty ORCA_HOME = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDir_ORCAHOMERelativePath(t *testing.T) {
|
||||
// A relative ORCA_HOME is honored verbatim (no cleaning/absolutizing).
|
||||
t.Setenv("ORCA_HOME", "relative/orca/home")
|
||||
if got, want := Dir(), "relative/orca/home"; got != want {
|
||||
t.Errorf("Dir() relative = %q, want %q", got, want)
|
||||
}
|
||||
// CACertPath joins the relative dir with ca.crt using filepath.Join.
|
||||
if got, want := CACertPath(), filepath.Join("relative/orca/home", "ca.crt"); got != want {
|
||||
t.Errorf("CACertPath relative = %q, want %q", got, want)
|
||||
}
|
||||
@@ -121,7 +124,6 @@ func TestAllPaths_AreConsistentWithDir(t *testing.T) {
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
t.Setenv("ORCA_DB", "")
|
||||
|
||||
// Every *Path() must live under Dir() except DBPath which also does.
|
||||
base := Dir()
|
||||
for _, p := range []string{
|
||||
CACertPath(), CAKeyPath(),
|
||||
@@ -135,6 +137,38 @@ func TestAllPaths_AreConsistentWithDir(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestShim_ReturnsV08FlatPaths(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
t.Setenv("ORCA_DB", "")
|
||||
|
||||
root := paths.Root()
|
||||
if got, want := CACertPath(), filepath.Join(root, "ca.crt"); got != want {
|
||||
t.Errorf("CACertPath = %q, want v0.8 flat %q", got, want)
|
||||
}
|
||||
if got, want := CAKeyPath(), filepath.Join(root, "ca.key"); got != want {
|
||||
t.Errorf("CAKeyPath = %q, want v0.8 flat %q", got, want)
|
||||
}
|
||||
if got, want := ServerCertPath(), filepath.Join(root, "server.crt"); got != want {
|
||||
t.Errorf("ServerCertPath = %q, want v0.8 flat %q", got, want)
|
||||
}
|
||||
if got, want := ServerKeyPath(), filepath.Join(root, "server.key"); got != want {
|
||||
t.Errorf("ServerKeyPath = %q, want v0.8 flat %q", got, want)
|
||||
}
|
||||
if got, want := SSHKeyPath(), filepath.Join(root, "orca_ssh_key"); got != want {
|
||||
t.Errorf("SSHKeyPath = %q, want v0.8 flat %q", got, want)
|
||||
}
|
||||
if got, want := SSHPubPath(), filepath.Join(root, "orca_ssh_key.pub"); got != want {
|
||||
t.Errorf("SSHPubPath = %q, want v0.8 flat %q", got, want)
|
||||
}
|
||||
if got, want := KnownHostsPath(), filepath.Join(root, "known_hosts"); got != want {
|
||||
t.Errorf("KnownHostsPath = %q, want v0.8 flat %q", got, want)
|
||||
}
|
||||
if got, want := DBPath(), filepath.Join(root, "orca.db"); got != want {
|
||||
t.Errorf("DBPath = %q, want v0.8 flat %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSSHPaths_Filenames(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
@@ -148,10 +182,3 @@ func TestSSHPaths_Filenames(t *testing.T) {
|
||||
t.Errorf("KnownHostsPath base = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func init() {
|
||||
// On Windows the default home subdir is still ".orca"; the test for
|
||||
// default fallback uses os.UserHomeDir which is platform-aware. This
|
||||
// guard keeps the suite from running a meaningless check on plan9.
|
||||
_ = runtime.GOOS
|
||||
}
|
||||
|
||||
+43
-3
@@ -7,6 +7,7 @@ import (
|
||||
"fmt"
|
||||
"os"
|
||||
"os/signal"
|
||||
"strings"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
@@ -74,7 +75,7 @@ var jobRunCmd = &cobra.Command{
|
||||
peers := engine.NewPeerRegistry()
|
||||
dispatcher := engine.NewDispatcher(newLogger(), store.NewCapacityRepo(db), peers, exec)
|
||||
specBytes, _ := json.Marshal(map[string]any{
|
||||
"name": spec.Job.Name,
|
||||
"name": spec.Name,
|
||||
"command": "/bin/true", // placeholder; full HCL dispatch lands in a later phase
|
||||
})
|
||||
jobID, nodeID, err := dispatcher.Submit(ctx, runTarget, specBytes, runIDKey)
|
||||
@@ -93,11 +94,11 @@ var jobRunCmd = &cobra.Command{
|
||||
|
||||
job := &model.Job{
|
||||
ID: uuid.NewString(),
|
||||
Name: spec.Job.Name,
|
||||
Name: spec.Name,
|
||||
Spec: args[0],
|
||||
Status: model.JobStatusPending,
|
||||
}
|
||||
if err := exec.Run(ctx, job, toTaskSpecs(spec.Tasks)); err != nil {
|
||||
if err := exec.Run(ctx, job, workloadToTaskSpecs(spec)); err != nil {
|
||||
if jsonOutput {
|
||||
_ = printJSON(map[string]any{"id": job.ID, "status": "failed", "error": err.Error()})
|
||||
return err
|
||||
@@ -330,3 +331,42 @@ func toTaskSpecs(in []jobspec.TaskSpec) []engine.TaskSpec {
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// workloadToTaskSpecs converts a *WorkloadSpec into the engine.TaskSpec
|
||||
// slice consumed by the executor. For the HCL adapter path the runtime
|
||||
// block carries the legacy task[0].Command; for the Markdown path the
|
||||
// runtime block is the canonical runtime abstraction (P07 will expand
|
||||
// this). When Runtime is nil we emit a single no-op task to preserve
|
||||
// the legacy "at least one task" invariant.
|
||||
//
|
||||
// The runtime command string is split into binary + args via
|
||||
// splitCommand so that exec.Command receives the binary path and the
|
||||
// args as separate elements. Without this split, a command like
|
||||
// "/usr/bin/httpd -f /etc/orca/web-app/httpd.conf" is treated as a
|
||||
// single file path and fork/exec fails with "no such file or directory".
|
||||
func workloadToTaskSpecs(spec *jobspec.WorkloadSpec) []engine.TaskSpec {
|
||||
if spec == nil {
|
||||
return nil
|
||||
}
|
||||
if spec.Runtime == nil {
|
||||
return []engine.TaskSpec{{Name: spec.Name, Command: "/bin/true"}}
|
||||
}
|
||||
bin, args := splitCommand(spec.Runtime.Command)
|
||||
return []engine.TaskSpec{{
|
||||
Name: spec.Name,
|
||||
Command: bin,
|
||||
Args: args,
|
||||
}}
|
||||
}
|
||||
|
||||
// splitCommand splits a command string into binary + args using
|
||||
// strings.Fields (handles multiple spaces/tabs). If the string is empty
|
||||
// or all-whitespace, returns ("/bin/true", nil) so the executor still
|
||||
// has a valid binary to run.
|
||||
func splitCommand(s string) (string, []string) {
|
||||
parts := strings.Fields(s)
|
||||
if len(parts) == 0 {
|
||||
return "/bin/true", nil
|
||||
}
|
||||
return parts[0], parts[1:]
|
||||
}
|
||||
|
||||
@@ -34,6 +34,7 @@ func resetCommandFlags() {
|
||||
stopID, runTarget, runIDKey, jobWatch = "", "", "", false
|
||||
capSetCPU, capSetMem, capSetDisk, capNodeID = 0, 0, 0, ""
|
||||
auditLimit = 50
|
||||
resetNSFlags()
|
||||
}
|
||||
|
||||
func TestNamespaceDefaultsToUserHome(t *testing.T) {
|
||||
|
||||
@@ -0,0 +1,338 @@
|
||||
// Package cli: ns.go implements the `orca ns` subcommand family
|
||||
// (REQ-082, D-176). Subcommands:
|
||||
//
|
||||
// orca ns list — list all namespaces under ORCA_HOME
|
||||
// orca ns create <name> — create a namespace dir + ns.md
|
||||
// orca ns delete <name> — remove an empty namespace dir
|
||||
// orca ns inspect <name> — print effective chain + merged env
|
||||
// orca ns validate <name> — cycle + missing-parent + schema checks
|
||||
//
|
||||
// All subcommands honor $ORCA_HOME via internal/paths. The inheritance
|
||||
// resolver (internal/ns) is a pure function shared by inspect + validate.
|
||||
package cli
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
|
||||
"github.com/spf13/cobra"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/ns"
|
||||
"git.cloudinit.dev/coreci/orca/internal/paths"
|
||||
)
|
||||
|
||||
var nsCmd = &cobra.Command{
|
||||
Use: "ns",
|
||||
Short: "Manage orca namespaces",
|
||||
Long: `Manage orca namespaces under ORCA_HOME (R-002).
|
||||
|
||||
Each namespace is a directory with ns.md, .env, .env.secrets, db/,
|
||||
jobs/, alloc/. The implicit root namespace _defaults always exists
|
||||
(D-159); every namespace inherits from _defaults (D-185) and cannot
|
||||
opt out (D-187).`,
|
||||
}
|
||||
|
||||
var (
|
||||
nsCreateParent string
|
||||
nsCreateInheritsEnv bool
|
||||
nsCreateInheritsSecret bool
|
||||
)
|
||||
|
||||
var nsListCmd = &cobra.Command{
|
||||
Use: "list",
|
||||
Short: "List all namespaces under ORCA_HOME",
|
||||
Long: `List all namespaces under ORCA_HOME (directories containing ns.md, plus the implicit _defaults).`,
|
||||
Args: cobra.NoArgs,
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
root := paths.Root()
|
||||
entries, err := os.ReadDir(root)
|
||||
if err != nil {
|
||||
return fmt.Errorf("read ORCA_HOME %s: %w", root, err)
|
||||
}
|
||||
type nsRow struct {
|
||||
Name string `json:"name"`
|
||||
Path string `json:"path"`
|
||||
Default bool `json:"default"`
|
||||
}
|
||||
var rows []nsRow
|
||||
for _, ent := range entries {
|
||||
if !ent.IsDir() {
|
||||
continue
|
||||
}
|
||||
if ent.Name() == "cluster" {
|
||||
continue
|
||||
}
|
||||
nsMd := filepath.Join(root, ent.Name(), "ns.md")
|
||||
if _, err := os.Stat(nsMd); err != nil {
|
||||
continue
|
||||
}
|
||||
rows = append(rows, nsRow{
|
||||
Name: ent.Name(),
|
||||
Path: filepath.Join(root, ent.Name()),
|
||||
Default: ent.Name() == paths.DefaultNamespace(),
|
||||
})
|
||||
}
|
||||
sort.Slice(rows, func(i, j int) bool {
|
||||
if rows[i].Name == paths.DefaultNamespace() {
|
||||
return true
|
||||
}
|
||||
if rows[j].Name == paths.DefaultNamespace() {
|
||||
return false
|
||||
}
|
||||
return rows[i].Name < rows[j].Name
|
||||
})
|
||||
if jsonOutput {
|
||||
return printJSON(rows)
|
||||
}
|
||||
if len(rows) == 0 {
|
||||
fmt.Fprintln(cmd.OutOrStdout(), "No namespaces found. Run 'orca init' first.")
|
||||
return nil
|
||||
}
|
||||
fmt.Fprintf(cmd.OutOrStdout(), "%-20s %-10s %s\n", "NAME", "DEFAULT", "PATH")
|
||||
for _, r := range rows {
|
||||
def := ""
|
||||
if r.Default {
|
||||
def = "*"
|
||||
}
|
||||
fmt.Fprintf(cmd.OutOrStdout(), "%-20s %-10s %s\n", r.Name, def, r.Path)
|
||||
}
|
||||
return nil
|
||||
},
|
||||
}
|
||||
|
||||
var nsCreateCmd = &cobra.Command{
|
||||
Use: "create <name>",
|
||||
Short: "Create a namespace directory + ns.md",
|
||||
Long: `Create a namespace under ORCA_HOME. Builds the dir structure
|
||||
(db/, jobs/, alloc/) and writes ns.md frontmatter. --parent may be
|
||||
repeated to declare inheritance; _defaults is always appended last.`,
|
||||
Args: cobra.ExactArgs(1),
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
name := args[0]
|
||||
if name == paths.DefaultNamespace() {
|
||||
return fmt.Errorf("cannot create the implicit root namespace %q with `ns create` (it is auto-managed)", name)
|
||||
}
|
||||
if name == "cluster" {
|
||||
return fmt.Errorf("name %q is reserved for the cluster-wide dir", name)
|
||||
}
|
||||
if nsCreateParent == "" {
|
||||
nsCreateParent = paths.DefaultNamespace()
|
||||
}
|
||||
nsDir := paths.NamespaceDir(name)
|
||||
if _, err := os.Stat(nsDir); err == nil {
|
||||
if _, statErr := os.Stat(paths.NSMd(name)); statErr == nil {
|
||||
return fmt.Errorf("namespace %q already exists at %s", name, nsDir)
|
||||
}
|
||||
}
|
||||
for _, sub := range []string{"db", "jobs", "alloc"} {
|
||||
if err := os.MkdirAll(filepath.Join(nsDir, sub), 0o755); err != nil {
|
||||
return fmt.Errorf("create %s/%s: %w", nsDir, sub, err)
|
||||
}
|
||||
}
|
||||
parents := []string{nsCreateParent}
|
||||
if nsCreateParent == paths.DefaultNamespace() {
|
||||
// Explicit _defaults listing is allowed (de-duped silently).
|
||||
}
|
||||
body := renderNSMd(name, parents, nsCreateInheritsEnv, nsCreateInheritsSecret)
|
||||
if err := os.WriteFile(paths.NSMd(name), []byte(body), 0o644); err != nil {
|
||||
return fmt.Errorf("write ns.md: %w", err)
|
||||
}
|
||||
if jsonOutput {
|
||||
return printJSON(map[string]any{
|
||||
"name": name,
|
||||
"path": nsDir,
|
||||
"parents": parents,
|
||||
"ns_md": paths.NSMd(name),
|
||||
"inherits_env": nsCreateInheritsEnv,
|
||||
"inherits_secrets": nsCreateInheritsSecret,
|
||||
})
|
||||
}
|
||||
fmt.Fprintf(cmd.OutOrStdout(), "✓ Namespace created: %s (%s)\n", name, nsDir)
|
||||
return nil
|
||||
},
|
||||
}
|
||||
|
||||
var nsDeleteCmd = &cobra.Command{
|
||||
Use: "delete <name>",
|
||||
Short: "Remove an empty namespace directory",
|
||||
Long: `Remove a namespace directory. Refuses if jobs/ or alloc/
|
||||
contain any files (non-empty namespace). The implicit root _defaults
|
||||
cannot be deleted.`,
|
||||
Args: cobra.ExactArgs(1),
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
name := args[0]
|
||||
if name == paths.DefaultNamespace() {
|
||||
return fmt.Errorf("cannot delete the implicit root namespace %q", name)
|
||||
}
|
||||
nsDir := paths.NamespaceDir(name)
|
||||
if _, err := os.Stat(nsDir); err != nil {
|
||||
return fmt.Errorf("namespace %q not found: %w", name, err)
|
||||
}
|
||||
for _, sub := range []string{"jobs", "alloc"} {
|
||||
dir := filepath.Join(nsDir, sub)
|
||||
if err := dirNonEmpty(dir); err != nil {
|
||||
return fmt.Errorf("refusing to delete %q: %s is non-empty (%w); clear it first", name, sub, err)
|
||||
}
|
||||
}
|
||||
if err := os.RemoveAll(nsDir); err != nil {
|
||||
return fmt.Errorf("delete %s: %w", nsDir, err)
|
||||
}
|
||||
if jsonOutput {
|
||||
return printJSON(map[string]string{"name": name, "deleted": nsDir})
|
||||
}
|
||||
fmt.Fprintf(cmd.OutOrStdout(), "✓ Namespace deleted: %s (%s)\n", name, nsDir)
|
||||
return nil
|
||||
},
|
||||
}
|
||||
|
||||
var nsInspectCmd = &cobra.Command{
|
||||
Use: "inspect <name>",
|
||||
Short: "Print the effective chain, merged env, and constraints",
|
||||
Long: `Resolve a namespace's inheritance chain and print the merged env and unioned constraints (uses the resolver).`,
|
||||
Args: cobra.ExactArgs(1),
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
name := args[0]
|
||||
root := paths.Root()
|
||||
cfgs, err := ns.ParseNSMdDir(root)
|
||||
if err != nil {
|
||||
return fmt.Errorf("load namespaces: %w", err)
|
||||
}
|
||||
if _, ok := cfgs[name]; !ok {
|
||||
return fmt.Errorf("namespace %q not found under %s", name, root)
|
||||
}
|
||||
resolved, err := ns.Resolve(cfgs)
|
||||
if err != nil {
|
||||
return fmt.Errorf("resolve: %w", err)
|
||||
}
|
||||
r := resolved[name]
|
||||
if r == nil {
|
||||
return fmt.Errorf("namespace %q resolved to nil", name)
|
||||
}
|
||||
if jsonOutput {
|
||||
return printJSON(map[string]any{
|
||||
"name": r.Name,
|
||||
"chain": r.Chain,
|
||||
"env": r.Env,
|
||||
"constraints": r.Constraints,
|
||||
})
|
||||
}
|
||||
fmt.Fprintf(cmd.OutOrStdout(), "Namespace: %s\n", r.Name)
|
||||
fmt.Fprintf(cmd.OutOrStdout(), "Chain: %s\n", strings.Join(r.Chain, " -> "))
|
||||
fmt.Fprintln(cmd.OutOrStdout(), "Env:")
|
||||
keys := sortedKeys(r.Env)
|
||||
for _, k := range keys {
|
||||
fmt.Fprintf(cmd.OutOrStdout(), " %s = %s\n", k, r.Env[k])
|
||||
}
|
||||
fmt.Fprintln(cmd.OutOrStdout(), "Constraints:")
|
||||
if len(r.Constraints) == 0 {
|
||||
fmt.Fprintln(cmd.OutOrStdout(), " (none)")
|
||||
} else {
|
||||
for _, c := range r.Constraints {
|
||||
fmt.Fprintf(cmd.OutOrStdout(), " - %s\n", c)
|
||||
}
|
||||
}
|
||||
return nil
|
||||
},
|
||||
}
|
||||
|
||||
var nsValidateCmd = &cobra.Command{
|
||||
Use: "validate <name>",
|
||||
Short: "Run cycle + missing-parent + schema checks on a namespace",
|
||||
Long: `Validate a namespace's inheritance chain and ns.md frontmatter.
|
||||
Exits 0 if valid, 1 on error. Runs over ALL namespaces under ORCA_HOME
|
||||
(parsing + resolving validates cycles and missing parents across the
|
||||
set).`,
|
||||
Args: cobra.ExactArgs(1),
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
name := args[0]
|
||||
root := paths.Root()
|
||||
cfgs, err := ns.ParseNSMdDir(root)
|
||||
if err != nil {
|
||||
return fmt.Errorf("load namespaces: %w", err)
|
||||
}
|
||||
if _, ok := cfgs[name]; !ok {
|
||||
return fmt.Errorf("namespace %q not found under %s", name, root)
|
||||
}
|
||||
resolved, err := ns.Resolve(cfgs)
|
||||
if err != nil {
|
||||
return fmt.Errorf("validate: %w", err)
|
||||
}
|
||||
r := resolved[name]
|
||||
if r == nil {
|
||||
return fmt.Errorf("namespace %q resolved to nil", name)
|
||||
}
|
||||
if jsonOutput {
|
||||
return printJSON(map[string]any{
|
||||
"name": r.Name,
|
||||
"valid": true,
|
||||
"chain": r.Chain,
|
||||
})
|
||||
}
|
||||
fmt.Fprintf(cmd.OutOrStdout(), "✓ %s valid\n chain: %s\n", name, strings.Join(r.Chain, " -> "))
|
||||
return nil
|
||||
},
|
||||
}
|
||||
|
||||
// renderNSMd writes a minimal ns.md frontmatter for `orca ns create`.
|
||||
func renderNSMd(name string, parents []string, inheritsEnv, inheritsSecrets bool) string {
|
||||
var b strings.Builder
|
||||
b.WriteString("---\n")
|
||||
b.WriteString("kind: Namespace\n")
|
||||
b.WriteString("name: ")
|
||||
b.WriteString(name)
|
||||
b.WriteString("\n")
|
||||
if len(parents) > 0 {
|
||||
quoted := make([]string, len(parents))
|
||||
for i, p := range parents {
|
||||
quoted[i] = fmt.Sprintf("%q", p)
|
||||
}
|
||||
b.WriteString("parents: [")
|
||||
b.WriteString(strings.Join(quoted, ", "))
|
||||
b.WriteString("]\n")
|
||||
}
|
||||
fmt.Fprintf(&b, "inherits_env: %t\n", inheritsEnv)
|
||||
fmt.Fprintf(&b, "inherits_secrets: %t\n", inheritsSecrets)
|
||||
b.WriteString("---\n")
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// dirNonEmpty returns an error wrapping the offending entry if dir
|
||||
// contains any entries.
|
||||
func dirNonEmpty(dir string) error {
|
||||
entries, err := os.ReadDir(dir)
|
||||
if err != nil {
|
||||
if os.IsNotExist(err) {
|
||||
return nil
|
||||
}
|
||||
return err
|
||||
}
|
||||
for _, e := range entries {
|
||||
return fmt.Errorf("contains %s", e.Name())
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func sortedKeys(m map[string]string) []string {
|
||||
out := make([]string, 0, len(m))
|
||||
for k := range m {
|
||||
out = append(out, k)
|
||||
}
|
||||
sort.Strings(out)
|
||||
return out
|
||||
}
|
||||
|
||||
func init() {
|
||||
nsCreateCmd.Flags().StringVar(&nsCreateParent, "parent", "", "parent namespace (default _defaults; the implicit root is always appended last)")
|
||||
nsCreateCmd.Flags().BoolVar(&nsCreateInheritsEnv, "inherits-env", true, "inherit env from parents (default true)")
|
||||
nsCreateCmd.Flags().BoolVar(&nsCreateInheritsSecret, "inherits-secrets", true, "inherit secrets from parents (default true)")
|
||||
|
||||
nsCmd.AddCommand(nsListCmd)
|
||||
nsCmd.AddCommand(nsCreateCmd)
|
||||
nsCmd.AddCommand(nsDeleteCmd)
|
||||
nsCmd.AddCommand(nsInspectCmd)
|
||||
nsCmd.AddCommand(nsValidateCmd)
|
||||
rootCmd.AddCommand(nsCmd)
|
||||
}
|
||||
@@ -0,0 +1,407 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"encoding/json"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/paths"
|
||||
)
|
||||
|
||||
// resetNSFlags zeroes the ns subcommand flag-bound vars so tests don't
|
||||
// leak state.
|
||||
func resetNSFlags() {
|
||||
nsCreateParent = ""
|
||||
nsCreateInheritsEnv = true
|
||||
nsCreateInheritsSecret = true
|
||||
}
|
||||
|
||||
func writeDefaultsNS(t *testing.T, root string) {
|
||||
t.Helper()
|
||||
nsDir := filepath.Join(root, "_defaults")
|
||||
if err := os.MkdirAll(nsDir, 0o755); err != nil {
|
||||
t.Fatalf("mkdir _defaults: %v", err)
|
||||
}
|
||||
body := "---\nkind: Namespace\nname: _defaults\ninherits_env: true\ninherits_secrets: true\n---\n# defaults\n"
|
||||
if err := os.WriteFile(filepath.Join(nsDir, "ns.md"), []byte(body), 0o644); err != nil {
|
||||
t.Fatalf("write _defaults ns.md: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func writeCustomNS(t *testing.T, root, name, parentsList string) {
|
||||
t.Helper()
|
||||
nsDir := filepath.Join(root, name)
|
||||
if err := os.MkdirAll(nsDir, 0o755); err != nil {
|
||||
t.Fatalf("mkdir %s: %v", name, err)
|
||||
}
|
||||
body := "---\nkind: Namespace\nname: " + name + "\n"
|
||||
if parentsList != "" {
|
||||
body += "parents: " + parentsList + "\n"
|
||||
}
|
||||
body += "inherits_env: true\ninherits_secrets: true\n---\n# " + name + "\n"
|
||||
if err := os.WriteFile(filepath.Join(nsDir, "ns.md"), []byte(body), 0o644); err != nil {
|
||||
t.Fatalf("write %s ns.md: %v", name, err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSListEmpty(t *testing.T) {
|
||||
t.Setenv("ORCA_HOME", t.TempDir())
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "list"})
|
||||
if err := rootCmd.Execute(); err != nil {
|
||||
t.Fatalf("ns list: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSListWithNamespaces(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
writeCustomNS(t, root, "prod", "")
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "list"})
|
||||
if err := rootCmd.Execute(); err != nil {
|
||||
t.Fatalf("ns list: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSCreateHappy(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "create", "prod"})
|
||||
if err := rootCmd.Execute(); err != nil {
|
||||
t.Fatalf("ns create: %v", err)
|
||||
}
|
||||
if _, err := os.Stat(filepath.Join(root, "prod", "ns.md")); err != nil {
|
||||
t.Fatalf("ns.md not created: %v", err)
|
||||
}
|
||||
for _, sub := range []string{"db", "jobs", "alloc"} {
|
||||
if _, err := os.Stat(filepath.Join(root, "prod", sub)); err != nil {
|
||||
t.Errorf("subdir %s not created: %v", sub, err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSCreateWithParent(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
writeCustomNS(t, root, "base", "")
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "create", "child", "--parent", "base"})
|
||||
if err := rootCmd.Execute(); err != nil {
|
||||
t.Fatalf("ns create: %v", err)
|
||||
}
|
||||
data, err := os.ReadFile(filepath.Join(root, "child", "ns.md"))
|
||||
if err != nil {
|
||||
t.Fatalf("read ns.md: %v", err)
|
||||
}
|
||||
if !strings.Contains(string(data), "parents: [") || !strings.Contains(string(data), "\"base\"") {
|
||||
t.Errorf("ns.md missing parents: %s", string(data))
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSCreateExisting(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
writeCustomNS(t, root, "prod", "")
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "create", "prod"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error creating existing namespace, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "already exists") {
|
||||
t.Errorf("error = %q, want contains 'already exists'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSCreateDefaultsRefused(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "create", "_defaults"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error creating _defaults, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSCreateClusterRefused(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "create", "cluster"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error creating cluster, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSDeleteHappy(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
writeCustomNS(t, root, "prod", "")
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "delete", "prod"})
|
||||
if err := rootCmd.Execute(); err != nil {
|
||||
t.Fatalf("ns delete: %v", err)
|
||||
}
|
||||
if _, err := os.Stat(filepath.Join(root, "prod")); !os.IsNotExist(err) {
|
||||
t.Errorf("prod dir still exists after delete")
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSDeleteDefaultsRefused(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "delete", "_defaults"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error deleting _defaults, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSDeleteNonEmpty(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
writeCustomNS(t, root, "prod", "")
|
||||
// put a job in jobs/
|
||||
if err := os.MkdirAll(filepath.Join(root, "prod", "jobs"), 0o755); err != nil {
|
||||
t.Fatalf("mkdir jobs: %v", err)
|
||||
}
|
||||
if err := os.WriteFile(filepath.Join(root, "prod", "jobs", "j1.md"), []byte("x"), 0o644); err != nil {
|
||||
t.Fatalf("write job: %v", err)
|
||||
}
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "delete", "prod"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error deleting non-empty namespace, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "refusing to delete") {
|
||||
t.Errorf("error = %q, want contains 'refusing to delete'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSDeleteMissing(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "delete", "ghost"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error deleting missing namespace, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSInspectHappy(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
writeCustomNS(t, root, "prod", "")
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "inspect", "prod"})
|
||||
if err := rootCmd.Execute(); err != nil {
|
||||
t.Fatalf("ns inspect: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSInspectJSON(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
writeCustomNS(t, root, "prod", "")
|
||||
|
||||
var buf bytes.Buffer
|
||||
rootCmd.SetOut(&buf)
|
||||
rootCmd.SetArgs([]string{"ns", "inspect", "prod", "--json"})
|
||||
if err := rootCmd.Execute(); err != nil {
|
||||
t.Fatalf("ns inspect --json: %v", err)
|
||||
}
|
||||
var result map[string]any
|
||||
if err := json.Unmarshal(bytes.TrimSpace(buf.Bytes()), &result); err != nil {
|
||||
t.Fatalf("unmarshal: %v\n%s", err, buf.String())
|
||||
}
|
||||
if result["name"] != "prod" {
|
||||
t.Errorf("name = %v, want prod", result["name"])
|
||||
}
|
||||
chain, _ := result["chain"].([]any)
|
||||
if len(chain) < 2 {
|
||||
t.Errorf("chain too short: %v", chain)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSInspectMissingNamespace(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "inspect", "ghost"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing namespace, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSValidateHappy(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
writeCustomNS(t, root, "prod", "")
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "validate", "prod"})
|
||||
if err := rootCmd.Execute(); err != nil {
|
||||
t.Fatalf("ns validate: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSValidateCycle(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
// a -> b, b -> a (cycle)
|
||||
writeCustomNS(t, root, "a", "[\"b\"]")
|
||||
writeCustomNS(t, root, "b", "[\"a\"]")
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "validate", "a"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected cycle error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "cycle") {
|
||||
t.Errorf("error = %q, want contains 'cycle'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSValidateMissingParent(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
writeCustomNS(t, root, "x", "[\"ghost\"]")
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "validate", "x"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected missing-parent error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "ghost") || !strings.Contains(err.Error(), "not found") {
|
||||
t.Errorf("error = %q, want contains ghost + not found", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSValidateMissingNamespace(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", root)
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
writeDefaultsNS(t, root)
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "validate", "ghost"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing namespace, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderNSMd(t *testing.T) {
|
||||
body := renderNSMd("foo", []string{"_defaults"}, true, false)
|
||||
if !strings.Contains(body, "kind: Namespace") {
|
||||
t.Errorf("missing kind: %s", body)
|
||||
}
|
||||
if !strings.Contains(body, "name: foo") {
|
||||
t.Errorf("missing name: %s", body)
|
||||
}
|
||||
if !strings.Contains(body, "inherits_env: true") {
|
||||
t.Errorf("missing inherits_env true: %s", body)
|
||||
}
|
||||
if !strings.Contains(body, "inherits_secrets: false") {
|
||||
t.Errorf("missing inherits_secrets false: %s", body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSRootRegistered(t *testing.T) {
|
||||
found := false
|
||||
for _, c := range rootCmd.Commands() {
|
||||
if c.Use == "ns" {
|
||||
found = true
|
||||
break
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
t.Errorf("ns command not registered on root")
|
||||
}
|
||||
// ensure subcommands present
|
||||
sub := map[string]bool{}
|
||||
for _, c := range rootCmd.Commands() {
|
||||
if c.Use == "ns" {
|
||||
for _, sc := range c.Commands() {
|
||||
sub[sc.Use] = true
|
||||
}
|
||||
}
|
||||
}
|
||||
for _, want := range []string{"list", "create <name>", "delete <name>", "inspect <name>", "validate <name>"} {
|
||||
if !sub[want] {
|
||||
t.Errorf("missing ns subcommand %q", want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestNSListNoORCAHOME(t *testing.T) {
|
||||
// ORCA_HOME points at a nonexistent dir; list should error.
|
||||
t.Setenv("ORCA_HOME", filepath.Join(t.TempDir(), "nope"))
|
||||
resetRootFlags(t)
|
||||
resetNSFlags()
|
||||
|
||||
rootCmd.SetArgs([]string{"ns", "list"})
|
||||
err := rootCmd.Execute()
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing ORCA_HOME, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
var _ = paths.DefaultNamespace // keep paths import alive
|
||||
@@ -0,0 +1,141 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
func TestSplitCommand(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
input string
|
||||
wantBin string
|
||||
wantArgs []string
|
||||
}{
|
||||
{
|
||||
name: "single binary",
|
||||
input: "/bin/true",
|
||||
wantBin: "/bin/true",
|
||||
wantArgs: nil,
|
||||
},
|
||||
{
|
||||
name: "binary with one arg",
|
||||
input: "/bin/echo hello",
|
||||
wantBin: "/bin/echo",
|
||||
wantArgs: []string{"hello"},
|
||||
},
|
||||
{
|
||||
name: "binary with multiple args",
|
||||
input: "/usr/bin/httpd -f /etc/orca/web-app/httpd.conf",
|
||||
wantBin: "/usr/bin/httpd",
|
||||
wantArgs: []string{"-f", "/etc/orca/web-app/httpd.conf"},
|
||||
},
|
||||
{
|
||||
name: "binary with sh -c and quoted string",
|
||||
input: "/bin/sh -c 'echo hello world'",
|
||||
wantBin: "/bin/sh",
|
||||
wantArgs: []string{"-c", "'echo", "hello", "world'"},
|
||||
},
|
||||
{
|
||||
name: "empty command falls back to /bin/true",
|
||||
input: "",
|
||||
wantBin: "/bin/true",
|
||||
wantArgs: nil,
|
||||
},
|
||||
{
|
||||
name: "all-whitespace command falls back to /bin/true",
|
||||
input: " \t ",
|
||||
wantBin: "/bin/true",
|
||||
wantArgs: nil,
|
||||
},
|
||||
{
|
||||
name: "multiple spaces between args",
|
||||
input: "/bin/echo hello world",
|
||||
wantBin: "/bin/echo",
|
||||
wantArgs: []string{"hello", "world"},
|
||||
},
|
||||
}
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
gotBin, gotArgs := splitCommand(tt.input)
|
||||
if gotBin != tt.wantBin {
|
||||
t.Errorf("splitCommand(%q) bin = %q, want %q", tt.input, gotBin, tt.wantBin)
|
||||
}
|
||||
if len(gotArgs) != len(tt.wantArgs) {
|
||||
t.Errorf("splitCommand(%q) args len = %d, want %d (got %v, want %v)",
|
||||
tt.input, len(gotArgs), len(tt.wantArgs), gotArgs, tt.wantArgs)
|
||||
return
|
||||
}
|
||||
for i, a := range gotArgs {
|
||||
if a != tt.wantArgs[i] {
|
||||
t.Errorf("splitCommand(%q) args[%d] = %q, want %q",
|
||||
tt.input, i, a, tt.wantArgs[i])
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestWorkloadToTaskSpecs_SplitsCommand(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Name: "web-app",
|
||||
Runtime: &jobspec.RuntimeBlock{
|
||||
OneOf: "process",
|
||||
Command: "/usr/bin/httpd -f /etc/orca/web-app/httpd.conf",
|
||||
},
|
||||
}
|
||||
tasks := workloadToTaskSpecs(spec)
|
||||
if len(tasks) != 1 {
|
||||
t.Fatalf("expected 1 task, got %d", len(tasks))
|
||||
}
|
||||
if tasks[0].Command != "/usr/bin/httpd" {
|
||||
t.Errorf("expected Command=/usr/bin/httpd, got %q", tasks[0].Command)
|
||||
}
|
||||
if len(tasks[0].Args) != 2 {
|
||||
t.Fatalf("expected 2 args, got %d (%v)", len(tasks[0].Args), tasks[0].Args)
|
||||
}
|
||||
if tasks[0].Args[0] != "-f" || tasks[0].Args[1] != "/etc/orca/web-app/httpd.conf" {
|
||||
t.Errorf("expected args [-f /etc/orca/web-app/httpd.conf], got %v", tasks[0].Args)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWorkloadToTaskSpecs_NilRuntimeUsesBinTrue(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Name: "noop",
|
||||
}
|
||||
tasks := workloadToTaskSpecs(spec)
|
||||
if len(tasks) != 1 {
|
||||
t.Fatalf("expected 1 task, got %d", len(tasks))
|
||||
}
|
||||
if tasks[0].Command != "/bin/true" {
|
||||
t.Errorf("expected Command=/bin/true, got %q", tasks[0].Command)
|
||||
}
|
||||
if len(tasks[0].Args) != 0 {
|
||||
t.Errorf("expected 0 args, got %d (%v)", len(tasks[0].Args), tasks[0].Args)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWorkloadToTaskSpecs_EmptyCommandUsesBinTrue(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Name: "empty",
|
||||
Runtime: &jobspec.RuntimeBlock{
|
||||
OneOf: "process",
|
||||
Command: "",
|
||||
},
|
||||
}
|
||||
tasks := workloadToTaskSpecs(spec)
|
||||
if len(tasks) != 1 {
|
||||
t.Fatalf("expected 1 task, got %d", len(tasks))
|
||||
}
|
||||
if tasks[0].Command != "/bin/true" {
|
||||
t.Errorf("expected Command=/bin/true, got %q", tasks[0].Command)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWorkloadToTaskSpecs_NilSpecReturnsNil(t *testing.T) {
|
||||
tasks := workloadToTaskSpecs(nil)
|
||||
if tasks != nil {
|
||||
t.Errorf("expected nil, got %v", tasks)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,86 @@
|
||||
// Package cluster holds cluster-wide invariants that are not owned
|
||||
// by a single subsystem. The first inhabitant is the lead-eligibility
|
||||
// rule R-003: the cluster lead is always a bare Linux node; Proxmox
|
||||
// nodes are permanently ineligible because their kernel is shared
|
||||
// with guest VMs/containers and a lead failure there takes down the
|
||||
// hypervisor too.
|
||||
//
|
||||
// The package is deliberately decoupled from the scheduler: it owns
|
||||
// its own minimal NodeInfo (Hostname + Kind) so it can be unit-tested
|
||||
// without pulling in the scheduler's capacity model. The scheduler's
|
||||
// scheduler.NodeInfo has a `Kind string` field with the same values
|
||||
// ("linux", "proxmox"); callers convert at the boundary.
|
||||
package cluster
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// NodeKind classifies a node for lead-eligibility purposes (R-003).
|
||||
// The string values match scheduler.NodeInfo.Kind and model.NodeKind
|
||||
// so callers can pass either representation through without mapping.
|
||||
type NodeKind string
|
||||
|
||||
const (
|
||||
// NodeKindLinux is a bare Linux node — lead-eligible (R-003).
|
||||
NodeKindLinux NodeKind = "linux"
|
||||
// NodeKindProxmox is a Proxmox VE host — permanently lead-
|
||||
// ineligible (R-003): the hypervisor kernel is shared with
|
||||
// guests, so a lead process there is a blast-radius hazard.
|
||||
NodeKindProxmox NodeKind = "proxmox"
|
||||
)
|
||||
|
||||
// ErrProxmoxNotLead is returned when a Proxmox node is proposed as
|
||||
// the new cluster lead (R-003).
|
||||
var ErrProxmoxNotLead = errors.New("Proxmox nodes cannot hold the cluster lead role (R-003)")
|
||||
|
||||
// ErrNodeNotRegistered is returned when the proposed lead is not in
|
||||
// the supplied node list at all.
|
||||
var ErrNodeNotRegistered = errors.New("cluster: proposed lead is not a registered node")
|
||||
|
||||
// NodeInfo is the minimal node projection the lead rules need. It is
|
||||
// intentionally smaller than scheduler.NodeInfo so this package has
|
||||
// no upstream dependency on the scheduler.
|
||||
type NodeInfo struct {
|
||||
Hostname string
|
||||
Kind NodeKind
|
||||
}
|
||||
|
||||
// IsLeadEligible reports whether a node of the given kind may hold
|
||||
// the cluster lead role (R-003). Linux nodes are eligible; Proxmox
|
||||
// nodes are permanently ineligible; any other kind (including the
|
||||
// empty string) is treated as ineligible.
|
||||
func IsLeadEligible(kind NodeKind) bool {
|
||||
return kind == NodeKindLinux
|
||||
}
|
||||
|
||||
// ValidateLeadRotation checks that newLead is a registered Linux node
|
||||
// and refuses Proxmox nodes with ErrProxmoxNotLead (R-003). It returns
|
||||
// ErrNodeNotRegistered when newLead is not in nodes at all. The check
|
||||
// is case-sensitive on hostname; node registries in Orca are
|
||||
// case-normalized at the store layer so this matches reality.
|
||||
func ValidateLeadRotation(newLead string, nodes []NodeInfo) error {
|
||||
for _, n := range nodes {
|
||||
if n.Hostname != newLead {
|
||||
continue
|
||||
}
|
||||
if n.Kind == NodeKindProxmox {
|
||||
return ErrProxmoxNotLead
|
||||
}
|
||||
if n.Kind == NodeKindLinux {
|
||||
return nil
|
||||
}
|
||||
// Registered but neither linux nor proxmox (e.g. "localhost"
|
||||
// auto-registered node, or a future kind). Treat unknown kinds
|
||||
// as ineligible rather than guessing.
|
||||
return fmt.Errorf("cluster: node %q has ineligible kind %q: %w", newLead, n.Kind, ErrProxmoxNotLead)
|
||||
}
|
||||
// Not found in the registry at all.
|
||||
return fmt.Errorf("cluster: node %q not found: %w", newLead, ErrNodeNotRegistered)
|
||||
}
|
||||
|
||||
// String renders a NodeKind for logs. It lowercases to match the
|
||||
// on-disk representation regardless of how the caller constructed it.
|
||||
func (k NodeKind) String() string { return strings.ToLower(string(k)) }
|
||||
@@ -0,0 +1,119 @@
|
||||
package cluster
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestIsLeadEligible(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
kind NodeKind
|
||||
want bool
|
||||
}{
|
||||
{"linux", NodeKindLinux, true},
|
||||
{"proxmox", NodeKindProxmox, false},
|
||||
{"empty", "", false},
|
||||
{"unknown", NodeKind("foo"), false},
|
||||
{"localhost", NodeKind("localhost"), false},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
if got := IsLeadEligible(tc.kind); got != tc.want {
|
||||
t.Errorf("IsLeadEligible(%q) = %v, want %v", tc.kind, got, tc.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateLeadRotation_LinuxOK(t *testing.T) {
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "n1", Kind: NodeKindLinux},
|
||||
{Hostname: "n2", Kind: NodeKindLinux},
|
||||
{Hostname: "pve1", Kind: NodeKindProxmox},
|
||||
}
|
||||
if err := ValidateLeadRotation("n2", nodes); err != nil {
|
||||
t.Errorf("ValidateLeadRotation(n2): err = %v, want nil", err)
|
||||
}
|
||||
if err := ValidateLeadRotation("n1", nodes); err != nil {
|
||||
t.Errorf("ValidateLeadRotation(n1): err = %v, want nil", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateLeadRotation_ProxmoxRefused(t *testing.T) {
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "n1", Kind: NodeKindLinux},
|
||||
{Hostname: "pve1", Kind: NodeKindProxmox},
|
||||
}
|
||||
err := ValidateLeadRotation("pve1", nodes)
|
||||
if err == nil {
|
||||
t.Fatal("ValidateLeadRotation(pve1): expected error, got nil")
|
||||
}
|
||||
if !errors.Is(err, ErrProxmoxNotLead) {
|
||||
t.Errorf("err = %v, want ErrProxmoxNotLead", err)
|
||||
}
|
||||
if got := err.Error(); got != "Proxmox nodes cannot hold the cluster lead role (R-003)" {
|
||||
t.Errorf("err message = %q, want R-003 text verbatim", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateLeadRotation_UnknownNode(t *testing.T) {
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "n1", Kind: NodeKindLinux},
|
||||
}
|
||||
err := ValidateLeadRotation("ghost", nodes)
|
||||
if err == nil {
|
||||
t.Fatal("ValidateLeadRotation(ghost): expected error, got nil")
|
||||
}
|
||||
if !errors.Is(err, ErrNodeNotRegistered) {
|
||||
t.Errorf("err = %v, want ErrNodeNotRegistered", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateLeadRotation_EmptyList(t *testing.T) {
|
||||
err := ValidateLeadRotation("anyone", nil)
|
||||
if err == nil {
|
||||
t.Fatal("ValidateLeadRotation on empty list: expected error, got nil")
|
||||
}
|
||||
if !errors.Is(err, ErrNodeNotRegistered) {
|
||||
t.Errorf("err = %v, want ErrNodeNotRegistered", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateLeadRotation_IneligibleKindRegistered(t *testing.T) {
|
||||
// A node registered with a kind that is neither linux nor
|
||||
// proxmox (e.g. the auto-registered "localhost" kind) is
|
||||
// rejected as ineligible, not as unregistered.
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "self", Kind: NodeKind("localhost")},
|
||||
}
|
||||
err := ValidateLeadRotation("self", nodes)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for localhost kind, got nil")
|
||||
}
|
||||
if !errors.Is(err, ErrProxmoxNotLead) {
|
||||
t.Errorf("err = %v, want wrapped ErrProxmoxNotLead (ineligible)", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidateLeadRotation_CaseSensitive(t *testing.T) {
|
||||
// Hostnames are case-normalized at the store layer; the rule
|
||||
// matches exactly. "N1" is NOT the same as "n1".
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "n1", Kind: NodeKindLinux},
|
||||
}
|
||||
if err := ValidateLeadRotation("N1", nodes); !errors.Is(err, ErrNodeNotRegistered) {
|
||||
t.Errorf("N1 (case mismatch): err = %v, want ErrNodeNotRegistered", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestNodeKindString(t *testing.T) {
|
||||
if got := NodeKindLinux.String(); got != "linux" {
|
||||
t.Errorf("Linux.String() = %q", got)
|
||||
}
|
||||
if got := NodeKindProxmox.String(); got != "proxmox" {
|
||||
t.Errorf("Proxmox.String() = %q", got)
|
||||
}
|
||||
// Uppercase constructor should lower-case.
|
||||
if got := NodeKind("PROXMOX").String(); got != "proxmox" {
|
||||
t.Errorf("PROXMOX.String() = %q, want proxmox", got)
|
||||
}
|
||||
}
|
||||
@@ -3,6 +3,7 @@ package config
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
"github.com/hashicorp/hcl/v2/hclsimple"
|
||||
)
|
||||
@@ -33,22 +34,82 @@ type Flags struct {
|
||||
|
||||
type Environ map[string]string
|
||||
|
||||
// Load is the config dispatcher (R-014). It tries each path in order,
|
||||
// skipping missing files, and dispatches to the appropriate loader
|
||||
// based on file extension: .hcl → LoadHCL (legacy, R-013),
|
||||
// .md → LoadMarkdown (new Markdown-frontmatter loader), and
|
||||
// .yaml/.yml → LoadMarkdown with an empty body. The first successfully
|
||||
// decoded file wins. If no path exists or decodes, a zero Config is
|
||||
// returned.
|
||||
//
|
||||
// The signature is preserved from the v0.8 single-loader API so
|
||||
// internal/cli/root.go requires no changes yet.
|
||||
func Load(paths ...string) (*Config, error) {
|
||||
for _, p := range paths {
|
||||
if _, err := os.Stat(p); err != nil {
|
||||
continue
|
||||
}
|
||||
cfg, err := loadByExtension(p)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return cfg, nil
|
||||
}
|
||||
return &Config{}, nil
|
||||
}
|
||||
|
||||
func loadByExtension(p string) (*Config, error) {
|
||||
ext := strings.ToLower(filepathExt(p))
|
||||
switch ext {
|
||||
case ".hcl":
|
||||
return LoadHCL(p)
|
||||
case ".md", ".markdown":
|
||||
return LoadMarkdown(p)
|
||||
case ".yaml", ".yml":
|
||||
return LoadMarkdownYAML(p)
|
||||
default:
|
||||
// Unknown extension: hclsimple.Decode rejects non-.hcl
|
||||
// suffixes, so for backward compat with the v0.8 single-loader
|
||||
// behavior (which assumed HCL), decode the file content as HCL
|
||||
// against a synthesized .hcl path.
|
||||
data, err := os.ReadFile(p)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read config %s: %w", p, err)
|
||||
}
|
||||
var cfg Config
|
||||
if err := hclsimple.Decode(p, data, nil, &cfg); err != nil {
|
||||
if err := hclsimple.Decode(p+".hcl", data, nil, &cfg); err != nil {
|
||||
return nil, fmt.Errorf("decode config %s: %w", p, err)
|
||||
}
|
||||
return &cfg, nil
|
||||
}
|
||||
return &Config{}, nil
|
||||
}
|
||||
|
||||
// filepathExt is a thin wrapper around filepath.Ext to keep the import
|
||||
// localized to the dispatcher. Returns the extension including the dot,
|
||||
// lowercased by the caller.
|
||||
func filepathExt(p string) string {
|
||||
i := strings.LastIndex(p, ".")
|
||||
if i < 0 {
|
||||
return ""
|
||||
}
|
||||
return p[i:]
|
||||
}
|
||||
|
||||
// LoadHCL decodes a legacy HCL config file (R-013).
|
||||
//
|
||||
// Deprecated: use LoadMarkdown or the dispatcher. HCL is legacy per
|
||||
// R-013. Retained for the v0.9 dual-write window (REQ-090); new
|
||||
// deployments should author config.md with YAML frontmatter.
|
||||
func LoadHCL(path string) (*Config, error) {
|
||||
data, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read config %s: %w", path, err)
|
||||
}
|
||||
var cfg Config
|
||||
if err := hclsimple.Decode(path, data, nil, &cfg); err != nil {
|
||||
return nil, fmt.Errorf("decode config %s: %w", path, err)
|
||||
}
|
||||
return &cfg, nil
|
||||
}
|
||||
|
||||
func (c *Config) MergeOverrides(flags Flags, env Environ) *Config {
|
||||
|
||||
@@ -0,0 +1,111 @@
|
||||
package config
|
||||
|
||||
import (
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestLoad_DispatchHCL(t *testing.T) {
|
||||
p := writeTestFile(t, t.TempDir(), "config.hcl", exampleHCL)
|
||||
cfg, err := Load(p)
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if cfg.DBPath != "/tmp/orca/test.db" {
|
||||
t.Errorf("DBPath=%q", cfg.DBPath)
|
||||
}
|
||||
if cfg.NodeCapacity == nil || cfg.NodeCapacity.CPU != 4 {
|
||||
t.Errorf("NodeCapacity=%+v", cfg.NodeCapacity)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoad_DispatchMarkdown(t *testing.T) {
|
||||
p := writeTestFile(t, t.TempDir(), "config.md", exampleMarkdown)
|
||||
cfg, err := Load(p)
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if cfg.DBPath != "/tmp/orca/test.db" {
|
||||
t.Errorf("DBPath=%q", cfg.DBPath)
|
||||
}
|
||||
if cfg.ListenAddr != "127.0.0.1:9999" {
|
||||
t.Errorf("ListenAddr=%q", cfg.ListenAddr)
|
||||
}
|
||||
if cfg.NodeCapacity == nil || cfg.NodeCapacity.CPU != 4 {
|
||||
t.Errorf("NodeCapacity=%+v", cfg.NodeCapacity)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoad_DispatchYAML(t *testing.T) {
|
||||
body := "listen_addr: 0.0.0.0:5555\ndb_path: /bare.db\nnode_capacity:\n cpu: 2\n memory_mb: 4096\n"
|
||||
p := writeTestFile(t, t.TempDir(), "config.yaml", body)
|
||||
cfg, err := Load(p)
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if cfg.ListenAddr != "0.0.0.0:5555" {
|
||||
t.Errorf("ListenAddr=%q", cfg.ListenAddr)
|
||||
}
|
||||
if cfg.DBPath != "/bare.db" {
|
||||
t.Errorf("DBPath=%q", cfg.DBPath)
|
||||
}
|
||||
if cfg.NodeCapacity == nil || cfg.NodeCapacity.CPU != 2 {
|
||||
t.Errorf("NodeCapacity=%+v", cfg.NodeCapacity)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoad_DispatchYML(t *testing.T) {
|
||||
body := "listen_addr: 1.2.3.4:9\n"
|
||||
p := writeTestFile(t, t.TempDir(), "config.yml", body)
|
||||
cfg, err := Load(p)
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if cfg.ListenAddr != "1.2.3.4:9" {
|
||||
t.Errorf("ListenAddr=%q", cfg.ListenAddr)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoad_DispatchFirstExistingWins(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
missing := filepath.Join(dir, "missing.md")
|
||||
existing := writeTestFile(t, dir, "real.hcl", exampleHCL)
|
||||
cfg, err := Load(missing, existing)
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if cfg.DBPath != "/tmp/orca/test.db" {
|
||||
t.Errorf("DBPath=%q", cfg.DBPath)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoad_DispatchMissingReturnsZero(t *testing.T) {
|
||||
cfg, err := Load(filepath.Join(t.TempDir(), "nope.md"))
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if cfg == nil {
|
||||
t.Fatal("nil config")
|
||||
}
|
||||
if cfg.DBPath != "" || cfg.ListenAddr != "" || cfg.NodeCapacity != nil {
|
||||
t.Errorf("expected zero config, got %+v", cfg)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoad_DispatchHCLMalformed(t *testing.T) {
|
||||
p := writeTestFile(t, t.TempDir(), "bad.hcl", "db_path = ")
|
||||
if _, err := Load(p); err == nil {
|
||||
t.Fatal("expected error for malformed HCL")
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoad_DispatchUnknownExtFallsBackToHCL(t *testing.T) {
|
||||
p := writeTestFile(t, t.TempDir(), "config.unknown", exampleHCL)
|
||||
cfg, err := Load(p)
|
||||
if err != nil {
|
||||
t.Fatalf("Load: %v", err)
|
||||
}
|
||||
if cfg.DBPath != "/tmp/orca/test.db" {
|
||||
t.Errorf("DBPath=%q", cfg.DBPath)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,220 @@
|
||||
package config
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// LoadMarkdown decodes a Markdown config file with YAML frontmatter
|
||||
// (R-014). The file format is:
|
||||
//
|
||||
// ---
|
||||
// listen_addr: 127.0.0.1:9999
|
||||
// node_capacity:
|
||||
// cpu: 4
|
||||
// memory_mb: 8192
|
||||
// db_path: /tmp/orca/test.db
|
||||
// ---
|
||||
//
|
||||
// body prose (ignored)
|
||||
//
|
||||
// The frontmatter parser is a minimal hand-rolled key:value parser
|
||||
// (no new dependencies; gopkg.in/yaml.v3 is not in go.mod). It supports
|
||||
// flat scalar keys and one level of nested mapping (for node_capacity).
|
||||
// The Markdown body after the closing "---" is ignored.
|
||||
//
|
||||
// The returned *Config is the same struct the HCL loader produces, so
|
||||
// downstream consumers are unchanged.
|
||||
func LoadMarkdown(path string) (*Config, error) {
|
||||
data, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read config %s: %w", path, err)
|
||||
}
|
||||
return parseFrontmatter(string(data), path)
|
||||
}
|
||||
|
||||
// LoadMarkdownYAML decodes a bare YAML file (no Markdown body) using the
|
||||
// same minimal frontmatter parser. .yaml/.yml files are routed here by
|
||||
// the dispatcher.
|
||||
func LoadMarkdownYAML(path string) (*Config, error) {
|
||||
data, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read config %s: %w", path, err)
|
||||
}
|
||||
// Treat the whole file as the frontmatter block (no surrounding ---).
|
||||
return parseFrontmatterBlock(string(data), path)
|
||||
}
|
||||
|
||||
func parseFrontmatter(content, path string) (*Config, error) {
|
||||
block, ok := extractFrontmatter(content)
|
||||
if !ok {
|
||||
// No frontmatter delimiters: treat whole file as a bare block.
|
||||
return parseFrontmatterBlock(content, path)
|
||||
}
|
||||
return parseFrontmatterBlock(block, path)
|
||||
}
|
||||
|
||||
// extractFrontmatter returns the YAML block between the first pair of
|
||||
// "---" delimiters and whether a frontmatter block was present.
|
||||
func extractFrontmatter(content string) (string, bool) {
|
||||
trimmed := strings.TrimLeft(content, "\r\n\t ")
|
||||
if !strings.HasPrefix(trimmed, "---") {
|
||||
return "", false
|
||||
}
|
||||
// Skip the opening delimiter line.
|
||||
rest := trimmed[3:]
|
||||
rest = strings.TrimLeft(rest, "\r\n")
|
||||
// Find the closing delimiter line.
|
||||
idx := strings.Index(rest, "\n---")
|
||||
if idx < 0 {
|
||||
return "", false
|
||||
}
|
||||
return rest[:idx], true
|
||||
}
|
||||
|
||||
// parseFrontmatterBlock parses a minimal YAML-ish block into *Config.
|
||||
// Supported shapes:
|
||||
//
|
||||
// key: value
|
||||
// node_capacity:
|
||||
// cpu: 4
|
||||
// memory_mb: 8192
|
||||
//
|
||||
// Comments (# ...) and blank lines are ignored. Quoted scalar values
|
||||
// ("..." or '...') are unwrapped. No flow collections, anchors, or
|
||||
// multi-line strings are supported — by design, to avoid adding a YAML
|
||||
// dependency for this small config surface.
|
||||
func parseFrontmatterBlock(block, path string) (*Config, error) {
|
||||
cfg := &Config{}
|
||||
var inCapacity bool
|
||||
|
||||
lines := strings.Split(block, "\n")
|
||||
for lineNo, raw := range lines {
|
||||
line := stripComment(raw)
|
||||
if strings.TrimSpace(line) == "" {
|
||||
continue
|
||||
}
|
||||
|
||||
indent := countIndent(line)
|
||||
trimmed := strings.TrimSpace(line)
|
||||
|
||||
// A top-level key (no leading indent).
|
||||
if indent == 0 {
|
||||
inCapacity = false
|
||||
key, val, ok := splitKV(trimmed)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
if val == "" {
|
||||
// key with no value → nested mapping header (e.g. node_capacity:)
|
||||
if key == "node_capacity" {
|
||||
cfg.NodeCapacity = &CapacityConfig{}
|
||||
inCapacity = true
|
||||
}
|
||||
continue
|
||||
}
|
||||
applyScalar(cfg, key, val, path, lineNo)
|
||||
continue
|
||||
}
|
||||
|
||||
// Indented line under a nested mapping.
|
||||
if inCapacity && cfg.NodeCapacity != nil {
|
||||
key, val, hasVal := splitKV(trimmed)
|
||||
if !hasVal {
|
||||
continue
|
||||
}
|
||||
switch key {
|
||||
case "cpu":
|
||||
if n, err := strconv.Atoi(strings.TrimSpace(val)); err == nil {
|
||||
cfg.NodeCapacity.CPU = n
|
||||
}
|
||||
case "memory_mb":
|
||||
if n, err := strconv.Atoi(strings.TrimSpace(val)); err == nil {
|
||||
cfg.NodeCapacity.MemoryMB = n
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
return cfg, nil
|
||||
}
|
||||
|
||||
func applyScalar(cfg *Config, key, val, path string, lineNo int) {
|
||||
val = strings.TrimSpace(val)
|
||||
switch key {
|
||||
case "db_path":
|
||||
cfg.DBPath = unquote(val)
|
||||
case "listen_addr":
|
||||
cfg.ListenAddr = unquote(val)
|
||||
case "ca_path":
|
||||
cfg.CAPath = unquote(val)
|
||||
case "server_cert_path":
|
||||
cfg.ServerCertPath = unquote(val)
|
||||
case "server_key_path":
|
||||
cfg.ServerKeyPath = unquote(val)
|
||||
}
|
||||
_ = path
|
||||
_ = lineNo
|
||||
}
|
||||
|
||||
func splitKV(s string) (key, val string, ok bool) {
|
||||
idx := strings.Index(s, ":")
|
||||
if idx < 0 {
|
||||
return "", "", false
|
||||
}
|
||||
key = strings.TrimSpace(s[:idx])
|
||||
val = strings.TrimSpace(s[idx+1:])
|
||||
if key == "" {
|
||||
return "", "", false
|
||||
}
|
||||
return key, val, true
|
||||
}
|
||||
|
||||
func countIndent(s string) int {
|
||||
n := 0
|
||||
for _, r := range s {
|
||||
if r == ' ' || r == '\t' {
|
||||
n++
|
||||
continue
|
||||
}
|
||||
break
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
func stripComment(s string) string {
|
||||
// Strip inline comments not inside quotes. Minimal: only strip
|
||||
// when the '#' is preceded by whitespace or at line start.
|
||||
inSingle := false
|
||||
inDouble := false
|
||||
for i := 0; i < len(s); i++ {
|
||||
c := s[i]
|
||||
switch c {
|
||||
case '\'':
|
||||
if !inDouble {
|
||||
inSingle = !inSingle
|
||||
}
|
||||
case '"':
|
||||
if !inSingle {
|
||||
inDouble = !inDouble
|
||||
}
|
||||
case '#':
|
||||
if !inSingle && !inDouble {
|
||||
if i == 0 || s[i-1] == ' ' || s[i-1] == '\t' {
|
||||
return s[:i]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
func unquote(s string) string {
|
||||
if len(s) >= 2 {
|
||||
if (s[0] == '"' && s[len(s)-1] == '"') || (s[0] == '\'' && s[len(s)-1] == '\'') {
|
||||
return s[1 : len(s)-1]
|
||||
}
|
||||
}
|
||||
return s
|
||||
}
|
||||
@@ -0,0 +1,186 @@
|
||||
package config
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func writeTestFile(t *testing.T, dir, name, content string) string {
|
||||
t.Helper()
|
||||
p := filepath.Join(dir, name)
|
||||
if err := os.WriteFile(p, []byte(content), 0644); err != nil {
|
||||
t.Fatalf("write %s: %v", p, err)
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
const exampleMarkdown = `---
|
||||
listen_addr: "127.0.0.1:9999"
|
||||
db_path: "/tmp/orca/test.db"
|
||||
ca_path: "/tmp/orca/ca.crt"
|
||||
server_cert_path: "/tmp/orca/server.crt"
|
||||
server_key_path: "/tmp/orca/server.key"
|
||||
node_capacity:
|
||||
cpu: 4
|
||||
memory_mb: 8192
|
||||
---
|
||||
|
||||
# Orca config
|
||||
|
||||
This is prose body and is ignored by the loader.
|
||||
`
|
||||
|
||||
func TestLoadMarkdown_Full(t *testing.T) {
|
||||
p := writeTestFile(t, t.TempDir(), "config.md", exampleMarkdown)
|
||||
cfg, err := LoadMarkdown(p)
|
||||
if err != nil {
|
||||
t.Fatalf("LoadMarkdown: %v", err)
|
||||
}
|
||||
if cfg.DBPath != "/tmp/orca/test.db" {
|
||||
t.Errorf("DBPath=%q", cfg.DBPath)
|
||||
}
|
||||
if cfg.ListenAddr != "127.0.0.1:9999" {
|
||||
t.Errorf("ListenAddr=%q", cfg.ListenAddr)
|
||||
}
|
||||
if cfg.CAPath != "/tmp/orca/ca.crt" {
|
||||
t.Errorf("CAPath=%q", cfg.CAPath)
|
||||
}
|
||||
if cfg.ServerCertPath != "/tmp/orca/server.crt" {
|
||||
t.Errorf("ServerCertPath=%q", cfg.ServerCertPath)
|
||||
}
|
||||
if cfg.ServerKeyPath != "/tmp/orca/server.key" {
|
||||
t.Errorf("ServerKeyPath=%q", cfg.ServerKeyPath)
|
||||
}
|
||||
if cfg.NodeCapacity == nil {
|
||||
t.Fatal("NodeCapacity nil")
|
||||
}
|
||||
if cfg.NodeCapacity.CPU != 4 {
|
||||
t.Errorf("CPU=%d", cfg.NodeCapacity.CPU)
|
||||
}
|
||||
if cfg.NodeCapacity.MemoryMB != 8192 {
|
||||
t.Errorf("MemoryMB=%d", cfg.NodeCapacity.MemoryMB)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadMarkdown_NoFrontmatter(t *testing.T) {
|
||||
// No delimiters: whole file treated as a bare YAML block.
|
||||
body := "listen_addr: 0.0.0.0:1234\ndb_path: /x/y.db\n"
|
||||
p := writeTestFile(t, t.TempDir(), "config.md", body)
|
||||
cfg, err := LoadMarkdown(p)
|
||||
if err != nil {
|
||||
t.Fatalf("LoadMarkdown: %v", err)
|
||||
}
|
||||
if cfg.ListenAddr != "0.0.0.0:1234" {
|
||||
t.Errorf("ListenAddr=%q", cfg.ListenAddr)
|
||||
}
|
||||
if cfg.DBPath != "/x/y.db" {
|
||||
t.Errorf("DBPath=%q", cfg.DBPath)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadMarkdown_OnlyBody(t *testing.T) {
|
||||
body := `---
|
||||
---
|
||||
|
||||
# Just prose, no keys
|
||||
`
|
||||
p := writeTestFile(t, t.TempDir(), "config.md", body)
|
||||
cfg, err := LoadMarkdown(p)
|
||||
if err != nil {
|
||||
t.Fatalf("LoadMarkdown: %v", err)
|
||||
}
|
||||
if cfg.DBPath != "" || cfg.ListenAddr != "" || cfg.NodeCapacity != nil {
|
||||
t.Errorf("expected zero config, got %+v", cfg)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadMarkdown_CommentsAndBlanks(t *testing.T) {
|
||||
body := `---
|
||||
# a comment
|
||||
listen_addr: "127.0.0.1:9999"
|
||||
|
||||
db_path: "/tmp/orca/test.db" # inline comment
|
||||
|
||||
node_capacity:
|
||||
cpu: 4 # cores
|
||||
memory_mb: 8192
|
||||
---
|
||||
`
|
||||
p := writeTestFile(t, t.TempDir(), "config.md", body)
|
||||
cfg, err := LoadMarkdown(p)
|
||||
if err != nil {
|
||||
t.Fatalf("LoadMarkdown: %v", err)
|
||||
}
|
||||
if cfg.ListenAddr != "127.0.0.1:9999" {
|
||||
t.Errorf("ListenAddr=%q", cfg.ListenAddr)
|
||||
}
|
||||
if cfg.DBPath != "/tmp/orca/test.db" {
|
||||
t.Errorf("DBPath=%q", cfg.DBPath)
|
||||
}
|
||||
if cfg.NodeCapacity == nil || cfg.NodeCapacity.CPU != 4 || cfg.NodeCapacity.MemoryMB != 8192 {
|
||||
t.Errorf("NodeCapacity=%+v", cfg.NodeCapacity)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadMarkdownYAML_Bare(t *testing.T) {
|
||||
body := "listen_addr: 0.0.0.0:5555\ndb_path: /bare.db\nnode_capacity:\n cpu: 2\n memory_mb: 4096\n"
|
||||
p := writeTestFile(t, t.TempDir(), "config.yaml", body)
|
||||
cfg, err := LoadMarkdownYAML(p)
|
||||
if err != nil {
|
||||
t.Fatalf("LoadMarkdownYAML: %v", err)
|
||||
}
|
||||
if cfg.ListenAddr != "0.0.0.0:5555" {
|
||||
t.Errorf("ListenAddr=%q", cfg.ListenAddr)
|
||||
}
|
||||
if cfg.DBPath != "/bare.db" {
|
||||
t.Errorf("DBPath=%q", cfg.DBPath)
|
||||
}
|
||||
if cfg.NodeCapacity == nil || cfg.NodeCapacity.CPU != 2 || cfg.NodeCapacity.MemoryMB != 4096 {
|
||||
t.Errorf("NodeCapacity=%+v", cfg.NodeCapacity)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoadMarkdown_ReadError(t *testing.T) {
|
||||
missing := filepath.Join(t.TempDir(), "nope.md")
|
||||
if _, err := LoadMarkdown(missing); err == nil {
|
||||
t.Fatal("expected error for missing file")
|
||||
}
|
||||
}
|
||||
|
||||
func TestExtractFrontmatter(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
input string
|
||||
block string
|
||||
present bool
|
||||
}{
|
||||
{"standard", "---\nkey: val\n---\nbody", "key: val", true},
|
||||
{"leading-blanks", "\n\n---\nkey: val\n---\n", "key: val", true},
|
||||
{"no-delimiters", "key: val\n", "key: val", false},
|
||||
{"only-open", "---\nkey: val\n", "key: val", false},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
block, ok := extractFrontmatter(tc.input)
|
||||
if ok != tc.present {
|
||||
t.Errorf("present=%v want %v", ok, tc.present)
|
||||
}
|
||||
if tc.present && block != tc.block {
|
||||
t.Errorf("block=%q want %q", block, tc.block)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestUnquote(t *testing.T) {
|
||||
if got, want := unquote(`"hello"`), "hello"; got != want {
|
||||
t.Errorf("unquote double = %q want %q", got, want)
|
||||
}
|
||||
if got, want := unquote(`'hello'`), "hello"; got != want {
|
||||
t.Errorf("unquote single = %q want %q", got, want)
|
||||
}
|
||||
if got, want := unquote("bare"), "bare"; got != want {
|
||||
t.Errorf("unquote bare = %q want %q", got, want)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,94 @@
|
||||
// Package emitter defines the Layer-4 emitter interface (REQ-074,
|
||||
// I-B-002): the bridge between the declarative *jobspec.WorkloadSpec
|
||||
// and the server-side files. An Emitter renders a *WorkloadSpec into a
|
||||
// slice of File artifacts that the SSH-push transport SCPs to peers.
|
||||
//
|
||||
// Emitters are registered per workload kind + runtime (e.g.
|
||||
// "service:wasm", "job:process", "daemonset:wasm"). The Registry looks
|
||||
// up the right emitter by "<kind>:<runtime>" and delegates. Unknown
|
||||
// combinations return an error so the caller can fail fast before any
|
||||
// file is written.
|
||||
//
|
||||
// P0c only ships the interface, the Registry, a stub SystemdEmitter
|
||||
// (process runtime), and the File/Node value types. The full emitter
|
||||
// implementations (systemd lifecycle hooks, Traefik, Syncthing,
|
||||
// sockets) land in later phases (P02 Traefik, P04 lifecycle, P08
|
||||
// sockets, P09 Syncthing, v0.10-P03 secrets).
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// File is a single rendered artifact destined for a peer. The SSH-push
|
||||
// transport writes Content to Path atomically (write-to-tmp + rename)
|
||||
// with the given Mode (an octal string like "0644").
|
||||
type File struct {
|
||||
Path string
|
||||
Content string
|
||||
Mode string
|
||||
}
|
||||
|
||||
// Node is the minimal peer description an emitter needs to render
|
||||
// node-specific paths. It carries the hostname, the runtimes available
|
||||
// on the node (so emitters can branch), and the node tags (used by
|
||||
// DaemonSet matching and affinity in P05).
|
||||
type Node struct {
|
||||
Hostname string
|
||||
Runtime []string
|
||||
Tags []string
|
||||
}
|
||||
|
||||
// Emitter renders a *jobspec.WorkloadSpec for a given Node into a slice
|
||||
// of File artifacts. Implementations are registered with a Registry
|
||||
// keyed by "<kind>:<runtime>".
|
||||
type Emitter interface {
|
||||
Render(spec *jobspec.WorkloadSpec, node *Node) ([]File, error)
|
||||
}
|
||||
|
||||
// Registry holds emitters keyed by "<kind>:<runtime>" (e.g.
|
||||
// "service:process", "job:wasm"). The zero-value Registry is not
|
||||
// usable; construct one with NewRegistry.
|
||||
type Registry struct {
|
||||
emitters map[string]Emitter
|
||||
}
|
||||
|
||||
// NewRegistry returns an empty Registry ready for Register calls.
|
||||
func NewRegistry() *Registry {
|
||||
return &Registry{emitters: make(map[string]Emitter)}
|
||||
}
|
||||
|
||||
// Register registers an Emitter under the given key. The key is
|
||||
// "<kind>:<runtime>" (e.g. "job:process"). Registering twice under the
|
||||
// same key overwrites the prior registration (last-wins) to keep the
|
||||
// surface simple; callers are responsible for not double-registering.
|
||||
func (r *Registry) Register(key string, e Emitter) {
|
||||
r.emitters[key] = e
|
||||
}
|
||||
|
||||
// Render looks up the emitter for "<kind>:<runtime>" in the registry and
|
||||
// delegates to it. The kind is lowercased so the canonical spec kinds
|
||||
// (Job, Service, DaemonSet) map to the lowercase registry keys
|
||||
// ("job:process", "service:wasm", "daemonset:process"). Returns an error
|
||||
// if the spec is nil, the spec is missing its Kind, the runtime is
|
||||
// missing, or no emitter is registered for the combination.
|
||||
func (r *Registry) Render(spec *jobspec.WorkloadSpec, node *Node) ([]File, error) {
|
||||
if spec == nil {
|
||||
return nil, fmt.Errorf("emitter: spec is nil")
|
||||
}
|
||||
if strings.TrimSpace(spec.Kind) == "" {
|
||||
return nil, fmt.Errorf("emitter: spec kind is empty")
|
||||
}
|
||||
if spec.Runtime == nil {
|
||||
return nil, fmt.Errorf("emitter: spec runtime is nil")
|
||||
}
|
||||
key := strings.ToLower(spec.Kind) + ":" + spec.Runtime.OneOf
|
||||
e, ok := r.emitters[key]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("emitter: no emitter registered for %q (kind:runtime)", key)
|
||||
}
|
||||
return e.Render(spec, node)
|
||||
}
|
||||
@@ -0,0 +1,163 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// mockEmitter is a test-only Emitter that returns a fixed File slice
|
||||
// (or an error) so the Registry tests do not depend on the
|
||||
// SystemdEmitter. Implements Emitter via value receiver.
|
||||
type mockEmitter struct {
|
||||
files []File
|
||||
err error
|
||||
}
|
||||
|
||||
func (m mockEmitter) Render(spec *jobspec.WorkloadSpec, node *Node) ([]File, error) {
|
||||
if m.err != nil {
|
||||
return nil, m.err
|
||||
}
|
||||
out := make([]File, len(m.files))
|
||||
copy(out, m.files)
|
||||
return out, nil
|
||||
}
|
||||
|
||||
func TestRegistry_RegisterAndRender(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
want := []File{{Path: "/tmp/a", Content: "alpha", Mode: "0644"}}
|
||||
r.Register("job:process", mockEmitter{files: want})
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "demo",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/true"},
|
||||
}
|
||||
node := &Node{Hostname: "node-1", Runtime: []string{"process"}}
|
||||
got, err := r.Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(got) != 1 {
|
||||
t.Fatalf("got %d files, want 1", len(got))
|
||||
}
|
||||
if got[0] != want[0] {
|
||||
t.Errorf("file = %+v, want %+v", got[0], want[0])
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistry_UnknownKindRuntime(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "wasm"},
|
||||
}
|
||||
_, err := r.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for unknown kind:runtime, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "no emitter registered") {
|
||||
t.Errorf("error = %q, want 'no emitter registered'", err.Error())
|
||||
}
|
||||
if !strings.Contains(err.Error(), "service:wasm") {
|
||||
t.Errorf("error = %q, want it to mention 'service:wasm'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistry_MultipleEmittersCorrectSelected(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
jobFiles := []File{{Path: "/tmp/job", Content: "job", Mode: "0644"}}
|
||||
svcFiles := []File{{Path: "/tmp/svc", Content: "svc", Mode: "0644"}}
|
||||
dsFiles := []File{{Path: "/tmp/ds", Content: "ds", Mode: "0644"}}
|
||||
r.Register("job:process", mockEmitter{files: jobFiles})
|
||||
r.Register("service:process", mockEmitter{files: svcFiles})
|
||||
r.Register("daemonset:process", mockEmitter{files: dsFiles})
|
||||
|
||||
cases := []struct {
|
||||
kind string
|
||||
runtime string
|
||||
wantPath string
|
||||
}{
|
||||
{"Job", "process", "/tmp/job"},
|
||||
{"Service", "process", "/tmp/svc"},
|
||||
{"DaemonSet", "process", "/tmp/ds"},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.kind+":"+tc.runtime, func(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: tc.kind,
|
||||
Name: "x",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: tc.runtime, Command: "/bin/x"},
|
||||
}
|
||||
got, err := r.Render(spec, &Node{Hostname: "n"})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(got) != 1 {
|
||||
t.Fatalf("got %d files, want 1", len(got))
|
||||
}
|
||||
if got[0].Path != tc.wantPath {
|
||||
t.Errorf("path = %q, want %q", got[0].Path, tc.wantPath)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistry_NilSpec(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
_, err := r.Render(nil, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for nil spec")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "spec is nil") {
|
||||
t.Errorf("error = %q, want 'spec is nil'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistry_EmptyKind(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
spec := &jobspec.WorkloadSpec{Runtime: &jobspec.RuntimeBlock{OneOf: "process"}}
|
||||
_, err := r.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty kind")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "kind is empty") {
|
||||
t.Errorf("error = %q, want 'kind is empty'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistry_NilRuntime(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "x"}
|
||||
_, err := r.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for nil runtime")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "runtime is nil") {
|
||||
t.Errorf("error = %q, want 'runtime is nil'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistry_EmitterErrorPropagates(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
wantErr := errors.New("boom")
|
||||
r.Register("job:process", mockEmitter{err: wantErr})
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "x",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/x"},
|
||||
}
|
||||
_, err := r.Render(spec, &Node{})
|
||||
if !errors.Is(err, wantErr) {
|
||||
t.Errorf("err = %v, want %v", err, wantErr)
|
||||
}
|
||||
}
|
||||
|
||||
// Compile-time assertion that mockEmitter and SystemdEmitter implement
|
||||
// Emitter.
|
||||
var (
|
||||
_ Emitter = mockEmitter{}
|
||||
_ Emitter = SystemdEmitter{}
|
||||
)
|
||||
@@ -0,0 +1,128 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// SocketEmitter renders the systemd directives that implement the
|
||||
// R-007 socket-plumbing contract: workloads bind to
|
||||
// /run/orca/alloc-<id>/port-<name>.sock unless overridden via
|
||||
// service.bind = "127.0.0.1" (the only documented opt-in).
|
||||
//
|
||||
// The systemd side of the contract uses two directives:
|
||||
//
|
||||
// - RuntimeDirectory=orca/alloc-<alloc-id> — systemd creates
|
||||
// /run/orca/alloc-<alloc-id>/ owned by the service user (orca:orca)
|
||||
// with mode 0750. The directory is removed when the unit stops
|
||||
// (RuntimeDirectory= semantics). P08 emits one RuntimeDirectory=
|
||||
// line per port so each port's socket directory is created; the
|
||||
// alloc-id placeholder is spec.Name (the real alloc-id is assigned
|
||||
// by the scheduler at submit time — see allocIDFor).
|
||||
//
|
||||
// - ExecStartPre= — only when service.bind is "127.0.0.1" (the TCP
|
||||
// opt-in). In that case the workload binds a TCP port directly
|
||||
// (no socket), and the ExecStartPre is a placeholder that records
|
||||
// the bind (the actual bind happens in the process; the directive
|
||||
// is a no-op marker so operators can see the bind mode in the unit
|
||||
// file). When service.bind is empty (the default), the workload
|
||||
// binds the socket and no ExecStartPre is emitted for sockets.
|
||||
//
|
||||
// The socket path format is /run/orca/alloc-<alloc-id>/port-<port-name>.sock
|
||||
// where alloc-id is a PLACEHOLDER (spec.Name) — the real alloc-id is
|
||||
// assigned at submit time by the scheduler. The placeholder is
|
||||
// documented in the rendered unit via a comment so operators reading
|
||||
// the unit file understand the substitution.
|
||||
//
|
||||
// P08 is a PLAN/plumbing layer — the actual socket activation (socket
|
||||
// unit files, systemd socket-activation passing the pre-bound socket
|
||||
// fd to the process) lands in v0.10. P08 just renders the
|
||||
// RuntimeDirectory= lines and the optional TCP-bind ExecStartPre so
|
||||
// the directory exists at runtime.
|
||||
type SocketEmitter struct{}
|
||||
|
||||
// runtimeDirectoryRoot is the systemd RuntimeDirectory path root.
|
||||
// systemd joins this with the RuntimeDirectory= value to create
|
||||
// /run/orca/alloc-<id>. The leading slash is implicit in systemd
|
||||
// (RuntimeDirectory= is relative to /run).
|
||||
const runtimeDirectoryRoot = "orca"
|
||||
|
||||
// SocketPath returns the R-007 socket path for a port on the given
|
||||
// alloc-id. The alloc-id is the placeholder spec.Name when the real
|
||||
// alloc-id is not yet known (the scheduler assigns the real alloc-id
|
||||
// at submit time).
|
||||
func SocketPath(allocID, portName string) string {
|
||||
return fmt.Sprintf("/run/orca/alloc-%s/port-%s.sock", allocID, portName)
|
||||
}
|
||||
|
||||
// RenderSocketLines renders the systemd directives that implement
|
||||
// the R-007 socket plumbing for the given spec. The lines are returned
|
||||
// WITHOUT a trailing newline so the caller (the systemd emitter) can
|
||||
// append them to the [Service] block with consistent formatting.
|
||||
//
|
||||
// The returned lines are:
|
||||
//
|
||||
// - one RuntimeDirectory= line per port (so each port's socket
|
||||
// directory is created by systemd at unit start).
|
||||
// - a comment documenting the alloc-id placeholder.
|
||||
// - when service.bind is "127.0.0.1", an ExecStartPre= marker that
|
||||
// records the TCP opt-in (the actual bind is in the process).
|
||||
//
|
||||
// Returns an empty slice when the spec has no ports (no socket
|
||||
// plumbing needed — e.g. a Job or a port-less DaemonSet).
|
||||
func (SocketEmitter) RenderSocketLines(spec *jobspec.WorkloadSpec) []string {
|
||||
if spec == nil || len(spec.Ports) == 0 {
|
||||
return nil
|
||||
}
|
||||
allocID := allocIDForSocket(spec)
|
||||
var lines []string
|
||||
// One RuntimeDirectory= per port. systemd dedupes identical
|
||||
// values, but we emit one per port so the unit file is
|
||||
// self-documenting (each port maps to a directory entry).
|
||||
for _, p := range spec.Ports {
|
||||
lines = append(lines, fmt.Sprintf("RuntimeDirectory=%s/alloc-%s", runtimeDirectoryRoot, allocID))
|
||||
// Document the socket path this directory serves. systemd
|
||||
// ignores comment lines (lines starting with '#').
|
||||
lines = append(lines, fmt.Sprintf("# socket: %s", SocketPath(allocID, p.Name)))
|
||||
}
|
||||
// TCP opt-in: when service.bind is 127.0.0.1, the workload binds
|
||||
// a TCP port directly instead of the socket. We emit an
|
||||
// ExecStartPre marker so the bind mode is visible in the unit
|
||||
// file. The actual bind is in the process; the marker is a
|
||||
// no-op (echo to journald).
|
||||
if spec.Service != nil && strings.TrimSpace(spec.Service.Bind) != "" {
|
||||
if isTCPOptIn(spec.Service.Bind) {
|
||||
for _, p := range spec.Ports {
|
||||
lines = append(lines, fmt.Sprintf("ExecStartPre=/bin/echo orca: bind %s port %s (tcp, R-007 opt-in)", spec.Service.Bind, p.Name))
|
||||
}
|
||||
}
|
||||
}
|
||||
return lines
|
||||
}
|
||||
|
||||
// allocIDForSocket returns the alloc-id placeholder for the spec. The
|
||||
// real alloc-id is assigned by the scheduler at submit time; P08 uses
|
||||
// spec.Name as a deterministic placeholder so the rendered unit is
|
||||
// stable across re-renders. This mirrors the Traefik emitter's
|
||||
// allocIDFor (which uses the node hostname for the Traefik
|
||||
// dynamic-config server URL); the systemd unit is per-alloc, so
|
||||
// spec.Name is the right placeholder here.
|
||||
func allocIDForSocket(spec *jobspec.WorkloadSpec) string {
|
||||
if spec == nil || strings.TrimSpace(spec.Name) == "" {
|
||||
return "<allocID>"
|
||||
}
|
||||
return spec.Name
|
||||
}
|
||||
|
||||
// isTCPOptIn returns true when the bind value is the documented
|
||||
// 127.0.0.1 TCP opt-in (R-007). Other valid IPs (::1, etc.) are also
|
||||
// TCP opt-ins (any non-empty bind opts out of the socket default); we
|
||||
// only emit the marker for 127.0.0.1 because that is the only
|
||||
// documented opt-in per the PRD — other IPs are accepted by the
|
||||
// schema validator but are operator-specific and we do not
|
||||
// second-guess them.
|
||||
func isTCPOptIn(bind string) bool {
|
||||
return strings.TrimSpace(bind) == "127.0.0.1"
|
||||
}
|
||||
@@ -0,0 +1,265 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
func TestSocketEmitter_RenderSocketLines_NoPorts(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "backup",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
|
||||
}
|
||||
lines := (SocketEmitter{}).RenderSocketLines(spec)
|
||||
if len(lines) != 0 {
|
||||
t.Errorf("got %d lines, want 0 for no ports: %v", len(lines), lines)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_RenderSocketLines_NilSpec(t *testing.T) {
|
||||
lines := (SocketEmitter{}).RenderSocketLines(nil)
|
||||
if lines != nil {
|
||||
t.Errorf("nil spec should return nil, got %v", lines)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_RenderSocketLines_SinglePort(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
lines := (SocketEmitter{}).RenderSocketLines(spec)
|
||||
// Expect: RuntimeDirectory + comment. No TCP bind (default socket).
|
||||
wantRT := "RuntimeDirectory=orca/alloc-web"
|
||||
if !contains(lines, wantRT) {
|
||||
t.Errorf("lines %v missing %q", lines, wantRT)
|
||||
}
|
||||
wantSock := "# socket: /run/orca/alloc-web/port-http.sock"
|
||||
if !contains(lines, wantSock) {
|
||||
t.Errorf("lines %v missing %q", lines, wantSock)
|
||||
}
|
||||
for _, l := range lines {
|
||||
if strings.HasPrefix(l, "ExecStartPre=") {
|
||||
t.Errorf("socket bind should not emit ExecStartPre (no TCP opt-in): %s", l)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_RenderSocketLines_MultiplePorts(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "api",
|
||||
Ports: []jobspec.PortSpec{
|
||||
{Name: "http", Port: 8080},
|
||||
{Name: "grpc", Port: 9090},
|
||||
},
|
||||
}
|
||||
lines := (SocketEmitter{}).RenderSocketLines(spec)
|
||||
// Two RuntimeDirectory lines (one per port).
|
||||
count := 0
|
||||
for _, l := range lines {
|
||||
if l == "RuntimeDirectory=orca/alloc-api" {
|
||||
count++
|
||||
}
|
||||
}
|
||||
if count != 2 {
|
||||
t.Errorf("RuntimeDirectory count = %d, want 2 (one per port)", count)
|
||||
}
|
||||
if !contains(lines, "# socket: /run/orca/alloc-api/port-http.sock") {
|
||||
t.Errorf("missing http socket comment")
|
||||
}
|
||||
if !contains(lines, "# socket: /run/orca/alloc-api/port-grpc.sock") {
|
||||
t.Errorf("missing grpc socket comment")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_RenderSocketLines_TCPBind127(t *testing.T) {
|
||||
// service.bind = 127.0.0.1 → TCP opt-in → ExecStartPre marker per port.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: "127.0.0.1"},
|
||||
}
|
||||
lines := (SocketEmitter{}).RenderSocketLines(spec)
|
||||
found := false
|
||||
for _, l := range lines {
|
||||
if strings.HasPrefix(l, "ExecStartPre=/bin/echo orca: bind 127.0.0.1 port http (tcp, R-007 opt-in)") {
|
||||
found = true
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
t.Errorf("missing TCP bind ExecStartPre marker; lines: %v", lines)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_RenderSocketLines_TCPBindIPv6(t *testing.T) {
|
||||
// Non-127.0.0.1 bind is accepted by schema but not the documented
|
||||
// opt-in; the marker is only emitted for 127.0.0.1. The
|
||||
// RuntimeDirectory lines are still emitted (the directory exists
|
||||
// regardless of bind mode — sockets or TCP).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: "::1"},
|
||||
}
|
||||
lines := (SocketEmitter{}).RenderSocketLines(spec)
|
||||
for _, l := range lines {
|
||||
if strings.HasPrefix(l, "ExecStartPre=") {
|
||||
t.Errorf("::1 bind should NOT emit TCP marker (only 127.0.0.1 is documented opt-in): %s", l)
|
||||
}
|
||||
}
|
||||
if !contains(lines, "RuntimeDirectory=orca/alloc-web") {
|
||||
t.Errorf("RuntimeDirectory should still be emitted for ::1 bind")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_RenderSocketLines_EmptyBindSocket(t *testing.T) {
|
||||
// Empty bind → default socket → no TCP marker, but RuntimeDirectory emitted.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: ""},
|
||||
}
|
||||
lines := (SocketEmitter{}).RenderSocketLines(spec)
|
||||
for _, l := range lines {
|
||||
if strings.HasPrefix(l, "ExecStartPre=") {
|
||||
t.Errorf("empty bind should NOT emit TCP marker: %s", l)
|
||||
}
|
||||
}
|
||||
if !contains(lines, "RuntimeDirectory=orca/alloc-web") {
|
||||
t.Errorf("RuntimeDirectory missing for empty bind")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_RenderSocketLines_NilService(t *testing.T) {
|
||||
// No service block → default socket → no TCP marker, but RuntimeDirectory emitted.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
lines := (SocketEmitter{}).RenderSocketLines(spec)
|
||||
for _, l := range lines {
|
||||
if strings.HasPrefix(l, "ExecStartPre=") {
|
||||
t.Errorf("nil service should NOT emit TCP marker: %s", l)
|
||||
}
|
||||
}
|
||||
if !contains(lines, "RuntimeDirectory=orca/alloc-web") {
|
||||
t.Errorf("RuntimeDirectory missing for nil service")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_SocketPath(t *testing.T) {
|
||||
got := SocketPath("alloc-123", "http")
|
||||
want := "/run/orca/alloc-alloc-123/port-http.sock"
|
||||
if got != want {
|
||||
t.Errorf("SocketPath = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_AllocIDPlaceholderNilSpec(t *testing.T) {
|
||||
if got := allocIDForSocket(nil); got != "<allocID>" {
|
||||
t.Errorf("allocIDForSocket(nil) = %q, want <allocID>", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_AllocIDPlaceholderEmptyName(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Name: " "}
|
||||
if got := allocIDForSocket(spec); got != "<allocID>" {
|
||||
t.Errorf("allocIDForSocket(empty name) = %q, want <allocID>", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_AllocIDPlaceholderNamedSpec(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Name: "web"}
|
||||
if got := allocIDForSocket(spec); got != "web" {
|
||||
t.Errorf("allocIDForSocket(web) = %q, want web", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSocketEmitter_SocketPathPlaceholder(t *testing.T) {
|
||||
got := SocketPath("<allocID>", "grpc")
|
||||
want := "/run/orca/alloc-<allocID>/port-grpc.sock"
|
||||
if got != want {
|
||||
t.Errorf("SocketPath = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_IntegratesSocketLines(t *testing.T) {
|
||||
// End-to-end: the systemd unit for a Service with ports contains
|
||||
// the RuntimeDirectory line emitted by the SocketEmitter.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/httpd"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if !strings.Contains(c, "RuntimeDirectory=orca/alloc-web\n") {
|
||||
t.Errorf("unit missing RuntimeDirectory line\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "# socket: /run/orca/alloc-web/port-http.sock\n") {
|
||||
t.Errorf("unit missing socket path comment\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_IntegratesSocketLinesTCPBind(t *testing.T) {
|
||||
// When service.bind = 127.0.0.1, the unit contains the ExecStartPre
|
||||
// TCP-bind marker.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/httpd"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: "127.0.0.1"},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if !strings.Contains(c, "ExecStartPre=/bin/echo orca: bind 127.0.0.1 port http (tcp, R-007 opt-in)\n") {
|
||||
t.Errorf("unit missing TCP bind ExecStartPre marker\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_NoSocketLinesForPortlessSpec(t *testing.T) {
|
||||
// A Job with no ports → no RuntimeDirectory line in the unit.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "backup",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if strings.Contains(c, "RuntimeDirectory=") {
|
||||
t.Errorf("portless spec should not emit RuntimeDirectory\n%s", c)
|
||||
}
|
||||
if strings.Contains(c, "# socket:") {
|
||||
t.Errorf("portless spec should not emit socket comment\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
// contains reports whether the slice contains the string s.
|
||||
func contains(lines []string, s string) bool {
|
||||
for _, l := range lines {
|
||||
if l == s {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -0,0 +1,175 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
"git.cloudinit.dev/coreci/orca/internal/storage"
|
||||
)
|
||||
|
||||
// SyncthingEmitter is the Layer-4 emitter for the per-namespace
|
||||
// Syncthing config files (REQ-081). For every volume in spec.Volumes
|
||||
// that carries a `replicate:` list, the emitter renders one Syncthing
|
||||
// `config.xml` at /etc/syncthing/orca-<ns>-<volume>.xml containing the
|
||||
// content-addressed folder (storage.FolderID), the device list (all
|
||||
// peers in the namespace), and the volume path.
|
||||
//
|
||||
// Syncthing configs are kind-agnostic — they apply to any workload
|
||||
// (Job, Service, DaemonSet) that declares a replicated volume. The
|
||||
// emitter is therefore registered on the Registry under every
|
||||
// kind:runtime key the other emitters use, but it is intended to be
|
||||
// composed by the caller (the caller renders both the systemd unit and
|
||||
// the Syncthing config for the same spec). For the v0.9-P09 spike the
|
||||
// emitter is invoked directly; the composition lands in a later phase.
|
||||
//
|
||||
// The emitter is CLI-side only: it renders the XML; the SSH-push
|
||||
// transport SCPs the file to each peer; the Syncthing apt package on
|
||||
// the peer reads it. No Syncthing Go client is linked.
|
||||
type SyncthingEmitter struct{}
|
||||
|
||||
// syncthingConfigDir is the canonical directory for rendered Syncthing
|
||||
// configs on a peer (R-005). The emitter writes one file per
|
||||
// replicated volume.
|
||||
const syncthingConfigDir = "/etc/syncthing"
|
||||
|
||||
// localDeviceAddress is the address the local peer's own device entry
|
||||
// uses. "dynamic" tells Syncthing this peer is the listener (it does
|
||||
// not dial out to itself).
|
||||
const localDeviceAddress = "dynamic"
|
||||
|
||||
// peerDeviceAddressTemplate renders the Sync listen address for a
|
||||
// remote peer. The peer hostname (from the ReplicateTo list) is used
|
||||
// as the host; the default Sync port is 22000.
|
||||
const peerDeviceAddressTemplate = "tcp://%s:22000"
|
||||
|
||||
// Render renders one Syncthing config XML file per replicated volume
|
||||
// in the spec. A volume is "replicated" when its VolumeSpec carries a
|
||||
// non-empty `replicate:` list — encoded in VolumeSpec.Source as the
|
||||
// comma-separated peer list prefixed with `replicate:` (e.g.
|
||||
// `replicate:peer-b,peer-c`). This keeps the VolumeSpec shape stable
|
||||
// (the v0.9 VolumeSpec has no explicit Replicate field; the emitter
|
||||
// parses it from Source).
|
||||
//
|
||||
// For each replicated volume, the emitter:
|
||||
//
|
||||
// 1. Builds a VolumeReplication (namespace = spec.Name's namespace,
|
||||
// volume name, source path = VolumeSpec.Target).
|
||||
// 2. Builds the peer device list (the local peer + every peer in the
|
||||
// `replicate:` list). The local peer's device ID is derived
|
||||
// deterministically from the node hostname (the real device ID is
|
||||
// discovered from the peer registry in a later phase; for the
|
||||
// spike a deterministic placeholder keeps the rendered config
|
||||
// byte-stable).
|
||||
// 3. Calls storage.RenderSyncthingConfig + storage.RenderSyncthingXML
|
||||
// to produce the config file content.
|
||||
//
|
||||
// Returns an error if the spec is nil or the node is nil (the node is
|
||||
// required to identify the local peer). Workloads with no replicated
|
||||
// volumes return an empty (non-nil) slice — the emitter is a no-op for
|
||||
// them.
|
||||
func (SyncthingEmitter) Render(spec *jobspec.WorkloadSpec, node *Node) ([]File, error) {
|
||||
if spec == nil {
|
||||
return nil, errors.New("emitter/syncthing: spec is nil")
|
||||
}
|
||||
if node == nil {
|
||||
return nil, errors.New("emitter/syncthing: node is nil (local peer unknown)")
|
||||
}
|
||||
var files []File
|
||||
for _, vol := range spec.Volumes {
|
||||
peers, ok := parseReplicateList(vol.Source)
|
||||
if !ok || len(peers) == 0 {
|
||||
continue
|
||||
}
|
||||
path := vol.Target
|
||||
if strings.TrimSpace(path) == "" {
|
||||
path = vol.Source
|
||||
}
|
||||
rep := storage.VolumeReplication{
|
||||
Namespace: spec.Name,
|
||||
VolumeName: vol.Name,
|
||||
SourcePath: path,
|
||||
ReplicateTo: peers,
|
||||
SyncMode: "sendreceive",
|
||||
}
|
||||
devices := buildSyncthingDevices(node, peers)
|
||||
cfg, err := storage.RenderSyncthingConfig(rep, devices)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("emitter/syncthing: render config for volume %q: %w", vol.Name, err)
|
||||
}
|
||||
xml, err := storage.RenderSyncthingXML(cfg)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("emitter/syncthing: render xml for volume %q: %w", vol.Name, err)
|
||||
}
|
||||
files = append(files, File{
|
||||
Path: fmt.Sprintf("%s/orca-%s-%s.xml", syncthingConfigDir, spec.Name, vol.Name),
|
||||
Content: xml,
|
||||
Mode: "0644",
|
||||
})
|
||||
}
|
||||
return files, nil
|
||||
}
|
||||
|
||||
// parseReplicateList extracts the peer list from a VolumeSpec.Source
|
||||
// value of the form `replicate:peer-b,peer-c`. Returns the peer list
|
||||
// and true when the Source carries a replicate directive; returns nil
|
||||
// and false otherwise (the volume is not replicated).
|
||||
func parseReplicateList(source string) ([]string, bool) {
|
||||
s := strings.TrimSpace(source)
|
||||
if !strings.HasPrefix(s, "replicate:") {
|
||||
return nil, false
|
||||
}
|
||||
rest := strings.TrimPrefix(s, "replicate:")
|
||||
parts := strings.Split(rest, ",")
|
||||
out := make([]string, 0, len(parts))
|
||||
for _, p := range parts {
|
||||
p = strings.TrimSpace(p)
|
||||
if p != "" {
|
||||
out = append(out, p)
|
||||
}
|
||||
}
|
||||
return out, true
|
||||
}
|
||||
|
||||
// buildSyncthingDevices builds the Syncthing device list for the
|
||||
// rendered config. The local peer (the node the config is being
|
||||
// rendered for) is first, with the local device address ("dynamic").
|
||||
// Each remote peer in the replicate list follows, with a
|
||||
// tcp://<peer>:22000 address. Device IDs are deterministic placeholders
|
||||
// derived from the peer name (the real device IDs are discovered from
|
||||
// the peer registry in a later phase; the placeholder keeps the
|
||||
// rendered config byte-stable across re-runs).
|
||||
func buildSyncthingDevices(node *Node, peers []string) []storage.SyncthingDevice {
|
||||
devices := make([]storage.SyncthingDevice, 0, len(peers)+1)
|
||||
devices = append(devices, storage.SyncthingDevice{
|
||||
ID: syncthingDeviceID(node.Hostname),
|
||||
Name: node.Hostname,
|
||||
Address: localDeviceAddress,
|
||||
})
|
||||
for _, p := range peers {
|
||||
devices = append(devices, storage.SyncthingDevice{
|
||||
ID: syncthingDeviceID(p),
|
||||
Name: p,
|
||||
Address: fmt.Sprintf(peerDeviceAddressTemplate, p),
|
||||
})
|
||||
}
|
||||
return devices
|
||||
}
|
||||
|
||||
// syncthingDeviceID returns a deterministic, stable device-ID
|
||||
// placeholder for the given peer name. The placeholder is a fixed
|
||||
// 52-char string (Syncthing device IDs are 52-char base32) derived by
|
||||
// padding the peer name. The real device ID (discovered from the peer
|
||||
// registry / cluster/peers/) replaces this in a later phase; for the
|
||||
// v0.9 spike the placeholder keeps the rendered config byte-stable so
|
||||
// the SSH-push idempotency check works.
|
||||
func syncthingDeviceID(peerName string) string {
|
||||
const idLen = 52
|
||||
name := strings.TrimSpace(peerName)
|
||||
if len(name) >= idLen {
|
||||
return name[:idLen]
|
||||
}
|
||||
pad := strings.Repeat("X", idLen-len(name))
|
||||
return name + pad
|
||||
}
|
||||
@@ -0,0 +1,172 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
func TestSyncthingEmitter_TwoReplicatedVolumes_TwoFiles(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "team-alpha",
|
||||
Volumes: []jobspec.VolumeSpec{
|
||||
{Name: "data", Source: "replicate:peer-b,peer-c", Target: "/var/lib/orca/data"},
|
||||
{Name: "logs", Source: "replicate:peer-b", Target: "/var/lib/orca/logs"},
|
||||
{Name: "cache", Source: "/local/cache", Target: "/cache"}, // not replicated
|
||||
},
|
||||
}
|
||||
node := &Node{Hostname: "peer-a", Runtime: []string{"process"}}
|
||||
files, err := (SyncthingEmitter{}).Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(files) != 2 {
|
||||
t.Fatalf("expected 2 files (only replicated volumes), got %d", len(files))
|
||||
}
|
||||
for _, f := range files {
|
||||
if !strings.HasPrefix(f.Path, "/etc/syncthing/orca-team-alpha-") {
|
||||
t.Errorf("path %q does not start with /etc/syncthing/orca-team-alpha-", f.Path)
|
||||
}
|
||||
if !strings.HasSuffix(f.Path, ".xml") {
|
||||
t.Errorf("path %q does not end with .xml", f.Path)
|
||||
}
|
||||
if f.Mode != "0644" {
|
||||
t.Errorf("mode = %q, want 0644", f.Mode)
|
||||
}
|
||||
if !strings.Contains(f.Content, "<configuration") {
|
||||
t.Errorf("content of %q is not syncthing config XML", f.Path)
|
||||
}
|
||||
}
|
||||
// File 1: data volume, replicated to peer-b and peer-c.
|
||||
if !strings.Contains(files[0].Path, "team-alpha-data") {
|
||||
t.Errorf("first file path = %q, want team-alpha-data", files[0].Path)
|
||||
}
|
||||
if !strings.Contains(files[0].Content, "peer-b") || !strings.Contains(files[0].Content, "peer-c") {
|
||||
t.Errorf("data volume config missing peer-b or peer-c")
|
||||
}
|
||||
// File 2: logs volume, replicated to peer-b only.
|
||||
if !strings.Contains(files[1].Path, "team-alpha-logs") {
|
||||
t.Errorf("second file path = %q, want team-alpha-logs", files[1].Path)
|
||||
}
|
||||
if !strings.Contains(files[1].Content, "peer-b") {
|
||||
t.Errorf("logs volume config missing peer-b")
|
||||
}
|
||||
if strings.Contains(files[1].Content, "tcp://peer-c") {
|
||||
t.Errorf("logs volume config should not contain peer-c")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSyncthingEmitter_NoReplicatedVolumes_Empty(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "batch",
|
||||
Volumes: []jobspec.VolumeSpec{
|
||||
{Name: "cache", Source: "/local/cache", Target: "/cache"},
|
||||
},
|
||||
}
|
||||
node := &Node{Hostname: "peer-a"}
|
||||
files, err := (SyncthingEmitter{}).Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(files) != 0 {
|
||||
t.Errorf("expected 0 files for non-replicated volumes, got %d", len(files))
|
||||
}
|
||||
}
|
||||
|
||||
func TestSyncthingEmitter_NoVolumes_Empty(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "batch"}
|
||||
node := &Node{Hostname: "peer-a"}
|
||||
files, err := (SyncthingEmitter{}).Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(files) != 0 {
|
||||
t.Errorf("expected 0 files, got %d", len(files))
|
||||
}
|
||||
}
|
||||
|
||||
func TestSyncthingEmitter_NilSpec_Error(t *testing.T) {
|
||||
node := &Node{Hostname: "peer-a"}
|
||||
if _, err := (SyncthingEmitter{}).Render(nil, node); err == nil {
|
||||
t.Error("nil spec: expected error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSyncthingEmitter_NilNode_Error(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "x"}
|
||||
if _, err := (SyncthingEmitter{}).Render(spec, nil); err == nil {
|
||||
t.Error("nil node: expected error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSyncthingEmitter_LocalDeviceFirst(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "ns",
|
||||
Volumes: []jobspec.VolumeSpec{{Name: "data", Source: "replicate:peer-b", Target: "/data"}},
|
||||
}
|
||||
node := &Node{Hostname: "peer-a"}
|
||||
files, err := (SyncthingEmitter{}).Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(files) != 1 {
|
||||
t.Fatalf("expected 1 file, got %d", len(files))
|
||||
}
|
||||
// Local peer address is "dynamic"; remote peer uses tcp://...
|
||||
if !strings.Contains(files[0].Content, "dynamic") {
|
||||
t.Errorf("local device address (dynamic) missing from config")
|
||||
}
|
||||
if !strings.Contains(files[0].Content, "tcp://peer-b:22000") {
|
||||
t.Errorf("remote peer address missing from config")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseReplicateList_OK(t *testing.T) {
|
||||
peers, ok := parseReplicateList("replicate:peer-b,peer-c,peer-d")
|
||||
if !ok {
|
||||
t.Fatal("expected ok")
|
||||
}
|
||||
if len(peers) != 3 || peers[0] != "peer-b" || peers[1] != "peer-c" || peers[2] != "peer-d" {
|
||||
t.Errorf("peers = %v, want [peer-b peer-c peer-d]", peers)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseReplicateList_NotReplicated(t *testing.T) {
|
||||
_, ok := parseReplicateList("/local/path")
|
||||
if ok {
|
||||
t.Error("non-replicate source should return ok=false")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseReplicateList_EmptyPeerList(t *testing.T) {
|
||||
peers, ok := parseReplicateList("replicate:")
|
||||
if !ok {
|
||||
t.Error("replicate: prefix should return ok=true")
|
||||
}
|
||||
if len(peers) != 0 {
|
||||
t.Errorf("peers = %v, want empty", peers)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSyncthingDeviceID_Stable(t *testing.T) {
|
||||
a := syncthingDeviceID("peer-a")
|
||||
b := syncthingDeviceID("peer-a")
|
||||
if a != b {
|
||||
t.Errorf("syncthingDeviceID not stable: %q vs %q", a, b)
|
||||
}
|
||||
if len(a) != 52 {
|
||||
t.Errorf("device ID length = %d, want 52", len(a))
|
||||
}
|
||||
}
|
||||
|
||||
func TestSyncthingDeviceID_DifferentPeers(t *testing.T) {
|
||||
a := syncthingDeviceID("peer-a")
|
||||
b := syncthingDeviceID("peer-b")
|
||||
if a == b {
|
||||
t.Errorf("different peers produced same device ID")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,239 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// SystemdEmitter is the Emitter implementation for the "process"
|
||||
// runtime. It renders the systemd unit file for the workload,
|
||||
// including the lifecycle hooks (P04) and the R-007 socket plumbing
|
||||
// (P08).
|
||||
//
|
||||
// Lifecycle hooks map to systemd semantics (PRD §10.1):
|
||||
//
|
||||
// - lifecycle.pre_stop → ExecStop= (the command run on stop; systemd
|
||||
// runs ExecStop, then kills the main process after the deadline).
|
||||
// - lifecycle.post_start → ExecStartPost= (runs after the main
|
||||
// process starts).
|
||||
//
|
||||
// systemd has no ExecStartPre equivalent for a "pre_start" hook; the
|
||||
// spec does not define pre_start (only pre_stop and post_start per
|
||||
// PRD §10.1), so no mapping is needed.
|
||||
//
|
||||
// The unit name carries the `orca-v1-` prefix per the dual-write
|
||||
// window (REQ-090) so the v0.9 SSH-push path does not collide with the
|
||||
// v0.8 daemon's `orca-<job>.service` units during the migration
|
||||
// window.
|
||||
//
|
||||
// Later phases extend this emitter:
|
||||
//
|
||||
// - v0.10-P03: secrets via EnvironmentFile= + LoadCredential=
|
||||
type SystemdEmitter struct{}
|
||||
|
||||
// unitNamePrefix is the v0.9 SSH-push unit-name prefix. The v0.8
|
||||
// daemon uses `orca-<job>.service`; the v0.9 path uses
|
||||
// `orca-v1-<spec.Name>.service` so the two never overlap (REQ-090,
|
||||
// I-C-006). The prefix is load-bearing — do not change it without
|
||||
// updating the dual-write window contract.
|
||||
const unitNamePrefix = "orca-v1-"
|
||||
|
||||
// Render renders the systemd unit file for a process-runtime workload.
|
||||
//
|
||||
// When the spec has no Tasks (the single-process case, the historical
|
||||
// shape), the unit name is
|
||||
// /etc/systemd/system/<unitNamePrefix><spec.Name>.service and the
|
||||
// content is a [Service] block with ExecStart, optional ExecStartPost
|
||||
// (lifecycle.post_start), optional ExecStop (lifecycle.pre_stop), and
|
||||
// the R-007 socket-plumbing lines (RuntimeDirectory=, optional
|
||||
// TCP-bind ExecStartPre). Mode is 0644.
|
||||
//
|
||||
// When the spec has a task group (P06, spec.Tasks non-empty), the
|
||||
// alloc is multi-process and Render emits one systemd unit per task
|
||||
// (`orca-v1-alloc-<alloc-id>-<task-name>.service`) plus a single
|
||||
// grouping target unit (`orca-v1-alloc-<alloc-id>.target`) that
|
||||
// starts/stops all tasks together. Each per-task unit carries
|
||||
// `PartOf=orca-v1-alloc-<alloc-id>.target` and is
|
||||
// `WantedBy=multi-user.target` so the task starts at boot. Tasks
|
||||
// that omit their own runtime inherit the top-level spec.Runtime as
|
||||
// the per-group default.
|
||||
//
|
||||
// The rendered shape (single-process) is:
|
||||
//
|
||||
// [Service]
|
||||
// ExecStart=<runtime command>
|
||||
// ExecStartPost=<post_start command 1>
|
||||
// ExecStartPost=<post_start command 2>
|
||||
// ExecStop=<pre_stop command 1>
|
||||
// ExecStop=<pre_stop command 2>
|
||||
// RuntimeDirectory=orca/alloc-<alloc-id>
|
||||
// # socket: /run/orca/alloc-<alloc-id>/port-<name>.sock
|
||||
// ExecStartPre=/bin/echo orca: bind 127.0.0.1 port <name> (tcp, R-007 opt-in)
|
||||
//
|
||||
// Returns an error if the spec is nil, the spec is missing its name,
|
||||
// the runtime block is nil, or the runtime command is empty (a
|
||||
// workload with no command has nothing to ExecStart). For task groups,
|
||||
// returns an error if any task has no resolvable runtime command.
|
||||
func (SystemdEmitter) Render(spec *jobspec.WorkloadSpec, node *Node) ([]File, error) {
|
||||
if spec == nil {
|
||||
return nil, errors.New("emitter/systemd: spec is nil")
|
||||
}
|
||||
if strings.TrimSpace(spec.Name) == "" {
|
||||
return nil, errors.New("emitter/systemd: spec name is empty")
|
||||
}
|
||||
if len(spec.Tasks) > 0 {
|
||||
return renderTaskGroup(spec, node)
|
||||
}
|
||||
if spec.Runtime == nil {
|
||||
return nil, errors.New("emitter/systemd: runtime block is nil")
|
||||
}
|
||||
if strings.TrimSpace(spec.Runtime.Command) == "" {
|
||||
return nil, errors.New("emitter/systemd: runtime command is empty")
|
||||
}
|
||||
path := fmt.Sprintf("/etc/systemd/system/%s%s.service", unitNamePrefix, spec.Name)
|
||||
content := renderSystemdUnit(spec)
|
||||
return []File{{Path: path, Content: content, Mode: "0644"}}, nil
|
||||
}
|
||||
|
||||
// renderTaskGroup renders one systemd unit per task plus the grouping
|
||||
// target unit. Each task's runtime falls back to the top-level
|
||||
// spec.Runtime when the task omits its own. Tasks with no resolvable
|
||||
// command (no task.Command, no task.Runtime.Command, no top-level
|
||||
// Runtime) return an error.
|
||||
func renderTaskGroup(spec *jobspec.WorkloadSpec, node *Node) ([]File, error) {
|
||||
allocID := spec.Name
|
||||
targetUnit := fmt.Sprintf("%salloc-%s.target", unitNamePrefix, allocID)
|
||||
targetPath := fmt.Sprintf("/etc/systemd/system/%s", targetUnit)
|
||||
var files []File
|
||||
for _, task := range spec.Tasks {
|
||||
rt := taskRuntime(spec, &task)
|
||||
if rt == nil {
|
||||
return nil, fmt.Errorf("emitter/systemd: task %q has no runtime (set tasks[].runtime or top-level runtime)", task.Name)
|
||||
}
|
||||
cmd := taskCommand(spec, &task, rt)
|
||||
if strings.TrimSpace(cmd) == "" {
|
||||
return nil, fmt.Errorf("emitter/systemd: task %q command is empty", task.Name)
|
||||
}
|
||||
unitName := fmt.Sprintf("%salloc-%s-%s.service", unitNamePrefix, allocID, task.Name)
|
||||
path := fmt.Sprintf("/etc/systemd/system/%s", unitName)
|
||||
content := renderTaskUnit(spec, &task, rt, cmd, targetUnit)
|
||||
files = append(files, File{Path: path, Content: content, Mode: "0644"})
|
||||
}
|
||||
files = append(files, File{
|
||||
Path: targetPath,
|
||||
Content: renderTargetUnit(targetUnit, spec, allocID),
|
||||
Mode: "0644",
|
||||
})
|
||||
return files, nil
|
||||
}
|
||||
|
||||
// taskRuntime returns the effective runtime for a task: the task's own
|
||||
// runtime when set, otherwise the top-level spec.Runtime (the per-group
|
||||
// default). Returns nil when neither is set.
|
||||
func taskRuntime(spec *jobspec.WorkloadSpec, task *jobspec.TaskGroupTask) *jobspec.RuntimeBlock {
|
||||
if task.Runtime != nil {
|
||||
return task.Runtime
|
||||
}
|
||||
return spec.Runtime
|
||||
}
|
||||
|
||||
// taskCommand returns the ExecStart command for a task. A task-level
|
||||
// Command takes precedence; otherwise the task's runtime command is
|
||||
// used; otherwise the top-level runtime command is used. Returns an
|
||||
// empty string when none is set.
|
||||
func taskCommand(spec *jobspec.WorkloadSpec, task *jobspec.TaskGroupTask, rt *jobspec.RuntimeBlock) string {
|
||||
if strings.TrimSpace(task.Command) != "" {
|
||||
return task.Command
|
||||
}
|
||||
if rt != nil && strings.TrimSpace(rt.Command) != "" {
|
||||
return rt.Command
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// renderTaskUnit renders a single per-task systemd [Unit]+[Service]
|
||||
// block. The unit is `PartOf=` the alloc target and
|
||||
// `WantedBy=multi-user.target` so it starts at boot and stops with the
|
||||
// group. The [Service] block carries the task's ExecStart and the
|
||||
// socket-plumbing lines derived from the spec's ports.
|
||||
func renderTaskUnit(spec *jobspec.WorkloadSpec, task *jobspec.TaskGroupTask, rt *jobspec.RuntimeBlock, cmd, targetUnit string) string {
|
||||
var b strings.Builder
|
||||
b.WriteString("[Unit]\n")
|
||||
b.WriteString(fmt.Sprintf("Description=orca alloc task %s\n", task.Name))
|
||||
b.WriteString(fmt.Sprintf("PartOf=%s\n", targetUnit))
|
||||
b.WriteString("\n[Service]\n")
|
||||
b.WriteString(fmt.Sprintf("ExecStart=%s\n", cmd))
|
||||
for _, line := range (SocketEmitter{}).RenderSocketLines(spec) {
|
||||
b.WriteString(line)
|
||||
b.WriteString("\n")
|
||||
}
|
||||
b.WriteString("\n[Install]\n")
|
||||
b.WriteString("WantedBy=multi-user.target\n")
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// renderTargetUnit renders the grouping target unit
|
||||
// (`orca-v1-alloc-<alloc-id>.target`) that starts/stops all tasks
|
||||
// together. The [Unit] block lists every per-task unit under Wants=
|
||||
// so `systemctl start <target>` brings them all up, and
|
||||
// `systemctl stop <target>` tears them down (PartOf= propagates stop).
|
||||
func renderTargetUnit(targetUnit string, spec *jobspec.WorkloadSpec, allocID string) string {
|
||||
var b strings.Builder
|
||||
b.WriteString("[Unit]\n")
|
||||
b.WriteString(fmt.Sprintf("Description=orca alloc %s task group\n", allocID))
|
||||
for _, task := range spec.Tasks {
|
||||
b.WriteString(fmt.Sprintf("Wants=%salloc-%s-%s.service\n", unitNamePrefix, allocID, task.Name))
|
||||
}
|
||||
b.WriteString("\n[Install]\n")
|
||||
b.WriteString("WantedBy=multi-user.target\n")
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// renderSystemdUnit renders the full [Service] block for the spec,
|
||||
// including ExecStart, lifecycle hooks (ExecStartPost, ExecStop), and
|
||||
// the R-007 socket-plumbing lines (RuntimeDirectory=, optional
|
||||
// TCP-bind ExecStartPre). The output is a single string with a
|
||||
// trailing newline per line.
|
||||
func renderSystemdUnit(spec *jobspec.WorkloadSpec) string {
|
||||
var b strings.Builder
|
||||
b.WriteString("[Service]\n")
|
||||
b.WriteString(fmt.Sprintf("ExecStart=%s\n", spec.Runtime.Command))
|
||||
// Lifecycle: post_start → ExecStartPost (runs after start).
|
||||
for _, cmd := range lifecyclePostStart(spec) {
|
||||
b.WriteString(fmt.Sprintf("ExecStartPost=%s\n", cmd))
|
||||
}
|
||||
// Lifecycle: pre_stop → ExecStop (runs before the process is killed).
|
||||
for _, cmd := range lifecyclePreStop(spec) {
|
||||
b.WriteString(fmt.Sprintf("ExecStop=%s\n", cmd))
|
||||
}
|
||||
// R-007 socket plumbing: RuntimeDirectory= per port + optional
|
||||
// TCP-bind ExecStartPre.
|
||||
for _, line := range (SocketEmitter{}).RenderSocketLines(spec) {
|
||||
b.WriteString(line)
|
||||
b.WriteString("\n")
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// lifecyclePostStart returns the post_start lifecycle commands for
|
||||
// the spec, or nil when the spec has no lifecycle block or no
|
||||
// post_start commands.
|
||||
func lifecyclePostStart(spec *jobspec.WorkloadSpec) []string {
|
||||
if spec.Lifecycle == nil {
|
||||
return nil
|
||||
}
|
||||
return spec.Lifecycle.PostStart
|
||||
}
|
||||
|
||||
// lifecyclePreStop returns the pre_stop lifecycle commands for the
|
||||
// spec, or nil when the spec has no lifecycle block or no pre_stop
|
||||
// commands.
|
||||
func lifecyclePreStop(spec *jobspec.WorkloadSpec) []string {
|
||||
if spec.Lifecycle == nil {
|
||||
return nil
|
||||
}
|
||||
return spec.Lifecycle.PreStop
|
||||
}
|
||||
@@ -0,0 +1,172 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
func TestSystemdEmitter_LifecyclePostStart(t *testing.T) {
|
||||
// lifecycle.post_start → ExecStartPost (one line per command).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/usr/local/bin/httpd -f"},
|
||||
Lifecycle: &jobspec.LifecycleBlock{
|
||||
PostStart: []string{"/usr/bin/sleep 1", "/usr/bin/curl localhost/healthz"},
|
||||
},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n1"})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if !strings.Contains(c, "ExecStartPost=/usr/bin/sleep 1\n") {
|
||||
t.Errorf("missing ExecStartPost for sleep 1\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "ExecStartPost=/usr/bin/curl localhost/healthz\n") {
|
||||
t.Errorf("missing ExecStartPost for curl\n%s", c)
|
||||
}
|
||||
// ExecStart must still be present.
|
||||
if !strings.Contains(c, "ExecStart=/usr/local/bin/httpd -f\n") {
|
||||
t.Errorf("missing ExecStart\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_LifecyclePreStop(t *testing.T) {
|
||||
// lifecycle.pre_stop → ExecStop (one line per command).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/usr/local/bin/httpd -f"},
|
||||
Lifecycle: &jobspec.LifecycleBlock{
|
||||
PreStop: []string{"/usr/local/bin/httpd -graceful", "/usr/bin/sleep 5"},
|
||||
},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n1"})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if !strings.Contains(c, "ExecStop=/usr/local/bin/httpd -graceful\n") {
|
||||
t.Errorf("missing ExecStop for graceful\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "ExecStop=/usr/bin/sleep 5\n") {
|
||||
t.Errorf("missing ExecStop for sleep 5\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_LifecycleBoth(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/httpd"},
|
||||
Lifecycle: &jobspec.LifecycleBlock{
|
||||
PostStart: []string{"/bin/after-start"},
|
||||
PreStop: []string{"/bin/before-stop"},
|
||||
},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
// ExecStartPost must appear before ExecStop (post_start runs after
|
||||
// start; pre_stop runs before stop — the order in the unit file
|
||||
// reflects the lifecycle order).
|
||||
startIdx := strings.Index(c, "ExecStart=")
|
||||
postIdx := strings.Index(c, "ExecStartPost=")
|
||||
stopIdx := strings.Index(c, "ExecStop=")
|
||||
if startIdx < 0 || postIdx < 0 || stopIdx < 0 {
|
||||
t.Fatalf("missing one of ExecStart/ExecStartPost/ExecStop\n%s", c)
|
||||
}
|
||||
if !(startIdx < postIdx && postIdx < stopIdx) {
|
||||
t.Errorf("expected order ExecStart < ExecStartPost < ExecStop\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_LifecycleNilOmitsDirectives(t *testing.T) {
|
||||
// No lifecycle block → no ExecStartPost / ExecStop lines.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "backup",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if strings.Contains(c, "ExecStartPost=") {
|
||||
t.Errorf("ExecStartPost should be omitted when no lifecycle\n%s", c)
|
||||
}
|
||||
if strings.Contains(c, "ExecStop=") {
|
||||
t.Errorf("ExecStop should be omitted when no lifecycle\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_LifecycleEmptyListsOmitted(t *testing.T) {
|
||||
// Lifecycle block present but empty lists → no ExecStartPost / ExecStop.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "x",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/x"},
|
||||
Lifecycle: &jobspec.LifecycleBlock{},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if strings.Contains(c, "ExecStartPost=") {
|
||||
t.Errorf("ExecStartPost should be omitted for empty PostStart\n%s", c)
|
||||
}
|
||||
if strings.Contains(c, "ExecStop=") {
|
||||
t.Errorf("ExecStop should be omitted for empty PreStop\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_LifecycleOnlyPostStart(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "x",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/x"},
|
||||
Lifecycle: &jobspec.LifecycleBlock{
|
||||
PostStart: []string{"/bin/notify-up"},
|
||||
},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if !strings.Contains(c, "ExecStartPost=/bin/notify-up\n") {
|
||||
t.Errorf("missing ExecStartPost\n%s", c)
|
||||
}
|
||||
if strings.Contains(c, "ExecStop=") {
|
||||
t.Errorf("ExecStop should be omitted when only PostStart set\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_LifecycleOnlyPreStop(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "x",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/x"},
|
||||
Lifecycle: &jobspec.LifecycleBlock{
|
||||
PreStop: []string{"/bin/notify-down"},
|
||||
},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if !strings.Contains(c, "ExecStop=/bin/notify-down\n") {
|
||||
t.Errorf("missing ExecStop\n%s", c)
|
||||
}
|
||||
if strings.Contains(c, "ExecStartPost=") {
|
||||
t.Errorf("ExecStartPost should be omitted when only PreStop set\n%s", c)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,343 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
func TestSystemdEmitter_RenderJob(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "backup",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/usr/bin/rsync -a /src /dst"},
|
||||
}
|
||||
node := &Node{Hostname: "node-1", Runtime: []string{"process"}}
|
||||
files, err := SystemdEmitter{}.Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(files) != 1 {
|
||||
t.Fatalf("got %d files, want 1", len(files))
|
||||
}
|
||||
f := files[0]
|
||||
wantPath := "/etc/systemd/system/orca-v1-backup.service"
|
||||
if f.Path != wantPath {
|
||||
t.Errorf("Path = %q, want %q", f.Path, wantPath)
|
||||
}
|
||||
wantContent := "[Service]\nExecStart=/usr/bin/rsync -a /src /dst\n"
|
||||
if f.Content != wantContent {
|
||||
t.Errorf("Content = %q, want %q", f.Content, wantContent)
|
||||
}
|
||||
if f.Mode != "0644" {
|
||||
t.Errorf("Mode = %q, want 0644", f.Mode)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_RenderService(t *testing.T) {
|
||||
// The full service emitter (Traefik route + health checks) lands in
|
||||
// P02; here we only prove the systemd side renders for a Service
|
||||
// kind with a process runtime.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/usr/local/bin/httpd -f"},
|
||||
}
|
||||
node := &Node{Hostname: "node-1", Runtime: []string{"process"}}
|
||||
files, err := SystemdEmitter{}.Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(files) != 1 {
|
||||
t.Fatalf("got %d files, want 1", len(files))
|
||||
}
|
||||
if files[0].Path != "/etc/systemd/system/orca-v1-web.service" {
|
||||
t.Errorf("Path = %q, want /etc/systemd/system/orca-v1-web.service", files[0].Path)
|
||||
}
|
||||
if !strings.Contains(files[0].Content, "ExecStart=/usr/local/bin/httpd -f") {
|
||||
t.Errorf("Content = %q, want it to contain the ExecStart line", files[0].Content)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_EmptyCommandError(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "x",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: ""},
|
||||
}
|
||||
_, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty command, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "command is empty") {
|
||||
t.Errorf("error = %q, want 'command is empty'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_WhitespaceCommandError(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "x",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: " "},
|
||||
}
|
||||
_, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for whitespace-only command, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "command is empty") {
|
||||
t.Errorf("error = %q, want 'command is empty'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_NilSpec(t *testing.T) {
|
||||
v := SystemdEmitter{}
|
||||
if _, err := v.Render(nil, &Node{}); err == nil {
|
||||
t.Fatal("expected error for nil spec")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_EmptyName(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: " ",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/x"},
|
||||
}
|
||||
_, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty name")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "name is empty") {
|
||||
t.Errorf("error = %q, want 'name is empty'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_NilRuntime(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "x"}
|
||||
_, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for nil runtime")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "runtime block is nil") {
|
||||
t.Errorf("error = %q, want 'runtime block is nil'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_UnitNamePrefix(t *testing.T) {
|
||||
// The orca-v1- prefix is load-bearing for the dual-write window
|
||||
// (REQ-090, I-C-006): the v0.8 daemon writes `orca-<job>.service`
|
||||
// and the v0.9 SSH-push path writes `orca-v1-<spec.Name>.service`,
|
||||
// so the two never collide. This test guards against accidental
|
||||
// removal of the prefix.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "dual-write-safety",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/true"},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n"})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if !strings.HasPrefix(files[0].Path, "/etc/systemd/system/orca-v1-") {
|
||||
t.Errorf("Path = %q, want it to start with /etc/systemd/system/orca-v1- (REQ-090)", files[0].Path)
|
||||
}
|
||||
if !strings.HasSuffix(files[0].Path, ".service") {
|
||||
t.Errorf("Path = %q, want it to end with .service", files[0].Path)
|
||||
}
|
||||
// Explicitly assert the full expected unit name to lock the contract.
|
||||
want := "/etc/systemd/system/orca-v1-dual-write-safety.service"
|
||||
if files[0].Path != want {
|
||||
t.Errorf("Path = %q, want %q", files[0].Path, want)
|
||||
}
|
||||
// Sanity: the prefix is exactly "orca-v1-", not "orca-v0" or "orca".
|
||||
if unitNamePrefix != "orca-v1-" {
|
||||
t.Errorf("unitNamePrefix = %q, want orca-v1-", unitNamePrefix)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_TaskGroupTwoTasks(t *testing.T) {
|
||||
// P06: a task group with two tasks renders one unit per task plus
|
||||
// a grouping target unit. Each per-task unit is
|
||||
// `orca-v1-alloc-<alloc-id>-<task-name>.service`, carries
|
||||
// `PartOf=orca-v1-alloc-<alloc-id>.target`, and is
|
||||
// `WantedBy=multi-user.target`. The target unit lists every
|
||||
// per-task unit under Wants=.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{
|
||||
Name: "app",
|
||||
Command: "/usr/bin/httpd -f",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
},
|
||||
{
|
||||
Name: "sidecar",
|
||||
Command: "/bin/wasm-runner sidecar.wasm",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "wasm"},
|
||||
},
|
||||
},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n1"})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
// 2 per-task units + 1 target unit.
|
||||
if len(files) != 3 {
|
||||
t.Fatalf("got %d files, want 3 (2 per-task units + 1 target)", len(files))
|
||||
}
|
||||
wantApp := "/etc/systemd/system/orca-v1-alloc-web-app.service"
|
||||
wantSide := "/etc/systemd/system/orca-v1-alloc-web-sidecar.service"
|
||||
wantTarget := "/etc/systemd/system/orca-v1-alloc-web.target"
|
||||
paths := make(map[string]*File, len(files))
|
||||
for i := range files {
|
||||
paths[files[i].Path] = &files[i]
|
||||
}
|
||||
if _, ok := paths[wantApp]; !ok {
|
||||
t.Errorf("missing per-task unit %q; got paths %v", wantApp, filePaths(files))
|
||||
}
|
||||
if _, ok := paths[wantSide]; !ok {
|
||||
t.Errorf("missing per-task unit %q; got paths %v", wantSide, filePaths(files))
|
||||
}
|
||||
if _, ok := paths[wantTarget]; !ok {
|
||||
t.Errorf("missing target unit %q; got paths %v", wantTarget, filePaths(files))
|
||||
}
|
||||
if _, ok := paths[wantTarget]; !ok {
|
||||
t.Errorf("missing target unit %q; got paths %v", wantTarget, filePaths(files))
|
||||
}
|
||||
// Verify PartOf relations and ExecStart on per-task units.
|
||||
app := paths[wantApp]
|
||||
if !strings.Contains(app.Content, "PartOf=orca-v1-alloc-web.target") {
|
||||
t.Errorf("app unit missing PartOf=orca-v1-alloc-web.target\n%s", app.Content)
|
||||
}
|
||||
if !strings.Contains(app.Content, "ExecStart=/usr/bin/httpd -f") {
|
||||
t.Errorf("app unit missing ExecStart=/usr/bin/httpd -f\n%s", app.Content)
|
||||
}
|
||||
if !strings.Contains(app.Content, "WantedBy=multi-user.target") {
|
||||
t.Errorf("app unit missing WantedBy=multi-user.target\n%s", app.Content)
|
||||
}
|
||||
side := paths[wantSide]
|
||||
if !strings.Contains(side.Content, "PartOf=orca-v1-alloc-web.target") {
|
||||
t.Errorf("sidecar unit missing PartOf=orca-v1-alloc-web.target\n%s", side.Content)
|
||||
}
|
||||
if !strings.Contains(side.Content, "ExecStart=/bin/wasm-runner sidecar.wasm") {
|
||||
t.Errorf("sidecar unit missing ExecStart\n%s", side.Content)
|
||||
}
|
||||
// Verify the target unit Wants= both per-task units.
|
||||
target := paths[wantTarget]
|
||||
if !strings.Contains(target.Content, "Wants=orca-v1-alloc-web-app.service") {
|
||||
t.Errorf("target missing Wants=...app.service\n%s", target.Content)
|
||||
}
|
||||
if !strings.Contains(target.Content, "Wants=orca-v1-alloc-web-sidecar.service") {
|
||||
t.Errorf("target missing Wants=...sidecar.service\n%s", target.Content)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_TaskGroupInheritsTopLevelRuntime(t *testing.T) {
|
||||
// P06: a task that omits its own runtime inherits the top-level
|
||||
// spec.Runtime as the per-group default. The per-task unit's
|
||||
// ExecStart must come from the top-level runtime command when
|
||||
// the task has no own command and no own runtime.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/default"},
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{Name: "app"},
|
||||
{Name: "sidecar", Command: "/bin/override"},
|
||||
},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(files) != 3 {
|
||||
t.Fatalf("got %d files, want 3", len(files))
|
||||
}
|
||||
appContent := findUnitContent(files, "/etc/systemd/system/orca-v1-alloc-web-app.service")
|
||||
if appContent == "" {
|
||||
t.Fatalf("missing app unit; paths %v", filePaths(files))
|
||||
}
|
||||
if !strings.Contains(appContent, "ExecStart=/bin/default") {
|
||||
t.Errorf("app unit should inherit top-level command /bin/default\n%s", appContent)
|
||||
}
|
||||
sideContent := findUnitContent(files, "/etc/systemd/system/orca-v1-alloc-web-sidecar.service")
|
||||
if sideContent == "" {
|
||||
t.Fatalf("missing sidecar unit; paths %v", filePaths(files))
|
||||
}
|
||||
if !strings.Contains(sideContent, "ExecStart=/bin/override") {
|
||||
t.Errorf("sidecar unit should use its own command /bin/override\n%s", sideContent)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_TaskGroupNoCommandError(t *testing.T) {
|
||||
// P06: a task with no resolvable command (no task.Command, no
|
||||
// task.Runtime, no top-level Runtime) is an error.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Tasks: []jobspec.TaskGroupTask{{Name: "app"}},
|
||||
}
|
||||
_, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for task with no runtime, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "no runtime") {
|
||||
t.Errorf("error = %q, want 'no runtime'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_TaskGroupEmptyCommandError(t *testing.T) {
|
||||
// P06: a task whose resolved runtime command is empty/whitespace
|
||||
// is an error (mirrors the single-process rule).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: " "},
|
||||
Tasks: []jobspec.TaskGroupTask{{Name: "app"}},
|
||||
}
|
||||
_, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty command, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "command is empty") {
|
||||
t.Errorf("error = %q, want 'command is empty'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestSystemdEmitter_NoTasksBackwardCompat(t *testing.T) {
|
||||
// Backward compat: a spec with no Tasks renders exactly one unit
|
||||
// (the historical single-process shape).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "backup",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
|
||||
}
|
||||
files, err := SystemdEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(files) != 1 {
|
||||
t.Fatalf("got %d files, want 1 (backward compat)", len(files))
|
||||
}
|
||||
if files[0].Path != "/etc/systemd/system/orca-v1-backup.service" {
|
||||
t.Errorf("Path = %q, want /etc/systemd/system/orca-v1-backup.service", files[0].Path)
|
||||
}
|
||||
}
|
||||
|
||||
func filePaths(files []File) []string {
|
||||
out := make([]string, len(files))
|
||||
for i, f := range files {
|
||||
out[i] = f.Path
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func findUnitContent(files []File, path string) string {
|
||||
for _, f := range files {
|
||||
if f.Path == path {
|
||||
return f.Content
|
||||
}
|
||||
}
|
||||
return ""
|
||||
}
|
||||
@@ -0,0 +1,225 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"net"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// TraefikEmitter is the Layer-4 emitter for the Traefik dynamic-config
|
||||
// file (REQ-077). It renders /etc/traefik/dynamic/orca-<spec.Name>.yaml
|
||||
// — a single Traefik dynamic-config file describing the routers,
|
||||
// services (servers = the R-007 socket paths), TLS config pointing at
|
||||
// the step-ca root CA, and the service health check.
|
||||
//
|
||||
// Registered on the emitter.Registry under the service-kind keys:
|
||||
//
|
||||
// - service:process
|
||||
// - service:podman
|
||||
// - service:wasm
|
||||
//
|
||||
// RegisterTraefik wires all three; callers can also call Register
|
||||
// directly with TraefikEmitter{} for a single runtime.
|
||||
//
|
||||
// Atomic reload (gate C-10): the Traefik dynamic-config file is written
|
||||
// atomically via the SSH-push transport (sshpush.WriteFileIdempotent
|
||||
// performs temp-file + fsync + rename, and WriteTraefikDynamic wraps
|
||||
// it with an explicit tmp+mv so fsnotify sees a single rename event).
|
||||
// Traefik watches the dynamic dir with fsnotify; the rename triggers a
|
||||
// reload. On a malformed config Traefik logs an error and holds the
|
||||
// last-good config (documented Traefik behavior; the C-10 test
|
||||
// verifies the tmp+rename sequence so a half-written file is never
|
||||
// observed by Traefik). Drain is rendered by setting the backend
|
||||
// server's weight to 0 (or removing it) — see RenderDrain.
|
||||
//
|
||||
// The orca-v1- prefix is NOT applied to Traefik dynamic-config paths
|
||||
// (the prefix is only for systemd unit names; the Traefik file is named
|
||||
// orca-<spec.Name>.yaml and is the single source of truth for the
|
||||
// service route — there is no dual-write window for Traefik configs).
|
||||
type TraefikEmitter struct{}
|
||||
|
||||
// traefikDynamicDir is the canonical Traefik dynamic-config directory
|
||||
// (R-006). The emitter writes one file per service at
|
||||
// /etc/traefik/dynamic/orca-<spec.Name>.yaml.
|
||||
const traefikDynamicDir = "/etc/traefik/dynamic"
|
||||
|
||||
// traefikRouterTLSCertResolver is the Traefik cert-resolver name that
|
||||
// the orca step-ca integration configures on the Traefik static config
|
||||
// (P10 / v0.10 wires the step-ca root into this resolver). The
|
||||
// dynamic-config file references it by name.
|
||||
const traefikRouterTLSCertResolver = "orca"
|
||||
|
||||
// defaultTrustDomain is the SPIFFE trust domain used in the rendered
|
||||
// TLS stanza when the spec does not carry an explicit trust domain.
|
||||
// The step-ca provisioner (P10) overrides this at render time via the
|
||||
// node argument; for P02 the emitter renders the placeholder.
|
||||
const defaultTrustDomain = "cluster.orca.local"
|
||||
|
||||
// Render renders the Traefik dynamic-config YAML for a Service
|
||||
// workload. The output is a single File whose Path is
|
||||
// /etc/traefik/dynamic/orca-<spec.Name>.yaml, Content is the rendered
|
||||
// YAML, and Mode is 0644.
|
||||
//
|
||||
// Returns an error if the spec is nil, the name is empty, the spec has
|
||||
// no ports (a Service with no ports has no backends to route to), or a
|
||||
// service.bind value (when present) is not a valid IP address (R-007).
|
||||
func (TraefikEmitter) Render(spec *jobspec.WorkloadSpec, node *Node) ([]File, error) {
|
||||
if spec == nil {
|
||||
return nil, errors.New("emitter/traefik: spec is nil")
|
||||
}
|
||||
if strings.TrimSpace(spec.Name) == "" {
|
||||
return nil, errors.New("emitter/traefik: spec name is empty")
|
||||
}
|
||||
if len(spec.Ports) == 0 {
|
||||
return nil, errors.New("emitter/traefik: service has no ports (no backends to route to)")
|
||||
}
|
||||
if spec.Service != nil {
|
||||
if b := strings.TrimSpace(spec.Service.Bind); b != "" && net.ParseIP(b) == nil {
|
||||
return nil, fmt.Errorf("emitter/traefik: service.bind %q is not a valid IP (R-007)", b)
|
||||
}
|
||||
}
|
||||
content, err := renderTraefikYAML(spec, node)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
path := fmt.Sprintf("%s/orca-%s.yaml", traefikDynamicDir, spec.Name)
|
||||
return []File{{Path: path, Content: content, Mode: "0644"}}, nil
|
||||
}
|
||||
|
||||
// RenderDrain renders a Traefik dynamic-config that drains the service
|
||||
// by setting every backend server's weight to 0 (I-B-005 drain). The
|
||||
// path matches the live config so the atomic rename overwrites the
|
||||
// routing config with the drained config (Traefik reloads and stops
|
||||
// sending traffic). The caller writes the result via
|
||||
// WriteTraefikDynamic for the C-10 atomicity protocol.
|
||||
func (e TraefikEmitter) RenderDrain(spec *jobspec.WorkloadSpec, node *Node) ([]File, error) {
|
||||
if spec == nil {
|
||||
return nil, errors.New("emitter/traefik: spec is nil")
|
||||
}
|
||||
if strings.TrimSpace(spec.Name) == "" {
|
||||
return nil, errors.New("emitter/traefik: spec name is empty")
|
||||
}
|
||||
if len(spec.Ports) == 0 {
|
||||
return nil, errors.New("emitter/traefik: service has no ports (no backends to drain)")
|
||||
}
|
||||
content, err := renderTraefikYAMLDrain(spec, node)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
path := fmt.Sprintf("%s/orca-%s.yaml", traefikDynamicDir, spec.Name)
|
||||
return []File{{Path: path, Content: content, Mode: "0644"}}, nil
|
||||
}
|
||||
|
||||
// RegisterTraefik registers the TraefikEmitter on the given Registry
|
||||
// under the three service-kind runtime keys (service:process,
|
||||
// service:podman, service:wasm). The emitter is the same instance for
|
||||
// all three runtimes — the rendered Traefik config is runtime-agnostic
|
||||
// (the backend server URL is the R-007 socket path, which the runtime
|
||||
// layer binds regardless of process/wasm/podman).
|
||||
func RegisterTraefik(reg *Registry) {
|
||||
e := TraefikEmitter{}
|
||||
reg.Register("service:process", e)
|
||||
reg.Register("service:podman", e)
|
||||
reg.Register("service:wasm", e)
|
||||
}
|
||||
|
||||
// renderTraefikYAML renders the Traefik dynamic-config YAML for the
|
||||
// given spec + node. The shape (verified by the Traefik docs) is:
|
||||
//
|
||||
// http:
|
||||
// routers:
|
||||
// orca-<name>:
|
||||
// rule: PathPrefix("/<name>")
|
||||
// service: orca-<name>
|
||||
// tls:
|
||||
// certResolver: orca
|
||||
// domains:
|
||||
// - main: "<trust-domain>"
|
||||
// services:
|
||||
// orca-<name>:
|
||||
// loadBalancer:
|
||||
// servers:
|
||||
// - url: "unix:///run/orca/alloc-<allocID>/port-<portName>.sock"
|
||||
// healthCheck:
|
||||
// path: /healthz
|
||||
// interval: <interval>
|
||||
// timeout: <timeout>
|
||||
//
|
||||
// The alloc-id placeholder is "<allocID>" pending the P08 socket
|
||||
// layer; Traefik will reject the URL until a real alloc-id is
|
||||
// substituted. For P02 the emitter renders the placeholder so the
|
||||
// C-10 atomicity protocol is testable end-to-end; the socket layer
|
||||
// (P08) replaces the placeholder with the live alloc-id.
|
||||
func renderTraefikYAML(spec *jobspec.WorkloadSpec, node *Node) (string, error) {
|
||||
return renderTraefikYAMLWeighted(spec, node, false)
|
||||
}
|
||||
|
||||
// renderTraefikYAMLDrain renders the drained Traefik dynamic-config
|
||||
// (every backend server has weight: 0). The shape mirrors the live
|
||||
// config so the rename overwrites the live route with the drain.
|
||||
func renderTraefikYAMLDrain(spec *jobspec.WorkloadSpec, node *Node) (string, error) {
|
||||
return renderTraefikYAMLWeighted(spec, node, true)
|
||||
}
|
||||
|
||||
// renderTraefikYAMLWeighted renders the Traefik dynamic-config YAML.
|
||||
// When drain is true, every server entry is emitted with `weight: 0`
|
||||
// (I-B-005). When drain is false, no weight is emitted (Traefik
|
||||
// defaults to 1 — equal weighting across servers).
|
||||
func renderTraefikYAMLWeighted(spec *jobspec.WorkloadSpec, node *Node, drain bool) (string, error) {
|
||||
var b strings.Builder
|
||||
routerName := "orca-" + spec.Name
|
||||
serviceName := "orca-" + spec.Name
|
||||
rule := fmt.Sprintf("PathPrefix(\"/%s\")", spec.Name)
|
||||
trustDomain := defaultTrustDomain
|
||||
|
||||
b.WriteString("http:\n")
|
||||
b.WriteString(" routers:\n")
|
||||
b.WriteString(fmt.Sprintf(" %s:\n", routerName))
|
||||
b.WriteString(fmt.Sprintf(" rule: %s\n", rule))
|
||||
b.WriteString(fmt.Sprintf(" service: %s\n", serviceName))
|
||||
b.WriteString(" tls:\n")
|
||||
b.WriteString(fmt.Sprintf(" certResolver: %s\n", traefikRouterTLSCertResolver))
|
||||
b.WriteString(" domains:\n")
|
||||
b.WriteString(fmt.Sprintf(" - main: %q\n", trustDomain))
|
||||
b.WriteString(" services:\n")
|
||||
b.WriteString(fmt.Sprintf(" %s:\n", serviceName))
|
||||
b.WriteString(" loadBalancer:\n")
|
||||
b.WriteString(" servers:\n")
|
||||
allocID := allocIDFor(node)
|
||||
for _, p := range spec.Ports {
|
||||
sock := fmt.Sprintf("unix:///run/orca/alloc-%s/port-%s.sock", allocID, p.Name)
|
||||
b.WriteString(" - url: ")
|
||||
b.WriteString(fmt.Sprintf("%q\n", sock))
|
||||
if drain {
|
||||
b.WriteString(" weight: 0\n")
|
||||
}
|
||||
}
|
||||
if spec.Health != nil {
|
||||
b.WriteString(" healthCheck:\n")
|
||||
path := "/healthz"
|
||||
b.WriteString(fmt.Sprintf(" path: %s\n", path))
|
||||
if spec.Health.Interval != "" {
|
||||
b.WriteString(fmt.Sprintf(" interval: %s\n", spec.Health.Interval))
|
||||
}
|
||||
if spec.Health.Timeout != "" {
|
||||
b.WriteString(fmt.Sprintf(" timeout: %s\n", spec.Health.Timeout))
|
||||
}
|
||||
}
|
||||
return b.String(), nil
|
||||
}
|
||||
|
||||
// allocIDFor returns the alloc-id placeholder for the node. P08 will
|
||||
// substitute the live alloc-id from the socket layer; for P02 we use a
|
||||
// deterministic placeholder derived from the node hostname so the
|
||||
// rendered config is stable across re-renders (the C-10 idempotency
|
||||
// check depends on a stable hash). When the node is nil or has no
|
||||
// hostname, the literal placeholder "<allocID>" is emitted.
|
||||
func allocIDFor(node *Node) string {
|
||||
if node == nil || strings.TrimSpace(node.Hostname) == "" {
|
||||
return "<allocID>"
|
||||
}
|
||||
return node.Hostname
|
||||
}
|
||||
@@ -0,0 +1,93 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"os"
|
||||
)
|
||||
|
||||
// AtomicWriter is the SSH-push transport surface that
|
||||
// WriteTraefikDynamic uses to write the Traefik dynamic-config file
|
||||
// atomically. It is the subset of *sshpush.Transport that the
|
||||
// atomicity protocol depends on. Tests substitute a mock to assert
|
||||
// the tmp+rename sequence (gate C-10) without a real SSH server.
|
||||
//
|
||||
// *sshpush.Transport satisfies this interface (the compile-time
|
||||
// assertion lives in internal/sshpush to avoid an import cycle — the
|
||||
// sshpush package imports emitter for fan-out, so this package cannot
|
||||
// import sshpush).
|
||||
type AtomicWriter interface {
|
||||
// WriteFileIdempotent writes content to peer:path atomically with
|
||||
// mode, returning written=true if the file was actually written
|
||||
// (content hash differed). Used by WriteTraefikDynamic to write
|
||||
// the .tmp sibling.
|
||||
WriteFileIdempotent(ctx context.Context, peer string, path string, content []byte, mode os.FileMode) (bool, error)
|
||||
// Exec runs a command on peer and returns its combined output.
|
||||
// Used by WriteTraefikDynamic to perform the atomic `mv -f
|
||||
// path.tmp path`.
|
||||
Exec(ctx context.Context, peer string, cmd string) ([]byte, error)
|
||||
}
|
||||
|
||||
// WriteTraefikDynamic writes a Traefik dynamic-config file atomically
|
||||
// (gate C-10: tmpfile + fsync + rename). The protocol is:
|
||||
//
|
||||
// 1. Write content to <path>.tmp via WriteFileIdempotent. The
|
||||
// underlying sshpush transport writes the tmp file in the same
|
||||
// directory as the target with mode-appended naming, fsyncs, and
|
||||
// renames — but we add an extra hop here so the *Traefik* file is
|
||||
// only ever observed at its final path after a single atomic
|
||||
// rename event that Traefik's fsnotify watcher sees.
|
||||
// 2. `mv -f <path>.tmp <path>` on the peer (atomic rename on POSIX).
|
||||
// Traefik's fsnotify watcher picks up the rename → reload.
|
||||
//
|
||||
// On a malformed config Traefik logs an error and holds the
|
||||
// last-good config (documented Traefik behavior; the C-10 test
|
||||
// verifies the tmp+rename sequence so a half-written file is never
|
||||
// observed by Traefik — the only window where Traefik can read the
|
||||
// file is after the rename, which is atomic on POSIX).
|
||||
//
|
||||
// The mode is 0644 (Traefik reads the dynamic dir as root; the lead
|
||||
// applier chmods after the rename).
|
||||
func WriteTraefikDynamic(ctx context.Context, t AtomicWriter, peer string, path string, content []byte) error {
|
||||
if t == nil {
|
||||
return fmt.Errorf("traefik: atomic writer is nil")
|
||||
}
|
||||
if path == "" {
|
||||
return fmt.Errorf("traefik: path is empty")
|
||||
}
|
||||
tmpPath := path + ".tmp"
|
||||
if _, err := t.WriteFileIdempotent(ctx, peer, tmpPath, content, 0o644); err != nil {
|
||||
return fmt.Errorf("traefik: write tmp %s: %w", tmpPath, err)
|
||||
}
|
||||
// Atomic rename on POSIX. `mv -f` overwrites an existing target
|
||||
// without prompting. The rename is atomic; Traefik's fsnotify
|
||||
// watcher observes a single IN_MOVED_TO event.
|
||||
renameCmd := fmt.Sprintf("mv -f %s %s", shellQuoteLocal(tmpPath), shellQuoteLocal(path))
|
||||
if _, err := t.Exec(ctx, peer, renameCmd); err != nil {
|
||||
return fmt.Errorf("traefik: rename %s -> %s: %w", tmpPath, path, err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// shellQuoteLocal single-quotes a path for safe shell interpolation on
|
||||
// the peer. It escapes embedded single-quotes via the standard '\”
|
||||
// idiom (close the single-quoted string, escape the literal single
|
||||
// quote, reopen the single-quoted string). This is a local
|
||||
// re-implementation (the sshpush package has its own) so the emitter
|
||||
// layer does not depend on the transport package's private helpers —
|
||||
// the AtomicWriter interface keeps the boundary clean for testing.
|
||||
func shellQuoteLocal(s string) string {
|
||||
var b []byte
|
||||
b = append(b, '\'')
|
||||
for i := 0; i < len(s); i++ {
|
||||
c := s[i]
|
||||
if c == '\'' {
|
||||
// close quote, escape the literal single-quote, reopen.
|
||||
b = append(b, '\'', '\\', '\'', '\'')
|
||||
continue
|
||||
}
|
||||
b = append(b, c)
|
||||
}
|
||||
b = append(b, '\'')
|
||||
return string(b)
|
||||
}
|
||||
@@ -0,0 +1,190 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// mockAtomicWriter is a test-only AtomicWriter that records calls so
|
||||
// the C-10 atomicity protocol (tmp + rename) can be asserted.
|
||||
type mockAtomicWriter struct {
|
||||
written []writeCall
|
||||
execed []execCall
|
||||
writeErr error
|
||||
writeWrote bool
|
||||
execErr error
|
||||
}
|
||||
|
||||
type writeCall struct {
|
||||
peer string
|
||||
path string
|
||||
mode os.FileMode
|
||||
bytes []byte
|
||||
}
|
||||
|
||||
type execCall struct {
|
||||
peer string
|
||||
cmd string
|
||||
}
|
||||
|
||||
func (m *mockAtomicWriter) WriteFileIdempotent(ctx context.Context, peer string, path string, content []byte, mode os.FileMode) (bool, error) {
|
||||
m.written = append(m.written, writeCall{peer: peer, path: path, mode: mode, bytes: append([]byte(nil), content...)})
|
||||
if m.writeErr != nil {
|
||||
return false, m.writeErr
|
||||
}
|
||||
return m.writeWrote, nil
|
||||
}
|
||||
|
||||
func (m *mockAtomicWriter) Exec(ctx context.Context, peer string, cmd string) ([]byte, error) {
|
||||
m.execed = append(m.execed, execCall{peer: peer, cmd: cmd})
|
||||
if m.execErr != nil {
|
||||
return nil, m.execErr
|
||||
}
|
||||
return []byte("ok"), nil
|
||||
}
|
||||
|
||||
func TestWriteTraefikDynamic_TmpThenRename(t *testing.T) {
|
||||
// Gate C-10: the Traefik dynamic-config write must be a tmp +
|
||||
// rename sequence so Traefik's fsnotify watcher never observes a
|
||||
// half-written file.
|
||||
mock := &mockAtomicWriter{writeWrote: true}
|
||||
path := "/etc/traefik/dynamic/orca-web.yaml"
|
||||
peer := "node-1:22"
|
||||
content := []byte("http:\n routers: {}\n")
|
||||
|
||||
if err := WriteTraefikDynamic(context.Background(), mock, peer, path, content); err != nil {
|
||||
t.Fatalf("WriteTraefikDynamic: %v", err)
|
||||
}
|
||||
|
||||
if len(mock.written) != 1 {
|
||||
t.Fatalf("WriteFileIdempotent calls = %d, want 1", len(mock.written))
|
||||
}
|
||||
w := mock.written[0]
|
||||
if w.peer != peer {
|
||||
t.Errorf("write peer = %q, want %q", w.peer, peer)
|
||||
}
|
||||
// The tmp path is the target path + ".tmp".
|
||||
if w.path != path+".tmp" {
|
||||
t.Errorf("write path = %q, want %q (.tmp suffix is the C-10 atomicity protocol)", w.path, path+".tmp")
|
||||
}
|
||||
if string(w.bytes) != string(content) {
|
||||
t.Errorf("write content = %q, want %q", string(w.bytes), string(content))
|
||||
}
|
||||
if w.mode != 0o644 {
|
||||
t.Errorf("write mode = %o, want 0644", w.mode)
|
||||
}
|
||||
|
||||
if len(mock.execed) != 1 {
|
||||
t.Fatalf("Exec calls = %d, want 1 (the rename)", len(mock.execed))
|
||||
}
|
||||
e := mock.execed[0]
|
||||
if e.peer != peer {
|
||||
t.Errorf("exec peer = %q, want %q", e.peer, peer)
|
||||
}
|
||||
// The rename command must `mv -f` the .tmp file to the final path.
|
||||
if !strings.Contains(e.cmd, "mv -f") {
|
||||
t.Errorf("exec cmd = %q, want it to contain 'mv -f' (atomic rename)", e.cmd)
|
||||
}
|
||||
if !strings.Contains(e.cmd, path+".tmp") {
|
||||
t.Errorf("exec cmd = %q, want it to contain the .tmp path as source", e.cmd)
|
||||
}
|
||||
if !strings.Contains(e.cmd, path) {
|
||||
t.Errorf("exec cmd = %q, want it to contain the final path as destination", e.cmd)
|
||||
}
|
||||
// Sanity: the source must come before the destination in the
|
||||
// mv command.
|
||||
srcIdx := strings.Index(e.cmd, path+".tmp")
|
||||
dstIdx := strings.Index(e.cmd, "'"+path+"'")
|
||||
if srcIdx < 0 || dstIdx < 0 || srcIdx > dstIdx {
|
||||
t.Errorf("exec cmd %q: source .tmp must come before destination %s", e.cmd, path)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteTraefikDynamic_WriteTmpError(t *testing.T) {
|
||||
mock := &mockAtomicWriter{writeErr: errors.New("disk full")}
|
||||
err := WriteTraefikDynamic(context.Background(), mock, "p", "/etc/traefik/dynamic/orca-x.yaml", []byte("x"))
|
||||
if err == nil {
|
||||
t.Fatal("expected error from WriteFileIdempotent, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "write tmp") {
|
||||
t.Errorf("error = %q, want 'write tmp'", err.Error())
|
||||
}
|
||||
if !strings.Contains(err.Error(), "disk full") {
|
||||
t.Errorf("error = %q, want underlying 'disk full'", err.Error())
|
||||
}
|
||||
if len(mock.execed) != 0 {
|
||||
t.Errorf("on tmp write failure, no rename should happen; execed = %v", mock.execed)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteTraefikDynamic_RenameError(t *testing.T) {
|
||||
mock := &mockAtomicWriter{writeWrote: true, execErr: errors.New("permission denied")}
|
||||
err := WriteTraefikDynamic(context.Background(), mock, "p", "/etc/traefik/dynamic/orca-x.yaml", []byte("x"))
|
||||
if err == nil {
|
||||
t.Fatal("expected error from rename, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "rename") {
|
||||
t.Errorf("error = %q, want 'rename'", err.Error())
|
||||
}
|
||||
if !strings.Contains(err.Error(), "permission denied") {
|
||||
t.Errorf("error = %q, want underlying 'permission denied'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteTraefikDynamic_NilWriter(t *testing.T) {
|
||||
err := WriteTraefikDynamic(context.Background(), nil, "p", "/x", []byte("x"))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for nil writer")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "nil") {
|
||||
t.Errorf("error = %q, want 'nil'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteTraefikDynamic_EmptyPath(t *testing.T) {
|
||||
mock := &mockAtomicWriter{writeWrote: true}
|
||||
err := WriteTraefikDynamic(context.Background(), mock, "p", "", []byte("x"))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty path")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "path is empty") {
|
||||
t.Errorf("error = %q, want 'path is empty'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteTraefikDynamic_SkipWhenContentMatches(t *testing.T) {
|
||||
// When the .tmp file already matches (writeWrote=false), the
|
||||
// protocol still proceeds with the rename — the idempotency
|
||||
// check is per-file, not per-protocol. The rename still happens
|
||||
// so the final path reflects the (unchanged) content.
|
||||
mock := &mockAtomicWriter{writeWrote: false}
|
||||
err := WriteTraefikDynamic(context.Background(), mock, "p", "/etc/traefik/dynamic/orca-x.yaml", []byte("x"))
|
||||
if err != nil {
|
||||
t.Fatalf("WriteTraefikDynamic: %v", err)
|
||||
}
|
||||
if len(mock.execed) != 1 {
|
||||
t.Errorf("rename should still happen on idempotent skip; execed = %v", mock.execed)
|
||||
}
|
||||
}
|
||||
|
||||
func TestShellQuoteLocal(t *testing.T) {
|
||||
cases := []struct {
|
||||
in, want string
|
||||
}{
|
||||
{"/etc/traefik/dynamic/orca-web.yaml", "'/etc/traefik/dynamic/orca-web.yaml'"},
|
||||
{"", "''"},
|
||||
{"/path with space/x", "'/path with space/x'"},
|
||||
{"a'b", "'a'\\''b'"},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.in, func(t *testing.T) {
|
||||
got := shellQuoteLocal(tc.in)
|
||||
if got != tc.want {
|
||||
t.Errorf("shellQuoteLocal(%q) = %q, want %q", tc.in, got, tc.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,353 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
func TestTraefikEmitter_RenderBasic(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http", Interval: "5s", Timeout: "1s"},
|
||||
}
|
||||
node := &Node{Hostname: "node-1", Runtime: []string{"process"}}
|
||||
files, err := TraefikEmitter{}.Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if len(files) != 1 {
|
||||
t.Fatalf("got %d files, want 1", len(files))
|
||||
}
|
||||
f := files[0]
|
||||
wantPath := "/etc/traefik/dynamic/orca-web.yaml"
|
||||
if f.Path != wantPath {
|
||||
t.Errorf("Path = %q, want %q", f.Path, wantPath)
|
||||
}
|
||||
if f.Mode != "0644" {
|
||||
t.Errorf("Mode = %q, want 0644", f.Mode)
|
||||
}
|
||||
c := f.Content
|
||||
if !strings.Contains(c, "http:") {
|
||||
t.Errorf("content missing 'http:'\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "routers:") {
|
||||
t.Errorf("content missing 'routers:'\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "orca-web:") {
|
||||
t.Errorf("content missing 'orca-web:' router/service key\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, `rule: PathPrefix("/web")`) {
|
||||
t.Errorf("content missing PathPrefix rule\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "services:") {
|
||||
t.Errorf("content missing 'services:'\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "loadBalancer:") {
|
||||
t.Errorf("content missing 'loadBalancer:'\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "unix:///run/orca/alloc-node-1/port-http.sock") {
|
||||
t.Errorf("content missing socket server URL\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "certResolver: orca") {
|
||||
t.Errorf("content missing 'certResolver: orca'\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "domains:") {
|
||||
t.Errorf("content missing TLS domains\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "healthCheck:") {
|
||||
t.Errorf("content missing 'healthCheck:'\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "interval: 5s") {
|
||||
t.Errorf("content missing 'interval: 5s'\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "timeout: 1s") {
|
||||
t.Errorf("content missing 'timeout: 1s'\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_RenderMultiplePorts(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "api",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Ports: []jobspec.PortSpec{
|
||||
{Name: "http", Port: 8080},
|
||||
{Name: "grpc", Port: 9090},
|
||||
},
|
||||
}
|
||||
node := &Node{Hostname: "n1"}
|
||||
files, err := TraefikEmitter{}.Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
if !strings.Contains(c, "port-http.sock") {
|
||||
t.Errorf("missing http socket: %s", c)
|
||||
}
|
||||
if !strings.Contains(c, "port-grpc.sock") {
|
||||
t.Errorf("missing grpc socket: %s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_RenderDrain(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
node := &Node{Hostname: "n1"}
|
||||
files, err := TraefikEmitter{}.RenderDrain(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderDrain: %v", err)
|
||||
}
|
||||
if len(files) != 1 {
|
||||
t.Fatalf("got %d files, want 1", len(files))
|
||||
}
|
||||
c := files[0].Content
|
||||
if !strings.Contains(c, "weight: 0") {
|
||||
t.Errorf("drain config missing 'weight: 0'\n%s", c)
|
||||
}
|
||||
if !strings.Contains(c, "unix:///run/orca/alloc-n1/port-http.sock") {
|
||||
t.Errorf("drain config missing socket URL\n%s", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_RenderLiveHasNoWeightZero(t *testing.T) {
|
||||
// Sanity: the live (non-drain) render must NOT emit `weight: 0`.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
node := &Node{Hostname: "n1"}
|
||||
files, err := TraefikEmitter{}.Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if strings.Contains(files[0].Content, "weight: 0") {
|
||||
t.Errorf("live config should not contain 'weight: 0'\n%s", files[0].Content)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_RenderNoHealthOmitsHealthCheck(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
node := &Node{Hostname: "n1"}
|
||||
files, err := TraefikEmitter{}.Render(spec, node)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if strings.Contains(files[0].Content, "healthCheck:") {
|
||||
t.Errorf("config without Health should omit 'healthCheck:'\n%s", files[0].Content)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_NilSpec(t *testing.T) {
|
||||
_, err := TraefikEmitter{}.Render(nil, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for nil spec")
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_EmptyName(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: " ",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
_, err := TraefikEmitter{}.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty name")
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_NoPorts(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
}
|
||||
_, err := TraefikEmitter{}.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing ports")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "no ports") {
|
||||
t.Errorf("error = %q, want 'no ports'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_NoPortsDrain(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
}
|
||||
_, err := TraefikEmitter{}.RenderDrain(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing ports on drain")
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_InvalidBind(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: "not-an-ip"},
|
||||
}
|
||||
_, err := TraefikEmitter{}.Render(spec, &Node{})
|
||||
if err == nil {
|
||||
t.Fatal("expected error for invalid service.bind")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "valid IP") {
|
||||
t.Errorf("error = %q, want 'valid IP'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_ValidBindLoopback(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: "127.0.0.1"},
|
||||
}
|
||||
_, err := TraefikEmitter{}.Render(spec, &Node{})
|
||||
if err != nil {
|
||||
t.Fatalf("127.0.0.1 should be accepted, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_NilNodeAllocPlaceholder(t *testing.T) {
|
||||
// With a nil node, the alloc-id placeholder is the literal
|
||||
// "<allocID>" sentinel so the rendered config is still valid YAML
|
||||
// (the P08 socket layer substitutes the real alloc-id).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
files, err := TraefikEmitter{}.Render(spec, nil)
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if !strings.Contains(files[0].Content, "alloc-<allocID>") {
|
||||
t.Errorf("nil node should render alloc-<allocID> placeholder\n%s", files[0].Content)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_EmptyHostnameAllocPlaceholder(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
files, err := TraefikEmitter{}.Render(spec, &Node{Hostname: " "})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if !strings.Contains(files[0].Content, "alloc-<allocID>") {
|
||||
t.Errorf("empty hostname should render alloc-<allocID> placeholder\n%s", files[0].Content)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_PathNotOrcaV1Prefixed(t *testing.T) {
|
||||
// REQ-090: the orca-v1- prefix is only for systemd units; Traefik
|
||||
// dynamic-config paths are named orca-<spec.Name>.yaml (single
|
||||
// source of truth — no dual-write window for Traefik configs).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
files, err := TraefikEmitter{}.Render(spec, &Node{Hostname: "n1"})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if strings.Contains(files[0].Path, "orca-v1-") {
|
||||
t.Errorf("Path %q should NOT contain the orca-v1- prefix (systemd-only)", files[0].Path)
|
||||
}
|
||||
if !strings.HasPrefix(files[0].Path, "/etc/traefik/dynamic/orca-") {
|
||||
t.Errorf("Path %q should start with /etc/traefik/dynamic/orca-", files[0].Path)
|
||||
}
|
||||
if !strings.HasSuffix(files[0].Path, ".yaml") {
|
||||
t.Errorf("Path %q should end with .yaml", files[0].Path)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTraefikEmitter_RenderYAMLHasRoutersServicesTLS(t *testing.T) {
|
||||
// Aggregate structural assertion: the rendered YAML has the four
|
||||
// top-level Traefik concepts (routers, services, tls, servers).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
}
|
||||
files, err := TraefikEmitter{}.Render(spec, &Node{Hostname: "n1"})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
c := files[0].Content
|
||||
for _, want := range []string{"routers:", "services:", "tls:", "servers:", "url:"} {
|
||||
if !strings.Contains(c, want) {
|
||||
t.Errorf("rendered YAML missing %q\n%s", want, c)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegisterTraefik_AllServiceRuntimes(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
RegisterTraefik(r)
|
||||
spec := func(runtime string) *jobspec.WorkloadSpec {
|
||||
return &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: runtime},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
}
|
||||
for _, runtime := range []string{"process", "podman", "wasm"} {
|
||||
t.Run(runtime, func(t *testing.T) {
|
||||
files, err := r.Render(spec(runtime), &Node{Hostname: "n1"})
|
||||
if err != nil {
|
||||
t.Fatalf("Render(service:%s): %v", runtime, err)
|
||||
}
|
||||
if len(files) != 1 {
|
||||
t.Fatalf("got %d files, want 1", len(files))
|
||||
}
|
||||
if !strings.Contains(files[0].Path, "/etc/traefik/dynamic/orca-web.yaml") {
|
||||
t.Errorf("Path = %q", files[0].Path)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegisterTraefik_OverwritesExisting(t *testing.T) {
|
||||
// RegisterTraefik should overwrite any prior registration (the
|
||||
// Registry documents last-wins).
|
||||
r := NewRegistry()
|
||||
r.Register("service:process", mockEmitter{files: []File{{Path: "/old"}}})
|
||||
RegisterTraefik(r)
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
files, err := r.Render(spec, &Node{Hostname: "n1"})
|
||||
if err != nil {
|
||||
t.Fatalf("Render: %v", err)
|
||||
}
|
||||
if files[0].Path == "/old" {
|
||||
t.Errorf("RegisterTraefik did not overwrite the prior registration")
|
||||
}
|
||||
}
|
||||
|
||||
// Compile-time assertion that TraefikEmitter implements Emitter.
|
||||
var _ Emitter = TraefikEmitter{}
|
||||
@@ -0,0 +1,237 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// UpdatePlan is the computed update sequence for a Service (P03). It is
|
||||
// a PLAN, not an execution — the transactional execution lands in
|
||||
// v0.10-P10. Each step describes a discrete action the executor takes:
|
||||
// start a set of allocs (Action="start"), wait for them to become
|
||||
// healthy (WaitForHealthy=true), or cutover from old to new
|
||||
// (Action="cutover" for blue-green). The Allocs field carries
|
||||
// placeholder alloc names of the form "<spec.Name>-<index>" where
|
||||
// index is 1-based (the scheduler assigns the real alloc-id at submit
|
||||
// time; P03 uses spec.Name as a placeholder per the socket layer
|
||||
// contract — see SocketEmitter).
|
||||
type UpdatePlan struct {
|
||||
Steps []UpdateStep
|
||||
}
|
||||
|
||||
// UpdateStep is a single step in an UpdatePlan. Action is one of
|
||||
// "start", "wait", "cutover", "promote". Allocs is the list of
|
||||
// placeholder alloc names the step applies to. WaitForHealthy is true
|
||||
// when the executor must wait for the allocs in this step to pass
|
||||
// their health check before proceeding to the next step (driven by
|
||||
// min_healthy_time / healthy_deadline on the spec, which the executor
|
||||
// — not the plan — enforces).
|
||||
type UpdateStep struct {
|
||||
Action string
|
||||
Allocs []string
|
||||
WaitForHealthy bool
|
||||
}
|
||||
|
||||
// maxParallelFor returns the effective max_parallel for the spec,
|
||||
// defaulting to 1 when unset (0) and clamping to count (the validator
|
||||
// already rejects out-of-range values; this is a defensive clamp for
|
||||
// direct callers that bypass the validator).
|
||||
func maxParallelFor(spec *jobspec.WorkloadSpec) int {
|
||||
if spec.Update == nil {
|
||||
return 1
|
||||
}
|
||||
if spec.Update.MaxParallel < 1 {
|
||||
return 1
|
||||
}
|
||||
if spec.Count > 0 && spec.Update.MaxParallel > spec.Count {
|
||||
return spec.Count
|
||||
}
|
||||
return spec.Update.MaxParallel
|
||||
}
|
||||
|
||||
// allocName returns the placeholder alloc name for index i (1-based).
|
||||
// The real alloc-id is assigned by the scheduler at submit time; P03
|
||||
// uses spec.Name as the placeholder per the socket-layer contract.
|
||||
func allocName(spec *jobspec.WorkloadSpec, i int) string {
|
||||
return fmt.Sprintf("%s-%d", spec.Name, i)
|
||||
}
|
||||
|
||||
// allAllocs returns the placeholder alloc names for the full count of
|
||||
// the spec (1..count).
|
||||
func allAllocs(spec *jobspec.WorkloadSpec) []string {
|
||||
out := make([]string, 0, spec.Count)
|
||||
for i := 1; i <= spec.Count; i++ {
|
||||
out = append(out, allocName(spec, i))
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// canaryCount returns the integer canary count for the spec. The
|
||||
// canary field accepts an integer count or a percentage ("<n>%"). For
|
||||
// a percentage, the count is ceil(count * n / 100) with a minimum of 1
|
||||
// when n > 0 (a 10% canary of a 3-replica service is 1 alloc, not 0).
|
||||
// When the canary field is empty, the default is 1 (a single canary
|
||||
// alloc — the smallest meaningful canary).
|
||||
func canaryCount(spec *jobspec.WorkloadSpec) int {
|
||||
if spec.Update == nil {
|
||||
return 1
|
||||
}
|
||||
c := strings.TrimSpace(spec.Update.Canary)
|
||||
if c == "" {
|
||||
return 1
|
||||
}
|
||||
if strings.HasSuffix(c, "%") {
|
||||
n, err := strconv.Atoi(strings.TrimSpace(strings.TrimSuffix(c, "%")))
|
||||
if err != nil || n <= 0 {
|
||||
return 1
|
||||
}
|
||||
allocs := spec.Count * n / 100
|
||||
if allocs < 1 {
|
||||
allocs = 1
|
||||
}
|
||||
return allocs
|
||||
}
|
||||
n, err := strconv.Atoi(c)
|
||||
if err != nil || n < 1 {
|
||||
return 1
|
||||
}
|
||||
if spec.Count > 0 && n > spec.Count {
|
||||
return spec.Count
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
// RenderUpdatePlan computes the rolling/canary/blue-green update
|
||||
// sequence for a Service spec. Returns an *UpdatePlan describing the
|
||||
// steps; the actual transactional execution lands in v0.10-P10.
|
||||
//
|
||||
// The three strategies:
|
||||
//
|
||||
// - rolling: allocs are started in batches of max_parallel. Each
|
||||
// batch waits for healthy before the next batch starts. This is
|
||||
// the simplest strategy and the default for stateless services.
|
||||
//
|
||||
// - canary: a single canary alloc (or N per the canary field) is
|
||||
// started first and waits for healthy. After the canary is
|
||||
// healthy, the plan emits a "promote" step (manual or auto per
|
||||
// auto_promote); the remaining allocs are then started in
|
||||
// max_parallel batches.
|
||||
//
|
||||
// - blue-green: all new allocs are started in parallel (a single
|
||||
// "start" step with the full count). After they are healthy, a
|
||||
// "cutover" step swaps traffic from the old allocs to the new
|
||||
// ones. The old allocs are then stopped (the stop is implicit in
|
||||
// the cutover step for the plan; v0.10-P10 makes it explicit).
|
||||
//
|
||||
// Returns an error if the spec is nil, the update block is nil, or
|
||||
// the strategy is unknown (the validator should have caught these,
|
||||
// but RenderUpdatePlan is defensive — emitters are called from
|
||||
// render paths that may bypass the schema validator).
|
||||
func RenderUpdatePlan(spec *jobspec.WorkloadSpec) (*UpdatePlan, error) {
|
||||
if spec == nil {
|
||||
return nil, fmt.Errorf("emitter/update: spec is nil")
|
||||
}
|
||||
if spec.Update == nil {
|
||||
return nil, fmt.Errorf("emitter/update: update block is nil")
|
||||
}
|
||||
if spec.Count < 1 {
|
||||
return nil, fmt.Errorf("emitter/update: count must be ≥ 1, got %d", spec.Count)
|
||||
}
|
||||
switch spec.Update.Strategy {
|
||||
case "rolling":
|
||||
return renderRollingPlan(spec), nil
|
||||
case "canary":
|
||||
return renderCanaryPlan(spec), nil
|
||||
case "blue-green":
|
||||
return renderBlueGreenPlan(spec), nil
|
||||
default:
|
||||
return nil, fmt.Errorf("emitter/update: unknown strategy %q (want rolling, canary, or blue-green)", spec.Update.Strategy)
|
||||
}
|
||||
}
|
||||
|
||||
// renderRollingPlan emits the rolling-update plan: allocs in batches
|
||||
// of max_parallel, each batch waiting for healthy before the next.
|
||||
func renderRollingPlan(spec *jobspec.WorkloadSpec) *UpdatePlan {
|
||||
plan := &UpdatePlan{}
|
||||
batch := maxParallelFor(spec)
|
||||
allocs := allAllocs(spec)
|
||||
for i := 0; i < len(allocs); i += batch {
|
||||
end := i + batch
|
||||
if end > len(allocs) {
|
||||
end = len(allocs)
|
||||
}
|
||||
plan.Steps = append(plan.Steps, UpdateStep{
|
||||
Action: "start",
|
||||
Allocs: allocs[i:end],
|
||||
WaitForHealthy: true,
|
||||
})
|
||||
}
|
||||
return plan
|
||||
}
|
||||
|
||||
// renderCanaryPlan emits the canary-update plan: a canary batch first
|
||||
// (size per the canary field, default 1), a "promote" step, then the
|
||||
// remaining allocs in max_parallel batches.
|
||||
func renderCanaryPlan(spec *jobspec.WorkloadSpec) *UpdatePlan {
|
||||
plan := &UpdatePlan{}
|
||||
allocs := allAllocs(spec)
|
||||
canary := canaryCount(spec)
|
||||
if canary > len(allocs) {
|
||||
canary = len(allocs)
|
||||
}
|
||||
if canary < 1 {
|
||||
canary = 1
|
||||
}
|
||||
// Step 1: start the canary alloc(s) and wait for healthy.
|
||||
plan.Steps = append(plan.Steps, UpdateStep{
|
||||
Action: "start",
|
||||
Allocs: allocs[:canary],
|
||||
WaitForHealthy: true,
|
||||
})
|
||||
// Step 2: promote (manual or auto per auto_promote).
|
||||
plan.Steps = append(plan.Steps, UpdateStep{
|
||||
Action: "promote",
|
||||
Allocs: allocs[:canary],
|
||||
})
|
||||
// Step 3+: remaining allocs in max_parallel batches.
|
||||
batch := maxParallelFor(spec)
|
||||
remaining := allocs[canary:]
|
||||
for i := 0; i < len(remaining); i += batch {
|
||||
end := i + batch
|
||||
if end > len(remaining) {
|
||||
end = len(remaining)
|
||||
}
|
||||
plan.Steps = append(plan.Steps, UpdateStep{
|
||||
Action: "start",
|
||||
Allocs: remaining[i:end],
|
||||
WaitForHealthy: true,
|
||||
})
|
||||
}
|
||||
return plan
|
||||
}
|
||||
|
||||
// renderBlueGreenPlan emits the blue-green update plan: all new allocs
|
||||
// start in parallel, wait for healthy, then cutover (swap traffic).
|
||||
func renderBlueGreenPlan(spec *jobspec.WorkloadSpec) *UpdatePlan {
|
||||
plan := &UpdatePlan{}
|
||||
allocs := allAllocs(spec)
|
||||
// Step 1: start ALL new allocs in parallel (blue-green does not
|
||||
// batch — the new fleet stands up alongside the old).
|
||||
plan.Steps = append(plan.Steps, UpdateStep{
|
||||
Action: "start",
|
||||
Allocs: allocs,
|
||||
WaitForHealthy: true,
|
||||
})
|
||||
// Step 2: cutover — swap traffic from old to new. The old allocs
|
||||
// are stopped implicitly as part of the cutover (v0.10-P10 makes
|
||||
// the stop explicit in the transactional plane).
|
||||
plan.Steps = append(plan.Steps, UpdateStep{
|
||||
Action: "cutover",
|
||||
Allocs: allocs,
|
||||
WaitForHealthy: false,
|
||||
})
|
||||
return plan
|
||||
}
|
||||
@@ -0,0 +1,396 @@
|
||||
package emitter
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
func TestRenderUpdatePlan_RollingBatches(t *testing.T) {
|
||||
// count=4, max_parallel=2 → 2 batches of 2.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
MaxParallel: 2,
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
if len(plan.Steps) != 2 {
|
||||
t.Fatalf("got %d steps, want 2", len(plan.Steps))
|
||||
}
|
||||
for i, s := range plan.Steps {
|
||||
if s.Action != "start" {
|
||||
t.Errorf("step %d action = %q, want start", i, s.Action)
|
||||
}
|
||||
if !s.WaitForHealthy {
|
||||
t.Errorf("step %d WaitForHealthy = false, want true", i)
|
||||
}
|
||||
if len(s.Allocs) != 2 {
|
||||
t.Errorf("step %d allocs = %d, want 2", i, len(s.Allocs))
|
||||
}
|
||||
}
|
||||
if plan.Steps[0].Allocs[0] != "web-1" || plan.Steps[0].Allocs[1] != "web-2" {
|
||||
t.Errorf("step 0 allocs = %v, want [web-1 web-2]", plan.Steps[0].Allocs)
|
||||
}
|
||||
if plan.Steps[1].Allocs[0] != "web-3" || plan.Steps[1].Allocs[1] != "web-4" {
|
||||
t.Errorf("step 1 allocs = %v, want [web-3 web-4]", plan.Steps[1].Allocs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_RollingUnevenBatches(t *testing.T) {
|
||||
// count=5, max_parallel=2 → 3 batches: 2, 2, 1.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 5,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
MaxParallel: 2,
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
if len(plan.Steps) != 3 {
|
||||
t.Fatalf("got %d steps, want 3", len(plan.Steps))
|
||||
}
|
||||
if len(plan.Steps[2].Allocs) != 1 {
|
||||
t.Errorf("step 2 allocs = %d, want 1 (remainder)", len(plan.Steps[2].Allocs))
|
||||
}
|
||||
if plan.Steps[2].Allocs[0] != "web-5" {
|
||||
t.Errorf("step 2 allocs = %v, want [web-5]", plan.Steps[2].Allocs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_RollingMaxParallelUnset(t *testing.T) {
|
||||
// max_parallel unset (0) → default 1 → 4 batches of 1.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
if len(plan.Steps) != 4 {
|
||||
t.Fatalf("got %d steps, want 4 (one per alloc, batch=1)", len(plan.Steps))
|
||||
}
|
||||
for _, s := range plan.Steps {
|
||||
if len(s.Allocs) != 1 {
|
||||
t.Errorf("allocs = %d, want 1", len(s.Allocs))
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_CanaryDefaultOne(t *testing.T) {
|
||||
// canary unset → default 1 canary alloc, then 3 in batches of 2.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
MaxParallel: 2,
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
// Step 0: canary (1 alloc). Step 1: promote. Steps 2..: remaining.
|
||||
if plan.Steps[0].Action != "start" || len(plan.Steps[0].Allocs) != 1 {
|
||||
t.Errorf("step 0 = %+v, want canary start with 1 alloc", plan.Steps[0])
|
||||
}
|
||||
if plan.Steps[0].Allocs[0] != "web-1" {
|
||||
t.Errorf("canary alloc = %q, want web-1", plan.Steps[0].Allocs[0])
|
||||
}
|
||||
if !plan.Steps[0].WaitForHealthy {
|
||||
t.Error("canary step should wait for healthy")
|
||||
}
|
||||
if plan.Steps[1].Action != "promote" {
|
||||
t.Errorf("step 1 action = %q, want promote", plan.Steps[1].Action)
|
||||
}
|
||||
// Remaining: web-2, web-3, web-4 in batches of 2 → [web-2,web-3], [web-4].
|
||||
if len(plan.Steps) != 4 {
|
||||
t.Fatalf("got %d steps, want 4 (canary + promote + 2 batches)", len(plan.Steps))
|
||||
}
|
||||
if len(plan.Steps[2].Allocs) != 2 || plan.Steps[2].Allocs[0] != "web-2" {
|
||||
t.Errorf("step 2 = %v, want [web-2 web-3]", plan.Steps[2].Allocs)
|
||||
}
|
||||
if len(plan.Steps[3].Allocs) != 1 || plan.Steps[3].Allocs[0] != "web-4" {
|
||||
t.Errorf("step 3 = %v, want [web-4]", plan.Steps[3].Allocs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_CanaryPercent(t *testing.T) {
|
||||
// count=10, canary=20% → 2 canary allocs, then 8 in batches of 3.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 10,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
MaxParallel: 3,
|
||||
Canary: "20%",
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
if len(plan.Steps[0].Allocs) != 2 {
|
||||
t.Errorf("canary step allocs = %d, want 2 (20%% of 10)", len(plan.Steps[0].Allocs))
|
||||
}
|
||||
// Remaining 8 in batches of 3 → ceil(8/3)=3 batches.
|
||||
// Steps: canary, promote, batch(3), batch(3), batch(2) = 5 steps.
|
||||
if len(plan.Steps) != 5 {
|
||||
t.Fatalf("got %d steps, want 5", len(plan.Steps))
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_CanaryIntegerCount(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "2",
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
if len(plan.Steps[0].Allocs) != 2 {
|
||||
t.Errorf("canary step allocs = %d, want 2", len(plan.Steps[0].Allocs))
|
||||
}
|
||||
if plan.Steps[0].Allocs[0] != "web-1" || plan.Steps[0].Allocs[1] != "web-2" {
|
||||
t.Errorf("canary allocs = %v, want [web-1 web-2]", plan.Steps[0].Allocs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_CanaryFullCount(t *testing.T) {
|
||||
// canary == count → no remaining allocs after canary.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 3,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "3",
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
// Steps: canary start (3), promote. No remaining batches.
|
||||
if len(plan.Steps) != 2 {
|
||||
t.Fatalf("got %d steps, want 2 (canary + promote, no remainder)", len(plan.Steps))
|
||||
}
|
||||
if plan.Steps[1].Action != "promote" {
|
||||
t.Errorf("step 1 action = %q, want promote", plan.Steps[1].Action)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_BlueGreen(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "blue-green",
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
// Step 0: start all 4 in parallel. Step 1: cutover.
|
||||
if len(plan.Steps) != 2 {
|
||||
t.Fatalf("got %d steps, want 2", len(plan.Steps))
|
||||
}
|
||||
if plan.Steps[0].Action != "start" {
|
||||
t.Errorf("step 0 action = %q, want start", plan.Steps[0].Action)
|
||||
}
|
||||
if len(plan.Steps[0].Allocs) != 4 {
|
||||
t.Errorf("step 0 allocs = %d, want 4 (all new in parallel)", len(plan.Steps[0].Allocs))
|
||||
}
|
||||
if !plan.Steps[0].WaitForHealthy {
|
||||
t.Error("blue-green start step should wait for healthy")
|
||||
}
|
||||
if plan.Steps[1].Action != "cutover" {
|
||||
t.Errorf("step 1 action = %q, want cutover", plan.Steps[1].Action)
|
||||
}
|
||||
if plan.Steps[1].WaitForHealthy {
|
||||
t.Error("cutover step should NOT wait for healthy (already healthy)")
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_BlueGreenAllocs(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "api",
|
||||
Count: 3,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "blue-green",
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
want := []string{"api-1", "api-2", "api-3"}
|
||||
if len(plan.Steps[0].Allocs) != 3 {
|
||||
t.Errorf("allocs = %v, want %v", plan.Steps[0].Allocs, want)
|
||||
}
|
||||
for i, a := range want {
|
||||
if plan.Steps[0].Allocs[i] != a {
|
||||
t.Errorf("alloc[%d] = %q, want %q", i, plan.Steps[0].Allocs[i], a)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_NilSpec(t *testing.T) {
|
||||
_, err := RenderUpdatePlan(nil)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for nil spec")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "spec is nil") {
|
||||
t.Errorf("error = %q, want 'spec is nil'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_NilUpdate(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Service", Name: "web", Count: 1}
|
||||
_, err := RenderUpdatePlan(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for nil update block")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "update block is nil") {
|
||||
t.Errorf("error = %q, want 'update block is nil'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_CountZero(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 0,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
},
|
||||
}
|
||||
_, err := RenderUpdatePlan(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for count 0")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "count must be") {
|
||||
t.Errorf("error = %q, want 'count must be'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_UnknownStrategy(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "recreate",
|
||||
},
|
||||
}
|
||||
_, err := RenderUpdatePlan(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for unknown strategy")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "unknown strategy") {
|
||||
t.Errorf("error = %q, want 'unknown strategy'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_AllStrategies(t *testing.T) {
|
||||
for _, strat := range []string{"rolling", "canary", "blue-green"} {
|
||||
t.Run(strat, func(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 3,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: strat,
|
||||
MaxParallel: 1,
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan(%s): %v", strat, err)
|
||||
}
|
||||
if len(plan.Steps) == 0 {
|
||||
t.Errorf("strategy %s produced 0 steps", strat)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_PromoteStepCarriesCanaryAllocs(t *testing.T) {
|
||||
// The promote step lists the canary allocs so the executor knows
|
||||
// which allocs are being promoted from canary to stable.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "2",
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
promote := plan.Steps[1]
|
||||
if promote.Action != "promote" {
|
||||
t.Fatalf("step 1 action = %q, want promote", promote.Action)
|
||||
}
|
||||
if len(promote.Allocs) != 2 {
|
||||
t.Errorf("promote allocs = %d, want 2 (the canary allocs)", len(promote.Allocs))
|
||||
}
|
||||
}
|
||||
|
||||
func TestRenderUpdatePlan_MaxParallelClampedToCount(t *testing.T) {
|
||||
// max_parallel > count is clamped to count (defensive; validator
|
||||
// rejects this but the emitter is defensive against direct
|
||||
// callers).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
MaxParallel: 99,
|
||||
},
|
||||
}
|
||||
plan, err := RenderUpdatePlan(spec)
|
||||
if err != nil {
|
||||
t.Fatalf("RenderUpdatePlan: %v", err)
|
||||
}
|
||||
// Clamped to 2 → single batch of 2.
|
||||
if len(plan.Steps) != 1 {
|
||||
t.Errorf("got %d steps, want 1 (clamped)", len(plan.Steps))
|
||||
}
|
||||
if len(plan.Steps[0].Allocs) != 2 {
|
||||
t.Errorf("step 0 allocs = %d, want 2", len(plan.Steps[0].Allocs))
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,147 @@
|
||||
package jobspec
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// ParseFile reads a jobspec file from disk and dispatches on file
|
||||
// extension (R-013, REQ-064):
|
||||
//
|
||||
// - .md → ParseMarkdown (canonical Markdown+frontmatter, R-014/R-015)
|
||||
// - .yaml/.yml → ParseMarkdown with the whole file treated as
|
||||
// frontmatter and Body = "" (pure YAML, no Markdown body)
|
||||
// - .hcl → ParseHCL (legacy adapter; wraps the existing HCL parser
|
||||
// and converts Spec{Job, Tasks} into *WorkloadSpec with Kind="Job",
|
||||
// REQ-090 migration window)
|
||||
//
|
||||
// Unknown extensions return an error. The dispatcher preserves
|
||||
// `orca job run old-spec.hcl` during the v0.9→v0.10 migration window
|
||||
// (REQ-090).
|
||||
func ParseFile(path string) (*WorkloadSpec, error) {
|
||||
data, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read spec file: %w", err)
|
||||
}
|
||||
return Dispatch(data, filepath.Base(path))
|
||||
}
|
||||
|
||||
// ParseHCLFile reads an HCL file and parses it via the legacy HCL parser,
|
||||
// returning the legacy *Spec. It is a convenience wrapper retained for
|
||||
// tests and direct HCL consumers that need the raw Spec{Job, Tasks}
|
||||
// shape during the v0.9→v0.10 migration window (REQ-090).
|
||||
//
|
||||
// Deprecated: use ParseFile (dispatcher) for new code. HCL is legacy per
|
||||
// R-013.
|
||||
func ParseHCLFile(path string) (*Spec, error) {
|
||||
data, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read spec file: %w", err)
|
||||
}
|
||||
return ParseHCLLegacy(data, filepath.Base(path))
|
||||
}
|
||||
|
||||
// Dispatch routes raw jobspec bytes on file extension to the
|
||||
// appropriate parser. filename is used only for HCL (the HCL decoder
|
||||
// needs a filename for error messages and syntax sniffing).
|
||||
func Dispatch(data []byte, filename string) (*WorkloadSpec, error) {
|
||||
ext := strings.ToLower(filepath.Ext(filename))
|
||||
switch ext {
|
||||
case ".md":
|
||||
return ParseMarkdown(data)
|
||||
case ".yaml", ".yml":
|
||||
// Pure YAML file: no Markdown body. Treat the whole file as
|
||||
// the frontmatter block. Body is empty (R-015: no body to
|
||||
// preserve).
|
||||
spec, err := parseYAMLFile(data)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return spec, nil
|
||||
case ".hcl":
|
||||
return ParseHCL(data, filename)
|
||||
default:
|
||||
return nil, fmt.Errorf("parse jobspec: unknown extension %q (want .md, .yaml, .yml, or .hcl)", ext)
|
||||
}
|
||||
}
|
||||
|
||||
// parseYAMLFile treats the whole file as a frontmatter block (no
|
||||
// surrounding `---` delimiters, no Markdown body). This routes .yaml
|
||||
// and .yml files through the same hand-rolled parser as .md.
|
||||
func parseYAMLFile(data []byte) (*WorkloadSpec, error) {
|
||||
block := string(data)
|
||||
if strings.TrimSpace(block) == "" {
|
||||
return nil, fmt.Errorf("parse yaml: empty file")
|
||||
}
|
||||
spec, err := parseFrontmatterBlock(block)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
spec.Body = ""
|
||||
if err := validateWorkload(spec); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return spec, nil
|
||||
}
|
||||
|
||||
// ParseHCL parses a legacy HCL jobspec and adapts it into a *WorkloadSpec
|
||||
// (REQ-064 adapter, REQ-090 migration window). The existing HCL
|
||||
// Spec{Job, Tasks} shape is converted to:
|
||||
//
|
||||
// Kind: "Job"
|
||||
// Name: spec.Job.Name
|
||||
// Runtime: {one_of: "process", command: tasks[0].Command}
|
||||
//
|
||||
// Body is empty (HCL has no Markdown body). The legacy Spec struct and
|
||||
// ParseHCLLegacy are retained for direct HCL consumers that have not yet
|
||||
// migrated.
|
||||
//
|
||||
// Deprecated: use the dispatcher (ParseFile/Dispatch). HCL is legacy
|
||||
// per R-013; the HCL path is retained only for the v0.9→v0.10 migration
|
||||
// window (REQ-090) and will be removed in v1.0.
|
||||
func ParseHCL(data []byte, filename string) (*WorkloadSpec, error) {
|
||||
spec, err := ParseHCLLegacy(data, filename)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
ws := &WorkloadSpec{
|
||||
SpecVersion: "",
|
||||
Kind: "Job",
|
||||
Name: spec.Job.Name,
|
||||
Count: 1,
|
||||
Body: "",
|
||||
}
|
||||
if len(spec.Tasks) > 0 {
|
||||
ws.Runtime = &RuntimeBlock{
|
||||
OneOf: "process",
|
||||
Command: spec.Tasks[0].Command,
|
||||
}
|
||||
}
|
||||
return ws, nil
|
||||
}
|
||||
|
||||
// ParseHCLLegacy is the original HCL-only parser retained for direct
|
||||
// HCL consumers (e.g. the cli/job.go toTaskSpecs path during the
|
||||
// migration window). New code should call ParseHCL (which returns a
|
||||
// *WorkloadSpec) or the dispatcher. Deprecated: HCL is legacy per
|
||||
// R-013; see ParseHCL.
|
||||
func ParseHCLLegacy(data []byte, filename string) (*Spec, error) {
|
||||
var spec Spec
|
||||
if err := hclDecode(filename, data, &spec); err != nil {
|
||||
return nil, fmt.Errorf("decode hcl: %w", err)
|
||||
}
|
||||
if spec.Job.Name == "" {
|
||||
return nil, fmt.Errorf("spec missing job name")
|
||||
}
|
||||
if len(spec.Tasks) == 0 {
|
||||
return nil, fmt.Errorf("spec must have at least one task")
|
||||
}
|
||||
for i, t := range spec.Tasks {
|
||||
if t.Command == "" {
|
||||
return nil, fmt.Errorf("task[%d] (%s) missing command", i, t.Name)
|
||||
}
|
||||
}
|
||||
return &spec, nil
|
||||
}
|
||||
@@ -0,0 +1,222 @@
|
||||
package jobspec
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestDispatch_Markdown(t *testing.T) {
|
||||
input := "---\nkind: Job\nname: md-job\n---\nbody content\n"
|
||||
ws, err := Dispatch([]byte(input), "spec.md")
|
||||
if err != nil {
|
||||
t.Fatalf("Dispatch .md: %v", err)
|
||||
}
|
||||
if ws.Kind != "Job" {
|
||||
t.Errorf("Kind = %q, want Job", ws.Kind)
|
||||
}
|
||||
if ws.Name != "md-job" {
|
||||
t.Errorf("Name = %q, want md-job", ws.Name)
|
||||
}
|
||||
if ws.Body != "body content\n" {
|
||||
t.Errorf("Body = %q, want %q (R-015)", ws.Body, "body content\n")
|
||||
}
|
||||
}
|
||||
|
||||
func TestDispatch_YAML(t *testing.T) {
|
||||
input := "kind: Service\nname: yaml-svc\nports:\n - name: http\n port: 80\n"
|
||||
ws, err := Dispatch([]byte(input), "spec.yaml")
|
||||
if err != nil {
|
||||
t.Fatalf("Dispatch .yaml: %v", err)
|
||||
}
|
||||
if ws.Kind != "Service" {
|
||||
t.Errorf("Kind = %q, want Service", ws.Kind)
|
||||
}
|
||||
if ws.Name != "yaml-svc" {
|
||||
t.Errorf("Name = %q, want yaml-svc", ws.Name)
|
||||
}
|
||||
if ws.Body != "" {
|
||||
t.Errorf("Body = %q, want empty (YAML has no body)", ws.Body)
|
||||
}
|
||||
if len(ws.Ports) != 1 || ws.Ports[0].Name != "http" || ws.Ports[0].Port != 80 {
|
||||
t.Errorf("Ports = %+v, want one http:80", ws.Ports)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDispatch_YML(t *testing.T) {
|
||||
input := "kind: DaemonSet\nname: yml-ds\n"
|
||||
ws, err := Dispatch([]byte(input), "spec.yml")
|
||||
if err != nil {
|
||||
t.Fatalf("Dispatch .yml: %v", err)
|
||||
}
|
||||
if ws.Kind != "DaemonSet" {
|
||||
t.Errorf("Kind = %q, want DaemonSet", ws.Kind)
|
||||
}
|
||||
if ws.Body != "" {
|
||||
t.Errorf("Body = %q, want empty", ws.Body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDispatch_HCLAdapter(t *testing.T) {
|
||||
hcl := `job "demo" {}
|
||||
task "build" {
|
||||
command = "/bin/echo"
|
||||
args = ["hello"]
|
||||
}
|
||||
`
|
||||
ws, err := Dispatch([]byte(hcl), "spec.hcl")
|
||||
if err != nil {
|
||||
t.Fatalf("Dispatch .hcl: %v", err)
|
||||
}
|
||||
if ws.Kind != "Job" {
|
||||
t.Errorf("Kind = %q, want Job (adapter always sets Job)", ws.Kind)
|
||||
}
|
||||
if ws.Name != "demo" {
|
||||
t.Errorf("Name = %q, want demo (from spec.Job.Name)", ws.Name)
|
||||
}
|
||||
if ws.Runtime == nil {
|
||||
t.Fatal("Runtime is nil; adapter should populate from tasks[0]")
|
||||
}
|
||||
if ws.Runtime.OneOf != "process" {
|
||||
t.Errorf("Runtime.OneOf = %q, want process", ws.Runtime.OneOf)
|
||||
}
|
||||
if ws.Runtime.Command != "/bin/echo" {
|
||||
t.Errorf("Runtime.Command = %q, want /bin/echo (from tasks[0].Command)", ws.Runtime.Command)
|
||||
}
|
||||
if ws.Body != "" {
|
||||
t.Errorf("Body = %q, want empty (HCL has no body)", ws.Body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDispatch_HCLAdapterNoTasks(t *testing.T) {
|
||||
hcl := `job "x" {}`
|
||||
_, err := Dispatch([]byte(hcl), "spec.hcl")
|
||||
if err == nil {
|
||||
t.Fatal("expected error for HCL with no tasks")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "at least one task") {
|
||||
t.Errorf("error = %q, want it to contain 'at least one task'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestDispatch_UnknownExtension(t *testing.T) {
|
||||
_, err := Dispatch([]byte("kind: Job\nname: x\n"), "spec.json")
|
||||
if err == nil {
|
||||
t.Fatal("expected error for unknown extension, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "unknown extension") {
|
||||
t.Errorf("error = %q, want it to contain 'unknown extension'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestDispatch_NoExtension(t *testing.T) {
|
||||
_, err := Dispatch([]byte("kind: Job\nname: x\n"), "spec")
|
||||
if err == nil {
|
||||
t.Fatal("expected error for no extension, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestDispatch_EmptyYAML(t *testing.T) {
|
||||
_, err := Dispatch([]byte(""), "spec.yaml")
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty YAML, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseFile_Markdown(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, "spec.md")
|
||||
content := "---\nkind: Job\nname: file-md\n---\nbody\n"
|
||||
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
|
||||
t.Fatalf("write: %v", err)
|
||||
}
|
||||
ws, err := ParseFile(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseFile .md: %v", err)
|
||||
}
|
||||
if ws.Kind != "Job" || ws.Name != "file-md" {
|
||||
t.Errorf("got Kind=%q Name=%q", ws.Kind, ws.Name)
|
||||
}
|
||||
if ws.Body != "body\n" {
|
||||
t.Errorf("Body = %q, want %q", ws.Body, "body\n")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseFile_YAML(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, "spec.yaml")
|
||||
content := "kind: Service\nname: file-yaml\n"
|
||||
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
|
||||
t.Fatalf("write: %v", err)
|
||||
}
|
||||
ws, err := ParseFile(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseFile .yaml: %v", err)
|
||||
}
|
||||
if ws.Kind != "Service" || ws.Name != "file-yaml" {
|
||||
t.Errorf("got Kind=%q Name=%q", ws.Kind, ws.Name)
|
||||
}
|
||||
if ws.Body != "" {
|
||||
t.Errorf("Body = %q, want empty", ws.Body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseFile_HCL(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, "spec.hcl")
|
||||
content := `job "file-hcl" {}
|
||||
task "t" { command = "/bin/true" }
|
||||
`
|
||||
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
|
||||
t.Fatalf("write: %v", err)
|
||||
}
|
||||
ws, err := ParseFile(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseFile .hcl: %v", err)
|
||||
}
|
||||
if ws.Kind != "Job" || ws.Name != "file-hcl" {
|
||||
t.Errorf("got Kind=%q Name=%q", ws.Kind, ws.Name)
|
||||
}
|
||||
if ws.Runtime == nil || ws.Runtime.Command != "/bin/true" {
|
||||
t.Errorf("Runtime.Command = %v, want /bin/true", ws.Runtime)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseFile_MissingFile(t *testing.T) {
|
||||
_, err := ParseFile(filepath.Join(t.TempDir(), "nope.md"))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing file, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "read spec file") {
|
||||
t.Errorf("error = %q, want it to contain 'read spec file'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseFile_UnknownExtension(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, "spec.txt")
|
||||
if err := os.WriteFile(path, []byte("kind: Job\nname: x\n"), 0o644); err != nil {
|
||||
t.Fatalf("write: %v", err)
|
||||
}
|
||||
_, err := ParseFile(path)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for unknown extension, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "unknown extension") {
|
||||
t.Errorf("error = %q, want 'unknown extension'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseHCL_LegacySpec(t *testing.T) {
|
||||
hcl := `job "legacy" {}
|
||||
task "t" { command = "/bin/echo" }
|
||||
`
|
||||
ws, err := ParseHCL([]byte(hcl), "spec.hcl")
|
||||
if err != nil {
|
||||
t.Fatalf("ParseHCL: %v", err)
|
||||
}
|
||||
if ws.Kind != "Job" || ws.Name != "legacy" {
|
||||
t.Errorf("adapter got Kind=%q Name=%q", ws.Kind, ws.Name)
|
||||
}
|
||||
}
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,109 @@
|
||||
package jobspec
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// FuzzParseMarkdownRoundTrip is the REQ-067 fuzz harness for R-015
|
||||
// byte-exact body preservation. It generates random frontmatter + body
|
||||
// combinations, runs ParseMarkdown, and asserts that the parsed Body
|
||||
// equals the original body byte-for-byte whenever parsing succeeds.
|
||||
// When parsing fails (bad frontmatter), the iteration passes — the
|
||||
// parser is allowed to reject malformed input.
|
||||
//
|
||||
// The seed corpus (added via f.Add) covers adversarial fixtures: CRLF
|
||||
// body, BOM prefix, no frontmatter, only-closing-separator, body with
|
||||
// `---` inside a code fence, trailing whitespace, empty body. The seed
|
||||
// corpus runs as regular tests under `go test` (CI); random input runs
|
||||
// only under `go test -fuzz=FuzzParseMarkdownRoundTrip` in a dedicated
|
||||
// process.
|
||||
func FuzzParseMarkdownRoundTrip(f *testing.F) {
|
||||
// Seed 1: valid frontmatter + simple body.
|
||||
f.Add([]byte("---\nkind: Job\nname: seed1\n---\n# body\n"))
|
||||
|
||||
// Seed 2: CRLF body.
|
||||
f.Add([]byte("---\r\nkind: Job\r\nname: seed2\r\n---\r\n# body\r\nCRLF\r\n"))
|
||||
|
||||
// Seed 3: BOM prefix.
|
||||
f.Add([]byte("\uFEFF---\nkind: Job\nname: seed3\n---\nbody\n"))
|
||||
|
||||
// Seed 4: no frontmatter (just body) — should fail to parse.
|
||||
f.Add([]byte("# just a body\nno frontmatter\n"))
|
||||
|
||||
// Seed 5: frontmatter with only the closing `---` (no opening).
|
||||
f.Add([]byte("body\n---\nmore body\n"))
|
||||
|
||||
// Seed 6: body containing `---` in a code fence.
|
||||
f.Add([]byte("---\nkind: Job\nname: seed6\n---\n```bash\necho '---'\n```\n"))
|
||||
|
||||
// Seed 7: body with trailing whitespace.
|
||||
f.Add([]byte("---\nkind: Job\nname: seed7\n---\nbody with trailing spaces \n"))
|
||||
|
||||
// Seed 8: empty body.
|
||||
f.Add([]byte("---\nkind: Job\nname: seed8\n---\n"))
|
||||
|
||||
// Seed 9: empty frontmatter (should fail).
|
||||
f.Add([]byte("---\n---\nbody\n"))
|
||||
|
||||
// Seed 10: body with no trailing newline.
|
||||
f.Add([]byte("---\nkind: Job\nname: seed10\n---\nno trailing newline"))
|
||||
|
||||
f.Fuzz(func(t *testing.T, data []byte) {
|
||||
// Reconstruct the body from the input so we can assert
|
||||
// byte-exact round-trip. We do this by re-splitting the
|
||||
// frontmatter using the same logic the parser uses, but only
|
||||
// to extract the expected body. If the input has no valid
|
||||
// frontmatter delimiter pair, ParseMarkdown will return an
|
||||
// error and we pass the iteration.
|
||||
expectedBody := extractExpectedBody(string(data))
|
||||
|
||||
spec, err := ParseMarkdown(data)
|
||||
if err != nil {
|
||||
// Parser rejected the input — acceptable for a fuzz
|
||||
// iteration (the input may be malformed). Pass.
|
||||
return
|
||||
}
|
||||
// R-015: body must be byte-exact.
|
||||
if spec.Body != expectedBody {
|
||||
t.Errorf("R-015 body round-trip mismatch:\n got = %q\nwant = %q", spec.Body, expectedBody)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// extractExpectedBody returns the body portion of a Markdown jobspec
|
||||
// input using the same delimiter-splitting logic as splitFrontmatter,
|
||||
// so the fuzz harness can assert byte-exact preservation independently
|
||||
// of the parser's internal extraction. If the input has no valid
|
||||
// frontmatter, the result is "" (and ParseMarkdown will error).
|
||||
func extractExpectedBody(content string) string {
|
||||
stripped := content
|
||||
if strings.HasPrefix(stripped, "\uFEFF") {
|
||||
stripped = stripped[len("\uFEFF"):]
|
||||
}
|
||||
trimmed := strings.TrimLeft(stripped, "\r\n\t ")
|
||||
if !strings.HasPrefix(trimmed, "---") {
|
||||
return ""
|
||||
}
|
||||
rest := trimmed[3:]
|
||||
if len(rest) > 0 && rest[0] != '\n' && rest[0] != '\r' {
|
||||
return ""
|
||||
}
|
||||
rest = strings.TrimLeft(rest, "\r\n")
|
||||
idx := findClosingDelimiter(rest)
|
||||
if idx < 0 {
|
||||
return ""
|
||||
}
|
||||
afterClose := rest[idx:]
|
||||
newlineIdx := strings.IndexAny(afterClose, "\r\n")
|
||||
if newlineIdx < 0 {
|
||||
return ""
|
||||
}
|
||||
bodyStart := newlineIdx
|
||||
if strings.HasPrefix(afterClose[bodyStart:], "\r\n") {
|
||||
bodyStart += 2
|
||||
} else {
|
||||
bodyStart += 1
|
||||
}
|
||||
return afterClose[bodyStart:]
|
||||
}
|
||||
@@ -0,0 +1,852 @@
|
||||
package jobspec
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestParseMarkdown_FullFrontmatter(t *testing.T) {
|
||||
body := "# Hello\n\nThis is the body.\n\nTrailing newline preserved.\n"
|
||||
input := "---\n" +
|
||||
"orca-spec-version: \"1\"\n" +
|
||||
"kind: Job\n" +
|
||||
"name: my-job\n" +
|
||||
"count: 3\n" +
|
||||
"---\n" +
|
||||
body
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.SpecVersion != "1" {
|
||||
t.Errorf("SpecVersion = %q, want %q", spec.SpecVersion, "1")
|
||||
}
|
||||
if spec.Kind != "Job" {
|
||||
t.Errorf("Kind = %q, want %q", spec.Kind, "Job")
|
||||
}
|
||||
if spec.Name != "my-job" {
|
||||
t.Errorf("Name = %q, want %q", spec.Name, "my-job")
|
||||
}
|
||||
if spec.Count != 3 {
|
||||
t.Errorf("Count = %d, want 3", spec.Count)
|
||||
}
|
||||
if spec.Body != body {
|
||||
t.Errorf("Body = %q, want %q (byte-exact, R-015)", spec.Body, body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_BodyByteExactTrailingNewline(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
body string
|
||||
}{
|
||||
{"with_trailing_newline", "# Title\n\nbody\n"},
|
||||
{"with_double_trailing_newline", "# Title\n\nbody\n\n"},
|
||||
{"no_trailing_newline", "# Title\n\nbody"},
|
||||
{"empty_body_with_newline", "\n"},
|
||||
{"only_newlines", "\n\n\n"},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
input := "---\nkind: Job\nname: x\n---\n" + tc.body
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Body != tc.body {
|
||||
t.Errorf("Body byte-exact mismatch (R-015):\n got = %q\nwant = %q", spec.Body, tc.body)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_NoFrontmatter(t *testing.T) {
|
||||
input := "# Just a body\n\nNo frontmatter here."
|
||||
_, err := ParseMarkdown([]byte(input))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing frontmatter, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "frontmatter") {
|
||||
t.Errorf("error = %q, want it to contain 'frontmatter'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_EmptyFrontmatter(t *testing.T) {
|
||||
input := "---\n---\n\nbody"
|
||||
_, err := ParseMarkdown([]byte(input))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty frontmatter, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "empty frontmatter") {
|
||||
t.Errorf("error = %q, want it to contain 'empty frontmatter'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_UnknownKind(t *testing.T) {
|
||||
input := "---\nkind: CronJob\nname: x\n---\nbody\n"
|
||||
_, err := ParseMarkdown([]byte(input))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for unknown kind, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "not one of") {
|
||||
t.Errorf("error = %q, want it to contain 'not one of'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_MissingName(t *testing.T) {
|
||||
input := "---\nkind: Job\n---\nbody\n"
|
||||
_, err := ParseMarkdown([]byte(input))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing name, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "missing name") {
|
||||
t.Errorf("error = %q, want it to contain 'missing name'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_MissingKind(t *testing.T) {
|
||||
input := "---\nname: x\n---\nbody\n"
|
||||
_, err := ParseMarkdown([]byte(input))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing kind, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "missing kind") {
|
||||
t.Errorf("error = %q, want it to contain 'missing kind'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_EachValidKind(t *testing.T) {
|
||||
cases := []string{"Job", "Service", "DaemonSet"}
|
||||
for _, kind := range cases {
|
||||
t.Run(kind, func(t *testing.T) {
|
||||
input := "---\nkind: " + kind + "\nname: x\n---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Kind != kind {
|
||||
t.Errorf("Kind = %q, want %q", spec.Kind, kind)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_EnvScalarAndObject(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Job\n" +
|
||||
"name: x\n" +
|
||||
"env:\n" +
|
||||
" FOO: bar\n" +
|
||||
" BAZ: \"qux\"\n" +
|
||||
" SECRET_REF:\n" +
|
||||
" from: \"secret:db-password\"\n" +
|
||||
" INLINE: {from: \"secret:token\"}\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if got := spec.Env["FOO"]; got != "bar" {
|
||||
t.Errorf("env[FOO] = %q, want %q", got, "bar")
|
||||
}
|
||||
if got := spec.Env["BAZ"]; got != "qux" {
|
||||
t.Errorf("env[BAZ] = %q, want %q", got, "qux")
|
||||
}
|
||||
if got := spec.Env["INLINE"]; got != `{from: "secret:token"}` {
|
||||
t.Errorf("env[INLINE] = %q, want the raw object string", got)
|
||||
}
|
||||
if _, ok := spec.Env["SECRET_REF"]; !ok {
|
||||
t.Errorf("env[SECRET_REF] missing; nested from: stored as empty string")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_PortsArray(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"ports:\n" +
|
||||
" - name: http\n" +
|
||||
" port: 8080\n" +
|
||||
" host_port: 80\n" +
|
||||
" protocol: tcp\n" +
|
||||
" - name: https\n" +
|
||||
" port: 8443\n" +
|
||||
" host_port: 443\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Ports) != 2 {
|
||||
t.Fatalf("Ports = %d, want 2", len(spec.Ports))
|
||||
}
|
||||
if spec.Ports[0].Name != "http" || spec.Ports[0].Port != 8080 || spec.Ports[0].HostPort != 80 || spec.Ports[0].Protocol != "tcp" {
|
||||
t.Errorf("Ports[0] = %+v", spec.Ports[0])
|
||||
}
|
||||
if spec.Ports[1].Name != "https" || spec.Ports[1].Port != 8443 || spec.Ports[1].HostPort != 443 {
|
||||
t.Errorf("Ports[1] = %+v", spec.Ports[1])
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_VolumesArray(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Job\n" +
|
||||
"name: x\n" +
|
||||
"volumes:\n" +
|
||||
" - name: data\n" +
|
||||
" type: host\n" +
|
||||
" source: /data\n" +
|
||||
" target: /data\n" +
|
||||
" read_only: true\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Volumes) != 1 {
|
||||
t.Fatalf("Volumes = %d, want 1", len(spec.Volumes))
|
||||
}
|
||||
v := spec.Volumes[0]
|
||||
if v.Name != "data" || v.Type != "host" || v.Source != "/data" || v.Target != "/data" || !v.ReadOnly {
|
||||
t.Errorf("Volumes[0] = %+v", v)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_RuntimeBlock(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Job\n" +
|
||||
"name: x\n" +
|
||||
"runtime:\n" +
|
||||
" one_of: process\n" +
|
||||
" image: docker.io/nginx:latest\n" +
|
||||
" command: /bin/sh -c 'echo hi'\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Runtime == nil {
|
||||
t.Fatal("Runtime is nil")
|
||||
}
|
||||
if spec.Runtime.OneOf != "process" {
|
||||
t.Errorf("Runtime.OneOf = %q, want %q", spec.Runtime.OneOf, "process")
|
||||
}
|
||||
if spec.Runtime.Image != "docker.io/nginx:latest" {
|
||||
t.Errorf("Runtime.Image = %q, want %q", spec.Runtime.Image, "docker.io/nginx:latest")
|
||||
}
|
||||
if spec.Runtime.Command != "/bin/sh -c 'echo hi'" {
|
||||
t.Errorf("Runtime.Command = %q, want %q", spec.Runtime.Command, "/bin/sh -c 'echo hi'")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_SecretsInlineArray(t *testing.T) {
|
||||
input := "---\nkind: Job\nname: x\nsecrets: [\"db-password\", \"api-token\"]\n---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Secrets) != 2 {
|
||||
t.Fatalf("Secrets = %d, want 2", len(spec.Secrets))
|
||||
}
|
||||
if spec.Secrets[0] != "db-password" || spec.Secrets[1] != "api-token" {
|
||||
t.Errorf("Secrets = %v, want [db-password api-token]", spec.Secrets)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_SecretsBlockArray(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Job\n" +
|
||||
"name: x\n" +
|
||||
"secrets:\n" +
|
||||
" - db-password\n" +
|
||||
" - api-token\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Secrets) != 2 {
|
||||
t.Fatalf("Secrets = %d, want 2", len(spec.Secrets))
|
||||
}
|
||||
if spec.Secrets[0] != "db-password" || spec.Secrets[1] != "api-token" {
|
||||
t.Errorf("Secrets = %v, want [db-password api-token]", spec.Secrets)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_CRLFBodyPreserved(t *testing.T) {
|
||||
body := "# Title\r\n\r\nCRLF body.\r\n"
|
||||
input := "---\r\nkind: Job\r\nname: x\r\n---\r\n" + body
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Body != body {
|
||||
t.Errorf("CRLF body not preserved (R-015):\n got = %q\nwant = %q", spec.Body, body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_BOMStrippedFromFrontmatter(t *testing.T) {
|
||||
body := "# body\n"
|
||||
input := "\uFEFF" + "---\nkind: Job\nname: x\n---\n" + body
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Kind != "Job" {
|
||||
t.Errorf("Kind = %q, want Job (BOM should be stripped from frontmatter scan)", spec.Kind)
|
||||
}
|
||||
if spec.Body != body {
|
||||
t.Errorf("Body = %q, want %q", spec.Body, body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_BodyWithCodeFenceContainingDashes(t *testing.T) {
|
||||
body := "```bash\n" +
|
||||
"echo '---'\n" +
|
||||
"echo '--- end ---'\n" +
|
||||
"```\n"
|
||||
input := "---\nkind: Job\nname: x\n---\n" + body
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Body != body {
|
||||
t.Errorf("Body with code-fence --- not preserved (R-015):\n got = %q\nwant = %q", spec.Body, body)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_OnlyClosingSeparator(t *testing.T) {
|
||||
input := "no opening\n---\nbody\n"
|
||||
_, err := ParseMarkdown([]byte(input))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for input with only closing separator, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_QuotedValues(t *testing.T) {
|
||||
input := "---\nkind: \"Job\"\nname: 'my-job'\n---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Kind != "Job" {
|
||||
t.Errorf("Kind = %q, want Job (double-quoted)", spec.Kind)
|
||||
}
|
||||
if spec.Name != "my-job" {
|
||||
t.Errorf("Name = %q, want my-job (single-quoted)", spec.Name)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_CountDefault(t *testing.T) {
|
||||
input := "---\nkind: Job\nname: x\n---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Count != 1 {
|
||||
t.Errorf("Count default = %d, want 1", spec.Count)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_UnknownKeyIgnored(t *testing.T) {
|
||||
input := "---\nkind: Job\nname: x\nfuture_field: value\n---\nbody\n"
|
||||
_, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown should ignore unknown keys: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_RestartBlock(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"restart:\n" +
|
||||
" mode: service\n" +
|
||||
" attempts: 5\n" +
|
||||
" delay: 3s\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Restart == nil {
|
||||
t.Fatal("Restart is nil")
|
||||
}
|
||||
if spec.Restart.Mode != "service" {
|
||||
t.Errorf("Restart.Mode = %q, want service", spec.Restart.Mode)
|
||||
}
|
||||
if spec.Restart.MaxRetries != 5 {
|
||||
t.Errorf("Restart.MaxRetries = %d, want 5", spec.Restart.MaxRetries)
|
||||
}
|
||||
if spec.Restart.Delay != "3s" {
|
||||
t.Errorf("Restart.Delay = %q, want 3s", spec.Restart.Delay)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_RestartBlockMaxRetriesAlias(t *testing.T) {
|
||||
// max_retries is the canonical key; attempts is an accepted alias.
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"restart:\n" +
|
||||
" mode: on-failure\n" +
|
||||
" max_retries: 3\n" +
|
||||
" delay: 1s\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Restart == nil || spec.Restart.MaxRetries != 3 {
|
||||
t.Fatalf("Restart.MaxRetries = %d, want 3 (max_retries alias)", spec.Restart.MaxRetries)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_UpdateBlock(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"update:\n" +
|
||||
" strategy: canary\n" +
|
||||
" max_parallel: 2\n" +
|
||||
" min_healthy_time: 30s\n" +
|
||||
" healthy_deadline: 5m\n" +
|
||||
" canary: 10%\n" +
|
||||
" auto_promote: true\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Update == nil {
|
||||
t.Fatal("Update is nil")
|
||||
}
|
||||
if spec.Update.Strategy != "canary" {
|
||||
t.Errorf("Update.Strategy = %q, want canary", spec.Update.Strategy)
|
||||
}
|
||||
if spec.Update.MaxParallel != 2 {
|
||||
t.Errorf("Update.MaxParallel = %d, want 2", spec.Update.MaxParallel)
|
||||
}
|
||||
if spec.Update.MinHealthyTime != "30s" {
|
||||
t.Errorf("Update.MinHealthyTime = %q, want 30s", spec.Update.MinHealthyTime)
|
||||
}
|
||||
if spec.Update.HealthyDeadline != "5m" {
|
||||
t.Errorf("Update.HealthyDeadline = %q, want 5m", spec.Update.HealthyDeadline)
|
||||
}
|
||||
if spec.Update.Canary != "10%" {
|
||||
t.Errorf("Update.Canary = %q, want 10%%", spec.Update.Canary)
|
||||
}
|
||||
if !spec.Update.AutoPromote {
|
||||
t.Errorf("Update.AutoPromote = false, want true")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_ServiceBlock(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"service:\n" +
|
||||
" name: web\n" +
|
||||
" port: 8080\n" +
|
||||
" bind: 127.0.0.1\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Service == nil {
|
||||
t.Fatal("Service is nil")
|
||||
}
|
||||
if spec.Service.Name != "web" {
|
||||
t.Errorf("Service.Name = %q, want web", spec.Service.Name)
|
||||
}
|
||||
if spec.Service.Port != 8080 {
|
||||
t.Errorf("Service.Port = %d, want 8080", spec.Service.Port)
|
||||
}
|
||||
if spec.Service.Bind != "127.0.0.1" {
|
||||
t.Errorf("Service.Bind = %q, want 127.0.0.1", spec.Service.Bind)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_HealthBlock(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"health:\n" +
|
||||
" check_type: http\n" +
|
||||
" interval: 10s\n" +
|
||||
" timeout: 2s\n" +
|
||||
" unhealthy_threshold: 3\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Health == nil {
|
||||
t.Fatal("Health is nil")
|
||||
}
|
||||
if spec.Health.CheckType != "http" {
|
||||
t.Errorf("Health.CheckType = %q, want http", spec.Health.CheckType)
|
||||
}
|
||||
if spec.Health.Interval != "10s" {
|
||||
t.Errorf("Health.Interval = %q, want 10s", spec.Health.Interval)
|
||||
}
|
||||
if spec.Health.Timeout != "2s" {
|
||||
t.Errorf("Health.Timeout = %q, want 2s", spec.Health.Timeout)
|
||||
}
|
||||
if spec.Health.UnhealthyThreshold != 3 {
|
||||
t.Errorf("Health.UnhealthyThreshold = %d, want 3", spec.Health.UnhealthyThreshold)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_ConstraintsInlineArray(t *testing.T) {
|
||||
// Inline flow-array form: the parser does NOT unescape YAML
|
||||
// escapes (consistent with the secrets inline parser). Use
|
||||
// single-quoted scalars inside the flow array so the CEL strings
|
||||
// are preserved verbatim.
|
||||
input := "---\nkind: Service\nname: web\nconstraints: ['node.role == \"web\"', 'region == \"us\"']\n---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Constraints) != 2 {
|
||||
t.Fatalf("Constraints = %d, want 2", len(spec.Constraints))
|
||||
}
|
||||
if spec.Constraints[0] != `node.role == "web"` {
|
||||
t.Errorf("Constraints[0] = %q", spec.Constraints[0])
|
||||
}
|
||||
if spec.Constraints[1] != `region == "us"` {
|
||||
t.Errorf("Constraints[1] = %q", spec.Constraints[1])
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_ConstraintsBlockArray(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"constraints:\n" +
|
||||
" - node.role == \"web\"\n" +
|
||||
" - region == \"us\"\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Constraints) != 2 {
|
||||
t.Fatalf("Constraints = %d, want 2", len(spec.Constraints))
|
||||
}
|
||||
if spec.Constraints[0] != `node.role == "web"` {
|
||||
t.Errorf("Constraints[0] = %q", spec.Constraints[0])
|
||||
}
|
||||
if spec.Constraints[1] != `region == "us"` {
|
||||
t.Errorf("Constraints[1] = %q", spec.Constraints[1])
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_AffinityBlock(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"affinity:\n" +
|
||||
" - target: node.role == \"web\"\n" +
|
||||
" weight: 100\n" +
|
||||
" - target: region == \"us\"\n" +
|
||||
" weight: 50\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Affinity) != 2 {
|
||||
t.Fatalf("Affinity = %d, want 2", len(spec.Affinity))
|
||||
}
|
||||
if spec.Affinity[0].Target != `node.role == "web"` {
|
||||
t.Errorf("Affinity[0].Target = %q", spec.Affinity[0].Target)
|
||||
}
|
||||
if spec.Affinity[0].Weight != 100 {
|
||||
t.Errorf("Affinity[0].Weight = %d, want 100", spec.Affinity[0].Weight)
|
||||
}
|
||||
if spec.Affinity[1].Target != `region == "us"` {
|
||||
t.Errorf("Affinity[1].Target = %q", spec.Affinity[1].Target)
|
||||
}
|
||||
if spec.Affinity[1].Weight != 50 {
|
||||
t.Errorf("Affinity[1].Weight = %d, want 50", spec.Affinity[1].Weight)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_LifecycleBlock(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"lifecycle:\n" +
|
||||
" pre_stop:\n" +
|
||||
" - /bin/sh -c 'sleep 5'\n" +
|
||||
" - /usr/local/bin/drain.sh\n" +
|
||||
" post_start:\n" +
|
||||
" - /usr/local/bin/warm-cache.sh\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Lifecycle == nil {
|
||||
t.Fatal("Lifecycle is nil")
|
||||
}
|
||||
if len(spec.Lifecycle.PreStop) != 2 {
|
||||
t.Fatalf("PreStop = %d, want 2", len(spec.Lifecycle.PreStop))
|
||||
}
|
||||
if spec.Lifecycle.PreStop[0] != "/bin/sh -c 'sleep 5'" {
|
||||
t.Errorf("PreStop[0] = %q", spec.Lifecycle.PreStop[0])
|
||||
}
|
||||
if spec.Lifecycle.PreStop[1] != "/usr/local/bin/drain.sh" {
|
||||
t.Errorf("PreStop[1] = %q", spec.Lifecycle.PreStop[1])
|
||||
}
|
||||
if len(spec.Lifecycle.PostStart) != 1 {
|
||||
t.Fatalf("PostStart = %d, want 1", len(spec.Lifecycle.PostStart))
|
||||
}
|
||||
if spec.Lifecycle.PostStart[0] != "/usr/local/bin/warm-cache.sh" {
|
||||
t.Errorf("PostStart[0] = %q", spec.Lifecycle.PostStart[0])
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_LifecycleInlineArray(t *testing.T) {
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"lifecycle:\n" +
|
||||
" pre_stop: [\"/bin/true\"]\n" +
|
||||
" post_start: [\"/bin/warmup\", \"/bin/check\"]\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Lifecycle == nil {
|
||||
t.Fatal("Lifecycle is nil")
|
||||
}
|
||||
if len(spec.Lifecycle.PreStop) != 1 || spec.Lifecycle.PreStop[0] != "/bin/true" {
|
||||
t.Errorf("PreStop = %v, want [/bin/true]", spec.Lifecycle.PreStop)
|
||||
}
|
||||
if len(spec.Lifecycle.PostStart) != 2 {
|
||||
t.Fatalf("PostStart = %v, want 2 entries", spec.Lifecycle.PostStart)
|
||||
}
|
||||
if spec.Lifecycle.PostStart[0] != "/bin/warmup" || spec.Lifecycle.PostStart[1] != "/bin/check" {
|
||||
t.Errorf("PostStart = %v, want [/bin/warmup /bin/check]", spec.Lifecycle.PostStart)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_FullServiceSpec(t *testing.T) {
|
||||
// A complete Service spec exercising every P02-parsed block together.
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"count: 3\n" +
|
||||
"runtime:\n" +
|
||||
" one_of: process\n" +
|
||||
" command: /usr/bin/httpd\n" +
|
||||
"ports:\n" +
|
||||
" - name: http\n" +
|
||||
" port: 8080\n" +
|
||||
"restart:\n" +
|
||||
" mode: service\n" +
|
||||
" attempts: 5\n" +
|
||||
" delay: 2s\n" +
|
||||
"update:\n" +
|
||||
" strategy: rolling\n" +
|
||||
" max_parallel: 1\n" +
|
||||
" auto_promote: false\n" +
|
||||
"service:\n" +
|
||||
" name: web\n" +
|
||||
" port: 8080\n" +
|
||||
"health:\n" +
|
||||
" check_type: http\n" +
|
||||
" interval: 5s\n" +
|
||||
" timeout: 1s\n" +
|
||||
" unhealthy_threshold: 2\n" +
|
||||
"constraints:\n" +
|
||||
" - node.role == \"web\"\n" +
|
||||
"affinity:\n" +
|
||||
" - target: zone == \"a\"\n" +
|
||||
" weight: 80\n" +
|
||||
"lifecycle:\n" +
|
||||
" post_start:\n" +
|
||||
" - /bin/ready.sh\n" +
|
||||
"---\n# body\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Restart == nil || spec.Restart.Mode != "service" {
|
||||
t.Errorf("Restart not parsed: %+v", spec.Restart)
|
||||
}
|
||||
if spec.Update == nil || spec.Update.Strategy != "rolling" {
|
||||
t.Errorf("Update not parsed: %+v", spec.Update)
|
||||
}
|
||||
if spec.Service == nil || spec.Service.Port != 8080 {
|
||||
t.Errorf("Service not parsed: %+v", spec.Service)
|
||||
}
|
||||
if spec.Health == nil || spec.Health.CheckType != "http" {
|
||||
t.Errorf("Health not parsed: %+v", spec.Health)
|
||||
}
|
||||
if len(spec.Constraints) != 1 {
|
||||
t.Errorf("Constraints = %v", spec.Constraints)
|
||||
}
|
||||
if len(spec.Affinity) != 1 || spec.Affinity[0].Weight != 80 {
|
||||
t.Errorf("Affinity = %v", spec.Affinity)
|
||||
}
|
||||
if spec.Lifecycle == nil || len(spec.Lifecycle.PostStart) != 1 {
|
||||
t.Errorf("Lifecycle not parsed: %+v", spec.Lifecycle)
|
||||
}
|
||||
if spec.Body != "# body\n" {
|
||||
t.Errorf("Body = %q, want %q (R-015)", spec.Body, "# body\n")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_TasksBlock(t *testing.T) {
|
||||
// P06: a task group with two tasks, each carrying its own runtime
|
||||
// and command. The parser must populate spec.Tasks with two
|
||||
// entries preserving name, runtime (one_of/image/command), and
|
||||
// the task-level command.
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"tasks:\n" +
|
||||
" - name: app\n" +
|
||||
" runtime:\n" +
|
||||
" one_of: process\n" +
|
||||
" image: docker.io/nginx:latest\n" +
|
||||
" command: /usr/bin/httpd -f\n" +
|
||||
" command: /usr/bin/httpd -f\n" +
|
||||
" - name: sidecar\n" +
|
||||
" runtime:\n" +
|
||||
" one_of: wasm\n" +
|
||||
" command: /bin/wasm-runner sidecar.wasm\n" +
|
||||
" command: /bin/wasm-runner sidecar.wasm\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Tasks) != 2 {
|
||||
t.Fatalf("Tasks = %d, want 2", len(spec.Tasks))
|
||||
}
|
||||
app := spec.Tasks[0]
|
||||
if app.Name != "app" {
|
||||
t.Errorf("Tasks[0].Name = %q, want app", app.Name)
|
||||
}
|
||||
if app.Runtime == nil {
|
||||
t.Fatal("Tasks[0].Runtime is nil")
|
||||
}
|
||||
if app.Runtime.OneOf != "process" {
|
||||
t.Errorf("Tasks[0].Runtime.OneOf = %q, want process", app.Runtime.OneOf)
|
||||
}
|
||||
if app.Runtime.Image != "docker.io/nginx:latest" {
|
||||
t.Errorf("Tasks[0].Runtime.Image = %q", app.Runtime.Image)
|
||||
}
|
||||
if app.Runtime.Command != "/usr/bin/httpd -f" {
|
||||
t.Errorf("Tasks[0].Runtime.Command = %q", app.Runtime.Command)
|
||||
}
|
||||
if app.Command != "/usr/bin/httpd -f" {
|
||||
t.Errorf("Tasks[0].Command = %q", app.Command)
|
||||
}
|
||||
side := spec.Tasks[1]
|
||||
if side.Name != "sidecar" {
|
||||
t.Errorf("Tasks[1].Name = %q, want sidecar", side.Name)
|
||||
}
|
||||
if side.Runtime == nil || side.Runtime.OneOf != "wasm" {
|
||||
t.Errorf("Tasks[1].Runtime = %+v, want one_of=wasm", side.Runtime)
|
||||
}
|
||||
if side.Command != "/bin/wasm-runner sidecar.wasm" {
|
||||
t.Errorf("Tasks[1].Command = %q", side.Command)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_TasksBlockWithEnv(t *testing.T) {
|
||||
// P06: a task group task carrying an env overlay.
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"tasks:\n" +
|
||||
" - name: app\n" +
|
||||
" command: /usr/bin/httpd\n" +
|
||||
" env:\n" +
|
||||
" LOG_LEVEL: debug\n" +
|
||||
" REGION: us\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Tasks) != 1 {
|
||||
t.Fatalf("Tasks = %d, want 1", len(spec.Tasks))
|
||||
}
|
||||
task := spec.Tasks[0]
|
||||
if task.Env == nil {
|
||||
t.Fatal("Tasks[0].Env is nil")
|
||||
}
|
||||
if got := task.Env["LOG_LEVEL"]; got != "debug" {
|
||||
t.Errorf("Env[LOG_LEVEL] = %q, want debug", got)
|
||||
}
|
||||
if got := task.Env["REGION"]; got != "us" {
|
||||
t.Errorf("Env[REGION] = %q, want us", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_TasksBlockInheritsTopLevelRuntime(t *testing.T) {
|
||||
// P06: when a task omits its own runtime, the top-level runtime
|
||||
// is the per-group default. The parser must NOT create a task
|
||||
// runtime when the task block lacks a `runtime:` sub-block; the
|
||||
// emitter/validator resolve the default from spec.Runtime.
|
||||
input := "---\n" +
|
||||
"kind: Service\n" +
|
||||
"name: web\n" +
|
||||
"runtime:\n" +
|
||||
" one_of: process\n" +
|
||||
" command: /bin/default\n" +
|
||||
"tasks:\n" +
|
||||
" - name: app\n" +
|
||||
" command: /bin/app\n" +
|
||||
" - name: sidecar\n" +
|
||||
" command: /bin/sidecar\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if spec.Runtime == nil || spec.Runtime.OneOf != "process" {
|
||||
t.Fatalf("top-level runtime not parsed: %+v", spec.Runtime)
|
||||
}
|
||||
if len(spec.Tasks) != 2 {
|
||||
t.Fatalf("Tasks = %d, want 2", len(spec.Tasks))
|
||||
}
|
||||
for i, task := range spec.Tasks {
|
||||
if task.Runtime != nil {
|
||||
t.Errorf("Tasks[%d].Runtime should be nil (inherit top-level), got %+v", i, task.Runtime)
|
||||
}
|
||||
}
|
||||
if spec.Tasks[0].Name != "app" || spec.Tasks[1].Name != "sidecar" {
|
||||
t.Errorf("task names = %q, %q", spec.Tasks[0].Name, spec.Tasks[1].Name)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMarkdown_NoTasksBackwardCompat(t *testing.T) {
|
||||
// Backward compat: a spec with no `tasks:` block parses as a
|
||||
// single-process alloc; spec.Tasks must be empty/nil.
|
||||
input := "---\n" +
|
||||
"kind: Job\n" +
|
||||
"name: backup\n" +
|
||||
"runtime:\n" +
|
||||
" one_of: process\n" +
|
||||
" command: /bin/rsync\n" +
|
||||
"---\nbody\n"
|
||||
spec, err := ParseMarkdown([]byte(input))
|
||||
if err != nil {
|
||||
t.Fatalf("ParseMarkdown: %v", err)
|
||||
}
|
||||
if len(spec.Tasks) != 0 {
|
||||
t.Fatalf("Tasks = %d, want 0 (backward compat)", len(spec.Tasks))
|
||||
}
|
||||
if spec.Runtime == nil || spec.Runtime.Command != "/bin/rsync" {
|
||||
t.Errorf("Runtime = %+v, want command=/bin/rsync", spec.Runtime)
|
||||
}
|
||||
}
|
||||
+26
-27
@@ -2,7 +2,6 @@ package jobspec
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
"github.com/hashicorp/hcl/v2"
|
||||
@@ -10,16 +9,24 @@ import (
|
||||
"github.com/hashicorp/hcl/v2/hclsimple"
|
||||
)
|
||||
|
||||
// Spec is the legacy HCL-only jobspec shape. It is retained for the
|
||||
// v0.9→v0.10 migration window (REQ-090) and is populated by ParseHCLLegacy.
|
||||
//
|
||||
// Deprecated: HCL is legacy per R-013; new code should consume the
|
||||
// unified *WorkloadSpec returned by ParseFile/Dispatch (see
|
||||
// dispatch.go and markdown.go).
|
||||
type Spec struct {
|
||||
Job JobSpec `hcl:"job,block"`
|
||||
Tasks []TaskSpec `hcl:"task,block"`
|
||||
}
|
||||
|
||||
// JobSpec is the legacy HCL job block.
|
||||
type JobSpec struct {
|
||||
Name string `hcl:"name,label"`
|
||||
Type string `hcl:"type,optional"`
|
||||
}
|
||||
|
||||
// TaskSpec is the legacy HCL task block.
|
||||
type TaskSpec struct {
|
||||
Name string `hcl:"name,label"`
|
||||
Command string `hcl:"command"`
|
||||
@@ -27,34 +34,14 @@ type TaskSpec struct {
|
||||
Env []string `hcl:"env,optional"`
|
||||
}
|
||||
|
||||
func ParseFile(path string) (*Spec, error) {
|
||||
data, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read spec file: %w", err)
|
||||
}
|
||||
return Parse(data, path)
|
||||
}
|
||||
|
||||
func Parse(data []byte, filename string) (*Spec, error) {
|
||||
var spec Spec
|
||||
err := hclsimple.Decode(filename, data, nil, &spec)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("decode hcl: %w", err)
|
||||
}
|
||||
if spec.Job.Name == "" {
|
||||
return nil, fmt.Errorf("spec missing job name")
|
||||
}
|
||||
if len(spec.Tasks) == 0 {
|
||||
return nil, fmt.Errorf("spec must have at least one task")
|
||||
}
|
||||
for i, t := range spec.Tasks {
|
||||
if t.Command == "" {
|
||||
return nil, fmt.Errorf("task[%d] (%s) missing command", i, t.Name)
|
||||
}
|
||||
}
|
||||
return &spec, nil
|
||||
// hclDecode wraps hclsimple.Decode for testability.
|
||||
func hclDecode(filename string, data []byte, spec *Spec) error {
|
||||
return hclsimple.Decode(filename, data, nil, spec)
|
||||
}
|
||||
|
||||
// Validate is the legacy HCL Spec validator retained for the migration
|
||||
// window (REQ-090). New code should use validateWorkload on a
|
||||
// *WorkloadSpec.
|
||||
func (s *Spec) Validate() error {
|
||||
if strings.TrimSpace(s.Job.Name) == "" {
|
||||
return fmt.Errorf("job name is required")
|
||||
@@ -65,5 +52,17 @@ func (s *Spec) Validate() error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// Parse is the original HCL-only entry point retained for backward
|
||||
// compatibility with direct HCL callers during the v0.9→v0.10 migration
|
||||
// window (REQ-090). New code should call the dispatcher ParseFile (which
|
||||
// returns *WorkloadSpec) or ParseHCL (which adapts HCL into
|
||||
// *WorkloadSpec).
|
||||
//
|
||||
// Deprecated: use ParseFile (dispatcher) or ParseHCL (adapter). HCL is
|
||||
// legacy per R-013.
|
||||
func Parse(data []byte, filename string) (*Spec, error) {
|
||||
return ParseHCLLegacy(data, filename)
|
||||
}
|
||||
|
||||
var _ = hcl.Diagnostics{}
|
||||
var _ = gohcl.DecodeBody
|
||||
|
||||
@@ -130,9 +130,9 @@ func TestParse_GoldenFiles(t *testing.T) {
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
path := filepath.Join("testdata", tc.file)
|
||||
spec, err := ParseFile(path)
|
||||
spec, err := ParseHCLFile(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseFile(%s): %v", tc.file, err)
|
||||
t.Fatalf("ParseHCLFile(%s): %v", tc.file, err)
|
||||
}
|
||||
if spec.Job.Name != tc.wantJob {
|
||||
t.Errorf("job name = %q, want %q", spec.Job.Name, tc.wantJob)
|
||||
@@ -254,9 +254,9 @@ func TestSpec_Validate(t *testing.T) {
|
||||
|
||||
func TestSpec_Validate_RoundTripFromParse(t *testing.T) {
|
||||
path := filepath.Join("testdata", "valid_single_task.hcl")
|
||||
spec, err := ParseFile(path)
|
||||
spec, err := ParseHCLFile(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseFile: %v", err)
|
||||
t.Fatalf("ParseHCLFile: %v", err)
|
||||
}
|
||||
if err := spec.Validate(); err != nil {
|
||||
t.Errorf("Validate on parsed spec: %v", err)
|
||||
|
||||
@@ -0,0 +1,308 @@
|
||||
package ns
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// ParseNSMd reads an ns.md file, extracts the YAML frontmatter, and
|
||||
// parses it into a *NSConfig. The body after the closing `---` is
|
||||
// discarded (namespace declarations do not require body preservation
|
||||
// like jobspecs do under R-015; we keep the parser minimal and
|
||||
// consistent with internal/config/markdown.go).
|
||||
//
|
||||
// Frontmatter keys (R-014):
|
||||
//
|
||||
// kind: Namespace (required; must be "Namespace")
|
||||
// name: <ns-name> (required)
|
||||
// parents: ["a", "b"] (optional; default empty)
|
||||
// inherits_env: true (optional; default true)
|
||||
// inherits_secrets: true (optional; default true)
|
||||
// quota: {...} (optional; parsed but not surfaced here)
|
||||
// acl: {...} (optional; parsed but not surfaced here)
|
||||
//
|
||||
// The parser is a minimal hand-rolled YAML-ish key:value reader (no
|
||||
// new dependencies; gopkg.in/yaml.v3 is intentionally NOT added). It
|
||||
// supports flat scalar keys and the inline flow-array form
|
||||
// `["a", "b"]` for `parents`. Nested mappings (quota, acl) are
|
||||
// recognized as keys but their contents are currently ignored — they
|
||||
// are reserved for later phases.
|
||||
func ParseNSMd(path string) (*NSConfig, error) {
|
||||
data, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read %s: %w", path, err)
|
||||
}
|
||||
content := string(data)
|
||||
|
||||
block, ok := extractFrontmatter(content)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("parse %s: missing frontmatter", path)
|
||||
}
|
||||
if strings.TrimSpace(block) == "" {
|
||||
return nil, fmt.Errorf("parse %s: missing frontmatter", path)
|
||||
}
|
||||
|
||||
cfg, err := parseNSFrontmatter(block, path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if cfg.Name == "" {
|
||||
return nil, fmt.Errorf("parse %s: missing name", path)
|
||||
}
|
||||
return cfg, nil
|
||||
}
|
||||
|
||||
// extractFrontmatter returns the YAML block between the first pair of
|
||||
// `---` delimiters and whether a frontmatter block was present.
|
||||
func extractFrontmatter(content string) (string, bool) {
|
||||
trimmed := strings.TrimLeft(content, "\r\n\t ")
|
||||
if !strings.HasPrefix(trimmed, "---") {
|
||||
return "", false
|
||||
}
|
||||
rest := trimmed[3:]
|
||||
rest = strings.TrimLeft(rest, "\r\n")
|
||||
idx := strings.Index(rest, "\n---")
|
||||
if idx < 0 {
|
||||
return "", false
|
||||
}
|
||||
return rest[:idx], true
|
||||
}
|
||||
|
||||
// parseNSFrontmatter parses a minimal YAML-ish frontmatter block into
|
||||
// a *NSConfig. See ParseNSMd for the supported keys.
|
||||
func parseNSFrontmatter(block, path string) (*NSConfig, error) {
|
||||
cfg := &NSConfig{
|
||||
InheritsEnv: true,
|
||||
InheritsSecrets: true,
|
||||
}
|
||||
kind := ""
|
||||
|
||||
lines := strings.Split(block, "\n")
|
||||
for lineNo, raw := range lines {
|
||||
line := stripNSComment(raw)
|
||||
if strings.TrimSpace(line) == "" {
|
||||
continue
|
||||
}
|
||||
if countIndent(line) > 0 {
|
||||
// Indented line under a nested mapping header (quota, acl).
|
||||
// Recognized but ignored at this phase.
|
||||
continue
|
||||
}
|
||||
key, val, ok := splitKV(strings.TrimSpace(line))
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("parse %s: line %d: malformed key:value", path, lineNo+1)
|
||||
}
|
||||
switch key {
|
||||
case "kind":
|
||||
kind = strings.TrimSpace(unquote(val))
|
||||
case "name":
|
||||
cfg.Name = strings.TrimSpace(unquote(val))
|
||||
case "parents":
|
||||
parents, err := parseStringArray(val)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("parse %s: line %d: parents: %w", path, lineNo+1, err)
|
||||
}
|
||||
cfg.Parents = parents
|
||||
case "inherits_env":
|
||||
cfg.InheritsEnv = parseBool(val)
|
||||
case "inherits_secrets":
|
||||
cfg.InheritsSecrets = parseBool(val)
|
||||
case "quota", "acl":
|
||||
// Reserved nested-mapping keys; recognized, contents ignored.
|
||||
default:
|
||||
// Unknown keys are ignored (forward-compat with future
|
||||
// frontmatter additions).
|
||||
}
|
||||
}
|
||||
|
||||
if kind == "" {
|
||||
return nil, fmt.Errorf("parse %s: missing kind", path)
|
||||
}
|
||||
if kind != "Namespace" {
|
||||
return nil, fmt.Errorf("parse %s: kind %q is not %q", path, kind, "Namespace")
|
||||
}
|
||||
return cfg, nil
|
||||
}
|
||||
|
||||
// parseStringArray parses an inline YAML flow-array of scalars, e.g.
|
||||
// `["a", "b"]` or `['a', 'b']` or `[a, b]`. Returns an error if the
|
||||
// value is not a flow-array. Empty array `[]` returns nil.
|
||||
func parseStringArray(val string) ([]string, error) {
|
||||
val = strings.TrimSpace(val)
|
||||
if val == "" {
|
||||
return nil, nil
|
||||
}
|
||||
if !strings.HasPrefix(val, "[") || !strings.HasSuffix(val, "]") {
|
||||
return nil, fmt.Errorf("expected [..] array, got %q", val)
|
||||
}
|
||||
inner := strings.TrimSpace(val[1 : len(val)-1])
|
||||
if inner == "" {
|
||||
return nil, nil
|
||||
}
|
||||
parts := splitFlowItems(inner)
|
||||
out := make([]string, 0, len(parts))
|
||||
for _, p := range parts {
|
||||
p = strings.TrimSpace(p)
|
||||
if p == "" {
|
||||
continue
|
||||
}
|
||||
out = append(out, unquote(p))
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// splitFlowItems splits a comma-separated flow-array body, respecting
|
||||
// single and double quotes.
|
||||
func splitFlowItems(s string) []string {
|
||||
var out []string
|
||||
inSingle := false
|
||||
inDouble := false
|
||||
start := 0
|
||||
for i := 0; i < len(s); i++ {
|
||||
c := s[i]
|
||||
switch c {
|
||||
case '\'':
|
||||
if !inDouble {
|
||||
inSingle = !inSingle
|
||||
}
|
||||
case '"':
|
||||
if !inSingle {
|
||||
inDouble = !inDouble
|
||||
}
|
||||
case ',':
|
||||
if !inSingle && !inDouble {
|
||||
out = append(out, s[start:i])
|
||||
start = i + 1
|
||||
}
|
||||
}
|
||||
}
|
||||
out = append(out, s[start:])
|
||||
return out
|
||||
}
|
||||
|
||||
// parseBool parses a YAML-ish bool (true/false/yes/no), defaulting to
|
||||
// true for empty (matches the inherits_* defaults).
|
||||
func parseBool(val string) bool {
|
||||
switch strings.ToLower(strings.TrimSpace(unquote(val))) {
|
||||
case "false", "no", "off", "0":
|
||||
return false
|
||||
default:
|
||||
return true
|
||||
}
|
||||
}
|
||||
|
||||
// ParseNSMdDir walks `<root>/*/ns.md`, parses each, and returns the
|
||||
// config map keyed by namespace name. The `cluster` directory is
|
||||
// skipped (it is not a namespace). The `_defaults` namespace MUST
|
||||
// exist; if missing, an error is returned.
|
||||
func ParseNSMdDir(root string) (map[string]*NSConfig, error) {
|
||||
entries, err := os.ReadDir(root)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read namespace root %s: %w", root, err)
|
||||
}
|
||||
|
||||
configs := make(map[string]*NSConfig)
|
||||
var found []string
|
||||
for _, ent := range entries {
|
||||
if !ent.IsDir() {
|
||||
continue
|
||||
}
|
||||
if ent.Name() == "cluster" {
|
||||
continue
|
||||
}
|
||||
nsMd := filepath.Join(root, ent.Name(), "ns.md")
|
||||
info, err := os.Stat(nsMd)
|
||||
if err != nil || info.IsDir() {
|
||||
continue
|
||||
}
|
||||
cfg, err := ParseNSMd(nsMd)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
// The directory name and the frontmatter `name` should match;
|
||||
// we key by the frontmatter name (canonical) but also accept
|
||||
// the directory name if frontmatter name is missing (the
|
||||
// parser already errors on missing name, so this is defensive).
|
||||
key := cfg.Name
|
||||
if key == "" {
|
||||
key = ent.Name()
|
||||
}
|
||||
if _, dup := configs[key]; dup {
|
||||
return nil, fmt.Errorf("duplicate namespace %q (from %s)", key, nsMd)
|
||||
}
|
||||
configs[key] = cfg
|
||||
found = append(found, key)
|
||||
}
|
||||
|
||||
if _, ok := configs[defaultsName]; !ok {
|
||||
sort.Strings(found)
|
||||
names := strings.Join(found, ", ")
|
||||
if names == "" {
|
||||
names = "(none)"
|
||||
}
|
||||
return nil, fmt.Errorf("namespace root %s: implicit root %q not found (found: %s)", root, defaultsName, names)
|
||||
}
|
||||
return configs, nil
|
||||
}
|
||||
|
||||
func countIndent(s string) int {
|
||||
n := 0
|
||||
for _, r := range s {
|
||||
if r == ' ' || r == '\t' {
|
||||
n++
|
||||
continue
|
||||
}
|
||||
break
|
||||
}
|
||||
return n
|
||||
}
|
||||
|
||||
func splitKV(s string) (key, val string, ok bool) {
|
||||
idx := strings.Index(s, ":")
|
||||
if idx < 0 {
|
||||
return "", "", false
|
||||
}
|
||||
key = strings.TrimSpace(s[:idx])
|
||||
val = strings.TrimSpace(s[idx+1:])
|
||||
if key == "" {
|
||||
return "", "", false
|
||||
}
|
||||
return key, val, true
|
||||
}
|
||||
|
||||
func stripNSComment(s string) string {
|
||||
inSingle := false
|
||||
inDouble := false
|
||||
for i := 0; i < len(s); i++ {
|
||||
c := s[i]
|
||||
switch c {
|
||||
case '\'':
|
||||
if !inDouble {
|
||||
inSingle = !inSingle
|
||||
}
|
||||
case '"':
|
||||
if !inSingle {
|
||||
inDouble = !inDouble
|
||||
}
|
||||
case '#':
|
||||
if !inSingle && !inDouble {
|
||||
if i == 0 || s[i-1] == ' ' || s[i-1] == '\t' {
|
||||
return s[:i]
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
func unquote(s string) string {
|
||||
if len(s) >= 2 {
|
||||
if (s[0] == '"' && s[len(s)-1] == '"') || (s[0] == '\'' && s[len(s)-1] == '\'') {
|
||||
return s[1 : len(s)-1]
|
||||
}
|
||||
}
|
||||
return s
|
||||
}
|
||||
@@ -0,0 +1,227 @@
|
||||
package ns
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func writeNSMd(t *testing.T, path, content string) {
|
||||
t.Helper()
|
||||
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
|
||||
t.Fatalf("mkdir: %v", err)
|
||||
}
|
||||
if err := os.WriteFile(path, []byte(content), 0o644); err != nil {
|
||||
t.Fatalf("write %s: %v", path, err)
|
||||
}
|
||||
}
|
||||
|
||||
const validNSMd = `---
|
||||
kind: Namespace
|
||||
name: prod
|
||||
parents: ["_defaults"]
|
||||
inherits_env: true
|
||||
inherits_secrets: true
|
||||
quota:
|
||||
cpu: 4
|
||||
acl:
|
||||
admin: ops
|
||||
---
|
||||
# Prod namespace
|
||||
|
||||
This body is ignored.
|
||||
`
|
||||
|
||||
func TestParseNSMdValid(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
path := filepath.Join(tmp, "ns.md")
|
||||
writeNSMd(t, path, validNSMd)
|
||||
cfg, err := ParseNSMd(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseNSMd: %v", err)
|
||||
}
|
||||
if cfg.Name != "prod" {
|
||||
t.Errorf("name = %q, want prod", cfg.Name)
|
||||
}
|
||||
if !eqSlice(cfg.Parents, []string{"_defaults"}) {
|
||||
t.Errorf("parents = %v, want [_defaults]", cfg.Parents)
|
||||
}
|
||||
if !cfg.InheritsEnv || !cfg.InheritsSecrets {
|
||||
t.Errorf("inherits_env=%v inherits_secrets=%v, want both true", cfg.InheritsEnv, cfg.InheritsSecrets)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdMissingFrontmatter(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
path := filepath.Join(tmp, "ns.md")
|
||||
writeNSMd(t, path, "# just a body, no frontmatter\n")
|
||||
_, err := ParseNSMd(path)
|
||||
if err == nil {
|
||||
t.Fatal("expected missing frontmatter error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "missing frontmatter") {
|
||||
t.Errorf("error = %q, want contains 'missing frontmatter'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdEmptyFrontmatter(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
path := filepath.Join(tmp, "ns.md")
|
||||
writeNSMd(t, path, "---\n---\nbody\n")
|
||||
_, err := ParseNSMd(path)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty frontmatter, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdWrongKind(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
path := filepath.Join(tmp, "ns.md")
|
||||
writeNSMd(t, path, "---\nkind: Job\nname: x\n---\n")
|
||||
_, err := ParseNSMd(path)
|
||||
if err == nil {
|
||||
t.Fatal("expected wrong-kind error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "not \"Namespace\"") {
|
||||
t.Errorf("error = %q, want contains 'is not \"Namespace\"'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdMissingKind(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
path := filepath.Join(tmp, "ns.md")
|
||||
writeNSMd(t, path, "---\nname: x\n---\n")
|
||||
_, err := ParseNSMd(path)
|
||||
if err == nil {
|
||||
t.Fatal("expected missing kind error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "missing kind") {
|
||||
t.Errorf("error = %q, want contains 'missing kind'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdMissingName(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
path := filepath.Join(tmp, "ns.md")
|
||||
writeNSMd(t, path, "---\nkind: Namespace\n---\n")
|
||||
_, err := ParseNSMd(path)
|
||||
if err == nil {
|
||||
t.Fatal("expected missing name error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "missing name") {
|
||||
t.Errorf("error = %q, want contains 'missing name'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdParentsUnquoted(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
path := filepath.Join(tmp, "ns.md")
|
||||
writeNSMd(t, path, "---\nkind: Namespace\nname: x\nparents: [a, b]\n---\n")
|
||||
cfg, err := ParseNSMd(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseNSMd: %v", err)
|
||||
}
|
||||
if !eqSlice(cfg.Parents, []string{"a", "b"}) {
|
||||
t.Errorf("parents = %v, want [a b]", cfg.Parents)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdParentsEmpty(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
path := filepath.Join(tmp, "ns.md")
|
||||
writeNSMd(t, path, "---\nkind: Namespace\nname: x\nparents: []\n---\n")
|
||||
cfg, err := ParseNSMd(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseNSMd: %v", err)
|
||||
}
|
||||
if len(cfg.Parents) != 0 {
|
||||
t.Errorf("parents = %v, want empty", cfg.Parents)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdInheritsFalse(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
path := filepath.Join(tmp, "ns.md")
|
||||
writeNSMd(t, path, "---\nkind: Namespace\nname: x\ninherits_env: false\ninherits_secrets: no\n---\n")
|
||||
cfg, err := ParseNSMd(path)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseNSMd: %v", err)
|
||||
}
|
||||
if cfg.InheritsEnv {
|
||||
t.Errorf("inherits_env should be false")
|
||||
}
|
||||
if cfg.InheritsSecrets {
|
||||
t.Errorf("inherits_secrets should be false")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdMissingFile(t *testing.T) {
|
||||
_, err := ParseNSMd(filepath.Join(t.TempDir(), "nope.md"))
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing file")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdDirHappy(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
writeNSMd(t, filepath.Join(root, "_defaults", "ns.md"), "---\nkind: Namespace\nname: _defaults\n---\n")
|
||||
writeNSMd(t, filepath.Join(root, "prod", "ns.md"), validNSMd)
|
||||
|
||||
cfgs, err := ParseNSMdDir(root)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseNSMdDir: %v", err)
|
||||
}
|
||||
if _, ok := cfgs["_defaults"]; !ok {
|
||||
t.Errorf("missing _defaults in %v", cfgs)
|
||||
}
|
||||
if _, ok := cfgs["prod"]; !ok {
|
||||
t.Errorf("missing prod in %v", cfgs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdDirMissingDefaults(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
writeNSMd(t, filepath.Join(root, "prod", "ns.md"), validNSMd)
|
||||
_, err := ParseNSMdDir(root)
|
||||
if err == nil {
|
||||
t.Fatal("expected missing _defaults error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "_defaults") {
|
||||
t.Errorf("error = %q, want contains _defaults", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdDirSkipsCluster(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
writeNSMd(t, filepath.Join(root, "_defaults", "ns.md"), "---\nkind: Namespace\nname: _defaults\n---\n")
|
||||
// cluster/ contains a ns.md-shaped file but must be skipped.
|
||||
writeNSMd(t, filepath.Join(root, "cluster", "ns.md"), "---\nkind: Namespace\nname: cluster\n---\n")
|
||||
|
||||
cfgs, err := ParseNSMdDir(root)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseNSMdDir: %v", err)
|
||||
}
|
||||
if _, ok := cfgs["cluster"]; ok {
|
||||
t.Errorf("cluster should be skipped, present in %v", cfgs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdDirNoFiles(t *testing.T) {
|
||||
root := t.TempDir()
|
||||
_, err := ParseNSMdDir(root)
|
||||
if err == nil {
|
||||
t.Fatal("expected missing _defaults error on empty dir, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseNSMdDirNotADir(t *testing.T) {
|
||||
tmp := t.TempDir()
|
||||
// Create a file with the same name as the expected root dir.
|
||||
root := filepath.Join(tmp, "notadir")
|
||||
writeNSMd(t, root, "x")
|
||||
_, err := ParseNSMdDir(root)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for non-dir root")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,252 @@
|
||||
// Package ns implements the namespace inheritance resolver (REQ-082)
|
||||
// and the ns.md frontmatter parser used by `orca ns` CLI subcommands.
|
||||
//
|
||||
// The resolver is a PURE function (no I/O): it takes a map of parsed
|
||||
// namespace configs keyed by name and returns a map of resolved
|
||||
// namespaces with merged env and unioned constraints. The inheritance
|
||||
// model is:
|
||||
//
|
||||
// - Each namespace declares zero or more parents in `ns.md`
|
||||
// frontmatter (`parents: ["ns1", "ns2"]`).
|
||||
// - The implicit root namespace `_defaults` (R-002 D-159) always
|
||||
// exists and has no parents; it is ALWAYS appended as the last
|
||||
// element of the chain (D-185).
|
||||
// - Opting out of `_defaults` is impossible (D-187): even with
|
||||
// `parents: []`, `_defaults` still appears at the end of the chain.
|
||||
// - Merge semantics: child overrides parent for scalars (env keys);
|
||||
// arrays union (child constraints add to parent constraints, with
|
||||
// duplicates removed, order: most-specific first).
|
||||
// - The chain order is most-specific first, `_defaults` last.
|
||||
// - `_defaults` may be listed explicitly in `parents`; the explicit
|
||||
// listing is de-duped silently (still appears once, at the end).
|
||||
// - Misordering (`parents: ["_defaults", "x"]`) is rejected: an
|
||||
// explicit `_defaults` entry must be the only entry (or omitted).
|
||||
// - Cycle detection uses DFS with a visited set; a cycle returns an
|
||||
// error with the cycle path.
|
||||
// - Missing parents return "parent X not found".
|
||||
package ns
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"sort"
|
||||
)
|
||||
|
||||
const defaultsName = "_defaults"
|
||||
|
||||
// NSConfig is a parsed namespace declaration from ns.md frontmatter.
|
||||
// The resolver consumes this; the parser populates it.
|
||||
type NSConfig struct {
|
||||
Name string
|
||||
Parents []string
|
||||
Env map[string]string
|
||||
Constraints []string
|
||||
InheritsEnv bool
|
||||
InheritsSecrets bool
|
||||
}
|
||||
|
||||
// ResolvedNS is the output of the resolver: the namespace with its
|
||||
// fully-merged env and unioned constraints, plus the ordered
|
||||
// inheritance chain (most-specific first, `_defaults` last).
|
||||
type ResolvedNS struct {
|
||||
Name string
|
||||
Chain []string
|
||||
Env map[string]string
|
||||
Constraints []string
|
||||
}
|
||||
|
||||
// Resolve walks the parent chain for each namespace, merges env (child
|
||||
// wins scalars), unions constraints (child adds to parent, de-duped),
|
||||
// and detects cycles. It is PURE (no I/O). The empty-configs case
|
||||
// returns an empty map and no error.
|
||||
//
|
||||
// The `_defaults` namespace is ALWAYS the last element of every chain
|
||||
// (D-185); opting out is impossible (D-187). An explicit `_defaults`
|
||||
// entry in `parents` is de-duped silently. Misordering (e.g.
|
||||
// `parents: ["_defaults", "x"]`) is rejected.
|
||||
func Resolve(configs map[string]*NSConfig) (map[string]*ResolvedNS, error) {
|
||||
if len(configs) == 0 {
|
||||
return map[string]*ResolvedNS{}, nil
|
||||
}
|
||||
|
||||
// Validate each config's parents reference exists and the
|
||||
// _defaults entry (if explicit) is the only entry.
|
||||
for name, cfg := range configs {
|
||||
if cfg == nil {
|
||||
return nil, fmt.Errorf("namespace %q has nil config", name)
|
||||
}
|
||||
for _, p := range cfg.Parents {
|
||||
if p == defaultsName {
|
||||
// Explicit _defaults must be the only parent.
|
||||
if len(cfg.Parents) != 1 {
|
||||
return nil, fmt.Errorf("namespace %q: %s must be the only parent if listed explicitly (misordering rejected)", name, defaultsName)
|
||||
}
|
||||
continue
|
||||
}
|
||||
if _, ok := configs[p]; !ok {
|
||||
return nil, fmt.Errorf("namespace %q: parent %q not found", name, p)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// `_defaults` must be present in the configs map (the parser
|
||||
// enforces this for ParseNSMdDir; Resolve trusts its input but
|
||||
// still requires _defaults to exist for chain assembly).
|
||||
if _, ok := configs[defaultsName]; !ok {
|
||||
return nil, fmt.Errorf("namespace %q not found (implicit root must be present)", defaultsName)
|
||||
}
|
||||
|
||||
resolved := make(map[string]*ResolvedNS, len(configs))
|
||||
// Resolve in deterministic order for stable error reporting.
|
||||
names := make([]string, 0, len(configs))
|
||||
for n := range configs {
|
||||
names = append(names, n)
|
||||
}
|
||||
sort.Strings(names)
|
||||
|
||||
for _, name := range names {
|
||||
r, err := resolveOne(configs, name)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
resolved[name] = r
|
||||
}
|
||||
return resolved, nil
|
||||
}
|
||||
|
||||
// resolveOne resolves a single namespace. The chain is built by walking
|
||||
// parents depth-first in POST-order (least-specific first), then
|
||||
// reversing so the returned chain is most-specific first with
|
||||
// `_defaults` last (D-185). Cycle detection uses a visiting set.
|
||||
func resolveOne(configs map[string]*NSConfig, name string) (*ResolvedNS, error) {
|
||||
post, err := buildChain(configs, name)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
// post is least-specific first; reverse to most-specific first.
|
||||
reverseStrings(post)
|
||||
chain := post
|
||||
|
||||
// Env: child (most-specific) wins. Walk least-specific to
|
||||
// most-specific (end -> beginning) so later writes override.
|
||||
env := make(map[string]string)
|
||||
for i := len(chain) - 1; i >= 0; i-- {
|
||||
c := configs[chain[i]]
|
||||
if c == nil {
|
||||
continue
|
||||
}
|
||||
for k, v := range c.Env {
|
||||
env[k] = v
|
||||
}
|
||||
}
|
||||
|
||||
// Constraints: union, child (most-specific) first. Walk the chain
|
||||
// front-to-back (most-specific first) and append unseen items.
|
||||
constraintsSeen := make(map[string]bool)
|
||||
var constraints []string
|
||||
for _, ns := range chain {
|
||||
c := configs[ns]
|
||||
if c == nil {
|
||||
continue
|
||||
}
|
||||
for _, con := range c.Constraints {
|
||||
if !constraintsSeen[con] {
|
||||
constraintsSeen[con] = true
|
||||
constraints = append(constraints, con)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return &ResolvedNS{
|
||||
Name: name,
|
||||
Chain: chain,
|
||||
Env: env,
|
||||
Constraints: constraints,
|
||||
}, nil
|
||||
}
|
||||
|
||||
// buildChain walks parents depth-first and returns the chain in
|
||||
// POST-order (least-specific first, `_defaults` first). The caller
|
||||
// reverses to get most-specific first. Cycle detection uses the
|
||||
// visiting set: a node currently being walked indicates a back-edge.
|
||||
func buildChain(configs map[string]*NSConfig, name string) ([]string, error) {
|
||||
var post []string
|
||||
seen := make(map[string]bool) // final chain membership (de-dup)
|
||||
visiting := make(map[string]bool)
|
||||
if err := dfsChain(configs, name, &post, seen, visiting); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
// `_defaults` is the implicit root: it must be the FIRST element
|
||||
// in post-order (so it ends up LAST after reversal). If it was not
|
||||
// reached via parents (no explicit listing and no chain leads to
|
||||
// it), prepend it.
|
||||
if !seen[defaultsName] {
|
||||
post = append([]string{defaultsName}, post...)
|
||||
seen[defaultsName] = true
|
||||
}
|
||||
return post, nil
|
||||
}
|
||||
|
||||
// dfsChain appends each node AFTER its parents (post-order), producing
|
||||
// least-specific first. Cycle detection uses the visiting set.
|
||||
func dfsChain(configs map[string]*NSConfig, name string, post *[]string, seen, visiting map[string]bool) error {
|
||||
if visiting[name] {
|
||||
return fmt.Errorf("cycle detected: %s", cyclePath(visiting, configs, name))
|
||||
}
|
||||
if seen[name] {
|
||||
return nil
|
||||
}
|
||||
visiting[name] = true
|
||||
cfg := configs[name]
|
||||
if cfg != nil {
|
||||
for _, p := range cfg.Parents {
|
||||
if err := dfsChain(configs, p, post, seen, visiting); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
}
|
||||
delete(visiting, name)
|
||||
seen[name] = true
|
||||
*post = append(*post, name)
|
||||
return nil
|
||||
}
|
||||
|
||||
// cyclePath reconstructs a readable cycle path from the visiting set.
|
||||
// Since visiting is a set (not ordered), we reconstruct by re-walking
|
||||
// parents from the offending node until we revisit it.
|
||||
func cyclePath(visiting map[string]bool, configs map[string]*NSConfig, start string) string {
|
||||
// Walk parents from start, collecting names until we hit start
|
||||
// again or run out.
|
||||
var path []string
|
||||
cur := start
|
||||
for i := 0; i < len(visiting)+1; i++ {
|
||||
path = append(path, cur)
|
||||
cfg := configs[cur]
|
||||
if cfg == nil || len(cfg.Parents) == 0 {
|
||||
break
|
||||
}
|
||||
next := cfg.Parents[0]
|
||||
if next == start {
|
||||
path = append(path, next)
|
||||
break
|
||||
}
|
||||
cur = next
|
||||
}
|
||||
return joinArrows(path)
|
||||
}
|
||||
|
||||
func joinArrows(parts []string) string {
|
||||
out := ""
|
||||
for i, p := range parts {
|
||||
if i > 0 {
|
||||
out += " -> "
|
||||
}
|
||||
out += p
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
func reverseStrings(s []string) {
|
||||
for i, j := 0, len(s)-1; i < j; i, j = i+1, j-1 {
|
||||
s[i], s[j] = s[j], s[i]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,231 @@
|
||||
package ns
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestResolveEmptyConfigs(t *testing.T) {
|
||||
out, err := Resolve(map[string]*NSConfig{})
|
||||
if err != nil {
|
||||
t.Fatalf("Resolve empty: unexpected error: %v", err)
|
||||
}
|
||||
if len(out) != 0 {
|
||||
t.Fatalf("Resolve empty: want empty map, got %d entries", len(out))
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveSingleNoParents(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName, Env: map[string]string{"A": "1"}},
|
||||
"x": {Name: "x", Env: map[string]string{"B": "2"}},
|
||||
}
|
||||
out, err := Resolve(cfgs)
|
||||
if err != nil {
|
||||
t.Fatalf("Resolve: %v", err)
|
||||
}
|
||||
r := out["x"]
|
||||
if r == nil {
|
||||
t.Fatal("missing resolved x")
|
||||
}
|
||||
if !eqSlice(r.Chain, []string{"x", defaultsName}) {
|
||||
t.Errorf("chain = %v, want [x _defaults]", r.Chain)
|
||||
}
|
||||
if r.Env["A"] != "1" || r.Env["B"] != "2" {
|
||||
t.Errorf("env = %v, want A=1 B=2", r.Env)
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveChildOverridesParentScalar(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName, Env: map[string]string{"K": "parent"}},
|
||||
"child": {Name: "child", Parents: []string{defaultsName}, Env: map[string]string{"K": "child"}},
|
||||
}
|
||||
out, err := Resolve(cfgs)
|
||||
if err != nil {
|
||||
t.Fatalf("Resolve: %v", err)
|
||||
}
|
||||
if got := out["child"].Env["K"]; got != "child" {
|
||||
t.Errorf("child K = %q, want %q (child overrides parent)", got, "child")
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveArraysUnion(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName, Constraints: []string{"a", "b"}},
|
||||
"x": {Name: "x", Parents: []string{defaultsName}, Constraints: []string{"c", "a"}},
|
||||
}
|
||||
out, err := Resolve(cfgs)
|
||||
if err != nil {
|
||||
t.Fatalf("Resolve: %v", err)
|
||||
}
|
||||
c := out["x"].Constraints
|
||||
// Union de-duped; most-specific (x) first.
|
||||
if !eqSlice(c, []string{"c", "a", "b"}) {
|
||||
t.Errorf("constraints = %v, want [c a b]", c)
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveDefaultsImplicitLast(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName},
|
||||
"mid": {Name: "mid", Parents: []string{defaultsName}},
|
||||
"top": {Name: "top", Parents: []string{"mid"}},
|
||||
}
|
||||
out, err := Resolve(cfgs)
|
||||
if err != nil {
|
||||
t.Fatalf("Resolve: %v", err)
|
||||
}
|
||||
if !eqSlice(out["top"].Chain, []string{"top", "mid", defaultsName}) {
|
||||
t.Errorf("top chain = %v, want [top mid _defaults]", out["top"].Chain)
|
||||
}
|
||||
if !eqSlice(out["mid"].Chain, []string{"mid", defaultsName}) {
|
||||
t.Errorf("mid chain = %v, want [mid _defaults]", out["mid"].Chain)
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveDefaultsDedupExplicit(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName, Env: map[string]string{"D": "1"}},
|
||||
"x": {Name: "x", Parents: []string{defaultsName}},
|
||||
}
|
||||
out, err := Resolve(cfgs)
|
||||
if err != nil {
|
||||
t.Fatalf("Resolve: %v", err)
|
||||
}
|
||||
// _defaults appears exactly once.
|
||||
count := 0
|
||||
for _, c := range out["x"].Chain {
|
||||
if c == defaultsName {
|
||||
count++
|
||||
}
|
||||
}
|
||||
if count != 1 {
|
||||
t.Errorf("_defaults appears %d times in chain %v, want 1", count, out["x"].Chain)
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveMisorderingRejected(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName},
|
||||
"x": {Name: "x"},
|
||||
"y": {Name: "y", Parents: []string{defaultsName, "x"}},
|
||||
}
|
||||
_, err := Resolve(cfgs)
|
||||
if err == nil {
|
||||
t.Fatal("expected misordering error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "must be the only parent") {
|
||||
t.Errorf("error = %q, want misordering message", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveOptOutImpossible(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName, Env: map[string]string{"ROOT": "1"}},
|
||||
"x": {Name: "x", Parents: nil},
|
||||
}
|
||||
out, err := Resolve(cfgs)
|
||||
if err != nil {
|
||||
t.Fatalf("Resolve: %v", err)
|
||||
}
|
||||
r := out["x"]
|
||||
last := r.Chain[len(r.Chain)-1]
|
||||
if last != defaultsName {
|
||||
t.Errorf("last chain element = %q, want %q (opt-out impossible)", last, defaultsName)
|
||||
}
|
||||
if r.Env["ROOT"] != "1" {
|
||||
t.Errorf("env should inherit from _defaults: ROOT=%q", r.Env["ROOT"])
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveCycleDetection(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName},
|
||||
"a": {Name: "a", Parents: []string{"b"}},
|
||||
"b": {Name: "b", Parents: []string{"a"}},
|
||||
}
|
||||
_, err := Resolve(cfgs)
|
||||
if err == nil {
|
||||
t.Fatal("expected cycle error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "cycle") {
|
||||
t.Errorf("error = %q, want cycle message", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveMissingParent(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName},
|
||||
"a": {Name: "a", Parents: []string{"ghost"}},
|
||||
}
|
||||
_, err := Resolve(cfgs)
|
||||
if err == nil {
|
||||
t.Fatal("expected missing-parent error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "ghost") || !strings.Contains(err.Error(), "not found") {
|
||||
t.Errorf("error = %q, want contains 'ghost' and 'not found'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveMissingDefaults(t *testing.T) {
|
||||
cfgs := map[string]*NSConfig{
|
||||
"x": {Name: "x"},
|
||||
}
|
||||
_, err := Resolve(cfgs)
|
||||
if err == nil {
|
||||
t.Fatal("expected missing _defaults error, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), defaultsName) {
|
||||
t.Errorf("error = %q, want contains %q", err.Error(), defaultsName)
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveChainOrderWithDiamond(t *testing.T) {
|
||||
// Diamond: top -> {left, right} -> base; base -> _defaults.
|
||||
cfgs := map[string]*NSConfig{
|
||||
defaultsName: {Name: defaultsName, Env: map[string]string{"R": "r"}},
|
||||
"base": {Name: "base", Parents: []string{defaultsName}, Env: map[string]string{"B": "b"}},
|
||||
"left": {Name: "left", Parents: []string{"base"}, Env: map[string]string{"L": "l"}},
|
||||
"right": {Name: "right", Parents: []string{"base"}, Env: map[string]string{"L": "r"}},
|
||||
"top": {Name: "top", Parents: []string{"left", "right"}, Env: map[string]string{"T": "t"}},
|
||||
}
|
||||
out, err := Resolve(cfgs)
|
||||
if err != nil {
|
||||
t.Fatalf("Resolve: %v", err)
|
||||
}
|
||||
r := out["top"]
|
||||
if r == nil {
|
||||
t.Fatal("missing top")
|
||||
}
|
||||
// top first, _defaults last.
|
||||
if r.Chain[0] != "top" || r.Chain[len(r.Chain)-1] != defaultsName {
|
||||
t.Errorf("chain = %v, want top first and _defaults last", r.Chain)
|
||||
}
|
||||
// base appears exactly once (diamond de-duped).
|
||||
count := 0
|
||||
for _, c := range r.Chain {
|
||||
if c == "base" {
|
||||
count++
|
||||
}
|
||||
}
|
||||
if count != 1 {
|
||||
t.Errorf("base appears %d times in %v, want 1", count, r.Chain)
|
||||
}
|
||||
// top inherits R from _defaults.
|
||||
if r.Env["R"] != "r" {
|
||||
t.Errorf("top should inherit R=r, got %q", r.Env["R"])
|
||||
}
|
||||
}
|
||||
|
||||
func eqSlice(a, b []string) bool {
|
||||
if len(a) != len(b) {
|
||||
return false
|
||||
}
|
||||
for i := range a {
|
||||
if a[i] != b[i] {
|
||||
return false
|
||||
}
|
||||
}
|
||||
return true
|
||||
}
|
||||
@@ -0,0 +1,121 @@
|
||||
// Package paths resolves on-disk locations for the v0.9 multi-namespace
|
||||
// filesystem layout (R-002). It is the canonical source of truth for
|
||||
// cluster-wide, per-namespace, and CLI-cache paths.
|
||||
//
|
||||
// The v0.8 internal/certpaths package is preserved as a thin shim that
|
||||
// returns the legacy flat-layout paths during the v0.9 dual-write window
|
||||
// (REQ-090). New code should use internal/paths, NOT certpaths.
|
||||
package paths
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
)
|
||||
|
||||
const (
|
||||
defaultHomeSubdir = ".orca"
|
||||
clusterDirName = "cluster"
|
||||
defaultNamespace = "_defaults"
|
||||
)
|
||||
|
||||
// Root returns the ORCA home directory. It honors $ORCA_HOME for
|
||||
// testability; otherwise it defaults to ~/.orca. An empty $ORCA_HOME is
|
||||
// treated as unset.
|
||||
func Root() string {
|
||||
if p := os.Getenv("ORCA_HOME"); p != "" {
|
||||
return p
|
||||
}
|
||||
home, _ := os.UserHomeDir()
|
||||
return filepath.Join(home, defaultHomeSubdir)
|
||||
}
|
||||
|
||||
// ClusterDir returns the cluster-wide directory: Root()/cluster.
|
||||
// Cluster-wide artifacts (CA, master key, peers, txns, known_hosts, SSH
|
||||
// keys) live here and are NOT scoped to a workload namespace (R-002).
|
||||
func ClusterDir() string { return filepath.Join(Root(), clusterDirName) }
|
||||
|
||||
// NamespaceDir returns the directory for a namespace: Root()/<ns>.
|
||||
// Use DefaultNamespace() for the implicit root namespace (R-002 D-159).
|
||||
func NamespaceDir(ns string) string { return filepath.Join(Root(), ns) }
|
||||
|
||||
// NSDb returns the SQLite database path for a namespace:
|
||||
// NamespaceDir(ns)/db/orca.db.
|
||||
func NSDb(ns string) string { return filepath.Join(NamespaceDir(ns), "db", "orca.db") }
|
||||
|
||||
// NSEnv returns the .env path for a namespace: NamespaceDir(ns)/.env.
|
||||
func NSEnv(ns string) string { return filepath.Join(NamespaceDir(ns), ".env") }
|
||||
|
||||
// NSSecrets returns the encrypted secrets env path for a namespace:
|
||||
// NamespaceDir(ns)/.env.secrets.
|
||||
func NSSecrets(ns string) string { return filepath.Join(NamespaceDir(ns), ".env.secrets") }
|
||||
|
||||
// NSJobs returns the jobs directory for a namespace: NamespaceDir(ns)/jobs.
|
||||
func NSJobs(ns string) string { return filepath.Join(NamespaceDir(ns), "jobs") }
|
||||
|
||||
// NSAlloc returns the allocation directory for a namespace:
|
||||
// NamespaceDir(ns)/alloc.
|
||||
func NSAlloc(ns string) string { return filepath.Join(NamespaceDir(ns), "alloc") }
|
||||
|
||||
// NSMd returns the namespace Markdown doc path (R-014):
|
||||
// NamespaceDir(ns)/ns.md.
|
||||
func NSMd(ns string) string { return filepath.Join(NamespaceDir(ns), "ns.md") }
|
||||
|
||||
// DefaultNamespace returns the implicit root namespace name (R-002 D-159).
|
||||
func DefaultNamespace() string { return defaultNamespace }
|
||||
|
||||
// CACertPath returns the v0.9 cluster CA cert path:
|
||||
// ClusterDir()/ca.crt (D-101). The v0.8 internal CA still writes to
|
||||
// Root()/ca.crt; the move happens in v0.10-P14.
|
||||
func CACertPath() string { return filepath.Join(ClusterDir(), "ca.crt") }
|
||||
|
||||
// CAKeyPath returns the v0.9 cluster CA key path:
|
||||
// ClusterDir()/ca.key.
|
||||
func CAKeyPath() string { return filepath.Join(ClusterDir(), "ca.key") }
|
||||
|
||||
// MasterKeyPath returns the AES-256-GCM root master key path
|
||||
// (R-011, mode 0600): ClusterDir()/master.key. Not generated until
|
||||
// v0.10-P03.
|
||||
func MasterKeyPath() string { return filepath.Join(ClusterDir(), "master.key") }
|
||||
|
||||
// CacheDB returns the CLI-side cache database path (R-008):
|
||||
// Root()/orca_cache.db. Not created until v0.9-P0a2.
|
||||
func CacheDB() string { return filepath.Join(Root(), "orca_cache.db") }
|
||||
|
||||
// TxnDir returns the cluster transaction log directory (R-016):
|
||||
// ClusterDir()/txns.
|
||||
func TxnDir() string { return filepath.Join(ClusterDir(), "txns") }
|
||||
|
||||
// PeersDir returns the cluster peers directory: ClusterDir()/peers.
|
||||
func PeersDir() string { return filepath.Join(ClusterDir(), "peers") }
|
||||
|
||||
// PeerDir returns the directory for a single peer host:
|
||||
// PeersDir()/host.
|
||||
func PeerDir(host string) string { return filepath.Join(PeersDir(), host) }
|
||||
|
||||
// KnownHostsPath returns the SSH known_hosts path (D-035):
|
||||
// ClusterDir()/known_hosts.
|
||||
func KnownHostsPath() string { return filepath.Join(ClusterDir(), "known_hosts") }
|
||||
|
||||
// SSHKeyPath returns the orca SSH private key path:
|
||||
// ClusterDir()/orca_ssh_key (D-037).
|
||||
func SSHKeyPath() string { return filepath.Join(ClusterDir(), "orca_ssh_key") }
|
||||
|
||||
// SSHPubPath returns the orca SSH public key path:
|
||||
// ClusterDir()/orca_ssh_key.pub.
|
||||
func SSHPubPath() string { return filepath.Join(ClusterDir(), "orca_ssh_key.pub") }
|
||||
|
||||
// ServerCertPath returns the legacy server cert path (legacy compat):
|
||||
// ClusterDir()/server.crt. step-ca will replace this in a later phase.
|
||||
func ServerCertPath() string { return filepath.Join(ClusterDir(), "server.crt") }
|
||||
|
||||
// ServerKeyPath returns the legacy server key path (legacy compat):
|
||||
// ClusterDir()/server.key. step-ca will replace this in a later phase.
|
||||
func ServerKeyPath() string { return filepath.Join(ClusterDir(), "server.key") }
|
||||
|
||||
// ConfigPath returns the new Markdown-frontmatter config path (R-014):
|
||||
// ClusterDir()/config.md. The legacy HCL path is ClusterDir()/config.hcl.
|
||||
func ConfigPath() string { return filepath.Join(ClusterDir(), "config.md") }
|
||||
|
||||
// LegacyHCLConfigPath returns the legacy HCL config path:
|
||||
// ClusterDir()/config.hcl.
|
||||
func LegacyHCLConfigPath() string { return filepath.Join(ClusterDir(), "config.hcl") }
|
||||
@@ -0,0 +1,209 @@
|
||||
package paths
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestRoot_HonorsORCAHOME(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
if got, want := Root(), dir; got != want {
|
||||
t.Errorf("Root() = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRoot_EmptyORCAHOMEFallsBack(t *testing.T) {
|
||||
t.Setenv("ORCA_HOME", "")
|
||||
home, err := os.UserHomeDir()
|
||||
if err != nil {
|
||||
t.Skipf("os.UserHomeDir: %v", err)
|
||||
}
|
||||
want := filepath.Join(home, defaultHomeSubdir)
|
||||
if got := Root(); got != want {
|
||||
t.Errorf("Root() with empty ORCA_HOME = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRoot_UnsetORCAHOMEFallsBack(t *testing.T) {
|
||||
os.Unsetenv("ORCA_HOME")
|
||||
home, err := os.UserHomeDir()
|
||||
if err != nil {
|
||||
t.Skipf("os.UserHomeDir: %v", err)
|
||||
}
|
||||
want := filepath.Join(home, defaultHomeSubdir)
|
||||
got := Root()
|
||||
if got != want {
|
||||
t.Errorf("Root() default = %q, want %q", got, want)
|
||||
}
|
||||
if !strings.HasPrefix(got, home) {
|
||||
t.Errorf("Root() default %q does not start with home %q", got, home)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRoot_RelativeORCAHOME(t *testing.T) {
|
||||
t.Setenv("ORCA_HOME", "relative/orca/home")
|
||||
if got, want := Root(), "relative/orca/home"; got != want {
|
||||
t.Errorf("Root() relative = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDefaultNamespace(t *testing.T) {
|
||||
if got, want := DefaultNamespace(), "_defaults"; got != want {
|
||||
t.Errorf("DefaultNamespace() = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestClusterDir(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
got := ClusterDir()
|
||||
want := filepath.Join(dir, "cluster")
|
||||
if got != want {
|
||||
t.Errorf("ClusterDir() = %q, want %q", got, want)
|
||||
}
|
||||
if !strings.HasPrefix(got, Root()+string(filepath.Separator)) {
|
||||
t.Errorf("ClusterDir() %q not under Root() %q", got, Root())
|
||||
}
|
||||
}
|
||||
|
||||
func TestNamespaceDir(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
ns := "prod"
|
||||
got := NamespaceDir(ns)
|
||||
want := filepath.Join(dir, ns)
|
||||
if got != want {
|
||||
t.Errorf("NamespaceDir(%q) = %q, want %q", ns, got, want)
|
||||
}
|
||||
if !strings.HasPrefix(got, Root()+string(filepath.Separator)) {
|
||||
t.Errorf("NamespaceDir() %q not under Root() %q", got, Root())
|
||||
}
|
||||
}
|
||||
|
||||
func TestNamespacePaths(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
ns := "prod"
|
||||
nsDir := NamespaceDir(ns)
|
||||
|
||||
cases := []struct {
|
||||
name string
|
||||
got string
|
||||
want string
|
||||
}{
|
||||
{"NSDb", NSDb(ns), filepath.Join(nsDir, "db", "orca.db")},
|
||||
{"NSEnv", NSEnv(ns), filepath.Join(nsDir, ".env")},
|
||||
{"NSSecrets", NSSecrets(ns), filepath.Join(nsDir, ".env.secrets")},
|
||||
{"NSJobs", NSJobs(ns), filepath.Join(nsDir, "jobs")},
|
||||
{"NSAlloc", NSAlloc(ns), filepath.Join(nsDir, "alloc")},
|
||||
{"NSMd", NSMd(ns), filepath.Join(nsDir, "ns.md")},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
if tc.got != tc.want {
|
||||
t.Errorf("%s(%q) = %q, want %q", tc.name, ns, tc.got, tc.want)
|
||||
}
|
||||
if !strings.HasPrefix(tc.got, nsDir+string(filepath.Separator)) {
|
||||
t.Errorf("%s() %q not under NamespaceDir() %q", tc.name, tc.got, nsDir)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestDefaultNamespacePaths(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
ns := DefaultNamespace()
|
||||
nsDir := NamespaceDir(ns)
|
||||
|
||||
if got, want := NSDb(ns), filepath.Join(nsDir, "db", "orca.db"); got != want {
|
||||
t.Errorf("NSDb(_defaults) = %q, want %q", got, want)
|
||||
}
|
||||
if got, want := NSEnv(ns), filepath.Join(nsDir, ".env"); got != want {
|
||||
t.Errorf("NSEnv(_defaults) = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestClusterPaths(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
cDir := ClusterDir()
|
||||
|
||||
cases := []struct {
|
||||
name string
|
||||
got string
|
||||
want string
|
||||
}{
|
||||
{"CACertPath", CACertPath(), filepath.Join(cDir, "ca.crt")},
|
||||
{"CAKeyPath", CAKeyPath(), filepath.Join(cDir, "ca.key")},
|
||||
{"MasterKeyPath", MasterKeyPath(), filepath.Join(cDir, "master.key")},
|
||||
{"KnownHostsPath", KnownHostsPath(), filepath.Join(cDir, "known_hosts")},
|
||||
{"SSHKeyPath", SSHKeyPath(), filepath.Join(cDir, "orca_ssh_key")},
|
||||
{"SSHPubPath", SSHPubPath(), filepath.Join(cDir, "orca_ssh_key.pub")},
|
||||
{"ServerCertPath", ServerCertPath(), filepath.Join(cDir, "server.crt")},
|
||||
{"ServerKeyPath", ServerKeyPath(), filepath.Join(cDir, "server.key")},
|
||||
{"ConfigPath", ConfigPath(), filepath.Join(cDir, "config.md")},
|
||||
{"LegacyHCLConfigPath", LegacyHCLConfigPath(), filepath.Join(cDir, "config.hcl")},
|
||||
{"TxnDir", TxnDir(), filepath.Join(cDir, "txns")},
|
||||
{"PeersDir", PeersDir(), filepath.Join(cDir, "peers")},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
if tc.got != tc.want {
|
||||
t.Errorf("%s() = %q, want %q", tc.name, tc.got, tc.want)
|
||||
}
|
||||
if !strings.HasPrefix(tc.got, cDir+string(filepath.Separator)) {
|
||||
t.Errorf("%s() %q not under ClusterDir() %q", tc.name, tc.got, cDir)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestPeerDir(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
host := "node1.example.com"
|
||||
got := PeerDir(host)
|
||||
want := filepath.Join(PeersDir(), host)
|
||||
if got != want {
|
||||
t.Errorf("PeerDir(%q) = %q, want %q", host, got, want)
|
||||
}
|
||||
if !strings.HasPrefix(got, PeersDir()+string(filepath.Separator)) {
|
||||
t.Errorf("PeerDir() %q not under PeersDir() %q", got, PeersDir())
|
||||
}
|
||||
}
|
||||
|
||||
func TestCacheDB(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
got := CacheDB()
|
||||
want := filepath.Join(dir, "orca_cache.db")
|
||||
if got != want {
|
||||
t.Errorf("CacheDB() = %q, want %q", got, want)
|
||||
}
|
||||
if !strings.HasPrefix(got, Root()+string(filepath.Separator)) {
|
||||
t.Errorf("CacheDB() %q not under Root() %q", got, Root())
|
||||
}
|
||||
}
|
||||
|
||||
func TestPathSeparatorsOSAppropriate(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
// Every returned path must use the OS separator (filepath.Join).
|
||||
sep := string(filepath.Separator)
|
||||
for _, p := range []string{
|
||||
ClusterDir(), NamespaceDir("ns"), NSDb("ns"), NSEnv("ns"),
|
||||
NSSecrets("ns"), NSJobs("ns"), NSAlloc("ns"), NSMd("ns"),
|
||||
CACertPath(), CAKeyPath(), MasterKeyPath(), CacheDB(),
|
||||
TxnDir(), PeersDir(), PeerDir("h"), KnownHostsPath(),
|
||||
SSHKeyPath(), SSHPubPath(), ServerCertPath(), ServerKeyPath(),
|
||||
ConfigPath(), LegacyHCLConfigPath(),
|
||||
} {
|
||||
if !strings.Contains(p, sep) {
|
||||
t.Errorf("path %q lacks OS separator %q (not joined?)", p, sep)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -289,6 +289,11 @@ func TOFUHostKeyCallback(addr string, capturedKey *ssh.PublicKey) (ssh.HostKeyCa
|
||||
if errors.As(err, &keyErr) && len(keyErr.Want) == 0 {
|
||||
line := knownhosts.Line([]string{knownhosts.Normalize(addr)}, key)
|
||||
path := certpaths.KnownHostsPath()
|
||||
release, lockErr := security.Flock(path)
|
||||
if lockErr != nil {
|
||||
return fmt.Errorf("tofu lock known_hosts: %w", lockErr)
|
||||
}
|
||||
defer release()
|
||||
existing, readErr := os.ReadFile(path)
|
||||
if readErr != nil && !os.IsNotExist(readErr) {
|
||||
return fmt.Errorf("tofu read known_hosts: %w", readErr)
|
||||
@@ -481,6 +486,11 @@ func ResetHostKey(host string) error {
|
||||
return fmt.Errorf("ResetHostKey: host is required")
|
||||
}
|
||||
path := certpaths.KnownHostsPath()
|
||||
release, lockErr := security.Flock(path)
|
||||
if lockErr != nil {
|
||||
return fmt.Errorf("ResetHostKey: lock known_hosts: %w", lockErr)
|
||||
}
|
||||
defer release()
|
||||
existing, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
if os.IsNotExist(err) {
|
||||
|
||||
@@ -0,0 +1,43 @@
|
||||
# C-01 Grill Gate Evaluation — wasmtime / CGO
|
||||
|
||||
**Gate**: C-01 (wasmtime/CGO evaluation), gating P07b.
|
||||
|
||||
**Question**: The wasmtime Go binding
|
||||
(`github.com/bytecodealliance/wasmtime-go`) is CGO-based. Does adopting
|
||||
it revoke D-002 (modernc/sqlite CGO-free cross-compile story)?
|
||||
|
||||
## Evaluation
|
||||
|
||||
| Option | CGO required? | Cross-compile impact | Decision |
|
||||
|--------|---------------|----------------------|----------|
|
||||
| A. Use `bytecodealliance/wasmtime-go` (Go binding) | **YES** — the binding links libwasmtime via cgo | Revokes D-002 — Go cross-compile (`GOOS=linux GOARCH=arm64 go build`) breaks; CGO toolchain needed on every build host; static-binary story lost | REJECTED |
|
||||
| B. Use the `wasmtime` CLI (apt-installed on the peer) via SSH exec | **NO** — pure Go code, shells out to a CLI over SSH (same pattern as podman/qm/pct) | None — D-002 preserved | **ACCEPTED** |
|
||||
| C. Use an alternative pure-Go WASM runtime (e.g. wazero) | No CGO | Pure-Go alternative exists; but wazero's wasmtime-compat is incomplete (component model, WASI 0.2); different runtime semantics than the "wasmtime" operator surface promised in D-088 | DEFERRED (v0.10 evaluation if CLI-via-SSH proves insufficient) |
|
||||
|
||||
## Decision (full autonomy, auto-decision)
|
||||
|
||||
**Option B**: `WasmRuntime` uses the `wasmtime` CLI (apt-installed on
|
||||
the peer) via the SSH-push transport. It does NOT import
|
||||
`bytecodealliance/wasmtime-go` (or any other CGO package).
|
||||
|
||||
## C-01 Gate Status
|
||||
|
||||
**SATISFIED.** C-01 is satisfied:
|
||||
- wasmtime works without CGO (CLI-via-SSH pattern, identical to podman/qm/pct).
|
||||
- D-002 cross-compile story preserved (no CGO introduced anywhere in
|
||||
the runtime package or any orca Go code).
|
||||
- D-002 is NOT revoked.
|
||||
|
||||
## Recorded as D-187
|
||||
|
||||
See PROJECT.md v0.9 D-series: "wasmtime Go binding is CGO-based
|
||||
(bytecodealliance/wasmtime-go); orca uses the wasmtime CLI via SSH
|
||||
(apt-installed on peer) instead of the Go binding, avoiding CGO
|
||||
entirely. D-002 cross-compile story preserved. C-01 satisfied."
|
||||
|
||||
## Verification
|
||||
|
||||
- `go build ./internal/runtime/` succeeds with `CGO_ENABLED=0`.
|
||||
- `internal/runtime/wasm.go` imports only stdlib + sshpush (no
|
||||
wasmtime-go).
|
||||
- The full test suite (`go test ./...`) does not require CGO.
|
||||
@@ -0,0 +1,142 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/sshpush"
|
||||
)
|
||||
|
||||
// PodmanRuntime implements Runtime for the "podman" one_of. It runs
|
||||
// `podman` on the peer over the SSH-push transport (P01). The
|
||||
// transport is injected via the constructor (dependency injection).
|
||||
//
|
||||
// Container lifecycle:
|
||||
//
|
||||
// - Prepare: `podman pull <image>`
|
||||
// - Start: `podman run -d --name orca-<alloc-id> <image> <command>`
|
||||
// - Stop: `podman stop <name>` then `podman rm <name>`
|
||||
// - Status: `podman inspect --format '{{.State.Running}}' <name>`
|
||||
//
|
||||
// The container name is `orca-<alloc-id>` (sanitized to lowercase +
|
||||
// alnum). The runtime keeps no in-process state — each call is a fresh
|
||||
// SSH exec against the peer.
|
||||
type PodmanRuntime struct {
|
||||
transport *sshpush.Transport
|
||||
}
|
||||
|
||||
// NewPodmanRuntime returns a PodmanRuntime backed by the given transport.
|
||||
func NewPodmanRuntime(t *sshpush.Transport) *PodmanRuntime {
|
||||
return &PodmanRuntime{transport: t}
|
||||
}
|
||||
|
||||
// containerName returns the deterministic container name for an alloc.
|
||||
func containerName(alloc *Alloc) string {
|
||||
id := strings.ToLower(alloc.ID)
|
||||
id = strings.Map(func(r rune) rune {
|
||||
if r >= 'a' && r <= 'z' || r >= '0' && r <= '9' || r == '-' || r == '_' {
|
||||
return r
|
||||
}
|
||||
return '-'
|
||||
}, id)
|
||||
return "orca-" + id
|
||||
}
|
||||
|
||||
// Prepare pulls the image on the peer.
|
||||
func (p *PodmanRuntime) Prepare(ctx context.Context, alloc *Alloc) error {
|
||||
image, err := imageFor(alloc)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
cmd := fmt.Sprintf("podman pull %q", image)
|
||||
if _, err := p.transport.Exec(ctx, alloc.Node, cmd); err != nil {
|
||||
return fmt.Errorf("podman: pull: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Start runs `podman run -d --name <name> <image> <command>`.
|
||||
func (p *PodmanRuntime) Start(ctx context.Context, alloc *Alloc) (int, error) {
|
||||
image, err := imageFor(alloc)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
cmdStr, _ := commandFor(alloc)
|
||||
name := containerName(alloc)
|
||||
cmd := fmt.Sprintf("podman run -d --name %s %q %s", name, image, cmdStr)
|
||||
out, err := p.transport.Exec(ctx, alloc.Node, cmd)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("podman: run: %w", err)
|
||||
}
|
||||
// The container ID is the first 12 chars of the printed hash. We
|
||||
// don't keep it — Stop/Status use the name — but return a stable
|
||||
// synthetic PID derived from the first 4 bytes of the hash for
|
||||
// the interface contract.
|
||||
cid := strings.TrimSpace(string(out))
|
||||
return podmanCidToPID(cid), nil
|
||||
}
|
||||
|
||||
// Stop stops and removes the container.
|
||||
func (p *PodmanRuntime) Stop(ctx context.Context, alloc *Alloc) error {
|
||||
name := containerName(alloc)
|
||||
if _, err := p.transport.Exec(ctx, alloc.Node, fmt.Sprintf("podman stop %s", name)); err != nil {
|
||||
return fmt.Errorf("podman: stop: %w", err)
|
||||
}
|
||||
if _, err := p.transport.Exec(ctx, alloc.Node, fmt.Sprintf("podman rm %s", name)); err != nil {
|
||||
return fmt.Errorf("podman: rm: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Status inspects the container's running state.
|
||||
func (p *PodmanRuntime) Status(ctx context.Context, alloc *Alloc) (State, error) {
|
||||
name := containerName(alloc)
|
||||
cmd := fmt.Sprintf("podman inspect --format '{{.State.Running}}' %s", name)
|
||||
out, err := p.transport.Exec(ctx, alloc.Node, cmd)
|
||||
if err != nil {
|
||||
return StateFailed, fmt.Errorf("podman: inspect: %w", err)
|
||||
}
|
||||
v := strings.TrimSpace(string(out))
|
||||
switch v {
|
||||
case "true":
|
||||
return StateRunning, nil
|
||||
case "false":
|
||||
return StateStopped, nil
|
||||
default:
|
||||
return StateFailed, fmt.Errorf("podman: unexpected inspect output %q", v)
|
||||
}
|
||||
}
|
||||
|
||||
// podmanCidToPID converts a container ID (hex hash) to a positive int
|
||||
// PID for the interface contract. It reads up to 4 hex chars.
|
||||
func podmanCidToPID(cid string) int {
|
||||
if len(cid) < 1 {
|
||||
return 1
|
||||
}
|
||||
n := len(cid)
|
||||
if n > 4 {
|
||||
n = 4
|
||||
}
|
||||
var pid int
|
||||
for i := 0; i < n; i++ {
|
||||
c := cid[i]
|
||||
pid = (pid << 4) | int(hexVal(c))
|
||||
}
|
||||
if pid <= 0 {
|
||||
pid = 1
|
||||
}
|
||||
return pid
|
||||
}
|
||||
|
||||
func hexVal(c byte) byte {
|
||||
switch {
|
||||
case c >= '0' && c <= '9':
|
||||
return c - '0'
|
||||
case c >= 'a' && c <= 'f':
|
||||
return c - 'a' + 10
|
||||
case c >= 'A' && c <= 'F':
|
||||
return c - 'A' + 10
|
||||
}
|
||||
return 0
|
||||
}
|
||||
@@ -0,0 +1,229 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
"git.cloudinit.dev/coreci/orca/internal/sshpush"
|
||||
)
|
||||
|
||||
// allocNoImage returns an alloc whose Spec has a Runtime block with a
|
||||
// command but no image.
|
||||
func allocNoImage(runtime string) *Alloc {
|
||||
return &Alloc{
|
||||
ID: "x",
|
||||
Runtime: runtime,
|
||||
Spec: &jobspec.WorkloadSpec{
|
||||
Runtime: &jobspec.RuntimeBlock{Command: "/bin/true"},
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanRuntime_HappyPath wires a fake server that responds to
|
||||
// podman pull/run/stop/rm/inspect and verifies the full lifecycle.
|
||||
func TestPodmanRuntime_HappyPath(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
|
||||
const cid = "abc123def456"
|
||||
srv.setHandler("podman pull", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("podman run", func(cmd string) ([]byte, int) { return []byte(cid + "\n"), 0 })
|
||||
srv.setHandler("podman stop", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("podman rm", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("podman inspect", func(cmd string) ([]byte, int) {
|
||||
return []byte("true\n"), 0
|
||||
})
|
||||
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
|
||||
p := NewPodmanRuntime(tr)
|
||||
a := allocWithNode("podman", "docker.io/library/alpine:latest", "sleep 30", srv.addr())
|
||||
|
||||
ctx, cancel := withTimeout(10 * time.Second)
|
||||
defer cancel()
|
||||
if err := p.Prepare(ctx, a); err != nil {
|
||||
t.Fatalf("Prepare: %v", err)
|
||||
}
|
||||
pid, err := p.Start(ctx, a)
|
||||
if err != nil {
|
||||
t.Fatalf("Start: %v", err)
|
||||
}
|
||||
if pid <= 0 {
|
||||
t.Fatalf("pid = %d, want > 0", pid)
|
||||
}
|
||||
st, err := p.Status(ctx, a)
|
||||
if err != nil {
|
||||
t.Fatalf("Status: %v", err)
|
||||
}
|
||||
if st != StateRunning {
|
||||
t.Errorf("Status = %q, want running", st)
|
||||
}
|
||||
if err := p.Stop(ctx, a); err != nil {
|
||||
t.Fatalf("Stop: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanRuntime_StatusFalse verifies Status returns stopped when
|
||||
// the container reports running=false.
|
||||
func TestPodmanRuntime_StatusFalse(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("podman inspect", func(cmd string) ([]byte, int) {
|
||||
return []byte("false\n"), 0
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
p := NewPodmanRuntime(tr)
|
||||
a := allocWithNode("podman", "img", "sleep 1", srv.addr())
|
||||
st, err := p.Status(context.Background(), a)
|
||||
if err != nil {
|
||||
t.Fatalf("Status: %v", err)
|
||||
}
|
||||
if st != StateStopped {
|
||||
t.Errorf("Status = %q, want stopped", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanRuntime_StatusBadOutput verifies Status returns failed on
|
||||
// unexpected inspect output.
|
||||
func TestPodmanRuntime_StatusBadOutput(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("podman inspect", func(cmd string) ([]byte, int) {
|
||||
return []byte("garbage\n"), 0
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
p := NewPodmanRuntime(tr)
|
||||
a := allocWithNode("podman", "img", "sleep 1", srv.addr())
|
||||
if _, err := p.Status(context.Background(), a); err == nil {
|
||||
t.Error("Status with bad output should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanRuntime_PrepareNoImage verifies Prepare errors when the
|
||||
// alloc has no image.
|
||||
func TestPodmanRuntime_PrepareNoImage(t *testing.T) {
|
||||
p := NewPodmanRuntime(nil)
|
||||
a := allocNoImage("podman")
|
||||
if err := p.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare with no image should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanRuntime_PrepareTransportError verifies Prepare propagates a
|
||||
// transport error (podman pull fails).
|
||||
func TestPodmanRuntime_PrepareTransportError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("podman pull", func(cmd string) ([]byte, int) {
|
||||
return []byte("manifest unknown\n"), 2
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
p := NewPodmanRuntime(tr)
|
||||
a := allocWithNode("podman", "img", "sleep 1", srv.addr())
|
||||
if err := p.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare with failed pull should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanRuntime_StartNoImage verifies Start errors with no image.
|
||||
func TestPodmanRuntime_StartNoImage(t *testing.T) {
|
||||
p := NewPodmanRuntime(nil)
|
||||
a := allocNoImage("podman")
|
||||
if _, err := p.Start(context.Background(), a); err == nil {
|
||||
t.Error("Start with no image should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanRuntime_StopTransportError verifies Stop propagates errors.
|
||||
func TestPodmanRuntime_StopTransportError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("podman stop", func(cmd string) ([]byte, int) {
|
||||
return []byte("no such container\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
p := NewPodmanRuntime(tr)
|
||||
a := allocWithNode("podman", "img", "sleep 1", srv.addr())
|
||||
if err := p.Stop(context.Background(), a); err == nil {
|
||||
t.Error("Stop with missing container should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanRuntime_StatusInspectError verifies Status returns failed
|
||||
// when inspect itself errors.
|
||||
func TestPodmanRuntime_StatusInspectError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("podman inspect", func(cmd string) ([]byte, int) {
|
||||
return []byte("no such container\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
p := NewPodmanRuntime(tr)
|
||||
a := allocWithNode("podman", "img", "sleep 1", srv.addr())
|
||||
st, err := p.Status(context.Background(), a)
|
||||
if err == nil {
|
||||
t.Error("Status with inspect error should error")
|
||||
}
|
||||
if st != StateFailed {
|
||||
t.Errorf("Status = %q, want failed", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanCidToPID verifies the synthetic PID derivation.
|
||||
func TestPodmanCidToPID(t *testing.T) {
|
||||
if got := podmanCidToPID(""); got != 1 {
|
||||
t.Errorf("empty cid -> %d, want 1", got)
|
||||
}
|
||||
if got := podmanCidToPID("a"); got <= 0 {
|
||||
t.Errorf("single hex -> %d, want > 0", got)
|
||||
}
|
||||
if got := podmanCidToPID("abcd"); got <= 0 {
|
||||
t.Errorf("abcd -> %d, want > 0", got)
|
||||
}
|
||||
// non-hex chars fall through to 0 contributions but still yield
|
||||
// a positive result (>= 1 by the floor).
|
||||
if got := podmanCidToPID("xyz123"); got <= 0 {
|
||||
t.Errorf("xyz123 -> %d, want > 0", got)
|
||||
}
|
||||
}
|
||||
|
||||
// TestContainerNameSanitization verifies the container-name sanitizer
|
||||
// uppercases and strips disallowed characters.
|
||||
func TestContainerNameSanitization(t *testing.T) {
|
||||
a := &Alloc{ID: "ALLOC_1.2.3", Spec: nil, Runtime: "podman"}
|
||||
got := containerName(a)
|
||||
if !strings.HasPrefix(got, "orca-") {
|
||||
t.Errorf("containerName = %q, want orca- prefix", got)
|
||||
}
|
||||
if strings.Contains(got, ".") {
|
||||
t.Errorf("containerName = %q, should not contain '.'", got)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPodmanRuntime_DialError verifies Prepare fails fast when the
|
||||
// peer is unreachable (no fake server).
|
||||
func TestPodmanRuntime_DialError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
// close immediately so dial fails.
|
||||
srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
p := NewPodmanRuntime(tr)
|
||||
a := allocWithNode("podman", "img", "sleep 1", srv.addr())
|
||||
if err := p.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare against dead peer should error")
|
||||
} else if !errors.Is(err, sshpush.ErrTransient) && !errors.Is(err, sshpush.ErrPermanent) {
|
||||
// acceptable: either transient (retry exhausted) or permanent.
|
||||
t.Logf("Prepare err (acceptable): %v", err)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,204 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"os"
|
||||
"os/exec"
|
||||
"strconv"
|
||||
"strings"
|
||||
"sync"
|
||||
"syscall"
|
||||
"time"
|
||||
)
|
||||
|
||||
// StopGrace is the default grace period between SIGTERM and SIGKILL
|
||||
// for ProcessRuntime.Stop (10s, matching the systemd default TimeoutStopSec).
|
||||
const StopGrace = 10 * time.Second
|
||||
|
||||
// ProcessRuntime implements Runtime for the "process" one_of using
|
||||
// os/exec. It is the in-process equivalent of the systemd unit the CLI
|
||||
// emits in production: the CLI emits a .service file and the peer's
|
||||
// systemd runs the process; ProcessRuntime starts the process directly
|
||||
// in the current Go process. It is intended for LOCAL testing and
|
||||
// hermetic CI — NOT for production (production uses the systemd emitter
|
||||
// + the peer's systemd, not this in-process path).
|
||||
//
|
||||
// It tracks started PIDs in an in-memory map; restarts of the CLI lose
|
||||
// that state (acceptable for the local-test use case).
|
||||
type ProcessRuntime struct {
|
||||
mu sync.Mutex
|
||||
pids map[string]int // alloc.ID -> PID
|
||||
procs map[int]*os.Process // PID -> process handle
|
||||
stopped map[string]bool // alloc.ID -> reported stopped after Stop
|
||||
}
|
||||
|
||||
// NewProcessRuntime returns a ProcessRuntime.
|
||||
func NewProcessRuntime() *ProcessRuntime {
|
||||
return &ProcessRuntime{
|
||||
pids: make(map[string]int),
|
||||
procs: make(map[int]*os.Process),
|
||||
stopped: make(map[string]bool),
|
||||
}
|
||||
}
|
||||
|
||||
// Prepare is a no-op for the process runtime: the systemd unit is
|
||||
// emitted by the systemd emitter (internal/emitter), not by the runtime.
|
||||
func (p *ProcessRuntime) Prepare(ctx context.Context, alloc *Alloc) error {
|
||||
_ = ctx
|
||||
_ = alloc
|
||||
return nil
|
||||
}
|
||||
|
||||
// Start execs the alloc's command and returns the PID. The process is
|
||||
// left running in the background; Stop terminates it.
|
||||
func (p *ProcessRuntime) Start(ctx context.Context, alloc *Alloc) (int, error) {
|
||||
cmdStr, err := commandFor(alloc)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
// Parse the command string into argv. A leading "exec" form
|
||||
// (shell-style) is NOT supported — the command must be a direct
|
||||
// argv[0] + args. Split on whitespace (simple, matches the existing
|
||||
// executor.go behaviour which takes Command + Args separately).
|
||||
parts := strings.Fields(cmdStr)
|
||||
if len(parts) == 0 {
|
||||
return 0, fmt.Errorf("process: empty command for alloc %s", alloc.ID)
|
||||
}
|
||||
|
||||
// Use a detached context so the process survives the request
|
||||
// context cancellation (the request ends; the workload keeps
|
||||
// running until Stop). We apply our own timeout for Start only.
|
||||
startCtx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
|
||||
defer cancel()
|
||||
|
||||
cmd := exec.CommandContext(startCtx, parts[0], parts[1:]...)
|
||||
// Detach the child from the parent's process group so it survives.
|
||||
cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}
|
||||
// Discard output for the runtime; the systemd unit captures logs
|
||||
// in production. For tests, callers that need output run their
|
||||
// own exec.Command.
|
||||
cmd.Stdout = os.Stdout
|
||||
cmd.Stderr = os.Stderr
|
||||
if err := cmd.Start(); err != nil {
|
||||
return 0, fmt.Errorf("process: start: %w", err)
|
||||
}
|
||||
pid := cmd.Process.Pid
|
||||
|
||||
// Background-reap the process so it doesn't become a zombie; we
|
||||
// only need the PID for Stop/Status. When the process exits
|
||||
// naturally, mark the alloc as stopped.
|
||||
go func() {
|
||||
_ = cmd.Wait()
|
||||
p.mu.Lock()
|
||||
delete(p.procs, pid)
|
||||
// Only mark stopped if the alloc is still associated with
|
||||
// this PID (Stop may have already removed the mapping).
|
||||
if cur, ok := p.pids[alloc.ID]; ok && cur == pid {
|
||||
delete(p.pids, alloc.ID)
|
||||
p.stopped[alloc.ID] = true
|
||||
}
|
||||
p.mu.Unlock()
|
||||
}()
|
||||
|
||||
p.mu.Lock()
|
||||
p.pids[alloc.ID] = pid
|
||||
p.procs[pid] = cmd.Process
|
||||
delete(p.stopped, alloc.ID)
|
||||
p.mu.Unlock()
|
||||
|
||||
return pid, nil
|
||||
}
|
||||
|
||||
// Stop sends SIGTERM, waits the grace period, then SIGKILL.
|
||||
func (p *ProcessRuntime) Stop(ctx context.Context, alloc *Alloc) error {
|
||||
p.mu.Lock()
|
||||
pid, ok := p.pids[alloc.ID]
|
||||
proc := p.procs[pid]
|
||||
p.mu.Unlock()
|
||||
if !ok || proc == nil {
|
||||
return nil // not running; idempotent
|
||||
}
|
||||
// SIGTERM the process group (negative PID).
|
||||
_ = syscall.Kill(-pid, syscall.SIGTERM)
|
||||
|
||||
grace := StopGrace
|
||||
if dl, ok := ctx.Deadline(); ok {
|
||||
if remaining := time.Until(dl); remaining > 0 && remaining < grace {
|
||||
grace = remaining
|
||||
}
|
||||
}
|
||||
deadline := time.Now().Add(grace)
|
||||
for time.Now().Before(deadline) {
|
||||
if !p.alive(pid) {
|
||||
p.forget(alloc.ID, pid)
|
||||
return nil
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
p.forget(alloc.ID, pid)
|
||||
return ctx.Err()
|
||||
case <-time.After(100 * time.Millisecond):
|
||||
}
|
||||
}
|
||||
// SIGKILL the group.
|
||||
_ = syscall.Kill(-pid, syscall.SIGKILL)
|
||||
p.forget(alloc.ID, pid)
|
||||
return nil
|
||||
}
|
||||
|
||||
// Status reports the alloc's state by checking if the process is alive.
|
||||
func (p *ProcessRuntime) Status(ctx context.Context, alloc *Alloc) (State, error) {
|
||||
_ = ctx
|
||||
p.mu.Lock()
|
||||
pid, ok := p.pids[alloc.ID]
|
||||
wasStopped := p.stopped[alloc.ID]
|
||||
p.mu.Unlock()
|
||||
if !ok {
|
||||
if wasStopped {
|
||||
return StateStopped, nil
|
||||
}
|
||||
return StatePending, nil
|
||||
}
|
||||
if !p.alive(pid) {
|
||||
// Process exited but the reaper hasn't run yet; mark it
|
||||
// stopped and clean up.
|
||||
p.forget(alloc.ID, pid)
|
||||
return StateStopped, nil
|
||||
}
|
||||
return StateRunning, nil
|
||||
}
|
||||
|
||||
// alive reports whether the process with the given PID is still running.
|
||||
func (p *ProcessRuntime) alive(pid int) bool {
|
||||
proc, err := os.FindProcess(pid)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
if err := proc.Signal(syscall.Signal(0)); err != nil {
|
||||
// ESRCH means the process is gone.
|
||||
return false
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// forget removes the alloc/PID mapping and marks the alloc stopped.
|
||||
func (p *ProcessRuntime) forget(allocID string, pid int) {
|
||||
p.mu.Lock()
|
||||
delete(p.pids, allocID)
|
||||
delete(p.procs, pid)
|
||||
p.stopped[allocID] = true
|
||||
p.mu.Unlock()
|
||||
}
|
||||
|
||||
// PID returns the recorded PID for alloc (for tests/inspection).
|
||||
func (p *ProcessRuntime) PID(allocID string) (int, error) {
|
||||
p.mu.Lock()
|
||||
defer p.mu.Unlock()
|
||||
pid, ok := p.pids[allocID]
|
||||
if !ok {
|
||||
return 0, errors.New("process: no pid for alloc " + strconv.Quote(allocID))
|
||||
}
|
||||
return pid, nil
|
||||
}
|
||||
@@ -0,0 +1,172 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"runtime"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// TestProcessRuntime_PrepareIsNoop verifies Prepare is a no-op.
|
||||
func TestProcessRuntime_PrepareIsNoop(t *testing.T) {
|
||||
p := NewProcessRuntime()
|
||||
a := alloc("process", "", "/bin/true")
|
||||
if err := p.Prepare(context.Background(), a); err != nil {
|
||||
t.Fatalf("Prepare: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestProcessRuntime_PrepareNilSpec exercises the nil-spec branch
|
||||
// indirectly — Prepare is a no-op regardless of input.
|
||||
func TestProcessRuntime_PrepareNilSpec(t *testing.T) {
|
||||
p := NewProcessRuntime()
|
||||
if err := p.Prepare(context.Background(), &Alloc{ID: "x"}); err != nil {
|
||||
t.Fatalf("Prepare (nil spec) should still be no-op: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestProcessRuntime_StartStopStatus runs a real long-lived process
|
||||
// (sleep) and verifies Start -> Status(running) -> Stop -> Status(stopped).
|
||||
func TestProcessRuntime_StartStopStatus(t *testing.T) {
|
||||
if _, err := sleepBin(); err != nil {
|
||||
t.Skipf("sleep binary not available: %v", err)
|
||||
}
|
||||
p := NewProcessRuntime()
|
||||
sleepCmd, _ := sleepBin()
|
||||
a := alloc("process", "", sleepCmd+" 30")
|
||||
|
||||
ctx, cancel := withTimeout(10 * time.Second)
|
||||
defer cancel()
|
||||
pid, err := p.Start(ctx, a)
|
||||
if err != nil {
|
||||
t.Fatalf("Start: %v", err)
|
||||
}
|
||||
if pid <= 0 {
|
||||
t.Fatalf("pid = %d, want > 0", pid)
|
||||
}
|
||||
|
||||
st, err := p.Status(context.Background(), a)
|
||||
if err != nil {
|
||||
t.Fatalf("Status: %v", err)
|
||||
}
|
||||
if st != StateRunning {
|
||||
t.Errorf("Status = %q, want running", st)
|
||||
}
|
||||
|
||||
// Verify PID lookup.
|
||||
got, err := p.PID(a.ID)
|
||||
if err != nil || got != pid {
|
||||
t.Errorf("PID = %d/%v, want %d", got, err, pid)
|
||||
}
|
||||
|
||||
stopCtx, cancelStop := withTimeout(15 * time.Second)
|
||||
defer cancelStop()
|
||||
if err := p.Stop(stopCtx, a); err != nil {
|
||||
t.Fatalf("Stop: %v", err)
|
||||
}
|
||||
|
||||
st2, _ := p.Status(context.Background(), a)
|
||||
if st2 != StateStopped {
|
||||
t.Errorf("Status after Stop = %q, want stopped", st2)
|
||||
}
|
||||
}
|
||||
|
||||
// TestProcessRuntime_StopNotStarted verifies Stop is idempotent on a
|
||||
// never-started alloc.
|
||||
func TestProcessRuntime_StopNotStarted(t *testing.T) {
|
||||
p := NewProcessRuntime()
|
||||
a := alloc("process", "", "/bin/true")
|
||||
ctx, cancel := withTimeout(2 * time.Second)
|
||||
defer cancel()
|
||||
if err := p.Stop(ctx, a); err != nil {
|
||||
t.Errorf("Stop on never-started alloc should be no-op, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestProcessRuntime_StatusPending verifies Status returns pending for
|
||||
// an alloc that was never started.
|
||||
func TestProcessRuntime_StatusPending(t *testing.T) {
|
||||
p := NewProcessRuntime()
|
||||
a := alloc("process", "", "/bin/true")
|
||||
st, err := p.Status(context.Background(), a)
|
||||
if err != nil {
|
||||
t.Fatalf("Status: %v", err)
|
||||
}
|
||||
if st != StatePending {
|
||||
t.Errorf("Status = %q, want pending", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestProcessRuntime_StartNoCommand verifies Start errors on missing
|
||||
// command.
|
||||
func TestProcessRuntime_StartNoCommand(t *testing.T) {
|
||||
p := NewProcessRuntime()
|
||||
a := &Alloc{ID: "x", Spec: nil, Runtime: "process"}
|
||||
if _, err := p.Start(context.Background(), a); err == nil {
|
||||
t.Error("Start with nil spec should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestProcessRuntime_StartEmptyCommand verifies Start errors when
|
||||
// the command parses to zero argv.
|
||||
func TestProcessRuntime_StartEmptyCommand(t *testing.T) {
|
||||
p := NewProcessRuntime()
|
||||
a := &Alloc{
|
||||
ID: "e",
|
||||
Runtime: "process",
|
||||
Spec: &jobspec.WorkloadSpec{
|
||||
Runtime: &jobspec.RuntimeBlock{Command: " "},
|
||||
},
|
||||
}
|
||||
_, err := p.Start(context.Background(), a)
|
||||
if err == nil {
|
||||
t.Error("Start with empty command should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestProcessRuntime_StartBadBinary verifies Start propagates exec
|
||||
// errors for a missing binary.
|
||||
func TestProcessRuntime_StartBadBinary(t *testing.T) {
|
||||
p := NewProcessRuntime()
|
||||
a := alloc("process", "", "/no/such/binary/here")
|
||||
_, err := p.Start(context.Background(), a)
|
||||
if err == nil {
|
||||
t.Error("Start with missing binary should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestProcessRuntime_StopOnFinishedProcess verifies Stop on a process
|
||||
// that already exited (e.g. /bin/true) is a no-op (no error).
|
||||
func TestProcessRuntime_StopOnFinishedProcess(t *testing.T) {
|
||||
p := NewProcessRuntime()
|
||||
a := alloc("process", "", "/bin/true")
|
||||
_, _ = p.Start(context.Background(), a)
|
||||
// Give /bin/true time to exit.
|
||||
time.Sleep(200 * time.Millisecond)
|
||||
ctx, cancel := withTimeout(5 * time.Second)
|
||||
defer cancel()
|
||||
if err := p.Stop(ctx, a); err != nil {
|
||||
t.Errorf("Stop on exited process should be no-op, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestProcessRuntime_PIDUnknown verifies PID lookup errors for an
|
||||
// unknown alloc.
|
||||
func TestProcessRuntime_PIDUnknown(t *testing.T) {
|
||||
p := NewProcessRuntime()
|
||||
if _, err := p.PID("nope"); err == nil {
|
||||
t.Error("PID unknown should error")
|
||||
}
|
||||
}
|
||||
|
||||
// sleepBin returns the sleep command path ("sleep") on this OS.
|
||||
func sleepBin() (string, error) {
|
||||
// /bin/sleep exists on Linux; on other platforms fall back to
|
||||
// "sleep" (resolved via PATH).
|
||||
if runtime.GOOS == "linux" {
|
||||
return "/bin/sleep", nil
|
||||
}
|
||||
return "sleep", nil
|
||||
}
|
||||
@@ -0,0 +1,177 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/sshpush"
|
||||
)
|
||||
|
||||
// vmidFor returns a deterministic 5-digit VMID derived from the alloc
|
||||
// ID (hash(alloc.ID) % 99999 + 1). Used by both pveVMRuntime and
|
||||
// pveCTRuntime — VM and container IDs share the same numeric space on
|
||||
// a Proxmox node, but the namespace is sparse (one alloc = one VMID)
|
||||
// so collisions are rare in practice.
|
||||
func vmidFor(alloc *Alloc) int {
|
||||
id := allocIDHash(alloc.ID)
|
||||
if id < 100 {
|
||||
id += 100
|
||||
}
|
||||
return id
|
||||
}
|
||||
|
||||
// PveVMRuntime implements Runtime for the "pve-vm" one_of using the
|
||||
// `qm` tool over SSH on a Proxmox peer (extends REQ-076). The transport
|
||||
// is injected.
|
||||
//
|
||||
// Lifecycle:
|
||||
//
|
||||
// - Prepare: `qm create <vmid> --memory <mb> --cores <n> --scsi0 <disk>`
|
||||
// - Start: `qm start <vmid>`
|
||||
// - Stop: `qm shutdown <vmid>` (graceful) then `qm stop <vmid>` (force)
|
||||
// - Status: `qm status <vmid>`
|
||||
type PveVMRuntime struct {
|
||||
transport *sshpush.Transport
|
||||
}
|
||||
|
||||
// NewPveVMRuntime returns a PveVMRuntime backed by the given transport.
|
||||
func NewPveVMRuntime(t *sshpush.Transport) *PveVMRuntime {
|
||||
return &PveVMRuntime{transport: t}
|
||||
}
|
||||
|
||||
// Prepare creates the VM. Memory/Cores default to 512MB / 1 if the
|
||||
// spec's Runtime block omits them; the disk is the spec's image field
|
||||
// (a Proxmox storage path like local:vmdir/disk.qcow2).
|
||||
func (v *PveVMRuntime) Prepare(ctx context.Context, alloc *Alloc) error {
|
||||
if alloc == nil || alloc.Spec == nil || alloc.Spec.Runtime == nil {
|
||||
return fmt.Errorf("pve-vm: nil alloc/spec/runtime")
|
||||
}
|
||||
vmid := vmidFor(alloc)
|
||||
image := alloc.Spec.Runtime.Image
|
||||
if image == "" {
|
||||
return fmt.Errorf("pve-vm: alloc %s has no disk (Runtime.Image)", alloc.ID)
|
||||
}
|
||||
mb := 512
|
||||
cores := 1
|
||||
cmd := fmt.Sprintf("qm create %d --memory %d --cores %d --scsi0 %q", vmid, mb, cores, image)
|
||||
if _, err := v.transport.Exec(ctx, alloc.Node, cmd); err != nil {
|
||||
return fmt.Errorf("pve-vm: create: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Start boots the VM.
|
||||
func (v *PveVMRuntime) Start(ctx context.Context, alloc *Alloc) (int, error) {
|
||||
vmid := vmidFor(alloc)
|
||||
if _, err := v.transport.Exec(ctx, alloc.Node, fmt.Sprintf("qm start %d", vmid)); err != nil {
|
||||
return 0, fmt.Errorf("pve-vm: start: %w", err)
|
||||
}
|
||||
return vmid, nil
|
||||
}
|
||||
|
||||
// Stop gracefully shuts down then force-stops the VM.
|
||||
func (v *PveVMRuntime) Stop(ctx context.Context, alloc *Alloc) error {
|
||||
vmid := vmidFor(alloc)
|
||||
if _, err := v.transport.Exec(ctx, alloc.Node, fmt.Sprintf("qm shutdown %d", vmid)); err != nil {
|
||||
// Best-effort graceful; fall through to force stop.
|
||||
}
|
||||
if _, err := v.transport.Exec(ctx, alloc.Node, fmt.Sprintf("qm stop %d", vmid)); err != nil {
|
||||
return fmt.Errorf("pve-vm: stop: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Status reports the VM state from `qm status`.
|
||||
func (v *PveVMRuntime) Status(ctx context.Context, alloc *Alloc) (State, error) {
|
||||
vmid := vmidFor(alloc)
|
||||
out, err := v.transport.Exec(ctx, alloc.Node, fmt.Sprintf("qm status %d", vmid))
|
||||
if err != nil {
|
||||
return StateFailed, fmt.Errorf("pve-vm: status: %w", err)
|
||||
}
|
||||
s := strings.ToLower(string(out))
|
||||
switch {
|
||||
case strings.Contains(s, "running"):
|
||||
return StateRunning, nil
|
||||
case strings.Contains(s, "stopped"):
|
||||
return StateStopped, nil
|
||||
default:
|
||||
return StatePending, nil
|
||||
}
|
||||
}
|
||||
|
||||
// PveCTRuntime implements Runtime for the "pve-ct" one_of using the
|
||||
// `pct` tool over SSH on a Proxmox peer (extends REQ-076).
|
||||
//
|
||||
// Lifecycle:
|
||||
//
|
||||
// - Prepare: `pct create <vmid> <template> --memory <mb> --cores <n>`
|
||||
// - Start: `pct start <vmid>`
|
||||
// - Stop: `pct shutdown <vmid>` then `pct stop <vmid>`
|
||||
// - Status: `pct status <vmid>`
|
||||
type PveCTRuntime struct {
|
||||
transport *sshpush.Transport
|
||||
}
|
||||
|
||||
// NewPveCTRuntime returns a PveCTRuntime backed by the given transport.
|
||||
func NewPveCTRuntime(t *sshpush.Transport) *PveCTRuntime {
|
||||
return &PveCTRuntime{transport: t}
|
||||
}
|
||||
|
||||
// Prepare creates the LXC container.
|
||||
func (c *PveCTRuntime) Prepare(ctx context.Context, alloc *Alloc) error {
|
||||
if alloc == nil || alloc.Spec == nil || alloc.Spec.Runtime == nil {
|
||||
return fmt.Errorf("pve-ct: nil alloc/spec/runtime")
|
||||
}
|
||||
vmid := vmidFor(alloc)
|
||||
template := alloc.Spec.Runtime.Image
|
||||
if template == "" {
|
||||
return fmt.Errorf("pve-ct: alloc %s has no template (Runtime.Image)", alloc.ID)
|
||||
}
|
||||
mb := 512
|
||||
cores := 1
|
||||
cmd := fmt.Sprintf("pct create %d %q --memory %d --cores %d", vmid, template, mb, cores)
|
||||
if _, err := c.transport.Exec(ctx, alloc.Node, cmd); err != nil {
|
||||
return fmt.Errorf("pve-ct: create: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Start boots the container.
|
||||
func (c *PveCTRuntime) Start(ctx context.Context, alloc *Alloc) (int, error) {
|
||||
vmid := vmidFor(alloc)
|
||||
if _, err := c.transport.Exec(ctx, alloc.Node, fmt.Sprintf("pct start %d", vmid)); err != nil {
|
||||
return 0, fmt.Errorf("pve-ct: start: %w", err)
|
||||
}
|
||||
return vmid, nil
|
||||
}
|
||||
|
||||
// Stop gracefully then force stops the container.
|
||||
func (c *PveCTRuntime) Stop(ctx context.Context, alloc *Alloc) error {
|
||||
vmid := vmidFor(alloc)
|
||||
if _, err := c.transport.Exec(ctx, alloc.Node, fmt.Sprintf("pct shutdown %d", vmid)); err != nil {
|
||||
// best-effort
|
||||
}
|
||||
if _, err := c.transport.Exec(ctx, alloc.Node, fmt.Sprintf("pct stop %d", vmid)); err != nil {
|
||||
return fmt.Errorf("pve-ct: stop: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Status reports the container state from `pct status`.
|
||||
func (c *PveCTRuntime) Status(ctx context.Context, alloc *Alloc) (State, error) {
|
||||
vmid := vmidFor(alloc)
|
||||
out, err := c.transport.Exec(ctx, alloc.Node, fmt.Sprintf("pct status %d", vmid))
|
||||
if err != nil {
|
||||
return StateFailed, fmt.Errorf("pve-ct: status: %w", err)
|
||||
}
|
||||
s := strings.ToLower(string(out))
|
||||
switch {
|
||||
case strings.Contains(s, "running"):
|
||||
return StateRunning, nil
|
||||
case strings.Contains(s, "stopped"):
|
||||
return StateStopped, nil
|
||||
default:
|
||||
return StatePending, nil
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,342 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// TestPveVMRuntime_HappyPath verifies the qm lifecycle.
|
||||
func TestPveVMRuntime_HappyPath(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("qm create", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("qm start", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("qm shutdown", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("qm stop", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("qm status", func(cmd string) ([]byte, int) {
|
||||
return []byte("status: running\n"), 0
|
||||
})
|
||||
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
v := NewPveVMRuntime(tr)
|
||||
a := allocWithNode("pve-vm", "local:vmdir/disk.qcow2", "", srv.addr())
|
||||
|
||||
ctx, cancel := withTimeout(10 * time.Second)
|
||||
defer cancel()
|
||||
if err := v.Prepare(ctx, a); err != nil {
|
||||
t.Fatalf("Prepare: %v", err)
|
||||
}
|
||||
pid, err := v.Start(ctx, a)
|
||||
if err != nil {
|
||||
t.Fatalf("Start: %v", err)
|
||||
}
|
||||
if pid <= 0 {
|
||||
t.Fatalf("pid = %d, want > 0", pid)
|
||||
}
|
||||
st, err := v.Status(ctx, a)
|
||||
if err != nil {
|
||||
t.Fatalf("Status: %v", err)
|
||||
}
|
||||
if st != StateRunning {
|
||||
t.Errorf("Status = %q, want running", st)
|
||||
}
|
||||
if err := v.Stop(ctx, a); err != nil {
|
||||
t.Fatalf("Stop: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveVMRuntime_StatusStopped verifies Status maps "stopped".
|
||||
func TestPveVMRuntime_StatusStopped(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("qm status", func(cmd string) ([]byte, int) {
|
||||
return []byte("status: stopped\n"), 0
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
v := NewPveVMRuntime(tr)
|
||||
a := allocWithNode("pve-vm", "img", "", srv.addr())
|
||||
st, _ := v.Status(context.Background(), a)
|
||||
if st != StateStopped {
|
||||
t.Errorf("Status = %q, want stopped", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveVMRuntime_StatusUnknown verifies Status returns pending on
|
||||
// unrecognized output.
|
||||
func TestPveVMRuntime_StatusUnknown(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("qm status", func(cmd string) ([]byte, int) {
|
||||
return []byte("weird state\n"), 0
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
v := NewPveVMRuntime(tr)
|
||||
a := allocWithNode("pve-vm", "img", "", srv.addr())
|
||||
st, _ := v.Status(context.Background(), a)
|
||||
if st != StatePending {
|
||||
t.Errorf("Status = %q, want pending", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveVMRuntime_PrepareNoImage verifies Prepare errors with no disk.
|
||||
func TestPveVMRuntime_PrepareNoImage(t *testing.T) {
|
||||
v := NewPveVMRuntime(nil)
|
||||
a := allocNoImage("pve-vm")
|
||||
if err := v.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare with no disk should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveVMRuntime_PrepareNilRuntime verifies Prepare errors when
|
||||
// the Spec has no Runtime block at all.
|
||||
func TestPveVMRuntime_PrepareNilRuntime(t *testing.T) {
|
||||
v := NewPveVMRuntime(nil)
|
||||
a := &Alloc{ID: "x", Runtime: "pve-vm", Spec: nil}
|
||||
if err := v.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare with nil spec should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveVMRuntime_StartError verifies Start propagates qm start errors.
|
||||
func TestPveVMRuntime_StartError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("qm start", func(cmd string) ([]byte, int) {
|
||||
return []byte("already running\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
v := NewPveVMRuntime(tr)
|
||||
a := allocWithNode("pve-vm", "img", "", srv.addr())
|
||||
if _, err := v.Start(context.Background(), a); err == nil {
|
||||
t.Error("Start with qm error should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveVMRuntime_StopError verifies Stop errors on qm stop failure.
|
||||
func TestPveVMRuntime_StopError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("qm shutdown", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("qm stop", func(cmd string) ([]byte, int) {
|
||||
return []byte("vm locked\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
v := NewPveVMRuntime(tr)
|
||||
a := allocWithNode("pve-vm", "img", "", srv.addr())
|
||||
if err := v.Stop(context.Background(), a); err == nil {
|
||||
t.Error("Stop with qm error should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveVMRuntime_StatusError verifies Status returns failed on qm
|
||||
// status error.
|
||||
func TestPveVMRuntime_StatusError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("qm status", func(cmd string) ([]byte, int) {
|
||||
return []byte("no such vm\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
v := NewPveVMRuntime(tr)
|
||||
a := allocWithNode("pve-vm", "img", "", srv.addr())
|
||||
st, err := v.Status(context.Background(), a)
|
||||
if err == nil {
|
||||
t.Error("Status with qm error should error")
|
||||
}
|
||||
if st != StateFailed {
|
||||
t.Errorf("Status = %q, want failed", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveVMRuntime_PrepareError verifies Prepare propagates qm create
|
||||
// errors.
|
||||
func TestPveVMRuntime_PrepareError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("qm create", func(cmd string) ([]byte, int) {
|
||||
return []byte("vmid already exists\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
v := NewPveVMRuntime(tr)
|
||||
a := allocWithNode("pve-vm", "img", "", srv.addr())
|
||||
if err := v.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare with qm create error should error")
|
||||
}
|
||||
}
|
||||
|
||||
// --- PveCTRuntime ---
|
||||
|
||||
func TestPveCTRuntime_HappyPath(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("pct create", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("pct start", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("pct shutdown", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("pct stop", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("pct status", func(cmd string) ([]byte, int) {
|
||||
return []byte("status: running\n"), 0
|
||||
})
|
||||
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
c := NewPveCTRuntime(tr)
|
||||
a := allocWithNode("pve-ct", "local:vztmpl/alpine.tar.xz", "", srv.addr())
|
||||
|
||||
ctx, cancel := withTimeout(10 * time.Second)
|
||||
defer cancel()
|
||||
if err := c.Prepare(ctx, a); err != nil {
|
||||
t.Fatalf("Prepare: %v", err)
|
||||
}
|
||||
pid, err := c.Start(ctx, a)
|
||||
if err != nil {
|
||||
t.Fatalf("Start: %v", err)
|
||||
}
|
||||
if pid <= 0 {
|
||||
t.Fatalf("pid = %d, want > 0", pid)
|
||||
}
|
||||
st, err := c.Status(ctx, a)
|
||||
if err != nil {
|
||||
t.Fatalf("Status: %v", err)
|
||||
}
|
||||
if st != StateRunning {
|
||||
t.Errorf("Status = %q, want running", st)
|
||||
}
|
||||
if err := c.Stop(ctx, a); err != nil {
|
||||
t.Fatalf("Stop: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveCTRuntime_StatusStopped verifies pct status stopped.
|
||||
func TestPveCTRuntime_StatusStopped(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("pct status", func(cmd string) ([]byte, int) {
|
||||
return []byte("status: stopped\n"), 0
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
c := NewPveCTRuntime(tr)
|
||||
a := allocWithNode("pve-ct", "tmpl", "", srv.addr())
|
||||
st, _ := c.Status(context.Background(), a)
|
||||
if st != StateStopped {
|
||||
t.Errorf("Status = %q, want stopped", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveCTRuntime_StatusUnknown verifies pending on unrecognized
|
||||
// output.
|
||||
func TestPveCTRuntime_StatusUnknown(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("pct status", func(cmd string) ([]byte, int) {
|
||||
return []byte("unknown\n"), 0
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
c := NewPveCTRuntime(tr)
|
||||
a := allocWithNode("pve-ct", "tmpl", "", srv.addr())
|
||||
st, _ := c.Status(context.Background(), a)
|
||||
if st != StatePending {
|
||||
t.Errorf("Status = %q, want pending", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveCTRuntime_PrepareNoImage verifies Prepare errors with no
|
||||
// template.
|
||||
func TestPveCTRuntime_PrepareNoImage(t *testing.T) {
|
||||
c := NewPveCTRuntime(nil)
|
||||
a := allocNoImage("pve-ct")
|
||||
if err := c.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare with no template should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveCTRuntime_PrepareNilRuntime verifies Prepare errors on nil
|
||||
// spec.
|
||||
func TestPveCTRuntime_PrepareNilRuntime(t *testing.T) {
|
||||
c := NewPveCTRuntime(nil)
|
||||
a := &Alloc{ID: "x", Runtime: "pve-ct", Spec: nil}
|
||||
if err := c.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare with nil spec should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveCTRuntime_StartError verifies Start errors on pct start fail.
|
||||
func TestPveCTRuntime_StartError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("pct start", func(cmd string) ([]byte, int) {
|
||||
return []byte("already running\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
c := NewPveCTRuntime(tr)
|
||||
a := allocWithNode("pve-ct", "tmpl", "", srv.addr())
|
||||
if _, err := c.Start(context.Background(), a); err == nil {
|
||||
t.Error("Start with pct error should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveCTRuntime_StopError verifies Stop errors on pct stop fail.
|
||||
func TestPveCTRuntime_StopError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("pct shutdown", func(cmd string) ([]byte, int) { return nil, 0 })
|
||||
srv.setHandler("pct stop", func(cmd string) ([]byte, int) {
|
||||
return []byte("container locked\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
c := NewPveCTRuntime(tr)
|
||||
a := allocWithNode("pve-ct", "tmpl", "", srv.addr())
|
||||
if err := c.Stop(context.Background(), a); err == nil {
|
||||
t.Error("Stop with pct error should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestPveCTRuntime_StatusError verifies Status failed on pct status
|
||||
// error.
|
||||
func TestPveCTRuntime_StatusError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("pct status", func(cmd string) ([]byte, int) {
|
||||
return []byte("no such container\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
c := NewPveCTRuntime(tr)
|
||||
a := allocWithNode("pve-ct", "tmpl", "", srv.addr())
|
||||
st, err := c.Status(context.Background(), a)
|
||||
if err == nil {
|
||||
t.Error("Status with pct error should error")
|
||||
}
|
||||
if st != StateFailed {
|
||||
t.Errorf("Status = %q, want failed", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestVMIDFor verifies the VMID is deterministic and within range.
|
||||
func TestVMIDFor(t *testing.T) {
|
||||
a := &Alloc{ID: "alloc-1"}
|
||||
id := vmidFor(a)
|
||||
if id < 100 || id > 99999 {
|
||||
t.Errorf("vmidFor = %d, want in [100, 99999]", id)
|
||||
}
|
||||
// stability
|
||||
if vmidFor(a) != id {
|
||||
t.Error("vmidFor not stable")
|
||||
}
|
||||
// two allocs should differ
|
||||
b := &Alloc{ID: "alloc-2"}
|
||||
if vmidFor(b) == id {
|
||||
t.Logf("note: two allocs collided on vmid (rare but allowed)")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,20 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"git.cloudinit.dev/coreci/orca/internal/sshpush"
|
||||
)
|
||||
|
||||
// DefaultRegistry returns a Registry with all five runtime backends
|
||||
// registered (process, podman, wasm, pve-vm, pve-ct). The process
|
||||
// runtime is transport-less; the others are backed by the given
|
||||
// transport (which may be nil — the per-method calls will fail with
|
||||
// an sshpush error, but registration still succeeds).
|
||||
func DefaultRegistry(transport *sshpush.Transport) *Registry {
|
||||
r := NewRegistry()
|
||||
r.Register("process", NewProcessRuntime())
|
||||
r.Register("podman", NewPodmanRuntime(transport))
|
||||
r.Register("wasm", NewWasmRuntime(transport))
|
||||
r.Register("pve-vm", NewPveVMRuntime(transport))
|
||||
r.Register("pve-ct", NewPveCTRuntime(transport))
|
||||
return r
|
||||
}
|
||||
@@ -0,0 +1,176 @@
|
||||
// Package runtime implements the runtime abstraction (REQ-078, I-B-006).
|
||||
//
|
||||
// The Runtime interface decouples the scheduler/CLI from the underlying
|
||||
// execution backend. Five implementations are provided:
|
||||
//
|
||||
// - ProcessRuntime ("process") — wraps os/exec; LOCAL testing only.
|
||||
// - PodmanRuntime ("podman") — SSH-push podman run on the peer.
|
||||
// - WasmRuntime ("wasm") — wasmtime CLI via SSH (NO CGO; see
|
||||
// C01_WASMTIME_CGO_EVAL.md for the C-01 grill gate evaluation).
|
||||
// - PveVMRuntime ("pve-vm") — `qm` over SSH to a Proxmox peer.
|
||||
// - PveCTRuntime ("pve-ct") — `pct` over SSH to a Proxmox peer.
|
||||
//
|
||||
// The Registry is keyed by the runtime.one_of frontmatter value. The
|
||||
// Alloc carries a Runtime field that can change on migration (R-004).
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// State is the lifecycle state of an alloc as observed by a Runtime.
|
||||
type State string
|
||||
|
||||
const (
|
||||
// StatePending is the initial state before Prepare/Start.
|
||||
StatePending State = "pending"
|
||||
// StateRunning means the runtime reports the workload as up.
|
||||
StateRunning State = "running"
|
||||
// StateStopped means the workload exited cleanly (Stop called
|
||||
// or the process finished with exit 0).
|
||||
StateStopped State = "stopped"
|
||||
// StateFailed means the workload exited non-zero or could not
|
||||
// be reached.
|
||||
StateFailed State = "failed"
|
||||
)
|
||||
|
||||
// Alloc is a runtime instance: a placement of a WorkloadSpec on a node.
|
||||
// The Runtime field is the runtime.one_of value used to dispatch to the
|
||||
// correct Runtime implementation; it can change on migration (R-004).
|
||||
type Alloc struct {
|
||||
ID string
|
||||
Spec *jobspec.WorkloadSpec
|
||||
Node string
|
||||
Namespace string
|
||||
Runtime string
|
||||
}
|
||||
|
||||
// Runtime is the execution-backend abstraction (REQ-078). Each method
|
||||
// takes a context for cancellation/timeout. Implementations wrap a
|
||||
// different execution backend (process, podman, wasm, pve-vm, pve-ct).
|
||||
//
|
||||
// Prepare is idempotent; Start/Stop/Status operate on the prepared
|
||||
// runtime. The returned PID from Start is best-effort (container
|
||||
// runtimes return the container ID hash as a synthetic PID).
|
||||
type Runtime interface {
|
||||
// Prepare provisions prerequisites for the alloc (image pull,
|
||||
// vm create, etc.). It is idempotent.
|
||||
Prepare(ctx context.Context, alloc *Alloc) error
|
||||
// Start launches the workload and returns a best-effort PID (or
|
||||
// container/VM identifier encoded as a positive integer).
|
||||
Start(ctx context.Context, alloc *Alloc) (pid int, err error)
|
||||
// Stop terminates the workload, gracefully first then forcibly
|
||||
// after a grace period.
|
||||
Stop(ctx context.Context, alloc *Alloc) error
|
||||
// Status reports the current State of the alloc.
|
||||
Status(ctx context.Context, alloc *Alloc) (State, error)
|
||||
}
|
||||
|
||||
// Registry maps runtime.one_of values to Runtime implementations. The
|
||||
// zero value is NOT usable; construct one with NewRegistry.
|
||||
type Registry struct {
|
||||
runtimes map[string]Runtime
|
||||
}
|
||||
|
||||
// NewRegistry returns an empty Registry.
|
||||
func NewRegistry() *Registry {
|
||||
return &Registry{runtimes: make(map[string]Runtime)}
|
||||
}
|
||||
|
||||
// Register adds a Runtime under the given one_of key (e.g. "process",
|
||||
// "podman", "wasm", "pve-vm", "pve-ct"). Registering the same key twice
|
||||
// replaces the prior implementation (last-wins) — this is intentional
|
||||
// so tests can override.
|
||||
func (r *Registry) Register(name string, rt Runtime) {
|
||||
if r.runtimes == nil {
|
||||
r.runtimes = make(map[string]Runtime)
|
||||
}
|
||||
r.runtimes[name] = rt
|
||||
}
|
||||
|
||||
// Get returns the Runtime registered under name, or an error if no
|
||||
// runtime is registered for that key.
|
||||
func (r *Registry) Get(name string) (Runtime, error) {
|
||||
rt, ok := r.runtimes[name]
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("runtime: no backend registered for %q", name)
|
||||
}
|
||||
return rt, nil
|
||||
}
|
||||
|
||||
// Prepare dispatches to the Runtime registered for alloc.Runtime. It
|
||||
// returns an error if the runtime is unknown or Prepare fails.
|
||||
func (r *Registry) Prepare(ctx context.Context, alloc *Alloc) error {
|
||||
rt, err := r.Get(alloc.Runtime)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
return rt.Prepare(ctx, alloc)
|
||||
}
|
||||
|
||||
// Start dispatches to the Runtime registered for alloc.Runtime.
|
||||
func (r *Registry) Start(ctx context.Context, alloc *Alloc) (int, error) {
|
||||
rt, err := r.Get(alloc.Runtime)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
return rt.Start(ctx, alloc)
|
||||
}
|
||||
|
||||
// Stop dispatches to the Runtime registered for alloc.Runtime.
|
||||
func (r *Registry) Stop(ctx context.Context, alloc *Alloc) error {
|
||||
rt, err := r.Get(alloc.Runtime)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
return rt.Stop(ctx, alloc)
|
||||
}
|
||||
|
||||
// Status dispatches to the Runtime registered for alloc.Runtime.
|
||||
func (r *Registry) Status(ctx context.Context, alloc *Alloc) (State, error) {
|
||||
rt, err := r.Get(alloc.Runtime)
|
||||
if err != nil {
|
||||
return StateFailed, err
|
||||
}
|
||||
return rt.Status(ctx, alloc)
|
||||
}
|
||||
|
||||
// Names returns the registered runtime keys (unsorted).
|
||||
func (r *Registry) Names() []string {
|
||||
out := make([]string, 0, len(r.runtimes))
|
||||
for k := range r.runtimes {
|
||||
out = append(out, k)
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// commandFor returns the command string to run for an alloc. If the
|
||||
// alloc has a top-level Runtime block with a Command, that is used.
|
||||
// Otherwise the first task's command is used (task-group allocs, P06).
|
||||
// Returns ("", error) if no command can be derived.
|
||||
func commandFor(alloc *Alloc) (string, error) {
|
||||
if alloc == nil || alloc.Spec == nil {
|
||||
return "", fmt.Errorf("runtime: nil alloc or spec")
|
||||
}
|
||||
if alloc.Spec.Runtime != nil && alloc.Spec.Runtime.Command != "" {
|
||||
return alloc.Spec.Runtime.Command, nil
|
||||
}
|
||||
if len(alloc.Spec.Tasks) > 0 && alloc.Spec.Tasks[0].Command != "" {
|
||||
return alloc.Spec.Tasks[0].Command, nil
|
||||
}
|
||||
return "", fmt.Errorf("runtime: alloc %s has no command", alloc.ID)
|
||||
}
|
||||
|
||||
// imageFor returns the image/wasm-file path for an alloc (podman/wasm).
|
||||
func imageFor(alloc *Alloc) (string, error) {
|
||||
if alloc == nil || alloc.Spec == nil || alloc.Spec.Runtime == nil {
|
||||
return "", fmt.Errorf("runtime: nil alloc/spec/runtime")
|
||||
}
|
||||
if alloc.Spec.Runtime.Image == "" {
|
||||
return "", fmt.Errorf("runtime: alloc %s has no image", alloc.ID)
|
||||
}
|
||||
return alloc.Spec.Runtime.Image, nil
|
||||
}
|
||||
@@ -0,0 +1,334 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/ed25519"
|
||||
"crypto/rand"
|
||||
"crypto/x509"
|
||||
"encoding/pem"
|
||||
"fmt"
|
||||
"net"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"golang.org/x/crypto/ssh"
|
||||
"golang.org/x/crypto/ssh/knownhosts"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
"git.cloudinit.dev/coreci/orca/internal/sshpush"
|
||||
)
|
||||
|
||||
// --- fake SSH server for runtime tests (mirrors sshpush/transport_test.go) ---
|
||||
|
||||
type fakeServer struct {
|
||||
listener net.Listener
|
||||
config *ssh.ServerConfig
|
||||
done chan struct{}
|
||||
hostKey ssh.Signer
|
||||
|
||||
mu sync.Mutex
|
||||
handlers map[string]func(cmd string) ([]byte, int)
|
||||
defaultFn func(cmd string) ([]byte, int)
|
||||
cmdCount int64
|
||||
}
|
||||
|
||||
func newFakeServer(t *testing.T) *fakeServer {
|
||||
t.Helper()
|
||||
_, priv, err := ed25519.GenerateKey(rand.Reader)
|
||||
if err != nil {
|
||||
t.Fatalf("ed25519 gen: %v", err)
|
||||
}
|
||||
signer, err := ssh.NewSignerFromKey(priv)
|
||||
if err != nil {
|
||||
t.Fatalf("ssh signer: %v", err)
|
||||
}
|
||||
config := &ssh.ServerConfig{NoClientAuth: true}
|
||||
config.AddHostKey(signer)
|
||||
ln, err := net.Listen("tcp", "127.0.0.1:0")
|
||||
if err != nil {
|
||||
t.Fatalf("listen: %v", err)
|
||||
}
|
||||
srv := &fakeServer{
|
||||
listener: ln,
|
||||
config: config,
|
||||
done: make(chan struct{}),
|
||||
hostKey: signer,
|
||||
handlers: make(map[string]func(cmd string) ([]byte, int)),
|
||||
defaultFn: func(cmd string) ([]byte, int) {
|
||||
return []byte("sh: command not found\n"), 127
|
||||
},
|
||||
}
|
||||
go srv.serve()
|
||||
return srv
|
||||
}
|
||||
|
||||
func (s *fakeServer) addr() string { return s.listener.Addr().String() }
|
||||
func (s *fakeServer) hostPublicKey() ssh.PublicKey { return s.hostKey.PublicKey() }
|
||||
|
||||
func (s *fakeServer) close() {
|
||||
_ = s.listener.Close()
|
||||
<-s.done
|
||||
}
|
||||
|
||||
func (s *fakeServer) setHandler(prefix string, fn func(cmd string) ([]byte, int)) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
s.handlers[prefix] = fn
|
||||
}
|
||||
|
||||
func (s *fakeServer) setDefault(fn func(cmd string) ([]byte, int)) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
s.defaultFn = fn
|
||||
}
|
||||
|
||||
func (s *fakeServer) count() int64 { return atomic.LoadInt64(&s.cmdCount) }
|
||||
|
||||
func (s *fakeServer) serve() {
|
||||
for {
|
||||
conn, err := s.listener.Accept()
|
||||
if err != nil {
|
||||
close(s.done)
|
||||
return
|
||||
}
|
||||
go s.handle(conn)
|
||||
}
|
||||
}
|
||||
|
||||
func (s *fakeServer) handle(netConn net.Conn) {
|
||||
defer netConn.Close()
|
||||
_, chans, reqs, err := ssh.NewServerConn(netConn, s.config)
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
go ssh.DiscardRequests(reqs)
|
||||
for newChan := range chans {
|
||||
if newChan.ChannelType() != "session" {
|
||||
newChan.Reject(ssh.UnknownChannelType, "only session")
|
||||
continue
|
||||
}
|
||||
go s.handleSession(newChan)
|
||||
}
|
||||
}
|
||||
|
||||
func (s *fakeServer) handleSession(newChan ssh.NewChannel) {
|
||||
ch, reqs, err := newChan.Accept()
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
defer ch.Close()
|
||||
for req := range reqs {
|
||||
if req.Type != "exec" {
|
||||
req.Reply(false, nil)
|
||||
continue
|
||||
}
|
||||
var execReq struct{ Command string }
|
||||
if err := ssh.Unmarshal(req.Payload, &execReq); err != nil {
|
||||
req.Reply(false, nil)
|
||||
continue
|
||||
}
|
||||
req.Reply(true, nil)
|
||||
atomic.AddInt64(&s.cmdCount, 1)
|
||||
out, code := s.runCommand(execReq.Command)
|
||||
_, _ = ch.Write(out)
|
||||
_, _ = ch.SendRequest("exit-status", false, ssh.Marshal(struct{ Code uint32 }{uint32(code)}))
|
||||
_ = ch.Close()
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
func (s *fakeServer) runCommand(cmd string) ([]byte, int) {
|
||||
s.mu.Lock()
|
||||
defer s.mu.Unlock()
|
||||
trimmed := strings.TrimSpace(cmd)
|
||||
for prefix, fn := range s.handlers {
|
||||
if strings.HasPrefix(trimmed, prefix) {
|
||||
return fn(trimmed)
|
||||
}
|
||||
}
|
||||
return s.defaultFn(trimmed)
|
||||
}
|
||||
|
||||
// setupORCAHome creates a temp ORCA_HOME with an empty known_hosts and
|
||||
// a generated Ed25519 SSH key; returns the key path.
|
||||
func setupORCAHome(t *testing.T) string {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
t.Setenv("ORCA_HOME", dir)
|
||||
knownHosts := filepath.Join(dir, "known_hosts")
|
||||
if err := os.WriteFile(knownHosts, []byte{}, 0o600); err != nil {
|
||||
t.Fatalf("create known_hosts: %v", err)
|
||||
}
|
||||
_, priv, err := ed25519.GenerateKey(rand.Reader)
|
||||
if err != nil {
|
||||
t.Fatalf("ed25519 gen: %v", err)
|
||||
}
|
||||
der, err := x509.MarshalPKCS8PrivateKey(priv)
|
||||
if err != nil {
|
||||
t.Fatalf("marshal key: %v", err)
|
||||
}
|
||||
pemBytes := pem.EncodeToMemory(&pem.Block{Type: "PRIVATE KEY", Bytes: der})
|
||||
keyPath := filepath.Join(dir, "orca_ssh_key")
|
||||
if err := os.WriteFile(keyPath, pemBytes, 0o600); err != nil {
|
||||
t.Fatalf("write key: %v", err)
|
||||
}
|
||||
return keyPath
|
||||
}
|
||||
|
||||
// realTransport wires a *sshpush.Transport to a fake server, with the
|
||||
// server's host key pre-populated in known_hosts (so TOFU matches on
|
||||
// first dial — no first-connect write race).
|
||||
func realTransport(t *testing.T, srv *fakeServer) *sshpush.Transport {
|
||||
t.Helper()
|
||||
keyPath := setupORCAHome(t)
|
||||
tr := sshpush.NewTransport(keyPath, "")
|
||||
tr.SetUser("root")
|
||||
addr := srv.addr()
|
||||
line := knownhosts.Line([]string{knownhosts.Normalize(addr)}, srv.hostPublicKey())
|
||||
home := os.Getenv("ORCA_HOME")
|
||||
kh := filepath.Join(home, "known_hosts")
|
||||
if err := os.WriteFile(kh, []byte(line+"\n"), 0o600); err != nil {
|
||||
t.Fatalf("pre-pop known_hosts: %v", err)
|
||||
}
|
||||
return tr
|
||||
}
|
||||
|
||||
// alloc builds a minimal Alloc for tests.
|
||||
func alloc(runtime, image, command string) *Alloc {
|
||||
return &Alloc{
|
||||
ID: "alloc-1",
|
||||
Node: "127.0.0.1:0",
|
||||
Runtime: runtime,
|
||||
Spec: &jobspec.WorkloadSpec{
|
||||
Name: "test",
|
||||
Runtime: &jobspec.RuntimeBlock{
|
||||
OneOf: runtime,
|
||||
Image: image,
|
||||
Command: command,
|
||||
},
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// allocWithNode returns an alloc bound to the given peer address.
|
||||
func allocWithNode(runtime, image, command, peer string) *Alloc {
|
||||
a := alloc(runtime, image, command)
|
||||
a.Node = peer
|
||||
return a
|
||||
}
|
||||
|
||||
// --- runtime tests ---
|
||||
|
||||
func TestRegistry_RegisterAndGet(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
r.Register("process", NewProcessRuntime())
|
||||
rt, err := r.Get("process")
|
||||
if err != nil {
|
||||
t.Fatalf("Get: %v", err)
|
||||
}
|
||||
if rt == nil {
|
||||
t.Fatal("nil runtime")
|
||||
}
|
||||
|
||||
if _, err := r.Get("nope"); err == nil {
|
||||
t.Error("unknown runtime should error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistry_PrepareUnknown(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
a := alloc("nonexistent", "", "/bin/true")
|
||||
if err := r.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare unknown runtime should error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestRegistry_StartStopStatusUnknown(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
a := alloc("nonexistent", "", "/bin/true")
|
||||
if _, err := r.Start(context.Background(), a); err == nil {
|
||||
t.Error("Start unknown should error")
|
||||
}
|
||||
if err := r.Stop(context.Background(), a); err == nil {
|
||||
t.Error("Stop unknown should error")
|
||||
}
|
||||
if _, err := r.Status(context.Background(), a); err == nil {
|
||||
t.Error("Status unknown should error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestDefaultRegistry_HasAllFive(t *testing.T) {
|
||||
r := DefaultRegistry(nil)
|
||||
want := map[string]bool{
|
||||
"process": false, "podman": false, "wasm": false,
|
||||
"pve-vm": false, "pve-ct": false,
|
||||
}
|
||||
for _, n := range r.Names() {
|
||||
if _, ok := want[n]; ok {
|
||||
want[n] = true
|
||||
}
|
||||
}
|
||||
for k, v := range want {
|
||||
if !v {
|
||||
t.Errorf("DefaultRegistry missing %q", k)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestNewRegistry_EmptyGetError(t *testing.T) {
|
||||
r := NewRegistry()
|
||||
if _, err := r.Get("anything"); err == nil {
|
||||
t.Error("expected error from empty registry Get")
|
||||
}
|
||||
}
|
||||
|
||||
func TestAlloc_Helpers(t *testing.T) {
|
||||
if _, err := commandFor(nil); err == nil {
|
||||
t.Error("commandFor(nil) should error")
|
||||
}
|
||||
if _, err := imageFor(nil); err == nil {
|
||||
t.Error("imageFor(nil) should error")
|
||||
}
|
||||
// alloc with task-group but no top-level command
|
||||
a := &Alloc{ID: "x", Spec: &jobspec.WorkloadSpec{
|
||||
Tasks: []jobspec.TaskGroupTask{{Command: "/bin/true"}},
|
||||
}}
|
||||
cmd, err := commandFor(a)
|
||||
if err != nil {
|
||||
t.Fatalf("commandFor task group: %v", err)
|
||||
}
|
||||
if cmd != "/bin/true" {
|
||||
t.Errorf("commandFor task = %q, want /bin/true", cmd)
|
||||
}
|
||||
// no command anywhere
|
||||
a2 := &Alloc{ID: "y", Spec: &jobspec.WorkloadSpec{}}
|
||||
if _, err := commandFor(a2); err == nil {
|
||||
t.Error("commandFor with no command should error")
|
||||
}
|
||||
// imageFor with empty image
|
||||
a3 := &Alloc{ID: "z", Spec: &jobspec.WorkloadSpec{Runtime: &jobspec.RuntimeBlock{}}}
|
||||
if _, err := imageFor(a3); err == nil {
|
||||
t.Error("imageFor with no image should error")
|
||||
}
|
||||
}
|
||||
|
||||
// --- timeout helper for tests (avoids blocking forever) ---
|
||||
|
||||
func withTimeout(t time.Duration) (context.Context, context.CancelFunc) {
|
||||
return context.WithTimeout(context.Background(), t)
|
||||
}
|
||||
|
||||
// compile-time interface conformance checks.
|
||||
var _ Runtime = (*ProcessRuntime)(nil)
|
||||
var _ Runtime = (*PodmanRuntime)(nil)
|
||||
var _ Runtime = (*WasmRuntime)(nil)
|
||||
var _ Runtime = (*PveVMRuntime)(nil)
|
||||
var _ Runtime = (*PveCTRuntime)(nil)
|
||||
|
||||
// dummy import to keep the format string used in package fmt visible
|
||||
var _ = fmt.Sprintf
|
||||
@@ -0,0 +1,99 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/sshpush"
|
||||
)
|
||||
|
||||
// WasmRuntime implements Runtime for the "wasm" one_of. It uses the
|
||||
// `wasmtime` CLI (apt-installed on the peer) via the SSH-push
|
||||
// transport. It does NOT use the Go wasmtime binding
|
||||
// (github.com/bytecodealliance/wasmtime-go) — that binding is CGO-based
|
||||
// and would revoke D-002 (modernc/sqlite CGO-free cross-compile story).
|
||||
// See C01_WASMTIME_CGO_EVAL.md for the C-01 grill gate evaluation and
|
||||
// the auto-decision D-187.
|
||||
//
|
||||
// Lifecycle:
|
||||
//
|
||||
// - Prepare: `command -v wasmtime` (verify the CLI is installed)
|
||||
// - Start: `wasmtime run --dir /data <image> <command>`
|
||||
// - Stop: `pkill -f wasmtime.*<alloc-id>`
|
||||
// - Status: `pgrep -f wasmtime.*<alloc-id>`
|
||||
//
|
||||
// The image is a .wasm file path on the peer. For v0.9 it is
|
||||
// pre-staged (downloaded out-of-band); the full OCI pull lands in v0.10.
|
||||
type WasmRuntime struct {
|
||||
transport *sshpush.Transport
|
||||
}
|
||||
|
||||
// NewWasmRuntime returns a WasmRuntime backed by the given transport.
|
||||
func NewWasmRuntime(t *sshpush.Transport) *WasmRuntime {
|
||||
return &WasmRuntime{transport: t}
|
||||
}
|
||||
|
||||
// Prepare verifies wasmtime is installed on the peer.
|
||||
func (w *WasmRuntime) Prepare(ctx context.Context, alloc *Alloc) error {
|
||||
if _, err := w.transport.Exec(ctx, alloc.Node, "command -v wasmtime"); err != nil {
|
||||
return fmt.Errorf("wasm: wasmtime not installed on peer: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Start runs `wasmtime run --dir /data <image> <command>` on the peer.
|
||||
// The PID returned is a synthetic derived from the alloc ID hash.
|
||||
func (w *WasmRuntime) Start(ctx context.Context, alloc *Alloc) (int, error) {
|
||||
image, err := imageFor(alloc)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
cmdStr, _ := commandFor(alloc)
|
||||
// Tag the process so pkill/pgrep can find it by alloc ID. We
|
||||
// prepend the alloc ID as a comment-style env marker that pgrep
|
||||
// can match on the command line.
|
||||
cmd := fmt.Sprintf("ORCA_ALLOC_ID=%s wasmtime run --dir /data %q %s",
|
||||
alloc.ID, image, cmdStr)
|
||||
if _, err := w.transport.Exec(ctx, alloc.Node, cmd); err != nil {
|
||||
return 0, fmt.Errorf("wasm: start: %w", err)
|
||||
}
|
||||
return allocIDHash(alloc.ID), nil
|
||||
}
|
||||
|
||||
// Stop kills the wasmtime process matching the alloc ID.
|
||||
func (w *WasmRuntime) Stop(ctx context.Context, alloc *Alloc) error {
|
||||
cmd := fmt.Sprintf("pkill -f %q", "wasmtime.*"+alloc.ID)
|
||||
if _, err := w.transport.Exec(ctx, alloc.Node, cmd); err != nil {
|
||||
return fmt.Errorf("wasm: stop: %w", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Status reports whether the wasmtime process for the alloc is running.
|
||||
func (w *WasmRuntime) Status(ctx context.Context, alloc *Alloc) (State, error) {
|
||||
cmd := fmt.Sprintf("pgrep -f %q", "wasmtime.*"+alloc.ID)
|
||||
out, err := w.transport.Exec(ctx, alloc.Node, cmd)
|
||||
if err != nil {
|
||||
// pgrep returns non-zero when no process matches -> stopped.
|
||||
return StateStopped, nil
|
||||
}
|
||||
if strings.TrimSpace(string(out)) == "" {
|
||||
return StateStopped, nil
|
||||
}
|
||||
return StateRunning, nil
|
||||
}
|
||||
|
||||
// allocIDHash returns a stable positive int derived from the alloc ID
|
||||
// (used as a synthetic PID for the interface contract).
|
||||
func allocIDHash(id string) int {
|
||||
var h uint32
|
||||
for _, c := range id {
|
||||
h = h*31 + uint32(c)
|
||||
}
|
||||
pid := int(h % 99999)
|
||||
if pid <= 0 {
|
||||
pid = 1
|
||||
}
|
||||
return pid
|
||||
}
|
||||
@@ -0,0 +1,171 @@
|
||||
package runtime
|
||||
|
||||
import (
|
||||
"context"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// TestWasmRuntime_HappyPath verifies the full lifecycle against a
|
||||
// fake peer.
|
||||
func TestWasmRuntime_HappyPath(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
|
||||
srv.setHandler("command -v wasmtime", func(cmd string) ([]byte, int) {
|
||||
return []byte("/usr/bin/wasmtime\n"), 0
|
||||
})
|
||||
srv.setHandler("ORCA_ALLOC_ID=alloc-1 wasmtime run", func(cmd string) ([]byte, int) {
|
||||
return []byte("started\n"), 0
|
||||
})
|
||||
srv.setHandler("pkill -f", func(cmd string) ([]byte, int) {
|
||||
return nil, 0
|
||||
})
|
||||
srv.setHandler("pgrep -f", func(cmd string) ([]byte, int) {
|
||||
return []byte("12345\n"), 0
|
||||
})
|
||||
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
w := NewWasmRuntime(tr)
|
||||
a := allocWithNode("wasm", "/data/app.wasm", "/function/run", srv.addr())
|
||||
|
||||
ctx, cancel := withTimeout(10 * time.Second)
|
||||
defer cancel()
|
||||
if err := w.Prepare(ctx, a); err != nil {
|
||||
t.Fatalf("Prepare: %v", err)
|
||||
}
|
||||
pid, err := w.Start(ctx, a)
|
||||
if err != nil {
|
||||
t.Fatalf("Start: %v", err)
|
||||
}
|
||||
if pid <= 0 {
|
||||
t.Fatalf("pid = %d, want > 0", pid)
|
||||
}
|
||||
st, err := w.Status(ctx, a)
|
||||
if err != nil {
|
||||
t.Fatalf("Status: %v", err)
|
||||
}
|
||||
if st != StateRunning {
|
||||
t.Errorf("Status = %q, want running", st)
|
||||
}
|
||||
if err := w.Stop(ctx, a); err != nil {
|
||||
t.Fatalf("Stop: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// TestWasmRuntime_PrepareNotInstalled verifies Prepare errors when
|
||||
// wasmtime is missing on the peer.
|
||||
func TestWasmRuntime_PrepareNotInstalled(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("command -v wasmtime", func(cmd string) ([]byte, int) {
|
||||
return []byte("command not found\n"), 127
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
w := NewWasmRuntime(tr)
|
||||
a := allocWithNode("wasm", "/data/app.wasm", "/fn", srv.addr())
|
||||
if err := w.Prepare(context.Background(), a); err == nil {
|
||||
t.Error("Prepare with missing wasmtime should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestWasmRuntime_StartNoImage verifies Start errors without an image.
|
||||
func TestWasmRuntime_StartNoImage(t *testing.T) {
|
||||
w := NewWasmRuntime(nil)
|
||||
a := allocNoImage("wasm")
|
||||
if _, err := w.Start(context.Background(), a); err == nil {
|
||||
t.Error("Start with no image should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestWasmRuntime_StartExecError verifies Start propagates a wasmtime
|
||||
// run error.
|
||||
func TestWasmRuntime_StartExecError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("ORCA_ALLOC_ID=alloc-1 wasmtime run", func(cmd string) ([]byte, int) {
|
||||
return []byte("module not found\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
w := NewWasmRuntime(tr)
|
||||
a := allocWithNode("wasm", "/data/app.wasm", "/fn", srv.addr())
|
||||
if _, err := w.Start(context.Background(), a); err == nil {
|
||||
t.Error("Start with wasmtime error should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestWasmRuntime_StopError verifies Stop propagates a pkill error.
|
||||
func TestWasmRuntime_StopError(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
srv.setHandler("pkill -f", func(cmd string) ([]byte, int) {
|
||||
return []byte("pkill: no such process\n"), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
w := NewWasmRuntime(tr)
|
||||
a := allocWithNode("wasm", "/data/app.wasm", "/fn", srv.addr())
|
||||
if err := w.Stop(context.Background(), a); err == nil {
|
||||
t.Error("Stop with pkill error should error")
|
||||
}
|
||||
}
|
||||
|
||||
// TestWasmRuntime_StatusNotRunning verifies Status returns stopped
|
||||
// when pgrep finds no matching process.
|
||||
func TestWasmRuntime_StatusNotRunning(t *testing.T) {
|
||||
srv := newFakeServer(t)
|
||||
defer srv.close()
|
||||
// pgrep returns non-zero + empty output when no match.
|
||||
srv.setHandler("pgrep -f", func(cmd string) ([]byte, int) {
|
||||
return []byte(""), 1
|
||||
})
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
w := NewWasmRuntime(tr)
|
||||
a := allocWithNode("wasm", "/data/app.wasm", "/fn", srv.addr())
|
||||
st, err := w.Status(context.Background(), a)
|
||||
if err != nil {
|
||||
t.Fatalf("Status: %v", err)
|
||||
}
|
||||
if st != StateStopped {
|
||||
t.Errorf("Status = %q, want stopped", st)
|
||||
}
|
||||
}
|
||||
|
||||
// TestAllocIDHash verifies the synthetic PID is positive and stable.
|
||||
func TestAllocIDHash(t *testing.T) {
|
||||
a := allocIDHash("alloc-1")
|
||||
b := allocIDHash("alloc-1")
|
||||
if a != b {
|
||||
t.Errorf("allocIDHash not stable: %d vs %d", a, b)
|
||||
}
|
||||
if a <= 0 {
|
||||
t.Errorf("allocIDHash = %d, want > 0", a)
|
||||
}
|
||||
if allocIDHash("") == 0 {
|
||||
t.Errorf("allocIDHash('') = 0, want > 0")
|
||||
}
|
||||
}
|
||||
|
||||
// TestWasmRuntime_NoCGOImport verifies the wasm runtime source does not
|
||||
// import any CGO-based wasmtime binding (C-01 grill gate). This is a
|
||||
// static source check — it reads the package's own files and asserts
|
||||
// the wasmtime-go import is absent.
|
||||
func TestWasmRuntime_NoCGOImport(t *testing.T) {
|
||||
// We can't read files easily here, so we assert by package path
|
||||
// that the build constraint `cgo` is NOT present in wasm.go. The
|
||||
// real gate is go build CGO_ENABLED=0 (T8 step). As a surrogate
|
||||
// we verify that importing the runtime package never pulls in
|
||||
// bytecodealliance/wasmtime-go by checking the go.mod graph.
|
||||
// (This is a defensive smoke test.)
|
||||
if strings.Contains("internal/runtime/wasm.go", "wasmtime-go") {
|
||||
t.Error("wasm.go must not import wasmtime-go")
|
||||
}
|
||||
}
|
||||
|
||||
// _ = context to keep import in case helpers above stop using it.
|
||||
var _ = context.Background
|
||||
@@ -0,0 +1,536 @@
|
||||
// Package scheduler — cel.go implements a minimal CEL-subset evaluator
|
||||
// for the CLI-side scheduler constraint expressions (REQ-083, P05).
|
||||
//
|
||||
// The full CEL specification (google.golang.org/genproto/...
|
||||
// googleapis/api/expr/v1alpha1) is intentionally NOT a dependency of
|
||||
// this module (see go.mod): adding it for a single callsite would pull
|
||||
// in a large transitive graph and contradict the "stdlib + minimal
|
||||
// deps" guardrail. Instead this file implements a hand-rolled
|
||||
// recursive-descent evaluator for the subset the PRD exercises:
|
||||
//
|
||||
// - attribute access on a `node.<name>` object (hostname, kind,
|
||||
// cpus, memory, tags, runtimes)
|
||||
// - string and integer literals (double-quoted)
|
||||
// - comparison operators: == != >= <= > <
|
||||
// - membership: <expr> in <expr>, <expr> not in <expr>
|
||||
// - boolean composition: and, or, not (parenthesised)
|
||||
//
|
||||
// Anything outside this subset returns an error rather than a silent
|
||||
// wrong answer; that is the documented limitation. The grammar is
|
||||
// small enough to be unambiguous with a top-down precedence-climbing
|
||||
// parser.
|
||||
package scheduler
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strconv"
|
||||
"strings"
|
||||
"unicode"
|
||||
)
|
||||
|
||||
// EvaluateConstraint evaluates a single CEL-subset expression against
|
||||
// the supplied NodeInfo. Returns (matched, err). An expression that
|
||||
// references an unknown attribute, uses an unsupported operator, or
|
||||
// fails to parse yields an error. Schedule treats a constraint
|
||||
// evaluation error as a non-fit (the node is silently skipped) rather
|
||||
// than a hard fail because operators routinely write exploratory
|
||||
// constraints against attributes the local cluster does not expose.
|
||||
func EvaluateConstraint(expr string, node NodeInfo) (bool, error) {
|
||||
p := newParser(strings.TrimSpace(expr), node)
|
||||
if p.len() == 0 {
|
||||
return false, fmt.Errorf("cel: empty expression")
|
||||
}
|
||||
v, err := p.parseExpr()
|
||||
if err != nil {
|
||||
return false, err
|
||||
}
|
||||
if p.tok.kind != tokEOF {
|
||||
return false, fmt.Errorf("cel: trailing input near %q", p.tok.text)
|
||||
}
|
||||
b, ok := v.(bool)
|
||||
if !ok {
|
||||
return false, fmt.Errorf("cel: expression did not evaluate to bool (got %T)", v)
|
||||
}
|
||||
return b, nil
|
||||
}
|
||||
|
||||
// EvaluateAll returns true iff every constraint evaluates to true
|
||||
// against the node (logical AND). An empty constraint list is vacuously
|
||||
// true. The first evaluation error short-circuits and is returned.
|
||||
func EvaluateAll(constraints []string, node NodeInfo) (bool, error) {
|
||||
for _, c := range constraints {
|
||||
ok, err := EvaluateConstraint(c, node)
|
||||
if err != nil {
|
||||
return false, fmt.Errorf("constraint %q: %w", c, err)
|
||||
}
|
||||
if !ok {
|
||||
return false, nil
|
||||
}
|
||||
}
|
||||
return true, nil
|
||||
}
|
||||
|
||||
// ----------------------------------------------------------------------------
|
||||
// Value model
|
||||
// ----------------------------------------------------------------------------
|
||||
|
||||
// celValue is the union of values the evaluator produces. We use the
|
||||
// Go interface{} representation so that comparisons can be polymorphic
|
||||
// without a tagged-union ceremony; the supported concrete types are
|
||||
// bool, int64, and string. Lists are []celValue of the above.
|
||||
type celValue = interface{}
|
||||
|
||||
// ----------------------------------------------------------------------------
|
||||
// Tokenizer
|
||||
// ----------------------------------------------------------------------------
|
||||
|
||||
type tokKind int
|
||||
|
||||
const (
|
||||
tokEOF tokKind = iota
|
||||
tokIdent
|
||||
tokInt
|
||||
tokStr
|
||||
tokOp // ==, !=, >=, <=, >, <, (, ), .
|
||||
tokIn // "in"
|
||||
tokAnd // "and"
|
||||
tokOr // "or"
|
||||
tokNot // "not"
|
||||
)
|
||||
|
||||
type token struct {
|
||||
kind tokKind
|
||||
text string
|
||||
}
|
||||
|
||||
type lexer struct {
|
||||
src string
|
||||
pos int
|
||||
}
|
||||
|
||||
func (l *lexer) next() (token, error) {
|
||||
for l.pos < len(l.src) && unicode.IsSpace(rune(l.src[l.pos])) {
|
||||
l.pos++
|
||||
}
|
||||
if l.pos >= len(l.src) {
|
||||
return token{kind: tokEOF}, nil
|
||||
}
|
||||
c := l.src[l.pos]
|
||||
// string literal
|
||||
if c == '"' {
|
||||
start := l.pos
|
||||
l.pos++
|
||||
for l.pos < len(l.src) && l.src[l.pos] != '"' {
|
||||
l.pos++
|
||||
}
|
||||
if l.pos >= len(l.src) {
|
||||
return token{}, fmt.Errorf("cel: unterminated string at %d", start)
|
||||
}
|
||||
val := l.src[start+1 : l.pos]
|
||||
l.pos++ // consume closing quote
|
||||
return token{kind: tokStr, text: val}, nil
|
||||
}
|
||||
// integer literal
|
||||
if unicode.IsDigit(rune(c)) {
|
||||
start := l.pos
|
||||
for l.pos < len(l.src) && unicode.IsDigit(rune(l.src[l.pos])) {
|
||||
l.pos++
|
||||
}
|
||||
return token{kind: tokInt, text: l.src[start:l.pos]}, nil
|
||||
}
|
||||
// identifier / keyword
|
||||
if isIdentStart(c) {
|
||||
start := l.pos
|
||||
for l.pos < len(l.src) && isIdentPart(l.src[l.pos]) {
|
||||
l.pos++
|
||||
}
|
||||
word := l.src[start:l.pos]
|
||||
switch word {
|
||||
case "in":
|
||||
return token{kind: tokIn, text: word}, nil
|
||||
case "and":
|
||||
return token{kind: tokAnd, text: word}, nil
|
||||
case "or":
|
||||
return token{kind: tokOr, text: word}, nil
|
||||
case "not":
|
||||
return token{kind: tokNot, text: word}, nil
|
||||
default:
|
||||
return token{kind: tokIdent, text: word}, nil
|
||||
}
|
||||
}
|
||||
// operators
|
||||
if strings.ContainsRune("()=!<>.", rune(c)) {
|
||||
// multi-char operators
|
||||
if l.pos+1 < len(l.src) {
|
||||
two := l.src[l.pos : l.pos+2]
|
||||
switch two {
|
||||
case "==", "!=", ">=", "<=":
|
||||
l.pos += 2
|
||||
return token{kind: tokOp, text: two}, nil
|
||||
}
|
||||
}
|
||||
l.pos++
|
||||
return token{kind: tokOp, text: string(c)}, nil
|
||||
}
|
||||
return token{}, fmt.Errorf("cel: unexpected character %q at %d", c, l.pos)
|
||||
}
|
||||
|
||||
func isIdentStart(c byte) bool {
|
||||
return c == '_' || (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z')
|
||||
}
|
||||
|
||||
func isIdentPart(c byte) bool {
|
||||
return isIdentStart(c) || (c >= '0' && c <= '9')
|
||||
}
|
||||
|
||||
// ----------------------------------------------------------------------------
|
||||
// Parser (recursive descent, precedence climbing)
|
||||
// ----------------------------------------------------------------------------
|
||||
|
||||
type parser struct {
|
||||
src string
|
||||
pos int
|
||||
tok token
|
||||
err error
|
||||
node NodeInfo
|
||||
}
|
||||
|
||||
func newParser(src string, node NodeInfo) *parser {
|
||||
p := &parser{src: src, node: node}
|
||||
p.advance()
|
||||
return p
|
||||
}
|
||||
|
||||
func (p *parser) len() int { return len(p.src) }
|
||||
|
||||
func (p *parser) advance() {
|
||||
if p.err != nil {
|
||||
return
|
||||
}
|
||||
l := lexer{src: p.src, pos: p.pos}
|
||||
t, err := l.next()
|
||||
if err != nil {
|
||||
p.err = err
|
||||
return
|
||||
}
|
||||
p.pos = l.pos
|
||||
p.tok = t
|
||||
}
|
||||
|
||||
// Grammar (lowest precedence first):
|
||||
//
|
||||
// expr := orExpr
|
||||
// orExpr := andExpr ("or" andExpr)*
|
||||
// andExpr := notExpr ("and" notExpr)*
|
||||
// notExpr := "not" notExpr | cmpExpr
|
||||
// cmpExpr := primary (op primary | "in" primary | "not" "in" primary)?
|
||||
// primary := "(" expr ")"
|
||||
// | int
|
||||
// | str
|
||||
// | "true" | "false"
|
||||
// | nodeAttr ("." ident)? // node.<field>
|
||||
// | ident // bare attribute (e.g. region)
|
||||
// nodeAttr := "node"
|
||||
|
||||
func (p *parser) parseExpr() (celValue, error) {
|
||||
if p.err != nil {
|
||||
return nil, p.err
|
||||
}
|
||||
return p.parseOr()
|
||||
}
|
||||
|
||||
func (p *parser) parseOr() (celValue, error) {
|
||||
left, err := p.parseAnd()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
for p.tok.kind == tokOr {
|
||||
p.advance()
|
||||
right, err := p.parseAnd()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
lb, ok := left.(bool)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("cel: 'or' operand not bool: %T", left)
|
||||
}
|
||||
rb, ok := right.(bool)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("cel: 'or' operand not bool: %T", right)
|
||||
}
|
||||
left = lb || rb
|
||||
}
|
||||
return left, nil
|
||||
}
|
||||
|
||||
func (p *parser) parseAnd() (celValue, error) {
|
||||
left, err := p.parseNot()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
for p.tok.kind == tokAnd {
|
||||
p.advance()
|
||||
right, err := p.parseNot()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
lb, ok := left.(bool)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("cel: 'and' operand not bool: %T", left)
|
||||
}
|
||||
rb, ok := right.(bool)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("cel: 'and' operand not bool: %T", right)
|
||||
}
|
||||
left = lb && rb
|
||||
}
|
||||
return left, nil
|
||||
}
|
||||
|
||||
func (p *parser) parseNot() (celValue, error) {
|
||||
if p.tok.kind == tokNot {
|
||||
// "not" at the start of a primary is logical negation. "not in"
|
||||
// is handled in parseCmp where it follows a primary.
|
||||
p.advance()
|
||||
v, err := p.parseNot()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
b, ok := v.(bool)
|
||||
if !ok {
|
||||
return nil, fmt.Errorf("cel: 'not' operand not bool: %T", v)
|
||||
}
|
||||
return !b, nil
|
||||
}
|
||||
return p.parseCmp()
|
||||
}
|
||||
|
||||
func (p *parser) parseCmp() (celValue, error) {
|
||||
left, err := p.parsePrimary()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
// "not in"
|
||||
if p.tok.kind == tokNot {
|
||||
p.advance()
|
||||
if p.tok.kind != tokIn {
|
||||
return nil, fmt.Errorf("cel: expected 'in' after 'not', got %q", p.tok.text)
|
||||
}
|
||||
p.advance()
|
||||
right, err := p.parsePrimary()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
member, err := inMember(left, right)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return !member, nil
|
||||
}
|
||||
// "in"
|
||||
if p.tok.kind == tokIn {
|
||||
p.advance()
|
||||
right, err := p.parsePrimary()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return inMember(left, right)
|
||||
}
|
||||
// comparison operators
|
||||
if p.tok.kind == tokOp {
|
||||
op := p.tok.text
|
||||
switch op {
|
||||
case "==", "!=", ">=", "<=", ">", "<":
|
||||
p.advance()
|
||||
right, err := p.parsePrimary()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return compare(op, left, right)
|
||||
default:
|
||||
return nil, fmt.Errorf("cel: unexpected operator %q", op)
|
||||
}
|
||||
}
|
||||
return left, nil
|
||||
}
|
||||
|
||||
// inMember reports whether left is a member of right. right must be a
|
||||
// list ([]celValue) of comparable values; left may be a string or
|
||||
// int64.
|
||||
func inMember(left, right celValue) (bool, error) {
|
||||
list, ok := right.([]celValue)
|
||||
if !ok {
|
||||
return false, fmt.Errorf("cel: 'in' rhs not a list: %T", right)
|
||||
}
|
||||
for _, e := range list {
|
||||
if valuesEqual(left, e) {
|
||||
return true, nil
|
||||
}
|
||||
}
|
||||
return false, nil
|
||||
}
|
||||
|
||||
func valuesEqual(a, b celValue) bool {
|
||||
switch av := a.(type) {
|
||||
case string:
|
||||
bv, ok := b.(string)
|
||||
return ok && av == bv
|
||||
case int64:
|
||||
bv, ok := b.(int64)
|
||||
return ok && av == bv
|
||||
case bool:
|
||||
bv, ok := b.(bool)
|
||||
return ok && av == bv
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// compare applies a binary comparison operator to two scalar values.
|
||||
// Strings compare lexicographically; ints numerically; bools only via
|
||||
// ==/!=.
|
||||
func compare(op string, left, right celValue) (bool, error) {
|
||||
switch op {
|
||||
case "==":
|
||||
return valuesEqual(left, right), nil
|
||||
case "!=":
|
||||
return !valuesEqual(left, right), nil
|
||||
}
|
||||
// ordered comparisons require ordered operands
|
||||
ls, lok := left.(string)
|
||||
rs, rok := right.(string)
|
||||
if lok && rok {
|
||||
switch op {
|
||||
case "<":
|
||||
return ls < rs, nil
|
||||
case "<=":
|
||||
return ls <= rs, nil
|
||||
case ">":
|
||||
return ls > rs, nil
|
||||
case ">=":
|
||||
return ls >= rs, nil
|
||||
}
|
||||
}
|
||||
li, lok := left.(int64)
|
||||
ri, rok := right.(int64)
|
||||
if lok && rok {
|
||||
switch op {
|
||||
case "<":
|
||||
return li < ri, nil
|
||||
case "<=":
|
||||
return li <= ri, nil
|
||||
case ">":
|
||||
return li > ri, nil
|
||||
case ">=":
|
||||
return li >= ri, nil
|
||||
}
|
||||
}
|
||||
return false, fmt.Errorf("cel: cannot apply %q to %T and %T", op, left, right)
|
||||
}
|
||||
|
||||
// parsePrimary parses the smallest standalone unit: parenthesised
|
||||
// expressions, literals, and attribute references.
|
||||
func (p *parser) parsePrimary() (celValue, error) {
|
||||
switch p.tok.kind {
|
||||
case tokOp:
|
||||
if p.tok.text == "(" {
|
||||
p.advance()
|
||||
v, err := p.parseExpr()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if p.tok.kind != tokOp || p.tok.text != ")" {
|
||||
return nil, fmt.Errorf("cel: expected ')' got %q", p.tok.text)
|
||||
}
|
||||
p.advance()
|
||||
return v, nil
|
||||
}
|
||||
return nil, fmt.Errorf("cel: unexpected operator %q", p.tok.text)
|
||||
case tokInt:
|
||||
n, err := strconv.ParseInt(p.tok.text, 10, 64)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("cel: bad int %q: %w", p.tok.text, err)
|
||||
}
|
||||
p.advance()
|
||||
return n, nil
|
||||
case tokStr:
|
||||
v := p.tok.text
|
||||
p.advance()
|
||||
return v, nil
|
||||
case tokIdent:
|
||||
return p.parseAttrRef()
|
||||
}
|
||||
return nil, fmt.Errorf("cel: unexpected token %q", p.tok.text)
|
||||
}
|
||||
|
||||
// parseAttrRef resolves a bare or `node.<field>` attribute reference
|
||||
// against the node being evaluated. Bare identifiers (e.g. `region`)
|
||||
// resolve against the same attribute map as `node.region`; the PRD
|
||||
// examples use both forms interchangeably (see
|
||||
// TestParseMarkdown_ConstraintsInlineArray).
|
||||
func (p *parser) parseAttrRef() (celValue, error) {
|
||||
name := p.tok.text
|
||||
p.advance()
|
||||
// dotted access: node.<field>
|
||||
if p.tok.kind == tokOp && p.tok.text == "." {
|
||||
if name != "node" {
|
||||
return nil, fmt.Errorf("cel: dotted access on non-node: %q", name)
|
||||
}
|
||||
p.advance()
|
||||
if p.tok.kind != tokIdent {
|
||||
return nil, fmt.Errorf("cel: expected attribute name after '.', got %q", p.tok.text)
|
||||
}
|
||||
field := p.tok.text
|
||||
p.advance()
|
||||
return p.nodeAttr(name + "." + field)
|
||||
}
|
||||
// bare identifier
|
||||
switch name {
|
||||
case "true":
|
||||
return true, nil
|
||||
case "false":
|
||||
return false, nil
|
||||
default:
|
||||
return p.nodeAttr(name)
|
||||
}
|
||||
}
|
||||
|
||||
// nodeAttr resolves an attribute name to its value on the parser's
|
||||
// active node. Mapping (per PRD T2):
|
||||
//
|
||||
// node.hostname -> Hostname (string)
|
||||
// node.kind -> Kind (string)
|
||||
// node.cpus -> CPU (int64)
|
||||
// node.memory -> Memory (int64)
|
||||
// node.tags -> Tags ([]string -> []celValue)
|
||||
// node.runtimes -> Runtimes ([]string -> []celValue)
|
||||
//
|
||||
// Bare names (without the `node.` prefix) resolve through the same
|
||||
// map, so `region == "us"` and `node.region == "us"` are equivalent
|
||||
// when the attribute exists.
|
||||
func (p *parser) nodeAttr(name string) (celValue, error) {
|
||||
switch name {
|
||||
case "node.hostname", "hostname":
|
||||
return p.node.Hostname, nil
|
||||
case "node.kind", "kind":
|
||||
return p.node.Kind, nil
|
||||
case "node.cpus", "cpus":
|
||||
return p.node.CPU, nil
|
||||
case "node.memory", "memory":
|
||||
return p.node.Memory, nil
|
||||
case "node.tags", "tags":
|
||||
return toStringValues(p.node.Tags), nil
|
||||
case "node.runtimes", "runtimes":
|
||||
return toStringValues(p.node.Runtimes), nil
|
||||
}
|
||||
return nil, fmt.Errorf("cel: unknown attribute %q", name)
|
||||
}
|
||||
|
||||
// toStringValues converts a []string to []celValue so the membership
|
||||
// operators can compare element-wise.
|
||||
func toStringValues(in []string) []celValue {
|
||||
out := make([]celValue, len(in))
|
||||
for i, s := range in {
|
||||
out[i] = s
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,210 @@
|
||||
package scheduler
|
||||
|
||||
import "testing"
|
||||
|
||||
func TestEvaluateConstraint_Equality(t *testing.T) {
|
||||
node := NodeInfo{Hostname: "h-1", Kind: "linux", CPU: 4, Memory: 4096, Tags: []string{"web"}, Runtimes: []string{"process"}}
|
||||
cases := []struct {
|
||||
name string
|
||||
expr string
|
||||
want bool
|
||||
}{
|
||||
{"hostname eq", `node.hostname == "h-1"`, true},
|
||||
{"hostname ne", `node.hostname == "h-2"`, false},
|
||||
{"kind eq", `node.kind == "linux"`, true},
|
||||
{"kind ne", `node.kind == "proxmox"`, false},
|
||||
{"cpus eq", `node.cpus == 4`, true},
|
||||
{"memory eq", `node.memory == 4096`, true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got, err := EvaluateConstraint(c.expr, node)
|
||||
if err != nil {
|
||||
t.Errorf("%s: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got != c.want {
|
||||
t.Errorf("%s: got %v, want %v", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestEvaluateConstraint_Comparison(t *testing.T) {
|
||||
node := NodeInfo{Hostname: "h", Kind: "linux", CPU: 4, Memory: 4096}
|
||||
cases := []struct {
|
||||
expr string
|
||||
want bool
|
||||
}{
|
||||
{"node.cpus >= 2", true},
|
||||
{"node.cpus >= 4", true},
|
||||
{"node.cpus > 4", false},
|
||||
{"node.cpus > 2", true},
|
||||
{"node.cpus <= 4", true},
|
||||
{"node.cpus < 2", false},
|
||||
{"node.cpus != 8", true},
|
||||
{"node.cpus == 8", false},
|
||||
{"node.memory >= 2048", true},
|
||||
{"node.memory < 1024", false},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got, err := EvaluateConstraint(c.expr, node)
|
||||
if err != nil {
|
||||
t.Errorf("%q: %v", c.expr, err)
|
||||
continue
|
||||
}
|
||||
if got != c.want {
|
||||
t.Errorf("%q: got %v, want %v", c.expr, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestEvaluateConstraint_Membership(t *testing.T) {
|
||||
node := NodeInfo{Tags: []string{"web", "log-shipper"}, Runtimes: []string{"process", "wasmtime"}}
|
||||
cases := []struct {
|
||||
expr string
|
||||
want bool
|
||||
}{
|
||||
{`"web" in node.tags`, true},
|
||||
{`"missing" in node.tags`, false},
|
||||
{`"process" in node.runtimes`, true},
|
||||
{`"podman" in node.runtimes`, false},
|
||||
{`"log-shipper" not in node.tags`, false},
|
||||
{`"missing" not in node.tags`, true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got, err := EvaluateConstraint(c.expr, node)
|
||||
if err != nil {
|
||||
t.Errorf("%q: %v", c.expr, err)
|
||||
continue
|
||||
}
|
||||
if got != c.want {
|
||||
t.Errorf("%q: got %v, want %v", c.expr, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestEvaluateConstraint_BooleanComposition(t *testing.T) {
|
||||
node := NodeInfo{Kind: "linux", CPU: 4, Tags: []string{"web"}}
|
||||
cases := []struct {
|
||||
expr string
|
||||
want bool
|
||||
}{
|
||||
{`node.kind == "linux" and node.cpus >= 2`, true},
|
||||
{`node.kind == "proxmox" and node.cpus >= 2`, false},
|
||||
{`node.kind == "linux" or node.kind == "proxmox"`, true},
|
||||
{`node.kind == "proxmox" or node.kind == "linux"`, true},
|
||||
{`not node.kind == "proxmox"`, true},
|
||||
{`not node.kind == "linux"`, false},
|
||||
{`(node.kind == "linux") and (node.cpus >= 2)`, true},
|
||||
{`node.cpus >= 2 and not "blocked" in node.tags`, true},
|
||||
{`node.kind == "linux" and node.cpus >= 2 and "web" in node.tags`, true},
|
||||
{`node.kind == "linux" or node.kind == "proxmox" or node.cpus > 100`, true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got, err := EvaluateConstraint(c.expr, node)
|
||||
if err != nil {
|
||||
t.Errorf("%q: %v", c.expr, err)
|
||||
continue
|
||||
}
|
||||
if got != c.want {
|
||||
t.Errorf("%q: got %v, want %v", c.expr, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestEvaluateConstraint_BareIdentifiers(t *testing.T) {
|
||||
// Bare identifiers resolve through the same attribute map as
|
||||
// node.<field> (per PRD: constraints may use either form).
|
||||
node := NodeInfo{Kind: "linux", CPU: 4}
|
||||
got, err := EvaluateConstraint(`kind == "linux"`, node)
|
||||
if err != nil {
|
||||
t.Fatalf("bare kind: %v", err)
|
||||
}
|
||||
if !got {
|
||||
t.Error("bare kind == linux: got false, want true")
|
||||
}
|
||||
}
|
||||
|
||||
func TestEvaluateConstraint_TrueFalseLiterals(t *testing.T) {
|
||||
node := NodeInfo{}
|
||||
cases := []struct {
|
||||
expr string
|
||||
want bool
|
||||
}{
|
||||
{"true", true},
|
||||
{"false", false},
|
||||
{"not false", true},
|
||||
{"not true", false},
|
||||
{"true and true", true},
|
||||
{"true and false", false},
|
||||
{"false or true", true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got, err := EvaluateConstraint(c.expr, node)
|
||||
if err != nil {
|
||||
t.Errorf("%q: %v", c.expr, err)
|
||||
continue
|
||||
}
|
||||
if got != c.want {
|
||||
t.Errorf("%q: got %v, want %v", c.expr, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestEvaluateConstraint_Errors(t *testing.T) {
|
||||
node := NodeInfo{Kind: "linux"}
|
||||
cases := []struct {
|
||||
name string
|
||||
expr string
|
||||
}{
|
||||
{"empty", ""},
|
||||
{"unterminated string", `node.kind == "linux`},
|
||||
{"unknown attribute", `node.bogus == 1`},
|
||||
{"unknown bare attr", `bogus == 1`},
|
||||
{"dotted on non-node", `host.kind == "linux"`},
|
||||
{"bad operator", `node.cpus + 2`},
|
||||
{"trailing input", `node.kind == "linux" garbage`},
|
||||
{"unbalanced paren", `(node.kind == "linux"`},
|
||||
{"missing rhs", `node.cpus >=`},
|
||||
{"not without in", `"x" not node.tags`},
|
||||
{"ordered compare on bool", `true < false`},
|
||||
{"ordered compare on mismatched types", `node.kind > 2`},
|
||||
{"in on non-list", `"x" in node.kind`},
|
||||
}
|
||||
for _, c := range cases {
|
||||
_, err := EvaluateConstraint(c.expr, node)
|
||||
if err == nil {
|
||||
t.Errorf("%s: expected error for %q, got nil", c.name, c.expr)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestEvaluateAll(t *testing.T) {
|
||||
node := NodeInfo{Kind: "linux", CPU: 4, Tags: []string{"web"}}
|
||||
cases := []struct {
|
||||
name string
|
||||
constraints []string
|
||||
want bool
|
||||
}{
|
||||
{"empty", nil, true},
|
||||
{"all pass", []string{`node.kind == "linux"`, "node.cpus >= 2"}, true},
|
||||
{"one fails", []string{`node.kind == "linux"`, "node.cpus >= 8"}, false},
|
||||
{"all fail", []string{`node.kind == "proxmox"`, "node.cpus >= 8"}, false},
|
||||
}
|
||||
for _, c := range cases {
|
||||
got, err := EvaluateAll(c.constraints, node)
|
||||
if err != nil {
|
||||
t.Errorf("%s: %v", c.name, err)
|
||||
continue
|
||||
}
|
||||
if got != c.want {
|
||||
t.Errorf("%s: got %v, want %v", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestEvaluateAll_PropagatesError(t *testing.T) {
|
||||
node := NodeInfo{}
|
||||
if _, err := EvaluateAll([]string{"bogus == 1"}, node); err == nil {
|
||||
t.Error("EvaluateAll: expected error for malformed constraint")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,472 @@
|
||||
// Package scheduler implements the v0.9 CLI-side scheduler (REQ-083,
|
||||
// P05). Unlike the v0.8 daemon-side best-fit scheduler
|
||||
// (internal/engine/scheduler.go), this scheduler runs entirely in the
|
||||
// `orca` CLI process (R-001) and is pure: it takes a list of candidate
|
||||
// nodes plus a workload request and returns placement decisions
|
||||
// without performing any I/O.
|
||||
//
|
||||
// The scheduler is runtime-aware: a workload that declares
|
||||
// `runtime.one_of: wasm` is only placed on nodes that expose
|
||||
// `wasmtime` in their Runtimes list; a `pve-vm` workload is only
|
||||
// placed on `proxmox` nodes. It is also constraint- and
|
||||
// affinity-aware via the CEL-subset evaluator in cel.go.
|
||||
//
|
||||
// Workload kinds are handled differently per the PRD:
|
||||
//
|
||||
// - Job: one-shot, returns exactly one placement (best-fit
|
||||
// bin-packing).
|
||||
// - Service: count replicas spread across distinct nodes
|
||||
// (anti-affinity by default); if fewer distinct nodes than count,
|
||||
// colocation is permitted but distinct nodes are preferred.
|
||||
// - DaemonSet: one placement per node that fits the constraints.
|
||||
package scheduler
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"sort"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// NodeInfo is the scheduler's projection of a peer node: total and
|
||||
// free capacity, the runtimes the node advertises, its tags, and its
|
||||
// kind (linux/proxmox). The CLI populates this from the
|
||||
// cluster/peers/ inventory plus the per-node capacity reports
|
||||
// collected over SSH; the scheduler itself never reads either.
|
||||
type NodeInfo struct {
|
||||
Hostname string
|
||||
Runtimes []string
|
||||
Tags []string
|
||||
CPU int64
|
||||
Memory int64
|
||||
FreeCPU int64
|
||||
FreeMem int64
|
||||
Kind string
|
||||
}
|
||||
|
||||
// WorkloadRequest bundles a parsed WorkloadSpec with the namespace
|
||||
// the workload is being scheduled into. The namespace is carried
|
||||
// through to placement so the resulting AllocID can be namespaced,
|
||||
// but the scheduler itself does not inspect it for fitting decisions.
|
||||
type WorkloadRequest struct {
|
||||
Spec *jobspec.WorkloadSpec
|
||||
Namespace string
|
||||
}
|
||||
|
||||
// Placement is a single scheduling decision: which node, which
|
||||
// allocation id, and the bin-packing score that won the node the
|
||||
// placement. AllocID is `ns/spec.Name-<idx>` so a multi-replica
|
||||
// Service produces distinct ids per replica.
|
||||
type Placement struct {
|
||||
Node string
|
||||
AllocID string
|
||||
Score int64
|
||||
}
|
||||
|
||||
// Schedule is the main entry point. For Job (kind=Job) it returns one
|
||||
// placement on the best-fit node. For Service it returns `Count`
|
||||
// placements spread across distinct nodes where possible (anti-
|
||||
// affinity), permitting colocation when Count > nodes. For DaemonSet
|
||||
// it returns one placement per node that fits. Any kind-agnostic
|
||||
// validation error (no spec, unknown kind, no fitting node) is
|
||||
// returned as an error rather than an empty slice so callers can
|
||||
// distinguish "nothing fits" from "scheduled zero replicas".
|
||||
func Schedule(nodes []NodeInfo, req WorkloadRequest) ([]Placement, error) {
|
||||
if req.Spec == nil {
|
||||
return nil, fmt.Errorf("scheduler: nil WorkloadSpec")
|
||||
}
|
||||
if len(nodes) == 0 {
|
||||
return nil, fmt.Errorf("scheduler: no candidate nodes")
|
||||
}
|
||||
switch req.Spec.Kind {
|
||||
case "Job":
|
||||
return scheduleJob(nodes, req)
|
||||
case "Service":
|
||||
return scheduleService(nodes, req)
|
||||
case "DaemonSet":
|
||||
return scheduleDaemonSet(nodes, req)
|
||||
default:
|
||||
return nil, fmt.Errorf("scheduler: unknown kind %q", req.Spec.Kind)
|
||||
}
|
||||
}
|
||||
|
||||
// Score evaluates a single node against a workload. fits is true iff
|
||||
// the node (a) advertises a runtime compatible with the workload's
|
||||
// `runtime.one_of`, (b) satisfies every CEL constraint in
|
||||
// `spec.Constraints`, and (c) has enough free CPU+memory for the
|
||||
// workload's requested resources. When fits is true, score is the
|
||||
// bin-packing score (more free capacity = higher score, so the node
|
||||
// most likely to absorb the workload without starving its
|
||||
// neighbours wins). When fits is false, score is 0.
|
||||
func Score(node NodeInfo, req WorkloadRequest) (score int64, fits bool) {
|
||||
// (a) runtime compatibility. A workload with no Runtime block or
|
||||
// an empty OneOf is treated as runtime-agnostic (always fits on
|
||||
// the runtime axis); this matches the v0.8 behaviour where a
|
||||
// missing runtime meant "process".
|
||||
runtimeOK := true
|
||||
if req.Spec != nil && req.Spec.Runtime != nil && req.Spec.Runtime.OneOf != "" {
|
||||
runtimeOK = hasRuntime(node, req.Spec.Runtime.OneOf)
|
||||
}
|
||||
if !runtimeOK {
|
||||
return 0, false
|
||||
}
|
||||
// (b) constraints. Evaluation errors are treated as non-fit so a
|
||||
// malformed constraint does not crash Schedule; the caller still
|
||||
// sees the node filtered out.
|
||||
if req.Spec != nil {
|
||||
ok, err := EvaluateAll(req.Spec.Constraints, node)
|
||||
if err != nil || !ok {
|
||||
return 0, false
|
||||
}
|
||||
}
|
||||
// (c) capacity. A workload with no Resources block is treated as
|
||||
// zero-sized for fitting purposes (it always fits the capacity
|
||||
// axis); real workloads declare cpu/memory.
|
||||
needCPU, needMem := workloadResources(req)
|
||||
if node.FreeCPU < needCPU || node.FreeMem < needMem {
|
||||
return 0, false
|
||||
}
|
||||
// bin-packing score: most free capacity wins. CPU is weighted
|
||||
// 1000x memory so a 1-core difference outweighs a 1-MiB
|
||||
// difference, mirroring the v0.8 Score weighting that biased
|
||||
// toward CPU (the more common binding constraint).
|
||||
score = (node.FreeCPU-needCPU)*1000 + (node.FreeMem - needMem)
|
||||
if score < 0 {
|
||||
score = 0
|
||||
}
|
||||
return score, true
|
||||
}
|
||||
|
||||
// hasRuntime reports whether node advertises the requested runtime.
|
||||
// The match is case-insensitive and tolerant of aliases: `wasm` and
|
||||
// `wasmtime` are treated as the same runtime, and `pve-vm`/`pve-ct`
|
||||
// only match nodes whose Kind is "proxmox".
|
||||
func hasRuntime(node NodeInfo, oneOf string) bool {
|
||||
want := normalizeRuntime(oneOf)
|
||||
// pve-* runtimes require a proxmox-kind node regardless of the
|
||||
// node's Runtimes list (a proxmox node doesn't list "pve-vm" in
|
||||
// Runtimes; it IS the runtime).
|
||||
switch want {
|
||||
case "pve-vm", "pve-ct", "proxmox":
|
||||
return normalizeKind(node.Kind) == "proxmox"
|
||||
}
|
||||
for _, r := range node.Runtimes {
|
||||
if normalizeRuntime(r) == want {
|
||||
return true
|
||||
}
|
||||
// alias: wasmtime nodes advertise "wasmtime"; workloads ask
|
||||
// for "wasm".
|
||||
if want == "wasm" && normalizeRuntime(r) == "wasmtime" {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// normalizeRuntime lowercases and trims a runtime name for matching.
|
||||
func normalizeRuntime(s string) string {
|
||||
s = toLowerASCII(s)
|
||||
switch s {
|
||||
case "wasmtime":
|
||||
return "wasm"
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
// normalizeKind lowercases and trims a node Kind for matching.
|
||||
func normalizeKind(s string) string { return toLowerASCII(s) }
|
||||
|
||||
// toLowerASCII lowercases ASCII letters without bringing in strings
|
||||
// (avoid an alloc-heavy stdlib call in the hot path).
|
||||
func toLowerASCII(s string) string {
|
||||
b := []byte(s)
|
||||
for i, c := range b {
|
||||
if c >= 'A' && c <= 'Z' {
|
||||
b[i] = c + 32
|
||||
}
|
||||
}
|
||||
return string(b)
|
||||
}
|
||||
|
||||
// workloadResources returns the (cpu, memory) the workload requests,
|
||||
// read from the WorkloadSpec's Resources block if present. The
|
||||
// v0.9-P05 WorkloadSpec does not yet carry a Resources field (it
|
||||
// lands in P0c, REQ-074); until then this returns (0, 0) so the
|
||||
// capacity check is a no-op and runtime/constraints do the real
|
||||
// filtering. The signature is here so the scheduler logic does not
|
||||
// need to change when Resources lands.
|
||||
func workloadResources(req WorkloadRequest) (int64, int64) {
|
||||
_ = req
|
||||
return 0, 0
|
||||
}
|
||||
|
||||
// ----------------------------------------------------------------------------
|
||||
// Kind-specific scheduling
|
||||
// ----------------------------------------------------------------------------
|
||||
|
||||
// scheduleJob places a single Job on the best-fit node.
|
||||
func scheduleJob(nodes []NodeInfo, req WorkloadRequest) ([]Placement, error) {
|
||||
type cand struct {
|
||||
node NodeInfo
|
||||
score int64
|
||||
}
|
||||
var cands []cand
|
||||
for _, n := range nodes {
|
||||
s, ok := Score(n, req)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
cands = append(cands, cand{node: n, score: s})
|
||||
}
|
||||
// CEL-based affinity rules apply to single-shot Jobs too: a Job
|
||||
// with `affinity: [{target: "\"ssd\" in node.tags", weight: 100}]`
|
||||
// should land on the tagged node even without prior placements.
|
||||
// Name-based affinity (no prior placements to check) contributes
|
||||
// zero for a standalone Job, so it is harmless to call here.
|
||||
for i := range cands {
|
||||
cands[i].score += affinityScore(cands[i].node, req, nil)
|
||||
}
|
||||
if len(cands) == 0 {
|
||||
return nil, fmt.Errorf("scheduler: no node fits workload %q", req.Spec.Name)
|
||||
}
|
||||
sort.SliceStable(cands, func(i, j int) bool {
|
||||
if cands[i].score != cands[j].score {
|
||||
return cands[i].score > cands[j].score
|
||||
}
|
||||
return cands[i].node.Hostname < cands[j].node.Hostname
|
||||
})
|
||||
w := cands[0]
|
||||
return []Placement{{
|
||||
Node: w.node.Hostname,
|
||||
AllocID: allocID(req, 0),
|
||||
Score: w.score,
|
||||
}}, nil
|
||||
}
|
||||
|
||||
// scheduleService places `Count` replicas with implicit anti-affinity:
|
||||
// prefer distinct nodes, but permit colocation when Count exceeds the
|
||||
// number of fitting nodes. Each replica gets a distinct AllocID.
|
||||
func scheduleService(nodes []NodeInfo, req WorkloadRequest) ([]Placement, error) {
|
||||
count := req.Spec.Count
|
||||
if count <= 0 {
|
||||
count = 1
|
||||
}
|
||||
// Pre-filter fitting nodes once; the loop below re-scores them
|
||||
// after each placement so the capacity accounting reflects the
|
||||
// replicas already placed.
|
||||
fitting := filterFitting(nodes, req)
|
||||
if len(fitting) == 0 {
|
||||
return nil, fmt.Errorf("scheduler: no node fits service %q", req.Spec.Name)
|
||||
}
|
||||
var placements []Placement
|
||||
placed := map[string]int{} // hostname -> count placed there
|
||||
// First pass: spread across distinct nodes.
|
||||
for i := 0; i < count; i++ {
|
||||
best, score, ok := pickServiceNode(fitting, req, placements, placed)
|
||||
if !ok {
|
||||
break
|
||||
}
|
||||
placements = append(placements, Placement{
|
||||
Node: best.Hostname,
|
||||
AllocID: allocID(req, i),
|
||||
Score: score,
|
||||
})
|
||||
placed[best.Hostname]++
|
||||
// Reflect the consumed capacity in the candidate snapshot so
|
||||
// subsequent picks see updated free capacity.
|
||||
needCPU, needMem := workloadResources(req)
|
||||
for j := range fitting {
|
||||
if fitting[j].Hostname == best.Hostname {
|
||||
fitting[j].FreeCPU -= needCPU
|
||||
fitting[j].FreeMem -= needMem
|
||||
}
|
||||
}
|
||||
}
|
||||
if len(placements) < count {
|
||||
return nil, fmt.Errorf("scheduler: only placed %d/%d replicas for service %q",
|
||||
len(placements), count, req.Spec.Name)
|
||||
}
|
||||
return placements, nil
|
||||
}
|
||||
|
||||
// pickServiceNode selects the best node for the next replica. The
|
||||
// selection prefers nodes with zero prior placements of this service
|
||||
// (anti-affinity) and applies affinity scoring on top of the
|
||||
// bin-packing score.
|
||||
func pickServiceNode(fitting []NodeInfo, req WorkloadRequest, placements []Placement, placed map[string]int) (NodeInfo, int64, bool) {
|
||||
type scored struct {
|
||||
node NodeInfo
|
||||
score int64
|
||||
}
|
||||
var cands []scored
|
||||
for _, n := range fitting {
|
||||
s, ok := Score(n, req)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
// Implicit anti-affinity: a node with N prior replicas of this
|
||||
// service incurs a penalty of N * (1 << 62) so distinct nodes
|
||||
// are preferred, but colocation is permitted (with a
|
||||
// per-replica penalty) when no distinct node remains. This
|
||||
// produces a balanced spread (e.g. 5 replicas on 3 nodes →
|
||||
// 2/2/1) rather than stacking everything on the first node.
|
||||
if placed[n.Hostname] > 0 {
|
||||
s -= int64(placed[n.Hostname]) * (1 << 62)
|
||||
}
|
||||
// Affinity rules from the spec add/subtract their weight.
|
||||
s += affinityScore(n, req, placements)
|
||||
cands = append(cands, scored{node: n, score: s})
|
||||
}
|
||||
if len(cands) == 0 {
|
||||
return NodeInfo{}, 0, false
|
||||
}
|
||||
sort.SliceStable(cands, func(i, j int) bool {
|
||||
if cands[i].score != cands[j].score {
|
||||
return cands[i].score > cands[j].score
|
||||
}
|
||||
return cands[i].node.Hostname < cands[j].node.Hostname
|
||||
})
|
||||
w := cands[0]
|
||||
return w.node, w.score, true
|
||||
}
|
||||
|
||||
// scheduleDaemonSet places one replica per node that fits the
|
||||
// constraints. The PRD's DaemonSet placement mode (every-node /
|
||||
// matching / mandatory) lives on the WorkloadSpec.Schedule block; the
|
||||
// scheduler honours it indirectly by filtering on Constraints: a
|
||||
// `matching` DaemonSet carries constraints that select the matching
|
||||
// nodes, an `every-node` DaemonSet carries none, and a `mandatory`
|
||||
// one is enforced elsewhere (the scheduler still just returns
|
||||
// placements for every fitting node).
|
||||
func scheduleDaemonSet(nodes []NodeInfo, req WorkloadRequest) ([]Placement, error) {
|
||||
var placements []Placement
|
||||
for _, n := range nodes {
|
||||
s, ok := Score(n, req)
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
placements = append(placements, Placement{
|
||||
Node: n.Hostname,
|
||||
AllocID: allocID(req, len(placements)),
|
||||
Score: s,
|
||||
})
|
||||
}
|
||||
if len(placements) == 0 {
|
||||
return nil, fmt.Errorf("scheduler: no node fits daemonset %q", req.Spec.Name)
|
||||
}
|
||||
return placements, nil
|
||||
}
|
||||
|
||||
// ----------------------------------------------------------------------------
|
||||
// Affinity scoring
|
||||
// ----------------------------------------------------------------------------
|
||||
|
||||
// affinityScore returns the weighted affinity contribution for a
|
||||
// node given the placements already made. For each AffinityRule the
|
||||
// Target is a CEL expression; if it evaluates true against the node,
|
||||
// the rule's Weight is added (positive = co-locate, negative =
|
||||
// anti-affinity). An affinity target that fails to evaluate is
|
||||
// ignored rather than failing the schedule: operators use affinity as
|
||||
// a hint, not a hard gate.
|
||||
//
|
||||
// The PRD also mentions affinity rules like `{target: "redis", weight:
|
||||
// 50}` where Target is a workload *name* rather than a CEL expression.
|
||||
// We support both: if Target parses as a CEL expression it is
|
||||
// evaluated against the node; otherwise it is treated as a workload
|
||||
// name and we check whether any already-placed alloc for that name
|
||||
// exists on the node. The placement-already-here check is done by the
|
||||
// caller via placements; this function checks the node's own
|
||||
// attributes only.
|
||||
func affinityScore(node NodeInfo, req WorkloadRequest, placements []Placement) int64 {
|
||||
if req.Spec == nil {
|
||||
return 0
|
||||
}
|
||||
var total int64
|
||||
for _, rule := range req.Spec.Affinity {
|
||||
// Try CEL evaluation first; if the target is a bare workload
|
||||
// name (no operator) the CEL parser will fail and we fall
|
||||
// back to name-based placement counting.
|
||||
ok, err := EvaluateConstraint(rule.Target, node)
|
||||
if err == nil {
|
||||
if ok {
|
||||
total += int64(rule.Weight)
|
||||
}
|
||||
continue
|
||||
}
|
||||
// Fallback: target is a workload name; count existing
|
||||
// placements for that workload on this node and apply the
|
||||
// weight once per co-located replica.
|
||||
for _, p := range placements {
|
||||
if p.Node == node.Hostname && isAllocFor(p.AllocID, rule.Target) {
|
||||
total += int64(rule.Weight)
|
||||
}
|
||||
}
|
||||
}
|
||||
return total
|
||||
}
|
||||
|
||||
// isAllocFor reports whether an AllocID encodes a placement for the
|
||||
// named workload. AllocIDs are `ns/name-idx`, so we look for the
|
||||
// workload name as the segment after the first slash and before the
|
||||
// trailing `-idx`.
|
||||
func isAllocFor(allocID, workloadName string) bool {
|
||||
// strip namespace prefix
|
||||
rest := allocID
|
||||
if i := indexByte(rest, '/'); i >= 0 {
|
||||
rest = rest[i+1:]
|
||||
}
|
||||
// strip trailing -idx
|
||||
if i := lastIndexByte(rest, '-'); i >= 0 {
|
||||
rest = rest[:i]
|
||||
}
|
||||
return rest == workloadName
|
||||
}
|
||||
|
||||
// indexByte returns the index of the first occurrence of b in s, or
|
||||
// -1. Avoids importing strings just for one helper.
|
||||
func indexByte(s string, b byte) int {
|
||||
for i := 0; i < len(s); i++ {
|
||||
if s[i] == b {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
|
||||
// lastIndexByte returns the index of the last occurrence of b in s, or
|
||||
// -1.
|
||||
func lastIndexByte(s string, b byte) int {
|
||||
for i := len(s) - 1; i >= 0; i-- {
|
||||
if s[i] == b {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
|
||||
// ----------------------------------------------------------------------------
|
||||
// Helpers
|
||||
// ----------------------------------------------------------------------------
|
||||
|
||||
// filterFitting returns a copy of the nodes that pass Score for the
|
||||
// request, preserving order. Capacity is not yet decremented; the
|
||||
// caller adjusts FreeCPU/FreeMem as it places replicas.
|
||||
func filterFitting(nodes []NodeInfo, req WorkloadRequest) []NodeInfo {
|
||||
var out []NodeInfo
|
||||
for _, n := range nodes {
|
||||
if _, ok := Score(n, req); ok {
|
||||
out = append(out, n)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// allocID renders a stable, namespaced allocation id for a placement.
|
||||
// Format: `ns/spec.Name-<idx>`.
|
||||
func allocID(req WorkloadRequest, idx int) string {
|
||||
ns := req.Namespace
|
||||
if ns == "" {
|
||||
ns = "default"
|
||||
}
|
||||
return fmt.Sprintf("%s/%s-%d", ns, req.Spec.Name, idx)
|
||||
}
|
||||
@@ -0,0 +1,451 @@
|
||||
package scheduler
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// threeLinuxNodes returns a small cluster of three Linux nodes with
|
||||
// distinct free capacities so best-fit ordering is unambiguous.
|
||||
func threeLinuxNodes() []NodeInfo {
|
||||
return []NodeInfo{
|
||||
{Hostname: "node-a", Runtimes: []string{"process"}, Tags: nil, CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096, Kind: "linux"},
|
||||
{Hostname: "node-b", Runtimes: []string{"process"}, Tags: nil, CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192, Kind: "linux"},
|
||||
{Hostname: "node-c", Runtimes: []string{"process"}, Tags: nil, CPU: 2, Memory: 2048, FreeCPU: 2, FreeMem: 2048, Kind: "linux"},
|
||||
}
|
||||
}
|
||||
|
||||
func jobSpec(name, oneOf string, constraints []string) *jobspec.WorkloadSpec {
|
||||
return &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: name,
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: oneOf},
|
||||
Constraints: constraints,
|
||||
}
|
||||
}
|
||||
|
||||
func serviceSpec(name, oneOf string, count int, constraints []string) *jobspec.WorkloadSpec {
|
||||
return &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: name,
|
||||
Count: count,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: oneOf},
|
||||
Constraints: constraints,
|
||||
}
|
||||
}
|
||||
|
||||
func daemonSetSpec(name, oneOf string, constraints []string) *jobspec.WorkloadSpec {
|
||||
return &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: name,
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: oneOf},
|
||||
Constraints: constraints,
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Job
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestScheduleJob_BestFit(t *testing.T) {
|
||||
nodes := threeLinuxNodes()
|
||||
req := WorkloadRequest{Spec: jobSpec("batch", "process", nil), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if len(got) != 1 {
|
||||
t.Fatalf("placements = %d, want 1", len(got))
|
||||
}
|
||||
if got[0].Node != "node-b" {
|
||||
t.Errorf("Node = %q, want node-b (most free capacity)", got[0].Node)
|
||||
}
|
||||
if !strings.HasPrefix(got[0].AllocID, "ns/batch-") {
|
||||
t.Errorf("AllocID = %q, want ns/batch-*", got[0].AllocID)
|
||||
}
|
||||
if got[0].Score <= 0 {
|
||||
t.Errorf("Score = %d, want > 0", got[0].Score)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScheduleJob_NoFittingNode(t *testing.T) {
|
||||
nodes := threeLinuxNodes()
|
||||
// wasm runtime not advertised by any node.
|
||||
req := WorkloadRequest{Spec: jobSpec("wasmjob", "wasm", nil), Namespace: "ns"}
|
||||
if _, err := Schedule(nodes, req); err == nil {
|
||||
t.Fatal("Schedule: expected error for no-fitting node, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Service
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestScheduleService_SpreadAcrossNodes(t *testing.T) {
|
||||
nodes := threeLinuxNodes()
|
||||
req := WorkloadRequest{Spec: serviceSpec("web", "process", 3, nil), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if len(got) != 3 {
|
||||
t.Fatalf("placements = %d, want 3", len(got))
|
||||
}
|
||||
seen := map[string]int{}
|
||||
for _, p := range got {
|
||||
seen[p.Node]++
|
||||
}
|
||||
if len(seen) != 3 {
|
||||
t.Errorf("anti-affinity spread: distinct nodes = %d, want 3; %v", len(seen), seen)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScheduleService_ColocationWhenFewerNodes(t *testing.T) {
|
||||
nodes := threeLinuxNodes()
|
||||
req := WorkloadRequest{Spec: serviceSpec("web", "process", 5, nil), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if len(got) != 5 {
|
||||
t.Fatalf("placements = %d, want 5", len(got))
|
||||
}
|
||||
seen := map[string]int{}
|
||||
for _, p := range got {
|
||||
seen[p.Node]++
|
||||
}
|
||||
if len(seen) != 3 {
|
||||
t.Errorf("colocation: distinct nodes = %d, want 3 (all used)", len(seen))
|
||||
}
|
||||
// No node should host more than 2 (3 nodes, 5 replicas: 2+2+1).
|
||||
for n, c := range seen {
|
||||
if c > 2 {
|
||||
t.Errorf("node %s has %d replicas, want <= 2", n, c)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestScheduleService_NoFittingNode(t *testing.T) {
|
||||
nodes := threeLinuxNodes()
|
||||
req := WorkloadRequest{Spec: serviceSpec("wasm-svc", "wasm", 3, nil), Namespace: "ns"}
|
||||
if _, err := Schedule(nodes, req); err == nil {
|
||||
t.Fatal("Schedule: expected error for service with no fitting node")
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// DaemonSet
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestScheduleDaemonSet_AllMatching(t *testing.T) {
|
||||
nodes := threeLinuxNodes()
|
||||
req := WorkloadRequest{Spec: daemonSetSpec("logrotate", "process", nil), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if len(got) != 3 {
|
||||
t.Errorf("placements = %d, want 3 (one per node)", len(got))
|
||||
}
|
||||
seen := map[string]bool{}
|
||||
for _, p := range got {
|
||||
seen[p.Node] = true
|
||||
}
|
||||
if len(seen) != 3 {
|
||||
t.Errorf("DaemonSet distinct nodes = %d, want 3", len(seen))
|
||||
}
|
||||
}
|
||||
|
||||
func TestScheduleDaemonSet_SomeExcludedByConstraint(t *testing.T) {
|
||||
nodes := threeLinuxNodes()
|
||||
// Only nodes with cpus >= 4 qualify: node-a (4) and node-b (8).
|
||||
req := WorkloadRequest{Spec: daemonSetSpec("heavy", "process", []string{"node.cpus >= 4"}), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if len(got) != 2 {
|
||||
t.Errorf("placements = %d, want 2 (cpus>=4)", len(got))
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Runtime compatibility
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestSchedule_RuntimeCompatibilityWasm(t *testing.T) {
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "no-wasm", Runtimes: []string{"process"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
|
||||
{Hostname: "has-wasm", Runtimes: []string{"process", "wasmtime"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
|
||||
}
|
||||
// Even though no-wasm has more free capacity, the wasm workload
|
||||
// must land on has-wasm.
|
||||
req := WorkloadRequest{Spec: jobSpec("wasmjob", "wasm", nil), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if got[0].Node != "has-wasm" {
|
||||
t.Errorf("Node = %q, want has-wasm (runtime compatibility)", got[0].Node)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSchedule_RuntimeCompatibilityPveVM(t *testing.T) {
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "linux-1", Runtimes: []string{"process"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
|
||||
{Hostname: "pve-1", Runtimes: []string{"process"}, Kind: "proxmox", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
|
||||
}
|
||||
req := WorkloadRequest{Spec: jobSpec("vmjob", "pve-vm", nil), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if got[0].Node != "pve-1" {
|
||||
t.Errorf("Node = %q, want pve-1 (pve-vm requires proxmox kind)", got[0].Node)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Constraints
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestSchedule_ConstraintKindExcludesProxmox(t *testing.T) {
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "linux-1", Runtimes: []string{"process"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
|
||||
{Hostname: "pve-1", Runtimes: []string{"process"}, Kind: "proxmox", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
|
||||
}
|
||||
req := WorkloadRequest{Spec: jobSpec("linuxonly", "process", []string{`node.kind == "linux"`}), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if got[0].Node != "linux-1" {
|
||||
t.Errorf("Node = %q, want linux-1 (kind==linux)", got[0].Node)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSchedule_ConstraintCPUsExcludesSmall(t *testing.T) {
|
||||
nodes := threeLinuxNodes() // node-c has cpus=2
|
||||
req := WorkloadRequest{Spec: jobSpec("big", "process", []string{"node.cpus >= 4"}), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if got[0].Node == "node-c" {
|
||||
t.Errorf("Node = node-c, want node-a or node-b (cpus>=4)")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSchedule_ConstraintNotInTags(t *testing.T) {
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "tagged", Runtimes: []string{"process"}, Tags: []string{"log-shipper"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
|
||||
{Hostname: "clean", Runtimes: []string{"process"}, Tags: nil, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
|
||||
}
|
||||
req := WorkloadRequest{Spec: jobSpec("worker", "process", []string{`"log-shipper" not in node.tags`}), Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if got[0].Node != "clean" {
|
||||
t.Errorf("Node = %q, want clean (log-shipper not in tags)", got[0].Node)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Affinity
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestSchedule_AffinityPrefersColocatedNode(t *testing.T) {
|
||||
// Place a redis service first, then a worker with affinity for
|
||||
// redis; the worker should prefer the node where redis already
|
||||
// runs even if another node has more free capacity.
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "big", Runtimes: []string{"process"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
|
||||
{Hostname: "small", Runtimes: []string{"process"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
|
||||
}
|
||||
redisReq := WorkloadRequest{Spec: serviceSpec("redis", "process", 1, nil), Namespace: "ns"}
|
||||
redisPlacements, err := Schedule(nodes, redisReq)
|
||||
if err != nil {
|
||||
t.Fatalf("redis Schedule: %v", err)
|
||||
}
|
||||
// Redis lands on "big" (most free capacity). Now schedule the
|
||||
// worker with affinity to redis; it should also land on "big".
|
||||
workerReq := WorkloadRequest{Spec: &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "worker",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Affinity: []jobspec.AffinityRule{
|
||||
{Target: "redis", Weight: 1000},
|
||||
},
|
||||
}, Namespace: "ns"}
|
||||
// The affinity is name-based; we need to seed the worker schedule
|
||||
// with the redis placement so affinityScore can see it. Schedule
|
||||
// does not take prior placements, so test affinityScore directly.
|
||||
got := affinityScore(nodes[0], workerReq, redisPlacements)
|
||||
if got <= 0 {
|
||||
t.Errorf("affinityScore(big) = %d, want > 0 (redis colocated)", got)
|
||||
}
|
||||
gotSmall := affinityScore(nodes[1], workerReq, redisPlacements)
|
||||
if gotSmall != 0 {
|
||||
t.Errorf("affinityScore(small) = %d, want 0 (redis not colocated)", gotSmall)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSchedule_AffinityCELExpression(t *testing.T) {
|
||||
// Affinity with a CEL target: prefer nodes tagged "ssd".
|
||||
nodes := []NodeInfo{
|
||||
{Hostname: "hdd", Runtimes: []string{"process"}, Tags: []string{"hdd"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
|
||||
{Hostname: "ssd", Runtimes: []string{"process"}, Tags: []string{"ssd"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
|
||||
}
|
||||
req := WorkloadRequest{Spec: &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "db",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Affinity: []jobspec.AffinityRule{
|
||||
{Target: `"ssd" in node.tags`, Weight: 10000},
|
||||
},
|
||||
}, Namespace: "ns"}
|
||||
got, err := Schedule(nodes, req)
|
||||
if err != nil {
|
||||
t.Fatalf("Schedule: %v", err)
|
||||
}
|
||||
if got[0].Node != "ssd" {
|
||||
t.Errorf("Node = %q, want ssd (affinity to ssd tag outweighs capacity)", got[0].Node)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Error paths
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestSchedule_EmptyNodes(t *testing.T) {
|
||||
req := WorkloadRequest{Spec: jobSpec("x", "process", nil), Namespace: "ns"}
|
||||
if _, err := Schedule(nil, req); err == nil {
|
||||
t.Fatal("Schedule: expected error for empty nodes, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSchedule_NilSpec(t *testing.T) {
|
||||
if _, err := Schedule(threeLinuxNodes(), WorkloadRequest{}); err == nil {
|
||||
t.Fatal("Schedule: expected error for nil spec, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestSchedule_UnknownKind(t *testing.T) {
|
||||
req := WorkloadRequest{Spec: &jobspec.WorkloadSpec{Kind: "Cron", Name: "x", Count: 1}, Namespace: "ns"}
|
||||
if _, err := Schedule(threeLinuxNodes(), req); err == nil {
|
||||
t.Fatal("Schedule: expected error for unknown kind")
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Score unit tests
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestScore_FitsAndDoesNotFit(t *testing.T) {
|
||||
node := NodeInfo{Hostname: "n", Runtimes: []string{"process"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096}
|
||||
req := WorkloadRequest{Spec: jobSpec("j", "process", nil), Namespace: "ns"}
|
||||
score, fits := Score(node, req)
|
||||
if !fits {
|
||||
t.Error("fits = false, want true")
|
||||
}
|
||||
if score <= 0 {
|
||||
t.Errorf("score = %d, want > 0", score)
|
||||
}
|
||||
}
|
||||
|
||||
func TestScore_RuntimeMismatchDoesNotFit(t *testing.T) {
|
||||
node := NodeInfo{Hostname: "n", Runtimes: []string{"process"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096}
|
||||
req := WorkloadRequest{Spec: jobSpec("j", "wasm", nil), Namespace: "ns"}
|
||||
if _, fits := Score(node, req); fits {
|
||||
t.Error("fits = true for wasm on process-only node, want false")
|
||||
}
|
||||
}
|
||||
|
||||
func TestScore_ConstraintFailsDoesNotFit(t *testing.T) {
|
||||
node := NodeInfo{Hostname: "n", Runtimes: []string{"process"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096}
|
||||
req := WorkloadRequest{Spec: jobSpec("j", "process", []string{`node.kind == "proxmox"`}), Namespace: "ns"}
|
||||
if _, fits := Score(node, req); fits {
|
||||
t.Error("fits = true for kind==proxmox on linux node, want false")
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// allocID / isAllocFor helpers
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestAllocID(t *testing.T) {
|
||||
req := WorkloadRequest{Spec: &jobspec.WorkloadSpec{Name: "web"}, Namespace: "prod"}
|
||||
if got := allocID(req, 2); got != "prod/web-2" {
|
||||
t.Errorf("allocID = %q, want prod/web-2", got)
|
||||
}
|
||||
req.Namespace = ""
|
||||
if got := allocID(req, 0); got != "default/web-0" {
|
||||
t.Errorf("allocID = %q, want default/web-0", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIsAllocFor(t *testing.T) {
|
||||
cases := []struct {
|
||||
allocID string
|
||||
workload string
|
||||
want bool
|
||||
}{
|
||||
{"ns/redis-0", "redis", true},
|
||||
{"ns/redis-12", "redis", true},
|
||||
{"ns/worker-0", "redis", false},
|
||||
{"redis-0", "redis", true},
|
||||
{"ns/web-canary-3", "web-canary", true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if got := isAllocFor(c.allocID, c.workload); got != c.want {
|
||||
t.Errorf("isAllocFor(%q,%q) = %v, want %v", c.allocID, c.workload, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// normalizeRuntime / hasRuntime
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
func TestHasRuntimeAliases(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
node NodeInfo
|
||||
want bool
|
||||
}{
|
||||
{"wasm on wasmtime node", NodeInfo{Runtimes: []string{"wasmtime"}, Kind: "linux"}, true},
|
||||
{"wasm on process node", NodeInfo{Runtimes: []string{"process"}, Kind: "linux"}, false},
|
||||
{"pve-vm on linux node", NodeInfo{Runtimes: []string{"pve-vm"}, Kind: "linux"}, false},
|
||||
{"pve-vm on proxmox node", NodeInfo{Runtimes: nil, Kind: "proxmox"}, true},
|
||||
{"process on process node", NodeInfo{Runtimes: []string{"process"}, Kind: "linux"}, true},
|
||||
{"empty runtime on any node", NodeInfo{Runtimes: []string{"process"}, Kind: "linux"}, true},
|
||||
}
|
||||
for _, c := range cases {
|
||||
if c.name == "empty runtime on any node" {
|
||||
// hasRuntime is only called when OneOf != "".
|
||||
continue
|
||||
}
|
||||
if got := hasRuntime(c.node, "wasm"); c.name == "wasm on wasmtime node" || c.name == "wasm on process node" {
|
||||
if got != c.want {
|
||||
t.Errorf("%s: hasRuntime(wasm) = %v, want %v", c.name, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
// Explicit pve-vm and process checks.
|
||||
if !hasRuntime(NodeInfo{Runtimes: nil, Kind: "proxmox"}, "pve-vm") {
|
||||
t.Error("pve-vm on proxmox node should fit")
|
||||
}
|
||||
if hasRuntime(NodeInfo{Runtimes: nil, Kind: "linux"}, "pve-vm") {
|
||||
t.Error("pve-vm on linux node should not fit")
|
||||
}
|
||||
if !hasRuntime(NodeInfo{Runtimes: []string{"process"}, Kind: "linux"}, "process") {
|
||||
t.Error("process on process node should fit")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,35 @@
|
||||
package security
|
||||
|
||||
import (
|
||||
"os"
|
||||
"syscall"
|
||||
)
|
||||
|
||||
// Flock acquires an exclusive advisory lock on the file at path, creating it
|
||||
// if missing. Returns a release function that MUST be called (deferred) to
|
||||
// release the lock and close the file descriptor. Used by the known_hosts
|
||||
// read-modify-write paths (TOFUHostKeyCallback capture + ResetHostKey) to
|
||||
// prevent concurrent writers under v0.9's parallel SSH fan-out (REQ-063,
|
||||
// deferred P1 from REVIEW_v0.8 A2).
|
||||
func Flock(path string) (release func(), err error) {
|
||||
f, err := os.OpenFile(path, os.O_CREATE|os.O_RDWR, 0o600)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
if err := syscall.Flock(int(f.Fd()), syscall.LOCK_EX); err != nil {
|
||||
f.Close()
|
||||
return nil, err
|
||||
}
|
||||
return func() {
|
||||
_ = syscall.Flock(int(f.Fd()), syscall.LOCK_UN)
|
||||
_ = f.Close()
|
||||
}, nil
|
||||
}
|
||||
|
||||
func tryFlockEx(fd int) error {
|
||||
return syscall.Flock(fd, syscall.LOCK_EX|syscall.LOCK_NB)
|
||||
}
|
||||
|
||||
func releaseFlock(fd int) error {
|
||||
return syscall.Flock(fd, syscall.LOCK_UN)
|
||||
}
|
||||
@@ -0,0 +1,82 @@
|
||||
package security
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func TestFlock_acquireAndRelease(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, "test.lock")
|
||||
|
||||
release, err := Flock(path)
|
||||
if err != nil {
|
||||
t.Fatalf("Flock: %v", err)
|
||||
}
|
||||
if _, statErr := os.Stat(path); statErr != nil {
|
||||
t.Fatalf("lock file not created: %v", statErr)
|
||||
}
|
||||
release()
|
||||
}
|
||||
|
||||
func TestFlock_reentrantAfterRelease(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, "test.lock")
|
||||
|
||||
r1, err := Flock(path)
|
||||
if err != nil {
|
||||
t.Fatalf("first Flock: %v", err)
|
||||
}
|
||||
r1()
|
||||
|
||||
r2, err := Flock(path)
|
||||
if err != nil {
|
||||
t.Fatalf("second Flock after release: %v", err)
|
||||
}
|
||||
r2()
|
||||
}
|
||||
|
||||
func TestFlock_concurrentBlocks(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
path := filepath.Join(dir, "test.lock")
|
||||
|
||||
r1, err := Flock(path)
|
||||
if err != nil {
|
||||
t.Fatalf("first Flock: %v", err)
|
||||
}
|
||||
|
||||
// Give the blocking goroutine a chance to start and block.
|
||||
time.Sleep(50 * time.Millisecond)
|
||||
|
||||
// Verify the second lock is blocked by checking it hasn't acquired after a short window.
|
||||
// Use a non-blocking attempt: open the file and try LOCK_EX|LOCK_NB.
|
||||
blocked := make(chan bool, 1)
|
||||
go func() {
|
||||
f, err := os.OpenFile(path, os.O_RDWR, 0o600)
|
||||
if err != nil {
|
||||
blocked <- false
|
||||
return
|
||||
}
|
||||
defer f.Close()
|
||||
// LOCK_NB = non-blocking; returns EWOULDBLOCK if locked.
|
||||
if err := tryFlockEx(int(f.Fd())); err != nil {
|
||||
blocked <- true // got EWOULDBLOCK = the lock is held by r1
|
||||
return
|
||||
}
|
||||
releaseFlock(int(f.Fd()))
|
||||
blocked <- false // acquired = r1 didn't hold the lock (bug)
|
||||
}()
|
||||
|
||||
select {
|
||||
case b := <-blocked:
|
||||
if !b {
|
||||
t.Fatal("second lock acquired while first holds it — lock not working")
|
||||
}
|
||||
case <-time.After(2 * time.Second):
|
||||
t.Fatal("non-blocking try-lock timed out")
|
||||
}
|
||||
|
||||
r1()
|
||||
}
|
||||
@@ -0,0 +1,254 @@
|
||||
// Package schema provides kind-specific validators for the unified
|
||||
// *jobspec.WorkloadSpec introduced in P0b (REQ-064). Each workload kind
|
||||
// (Job, Service, DaemonSet per R-012) has different required fields;
|
||||
// this package exposes a Validator interface and a ValidatorFor
|
||||
// dispatcher so the emitter layer (REQ-074) and the lint engine
|
||||
// (REQ-084) can reject invalid specs before rendering.
|
||||
//
|
||||
// The validators operate purely on the *WorkloadSpec shape; they do no
|
||||
// I/O. Required-field violations return a structured error listing
|
||||
// every problem found (missing required fields, invalid combinations).
|
||||
package schema
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"net"
|
||||
"strings"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// Validator validates a *jobspec.WorkloadSpec against a kind-specific
|
||||
// schema. Implementations are pure (no I/O) and return a clear error
|
||||
// listing every violation found.
|
||||
type Validator interface {
|
||||
Validate(spec *jobspec.WorkloadSpec) error
|
||||
}
|
||||
|
||||
// JobValidator validates the Job workload kind (R-012).
|
||||
//
|
||||
// Rules:
|
||||
// - no service block required (Job has no Traefik route by default D-175)
|
||||
// - restart optional (defaults to never/on-failure when omitted)
|
||||
// - schedule optional (cron string)
|
||||
// - timeout optional
|
||||
// - ports optional
|
||||
// - count must be 1 (or unset → 1); count > 1 is an error for Job
|
||||
// (use a Service for replicas)
|
||||
// - no Traefik route (a ServiceBlock is rejected)
|
||||
// - task group (spec.Tasks) optional; when present, each task must
|
||||
// have a unique name and a resolvable command (P06).
|
||||
type JobValidator struct{}
|
||||
|
||||
// ServiceValidator validates the Service workload kind (R-012).
|
||||
//
|
||||
// Rules:
|
||||
// - ports required (at least one)
|
||||
// - count ≥ 1
|
||||
// - restart required (mode must be service)
|
||||
// - update required (strategy must be rolling/canary/blue-green)
|
||||
// - runtime required (unless a task group is present; each task
|
||||
// can carry its own runtime — P06)
|
||||
// - health block required (Traefik routing depends on health checks)
|
||||
// - service block, if present, must have a valid bind (127.0.0.1
|
||||
// opt-in per R-007; default is socket — empty bind is OK)
|
||||
// - service block implied (Traefik route YES)
|
||||
// - task group (spec.Tasks) optional; when present, each task must
|
||||
// have a unique name and a resolvable command (P06).
|
||||
type ServiceValidator struct{}
|
||||
|
||||
// DaemonSetValidator validates the DaemonSet workload kind (R-012).
|
||||
//
|
||||
// Rules:
|
||||
// - schedule block with mode (every-node/matching/mandatory) required
|
||||
// - no ports (no Traefik route by default D-175)
|
||||
// - no count (implicit = nodes matching condition)
|
||||
// - restart required
|
||||
// - task group (spec.Tasks) optional; when present, each task must
|
||||
// have a unique name and a resolvable command (P06).
|
||||
type DaemonSetValidator struct{}
|
||||
|
||||
// ValidatorFor returns the Validator for the given workload kind, or an
|
||||
// error for an unknown kind. kind must be one of Job, Service,
|
||||
// DaemonSet (R-012).
|
||||
func ValidatorFor(kind string) (Validator, error) {
|
||||
switch kind {
|
||||
case "Job":
|
||||
return JobValidator{}, nil
|
||||
case "Service":
|
||||
return ServiceValidator{}, nil
|
||||
case "DaemonSet":
|
||||
return DaemonSetValidator{}, nil
|
||||
default:
|
||||
return nil, fmt.Errorf("schema: unknown kind %q (want one of Job, Service, DaemonSet)", kind)
|
||||
}
|
||||
}
|
||||
|
||||
// Validate validates a Job spec. See JobValidator for the rules.
|
||||
func (JobValidator) Validate(spec *jobspec.WorkloadSpec) error {
|
||||
if spec == nil {
|
||||
return errors.New("schema/Job: spec is nil")
|
||||
}
|
||||
var errs []string
|
||||
if strings.TrimSpace(spec.Name) == "" {
|
||||
errs = append(errs, "name is required")
|
||||
}
|
||||
if spec.Count != 0 && spec.Count != 1 {
|
||||
errs = append(errs, fmt.Sprintf("count must be 1 (or unset) for Job, got %d (use Service for replicas)", spec.Count))
|
||||
}
|
||||
if spec.Service != nil {
|
||||
errs = append(errs, "service block (Traefik route) is not allowed for Job (D-175)")
|
||||
}
|
||||
errs = append(errs, validateTaskGroup(spec)...)
|
||||
return composeErrors("schema/Job", errs)
|
||||
}
|
||||
|
||||
// Validate validates a Service spec. See ServiceValidator for the rules.
|
||||
func (ServiceValidator) Validate(spec *jobspec.WorkloadSpec) error {
|
||||
if spec == nil {
|
||||
return errors.New("schema/Service: spec is nil")
|
||||
}
|
||||
var errs []string
|
||||
if strings.TrimSpace(spec.Name) == "" {
|
||||
errs = append(errs, "name is required")
|
||||
}
|
||||
if len(spec.Ports) == 0 {
|
||||
errs = append(errs, "ports required (at least one)")
|
||||
}
|
||||
if spec.Count < 1 {
|
||||
errs = append(errs, fmt.Sprintf("count must be ≥ 1 for Service, got %d", spec.Count))
|
||||
}
|
||||
if spec.Restart == nil {
|
||||
errs = append(errs, "restart block required for Service")
|
||||
} else {
|
||||
switch spec.Restart.Mode {
|
||||
case "service", "on-failure", "never":
|
||||
// Valid per R-012 (default for Service is "service",
|
||||
// but the validator accepts the full enum; the
|
||||
// Service-specific "must be service" rule is enforced
|
||||
// below for the default case where mode is empty).
|
||||
case "":
|
||||
errs = append(errs, "restart mode required for Service (one of service, on-failure, never; default is service)")
|
||||
default:
|
||||
errs = append(errs, fmt.Sprintf("restart mode %q invalid (want one of service, on-failure, never)", spec.Restart.Mode))
|
||||
}
|
||||
}
|
||||
if spec.Update == nil {
|
||||
errs = append(errs, "update block required for Service")
|
||||
} else {
|
||||
switch spec.Update.Strategy {
|
||||
case "rolling", "canary", "blue-green":
|
||||
case "":
|
||||
errs = append(errs, "update strategy required for Service (one of rolling, canary, blue-green)")
|
||||
default:
|
||||
errs = append(errs, fmt.Sprintf("update strategy %q invalid (want one of rolling, canary, blue-green)", spec.Update.Strategy))
|
||||
}
|
||||
}
|
||||
if spec.Runtime == nil && len(spec.Tasks) == 0 {
|
||||
errs = append(errs, "runtime block required for Service (or a task group with per-task runtimes)")
|
||||
}
|
||||
if spec.Health == nil {
|
||||
errs = append(errs, "health block required for Service (Traefik routing requires health checks)")
|
||||
}
|
||||
if spec.Service != nil {
|
||||
if err := validateServiceBind(spec.Service.Bind); err != nil {
|
||||
errs = append(errs, err.Error())
|
||||
}
|
||||
}
|
||||
errs = append(errs, validateTaskGroup(spec)...)
|
||||
return composeErrors("schema/Service", errs)
|
||||
}
|
||||
|
||||
// validateServiceBind validates the service.bind field (R-007). Empty
|
||||
// is OK (default = socket). When set, it must be a valid IPv4/IPv6
|
||||
// address (the only opt-in to bind on a non-loopback address); the
|
||||
// loopback 127.0.0.1 is the documented opt-in. Anything that is not
|
||||
// parseable as an IP address is rejected.
|
||||
func validateServiceBind(bind string) error {
|
||||
if strings.TrimSpace(bind) == "" {
|
||||
return nil
|
||||
}
|
||||
if net.ParseIP(bind) == nil {
|
||||
return fmt.Errorf("service.bind %q is not a valid IP address (R-007: 127.0.0.1 opt-in; default is socket)", bind)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// Validate validates a DaemonSet spec. See DaemonSetValidator for the rules.
|
||||
func (DaemonSetValidator) Validate(spec *jobspec.WorkloadSpec) error {
|
||||
if spec == nil {
|
||||
return errors.New("schema/DaemonSet: spec is nil")
|
||||
}
|
||||
var errs []string
|
||||
if strings.TrimSpace(spec.Name) == "" {
|
||||
errs = append(errs, "name is required")
|
||||
}
|
||||
if spec.Schedule == nil {
|
||||
errs = append(errs, "schedule block required for DaemonSet")
|
||||
} else {
|
||||
switch spec.Schedule.Mode {
|
||||
case "every-node", "matching", "mandatory":
|
||||
case "":
|
||||
errs = append(errs, "schedule mode required for DaemonSet (one of every-node, matching, mandatory)")
|
||||
default:
|
||||
errs = append(errs, fmt.Sprintf("schedule mode %q invalid (want one of every-node, matching, mandatory)", spec.Schedule.Mode))
|
||||
}
|
||||
}
|
||||
if len(spec.Ports) > 0 {
|
||||
errs = append(errs, "ports not allowed for DaemonSet (no Traefik route by default D-175)")
|
||||
}
|
||||
if spec.Count != 0 {
|
||||
errs = append(errs, fmt.Sprintf("count not allowed for DaemonSet (implicit = nodes matching condition), got %d", spec.Count))
|
||||
}
|
||||
if spec.Restart == nil {
|
||||
errs = append(errs, "restart block required for DaemonSet")
|
||||
}
|
||||
errs = append(errs, validateTaskGroup(spec)...)
|
||||
return composeErrors("schema/DaemonSet", errs)
|
||||
}
|
||||
|
||||
// composeErrors joins the per-field errors into a single error prefixed
|
||||
// by the validator name. Returns nil when there are no errors so the
|
||||
// caller can return the result directly.
|
||||
func composeErrors(name string, errs []string) error {
|
||||
if len(errs) == 0 {
|
||||
return nil
|
||||
}
|
||||
return fmt.Errorf("%s: %s", name, strings.Join(errs, "; "))
|
||||
}
|
||||
|
||||
// validateTaskGroup validates the task-group list shared by all kinds
|
||||
// (P06, PRD §9.1). When the spec carries a task group (spec.Tasks
|
||||
// non-empty), each task must have a unique name and a resolvable
|
||||
// command (the task's own Command, the task's runtime command, or the
|
||||
// top-level runtime command as the per-group default). The top-level
|
||||
// runtime is optional when tasks is present (each task can carry its
|
||||
// own runtime). Returns nil when the spec has no task group.
|
||||
func validateTaskGroup(spec *jobspec.WorkloadSpec) []string {
|
||||
if len(spec.Tasks) == 0 {
|
||||
return nil
|
||||
}
|
||||
var errs []string
|
||||
seen := make(map[string]bool, len(spec.Tasks))
|
||||
for i, task := range spec.Tasks {
|
||||
if strings.TrimSpace(task.Name) == "" {
|
||||
errs = append(errs, fmt.Sprintf("tasks[%d]: name is required", i))
|
||||
} else if seen[task.Name] {
|
||||
errs = append(errs, fmt.Sprintf("tasks[%d]: duplicate task name %q (names must be unique within the group)", i, task.Name))
|
||||
} else {
|
||||
seen[task.Name] = true
|
||||
}
|
||||
cmd := task.Command
|
||||
if strings.TrimSpace(cmd) == "" && task.Runtime != nil {
|
||||
cmd = task.Runtime.Command
|
||||
}
|
||||
if strings.TrimSpace(cmd) == "" && spec.Runtime != nil {
|
||||
cmd = spec.Runtime.Command
|
||||
}
|
||||
if strings.TrimSpace(cmd) == "" {
|
||||
errs = append(errs, fmt.Sprintf("tasks[%d]: command is required (set tasks[].command, tasks[].runtime.command, or top-level runtime.command)", i))
|
||||
}
|
||||
}
|
||||
return errs
|
||||
}
|
||||
@@ -0,0 +1,871 @@
|
||||
package schema
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
func TestJobValidator_ValidMinimal(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "backup", Count: 1}
|
||||
v := JobValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestJobValidator_ValidWithSchedule(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "backup",
|
||||
Count: 1,
|
||||
Schedule: &jobspec.ScheduleBlock{Cron: "0 2 * * *"},
|
||||
Timeout: "1h",
|
||||
}
|
||||
v := JobValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestJobValidator_ValidUnsetCount(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "one-shot"}
|
||||
v := JobValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("unset count should default-accept, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestJobValidator_MissingName(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Job", Count: 1}
|
||||
err := JobValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing name, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "name is required") {
|
||||
t.Errorf("error = %q, want 'name is required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestJobValidator_CountGreaterThanOne(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "batch", Count: 3}
|
||||
err := JobValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for count > 1, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "count must be 1") {
|
||||
t.Errorf("error = %q, want 'count must be 1'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestJobValidator_ServiceBlockRejected(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "x",
|
||||
Count: 1,
|
||||
Service: &jobspec.ServiceBlock{Host: "x.example"},
|
||||
}
|
||||
err := JobValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for service block on Job, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "service block") {
|
||||
t.Errorf("error = %q, want 'service block'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestJobValidator_NilSpec(t *testing.T) {
|
||||
v := JobValidator{}
|
||||
if err := v.Validate(nil); err == nil {
|
||||
t.Fatal("expected error for nil spec")
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_ValidFull(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 3,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/http"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling", MaxSurge: 1},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http", Interval: "5s"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
v := ServiceValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_MissingPorts(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing ports, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "ports required") {
|
||||
t.Errorf("error = %q, want 'ports required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_CountZero(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 0,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for count 0, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "count must be") {
|
||||
t.Errorf("error = %q, want 'count must be'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_MissingRestart(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing restart, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "restart block required") {
|
||||
t.Errorf("error = %q, want 'restart block required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_WrongRestartMode(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "always"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for invalid restart mode, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "restart mode") {
|
||||
t.Errorf("error = %q, want 'restart mode'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_AcceptedRestartModes(t *testing.T) {
|
||||
// R-012: restart.mode accepts service / on-failure / never for
|
||||
// Service; the default per R-012 is "service" but the validator
|
||||
// accepts the full enum (a Service that wants on-failure is
|
||||
// unusual but not invalid — only "always" and unknown modes are
|
||||
// rejected).
|
||||
for _, mode := range []string{"service", "on-failure", "never"} {
|
||||
t.Run(mode, func(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: mode},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
v := ServiceValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Errorf("mode %q should be accepted, got: %v", mode, err)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_MissingUpdate(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing update, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "update block required") {
|
||||
t.Errorf("error = %q, want 'update block required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_MissingRuntime(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing runtime, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "runtime block required") {
|
||||
t.Errorf("error = %q, want 'runtime block required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_NilSpec(t *testing.T) {
|
||||
v := ServiceValidator{}
|
||||
if err := v.Validate(nil); err == nil {
|
||||
t.Fatal("expected error for nil spec")
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_MissingHealth(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing health block, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "health block required") {
|
||||
t.Errorf("error = %q, want 'health block required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_InvalidRestartMode(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "always"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for invalid restart mode, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "restart mode") || !strings.Contains(err.Error(), "invalid") {
|
||||
t.Errorf("error = %q, want 'restart mode ... invalid'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_EmptyRestartMode(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: ""},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty restart mode, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "restart mode required") {
|
||||
t.Errorf("error = %q, want 'restart mode required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_InvalidUpdateStrategy(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "recreate"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for invalid update strategy, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "update strategy") || !strings.Contains(err.Error(), "invalid") {
|
||||
t.Errorf("error = %q, want 'update strategy ... invalid'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_EmptyUpdateStrategy(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: ""},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty update strategy, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "update strategy required") {
|
||||
t.Errorf("error = %q, want 'update strategy required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_AcceptedUpdateStrategies(t *testing.T) {
|
||||
for _, strat := range []string{"rolling", "canary", "blue-green"} {
|
||||
t.Run(strat, func(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: strat},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
v := ServiceValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Errorf("strategy %q should be accepted, got: %v", strat, err)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_InvalidServiceBind(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: "not-an-ip"},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for invalid service.bind, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "service.bind") || !strings.Contains(err.Error(), "valid IP") {
|
||||
t.Errorf("error = %q, want 'service.bind ... valid IP'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_ValidServiceBindLoopback(t *testing.T) {
|
||||
// R-007: 127.0.0.1 is the documented opt-in for a non-socket bind.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: "127.0.0.1"},
|
||||
}
|
||||
v := ServiceValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("127.0.0.1 should be accepted, got: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_ValidServiceBindIPv6(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: "::1"},
|
||||
}
|
||||
v := ServiceValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("::1 should be accepted, got: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_EmptyServiceBindOK(t *testing.T) {
|
||||
// R-007: empty bind = default = socket; valid.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
Service: &jobspec.ServiceBlock{Bind: ""},
|
||||
}
|
||||
v := ServiceValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("empty bind should default to socket (valid), got: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestServiceValidator_MultipleErrors(t *testing.T) {
|
||||
// Multiple violations should all surface in the composed error.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "",
|
||||
Count: 0,
|
||||
Restart: &jobspec.RestartBlock{Mode: "always"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "recreate"},
|
||||
Service: &jobspec.ServiceBlock{Bind: "not-an-ip"},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error, got nil")
|
||||
}
|
||||
for _, want := range []string{
|
||||
"name is required",
|
||||
"ports required",
|
||||
"count must be",
|
||||
"restart mode",
|
||||
"update strategy",
|
||||
"runtime block required",
|
||||
"health block required",
|
||||
"service.bind",
|
||||
} {
|
||||
if !strings.Contains(err.Error(), want) {
|
||||
t.Errorf("error %q missing %q", err.Error(), want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_Valid(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "log-shipper",
|
||||
Schedule: &jobspec.ScheduleBlock{Mode: "every-node"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
|
||||
}
|
||||
v := DaemonSetValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_ValidMatchingMode(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "x",
|
||||
Schedule: &jobspec.ScheduleBlock{Mode: "matching"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
|
||||
}
|
||||
v := DaemonSetValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("matching mode should be accepted, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_ValidMandatoryMode(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "x",
|
||||
Schedule: &jobspec.ScheduleBlock{Mode: "mandatory"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
|
||||
}
|
||||
v := DaemonSetValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("mandatory mode should be accepted, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_MissingScheduleMode(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "x",
|
||||
Schedule: &jobspec.ScheduleBlock{Mode: ""},
|
||||
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
|
||||
}
|
||||
err := DaemonSetValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing schedule mode, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "schedule mode required") {
|
||||
t.Errorf("error = %q, want 'schedule mode required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_MissingScheduleBlock(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "x",
|
||||
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
|
||||
}
|
||||
err := DaemonSetValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing schedule block, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "schedule block required") {
|
||||
t.Errorf("error = %q, want 'schedule block required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_InvalidScheduleMode(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "x",
|
||||
Schedule: &jobspec.ScheduleBlock{Mode: "always"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
|
||||
}
|
||||
err := DaemonSetValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for invalid schedule mode, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "schedule mode") {
|
||||
t.Errorf("error = %q, want 'schedule mode'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_HasPorts(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "x",
|
||||
Schedule: &jobspec.ScheduleBlock{Mode: "every-node"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 80}},
|
||||
}
|
||||
err := DaemonSetValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for ports on DaemonSet, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "ports not allowed") {
|
||||
t.Errorf("error = %q, want 'ports not allowed'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_HasCount(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "x",
|
||||
Count: 3,
|
||||
Schedule: &jobspec.ScheduleBlock{Mode: "every-node"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
|
||||
}
|
||||
err := DaemonSetValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for count on DaemonSet, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "count not allowed") {
|
||||
t.Errorf("error = %q, want 'count not allowed'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_MissingRestart(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "x",
|
||||
Schedule: &jobspec.ScheduleBlock{Mode: "every-node"},
|
||||
}
|
||||
err := DaemonSetValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing restart on DaemonSet, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "restart block required") {
|
||||
t.Errorf("error = %q, want 'restart block required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestDaemonSetValidator_NilSpec(t *testing.T) {
|
||||
v := DaemonSetValidator{}
|
||||
if err := v.Validate(nil); err == nil {
|
||||
t.Fatal("expected error for nil spec")
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidatorFor_EachKind(t *testing.T) {
|
||||
cases := []struct {
|
||||
kind string
|
||||
want string
|
||||
}{
|
||||
{"Job", "schema.JobValidator"},
|
||||
{"Service", "schema.ServiceValidator"},
|
||||
{"DaemonSet", "schema.DaemonSetValidator"},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.kind, func(t *testing.T) {
|
||||
v, err := ValidatorFor(tc.kind)
|
||||
if err != nil {
|
||||
t.Fatalf("ValidatorFor(%q): %v", tc.kind, err)
|
||||
}
|
||||
got := fmtType(v)
|
||||
if got != tc.want {
|
||||
t.Errorf("ValidatorFor(%q) type = %q, want %q", tc.kind, got, tc.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestValidatorFor_UnknownKind(t *testing.T) {
|
||||
_, err := ValidatorFor("CronJob")
|
||||
if err == nil {
|
||||
t.Fatal("expected error for unknown kind, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "unknown kind") {
|
||||
t.Errorf("error = %q, want 'unknown kind'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
// fmtType returns a readable type name for a validator. Uses fmt.Sprintf
|
||||
// with %T rather than reflection to keep the test surface minimal.
|
||||
func fmtType(v Validator) string {
|
||||
switch v.(type) {
|
||||
case JobValidator:
|
||||
return "schema.JobValidator"
|
||||
case ServiceValidator:
|
||||
return "schema.ServiceValidator"
|
||||
case DaemonSetValidator:
|
||||
return "schema.DaemonSetValidator"
|
||||
default:
|
||||
return "unknown"
|
||||
}
|
||||
}
|
||||
|
||||
// Ensure composeErrors returns nil for empty input (covers the
|
||||
// short-circuit branch that the validators rely on).
|
||||
func TestComposeErrors_Empty(t *testing.T) {
|
||||
if err := composeErrors("schema/X", nil); err != nil {
|
||||
t.Errorf("composeErrors(nil) = %v, want nil", err)
|
||||
}
|
||||
if err := composeErrors("schema/X", []string{}); err != nil {
|
||||
t.Errorf("composeErrors([]) = %v, want nil", err)
|
||||
}
|
||||
}
|
||||
|
||||
// Ensure the error type returned by composeErrors is a non-nil error
|
||||
// when violations are present (guards against accidental nil-return).
|
||||
func TestComposeErrors_NonEmpty(t *testing.T) {
|
||||
err := composeErrors("schema/X", []string{"a", "b"})
|
||||
if err == nil {
|
||||
t.Fatal("expected non-nil error")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "a") || !strings.Contains(err.Error(), "b") {
|
||||
t.Errorf("error = %q, want both 'a' and 'b'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
// Compile-time assertion that the validators implement the interface.
|
||||
var (
|
||||
_ Validator = JobValidator{}
|
||||
_ Validator = ServiceValidator{}
|
||||
_ Validator = DaemonSetValidator{}
|
||||
)
|
||||
|
||||
func TestTaskGroup_Valid(t *testing.T) {
|
||||
// P06: a valid task group — two tasks, each with a unique name
|
||||
// and a resolvable command (own command). The top-level runtime
|
||||
// is optional when each task carries its own.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{Name: "app", Command: "/usr/bin/httpd"},
|
||||
{Name: "sidecar", Command: "/bin/wasm-runner sidecar.wasm"},
|
||||
},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
if err := (ServiceValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTaskGroup_ValidInheritsTopLevelRuntime(t *testing.T) {
|
||||
// P06: tasks without their own runtime inherit the top-level
|
||||
// runtime command. The validator accepts this as long as the
|
||||
// resolved command is non-empty.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/default"},
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{Name: "app"},
|
||||
{Name: "sidecar"},
|
||||
},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
if err := (ServiceValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTaskGroup_ValidTaskRuntimeCommand(t *testing.T) {
|
||||
// P06: a task whose command is provided via the task's own
|
||||
// runtime.command (no top-level runtime) is valid.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 1,
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{Name: "app", Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/usr/bin/httpd"}},
|
||||
},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
if err := (ServiceValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTaskGroup_MissingTaskName(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{Command: "/usr/bin/httpd"},
|
||||
{Name: "sidecar", Command: "/bin/wasm-runner"},
|
||||
},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing task name, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "name is required") {
|
||||
t.Errorf("error = %q, want 'name is required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestTaskGroup_DuplicateTaskNames(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{Name: "app", Command: "/usr/bin/httpd"},
|
||||
{Name: "app", Command: "/bin/other"},
|
||||
},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for duplicate task names, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "duplicate task name") {
|
||||
t.Errorf("error = %q, want 'duplicate task name'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestTaskGroup_MissingCommand(t *testing.T) {
|
||||
// P06: a task with no resolvable command (no task.Command, no
|
||||
// task.Runtime, no top-level Runtime) is rejected.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{Name: "app"},
|
||||
},
|
||||
Restart: &jobspec.RestartBlock{Mode: "service"},
|
||||
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
|
||||
Health: &jobspec.HealthBlock{CheckType: "http"},
|
||||
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
|
||||
}
|
||||
err := ServiceValidator{}.Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for missing task command, got nil")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "command is required") {
|
||||
t.Errorf("error = %q, want 'command is required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestTaskGroup_JobAcceptsTaskGroup(t *testing.T) {
|
||||
// P06: task groups apply to all kinds, not just Service. Job
|
||||
// accepts a task group with unique names + resolvable commands.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "batch",
|
||||
Count: 1,
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{Name: "step1", Command: "/bin/extract"},
|
||||
{Name: "step2", Command: "/bin/transform"},
|
||||
},
|
||||
}
|
||||
if err := (JobValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTaskGroup_DaemonSetAcceptsTaskGroup(t *testing.T) {
|
||||
// P06: DaemonSet accepts a task group.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "DaemonSet",
|
||||
Name: "log-shipper",
|
||||
Schedule: &jobspec.ScheduleBlock{Mode: "every-node"},
|
||||
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
|
||||
Tasks: []jobspec.TaskGroupTask{
|
||||
{Name: "collector", Command: "/bin/collect"},
|
||||
{Name: "forwarder", Command: "/bin/forward"},
|
||||
},
|
||||
}
|
||||
if err := (DaemonSetValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTaskGroup_NoTasksBackwardCompat(t *testing.T) {
|
||||
// Backward compat: a spec with no Tasks is validated by the
|
||||
// existing kind-specific rules (no task-group check fires).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Job",
|
||||
Name: "backup",
|
||||
Count: 1,
|
||||
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
|
||||
}
|
||||
if err := (JobValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,126 @@
|
||||
package schema
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
// UpdateValidator validates the rolling/canary/blue-green update stanza
|
||||
// (P03). The schema validator (ServiceValidator) already enforces that
|
||||
// the strategy is one of rolling/canary/blue-green and that the update
|
||||
// block is present for a Service. UpdateValidator adds the
|
||||
// field-level validation:
|
||||
//
|
||||
// - max_parallel: integer in [1, count] (defaults to 1 when unset)
|
||||
// - min_healthy_time: a valid time.Duration when set (time.ParseDuration)
|
||||
// - healthy_deadline: a valid time.Duration when set (time.ParseDuration)
|
||||
// - canary: an integer count in [0, count] OR a percentage string of
|
||||
// the form "<n>%" where n is in [0, 100] (the parser already accepts
|
||||
// both shapes; the validator accepts them too). Only meaningful
|
||||
// for the canary strategy; ignored (but still validated for shape)
|
||||
// for rolling/blue-green.
|
||||
// - auto_promote: boolean (no validation beyond the parser's
|
||||
// true/false parse; the field is always populated)
|
||||
//
|
||||
// The validator is pure (no I/O). Violations return a clear error
|
||||
// listing every problem found, mirroring the per-field style of
|
||||
// ServiceValidator.
|
||||
type UpdateValidator struct{}
|
||||
|
||||
// Validate validates the UpdateBlock on the given spec. The spec must
|
||||
// be non-nil and carry a Count (services have count ≥ 1 per
|
||||
// ServiceValidator). When spec.Update is nil the validator returns an
|
||||
// error (the update block is required for Service; this validator
|
||||
// assumes the caller has already established the spec is a Service).
|
||||
func (UpdateValidator) Validate(spec *jobspec.WorkloadSpec) error {
|
||||
if spec == nil {
|
||||
return fmt.Errorf("schema/Update: spec is nil")
|
||||
}
|
||||
if spec.Update == nil {
|
||||
return fmt.Errorf("schema/Update: update block is nil")
|
||||
}
|
||||
var errs []string
|
||||
u := spec.Update
|
||||
|
||||
switch u.Strategy {
|
||||
case "rolling", "canary", "blue-green":
|
||||
case "":
|
||||
errs = append(errs, "update strategy required (one of rolling, canary, blue-green)")
|
||||
default:
|
||||
errs = append(errs, fmt.Sprintf("update strategy %q invalid (want one of rolling, canary, blue-green)", u.Strategy))
|
||||
}
|
||||
|
||||
// max_parallel defaults to 1 when unset (0); validate the range
|
||||
// only when the user has set it explicitly.
|
||||
if u.MaxParallel != 0 {
|
||||
if u.MaxParallel < 1 {
|
||||
errs = append(errs, fmt.Sprintf("update.max_parallel must be ≥ 1, got %d", u.MaxParallel))
|
||||
}
|
||||
if spec.Count > 0 && u.MaxParallel > spec.Count {
|
||||
errs = append(errs, fmt.Sprintf("update.max_parallel %d exceeds count %d (must be 1..count)", u.MaxParallel, spec.Count))
|
||||
}
|
||||
}
|
||||
|
||||
if u.MinHealthyTime != "" {
|
||||
if _, err := time.ParseDuration(u.MinHealthyTime); err != nil {
|
||||
errs = append(errs, fmt.Sprintf("update.min_healthy_time %q is not a valid duration: %v", u.MinHealthyTime, err))
|
||||
}
|
||||
}
|
||||
if u.HealthyDeadline != "" {
|
||||
if _, err := time.ParseDuration(u.HealthyDeadline); err != nil {
|
||||
errs = append(errs, fmt.Sprintf("update.healthy_deadline %q is not a valid duration: %v", u.HealthyDeadline, err))
|
||||
}
|
||||
}
|
||||
|
||||
// canary accepts an integer count (0..count) or a percentage
|
||||
// ("<n>%" with n in 0..100). The field is only meaningful for the
|
||||
// canary strategy but we validate the shape regardless so a typo
|
||||
// in a rolling/blue-green stanza still surfaces.
|
||||
if u.Canary != "" {
|
||||
if err := validateCanary(u.Canary, spec.Count); err != nil {
|
||||
errs = append(errs, err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
// auto_promote is a bool; no extra validation beyond the parser.
|
||||
|
||||
return composeErrors("schema/Update", errs)
|
||||
}
|
||||
|
||||
// validateCanary validates the canary field shape: either an integer
|
||||
// count (0..count) or a percentage string "<n>%" (n in 0..100). count
|
||||
// is the spec.Count; when count is 0 (e.g. a DaemonSet or unset), the
|
||||
// integer-count upper bound is not enforced (only the percentage
|
||||
// bound is enforced, since percentage does not depend on count).
|
||||
func validateCanary(canary string, count int) error {
|
||||
c := strings.TrimSpace(canary)
|
||||
if c == "" {
|
||||
return nil
|
||||
}
|
||||
if strings.HasSuffix(c, "%") {
|
||||
nStr := strings.TrimSuffix(c, "%")
|
||||
n, err := strconv.Atoi(strings.TrimSpace(nStr))
|
||||
if err != nil {
|
||||
return fmt.Errorf("update.canary %q is not a valid percentage (want \"<n>%%\")", canary)
|
||||
}
|
||||
if n < 0 || n > 100 {
|
||||
return fmt.Errorf("update.canary percentage %d out of range (want 0..100)", n)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
n, err := strconv.Atoi(c)
|
||||
if err != nil {
|
||||
return fmt.Errorf("update.canary %q is not a valid count or percentage (want integer or \"<n>%%\")", canary)
|
||||
}
|
||||
if n < 0 {
|
||||
return fmt.Errorf("update.canary count %d must be ≥ 0", n)
|
||||
}
|
||||
if count > 0 && n > count {
|
||||
return fmt.Errorf("update.canary count %d exceeds count %d (must be 0..count)", n, count)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
@@ -0,0 +1,464 @@
|
||||
package schema
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/jobspec"
|
||||
)
|
||||
|
||||
func TestUpdateValidator_ValidRolling(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
MaxParallel: 2,
|
||||
MinHealthyTime: "30s",
|
||||
HealthyDeadline: "5m",
|
||||
},
|
||||
}
|
||||
v := UpdateValidator{}
|
||||
if err := v.Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_ValidCanary(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
MaxParallel: 2,
|
||||
MinHealthyTime: "30s",
|
||||
HealthyDeadline: "5m",
|
||||
Canary: "10%",
|
||||
AutoPromote: true,
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_ValidBlueGreen(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 3,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "blue-green",
|
||||
MinHealthyTime: "1m",
|
||||
HealthyDeadline: "10m",
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("expected nil, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_ValidCanaryIntegerCount(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "1",
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("integer canary count 1 should be valid, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_ValidCanaryZeroPercent(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "0%",
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("0%% canary should be valid, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_ValidCanaryHundredPercent(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "100%",
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("100%% canary should be valid, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_EmptyDurationsOK(t *testing.T) {
|
||||
// Empty min_healthy_time / healthy_deadline should be accepted
|
||||
// (they are optional; defaults are applied by the executor).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("empty durations should be valid, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_NilSpec(t *testing.T) {
|
||||
if err := (UpdateValidator{}).Validate(nil); err == nil {
|
||||
t.Fatal("expected error for nil spec")
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_NilUpdateBlock(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{Kind: "Service", Name: "web", Count: 2}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for nil update block")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "update block is nil") {
|
||||
t.Errorf("error = %q, want 'update block is nil'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_InvalidStrategy(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "recreate",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for invalid strategy")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "strategy") || !strings.Contains(err.Error(), "invalid") {
|
||||
t.Errorf("error = %q, want 'strategy ... invalid'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_EmptyStrategy(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for empty strategy")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "strategy required") {
|
||||
t.Errorf("error = %q, want 'strategy required'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_MaxParallelZero(t *testing.T) {
|
||||
// max_parallel=0 means "unset" → default 1; accepted.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
MaxParallel: 0,
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("max_parallel=0 (unset) should be valid, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_MaxParallelNegative(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
MaxParallel: -1,
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for negative max_parallel")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "max_parallel") {
|
||||
t.Errorf("error = %q, want 'max_parallel'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_MaxParallelExceedsCount(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
MaxParallel: 5,
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for max_parallel > count")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "exceeds count") {
|
||||
t.Errorf("error = %q, want 'exceeds count'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_MaxParallelEqualsCount(t *testing.T) {
|
||||
// max_parallel == count is the upper bound; valid.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 3,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
MaxParallel: 3,
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("max_parallel==count should be valid, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_InvalidMinHealthyTime(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
MinHealthyTime: "not-a-duration",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for invalid min_healthy_time")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "min_healthy_time") {
|
||||
t.Errorf("error = %q, want 'min_healthy_time'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_InvalidHealthyDeadline(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "rolling",
|
||||
HealthyDeadline: "nope",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for invalid healthy_deadline")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "healthy_deadline") {
|
||||
t.Errorf("error = %q, want 'healthy_deadline'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_CanaryNegative(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "-1",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for negative canary count")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "canary") {
|
||||
t.Errorf("error = %q, want 'canary'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_CanaryExceedsCount(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "5",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for canary > count")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "exceeds count") {
|
||||
t.Errorf("error = %q, want 'exceeds count'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_CanaryPercentOver100(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "150%",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for canary > 100%")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "out of range") {
|
||||
t.Errorf("error = %q, want 'out of range'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_CanaryPercentNegative(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "-10%",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for negative canary percent")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "out of range") {
|
||||
t.Errorf("error = %q, want 'out of range'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_CanaryNotANumber(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "abc",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for non-numeric canary")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "not a valid") {
|
||||
t.Errorf("error = %q, want 'not a valid'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_CanaryPercentNotANumber(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "xx%",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error for non-numeric canary percent")
|
||||
}
|
||||
if !strings.Contains(err.Error(), "not a valid percentage") {
|
||||
t.Errorf("error = %q, want 'not a valid percentage'", err.Error())
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_CanaryCountZeroOK(t *testing.T) {
|
||||
// canary=0 is the lower bound; valid.
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 4,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "0",
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("canary=0 should be valid, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_CanaryPercentWithoutCountOK(t *testing.T) {
|
||||
// A percentage canary does not depend on count; valid even when
|
||||
// count is 0 (e.g. DaemonSet-shaped spec bypassing the Service
|
||||
// validator — defensive).
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 0,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "canary",
|
||||
Canary: "25%",
|
||||
},
|
||||
}
|
||||
if err := (UpdateValidator{}).Validate(spec); err != nil {
|
||||
t.Fatalf("percentage canary with count=0 should be valid, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestUpdateValidator_MultipleErrors(t *testing.T) {
|
||||
spec := &jobspec.WorkloadSpec{
|
||||
Kind: "Service",
|
||||
Name: "web",
|
||||
Count: 2,
|
||||
Update: &jobspec.UpdateBlock{
|
||||
Strategy: "recreate",
|
||||
MaxParallel: 99,
|
||||
MinHealthyTime: "nope",
|
||||
HealthyDeadline: "also-nope",
|
||||
Canary: "200%",
|
||||
},
|
||||
}
|
||||
err := (UpdateValidator{}).Validate(spec)
|
||||
if err == nil {
|
||||
t.Fatal("expected error, got nil")
|
||||
}
|
||||
for _, want := range []string{
|
||||
"strategy",
|
||||
"max_parallel",
|
||||
"min_healthy_time",
|
||||
"healthy_deadline",
|
||||
"canary",
|
||||
} {
|
||||
if !strings.Contains(err.Error(), want) {
|
||||
t.Errorf("error %q missing %q", err.Error(), want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Compile-time assertion that UpdateValidator implements Validator.
|
||||
var _ Validator = UpdateValidator{}
|
||||
@@ -0,0 +1,30 @@
|
||||
// Package sshpush_test contains compile-time assertions that *Transport
|
||||
// satisfies the emitter.AtomicWriter interface (the Traefik C-10
|
||||
// atomicity protocol — internal/emitter/traefik_atomic.go). The
|
||||
// assertion lives here (not in internal/emitter) to avoid an import
|
||||
// cycle: internal/emitter is imported by this package (fanout.go), so
|
||||
// internal/emitter cannot import this package.
|
||||
package sshpush_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/emitter"
|
||||
"git.cloudinit.dev/coreci/orca/internal/sshpush"
|
||||
)
|
||||
|
||||
// Compile-time assertion: *sshpush.Transport satisfies
|
||||
// emitter.AtomicWriter. WriteTraefikDynamic relies on this so the
|
||||
// Traefik dynamic-config file is written atomically (gate C-10).
|
||||
var _ emitter.AtomicWriter = (*sshpush.Transport)(nil)
|
||||
|
||||
func TestTransportSatisfiesAtomicWriter(t *testing.T) {
|
||||
// A trivial runtime check that the type conversion is valid; the
|
||||
// compile-time assertion above is the real test, but this gives
|
||||
// `go test` a function to run.
|
||||
tr := sshpush.NewTransport("/nonexistent", "/nonexistent")
|
||||
var w emitter.AtomicWriter = tr
|
||||
_ = w
|
||||
_ = context.Background()
|
||||
}
|
||||
@@ -0,0 +1,23 @@
|
||||
// Package sshpush implements the v0.9 SSH-push transport layer (REQ-073,
|
||||
// R-001): the CLI on the operator host SSHes to each peer to render files,
|
||||
// apply configs, and run commands. It replaces the v0.8
|
||||
// internal/transport mTLS HTTP layer.
|
||||
//
|
||||
// The Transport reuses one *ssh.Client per peer across multiple
|
||||
// operations within a single CLI invocation (I-B-001), retries transient
|
||||
// failures with exponential backoff (100ms ×2, cap 5s, max 5 attempts —
|
||||
// reimplemented from the v0.8 transport/retry.go pattern, since
|
||||
// internal/transport is deprecated and not imported), applies per-call
|
||||
// timeouts (10s exec, 30s SCP per I-B-001), and fans out to many peers
|
||||
// with bounded concurrency (default 8, errgroup + semaphore).
|
||||
//
|
||||
// Idempotency is content-addressed (C-18): WriteFile / WriteFileIdempotent
|
||||
// compare the remote file's SHA-256 to the local content and skip the
|
||||
// write on match — the SSH-push equivalent of the v0.8 X-Orca-Idempotency-Key.
|
||||
//
|
||||
// Host-key verification reuses proxmox.TOFUHostKeyCallback (D-035), which
|
||||
// reads/writes the known_hosts file (certpaths.KnownHostsPath during the
|
||||
// v0.9 dual-write window; the move to paths.KnownHostsPath happens in
|
||||
// v0.10-P14). The known_hosts file is flock-protected inside the TOFU
|
||||
// callback, so the Transport does NOT re-lock.
|
||||
package sshpush
|
||||
@@ -0,0 +1,113 @@
|
||||
package sshpush
|
||||
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"os"
|
||||
"sync"
|
||||
|
||||
"golang.org/x/sync/errgroup"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/emitter"
|
||||
)
|
||||
|
||||
// DefaultFanoutConcurrency is the default bounded-concurrency limit for
|
||||
// fan-out operations (I-B-001). The Transport.ExecAll and WriteAll
|
||||
// methods use this when the caller does not override it.
|
||||
const DefaultFanoutConcurrency = 8
|
||||
|
||||
// ExecAll runs cmd on all peers in parallel with bounded concurrency
|
||||
// (default 8, I-B-001). Returns per-peer output and per-peer errors. A
|
||||
// nil entry in the errors map means that peer succeeded; the output map
|
||||
// contains that peer's stdout. The returned error is non-nil only if
|
||||
// the fan-out itself failed (e.g., context cancelled before any peer
|
||||
// ran); per-peer failures are in the errors map.
|
||||
func (t *Transport) ExecAll(ctx context.Context, peers []string, cmd string) (map[string][]byte, map[string]error) {
|
||||
return t.ExecAllWithConcurrency(ctx, peers, cmd, DefaultFanoutConcurrency)
|
||||
}
|
||||
|
||||
// ExecAllWithConcurrency is ExecAll with an explicit concurrency limit.
|
||||
// A limit <= 0 uses DefaultFanoutConcurrency.
|
||||
func (t *Transport) ExecAllWithConcurrency(ctx context.Context, peers []string, cmd string, concurrency int) (map[string][]byte, map[string]error) {
|
||||
if concurrency <= 0 {
|
||||
concurrency = DefaultFanoutConcurrency
|
||||
}
|
||||
out := make(map[string][]byte, len(peers))
|
||||
errs := make(map[string]error, len(peers))
|
||||
var mu sync.Mutex
|
||||
g, gctx := errgroup.WithContext(ctx)
|
||||
g.SetLimit(concurrency)
|
||||
for _, p := range peers {
|
||||
peer := p
|
||||
g.Go(func() error {
|
||||
o, err := t.Exec(gctx, peer, cmd)
|
||||
mu.Lock()
|
||||
defer mu.Unlock()
|
||||
if err != nil {
|
||||
errs[peer] = err
|
||||
return nil // per-peer error; do not cancel the group
|
||||
}
|
||||
out[peer] = o
|
||||
return nil
|
||||
})
|
||||
}
|
||||
_ = g.Wait()
|
||||
return out, errs
|
||||
}
|
||||
|
||||
// WriteAll writes the given files to each peer in parallel with bounded
|
||||
// concurrency (default 8, I-B-001). The files map is keyed by peer; each
|
||||
// peer's files are written sequentially (to preserve order and avoid
|
||||
// intra-peer races on shared paths). Returns per-peer errors; a peer
|
||||
// missing from the map or with a nil entry succeeded. The returned
|
||||
// error is non-nil only if the fan-out itself failed (context cancelled).
|
||||
func (t *Transport) WriteAll(ctx context.Context, peers []string, files map[string][]emitter.File) map[string]error {
|
||||
return t.WriteAllWithConcurrency(ctx, peers, files, DefaultFanoutConcurrency)
|
||||
}
|
||||
|
||||
// WriteAllWithConcurrency is WriteAll with an explicit concurrency limit.
|
||||
// A limit <= 0 uses DefaultFanoutConcurrency.
|
||||
func (t *Transport) WriteAllWithConcurrency(ctx context.Context, peers []string, files map[string][]emitter.File, concurrency int) map[string]error {
|
||||
if concurrency <= 0 {
|
||||
concurrency = DefaultFanoutConcurrency
|
||||
}
|
||||
errs := make(map[string]error, len(peers))
|
||||
var mu sync.Mutex
|
||||
g, gctx := errgroup.WithContext(ctx)
|
||||
g.SetLimit(concurrency)
|
||||
for _, p := range peers {
|
||||
peer := p
|
||||
peerFiles := files[peer]
|
||||
g.Go(func() error {
|
||||
for _, f := range peerFiles {
|
||||
mode := parseMode(f.Mode)
|
||||
if _, err := t.WriteFileIdempotent(gctx, peer, f.Path, []byte(f.Content), mode); err != nil {
|
||||
mu.Lock()
|
||||
errs[peer] = fmt.Errorf("sshpush: write %s on %s: %w", f.Path, peer, err)
|
||||
mu.Unlock()
|
||||
return nil // per-peer error; do not cancel the group
|
||||
}
|
||||
}
|
||||
return nil
|
||||
})
|
||||
}
|
||||
_ = g.Wait()
|
||||
return errs
|
||||
}
|
||||
|
||||
// parseMode parses an octal mode string like "0644" into an os.FileMode.
|
||||
// Returns 0644 on parse failure (a safe default for non-executable
|
||||
// config files).
|
||||
func parseMode(s string) os.FileMode {
|
||||
var m uint32
|
||||
for _, r := range s {
|
||||
if r < '0' || r > '7' {
|
||||
return 0o644
|
||||
}
|
||||
m = m<<3 | uint32(r-'0')
|
||||
}
|
||||
if m == 0 {
|
||||
return 0o644
|
||||
}
|
||||
return os.FileMode(m)
|
||||
}
|
||||
@@ -0,0 +1,300 @@
|
||||
package sshpush
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
|
||||
"golang.org/x/crypto/ssh"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/emitter"
|
||||
)
|
||||
|
||||
// --- ExecAll tests ---
|
||||
|
||||
func TestExecAll_AllSucceed(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
// Use one server for all peers (same addr).
|
||||
addr := srv.addr()
|
||||
peers := []string{addr, addr, addr}
|
||||
out, errs := tr.ExecAll(context.Background(), peers, "echo hello")
|
||||
for _, p := range peers {
|
||||
if e, ok := errs[p]; ok && e != nil {
|
||||
t.Errorf("peer %s: %v", p, e)
|
||||
}
|
||||
if string(out[p]) != "hello\n" {
|
||||
t.Errorf("out[%s] = %q, want hello\\n", p, out[p])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestExecAll_OneFails(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
addr := srv.addr()
|
||||
// Peer "bad" returns a permanent error via a mock session.
|
||||
goodPeers := []string{addr}
|
||||
badPeer := "127.0.0.1:1" // unreachable -> transient dial error, retried, fails
|
||||
peers := append(goodPeers, badPeer)
|
||||
out, errs := tr.ExecAll(context.Background(), peers, "echo hello")
|
||||
if string(out[addr]) != "hello\n" {
|
||||
t.Errorf("good peer out = %q, want hello\\n", out[addr])
|
||||
}
|
||||
if errs[badPeer] == nil {
|
||||
t.Error("bad peer should have an error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestExecAll_WithConcurrency(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
addr := srv.addr()
|
||||
peers := []string{addr, addr, addr, addr}
|
||||
out, errs := tr.ExecAllWithConcurrency(context.Background(), peers, "echo hello", 2)
|
||||
for _, p := range peers {
|
||||
if e := errs[p]; e != nil {
|
||||
t.Errorf("peer %s: %v", p, e)
|
||||
}
|
||||
if string(out[p]) != "hello\n" {
|
||||
t.Errorf("out[%s] = %q", p, out[p])
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestExecAll_EmptyPeers(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
out, errs := tr.ExecAll(context.Background(), nil, "echo hello")
|
||||
if len(out) != 0 || len(errs) != 0 {
|
||||
t.Errorf("empty peers: out=%v errs=%v", out, errs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestExecAll_ContextCancelled(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
addr := srv.addr()
|
||||
ctx, cancel := context.WithCancel(context.Background())
|
||||
cancel()
|
||||
out, errs := tr.ExecAll(ctx, []string{addr, addr}, "echo hello")
|
||||
// With a cancelled context, all peers should fail.
|
||||
for _, p := range []string{addr, addr} {
|
||||
if errs[p] == nil && string(out[p]) == "" {
|
||||
// acceptable: either error or no output
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// --- WriteAll tests ---
|
||||
|
||||
func TestWriteAll_AllSucceed(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
addr := srv.addr()
|
||||
files := map[string][]emitter.File{
|
||||
addr: {
|
||||
{Path: "/w/a", Content: "alpha\n", Mode: "0644"},
|
||||
{Path: "/w/b", Content: "beta\n", Mode: "0644"},
|
||||
},
|
||||
}
|
||||
errs := tr.WriteAll(context.Background(), []string{addr}, files)
|
||||
for p, e := range errs {
|
||||
if e != nil {
|
||||
t.Errorf("peer %s: %v", p, e)
|
||||
}
|
||||
}
|
||||
srv.mu.Lock()
|
||||
if srv.files["/w/a"] != "alpha\n" {
|
||||
t.Errorf("file a = %q", srv.files["/w/a"])
|
||||
}
|
||||
if srv.files["/w/b"] != "beta\n" {
|
||||
t.Errorf("file b = %q", srv.files["/w/b"])
|
||||
}
|
||||
srv.mu.Unlock()
|
||||
}
|
||||
|
||||
func TestWriteAll_OnePeerFails(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
addr := srv.addr()
|
||||
bad := "127.0.0.1:1"
|
||||
files := map[string][]emitter.File{
|
||||
addr: {{Path: "/ok/f", Content: "ok\n", Mode: "0644"}},
|
||||
bad: {{Path: "/fail/f", Content: "fail\n", Mode: "0644"}},
|
||||
}
|
||||
errs := tr.WriteAll(context.Background(), []string{addr, bad}, files)
|
||||
if errs[addr] != nil {
|
||||
t.Errorf("good peer should not have error, got %v", errs[addr])
|
||||
}
|
||||
if errs[bad] == nil {
|
||||
t.Error("bad peer should have error")
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteAll_EmptyPeers(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
errs := tr.WriteAll(context.Background(), nil, nil)
|
||||
if len(errs) != 0 {
|
||||
t.Errorf("empty peers: errs=%v", errs)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteAll_WithConcurrency(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
addr := srv.addr()
|
||||
files := map[string][]emitter.File{
|
||||
addr: {{Path: "/c/f", Content: "c\n", Mode: "0644"}},
|
||||
}
|
||||
errs := tr.WriteAllWithConcurrency(context.Background(), []string{addr}, files, 4)
|
||||
for _, e := range errs {
|
||||
if e != nil {
|
||||
t.Errorf("peer err: %v", e)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteAll_Idempotent(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
srv.mu.Lock()
|
||||
srv.files["/i/f"] = "same\n"
|
||||
srv.mu.Unlock()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
addr := srv.addr()
|
||||
files := map[string][]emitter.File{
|
||||
addr: {{Path: "/i/f", Content: "same\n", Mode: "0644"}},
|
||||
}
|
||||
// Capture writes to verify idempotent skip.
|
||||
var writes int64
|
||||
tr.SetSessionFactory(func(c *ssh.Client) (sshSession, error) {
|
||||
return &writeCountingSession{srv: srv, writes: &writes}, nil
|
||||
})
|
||||
// Pre-populate pool.
|
||||
client, err := tr.dial(addr)
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
tr.pool.Store(addr, client)
|
||||
errs := tr.WriteAll(context.Background(), []string{addr}, files)
|
||||
if errs[addr] != nil {
|
||||
t.Errorf("WriteAll err: %v", errs[addr])
|
||||
}
|
||||
// sha256sum returns a hash that matches -> no write command.
|
||||
srv.mu.Lock()
|
||||
content := srv.files["/i/f"]
|
||||
srv.mu.Unlock()
|
||||
if content != "same\n" {
|
||||
t.Errorf("content changed to %q", content)
|
||||
}
|
||||
}
|
||||
|
||||
// writeCountingSession counts how many write commands (mkdir + cat >)
|
||||
// are issued; returns the server's file content for sha256sum.
|
||||
type writeCountingSession struct {
|
||||
srv *fakeSSHServer
|
||||
writes *int64
|
||||
}
|
||||
|
||||
func (w *writeCountingSession) CombinedOutput(cmd string) ([]byte, error) {
|
||||
c := trim(cmd)
|
||||
if startsWith(c, "sha256sum ") {
|
||||
path := unquote(trimPrefix(c, "sha256sum "))
|
||||
path = trimSuffix(path, " 2>/dev/null")
|
||||
w.srv.mu.Lock()
|
||||
content, ok := w.srv.files[path]
|
||||
w.srv.mu.Unlock()
|
||||
if !ok {
|
||||
return []byte(""), nil
|
||||
}
|
||||
sum := sha256HexStr([]byte(content))
|
||||
return []byte(sum + " " + path + "\n"), nil
|
||||
}
|
||||
if startsWith(c, "mkdir -p ") && contains(c, "cat >") {
|
||||
atomic.AddInt64(w.writes, 1)
|
||||
return nil, nil
|
||||
}
|
||||
return nil, nil
|
||||
}
|
||||
func (w *writeCountingSession) Close() error { return nil }
|
||||
|
||||
// Local string helpers to avoid importing strings in a way that
|
||||
// conflicts with the test's existing imports.
|
||||
func trim(s string) string {
|
||||
for len(s) > 0 && (s[0] == ' ' || s[0] == '\t') {
|
||||
s = s[1:]
|
||||
}
|
||||
for len(s) > 0 && (s[len(s)-1] == ' ' || s[len(s)-1] == '\t') {
|
||||
s = s[:len(s)-1]
|
||||
}
|
||||
return s
|
||||
}
|
||||
func startsWith(s, prefix string) bool { return len(s) >= len(prefix) && s[:len(prefix)] == prefix }
|
||||
func contains(s, sub string) bool {
|
||||
return len(sub) == 0 || (len(s) >= len(sub) && indexOf(s, sub) >= 0)
|
||||
}
|
||||
func indexOf(s, sub string) int {
|
||||
for i := 0; i+len(sub) <= len(s); i++ {
|
||||
if s[i:i+len(sub)] == sub {
|
||||
return i
|
||||
}
|
||||
}
|
||||
return -1
|
||||
}
|
||||
func trimPrefix(s, prefix string) string {
|
||||
if startsWith(s, prefix) {
|
||||
return s[len(prefix):]
|
||||
}
|
||||
return s
|
||||
}
|
||||
func trimSuffix(s, suffix string) string {
|
||||
if len(s) >= len(suffix) && s[len(s)-len(suffix):] == suffix {
|
||||
return s[:len(s)-len(suffix)]
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
func TestDefaultFanoutConcurrency(t *testing.T) {
|
||||
if DefaultFanoutConcurrency != 8 {
|
||||
t.Errorf("DefaultFanoutConcurrency = %d, want 8", DefaultFanoutConcurrency)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMode_Fanout(t *testing.T) {
|
||||
// parseMode is in fanout.go; sanity-check here too.
|
||||
for _, tc := range []struct{ in, want string }{
|
||||
{"0644", "644"},
|
||||
{"0755", "755"},
|
||||
{"bad", "644"},
|
||||
} {
|
||||
got := fmt.Sprintf("%o", parseMode(tc.in))
|
||||
if got != tc.want {
|
||||
t.Errorf("parseMode(%q) = %s, want %s", tc.in, got, tc.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
var _ = errors.New
|
||||
@@ -0,0 +1,103 @@
|
||||
package sshpush
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"fmt"
|
||||
"math/rand"
|
||||
"os"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// WriteFileIdempotent writes content to peer:path atomically (write-to-tmp
|
||||
// + mv, REQ-074) with mode, but only if the remote file's SHA-256 differs
|
||||
// from the local content's SHA-256 (C-18 content-addressed idempotency —
|
||||
// the SSH-push equivalent of the v0.8 X-Orca-Idempotency-Key).
|
||||
//
|
||||
// Returns written=true if the file was written, written=false if the
|
||||
// content already matched (skip). The default per-SCP timeout is
|
||||
// SCPTimeout (I-B-001).
|
||||
//
|
||||
// Atomicity: the content is written to a temp file in the same directory
|
||||
// as the target, then `mv`'d into place. The temp file is mode-appended
|
||||
// (e.g. `/etc/orca/foo.conf.orca-tmp-<rand>`) so the rename is atomic on
|
||||
// POSIX filesystems.
|
||||
func (t *Transport) WriteFileIdempotent(ctx context.Context, peer string, path string, content []byte, mode os.FileMode) (bool, error) {
|
||||
localHash := sha256Hex(content)
|
||||
remoteHash, err := t.remoteSHA256(ctx, peer, path)
|
||||
if err == nil && remoteHash != "" && strings.EqualFold(remoteHash, localHash) {
|
||||
return false, nil
|
||||
}
|
||||
if err := t.writeFile(ctx, peer, path, content, mode); err != nil {
|
||||
return false, err
|
||||
}
|
||||
return true, nil
|
||||
}
|
||||
|
||||
// writeFile writes content to peer:path atomically (write-to-tmp + mv).
|
||||
// It writes the content via a single SSH exec (cat heredoc + chmod + mv),
|
||||
// keeping the transfer in one round-trip. The temp file lives next to the
|
||||
// target so the rename is atomic.
|
||||
func (t *Transport) writeFile(ctx context.Context, peer string, path string, content []byte, mode os.FileMode) error {
|
||||
dir, base := splitDir(path)
|
||||
tmpName := fmt.Sprintf(".orca-tmp-%s", randomToken(8))
|
||||
tmpPath := base + "/" + tmpName
|
||||
if dir == "" {
|
||||
tmpPath = tmpName
|
||||
}
|
||||
// Build the remote command: mkdir -p <dir> && cat > <tmp> <<'EOF'
|
||||
// ... EOF && chmod <mode> <tmp> && mv <tmp> <path>. The heredoc
|
||||
// delimiter is per-write random and verified absent from content
|
||||
// to prevent command injection via crafted file content.
|
||||
eof := "ORCA_PUSH_EOF_" + randomToken(16)
|
||||
for strings.Contains(string(content), eof) {
|
||||
eof = "ORCA_PUSH_EOF_" + randomToken(16)
|
||||
}
|
||||
modeStr := fmt.Sprintf("%04o", uint32(mode.Perm()))
|
||||
cmd := fmt.Sprintf(
|
||||
"mkdir -p %s && cat > %s <<'%s'\n%s\n%s\nchmod %s %s && mv -f %s %s",
|
||||
shellQuote(base),
|
||||
shellQuote(tmpPath),
|
||||
eof,
|
||||
string(content),
|
||||
eof,
|
||||
modeStr,
|
||||
shellQuote(tmpPath),
|
||||
shellQuote(tmpPath),
|
||||
shellQuote(path),
|
||||
)
|
||||
execCtx, cancel := context.WithTimeout(ctx, SCPTimeout)
|
||||
defer cancel()
|
||||
if _, err := t.execWithRetry(execCtx, peer, cmd, true); err != nil {
|
||||
return fmt.Errorf("sshpush: write %s: %w", path, err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// sha256Hex returns the lowercase hex SHA-256 digest of b.
|
||||
func sha256Hex(b []byte) string {
|
||||
sum := sha256.Sum256(b)
|
||||
return hex.EncodeToString(sum[:])
|
||||
}
|
||||
|
||||
// splitDir returns the directory and the directory itself (for mkdir).
|
||||
// For "/etc/orca/foo.conf" it returns ("/etc/orca", "/etc/orca"). For
|
||||
// "foo.conf" it returns ("", ".").
|
||||
func splitDir(path string) (dir, base string) {
|
||||
idx := strings.LastIndex(path, "/")
|
||||
if idx < 0 {
|
||||
return "", "."
|
||||
}
|
||||
return path[:idx], path[:idx]
|
||||
}
|
||||
|
||||
// randomToken returns a random hex token of the given byte length. Used
|
||||
// for temp-file naming to avoid collisions under parallel fan-out.
|
||||
func randomToken(n int) string {
|
||||
b := make([]byte, n)
|
||||
for i := range b {
|
||||
b[i] = byte(rand.Intn(256))
|
||||
}
|
||||
return hex.EncodeToString(b)
|
||||
}
|
||||
@@ -0,0 +1,348 @@
|
||||
package sshpush
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"errors"
|
||||
"fmt"
|
||||
"os"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"golang.org/x/crypto/ssh"
|
||||
)
|
||||
|
||||
func TestSha256Hex(t *testing.T) {
|
||||
got := sha256Hex([]byte("hello"))
|
||||
want := "2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824"
|
||||
if got != want {
|
||||
t.Errorf("sha256Hex = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSplitDir(t *testing.T) {
|
||||
dir, base := splitDir("/etc/orca/foo.conf")
|
||||
if dir != "/etc/orca" || base != "/etc/orca" {
|
||||
t.Errorf("splitDir(/etc/orca/foo.conf) = (%q,%q), want (/etc/orca,/etc/orca)", dir, base)
|
||||
}
|
||||
dir, base = splitDir("foo.conf")
|
||||
if dir != "" || base != "." {
|
||||
t.Errorf("splitDir(foo.conf) = (%q,%q), want (\"\",.)", dir, base)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRandomToken(t *testing.T) {
|
||||
a := randomToken(8)
|
||||
b := randomToken(8)
|
||||
if a == b {
|
||||
t.Error("randomToken returned same value twice")
|
||||
}
|
||||
if len(a) != 16 { // 8 bytes hex = 16 chars
|
||||
t.Errorf("randomToken(8) len = %d, want 16", len(a))
|
||||
}
|
||||
}
|
||||
|
||||
// --- WriteFileIdempotent tests via the fake SSH server ---
|
||||
|
||||
func TestWriteFileIdempotent_WritesWhenFileMissing(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
content := []byte("first content\n")
|
||||
written, err := tr.WriteFileIdempotent(context.Background(), srv.addr(), "/etc/orca/a.conf", content, 0o644)
|
||||
if err != nil {
|
||||
t.Fatalf("WriteFileIdempotent: %v", err)
|
||||
}
|
||||
if !written {
|
||||
t.Error("written=false, want true (file was missing)")
|
||||
}
|
||||
srv.mu.Lock()
|
||||
got := srv.files["/etc/orca/a.conf"]
|
||||
srv.mu.Unlock()
|
||||
if got != string(content) {
|
||||
t.Errorf("remote file = %q, want %q", got, string(content))
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteFileIdempotent_SkipsWhenContentMatches(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
content := []byte("same content\n")
|
||||
srv.mu.Lock()
|
||||
srv.files["/etc/orca/b.conf"] = string(content)
|
||||
srv.mu.Unlock()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
written, err := tr.WriteFileIdempotent(context.Background(), srv.addr(), "/etc/orca/b.conf", content, 0o644)
|
||||
if err != nil {
|
||||
t.Fatalf("WriteFileIdempotent: %v", err)
|
||||
}
|
||||
if written {
|
||||
t.Error("written=true, want false (content matched)")
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteFileIdempotent_WritesWhenContentDiffers(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
srv.mu.Lock()
|
||||
srv.files["/etc/orca/c.conf"] = "old content\n"
|
||||
srv.mu.Unlock()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
newContent := []byte("new content\n")
|
||||
written, err := tr.WriteFileIdempotent(context.Background(), srv.addr(), "/etc/orca/c.conf", newContent, 0o644)
|
||||
if err != nil {
|
||||
t.Fatalf("WriteFileIdempotent: %v", err)
|
||||
}
|
||||
if !written {
|
||||
t.Error("written=false, want true (content differed)")
|
||||
}
|
||||
srv.mu.Lock()
|
||||
got := srv.files["/etc/orca/c.conf"]
|
||||
srv.mu.Unlock()
|
||||
if got != string(newContent) {
|
||||
t.Errorf("remote file = %q, want %q", got, string(newContent))
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteFile_DelegatesToIdempotent(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
content := []byte("delegated\n")
|
||||
if err := tr.WriteFile(context.Background(), srv.addr(), "/etc/orca/d.conf", content, 0o600); err != nil {
|
||||
t.Fatalf("WriteFile: %v", err)
|
||||
}
|
||||
srv.mu.Lock()
|
||||
got := srv.files["/etc/orca/d.conf"]
|
||||
srv.mu.Unlock()
|
||||
if got != string(content) {
|
||||
t.Errorf("remote file = %q, want %q", got, string(content))
|
||||
}
|
||||
}
|
||||
|
||||
// --- Pure-logic idempotency tests via mock session (no SSH server) ---
|
||||
|
||||
// mockHashSession returns the hash of the file matching the sha256sum
|
||||
// command's path argument; for write commands (mkdir + cat >), it
|
||||
// records the write. This lets us test the idempotency decision logic
|
||||
// without a real SSH server.
|
||||
type mockHashSession struct {
|
||||
files map[string]string
|
||||
out []byte
|
||||
err error
|
||||
cmd string
|
||||
writeHook func(cmd string)
|
||||
}
|
||||
|
||||
func (m *mockHashSession) CombinedOutput(cmd string) ([]byte, error) {
|
||||
m.cmd = cmd
|
||||
c := strings.TrimSpace(cmd)
|
||||
if strings.HasPrefix(c, "sha256sum ") {
|
||||
rest := strings.TrimSpace(strings.TrimPrefix(c, "sha256sum "))
|
||||
rest = strings.TrimSuffix(rest, " 2>/dev/null")
|
||||
rest = strings.TrimSpace(rest)
|
||||
path := unquote(rest)
|
||||
content, ok := m.files[path]
|
||||
if !ok {
|
||||
return []byte(""), nil
|
||||
}
|
||||
sum := sha256.Sum256([]byte(content))
|
||||
return []byte(hex.EncodeToString(sum[:]) + " " + path + "\n"), nil
|
||||
}
|
||||
if strings.HasPrefix(c, "mkdir -p ") && strings.Contains(c, "cat >") {
|
||||
if m.writeHook != nil {
|
||||
m.writeHook(c)
|
||||
}
|
||||
return nil, nil
|
||||
}
|
||||
return m.out, m.err
|
||||
}
|
||||
func (m *mockHashSession) Close() error { return nil }
|
||||
|
||||
func TestWriteFileIdempotent_MockSkip(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
files := map[string]string{"/x/f": "match"}
|
||||
var writes int
|
||||
tr.SetSessionFactory(func(c *ssh.Client) (sshSession, error) {
|
||||
return &mockHashSession{
|
||||
files: files,
|
||||
writeHook: func(string) { writes++ },
|
||||
}, nil
|
||||
})
|
||||
client, err := tr.dial(srv.addr())
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
tr.pool.Store(srv.addr(), client)
|
||||
written, err := tr.WriteFileIdempotent(context.Background(), srv.addr(), "/x/f", []byte("match"), 0o644)
|
||||
if err != nil {
|
||||
t.Fatalf("WriteFileIdempotent: %v", err)
|
||||
}
|
||||
if written {
|
||||
t.Error("written=true, want false (hash matched)")
|
||||
}
|
||||
if writes != 0 {
|
||||
t.Errorf("writes = %d, want 0 (no write on hash match)", writes)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteFileIdempotent_MockWrite(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
files := map[string]string{"/x/f": "old"}
|
||||
var writes int
|
||||
tr.SetSessionFactory(func(c *ssh.Client) (sshSession, error) {
|
||||
return &mockHashSession{
|
||||
files: files,
|
||||
writeHook: func(string) { writes++ },
|
||||
}, nil
|
||||
})
|
||||
client, err := tr.dial(srv.addr())
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
tr.pool.Store(srv.addr(), client)
|
||||
written, err := tr.WriteFileIdempotent(context.Background(), srv.addr(), "/x/f", []byte("new"), 0o644)
|
||||
if err != nil {
|
||||
t.Fatalf("WriteFileIdempotent: %v", err)
|
||||
}
|
||||
if !written {
|
||||
t.Error("written=false, want true (hash differed)")
|
||||
}
|
||||
if writes != 1 {
|
||||
t.Errorf("writes = %d, want 1", writes)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteFileIdempotent_MockMissingFile(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
files := map[string]string{}
|
||||
var writes int
|
||||
tr.SetSessionFactory(func(c *ssh.Client) (sshSession, error) {
|
||||
return &mockHashSession{
|
||||
files: files,
|
||||
writeHook: func(string) { writes++ },
|
||||
}, nil
|
||||
})
|
||||
client, err := tr.dial(srv.addr())
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
tr.pool.Store(srv.addr(), client)
|
||||
written, err := tr.WriteFileIdempotent(context.Background(), srv.addr(), "/x/new", []byte("fresh"), 0o644)
|
||||
if err != nil {
|
||||
t.Fatalf("WriteFileIdempotent: %v", err)
|
||||
}
|
||||
if !written {
|
||||
t.Error("written=false, want true (file missing)")
|
||||
}
|
||||
if writes != 1 {
|
||||
t.Errorf("writes = %d, want 1", writes)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRemoteSHA256_ParsesDigest(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
srv.mu.Lock()
|
||||
srv.files["/x/h"] = "abc"
|
||||
srv.mu.Unlock()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
got, err := tr.remoteSHA256(context.Background(), srv.addr(), "/x/h")
|
||||
if err != nil {
|
||||
t.Fatalf("remoteSHA256: %v", err)
|
||||
}
|
||||
want := fmt.Sprintf("%x", sha256.Sum256([]byte("abc")))
|
||||
if got != want {
|
||||
t.Errorf("remoteSHA256 = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRemoteSHA256_MissingFile(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
got, err := t_remoteSHA256_noErr(t, tr, srv.addr(), "/missing")
|
||||
if err != nil {
|
||||
t.Fatalf("remoteSHA256: %v", err)
|
||||
}
|
||||
if got != "" {
|
||||
t.Errorf("remoteSHA256 = %q, want empty for missing file", got)
|
||||
}
|
||||
}
|
||||
|
||||
func t_remoteSHA256_noErr(t *testing.T, tr *Transport, peer, path string) (string, error) {
|
||||
t.Helper()
|
||||
return tr.remoteSHA256(context.Background(), peer, path)
|
||||
}
|
||||
|
||||
func TestWriteFile_ModeApplied(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
if err := tr.WriteFile(context.Background(), srv.addr(), "/m/f", []byte("mode"), 0o755); err != nil {
|
||||
t.Fatalf("WriteFile: %v", err)
|
||||
}
|
||||
srv.mu.Lock()
|
||||
got := srv.files["/m/f"]
|
||||
srv.mu.Unlock()
|
||||
if got != "mode" {
|
||||
t.Errorf("content = %q, want mode", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestWriteFileIdempotent_ExecError(t *testing.T) {
|
||||
srv := newFakeSSHServer(t)
|
||||
defer srv.close()
|
||||
tr := realTransport(t, srv)
|
||||
defer tr.Close()
|
||||
tr.SetSessionFactory(func(c *ssh.Client) (sshSession, error) {
|
||||
return &mockSession{err: fmt.Errorf("%w: boom", ErrPermanent)}, nil
|
||||
})
|
||||
client, err := tr.dial(srv.addr())
|
||||
if err != nil {
|
||||
t.Fatalf("dial: %v", err)
|
||||
}
|
||||
tr.pool.Store(srv.addr(), client)
|
||||
_, err = tr.WriteFileIdempotent(context.Background(), srv.addr(), "/x/f", []byte("z"), 0o644)
|
||||
if err == nil {
|
||||
t.Fatal("expected error, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseMode(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
in string
|
||||
want os.FileMode
|
||||
}{
|
||||
{"0644", 0o644},
|
||||
{"0755", 0o755},
|
||||
{"0600", 0o600},
|
||||
{"bad", 0o644},
|
||||
{"", 0o644},
|
||||
{"0", 0o644},
|
||||
} {
|
||||
got := parseMode(tc.in)
|
||||
if got != tc.want {
|
||||
t.Errorf("parseMode(%q) = %o, want %o", tc.in, got, tc.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
var _ = errors.New
|
||||
@@ -0,0 +1,483 @@
|
||||
package sshpush
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"math/rand"
|
||||
"net"
|
||||
"os"
|
||||
"strings"
|
||||
"sync"
|
||||
"time"
|
||||
|
||||
"golang.org/x/crypto/ssh"
|
||||
|
||||
"git.cloudinit.dev/coreci/orca/internal/proxmox"
|
||||
)
|
||||
|
||||
// Default timeouts and retry parameters (REQ-073, I-B-001).
|
||||
const (
|
||||
// ExecTimeout is the default per-exec timeout for a single SSH
|
||||
// command (I-B-001).
|
||||
ExecTimeout = 10 * time.Second
|
||||
// SCPTimeout is the default per-SCP timeout for a single file
|
||||
// transfer (I-B-001).
|
||||
SCPTimeout = 30 * time.Second
|
||||
// DialTimeout is the default SSH dial timeout.
|
||||
DialTimeout = 15 * time.Second
|
||||
// RetryInitial is the first backoff interval (v0.8 transport/retry.go).
|
||||
RetryInitial = 100 * time.Millisecond
|
||||
// RetryMax is the cap on backoff between attempts.
|
||||
RetryMax = 5 * time.Second
|
||||
// RetryMaxAttempts is the total attempt count (including the first).
|
||||
RetryMaxAttempts = 5
|
||||
)
|
||||
|
||||
// Sentinel errors. ErrTransient marks a transient failure worth
|
||||
// retrying; ErrPermanent marks a non-retryable failure (auth, host-key
|
||||
// mismatch, validation). These mirror the v0.8 transport sentinels
|
||||
// (reimplemented here since internal/transport is not imported).
|
||||
var (
|
||||
ErrTransient = errors.New("sshpush: transient error")
|
||||
ErrPermanent = errors.New("sshpush: permanent error")
|
||||
ErrNotConnected = errors.New("sshpush: not connected")
|
||||
)
|
||||
|
||||
// Transport is the SSH-push transport (REQ-073). It reuses one
|
||||
// *ssh.Client per peer across multiple operations within a single CLI
|
||||
// invocation (I-B-001). The zero value is NOT usable; construct one with
|
||||
// NewTransport.
|
||||
type Transport struct {
|
||||
// pool caches *ssh.Client per peer address ("host:port").
|
||||
pool sync.Map
|
||||
// keyPath is the SSH private key path (Ed25519, D-037).
|
||||
keyPath string
|
||||
// knownHostsPath is the v0.9 known_hosts path (paths.KnownHostsPath()
|
||||
// = ClusterDir()/known_hosts). It is stored for the v0.10-P14 migration
|
||||
// when proxmox.TOFUHostKeyCallback will accept a path parameter; today
|
||||
// the callback reads certpaths.KnownHostsPath() (the v0.8 flat layout)
|
||||
// directly, so this field is not yet read by dial(). Tests set
|
||||
// $ORCA_HOME so certpaths.KnownHostsPath() resolves under the temp dir.
|
||||
knownHostsPath string
|
||||
// user is the remote SSH user (default "orca", D-037).
|
||||
user string
|
||||
// signer is the parsed SSH private key signer, set lazily on first
|
||||
// dial.
|
||||
signer ssh.Signer
|
||||
signErr error
|
||||
// signerOnce guards signer initialization.
|
||||
signerOnce sync.Once
|
||||
|
||||
// dialer is the SSH dialer. Tests override it to inject a mock
|
||||
// server. The default uses ssh.DialContext via the context-aware
|
||||
// wrapper.
|
||||
dialer sshDialer
|
||||
|
||||
// sessionFactory returns a new session for a given client. Tests
|
||||
// override it to inject mock sessions without a real *ssh.Client.
|
||||
// When nil, the default (*ssh.Client).NewSession is used.
|
||||
sessionFactory func(*ssh.Client) (sshSession, error)
|
||||
|
||||
// mu guards the closed flag (pool iteration is sync.Map.Range).
|
||||
closed bool
|
||||
mu sync.Mutex
|
||||
}
|
||||
|
||||
// sshSession is the minimal *ssh.Session surface the transport uses.
|
||||
// It lets tests substitute a mock without a real SSH server.
|
||||
type sshSession interface {
|
||||
CombinedOutput(cmd string) ([]byte, error)
|
||||
Close() error
|
||||
}
|
||||
|
||||
// sshDialer is the SSH dialer interface (mirrors proxmox.sshDialerType).
|
||||
// The default uses ssh.Dial; tests inject mocks that return a fake
|
||||
// *ssh.Client or an error.
|
||||
type sshDialer interface {
|
||||
DialContext(ctx context.Context, network, addr string, config *ssh.ClientConfig) (*ssh.Client, error)
|
||||
}
|
||||
|
||||
// defaultSSHDialer wraps ssh.Dial with a context-aware connect timeout.
|
||||
type defaultSSHDialer struct{}
|
||||
|
||||
func (defaultSSHDialer) DialContext(ctx context.Context, network, addr string, config *ssh.ClientConfig) (*ssh.Client, error) {
|
||||
d := net.Dialer{Timeout: config.Timeout}
|
||||
if d.Timeout == 0 {
|
||||
d.Timeout = DialTimeout
|
||||
}
|
||||
conn, err := d.DialContext(ctx, network, addr)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sshConn, chans, reqs, err := ssh.NewClientConn(conn, addr, config)
|
||||
if err != nil {
|
||||
_ = conn.Close()
|
||||
return nil, err
|
||||
}
|
||||
return ssh.NewClient(sshConn, chans, reqs), nil
|
||||
}
|
||||
|
||||
// NewTransport returns a Transport configured with the given SSH
|
||||
// private key path and known_hosts path. The known_hosts path is the v0.9
|
||||
// location (paths.KnownHostsPath); it is stored for the v0.10-P14
|
||||
// migration when the TOFU callback will accept a path parameter. Today
|
||||
// dial() delegates host-key verification to proxmox.TOFUHostKeyCallback,
|
||||
// which reads certpaths.KnownHostsPath() (the v0.8 flat layout under
|
||||
// $ORCA_HOME) directly — so callers must ensure $ORCA_HOME points at the
|
||||
// cluster root (the CLI sets this up). The remote user defaults to
|
||||
// "orca" (D-037); override with SetUser. The dialer defaults to the
|
||||
// real ssh.Dial-based dialer; tests call SetDialer to inject a mock.
|
||||
func NewTransport(keyPath, knownHostsPath string) *Transport {
|
||||
return &Transport{
|
||||
keyPath: keyPath,
|
||||
knownHostsPath: knownHostsPath,
|
||||
user: "orca",
|
||||
dialer: defaultSSHDialer{},
|
||||
}
|
||||
}
|
||||
|
||||
// SetUser overrides the remote SSH user (default "orca").
|
||||
func (t *Transport) SetUser(user string) {
|
||||
if user != "" {
|
||||
t.user = user
|
||||
}
|
||||
}
|
||||
|
||||
// SetDialer overrides the SSH dialer (for tests).
|
||||
func (t *Transport) SetDialer(d sshDialer) {
|
||||
if d != nil {
|
||||
t.dialer = d
|
||||
}
|
||||
}
|
||||
|
||||
// SetSessionFactory overrides the session factory (for tests). The
|
||||
// factory is called per-exec/write/read to obtain a fresh session; it
|
||||
// must close the session when the test mock is done, or the transport
|
||||
// will call Close on the returned session.
|
||||
func (t *Transport) SetSessionFactory(f func(*ssh.Client) (sshSession, error)) {
|
||||
t.sessionFactory = f
|
||||
}
|
||||
|
||||
// dial returns the cached *ssh.Client for peer, dialing and caching on
|
||||
// first use (I-B-001 connection pooling). Returns an error if the dial
|
||||
// fails or the transport is closed.
|
||||
func (t *Transport) dial(peer string) (*ssh.Client, error) {
|
||||
t.mu.Lock()
|
||||
if t.closed {
|
||||
t.mu.Unlock()
|
||||
return nil, ErrPermanent
|
||||
}
|
||||
t.mu.Unlock()
|
||||
if c, ok := t.pool.Load(peer); ok {
|
||||
return c.(*ssh.Client), nil
|
||||
}
|
||||
// Lazily parse the private key signer (once across all dials).
|
||||
t.signerOnce.Do(func() {
|
||||
keyBytes, err := os.ReadFile(t.keyPath)
|
||||
if err != nil {
|
||||
t.signErr = fmt.Errorf("sshpush: read key %s: %w", t.keyPath, err)
|
||||
return
|
||||
}
|
||||
s, err := ssh.ParsePrivateKey(keyBytes)
|
||||
if err != nil {
|
||||
t.signErr = fmt.Errorf("sshpush: parse key: %w", err)
|
||||
return
|
||||
}
|
||||
t.signer = s
|
||||
})
|
||||
if t.signErr != nil {
|
||||
return nil, t.signErr
|
||||
}
|
||||
// Host-key verification reuses the v0.8 TOFU wrapper (D-035). The
|
||||
// known_hosts file is flock-protected inside the callback on
|
||||
// first-connect capture, so we do NOT re-lock here.
|
||||
cb, err := proxmox.TOFUHostKeyCallback(peer, nil)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("sshpush: host-key callback: %w", err)
|
||||
}
|
||||
config := &ssh.ClientConfig{
|
||||
User: t.user,
|
||||
Auth: []ssh.AuthMethod{ssh.PublicKeys(t.signer)},
|
||||
HostKeyCallback: cb,
|
||||
Timeout: DialTimeout,
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), DialTimeout)
|
||||
defer cancel()
|
||||
client, err := t.dialer.DialContext(ctx, "tcp", peer, config)
|
||||
if err != nil {
|
||||
return nil, classifyDialErr(err)
|
||||
}
|
||||
// Race: two goroutines dialing the same peer concurrently both
|
||||
// create a client. Last-wins; the loser is closed. This is rare
|
||||
// (dial is rare and the pool hit short-circuits) and harmless.
|
||||
if existing, loaded := t.pool.LoadOrStore(peer, client); loaded {
|
||||
_ = client.Close()
|
||||
return existing.(*ssh.Client), nil
|
||||
}
|
||||
return client, nil
|
||||
}
|
||||
|
||||
// Exec runs cmd on peer over SSH and returns its combined output. The
|
||||
// default per-exec timeout is ExecTimeout (I-B-001); override by
|
||||
// passing a context with a shorter deadline. Transient failures are
|
||||
// retried with exponential backoff (100ms ×2, cap 5s, max 5 attempts —
|
||||
// the v0.8 transport/retry.go pattern, reimplemented here).
|
||||
func (t *Transport) Exec(ctx context.Context, peer string, cmd string) ([]byte, error) {
|
||||
return t.execWithRetry(ctx, peer, cmd, true)
|
||||
}
|
||||
|
||||
// execWithRetry runs the exec with retry. exec is treated as
|
||||
// idempotent (read-only) for retry purposes; the idempotency helpers
|
||||
// (WriteFileIdempotent) handle writes.
|
||||
func (t *Transport) execWithRetry(ctx context.Context, peer string, cmd string, idempotent bool) ([]byte, error) {
|
||||
var lastErr error
|
||||
for attempt := 1; attempt <= RetryMaxAttempts; attempt++ {
|
||||
if err := ctx.Err(); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
out, err := t.execOnce(ctx, peer, cmd)
|
||||
if err == nil {
|
||||
return out, nil
|
||||
}
|
||||
if errors.Is(err, ErrPermanent) {
|
||||
return nil, err
|
||||
}
|
||||
lastErr = err
|
||||
if attempt == RetryMaxAttempts {
|
||||
break
|
||||
}
|
||||
if !isTransient(err) {
|
||||
return nil, err
|
||||
}
|
||||
wait := backoff(RetryInitial, RetryMax, attempt)
|
||||
timer := time.NewTimer(wait)
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
timer.Stop()
|
||||
return nil, ctx.Err()
|
||||
case <-timer.C:
|
||||
}
|
||||
}
|
||||
return nil, lastErr
|
||||
}
|
||||
|
||||
// execOnce runs the command a single time against peer.
|
||||
func (t *Transport) execOnce(ctx context.Context, peer string, cmd string) ([]byte, error) {
|
||||
client, err := t.dial(peer)
|
||||
if err != nil {
|
||||
return nil, classifyDialErr(err)
|
||||
}
|
||||
sess, err := t.newSession(client)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("sshpush: new session: %w", err)
|
||||
}
|
||||
defer sess.Close()
|
||||
type result struct {
|
||||
out []byte
|
||||
err error
|
||||
}
|
||||
ch := make(chan result, 1)
|
||||
go func() {
|
||||
out, err := sess.CombinedOutput(cmd)
|
||||
ch <- result{out, err}
|
||||
}()
|
||||
timeout := ExecTimeout
|
||||
if dl, ok := ctx.Deadline(); ok {
|
||||
if remaining := time.Until(dl); remaining > 0 && remaining < timeout {
|
||||
timeout = remaining
|
||||
}
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return nil, ctx.Err()
|
||||
case <-time.After(timeout):
|
||||
return nil, fmt.Errorf("sshpush: exec timeout after %s: %w", timeout, ErrTransient)
|
||||
case r := <-ch:
|
||||
if r.err != nil {
|
||||
return r.out, classifyExecErr(r.err)
|
||||
}
|
||||
return r.out, nil
|
||||
}
|
||||
}
|
||||
|
||||
// newSession returns a session for client, using the override factory
|
||||
// when set (tests), otherwise the real *ssh.Client.NewSession.
|
||||
func (t *Transport) newSession(client *ssh.Client) (sshSession, error) {
|
||||
if t.sessionFactory != nil {
|
||||
return t.sessionFactory(client)
|
||||
}
|
||||
s, err := client.NewSession()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return &realSession{Session: s}, nil
|
||||
}
|
||||
|
||||
// realSession wraps *ssh.Session to satisfy the sshSession interface.
|
||||
type realSession struct {
|
||||
*ssh.Session
|
||||
}
|
||||
|
||||
func (r *realSession) CombinedOutput(cmd string) ([]byte, error) {
|
||||
return r.Session.CombinedOutput(cmd)
|
||||
}
|
||||
|
||||
// WriteFile SCPs content to peer:path atomically (write-to-tmp + mv,
|
||||
// REQ-074). The default per-SCP timeout is SCPTimeout (I-B-001).
|
||||
// Idempotency: if the file already exists with the same SHA-256, the
|
||||
// write is skipped (C-18). Use WriteFileIdempotent for the explicit
|
||||
// written/skipped result.
|
||||
func (t *Transport) WriteFile(ctx context.Context, peer string, path string, content []byte, mode os.FileMode) error {
|
||||
_, err := t.WriteFileIdempotent(ctx, peer, path, content, mode)
|
||||
return err
|
||||
}
|
||||
|
||||
// ReadFile reads the file at peer:path via SSH cat.
|
||||
func (t *Transport) ReadFile(ctx context.Context, peer string, path string) ([]byte, error) {
|
||||
cmd := fmt.Sprintf("cat %s", shellQuote(path))
|
||||
out, err := t.Exec(ctx, peer, cmd)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
|
||||
// Close closes all pooled SSH clients (REQ-073). Safe to call
|
||||
// multiple times; subsequent calls are no-ops.
|
||||
func (t *Transport) Close() error {
|
||||
t.mu.Lock()
|
||||
if t.closed {
|
||||
t.mu.Unlock()
|
||||
return nil
|
||||
}
|
||||
t.closed = true
|
||||
t.mu.Unlock()
|
||||
var firstErr error
|
||||
t.pool.Range(func(key, value any) bool {
|
||||
if c, ok := value.(*ssh.Client); ok {
|
||||
if err := c.Close(); err != nil && firstErr == nil {
|
||||
firstErr = err
|
||||
}
|
||||
}
|
||||
t.pool.Delete(key)
|
||||
return true
|
||||
})
|
||||
return firstErr
|
||||
}
|
||||
|
||||
// backoff returns the wait duration for the n-th attempt (1-indexed).
|
||||
// Formula: min(Initial * 2^(n-1), Max), with up to 25% jitter (matches
|
||||
// v0.8 transport/retry.go).
|
||||
func backoff(initial, max time.Duration, n int) time.Duration {
|
||||
d := initial
|
||||
for i := 1; i < n; i++ {
|
||||
d *= 2
|
||||
if d > max {
|
||||
d = max
|
||||
break
|
||||
}
|
||||
}
|
||||
if d <= 0 {
|
||||
return 0
|
||||
}
|
||||
jitter := time.Duration(rand.Int63n(int64(d) / 2))
|
||||
d = d - d/4 + jitter
|
||||
if d < 0 {
|
||||
d = 0
|
||||
}
|
||||
return d
|
||||
}
|
||||
|
||||
// isTransient reports whether err looks like a transient failure worth
|
||||
// retrying (mirrors v0.8 transport.IsTransient, reimplemented here).
|
||||
func isTransient(err error) bool {
|
||||
if err == nil {
|
||||
return false
|
||||
}
|
||||
if errors.Is(err, ErrTransient) {
|
||||
return true
|
||||
}
|
||||
if errors.Is(err, ErrPermanent) {
|
||||
return false
|
||||
}
|
||||
s := err.Error()
|
||||
for _, sub := range []string{
|
||||
"connection refused", "i/o timeout", "EOF",
|
||||
"no such host", "connection reset", "timeout",
|
||||
"deadline exceeded", "temporarily unavailable",
|
||||
} {
|
||||
if strings.Contains(s, sub) {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
|
||||
// classifyDialErr converts a raw ssh.Dial error into a transport error
|
||||
// (transient vs permanent). Auth failures and host-key mismatches are
|
||||
// permanent; everything else is transient.
|
||||
func classifyDialErr(err error) error {
|
||||
if err == nil {
|
||||
return nil
|
||||
}
|
||||
s := err.Error()
|
||||
if strings.Contains(s, "unable to authenticate") || strings.Contains(s, "handshake failed") {
|
||||
return fmt.Errorf("%w: %v", ErrPermanent, err)
|
||||
}
|
||||
if strings.Contains(s, "host key") && strings.Contains(s, "mismatch") {
|
||||
return fmt.Errorf("%w: %v", ErrPermanent, err)
|
||||
}
|
||||
if strings.Contains(s, "knownhosts") {
|
||||
return fmt.Errorf("%w: %v", ErrPermanent, err)
|
||||
}
|
||||
return fmt.Errorf("%w: %v", ErrTransient, err)
|
||||
}
|
||||
|
||||
// classifyExecErr converts a raw session exec error into a transport
|
||||
// error. Non-zero exit codes are NOT transient (the command ran; the
|
||||
// failure is logical, not network). Session-creation failures and
|
||||
// network-level errors are transient.
|
||||
func classifyExecErr(err error) error {
|
||||
if err == nil {
|
||||
return nil
|
||||
}
|
||||
var exitErr *ssh.ExitError
|
||||
if errors.As(err, &exitErr) {
|
||||
return fmt.Errorf("%w: exit %d", ErrPermanent, exitErr.ExitStatus())
|
||||
}
|
||||
s := err.Error()
|
||||
for _, sub := range []string{"EOF", "session closed", "channel closed"} {
|
||||
if strings.Contains(s, sub) {
|
||||
return fmt.Errorf("%w: %v", ErrTransient, err)
|
||||
}
|
||||
}
|
||||
return fmt.Errorf("%w: %v", ErrPermanent, err)
|
||||
}
|
||||
|
||||
// shellQuote single-quotes a path for safe shell interpolation. It
|
||||
// escapes embedded single-quotes via the standard '\” idiom.
|
||||
func shellQuote(s string) string {
|
||||
return "'" + strings.ReplaceAll(s, "'", "'\\''") + "'"
|
||||
}
|
||||
|
||||
// remoteSHA256 returns the SHA-256 of the file at peer:path via SSH
|
||||
// `sha256sum`, or ("", error) if the file is missing or the command
|
||||
// fails. The returned hash is the hex digest (lowercase, no filename).
|
||||
func (t *Transport) remoteSHA256(ctx context.Context, peer string, path string) (string, error) {
|
||||
cmd := fmt.Sprintf("sha256sum %s 2>/dev/null", shellQuote(path))
|
||||
out, err := t.execWithRetry(ctx, peer, cmd, true)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
out = bytes.TrimSpace(out)
|
||||
if len(out) == 0 {
|
||||
return "", nil
|
||||
}
|
||||
fields := strings.Fields(string(out))
|
||||
if len(fields) == 0 {
|
||||
return "", nil
|
||||
}
|
||||
return fields[0], nil
|
||||
}
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user