Files
orca/docs/uat.md
T
Jon Chery 0424f8ce02 feat(init): interactive remote pre-staging via ssh-copy-id
orca init now interactively prompts for remote host addresses and runs
ssh-copy-id automatically (password prompt passes through to the
operator). This makes orca init the single entry point — no manual
pre-staging of SSH keys required.

- Interactive: enter host addresses (one per line, empty line to finish)
- ssh-copy-id deploys the orca public key to each host
- Skipped in --json mode (non-interactive)
- Idempotent: re-running init can stage additional hosts

Also fixed: install.sh defaults to /usr/local/bin (on PATH for all users).
Non-root without sudo falls back to ~/.local/bin + auto-adds to .bashrc.
2026-08-10 17:37:42 +00:00

361 lines
10 KiB
Markdown

# Orca User Acceptance Testing (UAT) Plan
**Version**: v0.13 (production hardening round 2)
**Gate**: v1.0.0 production-ready tag is deferred until this UAT passes
**Signoff**: run `scripts/uat-signoff.sh` on the lead node and paste the output back
## Prerequisites
### Hardware
| Role | OS | Requirements |
|------|-----|-------------|
| **lead** | Ubuntu 22.04 LTS | Operator laptop or VM; SSH key; `orca` binary (built from v0.13 tag) |
| **pve01** | Proxmox VE 8/9 | Bare-metal or nested; SSH root access; orca SSH key pre-staged |
| **worker01** | Ubuntu 22.04 LTS | VM or bare-metal; SSH root access; orca SSH key pre-staged |
### Alternative topology (3x Ubuntu, no Proxmox)
If a Proxmox host is unavailable, run the UAT with 3x Ubuntu hosts.
Use `--type linux` for all remote nodes. Proxmox-specific claims
(`doctor proxmox`, PVE role, sudoers) are **skipped** in this path.
The signoff script reports exercised vs. skipped claims.
## Step-by-step UAT
### Step 1: Install orca + initialize the cluster
Install orca (1-liner):
```sh
curl -fsSL https://git.cloudinit.dev/coreci/orca/raw/branch/main/scripts/install.sh | bash
```
Initialize the cluster:
```sh
export ORCA_HOME=~/orca-uat
orca init
```
**Expected**: orca init creates:
- CA cert + server cert
- SSH keypair (orca_ssh_key + orca_ssh_key.pub)
- known_hosts file (empty, for TOFU capture)
- Master key (for secrets encryption)
- Traefik data-plane ingress (binary + systemd unit + config)
- Localhost node registered
**Pre-staging remote nodes**: `orca init` interactively prompts for remote
host addresses and runs `ssh-copy-id` automatically (password prompt passes
through). Enter each host (pve01, worker01) when prompted, or press Enter to
skip. The orca public key is deployed to each host; TOFU host-key capture is
automatic on the first `orca node join` — no manual fingerprint pinning needed.
### Step 2: Onboard the Proxmox host
```sh
orca node join --type proxmox \
--host pve01 \
--ssh-user root
```
**Expected**: SSH bootstrap succeeds, orca user created, PVE role assigned, node registered as `ready` with `kind=proxmox`.
### Step 3: Onboard the Ubuntu worker
```sh
orca node join --type linux \
--host worker01 \
--ssh-user root
```
**Expected**: SSH bootstrap succeeds, orca user created, drift-events dir created, node registered as `ready` with `kind=linux`.
### Step 4: Verify nodes
```sh
orca node list
orca node list --json
```
**Expected**: 3 nodes listed (localhost + pve01 + worker01), all `ready`.
### Step 5: Set capacity on remote nodes
```sh
orca node capacity set --node pve01 --cpu 4 --memory 8192 --disk 100000
orca node capacity set --node worker01 --cpu 2 --memory 4096 --disk 50000
orca node capacity list
```
**Expected**: capacity shown for both remote nodes.
### Step 6: Create a namespace
```sh
orca ns create prod
orca ns list
```
**Expected**: `prod` namespace listed.
### Step 7: Deploy the full stack
Deploy each service from `examples/full-stack/`:
```sh
orca job run examples/full-stack/web-app.md --target pve01
orca job run examples/full-stack/api.md --target pve01
orca job run examples/full-stack/worker.md --target worker01
orca job run examples/full-stack/postgres.md --target pve01
orca job run examples/full-stack/log-shipper.md --target worker01
```
**Expected**: each job is scheduled on the target, systemd unit deployed via SSH-push, job status `running` or `complete`.
### Step 8: Verify deployment
```sh
orca job list
orca job list --json
```
**Expected**: all 5 jobs listed, with correct target nodes.
On each remote node:
```sh
ssh root@pve01 systemctl status 'orca-alloc-*'
ssh root@worker01 systemctl status 'orca-alloc-*'
```
### Step 9: Verify Traefik routes
```sh
ssh root@pve01 ls /etc/traefik/dynamic/
ssh root@worker01 ls /etc/traefik/dynamic/
```
**Expected**: `traefik-dynamic-*.yaml` files present on nodes where jobs were deployed.
### Step 10: Migrate between hosts
Migrate `web-app` from pve01 to worker01:
```sh
orca job migrate web-app --to worker01
```
**Expected**: job drained on pve01, rescheduled on worker01, new systemd unit deployed.
Verify:
```sh
orca job list
ssh root@worker01 systemctl status 'orca-alloc-*web-app*'
ssh root@pve01 systemctl status 'orca-alloc-*web-app*' # should be stopped
```
### Step 11: Aggregate logs
```sh
orca logs --all-nodes --job web-app --since 5m
```
**Expected**: log entries from multiple nodes.
### Step 12: ACL enforcement
```sh
orca acl grant operator-1 --namespace prod --permissions read,write
orca acl check operator-1 --namespace prod --permission read
orca acl check operator-1 --namespace prod --permission admin
```
**Expected**: read+write allowed, admin denied (not granted).
### Step 12b: Initialize the OIDC provider (for seal)
```sh
orca auth init-idp --rp-id orca.local
```
**Expected**: Dex config + systemd unit + Traefik route rendered. (Dex binary must be installed separately.)
### Step 13: Seal/unseal
```sh
orca cluster seal
orca cluster unseal
orca secrets set prod TEST_KEY=test-value
orca secrets get prod TEST_KEY
```
**Expected**: seal succeeds, unseal succeeds, secrets readable post-unseal.
### Step 14: Audit chain
```sh
orca doctor audit
```
**Expected**: chain head reported, no tamper detected.
### Step 15: Doctor modes
```sh
orca doctor modes
```
**Expected**: all file modes correct, exit 0.
### Step 16: OIDC health
```sh
orca doctor oidc
```
**Expected**: Dex unit active, issuer reachable (or WARN if Dex not installed).
### Step 17: Backup and restore
```sh
orca backup --out /tmp/uat-backup.tar.gz
orca restore --in /tmp/uat-backup.tar.gz --dry-run
```
**Expected**: backup succeeds, restore dry-run succeeds.
### Step 18: Drift detection
```sh
orca drift show
```
**Expected**: no error (empty drift is fine).
### Step 19: Transaction idempotency
```sh
orca txn apply <some-txn-dir>
orca txn apply <some-txn-dir> # re-run
```
**Expected**: second apply is idempotent (exit 5 or "already applied").
### Step 20: Metrics
```sh
orca metrics --addr :9100 &
sleep 3
curl -s http://localhost:9100/metrics | grep orca_
```
**Expected**: expanded metric set present (`orca_jobs_running`, `orca_audit_chain_head`, etc.).
### Step 21: Compat check
```sh
orca cluster compat-check
```
**Expected**: exit 0, all nodes compatible.
### Step 22: Run the signoff script
```sh
scripts/uat-signoff.sh
```
**Expected**: `UAT SIGNOFF: N/35 assertions passed`, exit 0 iff N==35.
## Claim Matrix
| # | Claim | UAT Step | Signoff Assertion |
|---|-------|----------|-------------------|
| 1 | Cluster initializes from scratch | Step 1 | `assert_orca_version` |
| 2 | Proxmox host onboards via SSH | Step 2 | `assert_proxmox_onboarded` |
| 3 | Ubuntu worker onboards via `--type linux` | Step 3 | `assert_linux_worker_onboarded` |
| 4 | Node list shows all nodes | Step 4 | `assert_cluster_initialized` |
| 5 | Capacity is set on remote nodes | Step 5 | `assert_capacity_set` |
| 6 | Namespace created | Step 6 | `assert_namespace_created` |
| 7 | Full stack deploys to remote nodes | Step 7 | `assert_full_stack_running` |
| 8 | Scheduler deploys to remote (not local) | Step 7 | `assert_job_deploys_to_remote` |
| 9 | Traefik routes present | Step 9 | `assert_traefik_routes` |
| 10 | Job migrates between hosts | Step 10 | `assert_migrate_worked` |
| 11 | Logs aggregate from multiple nodes | Step 11 | `assert_logs_aggregate` |
| 12 | ACL grant/check works | Step 12 | `assert_acl_enforced` |
| 13 | ACL deny-by-default | Step 12 | `assert_acl_deny_default` |
| 14 | acl.json mode 0600 | Step 12 | `assert_acl_file_mode` |
| 15 | Seal/unseal round-trip | Step 13 | `assert_seal_unseal_roundtrip` |
| 16 | Audit chain intact | Step 14 | `assert_audit_chain_intact` |
| 17 | Doctor modes passes | Step 15 | `assert_doctor_modes` |
| 18 | OIDC health check | Step 16 | `assert_oidc_health` |
| 19 | Backup works | Step 17 | `assert_backup_restore_dryrun` |
| 20 | Drift visible | Step 18 | `assert_drift_visible` |
| 21 | Txn idempotent | Step 19 | `assert_txn_idempotent` |
| 22 | Metrics expanded | Step 20 | `assert_metrics_expanded` |
| 23 | Compat check passes | Step 21 | `assert_compat_check_passes` |
| 24 | No `--password` in docs/examples | — | `assert_no_password_in_docs` |
| 25 | Go toolchain current | — | `assert_go_toolchain_current` |
| 26 | cli.md matches `orca --help` | — | `assert_cli_md_complete` |
| 27 | pprof not on all interfaces | — | `assert_no_pprof_on_all_interfaces` |
| 28 | WebAuthn registration requires auth | — | `assert_webauthn_reg_requires_auth` |
| 29 | Audit chain survives concurrency | — | `assert_audit_chain_concurrent` |
| 30 | Concurrent secrets no data loss | — | `assert_concurrent_secrets_no_loss` |
| 31 | Cache invalidated after write | — | `assert_cache_invalidated_after_write` |
| 32 | SQLite no lock under concurrency | — | `assert_sqlite_no_lock` |
| 33 | No injection in logs --job | — | `assert_no_injection_in_logs` |
| 34 | `--type linux` exists as subcommand | Step 3 | `assert_type_linux_available` |
| 35 | `orca status` deprecated | — | `assert_status_deprecated` |
## Signoff procedure
1. Run all steps above on the 3-host cluster
2. Run `scripts/uat-signoff.sh` on the lead
3. Paste the output back to the CI agent
4. The CI agent verifies `35/35 PASS` and cuts `v1.0.0`
## Troubleshooting
### ORCA_HOME not set
All orca commands use `$ORCA_HOME` (default `~/.orca`). If commands fail
with "no such file or directory", verify:
```sh
echo $ORCA_HOME
ls $ORCA_HOME/orca.db $ORCA_HOME/orca_ssh_key $ORCA_HOME/known_hosts $ORCA_HOME/cluster/master.key
```
### known_hosts missing
If SSH operations fail with "open .../known_hosts: no such file", the
known_hosts file was not created during `orca init`. Fix:
```sh
touch $ORCA_HOME/known_hosts
chmod 600 $ORCA_HOME/known_hosts
```
### Traefik not running
If Traefik routes are not deployed, verify Traefik is running:
```sh
systemctl status orca-traefik
ls /etc/traefik/dynamic/
```
If not installed, `orca init` should have installed it. Re-run `orca init`
or install manually from https://github.com/traefik/traefik/releases.
### SSH connection refused
If the orca SSH key is not pre-staged on the remote host:
```sh
ssh-copy-id -i ~/.orca/orca_ssh_key.pub root@<host>
```
### Job deployed but not visible in `job list`
The remote dispatch path now inserts a DB record (v0.12.16). If you
still don't see it, check:
```sh
orca job list --json
```
Look for the `"node"` field — it shows which node the job deployed to.
### Proxmox: process runtime rejected
Proxmox nodes require `one_of: pve-ct` or `one_of: pve-vm` in the
jobspec. `one_of: process` (systemd) is for Linux/Ubuntu workers only.