0424f8ce02
orca init now interactively prompts for remote host addresses and runs ssh-copy-id automatically (password prompt passes through to the operator). This makes orca init the single entry point — no manual pre-staging of SSH keys required. - Interactive: enter host addresses (one per line, empty line to finish) - ssh-copy-id deploys the orca public key to each host - Skipped in --json mode (non-interactive) - Idempotent: re-running init can stage additional hosts Also fixed: install.sh defaults to /usr/local/bin (on PATH for all users). Non-root without sudo falls back to ~/.local/bin + auto-adds to .bashrc.
361 lines
10 KiB
Markdown
361 lines
10 KiB
Markdown
# Orca User Acceptance Testing (UAT) Plan
|
|
|
|
**Version**: v0.13 (production hardening round 2)
|
|
**Gate**: v1.0.0 production-ready tag is deferred until this UAT passes
|
|
**Signoff**: run `scripts/uat-signoff.sh` on the lead node and paste the output back
|
|
|
|
## Prerequisites
|
|
|
|
### Hardware
|
|
|
|
| Role | OS | Requirements |
|
|
|------|-----|-------------|
|
|
| **lead** | Ubuntu 22.04 LTS | Operator laptop or VM; SSH key; `orca` binary (built from v0.13 tag) |
|
|
| **pve01** | Proxmox VE 8/9 | Bare-metal or nested; SSH root access; orca SSH key pre-staged |
|
|
| **worker01** | Ubuntu 22.04 LTS | VM or bare-metal; SSH root access; orca SSH key pre-staged |
|
|
|
|
### Alternative topology (3x Ubuntu, no Proxmox)
|
|
|
|
If a Proxmox host is unavailable, run the UAT with 3x Ubuntu hosts.
|
|
Use `--type linux` for all remote nodes. Proxmox-specific claims
|
|
(`doctor proxmox`, PVE role, sudoers) are **skipped** in this path.
|
|
The signoff script reports exercised vs. skipped claims.
|
|
|
|
## Step-by-step UAT
|
|
|
|
### Step 1: Install orca + initialize the cluster
|
|
|
|
Install orca (1-liner):
|
|
```sh
|
|
curl -fsSL https://git.cloudinit.dev/coreci/orca/raw/branch/main/scripts/install.sh | bash
|
|
```
|
|
|
|
Initialize the cluster:
|
|
```sh
|
|
export ORCA_HOME=~/orca-uat
|
|
orca init
|
|
```
|
|
|
|
**Expected**: orca init creates:
|
|
- CA cert + server cert
|
|
- SSH keypair (orca_ssh_key + orca_ssh_key.pub)
|
|
- known_hosts file (empty, for TOFU capture)
|
|
- Master key (for secrets encryption)
|
|
- Traefik data-plane ingress (binary + systemd unit + config)
|
|
- Localhost node registered
|
|
|
|
**Pre-staging remote nodes**: `orca init` interactively prompts for remote
|
|
host addresses and runs `ssh-copy-id` automatically (password prompt passes
|
|
through). Enter each host (pve01, worker01) when prompted, or press Enter to
|
|
skip. The orca public key is deployed to each host; TOFU host-key capture is
|
|
automatic on the first `orca node join` — no manual fingerprint pinning needed.
|
|
|
|
### Step 2: Onboard the Proxmox host
|
|
|
|
```sh
|
|
orca node join --type proxmox \
|
|
--host pve01 \
|
|
--ssh-user root
|
|
```
|
|
|
|
**Expected**: SSH bootstrap succeeds, orca user created, PVE role assigned, node registered as `ready` with `kind=proxmox`.
|
|
|
|
### Step 3: Onboard the Ubuntu worker
|
|
|
|
```sh
|
|
orca node join --type linux \
|
|
--host worker01 \
|
|
--ssh-user root
|
|
```
|
|
|
|
**Expected**: SSH bootstrap succeeds, orca user created, drift-events dir created, node registered as `ready` with `kind=linux`.
|
|
|
|
### Step 4: Verify nodes
|
|
|
|
```sh
|
|
orca node list
|
|
orca node list --json
|
|
```
|
|
|
|
**Expected**: 3 nodes listed (localhost + pve01 + worker01), all `ready`.
|
|
|
|
### Step 5: Set capacity on remote nodes
|
|
|
|
```sh
|
|
orca node capacity set --node pve01 --cpu 4 --memory 8192 --disk 100000
|
|
orca node capacity set --node worker01 --cpu 2 --memory 4096 --disk 50000
|
|
orca node capacity list
|
|
```
|
|
|
|
**Expected**: capacity shown for both remote nodes.
|
|
|
|
### Step 6: Create a namespace
|
|
|
|
```sh
|
|
orca ns create prod
|
|
orca ns list
|
|
```
|
|
|
|
**Expected**: `prod` namespace listed.
|
|
|
|
### Step 7: Deploy the full stack
|
|
|
|
Deploy each service from `examples/full-stack/`:
|
|
|
|
```sh
|
|
orca job run examples/full-stack/web-app.md --target pve01
|
|
orca job run examples/full-stack/api.md --target pve01
|
|
orca job run examples/full-stack/worker.md --target worker01
|
|
orca job run examples/full-stack/postgres.md --target pve01
|
|
orca job run examples/full-stack/log-shipper.md --target worker01
|
|
```
|
|
|
|
**Expected**: each job is scheduled on the target, systemd unit deployed via SSH-push, job status `running` or `complete`.
|
|
|
|
### Step 8: Verify deployment
|
|
|
|
```sh
|
|
orca job list
|
|
orca job list --json
|
|
```
|
|
|
|
**Expected**: all 5 jobs listed, with correct target nodes.
|
|
|
|
On each remote node:
|
|
```sh
|
|
ssh root@pve01 systemctl status 'orca-alloc-*'
|
|
ssh root@worker01 systemctl status 'orca-alloc-*'
|
|
```
|
|
|
|
### Step 9: Verify Traefik routes
|
|
|
|
```sh
|
|
ssh root@pve01 ls /etc/traefik/dynamic/
|
|
ssh root@worker01 ls /etc/traefik/dynamic/
|
|
```
|
|
|
|
**Expected**: `traefik-dynamic-*.yaml` files present on nodes where jobs were deployed.
|
|
|
|
### Step 10: Migrate between hosts
|
|
|
|
Migrate `web-app` from pve01 to worker01:
|
|
|
|
```sh
|
|
orca job migrate web-app --to worker01
|
|
```
|
|
|
|
**Expected**: job drained on pve01, rescheduled on worker01, new systemd unit deployed.
|
|
|
|
Verify:
|
|
```sh
|
|
orca job list
|
|
ssh root@worker01 systemctl status 'orca-alloc-*web-app*'
|
|
ssh root@pve01 systemctl status 'orca-alloc-*web-app*' # should be stopped
|
|
```
|
|
|
|
### Step 11: Aggregate logs
|
|
|
|
```sh
|
|
orca logs --all-nodes --job web-app --since 5m
|
|
```
|
|
|
|
**Expected**: log entries from multiple nodes.
|
|
|
|
### Step 12: ACL enforcement
|
|
|
|
```sh
|
|
orca acl grant operator-1 --namespace prod --permissions read,write
|
|
orca acl check operator-1 --namespace prod --permission read
|
|
orca acl check operator-1 --namespace prod --permission admin
|
|
```
|
|
|
|
**Expected**: read+write allowed, admin denied (not granted).
|
|
|
|
### Step 12b: Initialize the OIDC provider (for seal)
|
|
|
|
```sh
|
|
orca auth init-idp --rp-id orca.local
|
|
```
|
|
|
|
**Expected**: Dex config + systemd unit + Traefik route rendered. (Dex binary must be installed separately.)
|
|
|
|
### Step 13: Seal/unseal
|
|
|
|
```sh
|
|
orca cluster seal
|
|
orca cluster unseal
|
|
orca secrets set prod TEST_KEY=test-value
|
|
orca secrets get prod TEST_KEY
|
|
```
|
|
|
|
**Expected**: seal succeeds, unseal succeeds, secrets readable post-unseal.
|
|
|
|
### Step 14: Audit chain
|
|
|
|
```sh
|
|
orca doctor audit
|
|
```
|
|
|
|
**Expected**: chain head reported, no tamper detected.
|
|
|
|
### Step 15: Doctor modes
|
|
|
|
```sh
|
|
orca doctor modes
|
|
```
|
|
|
|
**Expected**: all file modes correct, exit 0.
|
|
|
|
### Step 16: OIDC health
|
|
|
|
```sh
|
|
orca doctor oidc
|
|
```
|
|
|
|
**Expected**: Dex unit active, issuer reachable (or WARN if Dex not installed).
|
|
|
|
### Step 17: Backup and restore
|
|
|
|
```sh
|
|
orca backup --out /tmp/uat-backup.tar.gz
|
|
orca restore --in /tmp/uat-backup.tar.gz --dry-run
|
|
```
|
|
|
|
**Expected**: backup succeeds, restore dry-run succeeds.
|
|
|
|
### Step 18: Drift detection
|
|
|
|
```sh
|
|
orca drift show
|
|
```
|
|
|
|
**Expected**: no error (empty drift is fine).
|
|
|
|
### Step 19: Transaction idempotency
|
|
|
|
```sh
|
|
orca txn apply <some-txn-dir>
|
|
orca txn apply <some-txn-dir> # re-run
|
|
```
|
|
|
|
**Expected**: second apply is idempotent (exit 5 or "already applied").
|
|
|
|
### Step 20: Metrics
|
|
|
|
```sh
|
|
orca metrics --addr :9100 &
|
|
sleep 3
|
|
curl -s http://localhost:9100/metrics | grep orca_
|
|
```
|
|
|
|
**Expected**: expanded metric set present (`orca_jobs_running`, `orca_audit_chain_head`, etc.).
|
|
|
|
### Step 21: Compat check
|
|
|
|
```sh
|
|
orca cluster compat-check
|
|
```
|
|
|
|
**Expected**: exit 0, all nodes compatible.
|
|
|
|
### Step 22: Run the signoff script
|
|
|
|
```sh
|
|
scripts/uat-signoff.sh
|
|
```
|
|
|
|
**Expected**: `UAT SIGNOFF: N/35 assertions passed`, exit 0 iff N==35.
|
|
|
|
## Claim Matrix
|
|
|
|
| # | Claim | UAT Step | Signoff Assertion |
|
|
|---|-------|----------|-------------------|
|
|
| 1 | Cluster initializes from scratch | Step 1 | `assert_orca_version` |
|
|
| 2 | Proxmox host onboards via SSH | Step 2 | `assert_proxmox_onboarded` |
|
|
| 3 | Ubuntu worker onboards via `--type linux` | Step 3 | `assert_linux_worker_onboarded` |
|
|
| 4 | Node list shows all nodes | Step 4 | `assert_cluster_initialized` |
|
|
| 5 | Capacity is set on remote nodes | Step 5 | `assert_capacity_set` |
|
|
| 6 | Namespace created | Step 6 | `assert_namespace_created` |
|
|
| 7 | Full stack deploys to remote nodes | Step 7 | `assert_full_stack_running` |
|
|
| 8 | Scheduler deploys to remote (not local) | Step 7 | `assert_job_deploys_to_remote` |
|
|
| 9 | Traefik routes present | Step 9 | `assert_traefik_routes` |
|
|
| 10 | Job migrates between hosts | Step 10 | `assert_migrate_worked` |
|
|
| 11 | Logs aggregate from multiple nodes | Step 11 | `assert_logs_aggregate` |
|
|
| 12 | ACL grant/check works | Step 12 | `assert_acl_enforced` |
|
|
| 13 | ACL deny-by-default | Step 12 | `assert_acl_deny_default` |
|
|
| 14 | acl.json mode 0600 | Step 12 | `assert_acl_file_mode` |
|
|
| 15 | Seal/unseal round-trip | Step 13 | `assert_seal_unseal_roundtrip` |
|
|
| 16 | Audit chain intact | Step 14 | `assert_audit_chain_intact` |
|
|
| 17 | Doctor modes passes | Step 15 | `assert_doctor_modes` |
|
|
| 18 | OIDC health check | Step 16 | `assert_oidc_health` |
|
|
| 19 | Backup works | Step 17 | `assert_backup_restore_dryrun` |
|
|
| 20 | Drift visible | Step 18 | `assert_drift_visible` |
|
|
| 21 | Txn idempotent | Step 19 | `assert_txn_idempotent` |
|
|
| 22 | Metrics expanded | Step 20 | `assert_metrics_expanded` |
|
|
| 23 | Compat check passes | Step 21 | `assert_compat_check_passes` |
|
|
| 24 | No `--password` in docs/examples | — | `assert_no_password_in_docs` |
|
|
| 25 | Go toolchain current | — | `assert_go_toolchain_current` |
|
|
| 26 | cli.md matches `orca --help` | — | `assert_cli_md_complete` |
|
|
| 27 | pprof not on all interfaces | — | `assert_no_pprof_on_all_interfaces` |
|
|
| 28 | WebAuthn registration requires auth | — | `assert_webauthn_reg_requires_auth` |
|
|
| 29 | Audit chain survives concurrency | — | `assert_audit_chain_concurrent` |
|
|
| 30 | Concurrent secrets no data loss | — | `assert_concurrent_secrets_no_loss` |
|
|
| 31 | Cache invalidated after write | — | `assert_cache_invalidated_after_write` |
|
|
| 32 | SQLite no lock under concurrency | — | `assert_sqlite_no_lock` |
|
|
| 33 | No injection in logs --job | — | `assert_no_injection_in_logs` |
|
|
| 34 | `--type linux` exists as subcommand | Step 3 | `assert_type_linux_available` |
|
|
| 35 | `orca status` deprecated | — | `assert_status_deprecated` |
|
|
|
|
## Signoff procedure
|
|
|
|
1. Run all steps above on the 3-host cluster
|
|
2. Run `scripts/uat-signoff.sh` on the lead
|
|
3. Paste the output back to the CI agent
|
|
4. The CI agent verifies `35/35 PASS` and cuts `v1.0.0`
|
|
|
|
|
|
## Troubleshooting
|
|
|
|
### ORCA_HOME not set
|
|
All orca commands use `$ORCA_HOME` (default `~/.orca`). If commands fail
|
|
with "no such file or directory", verify:
|
|
```sh
|
|
echo $ORCA_HOME
|
|
ls $ORCA_HOME/orca.db $ORCA_HOME/orca_ssh_key $ORCA_HOME/known_hosts $ORCA_HOME/cluster/master.key
|
|
```
|
|
|
|
### known_hosts missing
|
|
If SSH operations fail with "open .../known_hosts: no such file", the
|
|
known_hosts file was not created during `orca init`. Fix:
|
|
```sh
|
|
touch $ORCA_HOME/known_hosts
|
|
chmod 600 $ORCA_HOME/known_hosts
|
|
```
|
|
|
|
### Traefik not running
|
|
If Traefik routes are not deployed, verify Traefik is running:
|
|
```sh
|
|
systemctl status orca-traefik
|
|
ls /etc/traefik/dynamic/
|
|
```
|
|
If not installed, `orca init` should have installed it. Re-run `orca init`
|
|
or install manually from https://github.com/traefik/traefik/releases.
|
|
|
|
### SSH connection refused
|
|
If the orca SSH key is not pre-staged on the remote host:
|
|
```sh
|
|
ssh-copy-id -i ~/.orca/orca_ssh_key.pub root@<host>
|
|
```
|
|
|
|
### Job deployed but not visible in `job list`
|
|
The remote dispatch path now inserts a DB record (v0.12.16). If you
|
|
still don't see it, check:
|
|
```sh
|
|
orca job list --json
|
|
```
|
|
Look for the `"node"` field — it shows which node the job deployed to.
|
|
|
|
### Proxmox: process runtime rejected
|
|
Proxmox nodes require `one_of: pve-ct` or `one_of: pve-vm` in the
|
|
jobspec. `one_of: process` (systemd) is for Linux/Ubuntu workers only.
|