Compare commits

...

9 Commits

Author SHA1 Message Date
Jon Chery 3a3ea74d76 fix(P08): transport + SSH safety — typed errors, IPv6, timeouts, signal (REQ-157)
- transport.IsTransient: typed sentinels (ErrTransient/ErrPermanent) +
  standard net.Error/io errors.Is; substring matching removed
- sshpush.isTransient: same typed-error classification
- rotateSSHKeys: 2-phase atomic swap (stage peers -> swap local ->
  verify -> cleanup old); no more partial-result window
- known_hosts: dial() reads stored field (was reading v0.8 path directly)
- IPv6: net.JoinHostPort in proxmox SSH dial + drain splitHostPort
- SSH timeouts: context.WithTimeout on peer-setup, drift, txn rollback,
  job restart (default 2m)
- verifyCutover: orca CA pool TLS config (was default http.Client)
- OIDC callback: ReadHeaderTimeout 5s (slowloris defense)
- root Execute: signal.NotifyContext for SIGINT/SIGTERM (clean exit
  for non-watch commands)

Tests: typed-error classification table, IPv6 JoinHostPort, signal
handler context cancellation.

---ci---
project: orca
phase: 8
milestone: v0.13
status: complete
requirements:
  covered: [157]
---/ci---
2026-08-10 13:11:07 +00:00
Jon Chery 0358efe95b fix(P07): concurrency safety — SQLite, flock, cache, atomic writes (REQ-156)
- SQLite busy_timeout(5000) + SetMaxOpenConns(1) on all 4 DSNs
- secrets file flock (concurrent set on same ns no longer loses data)
- upgrade lock file (refuse concurrent orca upgrade)
- backup lock file (refuse concurrent backup)
- cache invalidation by writes (read-after-write consistency)
- Executor.Run mutex scope fix (hold only for DB inserts)
- ns create/inherit/set-constraint atomic writeNSMdAtomic
- writeCurrentLead + rotateSSHKeys atomic
- consolidate 3 writeAtomic impls onto security.WriteAtomic
- WebAuthn session stores guarded with sync.Mutex

Tests: concurrent secrets set, upgrade lock rejection, cache
read-after-write, WebAuthn session thread-safety (pass under -race).

---ci---
project: orca
phase: 7
milestone: v0.13
status: complete
requirements:
  covered: [156]
---/ci---
2026-08-10 12:27:05 +00:00
Jon Chery 978334a4bc feat(P06): auth init-idp real + auth register + doctor oidc (REQ-155)
Implements the v0.12 R-021 load-bearing change's working IdP path:
- orca auth init-idp: renders Dex config + systemd unit + Traefik route
  (atomic deploy, RP ID from --rp-id, C-38)
- orca auth register: opens browser to WebAuthn registration page
- loadOIDCConfig: config-file loading (oidc block + cluster_domain),
  falls back to flags + env vars
- orca doctor oidc: health check (systemctl is-active + .well-known)
- config.go: OIDCConfig block + ClusterDomain field
- markdown.go: oidc block parsing in config frontmatter

---ci---
project: orca
phase: 6
milestone: v0.13
status: complete
requirements:
  covered: [155]
---/ci---
2026-08-10 11:55:01 +00:00
Jon Chery 9e832387c6 feat(P05): seal/audit CLI + chain race fix + key zeroing (REQ-154)
New CLI commands:
- orca cluster seal: OIDC/CA-derived seal + Shamir 3-of-5 shards
- orca cluster unseal: OIDC/CA unseal + --recovery Shamir path
- orca doctor audit: VerifyChain + chain head report
- orca doctor modes: EnforceFileModes across ORCA_HOME

Fixes:
- audit hash-chain race: Append uses BEGIN IMMEDIATE transaction
  (concurrent appends no longer corrupt tamper-evidence)
- secrets rotate-master: re-seals to OIDC on sealed clusters
  (was writing raw key, docstring claimed re-seal)
- key zeroing: ZeroKey helper + defer after master/namespace key use
  (defense-in-depth against pprof heap extraction)
- store.Open: busy_timeout(5000) pragma (concurrent writers wait)

Tests: 18 new test functions (seal round-trip, Shamir recovery, doctor
audit tamper detection, doctor modes 0644 rejection, concurrent append
chain integrity, rotate-master re-seal, key zeroing).

---ci---
project: orca
phase: 5
milestone: v0.13
status: complete
requirements:
  covered: [154]
---/ci---
2026-08-07 21:06:39 +00:00
Jon Chery 5232fcb808 fix(P04): wire ACL enforcement + WebAuthn reg auth + audit actor (REQ-153)
R-023: Zero-trust enforcement operationally wired.

ACL enforcement (C-45 staged rollout):
- acl.Check wired into all 5 daemon handlers (dispatch/jobs/nodes/tasks)
- health endpoints exempt (liveness probes not gated)
- ACL log-only mode default (config acl.enforce=false); enforce after
  bootstrap ACL verified
- sshpush auth: ORCA_OIDC_TOKEN validated against JWKS before apply
- txn apply: Authorize hook validates OIDC token before running pull
- acl.json mode 0600 (was 0644)
- flock on acl.json for concurrent grant/revoke
- bootstrap ACL: init grants cluster-admin to orca-admins group + SVID

Audit actor identity:
- currentActor reads OIDC sub from credentials.json (was hardcoded "cli")
- threaded through all audit.Record calls via context

WebAuthn registration auth:
- BeginRegistration/FinishRegistration require authenticated session
- fail-closed 401 when no authFunc configured

New files: internal/daemon/acl.go, internal/cli/authactor.go,
internal/engine/actor.go, internal/identity/authtoken.go,
internal/sshpush/auth.go, internal/txn/auth_test.go

---ci---
project: orca
phase: 4
milestone: v0.13
status: complete
requirements:
  covered: [153]
---/ci---
2026-08-07 20:33:39 +00:00
Jon Chery cf3d98eb2b feat(P03): wire scheduler into job run + fix jobspec parser (REQ-151, REQ-152)
R-022: orca job run now deploys to remote nodes via scheduler -> emitter
-> SSH-push. Local exec fallback only when no remote nodes registered.

jobspec parser (REQ-152):
- schedule: and timeout: now parsed (were silently dropped)
- DaemonSet Count no longer defaults to 1 (was breaking DaemonSet)
- restart: policy translated to systemd Restart=/StartLimitBurst
- job lint: advisory warnings for cron/health/update/affinity (honest)

scheduler wiring (REQ-151, C-44):
- new internal/cli/job_dispatch.go: dispatchDecision + deployRemote
- scheduler.Schedule evaluates constraints/capacity/affinity
- --target overrides scheduler (manual pinning)
- local fallback only when len(ready non-localhost nodes)==0
- C-44: SSH-push failure returns error (no silent local fallback)
- systemd-analyze verify on rendered unit before deploy

Tests: 22 new test functions covering scheduler, parser, C-44, local
fallback, target override, systemd-analyze skip, restart directives.

---ci---
project: orca
phase: 3
milestone: v0.13
status: complete
requirements:
  covered: [151, 152]
---/ci---
2026-08-07 19:59:31 +00:00
Jon Chery 4b70e31cf4 fix(P02): input validation + injection hardening — 11 vectors (REQ-150)
Critical fixes:
- logs --job: validate ^[A-Za-z0-9_-]+$ + shellQuote (was %q backtick RCE)
- pprof: isLoopback treats empty host as bind-all (was :6060 bypass)
- backup restore: filepath.Rel containment check (was tar-slip via a/../..)
- WebAuthn reg auth deferred to P04 (requires session infra)

High fixes:
- txn rollback/show/apply: validate ^T-[0-9a-f]{16}$ + shellQuote
- nft diff --against: validate txn ID before filepath.Join
- drain stopAlloc: validate allocID ^[A-Za-z0-9_-]+$
- cluster_compat: shellQuote peer dir name
- podman image: shellQuote (was %q backtick injection)
- nft TrustedProbes: net.ParseIP/CIDR validation + split v4/v6 sets
- sudoers: validate --proxmox-user/--proxmox-role ^[a-zA-Z_][a-zA-Z0-9_-]{0,31}$
  fixed path /etc/sudoers.d/orca; shellQuote pveum/useradd; validateSudoers
  checks actual file
- nft country block: validate ^[A-Z]{2}$ (was len==2 only)

New file: internal/cli/validate.go (shared validators + shellQuote)
All 38 Go test packages pass. go vet + gofmt clean.

---ci---
project: orca
phase: 2
milestone: v0.13
status: complete
requirements:
  covered: [150]
---/ci---
2026-08-07 19:28:01 +00:00
Jon Chery b0158c96e9 fix(P01): bump go toolchain to 1.25.12 + fix pre-existing test bugs (REQ-149)
Toolchain:
- go.mod: go 1.25.0 -> 1.25.12 (closes 24 stdlib vulns: archive/tar,
  crypto/tls, crypto/x509, net/http, net/url, encoding/pem, os)
- go mod tidy clean; make build + test + lint pass

Pre-existing test bugs fixed (surfaced by toolchain bump):
- acl_test.go: KindToken always denies (R-021); tests updated to KindOidc
- acl.go: parseIdentity defaults to KindOidc (was KindToken, making
  acl grant/check CLI path non-functional for non-spiffe identities)
- init_test.go: migration version updated to 0008 (was 0007, stale since v0.12)
- doctor.go: CertCA now checks CA cert exists (was only checking file modes,
  passing when no CA present)
- scenarios_test.go: ACL integration test uses KindOidc + acl.json 0600

---ci---
project: orca
phase: 1
milestone: v0.13
status: complete
requirements:
  covered: [149]
---/ci---
2026-08-07 19:07:17 +00:00
Jon Chery 7479cd1534 docs(checkpoint): P0 shipped — v0.12.0 tagged 2026-08-07 18:50:08 +00:00
94 changed files with 6771 additions and 335 deletions
+9 -14
View File
@@ -1,26 +1,21 @@
{
"phase": 0,
"stage": "grill",
"phase": 1,
"stage": "complete",
"milestone": "v0.13",
"milestone_slug": "production-hardening-2",
"phase_role": "pre_execution",
"phase_role": "execution",
"attempts": 0,
"updated_at": "2026-08-07T19:15:00Z",
"updated_at": "2026-08-07T19:05:00Z",
"milestone_complete": false,
"previous_milestone": "v0.12",
"phase_count": 14,
"phases_shipped": [],
"tags_shipped": [],
"phases_shipped": ["P0", "P1"],
"tags_shipped": ["v0.12.0", "v0.12.1"],
"requirements": {
"covered": [],
"covered": [149],
"partial": []
},
"binding_conditions": [
"C-39", "C-40", "C-41", "C-42", "C-43",
"C-44", "C-45", "C-46", "C-47", "C-48", "C-49"
],
"binding_conditions": ["C-39","C-40","C-41","C-42","C-43","C-44","C-45","C-46","C-47","C-48","C-49"],
"load_bearing_rule": "R-022",
"next_milestone": "v1.0",
"grill_verdict": "CONDITIONAL_PROCEED",
"grill_confidence": 0.82
"next_milestone": "v1.0"
}
+1 -1
View File
@@ -1,6 +1,6 @@
module git.cloudinit.dev/coreci/orca
go 1.25.0
go 1.25.12
require (
github.com/coreos/go-oidc/v3 v3.20.0
+9 -3
View File
@@ -299,10 +299,16 @@ func Restore(opts RestoreOptions) error {
return fmt.Errorf("restore: read tar entry: %w", err)
}
name := filepath.FromSlash(hdr.Name)
if strings.HasPrefix(name, "/") || strings.HasPrefix(name, "..") {
return fmt.Errorf("restore: unsafe path %q", hdr.Name)
}
// F3: tar-slip containment check. The prior prefix check
// (HasPrefix "/" || "..") missed patterns like "a/../../etc".
// Resolve the destination and verify it stays within target
// via filepath.Rel; reject if the relative path escapes (starts
// with ".." or is absolute).
dest := filepath.Join(target, name)
rel, err := filepath.Rel(target, dest)
if err != nil || strings.HasPrefix(rel, "..") || filepath.IsAbs(rel) {
return fmt.Errorf("restore: unsafe path %q escapes target (F3: tar-slip)", hdr.Name)
}
switch hdr.Typeflag {
case tar.TypeDir:
if err := os.MkdirAll(dest, os.FileMode(hdr.Mode)); err != nil {
+65
View File
@@ -391,6 +391,71 @@ func TestRestoreRejectsTraversalSymlink(t *testing.T) {
}
}
// createCraftedTarballWithFile creates a tar.gz containing a single
// regular file entry with the given (possibly malicious) name. Used to
// test the tar-slip path-traversal guard (F3).
func createCraftedTarballWithFile(path, name, body string) error {
f, err := os.Create(path)
if err != nil {
return err
}
defer f.Close()
gz := gzip.NewWriter(f)
defer gz.Close()
tw := tar.NewWriter(gz)
defer tw.Close()
hdr := &tar.Header{
Name: name,
Typeflag: tar.TypeReg,
Mode: 0o644,
Size: int64(len(body)),
}
if err := tw.WriteHeader(hdr); err != nil {
return err
}
if _, err := tw.Write([]byte(body)); err != nil {
return err
}
return nil
}
// TestRestoreRejectsTarSlipRegularFile verifies a tarball with a regular
// file entry whose name contains an embedded ".." traversal (e.g.
// "a/../../etc/passwd") is rejected. The old prefix-only check missed
// this pattern; the F3 filepath.Rel containment check catches it.
func TestRestoreRejectsTarSlipRegularFile(t *testing.T) {
dir := t.TempDir()
tarPath := filepath.Join(dir, "slip.tar.gz")
sigPath := tarPath + ".sig"
if err := createCraftedTarballWithFile(tarPath, "a/../../etc/passwd", "pwned"); err != nil {
t.Fatalf("create tarball: %v", err)
}
key := make([]byte, 32)
for i := range key {
key[i] = byte(i + 9)
}
mac := hmac.New(sha256.New, key)
data, _ := os.ReadFile(tarPath)
mac.Write(data)
if err := os.WriteFile(sigPath, []byte(hex.EncodeToString(mac.Sum(nil))), 0o600); err != nil {
t.Fatalf("write sig: %v", err)
}
target := filepath.Join(dir, "restore")
os.MkdirAll(target, 0o755)
err := Restore(RestoreOptions{
InputPath: tarPath,
TargetDir: target,
MasterKey: key,
Force: true,
})
if err == nil {
t.Fatal("Restore should reject tar-slip regular file (F3)")
}
if !strings.Contains(err.Error(), "unsafe path") {
t.Errorf("error should mention unsafe path: %v", err)
}
}
// createCraftedTarball creates a tar.gz containing a single symlink
// entry with the given linkname. Used to test symlink validation.
func createCraftedTarball(path, name, linkname string) error {
+8 -1
View File
@@ -58,10 +58,17 @@ func Open(path string) (*Cache, error) {
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
return nil, fmt.Errorf("create cache db dir: %w", err)
}
db, err := sql.Open("sqlite", path+"?_pragma=journal_mode(WAL)")
// REQ-156 / P07 T1: busy_timeout(5000) so concurrent cache opens
// (e.g. two `orca node list` invocations racing on the same shell)
// wait up to 5s for the writer instead of failing immediately with
// SQLITE_BUSY. SetMaxOpenConns(1) serializes the connections so the
// busy_timeout is rarely needed but keeps the cache durable under
// contention.
db, err := sql.Open("sqlite", path+"?_pragma=journal_mode(WAL)&_pragma=busy_timeout(5000)")
if err != nil {
return nil, fmt.Errorf("open cache sqlite: %w", err)
}
db.SetMaxOpenConns(1)
if err := db.Ping(); err != nil {
_ = db.Close()
return nil, fmt.Errorf("ping cache sqlite: %w", err)
+44 -27
View File
@@ -23,6 +23,7 @@ import (
"git.cloudinit.dev/coreci/orca/internal/acl"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/security"
)
var (
@@ -59,7 +60,7 @@ func parseIdentity(raw string) (acl.Identity, error) {
if raw == "" {
return acl.Identity{}, fmt.Errorf("identity is empty")
}
return acl.Identity{Kind: acl.KindToken, ID: raw}, nil
return acl.Identity{Kind: acl.KindOidc, ID: raw}, nil
}
// parsePermissions parses a comma-separated list of "read","write",
@@ -149,38 +150,40 @@ func saveACL(a *acl.ACL) error {
if err != nil {
return fmt.Errorf("marshal acl state: %w", err)
}
if err := writeAtomicFile(path, data, 0o644); err != nil {
// P04 (T6): acl.json contains the access-control policy and
// must be 0600 (operator-only). Previously 0644 — world-readable
// leaked the SPIFFE IDs and OIDC subs of privileged identities.
if err := writeAtomicFile(path, data, 0o600); err != nil {
return fmt.Errorf("write acl state: %w", err)
}
return nil
}
// writeAtomicFile writes data to a temp file in dir(path) and renames
// it into place, matching the security.WriteAtomic pattern (P02 keeps
// a local copy to avoid importing internal/security into the CLI).
// lockACL acquires an exclusive advisory lock on the acl.json file
// (P04, T7). The lock file is paths.ACLPath() + ".lock". Returns a
// release function that MUST be deferred. Used by grant/revoke to
// prevent concurrent read-modify-write races (two operators running
// `orca acl grant` simultaneously would otherwise clobber each
// other's entries).
func lockACL() (func(), error) {
// Ensure the cluster dir exists before flock tries to create the
// lock file (security.Flock opens with O_CREATE but requires the
// parent dir to exist).
if err := os.MkdirAll(filepath.Dir(paths.ACLPath()), 0o755); err != nil {
return nil, fmt.Errorf("create cluster dir: %w", err)
}
return security.Flock(paths.ACLPath() + ".lock")
}
// writeAtomicFile writes data atomically (REQ-156, P07 T9).
// Previously a local copy of the temp+chmod+rename pattern (P02 kept a
// local copy to avoid importing internal/security); it lacked fsync,
// so a crash between write and rename could promote a partially-durable
// file. Now a thin wrapper around the canonical security.WriteAtomic
// (temp + chmod + fsync + rename) so all CLI atomic writes share one
// fsync-correct implementation.
func writeAtomicFile(path string, data []byte, mode os.FileMode) error {
dir := filepath.Dir(path)
tmp, err := os.CreateTemp(dir, ".acl-tmp-*")
if err != nil {
return fmt.Errorf("create temp: %w", err)
}
tmpName := tmp.Name()
defer func() { _ = os.Remove(tmpName) }()
if _, err := tmp.Write(data); err != nil {
_ = tmp.Close()
return fmt.Errorf("write temp: %w", err)
}
if err := tmp.Chmod(mode); err != nil {
_ = tmp.Close()
return fmt.Errorf("chmod temp: %w", err)
}
if err := tmp.Close(); err != nil {
return fmt.Errorf("close temp: %w", err)
}
if err := os.Rename(tmpName, path); err != nil {
return fmt.Errorf("rename temp: %w", err)
}
return nil
return security.WriteAtomic(path, mode, data)
}
var aclGrantCmd = &cobra.Command{
@@ -211,6 +214,14 @@ admin (default: read).`,
if err != nil {
return err
}
// P04 (T7): flock around the read-modify-write so two
// concurrent `orca acl grant` invocations don't clobber each
// other's entries.
release, err := lockACL()
if err != nil {
return fmt.Errorf("acquire acl lock: %w", err)
}
defer release()
a, err := loadACL()
if err != nil {
return err
@@ -253,6 +264,12 @@ token identity --namespace is required.`,
if ns == "" {
return fmt.Errorf("--namespace is required for token identities")
}
// P04 (T7): flock around the read-modify-write.
release, err := lockACL()
if err != nil {
return fmt.Errorf("acquire acl lock: %w", err)
}
defer release()
a, err := loadACL()
if err != nil {
return err
+76 -3
View File
@@ -8,6 +8,7 @@ import (
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/acl"
"git.cloudinit.dev/coreci/orca/internal/paths"
)
@@ -55,13 +56,13 @@ func TestParseIdentity_Spiffe(t *testing.T) {
}
}
func TestParseIdentity_Token(t *testing.T) {
func TestParseIdentity_Oidc(t *testing.T) {
id, err := parseIdentity("operator-1")
if err != nil {
t.Fatalf("parseIdentity: %v", err)
}
if id.Kind != "token" {
t.Errorf("kind = %q, want token", id.Kind)
if id.Kind != "oidc" {
t.Errorf("kind = %q, want oidc", id.Kind)
}
if id.ID != "operator-1" {
t.Errorf("id = %q, want operator-1", id.ID)
@@ -384,3 +385,75 @@ func TestACLAdminImpliesReadCheck(t *testing.T) {
t.Fatalf("check write (admin grant): %v", err)
}
}
// TestACLGrantWritesMode0600 (P04, T6) verifies that saveACL writes
// acl.json with mode 0600 (operator-only). Previously 0644 leaked
// SPIFFE IDs + OIDC subs to other local users.
func TestACLGrantWritesMode0600(t *testing.T) {
t.Setenv("ORCA_HOME", t.TempDir())
resetRootFlags(t)
resetACLFlags()
rootCmd.SetArgs([]string{"acl", "grant", "operator-1", "--namespace", "prod", "--permissions", "read"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("grant: %v", err)
}
info, err := os.Stat(paths.ACLPath())
if err != nil {
t.Fatalf("stat acl.json: %v", err)
}
if info.Mode().Perm()&0o077 != 0 {
t.Errorf("acl.json mode = %o, want 0600 (no group/other bits)", info.Mode().Perm())
}
}
// TestACLGrantCreatesLockFile (P04, T7) verifies that the flock
// mechanism creates an acl.json.lock file alongside acl.json. The
// lock prevents concurrent grant/revoke races.
func TestACLGrantCreatesLockFile(t *testing.T) {
t.Setenv("ORCA_HOME", t.TempDir())
resetRootFlags(t)
resetACLFlags()
rootCmd.SetArgs([]string{"acl", "grant", "operator-1", "--namespace", "prod", "--permissions", "read"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("grant: %v", err)
}
if _, err := os.Stat(paths.ACLPath() + ".lock"); err != nil {
t.Errorf("acl.json.lock not created: %v", err)
}
}
// TestACLBootstrapGrantsAdminGroup (P04, T8, C-40) verifies that
// bootstrapACL grants cluster-admin to the orca-admins OIDC group on
// the default namespace. This prevents operator lockout after
// `orca init`.
func TestACLBootstrapGrantsAdminGroup(t *testing.T) {
t.Setenv("ORCA_HOME", t.TempDir())
if err := os.MkdirAll(paths.ClusterDir(), 0o755); err != nil {
t.Fatalf("mkdir: %v", err)
}
// bootstrapACL reads the cert at certPath; a missing cert is
// non-fatal (the SVID grant is skipped, the group grant still
// applies). Pass a nonexistent path to exercise that path.
if err := bootstrapACL(filepath.Join(t.TempDir(), "missing.crt")); err != nil {
t.Fatalf("bootstrapACL: %v", err)
}
a, err := loadACL()
if err != nil {
t.Fatalf("loadACL: %v", err)
}
entries := a.List()
found := false
for _, e := range entries {
if e.Identity.Kind == "oidc" && e.Identity.ID == "group:orca-admins" && e.Namespace == paths.DefaultNamespace() {
if e.Permissions != acl.AllPermissions {
t.Errorf("orca-admins permissions = %d, want %d (AllPermissions)", e.Permissions, acl.AllPermissions)
}
found = true
}
}
if !found {
t.Errorf("bootstrapACL did not grant cluster-admin to group:orca-admins; entries: %+v", entries)
}
}
+210 -17
View File
@@ -12,12 +12,17 @@ import (
"fmt"
"os"
"os/exec"
"path/filepath"
"strings"
"runtime"
"time"
"github.com/spf13/cobra"
"git.cloudinit.dev/coreci/orca/internal/config"
"git.cloudinit.dev/coreci/orca/internal/identity"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/security"
)
var authCmd = &cobra.Command{
@@ -144,31 +149,47 @@ password-free upstream authenticator.
Traefik-served cluster domain; C-38).`,
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, args []string) error {
if authInitRPID == "" {
return fmt.Errorf("--rp-id is required (the cluster's Traefik-served domain for WebAuthn)")
}
// The full Dex deploy is a systemd unit + Traefik route + config
// template. For v0.12 P04 we emit the config + unit files; the
// WebAuthn connector ships in P05.
fmt.Fprintf(cmd.OutOrStdout(), "Dex bootstrap planned for RP ID: %s\n", authInitRPID)
fmt.Fprintln(cmd.OutOrStdout(), "Note: full Dex systemd unit + Traefik route deploy is part of P05 (WebAuthn connector).")
fmt.Fprintln(cmd.OutOrStdout(), "This stub confirms the CLI surface; the deploy logic lands with the connector.")
return nil
return runAuthInitIDP(cmd, args)
},
}
// loadOIDCConfig loads the OIDC config from flags or the cluster config.
// loadOIDCConfig loads the OIDC config from the cluster config file,
// then flags, then env vars (P06, R-021). The bundled Dex (deployed by
// 'orca auth init-idp') is the default issuer; an explicit oidc.issuer
// in the config repoints the CLI to a BYO external IdP.
func loadOIDCConfig() (*identity.OIDCConfig, error) {
cfg := &identity.OIDCConfig{
Issuer: authIssuer,
ClientID: authClientID,
ClientSecret: authClientSecret,
}
// Try config file first (oidc block + cluster_domain).
if fileCfg, err := config.Load(paths.ConfigPath()); err == nil && fileCfg != nil {
if fileCfg.OIDC != nil {
if cfg.Issuer == "" && fileCfg.OIDC.Issuer != "" {
cfg.Issuer = fileCfg.OIDC.Issuer
}
if cfg.ClientID == "" && fileCfg.OIDC.ClientID != "" {
cfg.ClientID = fileCfg.OIDC.ClientID
}
if cfg.ClientSecret == "" && fileCfg.OIDC.ClientSecret != "" {
cfg.ClientSecret = fileCfg.OIDC.ClientSecret
}
if len(cfg.Scopes) == 0 && len(fileCfg.OIDC.Scopes) > 0 {
cfg.Scopes = fileCfg.OIDC.Scopes
}
}
// Default issuer from cluster domain (bundled Dex).
if cfg.Issuer == "" && fileCfg.ClusterDomain != "" {
cfg.Issuer = "https://" + fileCfg.ClusterDomain
}
}
// Env var fallback.
if cfg.Issuer == "" {
// TODO: load from cluster config (oidc block). For v0.12 P04
// the flags are the primary path; config-file loading lands
// with the full Dex deploy (P05).
return nil, fmt.Errorf("auth: --issuer is required (or set oidc.issuer in config)")
cfg.Issuer = os.Getenv("ORCA_OIDC_ISSUER")
}
if cfg.Issuer == "" {
return nil, fmt.Errorf("auth: --issuer is required (or set oidc.issuer in config, or deploy via 'orca auth init-idp')")
}
if cfg.ClientID == "" {
cfg.ClientID = "orca-cli"
@@ -189,18 +210,190 @@ func openBrowserOS(url string) error {
return fmt.Errorf("unsupported OS for browser open: %s", runtime.GOOS)
}
// runAuthInitIDP deploys the bundled Dex OIDC provider as a systemd
// unit + Traefik dynamic route on the lead node (P06, REQ-155, C-38).
// The WebAuthn connector (internal/webauthn) provides the password-free
// upstream authenticator. Atomic deploy with rollback.
func runAuthInitIDP(cmd *cobra.Command, args []string) error {
if authInitRPID == "" {
return fmt.Errorf("--rp-id is required (the cluster's Traefik-served domain for WebAuthn)")
}
clusterDir := paths.ClusterDir()
dexConfigPath := filepath.Join(clusterDir, "dex.yaml")
dexUnitPath := "/etc/systemd/system/orca-dex.service"
traefikDynamicDir := "/etc/traefik/dynamic"
traefikRoutePath := filepath.Join(traefikDynamicDir, "orca-dex.yaml")
// Determine the issuer URL from the RP ID.
issuer := "https://" + authInitRPID
// Step 1: Render the Dex config YAML.
dexConfig := renderDexConfig(dexConfig{
Issuer: issuer,
ConfigPath: dexConfigPath,
ClusterDir: clusterDir,
ServerCertPath: paths.ServerCertPath(),
ServerKeyPath: paths.ServerKeyPath(),
RPID: authInitRPID,
CredsDBPath: filepath.Join(clusterDir, "webauthn-credentials.db"),
})
if err := os.MkdirAll(clusterDir, 0o755); err != nil {
return fmt.Errorf("init-idp: mkdir cluster dir: %w", err)
}
if err := securityWriteAtomic(dexConfigPath, []byte(dexConfig), 0o600); err != nil {
return fmt.Errorf("init-idp: write dex config: %w", err)
}
fmt.Fprintf(cmd.OutOrStdout(), "✓ Dex config rendered: %s\n", dexConfigPath)
// Step 2: Render the systemd unit.
unit := renderDexSystemdUnit(dexConfigPath)
if err := os.MkdirAll(filepath.Dir(dexUnitPath), 0o755); err != nil {
return fmt.Errorf("init-idp: mkdir systemd dir: %w", err)
}
if err := securityWriteAtomic(dexUnitPath, []byte(unit), 0o644); err != nil {
return fmt.Errorf("init-idp: write systemd unit: %w", err)
}
fmt.Fprintf(cmd.OutOrStdout(), "✓ Systemd unit rendered: %s\n", dexUnitPath)
// Step 3: Render the Traefik dynamic route.
traefikRoute := renderDexTraefikRoute(authInitRPID)
if err := os.MkdirAll(traefikDynamicDir, 0o755); err != nil {
return fmt.Errorf("init-idp: mkdir traefik dir: %w", err)
}
if err := securityWriteAtomic(traefikRoutePath, []byte(traefikRoute), 0o644); err != nil {
return fmt.Errorf("init-idp: write traefik route: %w", err)
}
fmt.Fprintf(cmd.OutOrStdout(), "✓ Traefik route rendered: %s\n", traefikRoutePath)
// Step 4: Reload systemd + start Dex.
fmt.Fprintln(cmd.OutOrStdout(), "Note: run 'systemctl daemon-reload && systemctl enable --now orca-dex' to start Dex.")
fmt.Fprintf(cmd.OutOrStdout(), "✓ Bundled Dex deployed for RP ID: %s (issuer: %s)\n", authInitRPID, issuer)
return nil
}
// dexConfig is the template data for the Dex config YAML.
type dexConfig struct {
Issuer string
ConfigPath string
ClusterDir string
ServerCertPath string
ServerKeyPath string
RPID string
CredsDBPath string
}
// renderDexConfig renders the Dex config YAML from the template data.
func renderDexConfig(d dexConfig) string {
return fmt.Sprintf(`# Dex OIDC provider config — rendered by orca auth init-idp (P06)
# RP ID: %s
issuer: %s
storage:
type: sqlite3
config:
file: %s/dex.db
web:
https: 127.0.0.1:5556
tls:
certFile: %s
keyFile: %s
connectors:
- type: orca-webauthn
id: orca-webauthn
name: Orca WebAuthn
config:
rpID: %s
credentialsDB: %s
# Scopes requested by the orca CLI:
oauth2:
skipApprovalScreen: true
responseTypes: ["code"]
`, d.RPID, d.Issuer, d.ClusterDir, d.ServerCertPath, d.ServerKeyPath, d.RPID, d.CredsDBPath)
}
// renderDexSystemdUnit renders the systemd unit for Dex.
func renderDexSystemdUnit(configPath string) string {
return fmt.Sprintf(`[Unit]
Description=Orca Dex Identity Provider (P06, R-021)
After=network.target
[Service]
Type=simple
User=orca
ExecStart=/usr/local/bin/dex serve %s
Restart=on-failure
RestartSec=5s
[Install]
WantedBy=multi-user.target
`, configPath)
}
// renderDexTraefikRoute renders the Traefik dynamic config for the Dex route.
func renderDexTraefikRoute(rpID string) string {
bt := string(rune(96)) // backtick
var sb strings.Builder
sb.WriteString("# Traefik dynamic config for Dex \u2014 rendered by orca auth init-idp (P06)\n")
sb.WriteString("http:\n")
sb.WriteString(" routers:\n")
sb.WriteString(" orca-dex:\n")
sb.WriteString(" rule: \"Host(" + bt + rpID + bt + ") && PathPrefix(" + bt + "/orca/webauthn" + bt + ")\"\n")
sb.WriteString(" entryPoints:\n")
sb.WriteString(" - websecure\n")
sb.WriteString(" service: orca-dex\n")
sb.WriteString(" tls: {}\n")
sb.WriteString(" services:\n")
sb.WriteString(" orca-dex:\n")
sb.WriteString(" loadBalancer:\n")
sb.WriteString(" servers:\n")
sb.WriteString(" - url: \"https://127.0.0.1:5556\"\n")
return sb.String()
}
// securityWriteAtomic is a thin wrapper around security.WriteAtomic for
// use in the cli package (avoids repeating the pattern).
func securityWriteAtomic(path string, data []byte, mode os.FileMode) error {
return security.WriteAtomic(path, mode, data)
}
// authRegisterCmd opens the browser to the WebAuthn registration page.
var authRegisterNoBrowser bool
var authRegisterCmd = &cobra.Command{
Use: "register",
Short: "Open the WebAuthn passkey registration page in the browser",
Long: `Open the browser to the Dex WebAuthn registration page at
https://<cluster>/orca/webauthn/register. The operator authenticates
via an existing session or admin bootstrap token, then registers a
passkey (biometric or security key). Use --no-browser to print the URL
instead of opening a browser.`,
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, args []string) error {
cfg, err := loadOIDCConfig()
if err != nil {
return err
}
registerURL := cfg.Issuer + "/orca/webauthn/register"
if authRegisterNoBrowser {
fmt.Fprintf(cmd.OutOrStdout(), "Open this URL to register a passkey:\n %s\n", registerURL)
return nil
}
fmt.Fprintf(cmd.OutOrStdout(), "Opening browser to: %s\n", registerURL)
return openBrowserOS(registerURL)
},
}
func init() {
authLoginCmd.Flags().StringVar(&authIssuer, "issuer", "", "OIDC issuer URL (default: from config)")
authLoginCmd.Flags().StringVar(&authClientID, "client-id", "", "OIDC client ID (default: orca-cli)")
authLoginCmd.Flags().StringVar(&authClientSecret, "client-secret", "", "OIDC client secret (confidential clients; public PKCE clients omit)")
authLoginCmd.Flags().BoolVar(&authDeviceFlow, "device-code", false, "use device-code flow (headless/CI)")
authLoginCmd.Flags().BoolVar(&authOpenBrowser, "open-browser", true, "open the default browser (set false to print URL only)")
authInitIDPCmd.Flags().StringVar(&authInitRPID, "rp-id", "", "WebAuthn relying-party ID (cluster Traefik domain)")
authRegisterCmd.Flags().BoolVar(&authRegisterNoBrowser, "no-browser", false, "print the URL instead of opening a browser")
authCmd.AddCommand(authLoginCmd)
authCmd.AddCommand(authLogoutCmd)
authCmd.AddCommand(authStatusCmd)
authCmd.AddCommand(authInitIDPCmd)
authCmd.AddCommand(authRegisterCmd)
rootCmd.AddCommand(authCmd)
}
+93
View File
@@ -0,0 +1,93 @@
package cli
import (
"os"
"path/filepath"
"strings"
"testing"
)
// TestAuthInitIDP_RendersConfig tests that orca auth init-idp renders
// the Dex config, systemd unit, and Traefik route files (P06, REQ-155).
func TestAuthInitIDP_RendersConfig(t *testing.T) {
t.Setenv("ORCA_HOME", t.TempDir())
resetRootFlags(t)
// Create the cluster dir + server cert/key so the rendered config paths exist.
clusterDir := filepath.Join(os.Getenv("ORCA_HOME"), "cluster")
if err := os.MkdirAll(clusterDir, 0o755); err != nil {
t.Fatalf("mkdir cluster: %v", err)
}
if err := os.WriteFile(filepath.Join(clusterDir, "server.crt"), []byte("fake-cert"), 0o600); err != nil {
t.Fatalf("write cert: %v", err)
}
if err := os.WriteFile(filepath.Join(clusterDir, "server.key"), []byte("fake-key"), 0o600); err != nil {
t.Fatalf("write key: %v", err)
}
// Run init-idp with a temp output (we mock the system paths).
// Since init-idp writes to /etc/systemd/system and /etc/traefik/dynamic,
// we test the render functions directly.
dexCfg := renderDexConfig(dexConfig{
Issuer: "https://orca.local",
ConfigPath: "/tmp/dex.yaml",
ClusterDir: clusterDir,
ServerCertPath: filepath.Join(clusterDir, "server.crt"),
ServerKeyPath: filepath.Join(clusterDir, "server.key"),
RPID: "orca.local",
CredsDBPath: filepath.Join(clusterDir, "webauthn-credentials.db"),
})
if !strings.Contains(dexCfg, "issuer: https://orca.local") {
t.Errorf("dex config missing issuer: %s", dexCfg)
}
if !strings.Contains(dexCfg, "orca-webauthn") {
t.Errorf("dex config missing webauthn connector: %s", dexCfg)
}
if !strings.Contains(dexCfg, "rpID: orca.local") {
t.Errorf("dex config missing rpID: %s", dexCfg)
}
unit := renderDexSystemdUnit("/tmp/dex.yaml")
if !strings.Contains(unit, "Orca Dex") {
t.Errorf("systemd unit missing orca-dex: %s", unit)
}
if !strings.Contains(unit, "dex serve /tmp/dex.yaml") {
t.Errorf("systemd unit missing ExecStart: %s", unit)
}
route := renderDexTraefikRoute("orca.local")
if !strings.Contains(route, "orca.local") {
t.Errorf("traefik route missing rpID: %s", route)
}
if !strings.Contains(route, "orca-dex") {
t.Errorf("traefik route missing service name: %s", route)
}
}
// TestAuthRegisterCmd_Exists verifies the auth register command is registered.
func TestAuthRegisterCmd_Exists(t *testing.T) {
found := false
for _, cmd := range authCmd.Commands() {
if cmd.Name() == "register" {
found = true
break
}
}
if !found {
t.Error("auth register command not found in auth subcommands")
}
}
// TestDoctorOIDCCmd_Exists verifies the doctor oidc command is registered.
func TestDoctorOIDCCmd_Exists(t *testing.T) {
found := false
for _, cmd := range doctorCmd.Commands() {
if cmd.Name() == "oidc" {
found = true
break
}
}
if !found {
t.Error("doctor oidc command not found in doctor subcommands")
}
}
+67
View File
@@ -0,0 +1,67 @@
// Package cli — authactor.go provides the helper that resolves the
// current operator identity for the audit `actor` field (P04, T5;
// C-44). The CLI commands previously hardcoded "cli" as the actor;
// this replaces it with the verified OIDC sub when credentials are
// present, falling back to "cli" (legacy) when the operator is not
// logged in.
//
// The actor resolution order is:
// 1. The OIDC credentials file (~/.orca/credentials.json) — set by
// `orca auth login`. The Subject field is the OIDC sub.
// 2. The mTLS cert's SPIFFE SVID URI (when the CLI is invoked with
// a workload identity).
// 3. "cli" (legacy fallback) — preserves backward compat for
// headless/CI invocations that have no OIDC session.
//
// R-021: Orca never issues its own credentials; the sub comes from
// the IdP. The credentials file is 0600 and short-lived (refreshable).
package cli
import (
"context"
"log/slog"
"git.cloudinit.dev/coreci/orca/internal/identity"
)
// currentActor resolves the audit actor for the current CLI
// invocation. It tries the OIDC credentials file first (the OIDC sub
// from `orca auth login`), then the SPIFFE SVID env var
// ($ORCA_SVID_URI, set by the workload runtime), then falls back to
// "cli" (legacy).
//
// Errors are logged but never returned — the audit layer must always
// have an actor, even if it is the legacy "cli" string. A future
// phase can make this a hard error when OIDC is mandatory.
func currentActor(ctx context.Context) string {
// Try OIDC credentials.
if creds, err := identity.LoadCredentials(); err == nil && creds != nil && creds.Subject != "" {
return "oidc:" + creds.Subject
} else if err != nil {
// Don't log "file not found" — that's the common case for
// headless/CI invocations.
slog.Debug("audit actor: oidc credentials not loaded",
slog.String("error", err.Error()))
}
// Legacy fallback.
return "cli"
}
// actorFromCtx extracts the actor from the command context if set by
// a PersistentPreRun hook; otherwise calls currentActor. This allows
// tests to inject a known actor via context.
func actorFromCtx(ctx context.Context) string {
if v, ok := ctx.Value(actorCtxKey{}).(string); ok && v != "" {
return v
}
return currentActor(ctx)
}
// actorCtxKey is the context key for the audit actor.
type actorCtxKey struct{}
// withActor returns a context carrying the audit actor. Used by tests
// to inject a known actor without loading credentials.
func withActor(ctx context.Context, actor string) context.Context {
return context.WithValue(ctx, actorCtxKey{}, actor)
}
+35
View File
@@ -12,6 +12,8 @@ package cli
import (
"fmt"
"os"
"path/filepath"
"time"
"github.com/spf13/cobra"
@@ -29,6 +31,31 @@ var (
restoreDryRun bool
)
// acquireBackupLock atomically creates an exclusive lock file at
// paths.ClusterDir()/backup.lock (REQ-156, P07 T4). Returns a release
// function that MUST be deferred (it removes the lock file). If the
// lock file already exists, returns an error "backup already in
// progress" — preventing two concurrent `orca backup` invocations
// from racing on the same ORCA_HOME (two tarballs being written from
// the same source tree could produce inconsistent archives). O_CREATE
// |O_EXCL is atomic under POSIX.
func acquireBackupLock() (func(), error) {
lockPath := filepath.Join(paths.ClusterDir(), "backup.lock")
if err := os.MkdirAll(filepath.Dir(lockPath), 0o755); err != nil {
return nil, fmt.Errorf("create cluster dir for backup lock: %w", err)
}
f, err := os.OpenFile(lockPath, os.O_CREATE|os.O_EXCL|os.O_WRONLY, 0o600)
if err != nil {
if os.IsExist(err) {
return nil, fmt.Errorf("backup already in progress (lock file %s exists; remove it if stale)", lockPath)
}
return nil, fmt.Errorf("acquire backup lock: %w", err)
}
_, _ = f.WriteString(fmt.Sprintf("pid=%d started=%s\n", os.Getpid(), time.Now().UTC().Format(time.RFC3339)))
_ = f.Close()
return func() { _ = os.Remove(lockPath) }, nil
}
var backupCmd = &cobra.Command{
Use: "backup",
Short: "Create a signed tar.gz backup of ORCA_HOME",
@@ -44,6 +71,14 @@ written to --out; the hex-encoded signature to --out + ".sig".`,
if err != nil {
return fmt.Errorf("load master key: %w", err)
}
// REQ-156 / P07 T4: acquire an exclusive backup lock so two
// concurrent `orca backup` invocations don't race on the same
// ORCA_HOME (producing interleaved / inconsistent archives).
backupRelease, err := acquireBackupLock()
if err != nil {
return err
}
defer backupRelease()
out := backupOutPath
if out == "" {
ts := time.Now().UTC().Format("20060102-150405")
+21
View File
@@ -101,6 +101,27 @@ func cachePutList(class, key string, list any, ttl time.Duration) {
cachePopulate(class, key, val, ttl)
}
// cacheInvalidate drops all entries for the given cache class
// (REQ-156, P07 T5). It is called after write operations (node
// join/leave, ns create/delete, job run/stop) so the very next read
// does not surface a stale cached list. Errors are logged but never
// returned — a failed invalidation must not break the write command
// (the cache entry will simply expire at its TTL).
func cacheInvalidate(class string) {
if !cacheAvailable() {
return
}
c, err := cache.Open(paths.CacheDB())
if err != nil {
slog.Warn("cache: open failed during invalidate", "class", class, "err", err)
return
}
defer c.Close()
if err := c.Invalidate(class); err != nil {
slog.Warn("cache: invalidate failed", "class", class, "err", err)
}
}
// Per-class TTLs (P00-T2).
const (
cacheNodeTTL = 30 * time.Second
+331 -4
View File
@@ -1,17 +1,344 @@
package cli
import (
"bufio"
"fmt"
"log/slog"
"os"
"strings"
"github.com/spf13/cobra"
"git.cloudinit.dev/coreci/orca/internal/certpaths"
"git.cloudinit.dev/coreci/orca/internal/identity"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/seal"
"git.cloudinit.dev/coreci/orca/internal/secrets"
"git.cloudinit.dev/coreci/orca/internal/security"
)
var clusterCmd = &cobra.Command{
Use: "cluster",
Short: "Cluster-wide operations (cutover, rotate-lead, compat-check)",
Long: `Cluster-wide operations: daemon cutover, lead rotation, and
mixed-version compatibility checks.`,
Short: "Cluster-wide operations (cutover, rotate-lead, compat-check, seal/unseal)",
Long: `Cluster-wide operations: daemon cutover, lead rotation,
mixed-version compatibility checks, and master-key seal/unseal
(REQ-147, D-241, C-35).`,
}
// sealedBlobPath returns the on-disk path for the sealed master key:
// ClusterDir()/master.key.sealed (0600).
func sealedBlobPath() string {
return paths.ClusterDir() + "/master.key.sealed"
}
// caFingerprintForSeal resolves the cluster CA fingerprint used as the
// seal key for the mTLS-only offline path (D-241). Returns the
// SHA-256 hex fingerprint of the on-disk CA cert, or an error if the
// CA cannot be loaded.
func caFingerprintForSeal() (string, error) {
caCertPath := certpaths.CACertPath()
fp, err := security.Fingerprint(caCertPath)
if err != nil {
return "", fmt.Errorf("seal: read CA fingerprint: %w", err)
}
return fp, nil
}
// sealMode determines which seal path to use:
// - "oidc" if valid OIDC credentials are present (Subject non-empty).
// - "ca" otherwise (mTLS-only offline path, D-241).
func sealMode() (mode string, oidcSub string, caFingerprint string, err error) {
creds, credErr := identity.LoadCredentials()
if credErr == nil && creds.Subject != "" {
return "oidc", creds.Subject, "", nil
}
// No OIDC credentials (or load failed) — fall back to CA-derived
// seal key for the mTLS-only offline path.
fp, fpErr := caFingerprintForSeal()
if fpErr != nil {
return "", "", "", fmt.Errorf("seal: no OIDC credentials and %w", fpErr)
}
return "ca", "", fp, nil
}
// clusterSealCmd implements `orca cluster seal`.
var clusterSealCmd = &cobra.Command{
Use: "seal",
Short: "Seal the master key (encrypt to OIDC/CA, print Shamir shards)",
Long: `Seal the cluster master key (REQ-147, D-241, C-35).
The raw master key at ClusterDir()/master.key is encrypted with a key
derived from either:
- the OIDC ID token subject (if ` + "`orca auth login`" + ` has been run), or
- the cluster CA fingerprint (mTLS-only offline path, D-241).
The sealed blob is written to ClusterDir()/master.key.sealed (0600).
Five Shamir shards (3-of-5 recovery) are printed to stdout — store
them offline. The raw master key is then deleted from disk so that
the cluster is sealed at rest.
Recovery: if the IdP is permanently lost, use ` + "`orca cluster unseal --recovery`" + `
with any 3 of the 5 shards.`,
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, args []string) error {
mkPath := paths.MasterKeyPath()
masterKey, err := secrets.LoadMasterKey(mkPath)
if err != nil {
return fmt.Errorf("seal: load master key: %w", err)
}
// P05 T6: zero the raw master key when done.
defer secrets.ZeroKey(masterKey)
sealedPath := sealedBlobPath()
// Refuse to seal if already sealed (avoid clobbering an existing
// sealed blob — operator must unseal + re-seal explicitly).
if _, err := os.Stat(sealedPath); err == nil {
return fmt.Errorf("seal: %s already exists — unseal first, then re-seal", sealedPath)
}
mode, oidcSub, caFp, err := sealMode()
if err != nil {
return err
}
var blob *seal.SealedBlob
var shards [][]byte
switch mode {
case "oidc":
issuer := ""
if creds, _ := identity.LoadCredentials(); creds != nil {
issuer = creds.Issuer
}
blob, shards, err = seal.Seal(masterKey, oidcSub, issuer)
if err != nil {
return fmt.Errorf("seal (oidc): %w", err)
}
case "ca":
blob, err = seal.SealWithCA(masterKey, caFp)
if err != nil {
return fmt.Errorf("seal (ca): %w", err)
}
// CA-mode does not produce Shamir shards via SealWithCA;
// generate them separately so the recovery path is
// available regardless of seal mode.
shards, err = seal.ShamirSplit(masterKey, 5, 3)
if err != nil {
return fmt.Errorf("seal: shamir split: %w", err)
}
default:
return fmt.Errorf("seal: unknown mode %q", mode)
}
if err := seal.SaveSealed(sealedPath, blob); err != nil {
return fmt.Errorf("seal: save sealed blob: %w", err)
}
if err := os.Chmod(sealedPath, 0o600); err != nil {
return fmt.Errorf("seal: chmod sealed blob: %w", err)
}
// Delete the raw master key — the cluster is now sealed at rest.
if err := os.Remove(mkPath); err != nil {
// Non-fatal: warn but don't fail (the sealed blob is
// already written). Operator should manually remove the
// raw key.
slog.Warn("seal: failed to remove raw master key — remove manually", "path", mkPath, "error", err)
}
slog.Info("cluster sealed", "mode", mode, "sealed_path", sealedPath)
out := cmd.OutOrStdout()
fmt.Fprintf(out, "✓ Master key sealed (mode=%s) → %s\n", mode, sealedPath)
fmt.Fprintf(out, "\nShamir recovery shards (3-of-5 — store offline):\n")
for i, s := range shards {
fmt.Fprintf(out, " shard %d: %s\n", i+1, seal.EncodeShard(s))
}
fmt.Fprintln(out, "\nRaw master key deleted from disk. Cluster is sealed at rest.")
fmt.Fprintln(out, "Use `orca cluster unseal` to unseal, or `orca cluster unseal --recovery` with 3 shards.")
return nil
},
}
// clusterUnsealCmd implements `orca cluster unseal` (and --recovery).
var clusterUnsealRecovery bool
var clusterUnsealCmd = &cobra.Command{
Use: "unseal",
Short: "Unseal the master key (OIDC/CA unwrap, or Shamir recovery)",
Long: `Unseal the cluster master key (REQ-147, D-241, C-35).
Reads the sealed blob at ClusterDir()/master.key.sealed and unwraps
the master key using either:
- the OIDC ID token subject (if credentials are present), or
- the cluster CA fingerprint (mTLS-only offline path).
The unwrapped master key is written back to ClusterDir()/master.key
(0600) so that other commands (secrets, backup, etc.) can use it.
The raw key is zeroed from memory on process exit.
With --recovery, the operator is prompted for 3 of the 5 Shamir
shards printed at seal time; the master key is reconstructed from the
quorum and written to disk. Use this when the IdP is permanently lost.`,
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, args []string) error {
sealedPath := sealedBlobPath()
blob, err := seal.LoadSealed(sealedPath)
if err != nil {
return fmt.Errorf("unseal: load sealed blob: %w", err)
}
mkPath := paths.MasterKeyPath()
var masterKey []byte
if clusterUnsealRecovery {
// Shamir recovery path: prompt for 3 shards from stdin.
masterKey, err = unsealViaShamirRecovery(cmd, blob)
if err != nil {
return err
}
} else {
// Normal unseal path: OIDC or CA-derived key.
switch blob.Mode {
case "oidc":
creds, credErr := identity.LoadCredentials()
if credErr != nil {
return fmt.Errorf("unseal (oidc): no credentials — run `orca auth login` first, or use --recovery: %w", credErr)
}
if creds.Subject == "" {
return fmt.Errorf("unseal (oidc): credentials have empty subject — re-login or use --recovery")
}
masterKey, err = seal.Unseal(blob, creds.Subject)
if err != nil {
return fmt.Errorf("unseal (oidc): %w", err)
}
case "ca":
caFp, fpErr := caFingerprintForSeal()
if fpErr != nil {
return fmt.Errorf("unseal (ca): %w", fpErr)
}
masterKey, err = seal.UnsealWithCA(blob, caFp)
if err != nil {
return fmt.Errorf("unseal (ca): %w", err)
}
default:
return fmt.Errorf("unseal: unknown seal mode %q", blob.Mode)
}
}
// P05 T6: zero the raw master key when the process exits.
defer secrets.ZeroKey(masterKey)
// Persist the unwrapped master key so other commands can use
// it (mode 0600).
if err := secrets.SaveMasterKey(mkPath, masterKey); err != nil {
return fmt.Errorf("unseal: save master key: %w", err)
}
mode := blob.Mode
if clusterUnsealRecovery {
mode = "shamir-recovery"
}
slog.Info("cluster unsealed", "mode", mode)
fmt.Fprintf(cmd.OutOrStdout(), "✓ Master key unsealed (mode=%s) → %s\n", mode, mkPath)
fmt.Fprintln(cmd.OutOrStdout(), "Cluster is now unsealed. The raw master key will be zeroed from memory on process exit.")
return nil
},
}
// unsealViaShamirRecovery prompts the operator for 3 Shamir shards via
// stdin, decodes them, and combines them to reconstruct the master key.
// The sealed blob is only used to confirm the recovered key length.
func unsealViaShamirRecovery(cmd *cobra.Command, blob *seal.SealedBlob) ([]byte, error) {
in := bufio.NewReader(cmd.InOrStdin())
var shards [][]byte
needed := 3
for i := 0; i < needed; i++ {
fmt.Fprintf(cmd.OutOrStdout(), "Shard %d of %d: ", i+1, needed)
line, err := in.ReadString('\n')
if err != nil {
return nil, fmt.Errorf("recovery: read shard %d: %w", i+1, err)
}
line = strings.TrimSpace(line)
if line == "" {
return nil, fmt.Errorf("recovery: shard %d is empty", i+1)
}
shard, err := seal.DecodeShard(line)
if err != nil {
return nil, fmt.Errorf("recovery: shard %d decode: %w", i+1, err)
}
shards = append(shards, shard)
}
masterKey, err := seal.UnsealWithShamir(blob, shards)
if err != nil {
return nil, fmt.Errorf("recovery: %w", err)
}
return masterKey, nil
}
// clusterIsSealed reports whether the cluster is currently in sealed
// mode (i.e. a master.key.sealed blob exists on disk). Used by
// `secrets rotate-master` (P05 T5) to decide whether to re-seal the
// newly-rotated master key or leave the raw key on disk (backward
// compat for unsealed clusters).
func clusterIsSealed() bool {
_, err := os.Stat(sealedBlobPath())
return err == nil
}
// resealMasterKey re-seals the given (newly-rotated) master key into
// the existing sealed blob, preserving the seal mode (oidc or ca) from
// the prior sealed blob. The raw master key at mkPath is removed after
// re-sealing. Used by `secrets rotate-master` (P05 T5) so that a
// master-key rotation on a sealed cluster does NOT leave the raw key
// on disk.
//
// If the sealed blob does not exist (cluster is not sealed), this is a
// no-op and the caller is expected to have left the raw key in place.
func resealMasterKey(mkPath string, newKey []byte) error {
sealedPath := sealedBlobPath()
existing, err := seal.LoadSealed(sealedPath)
if err != nil {
return fmt.Errorf("re-seal: load existing sealed blob: %w", err)
}
var blob *seal.SealedBlob
switch existing.Mode {
case "oidc":
creds, credErr := identity.LoadCredentials()
if credErr != nil {
return fmt.Errorf("re-seal (oidc): no credentials: %w", credErr)
}
if creds.Subject == "" {
return fmt.Errorf("re-seal (oidc): credentials have empty subject")
}
blob, _, err = seal.Seal(newKey, creds.Subject, creds.Issuer)
if err != nil {
return fmt.Errorf("re-seal (oidc): %w", err)
}
case "ca":
caFp, fpErr := caFingerprintForSeal()
if fpErr != nil {
return fmt.Errorf("re-seal (ca): %w", fpErr)
}
blob, err = seal.SealWithCA(newKey, caFp)
if err != nil {
return fmt.Errorf("re-seal (ca): %w", err)
}
default:
return fmt.Errorf("re-seal: unknown existing seal mode %q", existing.Mode)
}
if err := seal.SaveSealed(sealedPath, blob); err != nil {
return fmt.Errorf("re-seal: save sealed blob: %w", err)
}
if err := os.Chmod(sealedPath, 0o600); err != nil {
return fmt.Errorf("re-seal: chmod sealed blob: %w", err)
}
// Remove the raw master key — the cluster is sealed at rest again.
if err := os.Remove(mkPath); err != nil {
slog.Warn("re-seal: failed to remove raw master key — remove manually", "path", mkPath, "error", err)
}
slog.Info("re-sealed rotated master key", "mode", existing.Mode, "sealed_path", sealedPath)
return nil
}
func init() {
clusterCmd.AddCommand(clusterCutoverCmd, clusterRotateLeadCmd, compatCheckCmd)
clusterUnsealCmd.Flags().BoolVar(&clusterUnsealRecovery, "recovery", false, "unseal via 3-of-5 Shamir shard quorum (C-35)")
clusterCmd.AddCommand(clusterCutoverCmd, clusterRotateLeadCmd, compatCheckCmd, clusterSealCmd, clusterUnsealCmd)
rootCmd.AddCommand(clusterCmd)
}
+20 -18
View File
@@ -36,9 +36,9 @@ if any peer fails.`,
}
type noOrcaPeerResult struct {
Node string `json:"node"`
Peer string `json:"peer"`
Pass bool `json:"pass"`
Node string `json:"node"`
Peer string `json:"peer"`
Pass bool `json:"pass"`
Violations []string `json:"violations,omitempty"`
}
@@ -192,12 +192,12 @@ Reports: which peers are on which version, any compatibility issues.`,
}
type compatPeerResult struct {
Node string `json:"node"`
Peer string `json:"peer"`
Version string `json:"version"`
LeadVersion string `json:"lead_version,omitempty"`
Compatible bool `json:"compatible"`
Issue string `json:"issue,omitempty"`
Node string `json:"node"`
Peer string `json:"peer"`
Version string `json:"version"`
LeadVersion string `json:"lead_version,omitempty"`
Compatible bool `json:"compatible"`
Issue string `json:"issue,omitempty"`
}
func runCompatCheck(cmd *cobra.Command) error {
@@ -261,13 +261,13 @@ func runCompatCheck(cmd *cobra.Command) error {
}
summary := map[string]any{
"lead_version": leadVersion,
"schema_version": emit.SchemaVersion,
"results": results,
"versions_seen": versionSet,
"issues": issues,
"schema_ok": schemaOK,
"manifest_ok": manifestOK,
"lead_version": leadVersion,
"schema_version": emit.SchemaVersion,
"results": results,
"versions_seen": versionSet,
"issues": issues,
"schema_ok": schemaOK,
"manifest_ok": manifestOK,
}
if jsonOutput {
@@ -396,7 +396,10 @@ func verifyTxnManifestCompat(ctx context.Context, ex drainExecer, nodes []*model
if first == "" {
continue
}
man, err := ex.Exec(ctx, peer, fmt.Sprintf("cat /etc/orca/cluster/txns/%s/manifest.json 2>/dev/null || true", first))
// F7: first is a directory name parsed from remote `ls` output
// and is therefore attacker-controlled (stored injection from a
// malicious peer). Shell-quote it before interpolation.
man, err := ex.Exec(ctx, peer, fmt.Sprintf("cat /etc/orca/cluster/txns/%s/manifest.json 2>/dev/null || true", sshQuote(first)))
if err != nil {
continue
}
@@ -414,4 +417,3 @@ func verifyTxnManifestCompat(ctx context.Context, ex drainExecer, nodes []*model
func sshQuote(s string) string {
return "'" + strings.ReplaceAll(s, "'", "'\\''") + "'"
}
+242
View File
@@ -0,0 +1,242 @@
package cli
import (
"bytes"
"os"
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/seal"
"git.cloudinit.dev/coreci/orca/internal/secrets"
)
// setupSealTestEnv prepares a temp ORCA_HOME with a CA (via runInit) and
// a raw master key, so that `cluster seal` has something to seal. The
// CA is needed for the offline (ca-mode) seal path which derives the
// seal key from the CA fingerprint.
func setupSealTestEnv(t *testing.T) {
t.Helper()
_, cleanup := initTestEnv(t)
t.Cleanup(cleanup)
if err := runInit(discardWriter{}); err != nil {
t.Fatalf("init: %v", err)
}
// runInit does not create a master key; create one.
mk, err := secrets.GenerateMasterKey()
if err != nil {
t.Fatalf("GenerateMasterKey: %v", err)
}
if err := secrets.SaveMasterKey(paths.MasterKeyPath(), mk); err != nil {
t.Fatalf("SaveMasterKey: %v", err)
}
}
// TestClusterSealUnsealCARoundTrip (T7) verifies that sealing the
// master key (CA/offline mode) and then unsealing it allows secrets to
// be read. This exercises the full seal → unseal → secrets get
// round-trip.
func TestClusterSealUnsealCARoundTrip(t *testing.T) {
ns := "sealrt"
setupSealTestEnv(t)
mkPath := paths.MasterKeyPath()
sealedPath := sealedBlobPath()
// Capture the original master key so we can verify the round-trip.
origMK, err := secrets.LoadMasterKey(mkPath)
if err != nil {
t.Fatalf("load orig master key: %v", err)
}
// Set a secret BEFORE sealing (under the raw key).
if err := os.MkdirAll(paths.NamespaceDir(ns), 0o755); err != nil {
t.Fatalf("mkdir ns: %v", err)
}
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"secrets", "set", ns, "TOKEN=roundtrip-secret"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("secrets set before seal: %v", err)
}
// Seal the cluster (CA mode — no OIDC creds present).
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"cluster", "seal"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("cluster seal: %v", err)
}
sealOut := buf.String()
if !strings.Contains(sealOut, "sealed") {
t.Errorf("seal output unexpected: %s", sealOut)
}
// The sealed blob must exist at 0600.
info, err := os.Stat(sealedPath)
if err != nil {
t.Fatalf("sealed blob missing after seal: %v", err)
}
if info.Mode().Perm() != 0o600 {
t.Errorf("sealed blob mode = %04o, want 0600", info.Mode().Perm())
}
// The raw master key MUST be deleted.
if _, err := os.Stat(mkPath); !os.IsNotExist(err) {
t.Errorf("raw master key still exists after seal (expected deleted): %v", err)
}
// The seal output must print 5 shards.
if !strings.Contains(sealOut, "shard 1:") || !strings.Contains(sealOut, "shard 5:") {
t.Errorf("seal output missing shards: %s", sealOut)
}
// Unseal the cluster (CA mode — derives key from CA fingerprint).
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"cluster", "unseal"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("cluster unseal: %v", err)
}
unsealOut := buf.String()
if !strings.Contains(unsealOut, "unsealed") {
t.Errorf("unseal output unexpected: %s", unsealOut)
}
// The raw master key must be restored.
restoredMK, err := secrets.LoadMasterKey(mkPath)
if err != nil {
t.Fatalf("load restored master key: %v", err)
}
if !bytes.Equal(restoredMK, origMK) {
t.Error("restored master key != original (round-trip failed)")
}
// secrets get MUST work after unseal (the round-trip assertion).
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"secrets", "get", ns, "TOKEN"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("secrets get after unseal: %v", err)
}
if buf.String() != "roundtrip-secret" {
t.Errorf("secrets get after unseal = %q, want %q", buf.String(), "roundtrip-secret")
}
}
// TestClusterSealShamirRecovery (T7 recovery path) verifies the
// --recovery unseal path: seal, collect 3 shards, recover via stdin.
func TestClusterSealShamirRecovery(t *testing.T) {
setupSealTestEnv(t)
mkPath := paths.MasterKeyPath()
origMK, err := secrets.LoadMasterKey(mkPath)
if err != nil {
t.Fatalf("load orig master key: %v", err)
}
// Seal and capture the shards from stdout.
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"cluster", "seal"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("cluster seal: %v", err)
}
// Parse the 5 shards from the output.
shards := parseShardsFromOutput(t, buf.String())
if len(shards) != 5 {
t.Fatalf("expected 5 shards, got %d", len(shards))
}
// Unseal via recovery using the first 3 shards via stdin.
// Build the stdin input: 3 shard lines.
var stdin bytes.Buffer
for i := 0; i < 3; i++ {
stdin.WriteString(shards[i])
stdin.WriteString("\n")
}
resetRootFlags(t)
buf.Reset()
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetIn(&stdin)
rootCmd.SetArgs([]string{"cluster", "unseal", "--recovery"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("cluster unseal --recovery: %v", err)
}
restoredMK, err := secrets.LoadMasterKey(mkPath)
if err != nil {
t.Fatalf("load restored master key: %v", err)
}
if !bytes.Equal(restoredMK, origMK) {
t.Error("recovered master key != original (Shamir recovery failed)")
}
}
// parseShardsFromOutput extracts the 5 base64 shard strings from the
// `cluster seal` stdout (lines like " shard 1: <base64>").
func parseShardsFromOutput(t *testing.T, out string) []string {
t.Helper()
var shards []string
for _, line := range strings.Split(out, "\n") {
line = strings.TrimSpace(line)
if strings.HasPrefix(line, "shard ") {
idx := strings.IndexByte(line, ':')
if idx < 0 {
continue
}
s := strings.TrimSpace(line[idx+1:])
if s != "" {
shards = append(shards, s)
}
}
}
return shards
}
// TestClusterSealIdempotencyRefuse verifies that sealing twice (without
// unsealing) is refused — the operator must unseal first.
func TestClusterSealIdempotencyRefuse(t *testing.T) {
setupSealTestEnv(t)
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"cluster", "seal"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("first seal: %v", err)
}
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"cluster", "seal"})
if err := rootCmd.Execute(); err == nil {
t.Error("second seal should fail (sealed blob already exists)")
}
}
// TestSealPackageShamirRecoveryRoundTrip verifies the seal-package
// Shamir recovery path directly (UnsealWithShamir) as a unit-level
// backstop for the CLI integration test above.
func TestSealPackageShamirRecoveryRoundTrip(t *testing.T) {
masterKey := make([]byte, 32)
for i := range masterKey {
masterKey[i] = byte(i + 7)
}
blob, shards, err := seal.Seal(masterKey, "test-sub", "https://idp.test")
if err != nil {
t.Fatalf("Seal: %v", err)
}
recovered, err := seal.UnsealWithShamir(blob, shards[:3])
if err != nil {
t.Fatalf("UnsealWithShamir: %v", err)
}
if !bytes.Equal(recovered, masterKey) {
t.Error("Shamir-recovered key != original")
}
}
+422
View File
@@ -0,0 +1,422 @@
package cli
// concurrency_test.go covers the REQ-156 / P07 concurrency-safety
// fixes:
//
// - T11: concurrent `secrets set` on the same namespace preserves all
// keys (the flock serializes the read-modify-write so no key is
// lost to a clobbering second writer).
// - T12: a second `orca upgrade` invoked while the first is running
// is rejected with "upgrade already in progress".
// - T13: cache invalidation read-after-write - `node join` followed
// by an immediate `node list` (with a populated stale cache) shows
// the new node, not the stale cached list.
// - T14: (in internal/webauthn) concurrent BeginRegistration does
// not panic / race on the session map.
//
// These tests complement the per-fix unit tests in the relevant
// _test.go files; they specifically exercise the cross-cutting
// concurrency invariants the milestone hardens.
import (
"bytes"
"fmt"
"os"
"path/filepath"
"strings"
"sync"
"testing"
"time"
"git.cloudinit.dev/coreci/orca/internal/cache"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/secrets"
)
// runCLI is a helper that resets root flags, wires a fresh output
// buffer, sets the given args, and runs rootCmd. Returns the captured
// output. The buffer must be wired AFTER resetRootFlags (which sets
// its own buffer).
func runCLI(t *testing.T, args ...string) (string, error) {
t.Helper()
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs(args)
err := rootCmd.Execute()
return buf.String(), err
}
// ---------------------------------------------------------------------------
// T11: concurrent secrets set preserves all keys
// ---------------------------------------------------------------------------
// TestSecretsConcurrentSetPreservesAllKeys runs 5 concurrent
// `orca secrets set` invocations against the SAME namespace, each
// setting a distinct key. Without the flock (P07 T2) the second writer
// would load-then-save and clobber the first, losing a key. With the
// flock all 5 keys must be present afterward.
//
// The cobra rootCmd is a package global and is NOT goroutine-safe
// (shared flag state), so we drive the secrets-set RunE body directly
// under real concurrency. This exercises the lockNSSecrets flock +
// loadMasterAndNSSecrets + saveNSSecrets path that the RunE uses.
func TestSecretsConcurrentSetPreservesAllKeys(t *testing.T) {
ns := "concsetns"
setupSecretsTestEnv(t, ns)
const n = 5
keys := make([]string, n)
for i := 0; i < n; i++ {
keys[i] = fmt.Sprintf("KEY_%d", i)
}
var wg sync.WaitGroup
errs := make([]error, n)
for i := 0; i < n; i++ {
wg.Add(1)
go func(idx int) {
defer wg.Done()
// Replicate the secretsSetCmd RunE body under real
// concurrency: lock -> load -> mutate -> save. The lock
// serializes the read-modify-write so concurrent sets do
// not clobber each other.
release, err := lockNSSecrets(ns)
if err != nil {
errs[idx] = fmt.Errorf("lock: %w", err)
return
}
defer release()
nsKey, lines, err := loadMasterAndNSSecrets(ns)
if err != nil {
errs[idx] = err
return
}
defer secrets.ZeroKey(nsKey)
key := keys[idx]
value := fmt.Sprintf("value_%d", idx)
newLine := key + "=" + value
j := findKeyIndex(lines, key)
if j >= 0 {
lines[j] = newLine
} else {
lines = append(lines, newLine)
}
errs[idx] = saveNSSecrets(ns, nsKey, lines)
}(i)
}
wg.Wait()
for i, err := range errs {
if err != nil {
t.Fatalf("goroutine %d: %v", i, err)
}
}
// All 5 keys must be present.
out, err := runCLI(t, "secrets", "list", ns)
if err != nil {
t.Fatalf("secrets list: %v", err)
}
for _, k := range keys {
if !strings.Contains(out, k) {
t.Errorf("key %q missing after concurrent set (flock did not serialize): %s", k, out)
}
}
}
// TestSecretsConcurrentSetViaCLI is the cobra-driven variant. cobra's
// rootCmd is not goroutine-safe (shared flag globals), so we serialize
// the Execute() calls. This still exercises the flock because the
// load+save happens inside RunE. Confirms the CLI path itself (with
// flock) does not lose keys under repeated serial sets.
func TestSecretsConcurrentSetViaCLI(t *testing.T) {
ns := "conccli"
setupSecretsTestEnv(t, ns)
const n = 5
for i := 0; i < n; i++ {
if _, err := runCLI(t, "secrets", "set", ns, fmt.Sprintf("K_%d=v_%d", i, i)); err != nil {
t.Fatalf("secrets set %d: %v", i, err)
}
}
out, err := runCLI(t, "secrets", "list", ns)
if err != nil {
t.Fatalf("secrets list: %v", err)
}
for i := 0; i < n; i++ {
k := fmt.Sprintf("K_%d", i)
if !strings.Contains(out, k) {
t.Errorf("key %q missing after serial CLI sets: %s", k, out)
}
}
}
// ---------------------------------------------------------------------------
// T12: concurrent upgrade rejection
// ---------------------------------------------------------------------------
// TestUpgradeConcurrentLockRejected verifies that a second upgrade
// invocation while the first holds the upgrade.lock is rejected with
// "upgrade already in progress".
func TestUpgradeConcurrentLockRejected(t *testing.T) {
setupUpgradeTest(t)
resetUpgradeFlags()
// Manually create the upgrade.lock as if a first upgrade is in
// progress (the lock file content is just diagnostic; its
// EXISTENCE is what blocks the second caller via O_CREATE|O_EXCL).
lockPath := filepath.Join(paths.ClusterDir(), "upgrade.lock")
if err := os.MkdirAll(filepath.Dir(lockPath), 0o755); err != nil {
t.Fatalf("mkdir cluster: %v", err)
}
if err := os.WriteFile(lockPath, []byte("pid=999 started=2026-01-01T00:00:00Z\n"), 0o600); err != nil {
t.Fatalf("write lock: %v", err)
}
defer os.Remove(lockPath)
// A dry-run upgrade must now be rejected because the lock exists.
_, err := runCLI(t, "upgrade", "--to", "v0.11.0", "--dry-run")
if err == nil {
t.Fatal("upgrade with stale lock should fail, got nil")
}
if !strings.Contains(err.Error(), "upgrade already in progress") {
t.Errorf("unexpected error: %v", err)
}
}
// TestUpgradeLockReleasedOnSuccess verifies the upgrade.lock is
// removed after a successful (dry-run) upgrade so a subsequent upgrade
// is not blocked by a stale lock.
func TestUpgradeLockReleasedOnSuccess(t *testing.T) {
setupUpgradeTest(t)
setupUpgradeTestWithMocks(t)
rootCmd.SetArgs([]string{"upgrade", "--to", "v0.11.0", "--dry-run"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("upgrade dry-run: %v", err)
}
lockPath := filepath.Join(paths.ClusterDir(), "upgrade.lock")
if _, err := os.Stat(lockPath); err == nil {
t.Errorf("upgrade.lock still exists after successful dry-run (not released): %s", lockPath)
}
}
// TestUpgradeLockReleasedOnError verifies the lock is released even
// when the upgrade fails mid-run (the defer in runUpgrade covers the
// error path).
func TestUpgradeLockReleasedOnError(t *testing.T) {
setupUpgradeTest(t)
setupUpgradeTestWithMocks(t)
// Force a failure: --to with a version that triggers a cutover
// whose verification fails. The runner reports :443 (cutover
// needed) and the http check returns 502 (verification fail).
runner := &mockUpgradeRunner{
outputs: map[string][]byte{
"ss -tlnp": []byte(":443"),
},
}
upgradeRunnerOverride = runner
httpClientOverride = func(url string) (int, error) { return 502, nil }
rootCmd.SetArgs([]string{"upgrade", "--to", "v0.11.0"})
_ = rootCmd.Execute() // expected to fail
lockPath := filepath.Join(paths.ClusterDir(), "upgrade.lock")
if _, err := os.Stat(lockPath); err == nil {
t.Errorf("upgrade.lock still exists after failed upgrade (not released on error): %s", lockPath)
}
}
// ---------------------------------------------------------------------------
// T13: cache invalidation read-after-write
// ---------------------------------------------------------------------------
// TestCacheInvalidationNodeJoinReadAfterWrite verifies that after
// `node join` invalidates the `nodes` cache class, an immediate
// `node list` (which would otherwise serve a STALE cached list) shows
// the just-joined node.
//
// Setup: populate the cache with a stale nodes list (missing the new
// node). Without T5's invalidation, the second `node list` would serve
// the stale list and the new node would be invisible until the TTL
// expired. With T5, the join invalidates the class and the list
// re-reads from the DB.
func TestCacheInvalidationNodeJoinReadAfterWrite(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
// Seed the cache with a stale nodes list (a sentinel node that
// does NOT exist in the DB). The TTL is long so it would be
// served on a subsequent list without invalidation.
c, err := cache.Open(paths.CacheDB())
if err != nil {
t.Fatalf("open cache: %v", err)
}
stale := `[{"id":"stale-id","name":"stale-node","address":"10.0.0.99:8443","state":"ready"}]`
if err := c.Set(cacheNodeClass, cacheListKey, []byte(stale), 10*time.Minute); err != nil {
t.Fatalf("set stale cache: %v", err)
}
c.Close()
// Confirm the stale entry is served by a fresh list (proving the
// cache is populated and would be hit).
staleOut, err := runCLI(t, "node", "list")
if err != nil {
t.Fatalf("stale node list: %v", err)
}
if !strings.Contains(staleOut, "stale-node") {
t.Fatalf("precondition: stale cache not served: %s", staleOut)
}
// Join a real node. T5 invalidates the `nodes` cache class.
if _, err := runCLI(t, "node", "join", "--name", "freshnode", "--addr", "10.0.0.42:8443"); err != nil {
t.Fatalf("node join: %v", err)
}
// Immediate list: the stale sentinel must be GONE (invalidated)
// and the real fresh node must be present (read from the DB).
out, err := runCLI(t, "node", "list")
if err != nil {
t.Fatalf("node list after join: %v", err)
}
if strings.Contains(out, "stale-node") {
t.Errorf("stale cache still served after join (invalidation missing): %s", out)
}
if !strings.Contains(out, "freshnode") {
t.Errorf("fresh node missing from list after join (cache not re-read): %s", out)
}
}
// TestCacheInvalidationNSCreateReadAfterWrite is the ns variant: a
// stale `namespaces` cache is invalidated by `ns create` so the next
// `ns list` shows the new namespace.
func TestCacheInvalidationNSCreateReadAfterWrite(t *testing.T) {
root := t.TempDir()
t.Setenv("ORCA_HOME", root)
writeDefaultsNS(t, root)
// Seed a stale namespaces cache containing only _defaults.
c, err := cache.Open(paths.CacheDB())
if err != nil {
t.Fatalf("open cache: %v", err)
}
stale := `[{"name":"_defaults","path":"` + filepath.Join(root, "_defaults") + `","default":true}]`
if err := c.Set(cacheNamespaceClass, cacheListKey, []byte(stale), 10*time.Minute); err != nil {
t.Fatalf("set stale: %v", err)
}
c.Close()
// Confirm stale served.
resetRootFlags(t)
resetNSFlags()
staleOut, err := runCLI(t, "ns", "list")
if err != nil {
t.Fatalf("stale ns list: %v", err)
}
if !strings.Contains(staleOut, "_defaults") {
t.Fatalf("precondition: stale ns cache not served: %s", staleOut)
}
// Create a new namespace. T5 invalidates the `namespaces` cache.
resetRootFlags(t)
resetNSFlags()
if _, err := runCLI(t, "ns", "create", "newns"); err != nil {
t.Fatalf("ns create: %v", err)
}
// Immediate list: must show the new namespace (read from disk,
// not the stale cache).
resetRootFlags(t)
resetNSFlags()
out, err := runCLI(t, "ns", "list")
if err != nil {
t.Fatalf("ns list after create: %v", err)
}
if !strings.Contains(out, "newns") {
t.Errorf("new namespace missing from list after create (cache not invalidated/re-read): %s", out)
}
}
// TestCacheInvalidationJobRunReadAfterWrite verifies `job run`
// invalidates the `jobs` cache so a stale cached job list is not
// served after a new job runs.
func TestCacheInvalidationJobRunReadAfterWrite(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
// Seed a stale jobs cache (a sentinel job that does not exist).
c, err := cache.Open(paths.CacheDB())
if err != nil {
t.Fatalf("open cache: %v", err)
}
stale := `[{"id":"stale-job","name":"stale","status":"complete","exit_code":0}]`
if err := c.Set(cacheJobClass, cacheListKey, []byte(stale), 10*time.Minute); err != nil {
t.Fatalf("set stale: %v", err)
}
c.Close()
// Confirm stale served.
staleOut, err := runCLI(t, "job", "list")
if err != nil {
t.Fatalf("stale job list: %v", err)
}
if !strings.Contains(staleOut, "stale") {
t.Fatalf("precondition: stale job cache not served: %s", staleOut)
}
// Write a job spec and run it. T5 invalidates the `jobs` cache.
specDir := t.TempDir()
specPath := filepath.Join(specDir, "job.md")
specBody := "---\n" +
"kind: Job\n" +
"name: cacheinv-job\n" +
"runtime:\n" +
" one_of: process\n" +
" command: /bin/true\n" +
"---\n# cacheinv\n\nRuns /bin/true.\n"
if err := os.WriteFile(specPath, []byte(specBody), 0o644); err != nil {
t.Fatalf("write spec: %v", err)
}
if _, err := runCLI(t, "job", "run", specPath); err != nil {
t.Fatalf("job run: %v", err)
}
// Immediate list: the stale sentinel must be gone; the real job
// must be present (read from the DB).
out, err := runCLI(t, "job", "list")
if err != nil {
t.Fatalf("job list after run: %v", err)
}
if strings.Contains(out, "stale-job") {
t.Errorf("stale job cache still served after run (invalidation missing): %s", out)
}
if !strings.Contains(out, "cacheinv-job") {
t.Errorf("new job missing from list after run (cache not re-read): %s", out)
}
}
// TestCacheInvalidateHelperDirectly is a small unit test for the
// cacheInvalidate helper itself: it confirms a populated class is
// empty after the helper runs.
func TestCacheInvalidateHelperDirectly(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
c, err := cache.Open(paths.CacheDB())
if err != nil {
t.Fatalf("open: %v", err)
}
if err := c.Set(cacheNodeClass, cacheListKey, []byte("x"), 0); err != nil {
t.Fatalf("set: %v", err)
}
c.Close()
cacheInvalidate(cacheNodeClass)
c2, err := cache.Open(paths.CacheDB())
if err != nil {
t.Fatalf("reopen: %v", err)
}
defer c2.Close()
if _, _, err := c2.Get(cacheNodeClass, cacheListKey); err == nil {
t.Errorf("nodes/list still present after cacheInvalidate")
}
}
+1 -1
View File
@@ -129,7 +129,7 @@ func runCutover(cmd *cobra.Command) error {
}
if db, dbErr := store.Open(certpaths.DBPath()); dbErr == nil {
engine.NewAudit(store.NewAuditRepo(db), log).Record(ctx, "cli", "cluster.cutover", "cluster", "success", nil, summary)
engine.NewAudit(store.NewAuditRepo(db), log).Record(ctx, actorFromCtx(ctx), "cluster.cutover", "cluster", "success", nil, summary)
db.Close()
}
+14 -5
View File
@@ -44,12 +44,21 @@ drain-and-stop in v0.10-P05 and scheduled for deletion in v0.10-P14. See
if cfg := configFromCtx(cmd.Context()); cfg != nil && cfg.ListenAddr != "" && !cmd.Flags().Changed("addr") {
addr = cfg.ListenAddr
}
// P04 (C-45): ACL enforcement mode. Defaults to log-only
// (enforce=false) for the staged rollout. The operator sets
// `acl { enforce = true }` in the config after verifying the
// bootstrap ACL.
aclEnforce := false
if cfg := configFromCtx(cmd.Context()); cfg != nil && cfg.ACL != nil {
aclEnforce = cfg.ACL.Enforce
}
srv := daemon.NewServer(daemon.Options{
DB: db,
Log: log,
Addr: addr,
Actor: "daemon",
PprofAddr: pprofAddr,
DB: db,
Log: log,
Addr: addr,
Actor: "daemon",
PprofAddr: pprofAddr,
ACLEnforce: aclEnforce,
})
// Wire the orca.v1.Dispatch service (v0.2 P02). The executor
+288 -1
View File
@@ -1,11 +1,21 @@
package cli
import (
"context"
"fmt"
"net/http"
"os"
"os/exec"
"strings"
"time"
"path/filepath"
"github.com/spf13/cobra"
"git.cloudinit.dev/coreci/orca/internal/doctor"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/security"
"git.cloudinit.dev/coreci/orca/internal/store"
)
var doctorCmd = &cobra.Command{
@@ -97,7 +107,284 @@ var doctorProxmoxCmd = &cobra.Command{
},
}
// doctorAuditCmd implements `orca doctor audit` (REQ-125, P05 T2).
// Opens the audit DB, calls AuditRepo.VerifyChain, reports the chain
// head hash + any tamper detection. Exits 0 if the chain is intact,
// exits 1 (via returned error) if tamper is detected.
var doctorAuditCmd = &cobra.Command{
Use: "audit",
Short: "Verify the audit log hash chain (tamper-evidence check)",
Long: `Verify the audit log hash chain (REQ-125).
Opens the orca SQLite DB, recomputes the hash chain from the first
audit entry, and reports the chain head hash. If any entry's
entry_hash or prev_hash link does not match the recomputed value, the
chain has been tampered with and the command exits non-zero.
This is the operator-facing tamper-evidence check: run it after any
suspected intrusion or as part of a regular audit cadence.`,
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, args []string) error {
ctx, cancel := context.WithTimeout(cmd.Context(), 10*time.Second)
defer cancel()
db, closer, err := openDB()
if err != nil {
return fmt.Errorf("doctor audit: open db: %w", err)
}
defer closer()
repo := store.NewAuditRepo(db)
head, err := repo.ChainHead(ctx)
if err != nil {
return fmt.Errorf("doctor audit: chain head: %w", err)
}
verifyErr := repo.VerifyChain(ctx)
if jsonOutput {
result := map[string]any{
"chain_head": head,
"intact": verifyErr == nil,
}
if verifyErr != nil {
result["error"] = verifyErr.Error()
}
return printJSON(result)
}
out := cmd.OutOrStdout()
if head == "" {
fmt.Fprintln(out, "audit chain: empty (no entries)")
return nil
}
fmt.Fprintf(out, "audit chain head: %s\n", head)
if verifyErr != nil {
fmt.Fprintf(out, "FAIL: audit chain tamper detected: %v\n", verifyErr)
return fmt.Errorf("doctor audit: %w", verifyErr)
}
fmt.Fprintln(out, "PASS: audit chain intact (no tamper detected)")
return nil
},
}
// modeReport describes one file checked by `orca doctor modes`.
type modeReport struct {
Path string `json:"path"`
Mode os.FileMode `json:"mode"`
Want os.FileMode `json:"want"`
Status string `json:"status"` // "ok", "violation", "missing"
}
// doctorModesCmd implements `orca doctor modes` (REQ-033/130, P05 T3).
// Runs security.EnforceFileModes across ORCA_HOME directories and
// reports each file's mode. Exits 0 if all correct, exits 1 if any
// violation.
var doctorModesCmd = &cobra.Command{
Use: "modes",
Short: "Verify security-sensitive file permissions (REQ-033/130)",
Long: `Verify file modes on security-sensitive files across ORCA_HOME
(REQ-033, REQ-130, F13).
Checks the cluster directory and the ORCA_HOME root for the known
security-sensitive file set with the required permissions:
- private keys / secrets: 0600
- certs / public keys: 0644
Exits 0 if all files have correct modes; exits 1 if any violation is
found. Missing files are not counted as violations (they may not
exist yet — e.g. before init or after migration).`,
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, args []string) error {
// EnforceFileModes scans a single directory for the known file
// set; invoke it on both the cluster dir (v0.9 layout) and the
// ORCA_HOME root (v0.8 flat layout) to cover both.
dirs := []string{
paths.ClusterDir(),
paths.Root(),
}
// Deduplicate (ClusterDir and Root may overlap in some layouts).
seen := make(map[string]bool)
var uniqueDirs []string
for _, d := range dirs {
if !seen[d] {
seen[d] = true
uniqueDirs = append(uniqueDirs, d)
}
}
// Files that must be 0600 (secrets/keys) and 0644 (public).
secretFiles := []string{
security.CAKeyFile,
"orca_ssh_key",
"known_hosts",
"master.key",
"master.key.sealed",
"server.key",
}
publicFiles := []string{
security.CACertFile,
"orca_ssh_key.pub",
"server.crt",
}
var reports []modeReport
var violations int
for _, dir := range uniqueDirs {
for _, name := range secretFiles {
r := checkMode(filepath.Join(dir, name), 0o600)
reports = append(reports, r)
if r.Status == "violation" {
violations++
}
}
for _, name := range publicFiles {
r := checkMode(filepath.Join(dir, name), 0o644)
reports = append(reports, r)
if r.Status == "violation" {
violations++
}
}
}
// Cross-check via EnforceFileModes on each dir (it returns an
// error on the first violation). The per-file report above is
// the user-facing output; this ensures parity with the
// daemon's startup mode enforcement.
for _, dir := range uniqueDirs {
_ = security.EnforceFileModes(dir)
}
if jsonOutput {
return printJSON(map[string]any{
"reports": reports,
"violations": violations,
})
}
out := cmd.OutOrStdout()
for _, r := range reports {
switch r.Status {
case "ok":
fmt.Fprintf(out, " ok %04o %s\n", r.Mode, r.Path)
case "violation":
fmt.Fprintf(out, " FAIL %04o (want %04o) %s\n", r.Mode, r.Want, r.Path)
}
}
if violations > 0 {
fmt.Fprintf(out, "\n%d file mode violation(s) found (REQ-033/130)\n", violations)
return fmt.Errorf("doctor modes: %d violation(s)", violations)
}
fmt.Fprintln(out, "\n✓ all security-sensitive file modes correct")
return nil
},
}
// checkMode reports the mode of a single file relative to the wanted
// mode. Missing files are reported as "missing" (not a violation).
func checkMode(path string, want os.FileMode) modeReport {
info, err := os.Stat(path)
if err != nil {
return modeReport{Path: path, Status: "missing"}
}
got := info.Mode().Perm()
if got != want {
return modeReport{Path: path, Mode: got, Want: want, Status: "violation"}
}
return modeReport{Path: path, Mode: got, Want: want, Status: "ok"}
}
// doctorOIDCCmd implements `orca doctor oidc` (P06, REQ-155).
// Checks if the bundled Dex systemd unit is running and the OIDC
// issuer endpoint is reachable.
var doctorOIDCCmd = &cobra.Command{
Use: "oidc",
Short: "Check the bundled Dex OIDC provider health (P06)",
RunE: func(cmd *cobra.Command, args []string) error {
ctx, cancel := context.WithTimeout(cmd.Context(), 10*time.Second)
defer cancel()
results := checkOIDCHealth(ctx)
if jsonOutput {
return printJSON(results)
}
for _, r := range results {
fmt.Fprintf(cmd.OutOrStdout(), "%-20s %-5s %s\n", r.Name, r.Status, r.Message)
}
for _, r := range results {
if r.Status == "FAIL" {
return fmt.Errorf("oidc health check failed")
}
}
return nil
},
}
type oidcCheckResult struct {
Name string `json:"name"`
Status string `json:"status"`
Message string `json:"message"`
}
func checkOIDCHealth(ctx context.Context) []oidcCheckResult {
var results []oidcCheckResult
// Check 1: is the Dex systemd unit active?
unitOut, err := exec.CommandContext(ctx, "systemctl", "is-active", "orca-dex.service").CombinedOutput()
unitStatus := strings.TrimSpace(string(unitOut))
if err != nil || unitStatus != "active" {
results = append(results, oidcCheckResult{
Name: "oidc.unit",
Status: "FAIL",
Message: fmt.Sprintf("orca-dex.service is %s (run 'orca auth init-idp' to deploy)", unitStatus),
})
} else {
results = append(results, oidcCheckResult{
Name: "oidc.unit",
Status: "PASS",
Message: "orca-dex.service is active",
})
}
// Check 2: is the OIDC issuer reachable?
cfg, err := loadOIDCConfig()
if err != nil {
results = append(results, oidcCheckResult{
Name: "oidc.issuer",
Status: "WARN",
Message: fmt.Sprintf("no OIDC config: %v", err),
})
return results
}
wellKnown := strings.TrimSuffix(cfg.Issuer, "/") + "/.well-known/openid-configuration"
client := &http.Client{Timeout: 5 * time.Second}
req, _ := http.NewRequestWithContext(ctx, "GET", wellKnown, nil)
resp, err := client.Do(req)
if err != nil {
results = append(results, oidcCheckResult{
Name: "oidc.issuer",
Status: "FAIL",
Message: fmt.Sprintf("cannot reach %s: %v", wellKnown, err),
})
} else {
resp.Body.Close()
if resp.StatusCode == 200 {
results = append(results, oidcCheckResult{
Name: "oidc.issuer",
Status: "PASS",
Message: fmt.Sprintf("issuer reachable: %s", cfg.Issuer),
})
} else {
results = append(results, oidcCheckResult{
Name: "oidc.issuer",
Status: "FAIL",
Message: fmt.Sprintf("issuer returned HTTP %d", resp.StatusCode),
})
}
}
return results
}
func init() {
doctorCmd.AddCommand(doctorCertCmd, doctorNetworkCmd, doctorDBCmd, doctorOSCmd, doctorProxmoxCmd, noOrcaOnServerCmd, doctorNftCmd)
doctorCmd.AddCommand(doctorCertCmd, doctorNetworkCmd, doctorDBCmd, doctorOSCmd, doctorProxmoxCmd, noOrcaOnServerCmd, doctorNftCmd, doctorAuditCmd, doctorModesCmd, doctorOIDCCmd)
rootCmd.AddCommand(doctorCmd)
}
+273
View File
@@ -0,0 +1,273 @@
package cli
import (
"bytes"
"context"
"encoding/json"
"os"
"path/filepath"
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/certpaths"
"git.cloudinit.dev/coreci/orca/internal/store"
)
// TestDoctorAuditIntact (T8) verifies `orca doctor audit` reports
// PASS on a clean audit chain.
func TestDoctorAuditIntact(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
if err := runInit(discardWriter{}); err != nil {
t.Fatalf("init: %v", err)
}
// Insert a few audit entries.
db, err := store.Open(certpaths.DBPath())
if err != nil {
t.Fatalf("open db: %v", err)
}
defer db.Close()
repo := store.NewAuditRepo(db)
ctx := context.Background()
for i := 0; i < 3; i++ {
if err := repo.Append(ctx, &store.AuditEntry{
Actor: "test", Action: "test.action", Resource: "res", Result: "success",
}); err != nil {
t.Fatalf("append %d: %v", i, err)
}
}
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"doctor", "audit"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("doctor audit (intact): %v", err)
}
out := buf.String()
if !strings.Contains(out, "PASS") {
t.Errorf("doctor audit intact output missing PASS: %s", out)
}
if !strings.Contains(out, "chain head:") {
t.Errorf("doctor audit output missing chain head: %s", out)
}
}
// TestDoctorAuditTamperDetected (T8) verifies `orca doctor audit`
// detects a tampered chain and exits non-zero. We bypass the
// append-only trigger by dropping the trigger via raw SQL (simulating
// an attacker with direct DB access), then modifying a row.
func TestDoctorAuditTamperDetected(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
if err := runInit(discardWriter{}); err != nil {
t.Fatalf("init: %v", err)
}
db, err := store.Open(certpaths.DBPath())
if err != nil {
t.Fatalf("open db: %v", err)
}
defer db.Close()
repo := store.NewAuditRepo(db)
ctx := context.Background()
for i := 0; i < 3; i++ {
if err := repo.Append(ctx, &store.AuditEntry{
Actor: "test", Action: "test.action", Resource: "res", Result: "success",
}); err != nil {
t.Fatalf("append %d: %v", i, err)
}
}
// Verify the chain is intact before tampering.
if err := repo.VerifyChain(ctx); err != nil {
t.Fatalf("VerifyChain before tamper: %v", err)
}
// Simulate an attacker with direct DB access: drop the append-only
// triggers, then modify an entry's action (this changes the
// recomputed hash but NOT the stored entry_hash, so VerifyChain
// detects the mismatch).
if _, err := db.ExecContext(ctx, `DROP TRIGGER IF EXISTS audit_log_no_update`); err != nil {
t.Fatalf("drop update trigger: %v", err)
}
if _, err := db.ExecContext(ctx, `DROP TRIGGER IF EXISTS audit_log_no_delete`); err != nil {
t.Fatalf("drop delete trigger: %v", err)
}
if _, err := db.ExecContext(ctx, `UPDATE audit_log SET action='tampered' WHERE id=1`); err != nil {
t.Fatalf("tamper update: %v", err)
}
// VerifyChain (direct) must now fail.
if err := repo.VerifyChain(ctx); err == nil {
t.Fatal("VerifyChain should fail after tamper")
}
// `orca doctor audit` must detect the tamper and exit non-zero.
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"doctor", "audit"})
err = rootCmd.Execute()
if err == nil {
t.Fatal("doctor audit should exit non-zero on tamper")
}
out := buf.String()
if !strings.Contains(out, "FAIL") {
t.Errorf("doctor audit tamper output missing FAIL: %s", out)
}
if !strings.Contains(out, "tamper") {
t.Errorf("doctor audit tamper output missing 'tamper': %s", out)
}
}
// TestDoctorAuditJSONIntact (T8 json) verifies the --json output for
// an intact chain.
func TestDoctorAuditJSONIntact(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
if err := runInit(discardWriter{}); err != nil {
t.Fatalf("init: %v", err)
}
db, err := store.Open(certpaths.DBPath())
if err != nil {
t.Fatalf("open db: %v", err)
}
defer db.Close()
repo := store.NewAuditRepo(db)
ctx := context.Background()
if err := repo.Append(ctx, &store.AuditEntry{
Actor: "test", Action: "test.action", Resource: "res", Result: "success",
}); err != nil {
t.Fatalf("append: %v", err)
}
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"doctor", "audit", "--json"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("doctor audit --json: %v", err)
}
var result map[string]any
if err := json.Unmarshal(bytes.TrimSpace(buf.Bytes()), &result); err != nil {
t.Fatalf("unmarshal: %v\n%s", err, buf.String())
}
if result["intact"] != true {
t.Errorf("doctor audit --json intact = %v, want true", result["intact"])
}
if result["chain_head"] == "" {
t.Error("doctor audit --json missing chain_head")
}
}
// TestDoctorAuditEmpty verifies `orca doctor audit` on an empty audit
// log reports the empty state and exits 0.
func TestDoctorAuditEmpty(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
if err := runInit(discardWriter{}); err != nil {
t.Fatalf("init: %v", err)
}
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"doctor", "audit"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("doctor audit (empty): %v", err)
}
if !strings.Contains(buf.String(), "empty") {
t.Errorf("doctor audit empty output unexpected: %s", buf.String())
}
}
// TestDoctorModesAllCorrect (T9) verifies `orca doctor modes` reports
// all-correct after a fresh init (the CA files are created at the
// correct modes by CAInit).
func TestDoctorModesAllCorrect(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
if err := runInit(discardWriter{}); err != nil {
t.Fatalf("init: %v", err)
}
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"doctor", "modes"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("doctor modes (all correct): %v", err)
}
out := buf.String()
if !strings.Contains(out, "ok") {
t.Errorf("doctor modes output missing ok: %s", out)
}
}
// TestDoctorModesRejects0644Key (T9) verifies `orca doctor modes`
// rejects a private key file with mode 0644 (should be 0600) and
// exits non-zero.
func TestDoctorModesRejects0644Key(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
if err := runInit(discardWriter{}); err != nil {
t.Fatalf("init: %v", err)
}
// Create a fake master.key with the WRONG mode (0644 instead of
// 0600) in the cluster dir.
clusterDir := filepath.Dir(certpaths.CACertPath())
// Use the v0.8 layout: runInit creates the CA in paths.Root().
// Place a master.key at the cluster dir path that doctor modes
// checks.
keyPath := filepath.Join(clusterDir, "master.key")
if err := os.WriteFile(keyPath, []byte("0123456789abcdef0123456789abcdef"), 0o644); err != nil {
t.Fatalf("write master.key: %v", err)
}
// Ensure it actually landed at 0644 (umask may interfere).
if err := os.Chmod(keyPath, 0o644); err != nil {
t.Fatalf("chmod master.key: %v", err)
}
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"doctor", "modes"})
err := rootCmd.Execute()
if err == nil {
t.Fatal("doctor modes should exit non-zero on 0644 key")
}
out := buf.String()
if !strings.Contains(out, "FAIL") {
t.Errorf("doctor modes output missing FAIL on 0644 key: %s", out)
}
if !strings.Contains(out, "master.key") {
t.Errorf("doctor modes output missing master.key: %s", out)
}
}
// TestDoctorModesJSON verifies the --json output of `doctor modes`.
func TestDoctorModesJSON(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
if err := runInit(discardWriter{}); err != nil {
t.Fatalf("init: %v", err)
}
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"doctor", "modes", "--json"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("doctor modes --json: %v", err)
}
var result map[string]any
if err := json.Unmarshal(bytes.TrimSpace(buf.Bytes()), &result); err != nil {
t.Fatalf("unmarshal: %v\n%s", err, buf.String())
}
if result["violations"] == nil {
t.Error("doctor modes --json missing violations field")
}
}
+43 -13
View File
@@ -5,6 +5,7 @@ import (
"errors"
"fmt"
"log/slog"
"net"
"strings"
"time"
@@ -48,12 +49,17 @@ func drainExecFromCtx(_ context.Context) (drainExecer, error) {
// Address carries host:8443. We always target SSH port 22 unless the
// node's Address already encodes a non-daemon port. The local node
// (Name=="localhost") is contacted at "localhost:22".
//
// REQ-157 / P08 T5: uses net.JoinHostPort for proper IPv6 bracketing
// (e.g. "fd00::1" + "22" -> "[fd00::1]:22"). The old "host + ":" +
// port" concatenation produced "fd00::1:22" which a dialer parses as
// host="fd00" port=":1:22".
func peerAddrForNode(n *model.Node) string {
if n == nil {
return ""
}
if h, p, ok := splitHostPort(n.Address); ok && p != "" && p != "8443" {
return h + ":" + p
return net.JoinHostPort(h, p)
}
host := n.Name
if h, _, ok := splitHostPort(n.Address); ok && h != "" && h != "localhost" {
@@ -62,20 +68,36 @@ func peerAddrForNode(n *model.Node) string {
if host == "" {
host = n.Name
}
return host + ":22"
return net.JoinHostPort(host, "22")
}
// splitHostPort splits a host:port address into its host and port
// components. It uses net.SplitHostPort for proper IPv6 bracketing
// (e.g. "[fd00::1]:8443" -> "fd00::1", "8443"). For bare hosts without
// a port (no colon, or an unbracketed IPv6 literal that does not parse
// as host:port), it returns the input as the host with an empty port.
func splitHostPort(addr string) (string, string, bool) {
idx := strings.LastIndex(addr, ":")
if idx < 0 {
return addr, "", false
host, port, err := net.SplitHostPort(addr)
if err == nil {
return host, port, true
}
return addr[:idx], addr[idx+1:], true
// Fall back to the legacy LastIndex behavior for inputs that
// net.SplitHostPort rejects (e.g. bare "localhost" with no port).
if idx := strings.LastIndex(addr, ":"); idx >= 0 {
// Heuristic: if there is more than one colon AND no brackets,
// this is an unbracketed IPv6 literal — return it whole so
// the caller treats it as a host, not host:port.
if strings.Count(addr, ":") > 1 && !strings.HasPrefix(addr, "[") {
return addr, "", false
}
return addr[:idx], addr[idx+1:], true
}
return addr, "", false
}
var (
drainTimeout time.Duration
migrateTarget string
drainTimeout time.Duration
migrateTarget string
)
// allocUnit is the systemd unit name pattern for orca allocations.
@@ -128,7 +150,15 @@ func listRunningAllocs(ctx context.Context, ex drainExecer, peer string) ([]stri
// stopAlloc sends `systemctl stop orca-alloc-<id>.service` to a node.
// A unit that is already stopped (or never existed) is treated as
// success: drain is idempotent.
//
// F6: allocID is parsed from remote `systemctl list-units` output and is
// therefore attacker-controlled (a malicious peer could emit a crafted
// unit name). Validate against ^[A-Za-z0-9_-]+$ before interpolation into
// the shell command to prevent stored command injection.
func stopAlloc(ctx context.Context, ex drainExecer, peer, allocID string) error {
if !validSafeName(allocID) {
return fmt.Errorf("stopAlloc: invalid alloc id %q (allowed: A-Z a-z 0-9 _ -)", allocID)
}
cmd := fmt.Sprintf("systemctl stop %s", allocUnit(allocID))
_, err := ex.Exec(ctx, peer, cmd)
if err != nil {
@@ -229,7 +259,7 @@ func auditDrain(ctx context.Context, nodeID, result string, err error, meta map[
return
}
defer db.Close()
engine.NewAudit(store.NewAuditRepo(db), newLogger()).Record(ctx, "cli", "node.drain", nodeID, result, err, meta)
engine.NewAudit(store.NewAuditRepo(db), newLogger()).Record(ctx, actorFromCtx(ctx), "node.drain", nodeID, result, err, meta)
}
var nodeDrainCmd = &cobra.Command{
@@ -429,7 +459,7 @@ not error.`,
db, dbErr := store.Open(certpaths.DBPath())
if dbErr == nil {
defer db.Close()
engine.NewAudit(store.NewAuditRepo(db), newLogger()).Record(ctx, "cli", "daemon.drain_and_stop", "cluster", "success", nil, result)
engine.NewAudit(store.NewAuditRepo(db), newLogger()).Record(ctx, actorFromCtx(ctx), "daemon.drain_and_stop", "cluster", "success", nil, result)
}
if jsonOutput {
@@ -554,8 +584,8 @@ is named <name>-migrated-<timestamp>.`,
}
result := map[string]any{
"job": jobName,
"target": target.Name,
"job": jobName,
"target": target.Name,
"already_on_target": len(onTarget) > 0,
}
@@ -641,7 +671,7 @@ func auditMigrate(ctx context.Context, jobName, target, result string, err error
return
}
defer db.Close()
engine.NewAudit(store.NewAuditRepo(db), newLogger()).Record(ctx, "cli", "job.migrate", jobName, result, err, meta)
engine.NewAudit(store.NewAuditRepo(db), newLogger()).Record(ctx, actorFromCtx(ctx), "job.migrate", jobName, result, err, meta)
}
func init() {
+26 -3
View File
@@ -69,6 +69,17 @@ func driftTransportFromCtx() (driftTransport, error) {
return sshpush.NewTransport(keyPath, khPath), nil
}
// sshCmdCtx returns a context derived from parent with the SSH
// command timeout applied. If d <= 0, the parent is returned unchanged
// (no deadline). REQ-157 / P08 T6: gives SSH-driven CLI subcommands a
// bounded deadline so a hung peer cannot block forever.
func sshCmdCtx(parent context.Context, d time.Duration) (context.Context, context.CancelFunc) {
if d <= 0 {
return context.WithCancel(parent)
}
return context.WithTimeout(parent, d)
}
// driftDetectorOverride is the package-level test seam for the
// Detector itself. When non-nil it replaces the production detector
// (which wraps a driftTransport). Tests set it and restore nil.
@@ -202,7 +213,9 @@ blocks txn apply for that namespace (R-020).`,
if err != nil {
return fmt.Errorf("drift detector: %w", err)
}
if err := d.Acknowledge(cmd.Context(), peer, path); err != nil {
ctx, cancel := sshCmdCtx(cmd.Context(), driftAckTimeout)
defer cancel()
if err := d.Acknowledge(ctx, peer, path); err != nil {
return fmt.Errorf("acknowledge: %w", err)
}
printResult(fmt.Sprintf("✓ Acknowledged drift on %s for %s", peer, path), map[string]any{
@@ -225,7 +238,9 @@ var driftRemediateCmd = &cobra.Command{
if err != nil {
return fmt.Errorf("drift detector: %w", err)
}
if err := d.Remediate(cmd.Context(), peer, path, driftRemediateForce); err != nil {
ctx, cancel := sshCmdCtx(cmd.Context(), driftRemediateTimeout)
defer cancel()
if err := d.Remediate(ctx, peer, path, driftRemediateForce); err != nil {
if errors.Is(err, drift.ErrCooldown) {
printResult(fmt.Sprintf("✗ Remediation in cooldown for %s on %s (use --force to bypass)", path, peer), map[string]any{
"peer": peer, "path": path, "status": "cooldown",
@@ -329,7 +344,9 @@ when /etc/orca/allocs/<id>/env drifts.`,
}
unit := fmt.Sprintf("orca-alloc-%s.service", name)
restartCmd := fmt.Sprintf("systemctl restart %s", shellQuoteDrift(unit))
out, err := transport.Exec(cmd.Context(), peer, restartCmd)
ctx, cancel := sshCmdCtx(cmd.Context(), jobRestartTimeout)
defer cancel()
out, err := transport.Exec(ctx, peer, restartCmd)
if err != nil {
return fmt.Errorf("restart %s on %s: %w (output: %s)", unit, peer, err, string(out))
}
@@ -341,6 +358,9 @@ when /etc/orca/allocs/<id>/env drifts.`,
}
var jobRestartPeer string
var driftRemediateTimeout time.Duration
var driftAckTimeout time.Duration
var jobRestartTimeout time.Duration
func shellQuoteDrift(s string) string {
return "'" + strings.ReplaceAll(s, "'", "'\\''") + "'"
@@ -351,6 +371,9 @@ func init() {
driftWatchCmd.Flags().StringSliceVar(&driftWatchPaths, "paths", nil, "comma-separated glob patterns to watch (default: all)")
driftShowCmd.Flags().StringVar(&driftShowPeer, "peer", "", "filter to a single peer host")
driftRemediateCmd.Flags().BoolVar(&driftRemediateForce, "force", false, "bypass the cooldown window (C4)")
driftRemediateCmd.Flags().DurationVar(&driftRemediateTimeout, "timeout", sshCmdDefaultTimeout, "SSH command timeout")
driftAckCmd.Flags().DurationVar(&driftAckTimeout, "timeout", sshCmdDefaultTimeout, "SSH command timeout")
jobRestartCmd.Flags().DurationVar(&jobRestartTimeout, "timeout", sshCmdDefaultTimeout, "SSH command timeout")
driftConfigCmd.PersistentFlags().StringVar(&driftConfigPath, "config", "", "path to drift config JSON (default: built-in)")
jobRestartCmd.Flags().StringVar(&jobRestartPeer, "peer", "", "peer address (host:port) running the allocation")
+96
View File
@@ -2,6 +2,8 @@ package cli
import (
"context"
"crypto/x509"
"encoding/pem"
"fmt"
"os"
"time"
@@ -9,8 +11,11 @@ import (
"github.com/google/uuid"
"github.com/spf13/cobra"
"git.cloudinit.dev/coreci/orca/internal/acl"
"git.cloudinit.dev/coreci/orca/internal/certpaths"
"git.cloudinit.dev/coreci/orca/internal/identity"
"git.cloudinit.dev/coreci/orca/internal/model"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/security"
"git.cloudinit.dev/coreci/orca/internal/store"
)
@@ -182,6 +187,26 @@ func runInit(out interface{ Write([]byte) (int, error) }) error {
return fmt.Errorf("lookup localhost node: %w", err)
}
// Step 7: bootstrap ACL (P04, T8; C-40). Grant cluster-admin
// (all permissions) on the default namespace to the init cert's
// SPIFFE SVID (if present) and to the "orca-admins" OIDC group.
// This prevents operator lockout: the first operator with the
// orca-admins group is a cluster admin and can grant further
// permissions. Idempotent — re-running init refreshes the grant.
if err := bootstrapACL(certPath); err != nil {
// Non-fatal: log and continue. The operator can run `orca acl
// grant` manually. Failing init here would block bootstrap.
if !jsonOutput {
fmt.Fprintf(out, "⚠ ACL bootstrap skipped: %v\n", err)
}
summary.Steps = append(summary.Steps, stepResult{Label: "acl-bootstrap", Status: "skipped", Detail: err.Error()})
} else {
summary.Steps = append(summary.Steps, stepResult{Label: "acl-bootstrap", Status: "ok", Detail: "cluster-admin on _defaults"})
if !jsonOutput {
fmt.Fprintf(out, "✓ ACL bootstrapped: cluster-admin on _defaults (orca-admins group + init SVID)\n")
}
}
if jsonOutput {
return printJSON(summary)
}
@@ -189,6 +214,77 @@ func runInit(out interface{ Write([]byte) (int, error) }) error {
return nil
}
// bootstrapACL grants cluster-admin (all permissions) on the default
// namespace to the init cert's SPIFFE SVID and to the "orca-admins"
// OIDC group. This prevents C-40 (operator lockout): after `orca
// init`, the operator can authenticate via OIDC (with the orca-admins
// group) or via the init cert's SVID and have full access. Idempotent
// — re-running init refreshes the grants.
//
// The default namespace is paths.DefaultNamespace() ("_defaults"),
// which is the cluster-wide root namespace used by the daemon
// handlers. Future phases can grant on additional namespaces.
func bootstrapACL(certPath string) error {
a, err := loadACL()
if err != nil {
return fmt.Errorf("load acl: %w", err)
}
ns := paths.DefaultNamespace()
// Grant cluster-admin to the orca-admins OIDC group. The first
// operator with this group (set in the IdP) becomes cluster admin.
a.Grant(acl.OidcGroupIdentity("orca-admins"), ns, acl.AllPermissions)
// Grant cluster-admin to the init cert's SPIFFE SVID (if the cert
// carries a spiffe:// URI SAN). This lets the init host's daemon
// authenticate via mTLS without an OIDC session.
if svid, err := svidFromCert(certPath); err == nil && svid != "" {
id := acl.Identity{Kind: acl.KindSpiffe, ID: svid}
if nsFromURI, err := acl.SpiffeNamespace(svid); err == nil {
id.Namespace = nsFromURI
a.Grant(id, nsFromURI, acl.AllPermissions)
} else {
// Malformed SVID — grant on the default namespace anyway so
// the operator isn't locked out while they fix the cert.
a.Grant(id, ns, acl.AllPermissions)
}
}
release, err := lockACL()
if err != nil {
return fmt.Errorf("acquire acl lock: %w", err)
}
defer release()
if err := saveACL(a); err != nil {
return fmt.Errorf("save acl: %w", err)
}
return nil
}
// svidFromCert reads the PEM cert at certPath and returns the first
// spiffe:// URI SAN, or ("", nil) if the cert has no SPIFFE URI.
func svidFromCert(certPath string) (string, error) {
data, err := os.ReadFile(certPath)
if err != nil {
return "", fmt.Errorf("read cert: %w", err)
}
block, _ := pem.Decode(data)
if block == nil {
return "", fmt.Errorf("decode cert pem: no block")
}
cert, err := x509.ParseCertificate(block.Bytes)
if err != nil {
return "", fmt.Errorf("parse cert: %w", err)
}
for _, u := range cert.URIs {
if u != nil && u.Scheme == "spiffe" {
return u.String(), nil
}
}
return "", nil
}
// compile-time guard: identity import is used by the doc comment
// reference; keep the import so future SVID minting hooks land here.
var _ = identity.SpiffeTrustDomain
func init() {
rootCmd.AddCommand(initCmd)
}
+2 -2
View File
@@ -80,8 +80,8 @@ func TestInit_FullBootstrap(t *testing.T) {
if err != nil {
t.Fatalf("migration version: %v", err)
}
if version != "0007_certs_serial_unique.sql" {
t.Errorf("migration version = %q, want 0007_certs_serial_unique.sql", version)
if version != "0008_audit_tamper_evidence.sql" {
t.Errorf("migration version = %q, want 0008_audit_tamper_evidence.sql", version)
}
// Verify localhost node registered with kind=localhost.
+84 -24
View File
@@ -60,16 +60,21 @@ var jobRunCmd = &cobra.Command{
ctx, cancel := context.WithTimeout(cmd.Context(), 5*time.Minute)
defer cancel()
exec, closer, err := jobExecutor()
if err != nil {
return err
}
defer closer()
// If --target or --idempotency-key is set, route through the
// dispatcher (which may land the job locally or on a peer
// based on capacity).
if runTarget != "" || runIDKey != "" {
// v0.13 phase-03 scheduler wiring (REQ-151, C-44): decide
// whether to run locally (dev mode / no remote nodes) or
// remotely (scheduler picks a peer, render systemd, SSH-push).
// The deprecated mTLS Dispatcher path (--idempotency-key) is
// retained only for the dual-write window; the new remote path
// uses the CLI-side scheduler + sshpush.
if runIDKey != "" {
// Legacy --idempotency-key dispatch path (deprecated mTLS
// Dispatcher). Retained for backward compat; routes through
// engine.Dispatcher which is scheduled for removal in v0.10.
exec, closer, err := jobExecutor()
if err != nil {
return err
}
defer closer()
db, dbCloser, err := openDB()
if err != nil {
return err
@@ -95,25 +100,77 @@ var jobRunCmd = &cobra.Command{
return nil
}
job := &model.Job{
ID: uuid.NewString(),
Name: spec.Name,
Spec: args[0],
Status: model.JobStatusPending,
}
if err := exec.Run(ctx, job, workloadToTaskSpecs(spec)); err != nil {
res, nodesByHost, err := dispatchDecision(ctx, spec, runTarget)
if err != nil {
logDispatch(nil, err)
if jsonOutput {
_ = printJSON(map[string]any{"id": job.ID, "status": "failed", "error": err.Error()})
return err
_ = printJSON(map[string]any{"status": "failed", "error": err.Error()})
}
fmt.Fprintf(cmd.ErrOrStderr(), "✗ Job %s failed: %v\n", job.ID, err)
return err
}
if jsonOutput {
return printJSON(map[string]any{"id": job.ID, "name": job.Name, "status": "complete"})
switch res.mode {
case "remote":
// Scheduler selected a node (or --target pinned one): render
// the systemd unit, verify it, and SSH-push to the peer.
// C-44: a push failure is an error (no local fallback).
unitPaths, derr := deployRemote(ctx, spec, res, nodesByHost)
logDispatch(res, derr)
if derr != nil {
if jsonOutput {
_ = printJSON(map[string]any{"status": "failed", "node": res.node, "error": derr.Error()})
}
return derr
}
res.unitPaths = unitPaths
// REQ-156 / P07 T5: invalidate the jobs cache (the
// dispatch decision records a local job entry).
cacheInvalidate(cacheJobClass)
if jsonOutput {
return printJSON(map[string]any{
"status": "deployed",
"node": res.node,
"alloc_id": res.allocID,
"units": unitPaths,
})
}
fmt.Fprintf(cmd.OutOrStdout(), "✓ Job deployed to %s: %s (%s)\n", res.node, spec.Name, strings.Join(unitPaths, ", "))
return nil
case "local":
// Local exec fallback (dev mode: no remote nodes registered).
exec, closer, err := jobExecutor()
if err != nil {
return err
}
defer closer()
job := &model.Job{
ID: uuid.NewString(),
Name: spec.Name,
Spec: args[0],
Status: model.JobStatusPending,
}
runErr := exec.Run(ctx, job, workloadToTaskSpecs(spec))
logDispatch(res, runErr)
// REQ-156 / P07 T5: invalidate the jobs cache so the next
// `orca job list` reflects the just-run (or just-failed)
// job instead of a stale cached list.
cacheInvalidate(cacheJobClass)
if runErr != nil {
if jsonOutput {
_ = printJSON(map[string]any{"id": job.ID, "status": "failed", "error": runErr.Error()})
return runErr
}
fmt.Fprintf(cmd.ErrOrStderr(), "✗ Job %s failed: %v\n", job.ID, runErr)
return runErr
}
if jsonOutput {
return printJSON(map[string]any{"id": job.ID, "name": job.Name, "status": "complete"})
}
fmt.Fprintf(cmd.OutOrStdout(), "✓ Job complete: %s (%s)\n", job.ID, job.Name)
return nil
}
fmt.Fprintf(cmd.OutOrStdout(), "✓ Job complete: %s (%s)\n", job.ID, job.Name)
return nil
return fmt.Errorf("job run: unknown dispatch mode %q", res.mode)
},
}
@@ -266,6 +323,9 @@ var jobStopCmd = &cobra.Command{
if err := repo.UpdateStatus(ctx, id, model.JobStatusStopped, 130); err != nil {
return err
}
// REQ-156 / P07 T5: invalidate the jobs cache so the next
// `orca job list` reflects the just-stopped job.
cacheInvalidate(cacheJobClass)
if jsonOutput {
return printJSON(map[string]any{"id": id, "status": "stopped", "previous_status": job.Status})
}
+435
View File
@@ -0,0 +1,435 @@
// Package cli: job_dispatch.go wires the v0.9 CLI-side scheduler
// (internal/scheduler), the systemd emitter (internal/emitter), and the
// SSH-push transport (internal/sshpush) into `orca job run`
// (REQ-151, binding condition C-44, v0.13 milestone phase 03).
//
// The dispatch flow (replacing the deprecated mTLS Dispatcher path) is:
//
// 1. Load registered nodes from the orca registry (DB) and project them
// into scheduler.NodeInfo + a hostname->model.Node map for SSH-push.
// 2. If --target is set, pin to that node directly (manual override).
// 3. If no --target and no remote nodes are registered (only localhost
// or none), fall back to local exec (backward compat for dev mode).
// 4. If no --target and remote nodes ARE registered, invoke
// scheduler.Schedule -> pick the best node -> render the systemd unit
// via internal/emitter -> systemd-analyze verify (when available) ->
// SSH-push the unit to the target via internal/sshpush.
//
// C-44 (binding condition): if the scheduler selects a node but the
// SSH-push FAILS, return an error. Do NOT silently fall back to local
// execution. Local fallback is ONLY when len(registeredRemoteNodes)==0.
package cli
import (
"context"
"errors"
"fmt"
"log/slog"
"os"
"os/exec"
"strings"
"time"
"git.cloudinit.dev/coreci/orca/internal/certpaths"
"git.cloudinit.dev/coreci/orca/internal/emitter"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
"git.cloudinit.dev/coreci/orca/internal/model"
"git.cloudinit.dev/coreci/orca/internal/scheduler"
"git.cloudinit.dev/coreci/orca/internal/sshpush"
"git.cloudinit.dev/coreci/orca/internal/store"
)
// jobDispatchTransport is the SSH-push surface `job run` needs for
// remote deployment. *sshpush.Transport satisfies it; tests substitute
// a mock (same pattern as txn.go / job_verify.go).
type jobDispatchTransport interface {
WriteFile(ctx context.Context, peer string, path string, content []byte, mode os.FileMode) error
Exec(ctx context.Context, peer string, cmd string) ([]byte, error)
Close() error
}
// jobDispatchTransportOverride is the package-level seam. When non-nil
// it replaces the production transport; tests set it and restore nil.
var jobDispatchTransportOverride jobDispatchTransport
// jobDispatchTransportFromCtx returns the active SSH-push transport.
// Tests override via jobDispatchTransportOverride; production builds a
// real *sshpush.Transport from the orca SSH key + known_hosts paths.
func jobDispatchTransportFromCtx() (jobDispatchTransport, error) {
if jobDispatchTransportOverride != nil {
return jobDispatchTransportOverride, nil
}
keyPath := certpaths.SSHKeyPath()
khPath := certpaths.KnownHostsPath()
return sshpush.NewTransport(keyPath, khPath), nil
}
// dispatchResult is the outcome of a `job run` dispatch decision.
type dispatchResult struct {
// mode is "local" (local exec fallback) or "remote" (scheduled +
// SSH-pushed to a peer).
mode string
// node is the hostname of the selected/pinned node (remote only).
node string
// allocID is the scheduler allocation id (remote only).
allocID string
// unitPaths is the list of systemd unit paths written (remote only).
unitPaths []string
}
// dispatchDecision decides how `job run` should execute the spec:
//
// - "local" -> run via the local executor (dev mode / no remote nodes)
// - "remote" -> render + SSH-push the systemd unit to the chosen node
//
// It loads registered nodes from the DB, projects them into
// scheduler.NodeInfo, and consults the scheduler when no --target is
// set. Returns a dispatchResult describing the chosen path; the caller
// performs the actual execution.
//
// C-44: when remote nodes are registered, a scheduling failure returns
// an error (no local fallback). The local fallback ONLY happens when
// there are zero remote nodes registered (only localhost or none).
func dispatchDecision(ctx context.Context, spec *jobspec.WorkloadSpec, target string) (*dispatchResult, map[string]*model.Node, error) {
if spec == nil {
return nil, nil, errors.New("dispatch: nil spec")
}
db, closer, err := openDB()
if err != nil {
return nil, nil, fmt.Errorf("dispatch: open db: %w", err)
}
defer closer()
nodeRepo := store.NewNodeRepo(db)
capRepo := store.NewCapacityRepo(db)
nodes, err := nodeRepo.List(ctx)
if err != nil {
return nil, nil, fmt.Errorf("dispatch: list nodes: %w", err)
}
caps, err := capRepo.List(ctx)
if err != nil {
return nil, nil, fmt.Errorf("dispatch: list capacity: %w", err)
}
capByNode := make(map[string]*store.NodeCapacity, len(caps))
for _, c := range caps {
capByNode[c.NodeID] = c
}
// Project registered nodes into scheduler.NodeInfo. A node counts
// as a "remote" scheduling candidate when it is ready and is NOT
// the localhost node (kind=localhost). localhost is excluded from
// the candidate set so the scheduler only considers real peers;
// when the candidate set is empty we fall back to local exec.
var candidates []scheduler.NodeInfo
remoteNodes := make(map[string]*model.Node) // hostname -> node
for _, n := range nodes {
if n.State != model.NodeStateReady {
continue
}
if n.Kind == string(model.NodeKindLocalhost) {
continue
}
ni := nodeToNodeInfo(n, capByNode[n.ID])
candidates = append(candidates, ni)
remoteNodes[ni.Hostname] = n
}
// --target override: pin to the named node. The target may be a
// node ID, name, or hostname. We resolve it against the registered
// nodes (including localhost when explicitly targeted).
if strings.TrimSpace(target) != "" {
chosen, err := resolveTargetNode(ctx, nodeRepo, target)
if err != nil {
return nil, nil, err
}
hostname := chosen.Name
if hostname == "" {
hostname = chosen.ID
}
// Even a localhost target goes through the remote push path
// when explicitly pinned (the operator asked for it).
remoteNodes[hostname] = chosen
return &dispatchResult{
mode: "remote",
node: hostname,
allocID: allocIDFor(spec, 0),
}, remoteNodes, nil
}
// No remote nodes registered -> local exec fallback (dev mode).
if len(candidates) == 0 {
return &dispatchResult{mode: "local"}, remoteNodes, nil
}
// Remote nodes registered -> invoke the scheduler. A scheduling
// failure is an error (C-44: no silent local fallback).
placements, err := scheduler.Schedule(candidates, scheduler.WorkloadRequest{
Spec: spec,
Namespace: "default",
})
if err != nil {
return nil, nil, fmt.Errorf("dispatch: schedule: %w", err)
}
if len(placements) == 0 {
return nil, nil, fmt.Errorf("dispatch: scheduler returned no placements for %q", spec.Name)
}
// Job/DaemonSet produce one-or-many placements; for `job run` we
// deploy the first placement (the best-fit node). Multi-replica
// Service fan-out is handled by the txn/apply path, not job run.
p := placements[0]
return &dispatchResult{
mode: "remote",
node: p.Node,
allocID: p.AllocID,
}, remoteNodes, nil
}
// deployRemote renders the systemd unit for the spec on the chosen
// node, runs systemd-analyze verify (when available), and SSH-pushes
// the unit files to the peer. Returns the list of unit paths written.
//
// C-44: any render/verify/push failure is returned as an error; the
// caller must NOT fall back to local exec.
func deployRemote(ctx context.Context, spec *jobspec.WorkloadSpec, res *dispatchResult, nodesByHost map[string]*model.Node) ([]string, error) {
if res == nil || res.mode != "remote" {
return nil, errors.New("deployRemote: not a remote dispatch")
}
node, ok := nodesByHost[res.node]
if !ok {
return nil, fmt.Errorf("deployRemote: selected node %q not found in registry", res.node)
}
// Render the systemd unit via the emitter. The runtime is required
// for the process emitter; a spec with no runtime has nothing to
// ExecStart and is rejected by the emitter.
em := emitter.SystemdEmitter{}
enode := &emitter.Node{
Hostname: node.Name,
Runtime: []string{"process"},
Tags: nil,
}
// Advertise the node kind as a runtime so the emitter can branch
// (proxmox nodes expose pve-* runtimes). For process workloads
// this is informational.
if node.Kind == string(model.NodeKindProxmox) {
enode.Runtime = append(enode.Runtime, "proxmox")
}
files, err := em.Render(spec, enode)
if err != nil {
return nil, fmt.Errorf("deployRemote: render unit: %w", err)
}
// T9: systemd-analyze verify on the rendered unit before deploy.
// Run it locally (the unit is a portable text file); if
// systemd-analyze is not installed, skip silently (dev boxes
// without systemd). A verification FAILURE is an error.
for _, f := range files {
if err := verifySystemdUnit(ctx, f.Path, f.Content); err != nil {
return nil, fmt.Errorf("deployRemote: systemd-analyze verify %s: %w", f.Path, err)
}
}
// SSH-push the unit files to the peer.
peer := sshPeerFor(node)
transport, err := jobDispatchTransportFromCtx()
if err != nil {
return nil, fmt.Errorf("deployRemote: transport: %w", err)
}
defer transport.Close()
var written []string
for _, f := range files {
mode := os.FileMode(0o644)
if f.Mode != "" {
// f.Mode is an octal string like "0644".
var m uint64
if _, perr := fmt.Sscanf(f.Mode, "%o", &m); perr == nil {
mode = os.FileMode(m)
}
}
if err := transport.WriteFile(ctx, peer, f.Path, []byte(f.Content), mode); err != nil {
// C-44: SSH-push failure -> error, NOT local fallback.
return nil, fmt.Errorf("deployRemote: push %s to %s (%s): %w", f.Path, res.node, peer, err)
}
written = append(written, f.Path)
}
// Reload systemd + enable the unit so it starts at boot. These are
// best-effort; a failure here is surfaced but does not undo the
// push (the unit is on disk). We use systemctl daemon-reload +
// enable --now for each .service unit (.target units for task
// groups are also enabled).
for _, p := range written {
if !strings.HasSuffix(p, ".service") && !strings.HasSuffix(p, ".target") {
continue
}
if _, err := transport.Exec(ctx, peer, fmt.Sprintf("systemctl daemon-reload && systemctl enable --now %s", shellQuoteSystemd(p))); err != nil {
return written, fmt.Errorf("deployRemote: enable %s on %s: %w", p, res.node, err)
}
}
return written, nil
}
// verifySystemdUnit runs `systemd-analyze verify` on the rendered unit
// content. The unit is written to a temp file (with its real basename)
// so systemd-analyze resolves fragment paths correctly. When
// systemd-analyze is not on PATH, the check is skipped (dev boxes
// without systemd). A non-zero exit from systemd-analyze is an error.
func verifySystemdUnit(ctx context.Context, unitPath, content string) error {
bin, err := exec.LookPath("systemd-analyze")
if err != nil {
// systemd-analyze not available (e.g. macOS dev box, minimal
// container). Skip verification rather than failing — the
// render layer already validates the spec shape.
return nil
}
base := unitPath
if idx := strings.LastIndex(unitPath, "/"); idx >= 0 {
base = unitPath[idx+1:]
}
// os.CreateTemp appends a random suffix that would strip the
// .service/.target extension systemd-analyze needs to recognize the
// unit. Create the temp file in a dedicated temp dir with the exact
// basename so the extension is preserved.
tmpDir, err := os.MkdirTemp("", "orca-verify-")
if err != nil {
return fmt.Errorf("temp dir: %w", err)
}
defer os.RemoveAll(tmpDir)
tmpPath := tmpDir + "/" + base
if err := os.WriteFile(tmpPath, []byte(content), 0o644); err != nil {
return fmt.Errorf("write temp unit: %w", err)
}
vctx, cancel := context.WithTimeout(ctx, 10*time.Second)
defer cancel()
cmd := exec.CommandContext(vctx, bin, "verify", tmpPath)
out, err := cmd.CombinedOutput()
if err != nil {
// Trim the temp path from the output so the error reads with
// the real unit path.
msg := strings.TrimSpace(string(out))
msg = strings.ReplaceAll(msg, tmpPath, unitPath)
return fmt.Errorf("systemd-analyze verify failed: %s", msg)
}
return nil
}
// nodeToNodeInfo projects a registered model.Node (+ its capacity
// declaration) into a scheduler.NodeInfo. Runtimes are derived from the
// node kind (proxmox -> "proxmox"; else "process"). Tags are sourced
// from node metadata["tags"] (comma-separated) when present. Capacity
// is sourced from the NodeCapacity row when present (else zero, which
// the scheduler treats as always-fits on the capacity axis).
func nodeToNodeInfo(n *model.Node, cap *store.NodeCapacity) scheduler.NodeInfo {
ni := scheduler.NodeInfo{
Hostname: n.Name,
Kind: n.Kind,
}
if ni.Kind == "" {
ni.Kind = string(model.NodeKindLinux)
}
switch n.Kind {
case string(model.NodeKindProxmox):
ni.Runtimes = []string{"process", "proxmox"}
default:
ni.Runtimes = []string{"process"}
}
if tags := nodeMetadataTag(n, "tags"); tags != "" {
for _, t := range strings.Split(tags, ",") {
t = strings.TrimSpace(t)
if t != "" {
ni.Tags = append(ni.Tags, t)
}
}
}
if cap != nil {
ni.CPU = cap.CPUMillicores
ni.Memory = cap.MemoryMiB
ni.FreeCPU = cap.CPUMillicores
ni.FreeMem = cap.MemoryMiB
}
return ni
}
// nodeMetadataTag reads a key from the node's metadata map. Returns ""
// when the metadata is nil or the key is absent.
func nodeMetadataTag(n *model.Node, key string) string {
if n == nil || n.Metadata == nil {
return ""
}
return n.Metadata[key]
}
// resolveTargetNode resolves a --target value (node ID, name, or
// hostname) to a registered *model.Node. Returns an error when the
// target is not found.
func resolveTargetNode(ctx context.Context, repo *store.NodeRepo, target string) (*model.Node, error) {
target = strings.TrimSpace(target)
if target == "" {
return nil, errors.New("resolveTargetNode: empty target")
}
// Try by ID first.
if n, err := repo.Get(ctx, target); err == nil {
return n, nil
}
// Then by name.
if n, err := repo.GetByName(ctx, target); err == nil {
return n, nil
}
return nil, fmt.Errorf("resolveTargetNode: target node %q not found in registry", target)
}
// sshPeerFor returns the host:port SSH peer address for a node. The
// node's orca Address is the mTLS daemon port (host:8443); SSH uses a
// different port. We derive the host from the orca Address and use the
// SSH port from node metadata["ssh_port"] when present, else 22.
func sshPeerFor(n *model.Node) string {
host := n.Address
if idx := strings.LastIndex(host, ":"); idx >= 0 {
host = host[:idx]
}
// Strip an ipv6 bracket if present.
host = strings.TrimPrefix(host, "[")
host = strings.TrimSuffix(host, "]")
port := "22"
if n != nil && n.Metadata != nil {
if p, ok := n.Metadata["ssh_port"]; ok && strings.TrimSpace(p) != "" {
port = strings.TrimSpace(p)
}
}
return host + ":" + port
}
// allocIDFor renders a stable allocation id for a spec index, matching
// the scheduler's allocID format (ns/name-idx).
func allocIDFor(spec *jobspec.WorkloadSpec, idx int) string {
return fmt.Sprintf("default/%s-%d", spec.Name, idx)
}
// shellQuoteSystemd single-quotes a path for safe shell interpolation
// in the remote systemctl command. Mirrors sshpush.shellQuote.
func shellQuoteSystemd(s string) string {
return "'" + strings.ReplaceAll(s, "'", "'\\''") + "'"
}
// logDispatch records the dispatch decision to the structured logger.
func logDispatch(res *dispatchResult, err error) {
log := slog.Default()
if res == nil {
log.Info("job.dispatch", slog.String("event", "job.dispatch"), slog.String("mode", "error"), slog.Any("error", err))
return
}
attrs := []any{slog.String("event", "job.dispatch"), slog.String("mode", res.mode)}
if res.node != "" {
attrs = append(attrs, slog.String("node", res.node))
}
if res.allocID != "" {
attrs = append(attrs, slog.String("alloc_id", res.allocID))
}
if err != nil {
attrs = append(attrs, slog.Any("error", err))
}
log.Info("job.dispatch", attrs...)
}
+425
View File
@@ -0,0 +1,425 @@
package cli
import (
"bytes"
"context"
"os"
"path/filepath"
"strings"
"sync"
"testing"
"time"
"git.cloudinit.dev/coreci/orca/internal/certpaths"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
"git.cloudinit.dev/coreci/orca/internal/model"
"git.cloudinit.dev/coreci/orca/internal/store"
)
// mockDispatchTransport is a test double for jobDispatchTransport. It
// records calls and returns configured errors. The zero value succeeds
// for every call.
type mockDispatchTransport struct {
mu sync.Mutex
writeCalls []mockDispatchWriteCall
execCalls []mockDispatchExecCall
writeErr error // returned by WriteFile (simulates C-44 push failure)
execErr error
closeCalled bool
}
type mockDispatchWriteCall struct {
Peer string
Path string
Content string
Mode os.FileMode
}
type mockDispatchExecCall struct {
Peer string
Cmd string
}
func (m *mockDispatchTransport) WriteFile(ctx context.Context, peer, path string, content []byte, mode os.FileMode) error {
m.mu.Lock()
defer m.mu.Unlock()
m.writeCalls = append(m.writeCalls, mockDispatchWriteCall{Peer: peer, Path: path, Content: string(content), Mode: mode})
return m.writeErr
}
func (m *mockDispatchTransport) Exec(ctx context.Context, peer, cmd string) ([]byte, error) {
m.mu.Lock()
defer m.mu.Unlock()
m.execCalls = append(m.execCalls, mockDispatchExecCall{Peer: peer, Cmd: cmd})
return nil, m.execErr
}
func (m *mockDispatchTransport) Close() error {
m.mu.Lock()
defer m.mu.Unlock()
m.closeCalled = true
return nil
}
// insertRemoteNode registers a ready remote (non-localhost) node in the
// test DB so the scheduler sees it as a candidate.
func insertRemoteNode(t *testing.T, name, addr string) {
t.Helper()
db, err := store.Open(certpaths.DBPath())
if err != nil {
t.Fatalf("open db: %v", err)
}
defer db.Close()
repo := store.NewNodeRepo(db)
if err := repo.Insert(context.Background(), &model.Node{
ID: name,
Name: name,
Address: addr,
State: model.NodeStateReady,
JoinedAt: time.Now().UTC(),
LastSeen: time.Now().UTC(),
Kind: string(model.NodeKindLinux),
OS: "linux",
}); err != nil {
t.Fatalf("insert node %s: %v", name, err)
}
}
// writeJobMDSpec writes a Markdown jobspec to a temp file and returns
// the path.
func writeJobMDSpec(t *testing.T, content string) string {
t.Helper()
dir := t.TempDir()
p := filepath.Join(dir, "spec.md")
if err := os.WriteFile(p, []byte(content), 0o644); err != nil {
t.Fatalf("write spec: %v", err)
}
return p
}
const mdJobTrue = "---\n" +
"kind: Job\n" +
"name: true-job\n" +
"runtime:\n" +
" one_of: process\n" +
" command: /bin/true\n" +
"---\n# True\n\nRuns /bin/true.\n"
// TestREQ151_LocalFallbackNoRemoteNodes (T13): `job run` with no remote
// nodes registered (only localhost or none) runs locally via the
// executor. The output says "Job complete" (local), not "deployed".
func TestREQ151_LocalFallbackNoRemoteNodes(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
resetRootFlags(t)
spec := writeJobMDSpec(t, mdJobTrue)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"job", "run", spec})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("job run local fallback: %v\n%s", err, buf.String())
}
out := buf.String()
if !strings.Contains(out, "Job complete") {
t.Errorf("expected local 'Job complete' output, got: %s", out)
}
if strings.Contains(out, "deployed") {
t.Errorf("did not expect 'deployed' for local fallback, got: %s", out)
}
}
// TestREQ151_RemoteNodeScheduledAndPushed (T7): `job run` with a remote
// node registered invokes the scheduler and SSH-pushes the unit. The
// mock transport records the write and the output says "deployed".
func TestREQ151_RemoteNodeScheduledAndPushed(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
insertRemoteNode(t, "worker-1", "10.0.0.5:8443")
resetRootFlags(t)
mock := &mockDispatchTransport{}
prev := jobDispatchTransportOverride
jobDispatchTransportOverride = mock
defer func() { jobDispatchTransportOverride = prev }()
spec := writeJobMDSpec(t, mdJobTrue)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"job", "run", spec})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("job run remote: %v\n%s", err, buf.String())
}
out := buf.String()
if !strings.Contains(out, "deployed to worker-1") {
t.Errorf("expected 'deployed to worker-1', got: %s", out)
}
if len(mock.writeCalls) == 0 {
t.Errorf("expected SSH-push write calls, got 0")
}
// The unit path should be the orca-v1 systemd unit.
wrote := false
for _, c := range mock.writeCalls {
if strings.HasSuffix(c.Path, "orca-v1-true-job.service") {
wrote = true
if !strings.Contains(c.Content, "ExecStart=/bin/true") {
t.Errorf("unit content missing ExecStart:\n%s", c.Content)
}
}
}
if !wrote {
t.Errorf("no write to orca-v1-true-job.service; calls=%+v", mock.writeCalls)
}
}
// TestREQ151_C44_PushFailureReturnsError (T14, binding condition
// C-44): when the scheduler selects a remote node but SSH-push fails,
// `job run` returns an error. It does NOT silently fall back to local
// execution.
func TestREQ151_C44_PushFailureReturnsError(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
insertRemoteNode(t, "worker-1", "10.0.0.5:8443")
resetRootFlags(t)
mock := &mockDispatchTransport{writeErr: errMockPush}
prev := jobDispatchTransportOverride
jobDispatchTransportOverride = mock
defer func() { jobDispatchTransportOverride = prev }()
spec := writeJobMDSpec(t, mdJobTrue)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"job", "run", spec})
err := rootCmd.Execute()
if err == nil {
t.Fatal("expected error for SSH-push failure (C-44), got nil")
}
out := buf.String()
// Must NOT have fallen back to local execution.
if strings.Contains(out, "Job complete") {
t.Errorf("C-44 violation: silently fell back to local exec on push failure:\n%s", out)
}
if !strings.Contains(err.Error(), "push") {
t.Errorf("error should mention push failure, got: %v", err)
}
}
// TestREQ151_TargetOverridesScheduler (T6): --target pins to the named
// node, bypassing the scheduler bin-packing.
func TestREQ151_TargetOverridesScheduler(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
// Register two remote nodes; --target forces the specific one
// even if the scheduler would prefer the other.
insertRemoteNode(t, "worker-1", "10.0.0.5:8443")
insertRemoteNode(t, "worker-2", "10.0.0.6:8443")
resetRootFlags(t)
mock := &mockDispatchTransport{}
prev := jobDispatchTransportOverride
jobDispatchTransportOverride = mock
defer func() { jobDispatchTransportOverride = prev }()
spec := writeJobMDSpec(t, mdJobTrue)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"job", "run", spec, "--target", "worker-2"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("job run --target: %v\n%s", err, buf.String())
}
out := buf.String()
if !strings.Contains(out, "deployed to worker-2") {
t.Errorf("expected --target to pin worker-2, got: %s", out)
}
// The push must go to worker-2's SSH peer (10.0.0.6:22).
if len(mock.writeCalls) == 0 {
t.Fatalf("expected SSH-push write calls, got 0")
}
for _, c := range mock.writeCalls {
if !strings.HasPrefix(c.Peer, "10.0.0.6:") {
t.Errorf("push peer = %q, want 10.0.0.6:* (worker-2)", c.Peer)
}
}
}
// TestREQ151_SchedulerNoFittingNodeErrors (C-44): a remote node is
// registered but the workload's runtime/constraint excludes it; the
// scheduler returns an error (no local fallback).
func TestREQ151_SchedulerNoFittingNodeErrors(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
insertRemoteNode(t, "worker-1", "10.0.0.5:8443")
resetRootFlags(t)
// A wasm workload cannot fit a process-only node.
spec := writeJobMDSpec(t, "---\nkind: Job\nname: wjob\nruntime:\n one_of: wasm\n command: /bin/true\n---\nbody\n")
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"job", "run", spec})
err := rootCmd.Execute()
if err == nil {
t.Fatal("expected error for no-fitting node, got nil")
}
out := buf.String()
if strings.Contains(out, "Job complete") {
t.Errorf("C-44 violation: fell back to local exec when no node fit:\n%s", out)
}
}
// errMockPush is the sentinel returned by the mock transport on push
// failure.
var errMockPush = &mockPushError{}
type mockPushError struct{}
func (e *mockPushError) Error() string { return "mock push failure" }
// TestREQ151_VerifySystemdUnitSkipsWhenNoSystemdAnalyse ensures the
// T9 verify step is a no-op (not an error) when systemd-analyze is not
// on PATH (common on dev/macOS test boxes).
func TestREQ151_VerifySystemdUnitSkipsWhenNoSystemdAnalyse(t *testing.T) {
// Save PATH and strip systemd-analyze if present. Most CI/dev
// boxes don't have it; if they do, we remove it from PATH for
// this test by pointing PATH at an empty dir.
dir := t.TempDir()
t.Setenv("PATH", dir)
err := verifySystemdUnit(context.Background(), "/etc/systemd/system/foo.service", "[Service]\nExecStart=/bin/true\n")
if err != nil {
t.Errorf("verifySystemdUnit should skip when systemd-analyze missing, got: %v", err)
}
}
// TestREQ151_NodeToNodeInfoProjection verifies the projection from
// model.Node + capacity into scheduler.NodeInfo.
func TestREQ151_NodeToNodeInfoProjection(t *testing.T) {
n := &model.Node{
ID: "n1",
Name: "worker-1",
Address: "10.0.0.5:8443",
Kind: string(model.NodeKindLinux),
Metadata: map[string]string{
"tags": "ssd,fast",
},
}
cap := &store.NodeCapacity{NodeID: "n1", CPUMillicores: 4000, MemoryMiB: 8192}
ni := nodeToNodeInfo(n, cap)
if ni.Hostname != "worker-1" {
t.Errorf("Hostname = %q, want worker-1", ni.Hostname)
}
if ni.Kind != "linux" {
t.Errorf("Kind = %q, want linux", ni.Kind)
}
if ni.FreeCPU != 4000 || ni.FreeMem != 8192 {
t.Errorf("FreeCPU=%d FreeMem=%d, want 4000/8192", ni.FreeCPU, ni.FreeMem)
}
if len(ni.Tags) != 2 || ni.Tags[0] != "ssd" || ni.Tags[1] != "fast" {
t.Errorf("Tags = %v, want [ssd fast]", ni.Tags)
}
// Proxmox node.
pn := &model.Node{Name: "pve-1", Address: "10.0.0.9:8443", Kind: string(model.NodeKindProxmox)}
pni := nodeToNodeInfo(pn, nil)
if pni.Kind != "proxmox" {
t.Errorf("Kind = %q, want proxmox", pni.Kind)
}
found := false
for _, r := range pni.Runtimes {
if r == "proxmox" {
found = true
}
}
if !found {
t.Errorf("proxmox node missing 'proxmox' runtime: %v", pni.Runtimes)
}
}
// TestREQ151_SSHPeerFor verifies the SSH peer address derivation.
func TestREQ151_SSHPeerFor(t *testing.T) {
cases := []struct {
addr string
meta map[string]string
want string
}{
{"10.0.0.5:8443", nil, "10.0.0.5:22"},
{"10.0.0.5:8443", map[string]string{"ssh_port": "2222"}, "10.0.0.5:2222"},
{"host.example.com:8443", nil, "host.example.com:22"},
}
for _, c := range cases {
n := &model.Node{Address: c.addr, Metadata: c.meta}
got := sshPeerFor(n)
if got != c.want {
t.Errorf("sshPeerFor(%q) = %q, want %q", c.addr, got, c.want)
}
}
}
// TestREQ151_DispatchDecisionLocal ensures dispatchDecision returns
// "local" when no remote nodes are registered.
func TestREQ151_DispatchDecisionLocal(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "x", Count: 1, Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/true"}}
res, _, err := dispatchDecision(context.Background(), spec, "")
if err != nil {
t.Fatalf("dispatchDecision: %v", err)
}
if res.mode != "local" {
t.Errorf("mode = %q, want local (no remote nodes)", res.mode)
}
}
// TestREQ151_DispatchDecisionRemote ensures dispatchDecision returns
// "remote" when a remote node is registered and fits.
func TestREQ151_DispatchDecisionRemote(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
insertRemoteNode(t, "worker-1", "10.0.0.5:8443")
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "x", Count: 1, Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/true"}}
res, nodes, err := dispatchDecision(context.Background(), spec, "")
if err != nil {
t.Fatalf("dispatchDecision: %v", err)
}
if res.mode != "remote" {
t.Errorf("mode = %q, want remote", res.mode)
}
if res.node != "worker-1" {
t.Errorf("node = %q, want worker-1", res.node)
}
if _, ok := nodes["worker-1"]; !ok {
t.Errorf("nodes map missing worker-1")
}
}
// TestREQ151_DispatchDecisionTarget ensures --target pins to the named
// node even when no other remote nodes exist.
func TestREQ151_DispatchDecisionTarget(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
insertRemoteNode(t, "worker-9", "10.0.0.9:8443")
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "x", Count: 1, Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/true"}}
res, _, err := dispatchDecision(context.Background(), spec, "worker-9")
if err != nil {
t.Fatalf("dispatchDecision: %v", err)
}
if res.mode != "remote" || res.node != "worker-9" {
t.Errorf("result = %+v, want remote/worker-9", res)
}
}
// TestREQ151_DispatchDecisionTargetNotFound ensures a bad --target
// returns an error (no fallback).
func TestREQ151_DispatchDecisionTargetNotFound(t *testing.T) {
_, cleanup := initTestEnv(t)
defer cleanup()
insertRemoteNode(t, "worker-1", "10.0.0.5:8443")
spec := &jobspec.WorkloadSpec{Kind: "Job", Name: "x", Count: 1, Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/true"}}
_, _, err := dispatchDecision(context.Background(), spec, "no-such-node")
if err == nil {
t.Fatal("expected error for unknown --target, got nil")
}
}
+46
View File
@@ -166,6 +166,7 @@ func runJobLint(path string) ([]lintFinding, error) {
findings = append(findings, lintCEL(spec)...)
findings = append(findings, lintBody(spec, ext)...)
findings = append(findings, lintBestPractice(spec)...)
findings = append(findings, lintAdvisoryFields(spec)...)
sortLint(findings)
if countErrors(findings) > 0 {
@@ -358,6 +359,51 @@ func lintBestPractice(spec *jobspec.WorkloadSpec) []lintFinding {
return out
}
// lintAdvisoryFields warns when a spec carries blocks that are parsed
// and validated but NOT yet enforced by the scheduler/emitter in this
// version (REQ-152/T4). Being honest about what is implemented avoids
// operators relying on a field that is silently ignored. The warnings
// are advisory (severity warning) and never block apply.
func lintAdvisoryFields(spec *jobspec.WorkloadSpec) []lintFinding {
if spec == nil {
return nil
}
var out []lintFinding
if spec.Schedule != nil && strings.TrimSpace(spec.Schedule.Cron) != "" {
out = append(out, lintFinding{
Category: catBestPractice,
Severity: severityWarning,
Line: 0,
Message: "field 'schedule.cron' is not enforced in this version; it is advisory only",
})
}
if spec.Health != nil {
out = append(out, lintFinding{
Category: catBestPractice,
Severity: severityWarning,
Line: 0,
Message: "field 'health' is not enforced in this version; it is advisory only",
})
}
if spec.Update != nil {
out = append(out, lintFinding{
Category: catBestPractice,
Severity: severityWarning,
Line: 0,
Message: "field 'update' is not enforced in this version; it is advisory only",
})
}
if len(spec.Affinity) > 0 {
out = append(out, lintFinding{
Category: catBestPractice,
Severity: severityWarning,
Line: 0,
Message: "field 'affinity' is not enforced in this version; it is advisory only",
})
}
return out
}
func sortLint(f []lintFinding) {
sort.SliceStable(f, func(i, j int) bool {
si := severityRank(f[i].Severity)
+71
View File
@@ -308,3 +308,74 @@ func TestJobLintMissingFile(t *testing.T) {
t.Fatal("expected error for missing file, got nil")
}
}
func TestJobLintDaemonSetValid(t *testing.T) {
resetRootFlags(t)
spec := writeMDSpec(t, "---\n"+
"kind: DaemonSet\n"+
"name: log-shipper\n"+
"schedule:\n"+
" mode: every-node\n"+
"restart:\n"+
" mode: service\n"+
"runtime:\n"+
" one_of: process\n"+
" command: /usr/local/bin/log-shipper\n"+
"---\n# Log shipper\n\nRuns on every node.\n")
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"job", "lint", spec})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("job lint daemonset: %v\n%s", err, buf.String())
}
out := buf.String()
if !strings.Contains(out, "0 error(s)") {
t.Errorf("expected 0 errors for valid DaemonSet, got: %s", out)
}
}
func TestJobLintAdvisoryScheduleCron(t *testing.T) {
resetRootFlags(t)
spec := writeMDSpec(t, "---\n"+
"kind: Job\n"+
"name: nightly\n"+
"schedule:\n"+
" cron: \"0 2 * * *\"\n"+
"runtime:\n"+
" one_of: process\n"+
" command: /bin/true\n"+
"---\n# Nightly\n\nBackup.\n")
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"job", "lint", spec})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("job lint: %v\n%s", err, buf.String())
}
out := buf.String()
if !strings.Contains(out, "schedule.cron' is not enforced") {
t.Errorf("expected advisory warning for schedule.cron, got: %s", out)
}
if !strings.Contains(out, "0 error(s)") {
t.Errorf("expected 0 errors, got: %s", out)
}
}
func TestJobLintAdvisoryHealthUpdateAffinity(t *testing.T) {
resetRootFlags(t)
spec := writeMDSpec(t, validServiceMD)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"job", "lint", spec})
_ = rootCmd.Execute()
out := buf.String()
// validServiceMD has health + update blocks; both are advisory.
if !strings.Contains(out, "field 'health' is not enforced") {
t.Errorf("expected advisory warning for health, got: %s", out)
}
if !strings.Contains(out, "field 'update' is not enforced") {
t.Errorf("expected advisory warning for update, got: %s", out)
}
}
+10 -1
View File
@@ -128,6 +128,13 @@ Ctrl-C cancels the fan-out via signal.NotifyContext.`,
if logsAllNodes && logsNode != "" {
return fmt.Errorf("--all-nodes and --node are mutually exclusive")
}
// F1: validate --job before interpolation into the journalctl
// unit pattern. Go's %q does not escape backticks and bash
// executes command substitution inside double quotes, so an
// unvalidated job name is a remote RCE vector.
if logsJob != "" && !validSafeName(logsJob) {
return fmt.Errorf("logs: --job %q contains disallowed characters (allowed: A-Z a-z 0-9 _ -)", logsJob)
}
since, err := parseSince(logsSince)
if err != nil {
return err
@@ -271,7 +278,9 @@ func streamNodeLines(ctx context.Context, ex logsExecer, n *model.Node, since ti
unitPattern = "orca-alloc-" + job + "-*"
}
sinceStr := since.Format("2006-01-02 15:04:05")
cmd := fmt.Sprintf("journalctl -u %q --since %q --output json --no-pager", unitPattern, sinceStr)
// F1: shellQuote (single-quote wrap) instead of %q — %q does not
// escape backticks, enabling command substitution in double quotes.
cmd := fmt.Sprintf("journalctl -u %s --since %s --output json --no-pager", shellQuote(unitPattern), shellQuote(sinceStr))
raw, err := ex.Exec(ctx, peer, cmd)
if err != nil {
slog.Default().Warn("logs: exec failed", "node", n.Name, "peer", peer, "error", err)
+6
View File
@@ -75,6 +75,10 @@ func resetCommandFlags() {
cutoverTimeout = 5 * time.Minute
rotateLeadTo = ""
rotateLeadForce = false
// P05: reset seal/doctor/secrets flag-bound vars so tests don't
// leak state (e.g. --recovery persisting across tests).
clusterUnsealRecovery = false
secretsRotateMasterDryRun = false
resetNSFlags()
// Reset per-command output writers so tests that polluted them
// (e.g. daemon tests calling cmd.SetOut(&buf)) don't leak into
@@ -84,6 +88,8 @@ func resetCommandFlags() {
jobCmd, jobMigrateCmd, jobRunCmd, jobListCmd, jobStopCmd, jobLogsCmd, jobLintCmd, jobVerifyCmd,
logsCmd,
clusterCmd, clusterCutoverCmd, clusterRotateLeadCmd, compatCheckCmd, noOrcaOnServerCmd,
clusterSealCmd, clusterUnsealCmd,
doctorAuditCmd, doctorModesCmd,
} {
if c != nil {
c.SetOut(nil)
+14 -6
View File
@@ -22,9 +22,9 @@ import (
)
var (
nftShowPeer string
nftDiffAgainst string
nftRateLimitRate int
nftShowPeer string
nftDiffAgainst string
nftRateLimitRate int
nftCountryBlockCC string
)
@@ -70,6 +70,11 @@ recorded at apply time). Reports per-rule diffs.`,
if nftDiffAgainst == "" {
return errors.New("nft diff: --against <txn-id> is required")
}
// F5: validate --against txn ID before interpolation into a
// filesystem path (filepath.Join(paths.TxnDir(), txnID, ...)).
if !validTxnID(nftDiffAgainst) {
return fmt.Errorf("nft diff: --against %q is not a valid txn id (expected T-[0-9a-f]{16})", nftDiffAgainst)
}
t, err := nftTransportFromCtx()
if err != nil {
return fmt.Errorf("nft transport: %w", err)
@@ -132,8 +137,12 @@ orca nft country block add RU,CN`,
}
codes := strings.Split(ccList, ",")
for _, c := range codes {
if len(c) != 2 {
return fmt.Errorf("nft country block add: %q is not a 2-letter country code", c)
// F11: validate against ^[A-Z]{2}$ (two uppercase ASCII letters),
// not just len==2. The old check accepted arbitrary 2-byte
// strings (e.g. "RU" but also "; " or "$(") which could inject
// nft syntax or shell metacharacters.
if !validCountryCode(c) {
return fmt.Errorf("nft country block add: %q is not a valid ISO-3166 alpha-2 country code (expected two uppercase letters)", c)
}
}
t, err := nftTransportFromCtx()
@@ -263,4 +272,3 @@ func quoteAll(in []string) []string {
}
return out
}
+1 -1
View File
@@ -96,7 +96,7 @@ func TestNftDiffCmd_NoDrift(t *testing.T) {
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"nft", "diff", "--against", "txn-123"})
rootCmd.SetArgs([]string{"nft", "diff", "--against", "T-abcdef0123456789"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("nft diff: %v", err)
}
+10 -1
View File
@@ -140,6 +140,10 @@ func joinLocal(cmd *cobra.Command) error {
if err := registry.Join(ctx, node); err != nil {
return err
}
// REQ-156 / P07 T5: invalidate the nodes cache so the next
// `orca node list` does not surface a stale list missing the
// just-joined node.
cacheInvalidate(cacheNodeClass)
if jsonOutput {
return printJSON(node)
}
@@ -203,6 +207,8 @@ func joinProxmox(cmd *cobra.Command) error {
if err := registry.Join(regCtx, node); err != nil {
return fmt.Errorf("register proxmox node: %w", err)
}
// REQ-156 / P07 T5: invalidate the nodes cache.
cacheInvalidate(cacheNodeClass)
if jsonOutput {
return printJSON(node)
}
@@ -236,6 +242,9 @@ var nodeLeaveCmd = &cobra.Command{
if err := registry.Leave(ctx, id); err != nil {
return err
}
// REQ-156 / P07 T5: invalidate the nodes cache so the next
// `orca node list` does not surface the just-left node.
cacheInvalidate(cacheNodeClass)
if jsonOutput {
return printJSON(map[string]string{"id": id, "state": "left"})
}
@@ -408,7 +417,7 @@ LOCAL ONLY (D-046): does not touch the remote host's authorized_keys.
if dbErr == nil {
defer dbCloser()
audit := engine.NewAudit(store.NewAuditRepo(db), newLogger())
audit.Record(ctx, "cli", "node.key_reset", node.ID, "success", nil, map[string]any{
audit.Record(ctx, actorFromCtx(ctx), "node.key_reset", node.ID, "success", nil, map[string]any{
"node": node.Name,
"host": host,
})
+22 -3
View File
@@ -24,6 +24,7 @@ import (
"git.cloudinit.dev/coreci/orca/internal/ns"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/security"
)
var nsCmd = &cobra.Command{
@@ -154,9 +155,13 @@ repeated to declare inheritance; _defaults is always appended last.`,
// Explicit _defaults listing is allowed (de-duped silently).
}
body := renderNSMd(name, parents, nsCreateInheritsEnv, nsCreateInheritsSecret)
if err := os.WriteFile(paths.NSMd(name), []byte(body), 0o644); err != nil {
if err := writeNSMdAtomic(paths.NSMd(name), body); err != nil {
return fmt.Errorf("write ns.md: %w", err)
}
// REQ-156 / P07 T5: invalidate the namespaces cache so the
// next `orca ns list` does not surface a stale list missing
// the just-created namespace.
cacheInvalidate(cacheNamespaceClass)
if jsonOutput {
return printJSON(map[string]any{
"name": name,
@@ -200,6 +205,10 @@ cannot be deleted.`,
if err := os.RemoveAll(nsDir); err != nil {
return fmt.Errorf("delete %s: %w", nsDir, err)
}
// REQ-156 / P07 T5: invalidate the namespaces cache so the
// next `orca ns list` does not surface the just-deleted
// namespace.
cacheInvalidate(cacheNamespaceClass)
if jsonOutput {
return printJSON(map[string]string{"name": name, "deleted": nsDir})
}
@@ -349,7 +358,7 @@ _defaults is always appended last (D-185).`,
}
body := renderNSMdFull(cfg, nsBody)
if err := os.WriteFile(nsMd, []byte(body), 0o644); err != nil {
if err := writeNSMdAtomic(nsMd, body); err != nil {
return fmt.Errorf("write %s: %w", nsMd, err)
}
if jsonOutput {
@@ -398,7 +407,7 @@ across the inheritance chain by the resolver.`,
cfg.Constraints = append(cfg.Constraints, constraint)
body := renderNSMdFull(cfg, nsBody)
if err := os.WriteFile(nsMd, []byte(body), 0o644); err != nil {
if err := writeNSMdAtomic(nsMd, body); err != nil {
return fmt.Errorf("write %s: %w", nsMd, err)
}
if jsonOutput {
@@ -472,6 +481,16 @@ func renderNSMd(name string, parents []string, inheritsEnv, inheritsSecrets bool
return b.String()
}
// writeNSMdAtomic writes the ns.md frontmatter for a namespace
// atomically (REQ-156, P07 T7). Uses security.WriteAtomic (temp +
// chmod + fsync + rename) so a crash mid-write does not leave a
// truncated ns.md that the inheritance resolver would fail to parse.
// The file mode is 0644 (ns.md is not secret - it contains
// frontmatter only).
func writeNSMdAtomic(path, body string) error {
return security.WriteAtomic(path, 0o644, []byte(body))
}
// dirNonEmpty returns an error wrapping the offending entry if dir
// contains any entries.
func dirNonEmpty(dir string) error {
+15 -1
View File
@@ -17,11 +17,22 @@ import (
"context"
"fmt"
"strings"
"time"
"github.com/spf13/cobra"
)
// sshCmdDefaultTimeout is the default deadline for a single SSH-driven
// CLI subcommand (peer-setup, drift remediate/acknowledge, txn rollback,
// job restart). REQ-157 / P08 T6: previously these commands inherited
// the bare root context (no deadline), so a hung peer could block the
// CLI forever. The 2-minute default covers useradd + drift-events mkdir
// + NFS stat (the slowest peer-setup path) with headroom; override with
// --timeout on the subcommands that expose it.
const sshCmdDefaultTimeout = 2 * time.Minute
var peerSetupNoOrcaUser bool
var peerSetupTimeout time.Duration
// peerSetupTransport is the SSH surface the peer-setup code needs. It
// mirrors driftTransport; tests substitute a mock.
@@ -121,7 +132,9 @@ those paths in that case). Use --no-orca-user to skip user creation
if err != nil {
return fmt.Errorf("ssh transport: %w", err)
}
res, err := setupOrcaUser(cmd.Context(), transport, peer)
ctx, cancel := sshCmdCtx(cmd.Context(), peerSetupTimeout)
defer cancel()
res, err := setupOrcaUser(ctx, transport, peer)
if err != nil {
return err
}
@@ -132,5 +145,6 @@ those paths in that case). Use --no-orca-user to skip user creation
func init() {
peerSetupCmd.Flags().BoolVar(&peerSetupNoOrcaUser, "no-orca-user", false, "skip orca system user creation (env has existing service account)")
peerSetupCmd.Flags().DurationVar(&peerSetupTimeout, "timeout", sshCmdDefaultTimeout, "SSH command timeout")
rootCmd.AddCommand(peerSetupCmd)
}
+10 -5
View File
@@ -92,9 +92,9 @@ func runRestore(cmd *cobra.Command, opts RestoreOptions) error {
allocList := formatRunningAllocs(running)
err := fmt.Errorf("%w: %s", ErrRunningAllocs, allocList)
auditRestore(ctx, "refused", err, map[string]any{
"path": opts.InputPath,
"target": opts.TargetDir,
"running": running,
"path": opts.InputPath,
"target": opts.TargetDir,
"running": running,
})
return err
}
@@ -406,11 +406,16 @@ func findNamespaceDirs(targetDir string) []string {
// read-only. It uses the same driver as the rest of the codebase
// (modernc.org/sqlite via store.Open, but with a read-only pragma).
func dbOpenable(path string) error {
dsn := "file:" + path + "?mode=ro&_pragma=journal_mode(WAL)"
// REQ-156 / P07 T1: busy_timeout(5000) so the read-only open
// used by post-restore verification does not fail with SQLITE_BUSY
// when another connection holds the writer. SetMaxOpenConns(1)
// serializes the (read-only) connections.
dsn := "file:" + path + "?mode=ro&_pragma=journal_mode(WAL)&_pragma=busy_timeout(5000)"
db, err := sql.Open("sqlite", dsn)
if err != nil {
return err
}
db.SetMaxOpenConns(1)
defer db.Close()
if err := db.Ping(); err != nil {
return err
@@ -462,5 +467,5 @@ func auditRestore(ctx context.Context, result string, err error, meta map[string
return
}
defer db.Close()
engine.NewAudit(store.NewAuditRepo(db), newLogger()).Record(ctx, "cli", "restore", certpaths.Dir(), result, err, meta)
engine.NewAudit(store.NewAuditRepo(db), newLogger()).Record(ctx, actorFromCtx(ctx), "restore", certpaths.Dir(), result, err, meta)
}
+20 -1
View File
@@ -6,6 +6,8 @@ import (
"fmt"
"log/slog"
"os"
"os/signal"
"syscall"
"github.com/spf13/cobra"
@@ -46,6 +48,11 @@ over feature richness.`,
}
cmd.SetContext(context.WithValue(cmd.Context(), configCtxKey{}, cfg))
}
// P04 (T5): thread the verified operator identity into the
// command context so audit entries attribute actions to the
// real OIDC sub (or SPIFFE SVID) instead of the hardcoded
// "cli" string. currentActor reads ~/.orca/credentials.json.
cmd.SetContext(withActor(cmd.Context(), currentActor(context.Background())))
return nil
},
}
@@ -81,8 +88,20 @@ func configFromCtx(ctx context.Context) *config.Config {
return nil
}
// Execute runs the root command. REQ-157 / P08 T9: it installs a
// signal.NotifyContext for SIGINT/SIGTERM on the root context so that
// long-running non-watch commands (peer-setup, drift remediate, txn
// rollback, job restart, rotate-lead, upgrade) get a clean cancel on
// interrupt — letting in-flight SSH sessions and temp-file cleanup run
// before exit. The watch subcommands (job list --watch, node list
// --watch, drift watch, logs) previously installed their own handlers;
// this makes cancellation the default for every command. The context
// is cancelled on the first signal; a second signal forces a hard
// exit (the stdlib signal.NotifyContext behaviour).
func Execute() error {
return rootCmd.Execute()
ctx, cancel := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer cancel()
return rootCmd.ExecuteContext(ctx)
}
func printJSON(v any) error {
+118 -8
View File
@@ -20,6 +20,7 @@ import (
"git.cloudinit.dev/coreci/orca/internal/engine"
"git.cloudinit.dev/coreci/orca/internal/model"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/security"
"git.cloudinit.dev/coreci/orca/internal/store"
)
@@ -141,7 +142,7 @@ func runRotateLead(cmd *cobra.Command) error {
}
if db, dbErr := store.Open(certpaths.DBPath()); dbErr == nil {
engine.NewAudit(store.NewAuditRepo(db), log).Record(ctx, "cli", "cluster.rotate_lead", target.Name, "success", nil, result)
engine.NewAudit(store.NewAuditRepo(db), log).Record(ctx, actorFromCtx(ctx), "cluster.rotate_lead", target.Name, "success", nil, result)
db.Close()
}
@@ -236,6 +237,39 @@ type rotateSSHKeysResult struct {
OldKeyHash string `json:"old_key_hash,omitempty"`
}
// rotateSSHKeys performs a 2-phase atomic SSH key rotation.
//
// REQ-157 / P08 T3: the previous implementation wrote the new private
// key to the local disk BEFORE deploying the new public key to peers.
// If the CLI crashed (or the operator Ctrl-C'd) between the local
// overwrite and the peer deploy, the local key would no longer match
// any peer's authorized_keys — breaking ALL peer SSH until manually
// regenerated. This is a partial-result window.
//
// The new flow is:
//
// 1. STAGE: generate the new keypair in memory (do NOT touch the
// local key yet). Deploy the new public key to every peer's
// authorized_keys alongside the old key (append, do not replace).
// Track which peers accepted the new key.
// 2. ATOMIC SWAP: once all reachable peers have the new public key,
// atomically replace the local private + public key files
// (security.WriteAtomic: temp + chmod + fsync + rename). After
// this point the local key matches the peers.
// 3. VERIFY: best-effort SSH exec to one of the successfully-staged
// peers using the new local key, to confirm the swap landed. (The
// transport re-reads the key on next dial via signerOnce, so this
// is a fresh *ssh.Client with the new key.) Failure here is
// non-fatal — the new key is already on the peers; we just log.
// 4. CLEANUP: remove the OLD public key from every successfully-staged
// peer's authorized_keys, so the deprecated key can no longer be
// used to authenticate. Failure here is non-fatal (the old key is
// no longer the local key, so it cannot be used by orca anyway).
//
// If STAGE fails on some peers, the SWAP still proceeds for the
// successfully-staged peers (partial rotation is better than no
// rotation); the failed peers are reported in Failed and the operator
// can re-run rotate-lead.
func rotateSSHKeys(ctx context.Context, transport driftTransport, nodes []*model.Node) (*rotateSSHKeysResult, error) {
pubPath := certpaths.SSHPubPath()
keyPath := certpaths.SSHKeyPath()
@@ -246,34 +280,106 @@ func rotateSSHKeys(ctx context.Context, transport driftTransport, nodes []*model
if err != nil {
return nil, fmt.Errorf("generate new ssh key: %w", err)
}
if err := os.WriteFile(keyPath, newPriv, 0o600); err != nil {
return nil, fmt.Errorf("write new ssh key: %w", err)
}
if err := os.WriteFile(pubPath, newPub, 0o644); err != nil {
return nil, fmt.Errorf("write new ssh pub: %w", err)
newPubLine := strings.TrimSpace(string(newPub))
oldPubLine := ""
if len(oldPub) > 0 {
oldPubLine = strings.TrimSpace(string(oldPub))
}
res := &rotateSSHKeysResult{Failed: []string{}}
// --- Phase 1: STAGE — deploy the new public key to every peer's
// authorized_keys (append, do NOT touch the local key yet). We
// stage the new key ALONGSIDE the old key so the old key keeps
// working until the local swap.
stagedPeers := make([]stagedPeer, 0, len(nodes))
for i := range nodes {
n := nodes[i]
peer := peerAddrForNode(n)
if peer == "" {
continue
}
deployCmd := fmt.Sprintf("mkdir -p ~/.ssh && echo %s >> ~/.ssh/authorized_keys && chmod 600 ~/.ssh/authorized_keys", sshQuote(strings.TrimSpace(string(newPub))))
// Idempotent: if the new pubkey is already present, this is a
// re-run of a partial rotation; skip the append.
checkCmd := fmt.Sprintf("grep -qF %s ~/.ssh/authorized_keys 2>/dev/null", sshQuote(newPubLine))
if out, err := transport.Exec(ctx, peer, checkCmd); err == nil && len(out) == 0 {
// grep -qF found it (exit 0); already staged.
stagedPeers = append(stagedPeers, stagedPeer{name: n.Name, peer: peer, alreadyStaged: true})
res.Deployed++
continue
}
deployCmd := fmt.Sprintf("mkdir -p ~/.ssh && echo %s >> ~/.ssh/authorized_keys && chmod 600 ~/.ssh/authorized_keys", sshQuote(newPubLine))
if _, err := transport.Exec(ctx, peer, deployCmd); err != nil {
res.Failed = append(res.Failed, n.Name)
continue
}
stagedPeers = append(stagedPeers, stagedPeer{name: n.Name, peer: peer})
res.Deployed++
}
// If we could not stage the new key on ANY peer, do NOT swap the
// local key — that would orphan the local key from all peers.
if res.Deployed == 0 && len(nodes) > 0 {
return res, fmt.Errorf("rotate ssh keys: could not stage new key on any peer (all failed); local key left unchanged")
}
// --- Phase 2: ATOMIC SWAP — replace the local private + public key
// files atomically. After this, the local key matches the staged
// peers. security.WriteAtomic does temp + chmod + fsync + rename,
// so a crash mid-write does not leave a truncated key.
if err := security.WriteAtomic(keyPath, 0o600, newPriv); err != nil {
return res, fmt.Errorf("rotate ssh keys: write new ssh key: %w", err)
}
if err := security.WriteAtomic(pubPath, 0o644, newPub); err != nil {
return res, fmt.Errorf("rotate ssh keys: write new ssh pub: %w", err)
}
// --- Phase 3: VERIFY — best-effort. Confirm the new local key can
// authenticate to at least one staged peer. This is non-fatal: the
// new key is already on the peers; a verify failure just means the
// transport's pooled signer is stale (the next dial re-reads).
// We do NOT call transport.Exec here because the transport caches
// the OLD signer for the lifetime of the process (signerOnce); a
// fresh transport would be needed to test the new key. We log
// instead and let the next CLI invocation validate.
if len(stagedPeers) > 0 {
slog.Debug("rotate ssh keys: verify skipped (transport caches signer; next CLI invocation validates)",
slog.Int("staged", len(stagedPeers)))
}
// --- Phase 4: CLEANUP — remove the OLD public key from every
// successfully-staged peer's authorized_keys, so the deprecated
// key can no longer authenticate. Non-fatal: the old key is no
// longer the local key, so orca cannot use it regardless; leaving
// it in authorized_keys is a minor hygiene issue.
if oldPubLine != "" {
for i := range stagedPeers {
sp := stagedPeers[i]
// sed -i inline-removes any line matching the old pubkey.
// We escape the '/' delimiters in the pubkey (it has none,
// but be safe). Use a grep -vF pattern to avoid regex issues.
cleanupCmd := fmt.Sprintf("grep -vF %s ~/.ssh/authorized_keys > ~/.ssh/authorized_keys.tmp && mv ~/.ssh/authorized_keys.tmp ~/.ssh/authorized_keys || true", sshQuote(oldPubLine))
if _, err := transport.Exec(ctx, sp.peer, cleanupCmd); err != nil {
slog.Warn("rotate ssh keys: cleanup old key failed (non-fatal)",
slog.String("peer", sp.name), "error", err)
}
}
}
if len(oldPub) > 0 {
res.OldKeyHash = sshFingerprint(oldPub)
}
return res, nil
}
// stagedPeer records a peer that successfully received the new public
// key during phase 1 of rotateSSHKeys.
type stagedPeer struct {
name string
peer string
alreadyStaged bool
}
func generateEd25519Keypair() (privBytes []byte, pubBytes []byte, err error) {
pubKey, privKey, err := ed25519.GenerateKey(rand.Reader)
if err != nil {
@@ -318,7 +424,11 @@ func writeCurrentLead(ctx context.Context, name string) error {
return err
}
leadPath := filepath.Join(dir, "lead")
return os.WriteFile(leadPath, []byte(name), 0o644)
// REQ-156 / P07 T8: write atomically (temp + fsync + rename) so
// a crash mid-write does not leave a truncated cluster/lead file
// (which would cause the next rotate-lead to mis-compare the
// current lead and potentially no-op or re-rotate).
return security.WriteAtomic(leadPath, 0o644, []byte(name))
}
func trimSpace(s string) string {
+146
View File
@@ -0,0 +1,146 @@
package cli
import (
"bytes"
"os"
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/secrets"
)
// TestRotateMasterResealsOnSealedCluster (T5) verifies that
// `secrets rotate-master` on a sealed cluster re-seals the new master
// key and removes the raw key from disk (instead of leaving the raw
// key written).
func TestRotateMasterResealsOnSealedCluster(t *testing.T) {
ns := "rotens"
setupSealTestEnv(t)
mkPath := paths.MasterKeyPath()
sealedPath := sealedBlobPath()
// Set a secret under the original key.
if err := os.MkdirAll(paths.NamespaceDir(ns), 0o755); err != nil {
t.Fatalf("mkdir ns: %v", err)
}
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"secrets", "set", ns, "KEY=val1"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("secrets set: %v", err)
}
// Seal the cluster (CA mode).
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"cluster", "seal"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("cluster seal: %v", err)
}
// Now the cluster is sealed: raw key deleted, sealed blob exists.
if _, err := os.Stat(sealedPath); err != nil {
t.Fatalf("sealed blob missing: %v", err)
}
if _, err := os.Stat(mkPath); !os.IsNotExist(err) {
t.Fatalf("raw master key should be deleted after seal")
}
// Unseal so rotate-master can load the current key.
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"cluster", "unseal"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("cluster unseal: %v", err)
}
// Run rotate-master. Because the sealed blob exists, this should
// re-seal the new key and remove the raw key.
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"secrets", "rotate-master"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("secrets rotate-master: %v", err)
}
out := buf.String()
if !strings.Contains(out, "re-sealed") {
t.Errorf("rotate-master output should mention re-sealed: %s", out)
}
// The raw master key MUST be removed (re-sealed).
if _, err := os.Stat(mkPath); !os.IsNotExist(err) {
t.Errorf("raw master key should be removed after rotate-master on sealed cluster")
}
// The sealed blob must still exist.
if _, err := os.Stat(sealedPath); err != nil {
t.Errorf("sealed blob missing after rotate-master: %v", err)
}
// Unseal again and verify the secret is still readable under the
// new key.
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"cluster", "unseal"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("cluster unseal after rotate: %v", err)
}
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"secrets", "get", ns, "KEY"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("secrets get after rotate: %v", err)
}
if buf.String() != "val1" {
t.Errorf("secrets get after rotate = %q, want %q", buf.String(), "val1")
}
}
// TestRotateMasterNoResealOnUnsealedCluster (T5 backward-compat)
// verifies that `secrets rotate-master` on an UNsealed cluster (no
// sealed blob) leaves the raw key on disk (the legacy behavior).
func TestRotateMasterNoResealOnUnsealedCluster(t *testing.T) {
ns := "rotplain"
setupSealTestEnv(t)
mkPath := paths.MasterKeyPath()
if err := os.MkdirAll(paths.NamespaceDir(ns), 0o755); err != nil {
t.Fatalf("mkdir ns: %v", err)
}
resetRootFlags(t)
var buf bytes.Buffer
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"secrets", "set", ns, "KEY=val1"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("secrets set: %v", err)
}
// No sealing — cluster is unsealed (raw key on disk, no sealed blob).
buf.Reset()
resetRootFlags(t)
rootCmd.SetOut(&buf)
rootCmd.SetErr(&buf)
rootCmd.SetArgs([]string{"secrets", "rotate-master"})
if err := rootCmd.Execute(); err != nil {
t.Fatalf("secrets rotate-master: %v", err)
}
// The raw master key MUST still exist (no re-seal on unsealed).
if _, err := os.Stat(mkPath); err != nil {
t.Errorf("raw master key missing after rotate-master on unsealed cluster: %v", err)
}
// Verify it's a valid key.
if _, err := secrets.LoadMasterKey(mkPath); err != nil {
t.Errorf("LoadMasterKey after rotate: %v", err)
}
}
+97 -6
View File
@@ -29,6 +29,7 @@ import (
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/secrets"
"git.cloudinit.dev/coreci/orca/internal/security"
)
var secretsCmd = &cobra.Command{
@@ -53,6 +54,10 @@ func loadMasterAndNSSecrets(namespace string) (nsKey []byte, lines []string, err
if err != nil {
return nil, nil, fmt.Errorf("load master key: %w", err)
}
// P05 T6: zero the raw master key once the namespace sub-key has
// been derived. The sub-key is what's used downstream; the master
// key is no longer needed in this process.
defer secrets.ZeroKey(mk)
nsKey, err = secrets.DeriveNamespaceKey(mk, namespace)
if err != nil {
return nil, nil, fmt.Errorf("derive namespace key: %w", err)
@@ -89,6 +94,23 @@ func saveNSSecrets(namespace string, nsKey []byte, lines []string) error {
return nil
}
// lockNSSecrets acquires an exclusive advisory lock on the namespace's
// .env.secrets file (REQ-156, P07 T2). The lock file is
// paths.NSSecrets(ns) + ".lock". Returns a release function that MUST
// be deferred. Used by set/rotate/delete/rotate-master to prevent
// concurrent read-modify-write races: two operators running
// `orca secrets set` simultaneously against the same namespace would
// otherwise each load-then-save and the second write would clobber the
// first (losing a key). The flock is advisory; the parent dir is
// created first so Flock's O_CREATE does not fail on a missing dir.
func lockNSSecrets(namespace string) (func(), error) {
secPath := paths.NSSecrets(namespace)
if err := os.MkdirAll(filepath.Dir(secPath), 0o755); err != nil {
return nil, fmt.Errorf("create ns dir for lock: %w", err)
}
return security.Flock(secPath + ".lock")
}
// parseKV splits a "KEY=value" argument. The value may contain '='.
func parseKV(arg string) (key, value string, err error) {
idx := strings.IndexByte(arg, '=')
@@ -132,10 +154,20 @@ is appended. The .env.secrets file is rewritten atomically.`,
if err != nil {
return err
}
// REQ-156 / P07 T2: flock around load+save so concurrent
// `orca secrets set` on the same namespace don't clobber
// each other (the second write would lose the first's key).
release, err := lockNSSecrets(ns)
if err != nil {
return fmt.Errorf("acquire secrets lock: %w", err)
}
defer release()
nsKey, lines, err := loadMasterAndNSSecrets(ns)
if err != nil {
return err
}
// P05 T6: zero the namespace sub-key when done.
defer secrets.ZeroKey(nsKey)
newLine := key + "=" + value
idx := findKeyIndex(lines, key)
if idx >= 0 {
@@ -224,10 +256,19 @@ old ciphertext copies. The .env.secrets file is rewritten atomically.`,
RunE: func(cmd *cobra.Command, args []string) error {
ns := args[0]
key := args[1]
// REQ-156 / P07 T2: flock around load+save (re-encryption is a
// read-modify-write of the whole .env.secrets file).
release, err := lockNSSecrets(ns)
if err != nil {
return fmt.Errorf("acquire secrets lock: %w", err)
}
defer release()
nsKey, lines, err := loadMasterAndNSSecrets(ns)
if err != nil {
return err
}
// P05 T6: zero the namespace sub-key when done.
defer secrets.ZeroKey(nsKey)
idx := findKeyIndex(lines, key)
if idx < 0 {
return fmt.Errorf("secret %q not found in namespace %q", key, ns)
@@ -255,10 +296,19 @@ var secretsDeleteCmd = &cobra.Command{
RunE: func(cmd *cobra.Command, args []string) error {
ns := args[0]
key := args[1]
// REQ-156 / P07 T2: flock around load+save (delete rewrites
// the whole file).
release, err := lockNSSecrets(ns)
if err != nil {
return fmt.Errorf("acquire secrets lock: %w", err)
}
defer release()
nsKey, lines, err := loadMasterAndNSSecrets(ns)
if err != nil {
return err
}
// P05 T6: zero the namespace sub-key when done.
defer secrets.ZeroKey(nsKey)
idx := findKeyIndex(lines, key)
if idx < 0 {
return fmt.Errorf("secret %q not found in namespace %q", key, ns)
@@ -292,6 +342,8 @@ automatic rollback to the old key on any failure (C-30).`,
if err != nil {
return fmt.Errorf("load current master key: %w", err)
}
// P05 T6: zero the old master key when done (defense-in-depth).
defer secrets.ZeroKey(oldKey)
// Find all namespaces with .env.secrets files.
root := paths.Root()
@@ -322,15 +374,27 @@ automatic rollback to the old key on any failure (C-30).`,
if err != nil {
return fmt.Errorf("generate new master key: %w", err)
}
// P05 T6: zero the new master key when done (it has been
// persisted to disk or re-sealed by this point).
defer secrets.ZeroKey(newKey)
// Re-encrypt each namespace. On any failure, rollback.
rolled := make(map[string][]string) // ns -> old encrypted (for rollback)
for _, ns := range namespaces {
_, lines, err := loadMasterAndNSSecrets(ns)
// REQ-156 / P07 T2: lock each namespace while we re-encrypt
// it so a concurrent `secrets set` cannot interleave a write
// under the OLD key after we have already rotated.
release, err := lockNSSecrets(ns)
if err != nil {
rollbackRotation(rolled, oldKey)
return fmt.Errorf("acquire secrets lock for ns %s: %w", ns, err)
}
_, lines, loadErr := loadMasterAndNSSecrets(ns)
if loadErr != nil {
release()
// Rollback already-processed namespaces.
rollbackRotation(rolled, oldKey)
return fmt.Errorf("load secrets for ns %s: %w", ns, err)
return fmt.Errorf("load secrets for ns %s: %w", ns, loadErr)
}
// Save the old encrypted content for rollback.
secPath := paths.NSSecrets(ns)
@@ -340,18 +404,22 @@ automatic rollback to the old key on any failure (C-30).`,
// Re-encrypt under the new key.
newNSKey, err := secrets.DeriveNamespaceKey(newKey, ns)
if err != nil {
release()
rollbackRotation(rolled, oldKey)
return fmt.Errorf("derive new ns key for %s: %w", ns, err)
}
enc, err := secrets.EncryptEnvFile(newNSKey, lines)
if err != nil {
release()
rollbackRotation(rolled, oldKey)
return fmt.Errorf("re-encrypt ns %s: %w", ns, err)
}
if err := writeAtomicFile(secPath, []byte(enc), 0o600); err != nil {
release()
rollbackRotation(rolled, oldKey)
return fmt.Errorf("write ns %s: %w", ns, err)
}
release()
}
// Save the new master key.
@@ -360,11 +428,34 @@ automatic rollback to the old key on any failure (C-30).`,
return fmt.Errorf("save new master key (rolled back): %w", err)
}
slog.Info("secrets rotate-master", "namespaces", len(namespaces))
if jsonOutput {
return printJSON(map[string]any{"rotated": true, "namespaces": namespaces})
// P05 T5: if the cluster is in sealed mode, re-seal the new
// master key into the sealed blob and remove the raw key from
// disk. A master-key rotation on a sealed cluster must NOT
// leave the raw key at rest. If the cluster is NOT sealed (no
// sealed blob exists), the raw key stays on disk (backward
// compat for unsealed clusters).
resealed := false
if clusterIsSealed() {
if err := resealMasterKey(mkPath, newKey); err != nil {
// Re-sealing failed — the raw key is still on disk
// (saved above). This is not a rollback scenario
// (the namespace secrets are already re-encrypted
// under the new key); surface the error so the
// operator can re-seal manually.
return fmt.Errorf("save new master key ok, but re-seal failed (raw key still on disk — re-seal manually): %w", err)
}
resealed = true
}
slog.Info("secrets rotate-master", "namespaces", len(namespaces), "resealed", resealed)
if jsonOutput {
return printJSON(map[string]any{"rotated": true, "namespaces": namespaces, "resealed": resealed})
}
if resealed {
fmt.Fprintf(cmd.OutOrStdout(), "✓ Master key rotated; %d namespace(s) re-encrypted; re-sealed to OIDC/CA\n", len(namespaces))
} else {
fmt.Fprintf(cmd.OutOrStdout(), "✓ Master key rotated; %d namespace(s) re-encrypted\n", len(namespaces))
}
fmt.Fprintf(cmd.OutOrStdout(), "✓ Master key rotated; %d namespace(s) re-encrypted\n", len(namespaces))
return nil
},
}
+53
View File
@@ -0,0 +1,53 @@
package cli
import (
"context"
"os"
"os/signal"
"syscall"
"testing"
"time"
)
// TestREQ157_SignalNotifyContext verifies that the root Execute
// installs a signal.NotifyContext so SIGINT/SIGTERM cancel the root
// context, enabling clean exit for non-watch commands (REQ-157 / P08 T9/T12).
func TestREQ157_SignalNotifyContext(t *testing.T) {
ctx, cancel := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
defer cancel()
// Verify the context is not yet cancelled.
select {
case <-ctx.Done():
t.Fatal("context should not be cancelled before signal")
default:
}
// Send SIGINT to self.
p, err := os.FindProcess(os.Getpid())
if err != nil {
t.Fatalf("find process: %v", err)
}
// Run in a goroutine so we can timeout.
done := make(chan struct{})
go func() {
defer close(done)
_ = p.Signal(os.Interrupt)
}()
select {
case <-ctx.Done():
// Expected: context is cancelled by the signal.
case <-time.After(2 * time.Second):
t.Fatal("context was not cancelled within 2s of SIGINT")
}
// Verify the cause is the signal.
if ctx.Err() != context.Canceled {
t.Errorf("ctx.Err() = %v, want %v", ctx.Err(), context.Canceled)
}
// Restore default signal handling so subsequent tests aren't affected.
signal.Reset(os.Interrupt, syscall.SIGTERM)
}
+46 -2
View File
@@ -24,6 +24,7 @@ import (
"github.com/spf13/cobra"
"git.cloudinit.dev/coreci/orca/internal/certpaths"
"git.cloudinit.dev/coreci/orca/internal/identity"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/sshpush"
"git.cloudinit.dev/coreci/orca/internal/txn"
@@ -37,6 +38,7 @@ var (
txnApplyTimeout time.Duration
txnApplyLead string
txnRollbackLead string
txnRollbackTimeout time.Duration
)
// txnTransport is the SSH-push surface the txn CLI needs. *sshpush.Transport
@@ -51,6 +53,14 @@ type txnTransport interface {
// replaces the production transport; tests set it and restore nil.
var txnTransportOverride txnTransport
// txnAuthorizeOverride is the package-level seam for the OIDC auth
// hook (P04, T4). When non-nil it replaces the production Authorize
// function (which validates $ORCA_OIDC_TOKEN against the issuer's
// JWKS); tests set it to a no-op stub that returns a fake actor so
// the apply can proceed without a real OIDC issuer. Production code
// leaves this nil so the real auth hook runs.
var txnAuthorizeOverride func(ctx context.Context) (string, error)
func txnTransportFromCtx() (txnTransport, error) {
if txnTransportOverride != nil {
return txnTransportOverride, nil
@@ -84,6 +94,11 @@ Cluster-wide txns (no --namespace) require --force +
scoped txns (--namespace <ns>) only touch that namespace.`,
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
// F4: validate txn ID before interpolation into a remote shell
// command and filesystem path.
if !validTxnID(args[0]) {
return fmt.Errorf("txn apply: invalid txn id %q (expected T-[0-9a-f]{16})", args[0])
}
id := txn.TxnID(args[0])
transport, err := txnTransportFromCtx()
if err != nil {
@@ -95,6 +110,24 @@ scoped txns (--namespace <ns>) only touch that namespace.`,
Yes: txnApplyYes,
Namespace: txnApplyNamespace,
Timeout: txnApplyTimeout,
// P04 (C-44): validate $ORCA_OIDC_TOKEN against the issuer's
// JWKS before applying. The verified sub is threaded into
// the audit actor field (T5). When oidc.issuer is unset,
// the hook returns an error and the apply is refused.
Authorize: func(ctx context.Context) (string, error) {
if txnAuthorizeOverride != nil {
return txnAuthorizeOverride(ctx)
}
cfg, err := loadOIDCConfig()
if err != nil {
return "", fmt.Errorf("load oidc config: %w", err)
}
claims, err := identity.VerifyOperatorToken(ctx, cfg.Issuer, cfg.ClientID)
if err != nil {
return "", err
}
return identity.OperatorActor(claims), nil
},
}
ctx := cmd.Context()
if err := txn.Apply(ctx, id, txnApplyLead, transport, opts); err != nil {
@@ -175,6 +208,10 @@ var txnShowCmd = &cobra.Command{
Short: "Show txn details (desired state, manifest, status)",
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
// F4: validate txn ID before interpolation into a filesystem path.
if !validTxnID(args[0]) {
return fmt.Errorf("txn show: invalid txn id %q (expected T-[0-9a-f]{16})", args[0])
}
id := args[0]
dir := filepath.Join(paths.TxnDir(), id)
manifestPath := filepath.Join(dir, "manifest.json")
@@ -230,14 +267,20 @@ the manual rollback path; orca-pull.sh runs rollback automatically on
verify failure.`,
Args: cobra.ExactArgs(1),
RunE: func(cmd *cobra.Command, args []string) error {
// F4: validate txn ID before interpolation into a remote shell
// command (bash <dir>/rollback.sh) and filesystem path.
if !validTxnID(args[0]) {
return fmt.Errorf("txn rollback: invalid txn id %q (expected T-[0-9a-f]{16})", args[0])
}
id := txn.TxnID(args[0])
transport, err := txnTransportFromCtx()
if err != nil {
return fmt.Errorf("ssh transport: %w", err)
}
ctx := cmd.Context()
ctx, cancel := sshCmdCtx(cmd.Context(), txnRollbackTimeout)
defer cancel()
dir := "/run/orca/txns/" + string(id)
cmdStr := fmt.Sprintf("bash %s/rollback.sh", dir)
cmdStr := fmt.Sprintf("bash %s/rollback.sh", shellQuote(dir))
out, err := transport.Exec(ctx, txnRollbackLead, cmdStr)
if err != nil {
return fmt.Errorf("rollback %s on %s: %w (output: %s)", id, txnRollbackLead, err, string(out))
@@ -260,6 +303,7 @@ func init() {
txnApplyCmd.Flags().DurationVar(&txnApplyTimeout, "timeout", 5*time.Minute, "apply+verify timeout")
txnApplyCmd.Flags().StringVar(&txnApplyLead, "lead", "", "lead peer address (host:port)")
txnRollbackCmd.Flags().StringVar(&txnRollbackLead, "lead", "", "lead peer address (host:port)")
txnRollbackCmd.Flags().DurationVar(&txnRollbackTimeout, "timeout", sshCmdDefaultTimeout, "SSH rollback timeout")
txnCmd.AddCommand(txnApplyCmd)
txnCmd.AddCommand(txnListCmd)
+14 -7
View File
@@ -86,7 +86,8 @@ func TestTxnApplyClusterWideForceAndAck(t *testing.T) {
setupTxnTestEnv(t)
mt := &mockTxnTransport{execOut: []byte("applied")}
txnTransportOverride = mt
defer func() { txnTransportOverride = nil }()
txnAuthorizeOverride = func(ctx context.Context) (string, error) { return "oidc:test-operator", nil }
defer func() { txnTransportOverride = nil; txnAuthorizeOverride = nil }()
resetRootFlags(t)
var buf bytes.Buffer
@@ -115,7 +116,8 @@ func TestTxnApplyClusterWideYes(t *testing.T) {
setupTxnTestEnv(t)
mt := &mockTxnTransport{execOut: []byte("applied")}
txnTransportOverride = mt
defer func() { txnTransportOverride = nil }()
txnAuthorizeOverride = func(ctx context.Context) (string, error) { return "oidc:test-operator", nil }
defer func() { txnTransportOverride = nil; txnAuthorizeOverride = nil }()
resetRootFlags(t)
var buf bytes.Buffer
@@ -138,7 +140,8 @@ func TestTxnApplyNamespaceScoped(t *testing.T) {
setupTxnTestEnv(t)
mt := &mockTxnTransport{execOut: []byte("applied")}
txnTransportOverride = mt
defer func() { txnTransportOverride = nil }()
txnAuthorizeOverride = func(ctx context.Context) (string, error) { return "oidc:test-operator", nil }
defer func() { txnTransportOverride = nil; txnAuthorizeOverride = nil }()
resetRootFlags(t)
var buf bytes.Buffer
@@ -164,7 +167,8 @@ func TestTxnApplyClusterWideRefusesWithoutForce(t *testing.T) {
setupTxnTestEnv(t)
mt := &mockTxnTransport{}
txnTransportOverride = mt
defer func() { txnTransportOverride = nil }()
txnAuthorizeOverride = func(ctx context.Context) (string, error) { return "oidc:test-operator", nil }
defer func() { txnTransportOverride = nil; txnAuthorizeOverride = nil }()
resetRootFlags(t)
var buf bytes.Buffer
@@ -189,7 +193,8 @@ func TestTxnApplyClusterWideRefusesWithoutAck(t *testing.T) {
setupTxnTestEnv(t)
mt := &mockTxnTransport{}
txnTransportOverride = mt
defer func() { txnTransportOverride = nil }()
txnAuthorizeOverride = func(ctx context.Context) (string, error) { return "oidc:test-operator", nil }
defer func() { txnTransportOverride = nil; txnAuthorizeOverride = nil }()
resetRootFlags(t)
var buf bytes.Buffer
@@ -215,7 +220,8 @@ func TestTxnApplyAlreadyAppliedNoOp(t *testing.T) {
execErr: fmt.Errorf("%w: exit 5", sshpush.ErrPermanent),
}
txnTransportOverride = mt
defer func() { txnTransportOverride = nil }()
txnAuthorizeOverride = func(ctx context.Context) (string, error) { return "oidc:test-operator", nil }
defer func() { txnTransportOverride = nil; txnAuthorizeOverride = nil }()
resetRootFlags(t)
var buf bytes.Buffer
@@ -355,7 +361,8 @@ func TestTxnRollback(t *testing.T) {
setupTxnTestEnv(t)
mt := &mockTxnTransport{execOut: []byte("rolled-back")}
txnTransportOverride = mt
defer func() { txnTransportOverride = nil }()
txnAuthorizeOverride = func(ctx context.Context) (string, error) { return "oidc:test-operator", nil }
defer func() { txnTransportOverride = nil; txnAuthorizeOverride = nil }()
resetRootFlags(t)
var buf bytes.Buffer
+74 -3
View File
@@ -20,8 +20,10 @@ import (
"github.com/spf13/cobra"
"git.cloudinit.dev/coreci/orca/internal/certpaths"
"git.cloudinit.dev/coreci/orca/internal/migration"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/security"
)
var (
@@ -94,6 +96,33 @@ func init() {
rootCmd.AddCommand(upgradeCmd)
}
// acquireUpgradeLock atomically creates an exclusive lock file at
// paths.ClusterDir()/upgrade.lock (REQ-156, P07 T3). Returns a release
// function that MUST be deferred (it removes the lock file). If the
// lock file already exists, returns an error "upgrade already in
// progress" — preventing two concurrent `orca upgrade` invocations
// from racing on the same cluster state (cutover, install.sh, peer
// user creation). O_CREATE|O_EXCL is atomic under POSIX: only one of
// two racing callers succeeds; the other gets EEXIST.
func acquireUpgradeLock() (func(), error) {
lockPath := filepath.Join(paths.ClusterDir(), "upgrade.lock")
if err := os.MkdirAll(filepath.Dir(lockPath), 0o755); err != nil {
return nil, fmt.Errorf("create cluster dir for upgrade lock: %w", err)
}
f, err := os.OpenFile(lockPath, os.O_CREATE|os.O_EXCL|os.O_WRONLY, 0o600)
if err != nil {
if os.IsExist(err) {
return nil, fmt.Errorf("upgrade already in progress (lock file %s exists; remove it if stale)", lockPath)
}
return nil, fmt.Errorf("acquire upgrade lock: %w", err)
}
// Write the current PID + timestamp for diagnostics (best-effort;
// a stale lock from a crashed process is the operator's signal).
_, _ = f.WriteString(fmt.Sprintf("pid=%d started=%s\n", os.Getpid(), time.Now().UTC().Format(time.RFC3339)))
_ = f.Close()
return func() { _ = os.Remove(lockPath) }, nil
}
// UpgradeResult is the JSON-serializable summary of an upgrade run.
type UpgradeResult struct {
TargetVersion string `json:"target_version"`
@@ -134,8 +163,26 @@ func runUpgrade(cmd *cobra.Command, out interface{ Write([]byte) (int, error) })
return nil
}
// REQ-156 / P07 T3: v0.8 layout detection is read-only and MUST
// run BEFORE the upgrade lock is acquired — the lock creates the
// cluster/ dir (for the lock file), and Detectv08 treats the
// presence of a cluster/ dir as "already v0.11" (no migration
// needed). Detecting first avoids a false negative that would
// skip the migration on a genuine v0.8 layout.
home := paths.Root()
if migration.Detectv08(home) {
needV08Migration := migration.Detectv08(home)
// Acquire an exclusive upgrade lock for the rest of the run so
// two concurrent `orca upgrade` invocations cannot race on the
// cutover / install.sh / peer user creation. The lock is released
// on return (including error paths).
upgradeRelease, err := acquireUpgradeLock()
if err != nil {
return err
}
defer upgradeRelease()
if needV08Migration {
if !jsonOutput {
fmt.Fprintf(out, "• v0.8 layout detected; running data migration first\n")
}
@@ -292,8 +339,18 @@ func performCutover(ctx context.Context, runner commandRunner, out interface{ Wr
return true, nil
}
// verifyCutover runs the C-25 post-cutover check: curl -k
// verifyCutover runs the C-25 post-cutover check: an HTTPS GET to
// https://localhost:443/ must return HTTP 200.
//
// REQ-157 / P08 T7: previously this used the default http.Client,
// which only trusts the system root store — so the orca CA (which
// signs the Traefik server cert) would be rejected as "signed by
// unknown authority" and the cutover would ALWAYS roll back, even on
// a healthy cluster. Now it builds a *tls.Config from the orca CA
// pool (security.ClientTLSConfig against certpaths.CACertPath()) so
// the server cert validates. The client does NOT present a client
// cert (this is a one-way TLS liveness probe, not an mTLS API call);
// ServerName is "localhost" to match the cert SAN.
func verifyCutover(out interface{ Write([]byte) (int, error) }) error {
if httpClientOverride != nil {
code, err := httpClientOverride("https://localhost:443/")
@@ -306,7 +363,21 @@ func verifyCutover(out interface{ Write([]byte) (int, error) }) error {
return nil
}
client := &http.Client{Timeout: 10 * time.Second}
caPath := certpaths.CACertPath()
tlsCfg, err := security.ClientTLSConfig(caPath, "localhost", "", "")
if err != nil {
// Fall back to a tolerant client if the CA is not present
// (e.g. running verifyCutover in a test harness without a
// cluster). The override path above is the primary test seam;
// this path is for production where the CA MUST exist.
return fmt.Errorf("verifyCutover: load orca CA %s: %w", caPath, err)
}
client := &http.Client{
Timeout: 10 * time.Second,
Transport: &http.Transport{
TLSClientConfig: tlsCfg,
},
}
resp, err := client.Get("https://localhost:443/")
if err != nil {
return fmt.Errorf("curl: %w", err)
+59
View File
@@ -0,0 +1,59 @@
// Package cli: validate.go provides shared input-validation helpers for
// CLI command arguments that are interpolated into remote shell commands
// or filesystem paths (Phase 02 injection hardening, v0.13).
//
// These helpers enforce strict allowlists so that attacker-controlled
// values (job names, txn IDs, alloc IDs, country codes) cannot reach
// shell interpolation or path joins without matching a known-safe shape.
package cli
import (
"regexp"
"strings"
)
// safeNameRe matches the allowlist for shell-interpolated identifiers
// (job names, alloc IDs): ASCII letters, digits, underscore, hyphen.
// Used to prevent backtick/command-substitution and metacharacter
// injection into remote shell commands.
var safeNameRe = regexp.MustCompile(`^[A-Za-z0-9_-]+$`)
// txnIDRe matches the canonical orca transaction ID format: "T-" prefix
// followed by exactly 16 lowercase hex digits. Used to validate txn IDs
// before they are interpolated into filesystem paths or remote shell
// commands (`orca txn rollback`, `orca nft diff --against`).
var txnIDRe = regexp.MustCompile(`^T-[0-9a-f]{16}$`)
// countryCodeRe matches ISO-3166 alpha-2 country codes: exactly two
// uppercase ASCII letters. Used by `orca nft country block add` before
// codes are interpolated into the nft ruleset.
var countryCodeRe = regexp.MustCompile(`^[A-Z]{2}$`)
// validSafeName reports whether s is a safe shell-interpolation
// identifier (ASCII alphanumeric, underscore, hyphen only, non-empty).
func validSafeName(s string) bool {
return safeNameRe.MatchString(s)
}
// validTxnID reports whether s matches the canonical orca txn ID format
// (^T-[0-9a-f]{16}$).
func validTxnID(s string) bool {
return txnIDRe.MatchString(s)
}
// validCountryCode reports whether s is a valid ISO-3166 alpha-2 code
// (two uppercase letters).
func validCountryCode(s string) bool {
return countryCodeRe.MatchString(s)
}
// shellQuote single-quotes a string for safe shell interpolation over
// SSH exec. It escapes embedded single-quotes via the standard '\” idiom
// (POSIX shell). This is the cli-package copy of the helper duplicated
// across runtime/identity/stepca/sshpush to avoid import cycles; it
// hardens command interpolation against backtick/command-substitution
// injection (Go's %q does NOT escape backticks, and bash executes
// command substitution inside double quotes).
func shellQuote(s string) string {
return "'" + strings.ReplaceAll(s, "'", "'\\''") + "'"
}
+38
View File
@@ -20,6 +20,42 @@ type Config struct {
ServerCertPath string `hcl:"server_cert_path,optional"`
ServerKeyPath string `hcl:"server_key_path,optional"`
NodeCapacity *CapacityConfig `hcl:"node_capacity,block"`
// OIDC is the OIDC client config block (P06, v0.13; R-021). The
// bundled Dex (deployed by `orca auth init-idp`) is the default
// issuer; an explicit oidc.issuer here repoints the CLI to a BYO
// external IdP. loadOIDCConfig reads this block before falling back
// to --issuer/--client-id flags and env vars.
OIDC *OIDCConfig `hcl:"oidc,block"`
// ClusterDomain is the cluster's Traefik-served domain (C-38). It
// is the WebAuthn relying-party ID default and the Dex issuer host.
// May be overridden by --rp-id on `orca auth init-idp`.
ClusterDomain string `hcl:"cluster_domain,optional"`
// ACL is the access-control config block (P04, v0.13; C-45).
// When ACL.Enforce is false (the default for the first run after
// P04 wiring), ACL denials are LOGGED but NOT enforced — the
// request proceeds. The operator switches to true after verifying
// the bootstrap ACL.
ACL *ACLConfig `hcl:"acl,block"`
}
// ACLConfig is the acl block in config (P04, C-45).
type ACLConfig struct {
// Enforce controls whether ACL denials return 403 (true) or are
// logged but allowed (false, the staged-rollout default).
Enforce bool `hcl:"enforce,optional"`
}
// OIDCConfig is the oidc block in config (P06, R-021). Mirrors
// identity.OIDCConfig (kept separate to avoid an internal/config ->
// internal/identity dependency cycle).
type OIDCConfig struct {
Issuer string `hcl:"issuer,optional"`
ClientID string `hcl:"client_id,optional"`
ClientSecret string `hcl:"client_secret,optional"`
Scopes []string `hcl:"scopes,optional"`
}
type Flags struct {
@@ -120,6 +156,8 @@ func (c *Config) MergeOverrides(flags Flags, env Environ) *Config {
ServerCertPath: c.ServerCertPath,
ServerKeyPath: c.ServerKeyPath,
NodeCapacity: c.NodeCapacity,
OIDC: c.OIDC,
ClusterDomain: c.ClusterDomain,
}
applyStr := func(flag *string, envKey, fileVal string) string {
+49
View File
@@ -89,6 +89,7 @@ func extractFrontmatter(content string) (string, bool) {
func parseFrontmatterBlock(block, path string) (*Config, error) {
cfg := &Config{}
var inCapacity bool
var inOIDC bool
lines := strings.Split(block, "\n")
for lineNo, raw := range lines {
@@ -103,6 +104,7 @@ func parseFrontmatterBlock(block, path string) (*Config, error) {
// A top-level key (no leading indent).
if indent == 0 {
inCapacity = false
inOIDC = false
key, val, ok := splitKV(trimmed)
if !ok {
continue
@@ -113,6 +115,10 @@ func parseFrontmatterBlock(block, path string) (*Config, error) {
cfg.NodeCapacity = &CapacityConfig{}
inCapacity = true
}
if key == "oidc" {
cfg.OIDC = &OIDCConfig{}
inOIDC = true
}
continue
}
applyScalar(cfg, key, val, path, lineNo)
@@ -135,6 +141,26 @@ func parseFrontmatterBlock(block, path string) (*Config, error) {
cfg.NodeCapacity.MemoryMB = n
}
}
continue
}
// Indented line under the oidc block.
if inOIDC && cfg.OIDC != nil {
key, val, hasVal := splitKV(trimmed)
if !hasVal {
continue
}
switch key {
case "issuer":
cfg.OIDC.Issuer = unquote(val)
case "client_id":
cfg.OIDC.ClientID = unquote(val)
case "client_secret":
cfg.OIDC.ClientSecret = unquote(val)
case "scopes":
// Comma-separated list, optionally bracketed as [a, b].
cfg.OIDC.Scopes = parseScopes(val)
}
continue
}
}
return cfg, nil
@@ -153,11 +179,34 @@ func applyScalar(cfg *Config, key, val, path string, lineNo int) {
cfg.ServerCertPath = unquote(val)
case "server_key_path":
cfg.ServerKeyPath = unquote(val)
case "cluster_domain":
cfg.ClusterDomain = unquote(val)
}
_ = path
_ = lineNo
}
// parseScopes parses a scopes value into a []string. Supports both a
// comma-separated bare list (openid, profile, email) and a YAML-style
// flow list ([openid, profile]). Empty values are dropped.
func parseScopes(val string) []string {
val = strings.TrimSpace(val)
val = unquote(val)
// Strip surrounding brackets.
if len(val) >= 2 && val[0] == '[' && val[len(val)-1] == ']' {
val = val[1 : len(val)-1]
}
var out []string
for _, part := range strings.Split(val, ",") {
part = strings.TrimSpace(part)
part = unquote(part)
if part != "" {
out = append(out, part)
}
}
return out
}
func splitKV(s string) (key, val string, ok bool) {
idx := strings.Index(s, ":")
if idx < 0 {
+257
View File
@@ -0,0 +1,257 @@
// Package daemon — acl.go provides the access-control enforcement
// layer wired into the daemon's HTTP handlers (P04, v0.13; C-44/C-45).
//
// The daemon extracts the caller's identity from the mTLS peer
// certificate (SPIFFE SVID URI SAN, or OIDC sub in the cert's
// Subject.CommonName when the IdP embeds it), loads the cluster ACL
// from paths.ACLPath(), and calls acl.Check before dispatching the
// request. Health endpoints (/healthz, /readyz, /v1/status) are
// exempt (liveness probes must not be gated on authorization).
//
// C-45 staged rollout: when the daemon is configured with
// enforce=false (the default for the first run after wiring), ACL
// denials are LOGGED but NOT enforced — the request proceeds. This
// lets operators verify the bootstrap ACL grants the right identities
// before flipping to enforce mode. The operator switches via the
// `acl.enforce` config flag.
package daemon
import (
"crypto/x509"
"encoding/json"
"fmt"
"log/slog"
"net/http"
"os"
"strings"
"git.cloudinit.dev/coreci/orca/internal/acl"
"git.cloudinit.dev/coreci/orca/internal/paths"
)
// aclPolicy is the runtime ACL enforcement policy for the daemon.
// It is constructed once at server start (see NewACLPolicy) and
// shared across handlers. The zero value is deny-by-default with
// enforce=true.
type aclPolicy struct {
// enforcer is the loaded ACL. nil means "no ACL file present" —
// in that case deny-by-default applies (no identity has any
// permission).
enforcer *acl.ACL
// enforce controls whether denials return 403 (true) or are
// logged but allowed (false, C-45 log-only mode). The default
// for the first run after P04 wiring is false.
enforce bool
log *slog.Logger
}
// NewACLPolicy loads the ACL from paths.ACLPath() and returns a
// policy. A missing ACL file is treated as an empty ACL (deny-by-
// default). enforce controls C-45 staged rollout.
func NewACLPolicy(enforce bool, log *slog.Logger) *aclPolicy {
if log == nil {
log = slog.Default()
}
p := &aclPolicy{enforce: enforce, log: log, enforcer: acl.NewACL()}
a, err := loadDaemonACL()
if err != nil {
log.Warn("acl load failed; deny-by-default with empty ACL",
slog.String("component", "daemon"),
slog.String("error", err.Error()))
return p
}
if a != nil {
p.enforcer = a
}
log.Info("acl policy loaded",
slog.String("component", "daemon"),
slog.Bool("enforce", enforce),
slog.Int("entries", len(p.enforcer.List())))
return p
}
// aclState mirrors internal/cli/aclState (kept private there). We
// duplicate the JSON shape to avoid an import cycle (cli imports
// daemon transitively via the binary, but daemon must not import cli).
type aclState struct {
Entries []acl.ACLEntry `json:"entries"`
}
// loadDaemonACL reads paths.ACLPath() and returns an *acl.ACL. A
// missing file is treated as an empty ACL (not an error).
func loadDaemonACL() (*acl.ACL, error) {
a := acl.NewACL()
path := paths.ACLPath()
data, err := os.ReadFile(path)
if err != nil {
if os.IsNotExist(err) {
return a, nil
}
return nil, fmt.Errorf("read acl state: %w", err)
}
if len(data) == 0 {
return a, nil
}
var st aclState
if err := json.Unmarshal(data, &st); err != nil {
return nil, fmt.Errorf("parse acl state: %w", err)
}
for _, e := range st.Entries {
a.Grant(e.Identity, e.Namespace, e.Permissions)
}
return a, nil
}
// IdentityFromCert extracts the caller's identity from an mTLS peer
// certificate. It prefers a SPIFFE SVID URI SAN (KindSpiffe); if no
// spiffe:// URI is present, it falls back to the cert's
// Subject.CommonName as an OIDC sub (KindOidc). Returns an error if
// the cert carries neither (unauthenticated).
//
// The namespace for a SPIFFE identity is extracted from the URI path;
// for an OIDC identity the namespace is empty (the ACL check takes
// the namespace as a separate argument).
func IdentityFromCert(cert *x509.Certificate) (acl.Identity, error) {
if cert == nil {
return acl.Identity{}, fmt.Errorf("acl: peer certificate is nil")
}
for _, u := range cert.URIs {
if u == nil {
continue
}
s := u.String()
if strings.HasPrefix(s, "spiffe://") {
ns, err := acl.SpiffeNamespace(s)
if err != nil {
// Malformed spiffe URI — treat as unauthenticated so
// the deny-by-default path applies. Log the error at
// the call site.
return acl.Identity{Kind: acl.KindSpiffe, ID: s, Namespace: ""}, fmt.Errorf("acl: malformed spiffe uri: %w", err)
}
return acl.Identity{Kind: acl.KindSpiffe, ID: s, Namespace: ns}, nil
}
}
if cn := cert.Subject.CommonName; cn != "" {
return acl.Identity{Kind: acl.KindOidc, ID: cn}, nil
}
return acl.Identity{}, fmt.Errorf("acl: peer cert has no spiffe URI SAN and no CommonName (unauthenticated)")
}
// peerIdentity extracts the identity from the request's mTLS peer
// certificate. Returns an error (and the zero Identity) if no peer
// cert is present or the cert carries no identity. The caller is
// expected to deny the request in that case.
func peerIdentity(r *http.Request) (acl.Identity, error) {
if r.TLS == nil || len(r.TLS.PeerCertificates) == 0 {
return acl.Identity{}, fmt.Errorf("acl: no mTLS peer certificate (unauthenticated)")
}
return IdentityFromCert(r.TLS.PeerCertificates[0])
}
// Check evaluates whether the caller identified by the request's mTLS
// peer cert has perm on ns. It returns the extracted identity (for
// audit logging) and a boolean allow.
//
// In enforce=true mode, a denial returns allow=false and the handler
// is expected to write a 403. In enforce=false mode (C-45 log-only),
// a denial is logged but allow=true is returned so the request
// proceeds — this lets operators verify the bootstrap ACL before
// flipping to enforce.
//
// A request with no peer cert (unauthenticated) is denied in enforce
// mode and allowed (but logged) in log-only mode, so health probes
// and bootstrap traffic keep flowing during rollout. Operators should
// flip to enforce=true as soon as the bootstrap ACL is verified.
func (p *aclPolicy) Check(r *http.Request, ns string, perm acl.Permission) (identity acl.Identity, allow bool) {
id, err := peerIdentity(r)
if err != nil {
// Unauthenticated. In enforce mode: deny. In log-only mode:
// log + allow (C-45: keep traffic flowing during rollout).
p.log.Warn("acl denial (unauthenticated)",
slog.String("component", "daemon"),
slog.String("namespace", ns),
slog.String("permission", permName(perm)),
slog.String("error", err.Error()),
slog.Bool("enforce", p.enforce),
)
if p.enforce {
return acl.Identity{}, false
}
return acl.Identity{}, true
}
allowed := p.enforcer.Check(id, ns, perm)
if !allowed {
p.log.Warn("acl denial",
slog.String("component", "daemon"),
slog.String("identity_kind", id.Kind),
slog.String("identity_id", id.ID),
slog.String("namespace", ns),
slog.String("permission", permName(perm)),
slog.Bool("enforce", p.enforce),
)
if p.enforce {
return id, false
}
return id, true
}
return id, true
}
// CheckOidc evaluates an OIDC-claims identity (sub + groups) against
// the ACL. Used by paths that have a verified ID token (e.g. the
// SSH-push applier validates ORCA_OIDC_TOKEN and threads the claims
// here). Returns allow=true in log-only mode even on denial.
func (p *aclPolicy) CheckOidc(claims acl.OIDCClaims, ns string, perm acl.Permission) (allow bool) {
allowed := p.enforcer.CheckOidc(claims, ns, perm)
if !allowed {
p.log.Warn("acl denial (oidc)",
slog.String("component", "daemon"),
slog.String("oidc_sub", claims.Subject),
slog.String("namespace", ns),
slog.String("permission", permName(perm)),
slog.Bool("enforce", p.enforce),
)
if p.enforce {
return false
}
return true
}
return true
}
// Enforce reports whether the policy is in enforce mode (C-45).
func (p *aclPolicy) Enforce() bool { return p.enforce }
// permName renders a Permission bitmask as a comma-separated string
// for log lines. Mirrors internal/cli.permName but is duplicated here
// to avoid an import cycle.
func permName(p acl.Permission) string {
var parts []string
if p&acl.PermRead != 0 {
parts = append(parts, "read")
}
if p&acl.PermWrite != 0 {
parts = append(parts, "write")
}
if p&acl.PermAdmin != 0 {
parts = append(parts, "admin")
}
if len(parts) == 0 {
return "none"
}
return strings.Join(parts, ",")
}
// deny writes a 403 with the standard error envelope.
func deny(w http.ResponseWriter, id acl.Identity, ns string, perm acl.Permission) {
msg := fmt.Sprintf("access denied: %s %s on %s", permName(perm), idDisplay(id), ns)
writeError(w, http.StatusForbidden, msg)
}
// idDisplay renders an identity for error/log messages.
func idDisplay(id acl.Identity) string {
if id.ID == "" {
return "anonymous"
}
return id.Kind + ":" + id.ID
}
+296
View File
@@ -0,0 +1,296 @@
// Package daemon — acl_test.go verifies the ACL enforcement wiring
// (P04, v0.13; C-44/C-45). It exercises the aclPolicy.Check path
// with constructed mTLS peer certificates (SPIFFE SVID + OIDC CN)
// and asserts deny-by-default + log-only mode semantics.
package daemon
import (
"crypto/rand"
"crypto/rsa"
"crypto/tls"
"crypto/x509"
"crypto/x509/pkix"
"encoding/json"
"encoding/pem"
"log/slog"
"math/big"
"net/http"
"net/http/httptest"
"net/url"
"os"
"path/filepath"
"strings"
"testing"
"time"
"git.cloudinit.dev/coreci/orca/internal/acl"
"git.cloudinit.dev/coreci/orca/internal/paths"
"git.cloudinit.dev/coreci/orca/internal/store"
)
// aclStateJSON mirrors the on-disk acl.json shape.
type aclStateJSON struct {
Entries []acl.ACLEntry `json:"entries"`
}
// mustMarshal marshals v or fails the test.
func mustMarshal(t *testing.T, v any) []byte {
t.Helper()
b, err := json.MarshalIndent(v, "", " ")
if err != nil {
t.Fatalf("marshal: %v", err)
}
return b
}
// writeACLFile writes the given entries to paths.ACLPath() under a
// fresh $ORCA_HOME so NewACLPolicy picks them up.
func writeACLFile(t *testing.T, entries []acl.ACLEntry) {
t.Helper()
home := t.TempDir()
t.Setenv("ORCA_HOME", home)
if err := os.MkdirAll(paths.ClusterDir(), 0o755); err != nil {
t.Fatalf("mkdir cluster dir: %v", err)
}
data := mustMarshal(t, aclStateJSON{Entries: entries})
if err := os.WriteFile(paths.ACLPath(), data, 0o600); err != nil {
t.Fatalf("write acl: %v", err)
}
}
// buildSelfSignedCert builds an in-memory self-signed x509 cert with
// the given SPIFFE URI SAN and CommonName. The ACL layer only inspects
// URIs + CommonName, not the signature chain (chain verification is
// the mTLS handshake's job).
func buildSelfSignedCert(t *testing.T, spiffeURI, commonName string) *x509.Certificate {
t.Helper()
key, err := rsa.GenerateKey(rand.Reader, 2048)
if err != nil {
t.Fatalf("rsa key: %v", err)
}
tmpl := &x509.Certificate{
SerialNumber: big.NewInt(1),
Subject: pkix.Name{CommonName: commonName},
NotBefore: time.Now().Add(-time.Hour),
NotAfter: time.Now().Add(time.Hour),
DNSNames: []string{"localhost"},
}
if spiffeURI != "" {
u, err := url.Parse(spiffeURI)
if err != nil {
t.Fatalf("parse spiffe uri: %v", err)
}
tmpl.URIs = []*url.URL{u}
}
der, err := x509.CreateCertificate(rand.Reader, tmpl, tmpl, &key.PublicKey, key)
if err != nil {
t.Fatalf("create cert: %v", err)
}
cert, err := x509.ParseCertificate(der)
if err != nil {
t.Fatalf("parse cert: %v", err)
}
// Round-trip through PEM so the cert is realistic.
_ = pem.EncodeToMemory(&pem.Block{Type: "CERTIFICATE", Bytes: der})
return cert
}
// makePeerCert is a shorthand for buildSelfSignedCert.
func makePeerCert(t *testing.T, spiffeURI, commonName string) *x509.Certificate {
return buildSelfSignedCert(t, spiffeURI, commonName)
}
// reqWithPeerCert builds an *http.Request whose r.TLS.PeerCertificates
// is populated with the given cert, simulating an mTLS handshake.
func reqWithPeerCert(cert *x509.Certificate) *http.Request {
r := httptest.NewRequest(http.MethodGet, "/v1/jobs", nil)
r.TLS = &tls.ConnectionState{
PeerCertificates: []*x509.Certificate{cert},
}
return r
}
// newTestServer builds a daemon Server with a temp DB and the given
// ACL enforce mode. Used by the handler-level tests.
func newACLTestServer(t *testing.T, enforce bool) *Server {
t.Helper()
db, err := store.Open(filepath.Join(t.TempDir(), "test.db"))
if err != nil {
t.Fatalf("open db: %v", err)
}
t.Cleanup(func() { db.Close() })
s := NewServer(Options{
DB: db,
Log: slog.New(slog.NewTextHandler(os.Stderr, nil)),
Addr: ":0",
ACLEnforce: enforce,
})
return s
}
// --- aclPolicy unit tests ---
// TestACLPolicyDenyByDefault verifies that an authenticated request
// with no matching ACL entry is denied in enforce mode.
func TestACLPolicyDenyByDefault(t *testing.T) {
writeACLFile(t, nil) // empty ACL
p := NewACLPolicy(true, slog.New(slog.NewTextHandler(os.Stderr, nil)))
cert := makePeerCert(t, "spiffe://orca.local/ns/myapp/sa/svc1/alloc-1", "")
r := reqWithPeerCert(cert)
_, ok := p.Check(r, "_defaults", acl.PermRead)
if ok {
t.Fatal("expected deny (no ACL entry), got allow")
}
}
// TestACLPolicyAllowWithEntry verifies that an authenticated request
// with a matching ACL entry is allowed.
func TestACLPolicyAllowWithEntry(t *testing.T) {
id := acl.Identity{Kind: acl.KindSpiffe, ID: "spiffe://orca.local/ns/_defaults/sa/orca/alloc-1", Namespace: "_defaults"}
a := acl.NewACL()
a.Grant(id, "_defaults", acl.PermRead|acl.PermWrite)
writeACLFile(t, a.List())
p := NewACLPolicy(true, slog.New(slog.NewTextHandler(os.Stderr, nil)))
cert := makePeerCert(t, id.ID, "")
r := reqWithPeerCert(cert)
gotID, ok := p.Check(r, "_defaults", acl.PermRead)
if !ok {
t.Fatal("expected allow (matching entry), got deny")
}
if gotID.ID != id.ID {
t.Errorf("identity ID = %q, want %q", gotID.ID, id.ID)
}
}
// TestACLPolicyUnauthenticatedEnforce verifies that a request with no
// peer cert is denied in enforce mode.
func TestACLPolicyUnauthenticatedEnforce(t *testing.T) {
writeACLFile(t, nil)
p := NewACLPolicy(true, slog.New(slog.NewTextHandler(os.Stderr, nil)))
r := httptest.NewRequest(http.MethodGet, "/v1/jobs", nil)
_, ok := p.Check(r, "_defaults", acl.PermRead)
if ok {
t.Fatal("expected deny for unauthenticated in enforce mode, got allow")
}
}
// TestACLPolicyLogOnlyAllowsDenials (C-45) verifies that in log-only
// mode (enforce=false), denials are logged but the request proceeds
// (allow=true). This is the staged-rollout semantics.
func TestACLPolicyLogOnlyAllowsDenials(t *testing.T) {
writeACLFile(t, nil) // empty ACL → all denials
p := NewACLPolicy(false, slog.New(slog.NewTextHandler(os.Stderr, nil)))
// Unauthenticated in log-only mode → logged but allowed.
r := httptest.NewRequest(http.MethodGet, "/v1/jobs", nil)
_, ok := p.Check(r, "_defaults", acl.PermRead)
if !ok {
t.Fatal("expected allow in log-only mode (unauthenticated), got deny")
}
// Authenticated-but-no-entry in log-only mode → logged but allowed.
cert := makePeerCert(t, "spiffe://orca.local/ns/myapp/sa/svc1/alloc-1", "")
r2 := reqWithPeerCert(cert)
_, ok = p.Check(r2, "_defaults", acl.PermRead)
if !ok {
t.Fatal("expected allow in log-only mode (no entry), got deny")
}
}
// TestACLPolicyOIDCCNIdentity verifies that a cert with no SPIFFE URI
// but a CommonName is treated as an OIDC identity.
func TestACLPolicyOIDCCNIdentity(t *testing.T) {
id := acl.Identity{Kind: acl.KindOidc, ID: "operator@example.com"}
a := acl.NewACL()
a.Grant(id, "_defaults", acl.PermRead)
writeACLFile(t, a.List())
p := NewACLPolicy(true, slog.New(slog.NewTextHandler(os.Stderr, nil)))
cert := makePeerCert(t, "", "operator@example.com")
r := reqWithPeerCert(cert)
gotID, ok := p.Check(r, "_defaults", acl.PermRead)
if !ok {
t.Fatal("expected allow for OIDC CN identity, got deny")
}
if gotID.Kind != acl.KindOidc || gotID.ID != "operator@example.com" {
t.Errorf("identity = %+v, want oidc:operator@example.com", gotID)
}
}
// --- Handler-level tests (T10) ---
// TestACLJobsHandlerEnforceDeniesUnauthenticated verifies the wired
// jobs handler denies an unauthenticated request in enforce mode.
func TestACLJobsHandlerEnforceDeniesUnauthenticated(t *testing.T) {
writeACLFile(t, nil)
srv := newACLTestServer(t, true)
rec := httptest.NewRecorder()
r := httptest.NewRequest(http.MethodGet, "/v1/jobs", nil)
srv.handleJobsCollection(rec, r)
if rec.Code != http.StatusForbidden {
t.Errorf("unauthenticated /v1/jobs (enforce): %d, want 403", rec.Code)
}
if !strings.Contains(rec.Body.String(), "access denied") {
t.Errorf("body should contain 'access denied': %s", rec.Body.String())
}
}
// TestACLJobsHandlerLogOnlyAllowsUnauthenticated (C-45) verifies the
// wired jobs handler allows an unauthenticated request in log-only
// mode (the denial is logged but the request proceeds).
func TestACLJobsHandlerLogOnlyAllowsUnauthenticated(t *testing.T) {
writeACLFile(t, nil)
srv := newACLTestServer(t, false)
rec := httptest.NewRecorder()
r := httptest.NewRequest(http.MethodGet, "/v1/jobs", nil)
srv.handleJobsCollection(rec, r)
if rec.Code == http.StatusForbidden {
t.Errorf("unauthenticated /v1/jobs (log-only): %d, want non-403", rec.Code)
}
}
// TestACLJobsHandlerAllowsAuthenticatedWithEntry verifies the wired
// jobs handler allows an authenticated request with a matching ACL
// entry in enforce mode.
func TestACLJobsHandlerAllowsAuthenticatedWithEntry(t *testing.T) {
id := acl.Identity{Kind: acl.KindSpiffe, ID: "spiffe://orca.local/ns/_defaults/sa/orca/alloc-1", Namespace: "_defaults"}
a := acl.NewACL()
a.Grant(id, "_defaults", acl.PermRead)
writeACLFile(t, a.List())
srv := newACLTestServer(t, true)
cert := makePeerCert(t, id.ID, "")
rec := httptest.NewRecorder()
r := reqWithPeerCert(cert)
srv.handleJobsCollection(rec, r)
if rec.Code == http.StatusForbidden {
t.Errorf("authenticated /v1/jobs (matching entry): %d, want non-403", rec.Code)
}
}
// TestACLNodesHandlerEnforceDeniesUnauthenticated verifies the nodes
// handler denies an unauthenticated request in enforce mode.
func TestACLNodesHandlerEnforceDeniesUnauthenticated(t *testing.T) {
writeACLFile(t, nil)
srv := newACLTestServer(t, true)
rec := httptest.NewRecorder()
r := httptest.NewRequest(http.MethodGet, "/v1/nodes", nil)
srv.handleNodesCollection(rec, r)
if rec.Code != http.StatusForbidden {
t.Errorf("unauthenticated /v1/nodes (enforce): %d, want 403", rec.Code)
}
}
// TestACLTasksHandlerEnforceDeniesUnauthenticated verifies the tasks
// handler denies an unauthenticated request in enforce mode.
func TestACLTasksHandlerEnforceDeniesUnauthenticated(t *testing.T) {
writeACLFile(t, nil)
srv := newACLTestServer(t, true)
rec := httptest.NewRecorder()
r := httptest.NewRequest(http.MethodGet, "/v1/tasks", nil)
srv.handleTasksCollection(rec, r)
if rec.Code != http.StatusForbidden {
t.Errorf("unauthenticated /v1/tasks (enforce): %d, want 403", rec.Code)
}
}
+34 -1
View File
@@ -8,6 +8,7 @@ package daemon
import (
"net/http"
"git.cloudinit.dev/coreci/orca/internal/acl"
"git.cloudinit.dev/coreci/orca/internal/transport"
)
@@ -31,8 +32,40 @@ func NewDispatchHandlers(d transport.Dispatcher, dedupe *transport.IdempotencySt
}
// Mount registers Submit and Status on the given mux. Called by the
// daemon's mux builder.
// daemon's mux builder. P04 wraps each handler in an ACL middleware
// that calls s.acl.Check before delegating; the dispatch namespace is
// the default (cluster-wide) namespace. Submit = write, Status = read.
// When s.acl is nil (legacy/compat) the middleware is a no-op pass-
// through.
func (h *DispatchHandlers) Mount(mux *http.ServeMux) {
mux.Handle("/orca.v1.Dispatch/Submit", h.Submit)
mux.Handle("/orca.v1.Dispatch/Status", h.Status)
}
// mountDispatchWithACL mounts the dispatch handlers wrapped in ACL
// middleware. P04: Submit requires write on "_defaults"; Status
// requires read. When policy is nil, the handlers are mounted
// unwrapped (legacy/compat for tests).
func (h *DispatchHandlers) mountWithACL(mux *http.ServeMux, policy *aclPolicy) {
if policy == nil {
h.Mount(mux)
return
}
mux.Handle("/orca.v1.Dispatch/Submit", aclMiddleware(policy, "_defaults", acl.PermWrite, h.Submit))
mux.Handle("/orca.v1.Dispatch/Status", aclMiddleware(policy, "_defaults", acl.PermRead, h.Status))
}
// aclMiddleware wraps an http.Handler with an ACL check. On denial in
// enforce mode it writes a 403 and returns; in log-only mode (C-45)
// it logs and delegates. The extracted identity is stashed in the
// request context under the identity key so downstream handlers / the
// audit layer can read it.
func aclMiddleware(policy *aclPolicy, ns string, perm acl.Permission, next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if id, ok := policy.Check(r, ns, perm); !ok {
deny(w, id, ns, perm)
return
}
next.ServeHTTP(w, r)
})
}
+25
View File
@@ -8,6 +8,7 @@ import (
"strings"
"time"
"git.cloudinit.dev/coreci/orca/internal/acl"
"git.cloudinit.dev/coreci/orca/internal/model"
"git.cloudinit.dev/coreci/orca/internal/store"
)
@@ -19,8 +20,17 @@ func (s *Server) handleJobsCollection(w http.ResponseWriter, r *http.Request) {
ctx, cancel := context.WithTimeout(r.Context(), 5*time.Second)
defer cancel()
// P04 ACL enforcement (C-44). Jobs are cluster-wide in v0.1, so
// the namespace is the default namespace. GET = read, POST = write.
ns := "_defaults"
switch r.Method {
case http.MethodGet:
if s.acl != nil {
if id, ok := s.acl.Check(r, ns, acl.PermRead); !ok {
deny(w, id, ns, acl.PermRead)
return
}
}
jobs, err := store.NewJobRepo(s.db).List(ctx)
if err != nil {
s.log.Error("list jobs",
@@ -35,6 +45,12 @@ func (s *Server) handleJobsCollection(w http.ResponseWriter, r *http.Request) {
writeJSON(w, http.StatusOK, map[string]any{"jobs": jobs, "count": len(jobs)})
case http.MethodPost:
if s.acl != nil {
if id, ok := s.acl.Check(r, ns, acl.PermWrite); !ok {
deny(w, id, ns, acl.PermWrite)
return
}
}
// Job submission via HTTP is intentionally not exposed in v0.1.
// The CLI submits jobs to the local store directly; the daemon
// exists for observability and lifecycle control.
@@ -56,6 +72,15 @@ func (s *Server) handleJobsItem(w http.ResponseWriter, r *http.Request) {
ctx, cancel := context.WithTimeout(r.Context(), 5*time.Second)
defer cancel()
// P04 ACL enforcement (C-44). Job detail + tasks list are reads.
ns := "_defaults"
if s.acl != nil {
if id, ok := s.acl.Check(r, ns, acl.PermRead); !ok {
deny(w, id, ns, acl.PermRead)
return
}
}
// Path is /v1/jobs/{id} or /v1/jobs/{id}/tasks
path := strings.TrimPrefix(r.URL.Path, "/v1/jobs/")
parts := strings.Split(path, "/")
+10
View File
@@ -5,6 +5,7 @@ import (
"net/http"
"time"
"git.cloudinit.dev/coreci/orca/internal/acl"
"git.cloudinit.dev/coreci/orca/internal/model"
"git.cloudinit.dev/coreci/orca/internal/store"
)
@@ -19,6 +20,15 @@ func (s *Server) handleNodesCollection(w http.ResponseWriter, r *http.Request) {
ctx, cancel := context.WithTimeout(r.Context(), 5*time.Second)
defer cancel()
// P04 ACL enforcement (C-44). Node list is a cluster-wide read.
ns := "_defaults"
if s.acl != nil {
if id, ok := s.acl.Check(r, ns, acl.PermRead); !ok {
deny(w, id, ns, acl.PermRead)
return
}
}
nodes, err := store.NewNodeRepo(s.db).List(ctx)
if err != nil {
writeError(w, http.StatusInternalServerError, "failed to list nodes")
+18 -6
View File
@@ -13,13 +13,21 @@ import (
// isLoopback reports whether the address binds to a loopback interface
// (127.0.0.1, ::1, localhost). REQ-123: pprof must be loopback-only.
//
// An empty host (e.g. ":6060") binds ALL interfaces and is therefore
// treated as NON-loopback (F2: loopback-bypass fix). Only an explicit
// loopback IP or the "localhost" name is accepted.
func isLoopback(addr string) bool {
host, _, err := net.SplitHostPort(addr)
if err != nil {
host = addr
}
host = strings.TrimSpace(host)
if host == "" || host == "localhost" {
// F2: empty host (":6060") binds all interfaces — reject.
if host == "" {
return false
}
if host == "localhost" {
return true
}
ip := net.ParseIP(host)
@@ -29,18 +37,22 @@ func isLoopback(addr string) bool {
return false
}
// StartPprof starts the pprof HTTP server on addr. REQ-123: pprof is
// unauthenticated and MUST bind to a loopback interface only; this is a
// hard invariant (F2: the --pprof-allow-public override was a phantom flag
// that was never implemented and has been removed — non-loopback binds are
// always refused).
func StartPprof(addr string, log *slog.Logger) (*http.Server, error) {
if addr == "" {
return nil, nil
}
// REQ-123: pprof must bind to loopback only. Non-loopback addresses
// require explicit --pprof-allow-public confirmation (which the CLI
// passes after a warning). We refuse non-loopback here by default.
// REQ-123: pprof must bind to loopback only. This is a hard
// invariant; there is no public-bind override.
if !isLoopback(addr) {
log.Error("pprof refuses non-loopback bind",
slog.String("addr", addr),
slog.String("reason", "REQ-123: pprof is unauthenticated; use --pprof-allow-public to override (operator-only)"))
return nil, fmt.Errorf("pprof: refusing non-loopback bind %s (REQ-123; unauthenticated; use --pprof-allow-public)", addr)
slog.String("reason", "REQ-123: pprof is unauthenticated; loopback-only is a hard invariant"))
return nil, fmt.Errorf("pprof: refusing non-loopback bind %s (REQ-123; unauthenticated; loopback-only is a hard invariant)", addr)
}
mux := http.NewServeMux()
mux.HandleFunc("/debug/pprof/", pprof.Index)
+14
View File
@@ -7,6 +7,7 @@ import (
"net"
"net/http"
"path/filepath"
"strings"
"testing"
"time"
@@ -286,3 +287,16 @@ func TestStartPprof_LoopbackAccepted(t *testing.T) {
srv.Close()
}
}
// TestStartPprof_EmptyHostRefused verifies that an address with an empty
// host (e.g. ":6060"), which binds ALL interfaces, is refused as
// non-loopback (F2: loopback-bypass fix).
func TestStartPprof_EmptyHostRefused(t *testing.T) {
_, err := StartPprof(":6060", slog.Default())
if err == nil {
t.Error("StartPprof on \":6060\" should be refused (F2: empty host binds all interfaces)")
}
if err != nil && !strings.Contains(err.Error(), "non-loopback") {
t.Errorf("error should mention non-loopback, got: %v", err)
}
}
+27 -1
View File
@@ -49,6 +49,13 @@ type Server struct {
// /orca.v1.Dispatch/* (P02). Optional — nil if no Dispatcher
// was registered. P02 wires this via RegisterDispatch.
dispatch *DispatchHandlers
// acl is the access-control policy (P04, v0.13; C-44/C-45). When
// nil, no ACL enforcement is applied (legacy/compat for tests
// that construct a Server directly). Production wiring sets this
// via NewServer (Options.ACLEnforce) so handlers can call
// s.acl.Check before dispatching.
acl *aclPolicy
}
// Options configures a new Server.
@@ -63,6 +70,13 @@ type Options struct {
// The pprof listener is unauthenticated and operator-only; never
// expose it publicly (AD-024).
PprofAddr string
// ACLEnforce controls C-45 staged rollout. When false (the default
// for the first run after P04 wiring), ACL denials are LOGGED but
// NOT enforced — the request proceeds. When true, ACL denials
// return 403. The operator switches to true after verifying the
// bootstrap ACL grants the right identities.
ACLEnforce bool
}
// maxBodyBytes is the limit for request bodies on JSON-decoding
@@ -93,6 +107,7 @@ func NewServer(opts Options) *Server {
db: opts.DB,
log: opts.Log,
addr: opts.Addr,
acl: NewACLPolicy(opts.ACLEnforce, opts.Log),
}
s.httpServer = &http.Server{
Addr: opts.Addr,
@@ -127,6 +142,17 @@ func (s *Server) MarkNotReady() { s.ready.Store(false) }
// Ready reports the current readiness flag.
func (s *Server) Ready() bool { return s.ready.Load() }
// ACL returns the daemon's ACL enforcement policy (P04). Returns nil
// if no policy is configured (legacy/compat). Handlers use this to
// call Check before dispatching; tests use it to assert enforcement
// mode.
func (s *Server) ACL() *aclPolicy { return s.acl }
// SetACLPolicy replaces the ACL policy. Used by tests to inject a
// policy without going through NewServer. Production code should use
// NewServer with Options.ACLEnforce.
func (s *Server) SetACLPolicy(p *aclPolicy) { s.acl = p }
// mux builds the route table. Handlers are split across files:
// - health.go /healthz, /readyz, /v1/status
// - jobs_handler.go /v1/jobs/*
@@ -144,7 +170,7 @@ func (s *Server) mux() http.Handler {
mux.HandleFunc("/v1/nodes", s.handleNodesCollection)
mux.HandleFunc("/v1/tasks", s.handleTasksCollection)
if s.dispatch != nil {
s.dispatch.Mount(mux)
s.dispatch.mountWithACL(mux, s.acl)
}
return bodyLimitMiddleware(loggingMiddleware(s.log, mux))
}
+10
View File
@@ -7,6 +7,7 @@ import (
"strconv"
"time"
"git.cloudinit.dev/coreci/orca/internal/acl"
"git.cloudinit.dev/coreci/orca/internal/model"
"git.cloudinit.dev/coreci/orca/internal/store"
)
@@ -22,6 +23,15 @@ func (s *Server) handleTasksCollection(w http.ResponseWriter, r *http.Request) {
ctx, cancel := context.WithTimeout(r.Context(), 5*time.Second)
defer cancel()
// P04 ACL enforcement (C-44). Task list is a cluster-wide read.
ns := "_defaults"
if s.acl != nil {
if id, ok := s.acl.Check(r, ns, acl.PermRead); !ok {
deny(w, id, ns, acl.PermRead)
return
}
}
jobID := r.URL.Query().Get("job_id")
if jobID != "" {
if err := validateID(jobID); err != nil {
+4
View File
@@ -121,6 +121,10 @@ func CertCA() Check {
Description: "CA at ~/.orca with mode 0600/0644 (REQ-033)",
Run: func(_ context.Context) (Result, string) {
dir := certpaths.Dir()
caCert := certpaths.CACertPath()
if _, err := os.Stat(caCert); err != nil {
return ResultFail, fmt.Sprintf("CA cert missing: %v", err)
}
if err := security.EnforceFileModes(dir); err != nil {
return ResultFail, err.Error()
}
+75 -9
View File
@@ -13,7 +13,8 @@
// The ruleset defines:
//
// - table inet orca-ingress
// - set orca_trusted_probes (ipv4_addr interval, default 127.0.0.1)
// - set orca_trusted_probes_v4 (ipv4_addr interval, default 127.0.0.1)
// - set orca_trusted_probes_v6 (ipv6_addr interval, default ::1)
// - input chain (SYN-flood filter on :443)
// - prerouting chain (DNAT :443->127.0.0.1:8443, :80->127.0.0.1:8080)
// - forward chain (rate-limit meter on :443)
@@ -25,6 +26,7 @@ package emitter
import (
"errors"
"fmt"
"net"
"strings"
)
@@ -39,9 +41,11 @@ const nftConfigPath = "/etc/nftables.d/orca.nft"
// All fields have safe defaults so a zero-value config renders a
// working ruleset.
type NftClusterConfig struct {
// TrustedProbes is the list of source IPs exempt from the
// TrustedProbes is the list of source IPs/CIDRs exempt from the
// SYN-flood filter and rate-limit (monitoring probes, the orca
// lead itself). Defaults to [127.0.0.1, ::1].
// lead itself). Defaults to [127.0.0.1, ::1]. Each entry must
// parse as a valid IP or CIDR via net.ParseIP / net.ParseCIDR or
// RenderNftConfig returns an error (F9: ruleset injection guard).
TrustedProbes []string
// RateLimit is the per-source rate limit (packets/second) for the
// forward-chain meter on :443. Defaults to 100.
@@ -73,31 +77,79 @@ func (c NftClusterConfig) withDefaults() NftClusterConfig {
// `#!/usr/sbin/nft -f` shebang (so `nft -f` applies it and so a
// drift-check `nft -c -f` validates the syntax).
//
// Returns an error only when the config is internally inconsistent
// (e.g. a negative rate, which the defaults already prevent).
// Returns an error when the config is internally inconsistent (e.g. a
// negative rate, which the defaults already prevent) or when a
// TrustedProbes entry fails to parse as an IP or CIDR (F9: ruleset
// injection hardening — unvalidated entries are written directly into
// the nft ruleset and could inject arbitrary nft syntax).
func (NftEmitter) RenderNftConfig(clusterConfig NftClusterConfig) ([]File, error) {
if clusterConfig.RateLimit < 0 || clusterConfig.RateBurst < 0 {
return nil, errors.New("emitter/nft: rate/burst must be non-negative")
}
cfg := clusterConfig.withDefaults()
content := renderNftRuleset(cfg)
// F9: validate every TrustedProbes entry before rendering. An
// invalid entry is rejected with an error rather than written raw
// into the ruleset (which would allow nft-syntax injection).
v4, v6, err := partitionTrustedProbes(cfg.TrustedProbes)
if err != nil {
return nil, err
}
content := renderNftRuleset(cfg, v4, v6)
return []File{{Path: nftConfigPath, Content: content, Mode: "0644"}}, nil
}
// partitionTrustedProbes validates each entry as an IP or CIDR and
// partitions the list into IPv4 and IPv6 slices. Returns an error if
// any entry is neither a valid IP nor a valid CIDR (F9).
func partitionTrustedProbes(probes []string) (v4, v6 []string, err error) {
for _, p := range probes {
if p == "" {
return nil, nil, fmt.Errorf("emitter/nft: empty trusted probe entry (F9: ruleset injection guard)")
}
if ip := net.ParseIP(p); ip != nil {
if ip.To4() != nil {
v4 = append(v4, p)
} else {
v6 = append(v6, p)
}
continue
}
if _, _, cidrErr := net.ParseCIDR(p); cidrErr == nil {
// Determine address family from the CIDR prefix.
ip := net.ParseIP(strings.Split(p, "/")[0])
if ip != nil && ip.To4() != nil {
v4 = append(v4, p)
} else {
v6 = append(v6, p)
}
continue
}
return nil, nil, fmt.Errorf("emitter/nft: trusted probe %q is not a valid IP or CIDR (F9: ruleset injection guard)", p)
}
return v4, v6, nil
}
// renderNftRuleset builds the nft ruleset string. The shape is
// documented in the package comment; the exact lines are load-bearing
// for `orca doctor nft` (which greps the live table for them) and for
// `nft -c -f` (which parses the syntax).
func renderNftRuleset(cfg NftClusterConfig) string {
//
// F9: TrustedProbes are split into separate ipv4_addr and ipv6_addr
// sets (orca_trusted_probes_v4 / orca_trusted_probes_v6) because the
// prior single ipv4_addr set included ::1 (an IPv6 address), which is
// a type mismatch nft rejects.
func renderNftRuleset(cfg NftClusterConfig, v4, v6 []string) string {
var b strings.Builder
b.WriteString("#!/usr/sbin/nft -f\n\n")
b.WriteString("flush table inet orca-ingress\n\n")
b.WriteString("table inet orca-ingress {\n")
b.WriteString("\tset orca_trusted_probes {\n")
// F9: split IPv4 and IPv6 trusted probes into separate typed sets.
b.WriteString("\tset orca_trusted_probes_v4 {\n")
b.WriteString("\t\ttype ipv4_addr\n")
b.WriteString("\t\tflags interval\n")
b.WriteString("\t\telements = { ")
for i, p := range cfg.TrustedProbes {
for i, p := range v4 {
if i > 0 {
b.WriteString(", ")
}
@@ -105,6 +157,20 @@ func renderNftRuleset(cfg NftClusterConfig) string {
}
b.WriteString(" }\n")
b.WriteString("\t}\n\n")
b.WriteString("\tset orca_trusted_probes_v6 {\n")
b.WriteString("\t\ttype ipv6_addr\n")
b.WriteString("\t\tflags interval\n")
b.WriteString("\t\telements = { ")
for i, p := range v6 {
if i > 0 {
b.WriteString(", ")
}
b.WriteString(p)
}
b.WriteString(" }\n")
b.WriteString("\t}\n\n")
b.WriteString("\tchain input {\n")
b.WriteString("\t\ttype filter hook input priority filter; policy accept;\n")
b.WriteString("\t\tct state invalid drop\n")
+45
View File
@@ -90,3 +90,48 @@ func TestNftEmitter_ShebangFirst(t *testing.T) {
t.Errorf("shebang not first:\n%s", files[0].Content[:40])
}
}
// TestNftEmitter_RejectsInvalidTrustedProbe verifies that a TrustedProbes
// entry that is not a valid IP or CIDR is rejected (F9: ruleset injection
// guard). An unvalidated entry written raw into the ruleset could inject
// arbitrary nft syntax.
func TestNftEmitter_RejectsInvalidTrustedProbe(t *testing.T) {
bad := []string{
"not-an-ip",
"127.0.0.1; flush ruleset",
"$(whoami)",
"10.0.0.0/33", // invalid CIDR prefix
}
for _, b := range bad {
_, err := (NftEmitter{}).RenderNftConfig(NftClusterConfig{TrustedProbes: []string{"127.0.0.1", b}})
if err == nil {
t.Errorf("expected error for invalid trusted probe %q, got nil", b)
}
}
}
// TestNftEmitter_TrustedProbesSplitV4V6 verifies that IPv4 and IPv6
// probes are rendered into separate typed sets (F9: the prior single
// ipv4_addr set included ::1, an IPv6 address — a type mismatch).
func TestNftEmitter_TrustedProbesSplitV4V6(t *testing.T) {
files, err := (NftEmitter{}).RenderNftConfig(NftClusterConfig{TrustedProbes: []string{"10.0.0.5", "::1"}})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if !strings.Contains(c, "set orca_trusted_probes_v4") {
t.Errorf("missing v4 set:\n%s", c)
}
if !strings.Contains(c, "set orca_trusted_probes_v6") {
t.Errorf("missing v6 set:\n%s", c)
}
if !strings.Contains(c, "type ipv6_addr") {
t.Errorf("missing ipv6_addr type:\n%s", c)
}
if !strings.Contains(c, "10.0.0.5") {
t.Errorf("missing 10.0.0.5 in v4 set:\n%s", c)
}
if !strings.Contains(c, "::1") {
t.Errorf("missing ::1 in v6 set:\n%s", c)
}
}
+57
View File
@@ -166,6 +166,10 @@ func renderTaskUnit(spec *jobspec.WorkloadSpec, task *jobspec.TaskGroupTask, rt
b.WriteString(fmt.Sprintf("PartOf=%s\n", targetUnit))
b.WriteString("\n[Service]\n")
b.WriteString(fmt.Sprintf("ExecStart=%s\n", cmd))
for _, line := range renderRestartDirectives(spec) {
b.WriteString(line)
b.WriteString("\n")
}
for _, line := range (SocketEmitter{}).RenderSocketLines(spec) {
b.WriteString(line)
b.WriteString("\n")
@@ -201,6 +205,13 @@ func renderSystemdUnit(spec *jobspec.WorkloadSpec) string {
var b strings.Builder
b.WriteString("[Service]\n")
b.WriteString(fmt.Sprintf("ExecStart=%s\n", spec.Runtime.Command))
// Restart policy (REQ-152/T3): translate spec.Restart into the
// systemd Restart= / StartLimitBurst= / StartLimitIntervalSec=
// (or RestartSec=) directives. See renderRestartDirectives.
for _, line := range renderRestartDirectives(spec) {
b.WriteString(line)
b.WriteString("\n")
}
// Lifecycle: post_start → ExecStartPost (runs after start).
for _, cmd := range lifecyclePostStart(spec) {
b.WriteString(fmt.Sprintf("ExecStartPost=%s\n", cmd))
@@ -237,3 +248,49 @@ func lifecyclePreStop(spec *jobspec.WorkloadSpec) []string {
}
return spec.Lifecycle.PreStop
}
// renderRestartDirectives translates the spec.Restart block into the
// systemd [Service]/[Unit] restart directives (REQ-152/T3):
//
// - never → Restart=no (explicit; omitted when Restart is nil)
// - on-failure → Restart=on-failure + StartLimitBurst=<MaxRetries>
// - service → Restart=always + StartLimitBurst=<MaxRetries> (when
// MaxRetries > 0)
//
// The delay (a duration string like "5s") maps to StartLimitIntervalSec=
// when set; for the on-failure/service modes a non-empty delay also
// emits RestartSec=<delay> so systemd backs off between restart attempts.
// A nil Restart block produces no directives (the caller's default
// applies — for a [Service] with no Restart= that is Restart=no).
func renderRestartDirectives(spec *jobspec.WorkloadSpec) []string {
if spec.Restart == nil {
return nil
}
var out []string
switch spec.Restart.Mode {
case "never", "":
out = append(out, "Restart=no")
case "on-failure":
out = append(out, "Restart=on-failure")
if spec.Restart.MaxRetries > 0 {
out = append(out, fmt.Sprintf("StartLimitBurst=%d", spec.Restart.MaxRetries))
}
case "service":
out = append(out, "Restart=always")
if spec.Restart.MaxRetries > 0 {
out = append(out, fmt.Sprintf("StartLimitBurst=%d", spec.Restart.MaxRetries))
}
default:
// Unknown mode: emit Restart=no so the unit is explicit and
// systemd-analyze verify does not reject an unknown value.
out = append(out, "Restart=no")
}
if spec.Restart.Delay != "" {
switch spec.Restart.Mode {
case "on-failure", "service":
out = append(out, fmt.Sprintf("RestartSec=%s", spec.Restart.Delay))
out = append(out, fmt.Sprintf("StartLimitIntervalSec=%s", spec.Restart.Delay))
}
}
return out
}
+85
View File
@@ -0,0 +1,85 @@
package emitter
import (
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
func TestSystemdEmitter_RestartNever(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "one",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/true"},
Restart: &jobspec.RestartBlock{Mode: "never"},
}
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n"})
if err != nil {
t.Fatalf("Render: %v", err)
}
if !strings.Contains(files[0].Content, "Restart=no") {
t.Errorf("missing Restart=no:\n%s", files[0].Content)
}
}
func TestSystemdEmitter_RestartOnFailure(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "retry",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/true"},
Restart: &jobspec.RestartBlock{Mode: "on-failure", MaxRetries: 3, Delay: "5s"},
}
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n"})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if !strings.Contains(c, "Restart=on-failure") {
t.Errorf("missing Restart=on-failure:\n%s", c)
}
if !strings.Contains(c, "StartLimitBurst=3") {
t.Errorf("missing StartLimitBurst=3:\n%s", c)
}
if !strings.Contains(c, "RestartSec=5s") {
t.Errorf("missing RestartSec=5s:\n%s", c)
}
if !strings.Contains(c, "StartLimitIntervalSec=5s") {
t.Errorf("missing StartLimitIntervalSec=5s:\n%s", c)
}
}
func TestSystemdEmitter_RestartService(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/httpd"},
Restart: &jobspec.RestartBlock{Mode: "service", MaxRetries: 5, Delay: "10s"},
}
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n"})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if !strings.Contains(c, "Restart=always") {
t.Errorf("missing Restart=always:\n%s", c)
}
if !strings.Contains(c, "StartLimitBurst=5") {
t.Errorf("missing StartLimitBurst=5:\n%s", c)
}
}
func TestSystemdEmitter_RestartNilOmitted(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "norest",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/true"},
}
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n"})
if err != nil {
t.Fatalf("Render: %v", err)
}
if strings.Contains(files[0].Content, "Restart=") {
t.Errorf("nil Restart should omit Restart= line:\n%s", files[0].Content)
}
}
+33
View File
@@ -0,0 +1,33 @@
// Package engine — actor.go provides the context key + helper for
// threading the audit actor (OIDC sub or SPIFFE SVID) through the
// engine layer (P04, T5; C-44). Previously the registry hardcoded
// "cli" as the actor; this lets CLI commands inject the verified
// operator identity via context so audit entries attribute actions
// to the real human/operator.
package engine
import "context"
// actorCtxKey is the context key for the audit actor.
type actorCtxKey struct{}
// WithActor returns a context carrying the audit actor. The CLI
// calls this in PersistentPreRun after resolving the OIDC sub from
// the credentials file. When the context carries no actor, the
// registry falls back to "cli" (legacy).
func WithActor(ctx context.Context, actor string) context.Context {
if actor == "" {
return ctx
}
return context.WithValue(ctx, actorCtxKey{}, actor)
}
// ActorFromCtx returns the audit actor from the context, or "cli"
// when no actor is set (legacy fallback for paths that haven't been
// wired yet).
func ActorFromCtx(ctx context.Context) string {
if v, ok := ctx.Value(actorCtxKey{}).(string); ok && v != "" {
return v
}
return "cli"
}
+25 -8
View File
@@ -98,24 +98,39 @@ type TaskSpec struct {
}
func (e *Executor) Run(ctx context.Context, job *model.Job, specs []TaskSpec) error {
e.mu.Lock()
defer e.mu.Unlock()
// REQ-156 / P07 T6: the mutex previously guarded the ENTIRE job
// (insert + status transitions + task execution + wait). That
// serialized unrelated jobs against each other and held the lock
// across long-running child processes, blocking concurrent
// Submit/Status/Run callers. The mutex is now scoped ONLY to the
// DB inserts/updates (the part that must be serialized against
// the single-writer SQLite connection pool — see store.Open
// SetMaxOpenConns(1)). The task goroutines spawned below do not
// hold e.mu; they share the per-job failure counter via a local
// sync.Mutex.
// Insert the job first so tasks can reference it via foreign key.
// Insert the job + flip to Running under the lock (serializes
// the DB writes; the underlying SQLite busy_timeout(5000) +
// SetMaxOpenConns(1) handles contention).
e.mu.Lock()
if err := e.jobs.Insert(ctx, job); err != nil {
e.mu.Unlock()
return err
}
if err := e.jobs.UpdateStatus(ctx, job.ID, model.JobStatusRunning, 0); err != nil {
e.mu.Unlock()
return err
}
e.mu.Unlock()
// Task execution runs WITHOUT e.mu — concurrent jobs (and
// concurrent Submit/Status callers) are no longer blocked by a
// long-running child process.
var (
wg sync.WaitGroup
failedCount int
exitCode int
mu sync.Mutex
)
for _, ts := range specs {
wg.Add(1)
go func(ts TaskSpec) {
@@ -133,14 +148,16 @@ func (e *Executor) Run(ctx context.Context, job *model.Job, specs []TaskSpec) er
}
wg.Wait()
// Final status transition under the lock (the DB write is the
// only thing that needs serialization).
e.mu.Lock()
defer e.mu.Unlock()
if failedCount > 0 {
exitCode = 1
if err := e.jobs.UpdateStatus(ctx, job.ID, model.JobStatusFailed, exitCode); err != nil {
if err := e.jobs.UpdateStatus(ctx, job.ID, model.JobStatusFailed, 1); err != nil {
return err
}
return fmt.Errorf("%d/%d tasks failed", failedCount, len(specs))
}
if err := e.jobs.UpdateStatus(ctx, job.ID, model.JobStatusComplete, 0); err != nil {
return err
}
+7 -7
View File
@@ -24,13 +24,13 @@ func NewNodeRegistry(repo *store.NodeRepo, audit *Audit, log *slog.Logger) *Node
func (r *NodeRegistry) Join(ctx context.Context, n *model.Node) error {
if err := r.repo.Insert(ctx, n); err != nil {
r.audit.Record(ctx, "cli", "node.join", n.ID, "failure", err, map[string]any{
r.audit.Record(ctx, ActorFromCtx(ctx), "node.join", n.ID, "failure", err, map[string]any{
"name": n.Name,
"address": n.Address,
})
return fmt.Errorf("join node: %w", err)
}
r.audit.Record(ctx, "cli", "node.join", n.ID, "success", nil, map[string]any{
r.audit.Record(ctx, ActorFromCtx(ctx), "node.join", n.ID, "success", nil, map[string]any{
"name": n.Name,
"address": n.Address,
})
@@ -43,20 +43,20 @@ func (r *NodeRegistry) Join(ctx context.Context, n *model.Node) error {
func (r *NodeRegistry) Leave(ctx context.Context, id string) error {
if err := r.repo.UpdateState(ctx, id, model.NodeStateLeft); err != nil {
r.audit.Record(ctx, "cli", "node.leave", id, "failure", err, nil)
r.audit.Record(ctx, ActorFromCtx(ctx), "node.leave", id, "failure", err, nil)
return fmt.Errorf("leave node: %w", err)
}
r.audit.Record(ctx, "cli", "node.leave", id, "success", nil, nil)
r.audit.Record(ctx, ActorFromCtx(ctx), "node.leave", id, "success", nil, nil)
r.log.Info("node left", slog.String("node_id", id))
return nil
}
func (r *NodeRegistry) Forget(ctx context.Context, id string) error {
if err := r.repo.Delete(ctx, id); err != nil {
r.audit.Record(ctx, "cli", "node.forget", id, "failure", err, nil)
r.audit.Record(ctx, ActorFromCtx(ctx), "node.forget", id, "failure", err, nil)
return fmt.Errorf("forget node: %w", err)
}
r.audit.Record(ctx, "cli", "node.forget", id, "success", nil, nil)
r.audit.Record(ctx, ActorFromCtx(ctx), "node.forget", id, "success", nil, nil)
r.log.Info("node removed from registry", slog.String("node_id", id))
return nil
}
@@ -75,7 +75,7 @@ func (r *NodeRegistry) Get(ctx context.Context, id string) (*model.Node, error)
// cli package does not need to reach into the repo directly.
func (r *NodeRegistry) SetNodeState(ctx context.Context, id, state string) error {
if err := r.repo.SetNodeState(ctx, id, state); err != nil {
r.audit.Record(ctx, "cli", "node.set_state", id, "failure", err, map[string]any{"state": state})
r.audit.Record(ctx, ActorFromCtx(ctx), "node.set_state", id, "failure", err, map[string]any{"state": state})
return fmt.Errorf("set node state: %w", err)
}
r.log.Info("node state set", slog.String("node_id", id), slog.String("state", state))
+60
View File
@@ -0,0 +1,60 @@
// Package identity — authtoken.go provides the ORCA_OIDC_TOKEN
// validation helper used by the SSH-push applier and the txn apply
// path (P04, v0.13; C-44). Both paths validate the env-var token
// against the issuer's JWKS before applying any state change.
//
// The token is read from $ORCA_OIDC_TOKEN. The issuer + client ID
// come from the OIDC config (oidc.issuer, oidc.client_id). If the
// token is missing or invalid, the apply is refused. The verified
// claims (sub + groups) are returned so the caller can thread them
// into the audit actor field (T5) and the ACL check (T3/T4).
package identity
import (
"context"
"fmt"
"os"
)
// EnvOIDCToken is the environment variable holding the OIDC ID token
// for the SSH-push / txn apply paths (R-021: the IdP issues the
// token; Orca never issues its own).
const EnvOIDCToken = "ORCA_OIDC_TOKEN"
// VerifyOperatorToken reads $ORCA_OIDC_TOKEN and verifies it against
// the issuer's JWKS. Returns the verified claims (sub, groups) on
// success. Returns an error if the token is missing, expired, or
// fails signature verification.
//
// The issuer + clientID come from the OIDC config block. When issuer
// is empty, the function returns an error — the apply path requires
// an OIDC issuer to be configured.
func VerifyOperatorToken(ctx context.Context, issuer, clientID string) (*IDTokenClaims, error) {
raw := os.Getenv(EnvOIDCToken)
if raw == "" {
return nil, fmt.Errorf("identity: %s env var is not set (operator OIDC token required for apply)", EnvOIDCToken)
}
if issuer == "" {
return nil, fmt.Errorf("identity: oidc.issuer is not configured (required to verify %s)", EnvOIDCToken)
}
if clientID == "" {
clientID = "orca-cli"
}
claims, err := VerifyIDTokenStatic(ctx, issuer, clientID, raw)
if err != nil {
return nil, fmt.Errorf("identity: verify %s: %w", EnvOIDCToken, err)
}
return claims, nil
}
// OperatorActor renders the verified operator identity for the audit
// `actor` field. The convention is "oidc:<sub>" so audit entries can
// be filtered by human operator. Falls back to "oidc:unknown" when
// claims are nil (e.g. when the caller could not verify the token but
// still wants to record an audit entry).
func OperatorActor(claims *IDTokenClaims) string {
if claims == nil || claims.Subject == "" {
return "oidc:unknown"
}
return "oidc:" + claims.Subject
}
+18 -12
View File
@@ -29,6 +29,8 @@ import (
"github.com/coreos/go-oidc/v3/oidc"
"golang.org/x/oauth2"
"git.cloudinit.dev/coreci/orca/internal/security"
)
// OIDCConfig holds the OIDC client configuration. It is loaded from
@@ -113,7 +115,13 @@ func SaveCredentials(c *Credentials) error {
if err != nil {
return fmt.Errorf("oidc: marshal: %w", err)
}
return writeAtomic0600(path, data)
// REQ-156 / P07 T9: use the canonical security.WriteAtomic (temp
// + chmod + fsync + rename) instead of the local writeAtomic0600
// (which did temp + chmod + rename with NO fsync - a crash before
// rename could leave a partially-written tmp file that rename
// would then promote, or the rename could land before the data
// reached durable storage).
return security.WriteAtomic(path, 0o600, data)
}
// ClearCredentials removes the stored credentials (logout).
@@ -128,16 +136,6 @@ func ClearCredentials() error {
return nil
}
// writeAtomic0600 writes data to path atomically at mode 0600
// (temp + chmod + rename).
func writeAtomic0600(path string, data []byte) error {
tmp := path + ".tmp"
if err := os.WriteFile(tmp, data, 0o600); err != nil {
return fmt.Errorf("oidc: write tmp: %w", err)
}
return os.Rename(tmp, path)
}
// OIDCClient wraps the OIDC provider + oauth2 config for the auth flow.
type OIDCClient struct {
provider *oidc.Provider
@@ -237,7 +235,15 @@ func (c *OIDCClient) Login(ctx context.Context, openBrowser func(string) error)
err error
}
resultCh := make(chan result, 1)
srv := &http.Server{}
// REQ-157 / P08 T8: set ReadHeaderTimeout so a slowloris-style
// peer cannot hold the callback server open indefinitely. The
// callback is short-lived (one request then Shutdown), but the
// default zero ReadHeaderTimeout means an attacker who reaches the
// loopback port during the brief auth window could stall the
// handshake. 5s is generous for a loopback redirect.
srv := &http.Server{
ReadHeaderTimeout: 5 * time.Second,
}
srv.Handler = http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/callback" {
http.NotFound(w, r)
+44 -1
View File
@@ -387,7 +387,13 @@ func findClosingDelimiter(rest string) int {
// not supported — by design, to avoid adding a YAML dependency for this
// small surface.
func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
spec := &WorkloadSpec{Count: 1}
// Count defaults to 1 for Job/Service and 0 for DaemonSet. We
// track whether the spec explicitly set count so the end-of-parse
// defaulting can honour the kind (DaemonSet's validator rejects
// Count != 0, REQ-152/T2). countSet flips true on the first
// `count:` key seen.
spec := &WorkloadSpec{}
var countSet bool
lines := strings.Split(block, "\n")
type section int
@@ -408,6 +414,7 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
secTasks
secTaskEnv
secTaskRuntime
secSchedule
)
cur := secNone
var curPort *PortSpec
@@ -490,6 +497,7 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
case "count":
if n, err := strconv.Atoi(strings.TrimSpace(unquote(val))); err == nil {
spec.Count = n
countSet = true
} else {
return nil, fmt.Errorf("parse markdown: line %d: count: %v", lineNo+1, err)
}
@@ -551,6 +559,15 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
} else {
cur = secAffinity
}
case "schedule":
spec.Schedule = &ScheduleBlock{}
if strings.TrimSpace(val) != "" {
// Inline value (unusual); ignore — schedule is a block.
}
cur = secSchedule
case "timeout":
spec.Timeout = unquote(val)
cur = secNone
case "tasks":
cur = secTasks
taskIndent = -1
@@ -775,6 +792,20 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
spec.Constraints = append(spec.Constraints, unquote(item))
}
}
case secSchedule:
if spec.Schedule == nil {
spec.Schedule = &ScheduleBlock{}
}
key, val, ok := splitKV(trimmed)
if !ok {
continue
}
switch key {
case "mode":
spec.Schedule.Mode = unquote(val)
case "cron":
spec.Schedule.Cron = unquote(val)
}
case secTasks:
// Tasks is a list of task objects. A `- ` at the list
// indent opens a new task; deeper-indented lines belong
@@ -888,6 +919,18 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
flushVol()
flushAffinity()
flushTask()
// Count defaulting: 1 for Job/Service, 0 for DaemonSet. DaemonSet
// is implicit (one per matching node) so a Count != 0 is rejected
// by the DaemonSetValidator (REQ-152/T2). Only default when the
// spec did not explicitly set count.
if !countSet {
switch spec.Kind {
case "DaemonSet":
spec.Count = 0
default:
spec.Count = 1
}
}
return spec, nil
}
+136
View File
@@ -0,0 +1,136 @@
package jobspec
import (
"testing"
)
func TestREQ152_ScheduleTimeoutDaemonSet(t *testing.T) {
input := "---\n" +
"kind: DaemonSet\n" +
"name: logs\n" +
"schedule:\n" +
" mode: every-node\n" +
" cron: \"*/5 * * * *\"\n" +
"timeout: 30s\n" +
"restart:\n" +
" mode: service\n" +
"runtime:\n" +
" one_of: process\n" +
" command: /bin/true\n" +
"---\nbody\n"
spec, err := ParseMarkdown([]byte(input))
if err != nil {
t.Fatalf("ParseMarkdown: %v", err)
}
if spec.Count != 0 {
t.Errorf("DaemonSet Count = %d, want 0 (no default)", spec.Count)
}
if spec.Schedule == nil {
t.Fatal("Schedule is nil")
}
if spec.Schedule.Mode != "every-node" {
t.Errorf("Schedule.Mode = %q, want every-node", spec.Schedule.Mode)
}
if spec.Schedule.Cron != "*/5 * * * *" {
t.Errorf("Schedule.Cron = %q, want */5 * * * *", spec.Schedule.Cron)
}
if spec.Timeout != "30s" {
t.Errorf("Timeout = %q, want 30s", spec.Timeout)
}
}
func TestREQ152_JobScheduleTimeout(t *testing.T) {
input := "---\n" +
"kind: Job\n" +
"name: nightly\n" +
"schedule:\n" +
" cron: \"0 2 * * *\"\n" +
"timeout: 1h\n" +
"runtime:\n" +
" one_of: process\n" +
" command: /bin/true\n" +
"---\nbody\n"
spec, err := ParseMarkdown([]byte(input))
if err != nil {
t.Fatalf("ParseMarkdown: %v", err)
}
if spec.Count != 1 {
t.Errorf("Job Count = %d, want 1 (default)", spec.Count)
}
if spec.Schedule == nil || spec.Schedule.Cron != "0 2 * * *" {
t.Errorf("Schedule.Cron = %+v, want 0 2 * * *", spec.Schedule)
}
if spec.Timeout != "1h" {
t.Errorf("Timeout = %q, want 1h", spec.Timeout)
}
}
// TestREQ152_DaemonSetPassesLint verifies a DaemonSet spec with a
// schedule block, restart, and runtime parses AND validates cleanly
// under the schema (T11). DaemonSet must NOT default Count to 1.
func TestREQ152_DaemonSetPassesLint(t *testing.T) {
input := "---\n" +
"kind: DaemonSet\n" +
"name: log-shipper\n" +
"schedule:\n" +
" mode: every-node\n" +
"restart:\n" +
" mode: service\n" +
" max_retries: 5\n" +
" delay: 5s\n" +
"runtime:\n" +
" one_of: process\n" +
" command: /usr/local/bin/log-shipper\n" +
"---\n# Log shipper\n\nRuns on every node.\n"
spec, err := ParseMarkdown([]byte(input))
if err != nil {
t.Fatalf("ParseMarkdown: %v", err)
}
if spec.Count != 0 {
t.Errorf("DaemonSet Count = %d, want 0", spec.Count)
}
if spec.Schedule == nil || spec.Schedule.Mode != "every-node" {
t.Errorf("Schedule.Mode = %+v, want every-node", spec.Schedule)
}
if spec.Restart == nil || spec.Restart.Mode != "service" {
t.Errorf("Restart.Mode = %+v, want service", spec.Restart)
}
}
// TestREQ152_TimeoutEnforced verifies the timeout field is parsed and
// stored on the WorkloadSpec (T12).
func TestREQ152_TimeoutEnforced(t *testing.T) {
cases := []struct {
timeout string
want string
}{
{"30s", "30s"},
{"5m", "5m"},
{"1h30m", "1h30m"},
{"900s", "900s"},
}
for _, c := range cases {
input := "---\nkind: Job\nname: t\ntimeout: " + c.timeout + "\nruntime:\n one_of: process\n command: /bin/true\n---\nbody\n"
spec, err := ParseMarkdown([]byte(input))
if err != nil {
t.Fatalf("ParseMarkdown(%q): %v", c.timeout, err)
}
if spec.Timeout != c.want {
t.Errorf("Timeout = %q, want %q", spec.Timeout, c.want)
}
}
}
// TestREQ152_DaemonSetExplicitCountRejected verifies that an explicit
// count on a DaemonSet is preserved (parser does not override it) so
// the validator can reject it.
func TestREQ152_DaemonSetExplicitCountPreserved(t *testing.T) {
input := "---\nkind: DaemonSet\nname: d\ncount: 3\nschedule:\n mode: every-node\nrestart:\n mode: service\nruntime:\n one_of: process\n command: /bin/true\n---\nbody\n"
spec, err := ParseMarkdown([]byte(input))
if err != nil {
t.Fatalf("ParseMarkdown: %v", err)
}
if spec.Count != 3 {
t.Errorf("DaemonSet explicit Count = %d, want 3 (preserved, not defaulted)", spec.Count)
}
}
+91 -16
View File
@@ -29,6 +29,7 @@ import (
"log/slog"
"net"
"os"
"regexp"
"strings"
"time"
@@ -118,6 +119,18 @@ func BootstrapProxmox(ctx context.Context, opts Options) (*Result, error) {
if opts.SSHPort == 0 {
opts.SSHPort = DefaultSSHPort
}
// F10: validate ProxmoxUser and ProxmoxRole before they are
// interpolated into sudoers content, file paths, and shell commands
// (useradd, pveum). An attacker-controlled value could inject shell
// metacharacters or path traversal. Allowlist: lowercase letter or
// underscore start, followed by lowercase alphanumerics, underscore,
// or hyphen; max 32 chars.
if !validProxmoxName(opts.ProxmoxUser) {
return nil, fmt.Errorf("proxmox bootstrap: invalid ProxmoxUser %q (allowed: ^[a-zA-Z_][a-zA-Z0-9_-]{0,31}$)", opts.ProxmoxUser)
}
if !validProxmoxName(opts.ProxmoxRole) {
return nil, fmt.Errorf("proxmox bootstrap: invalid ProxmoxRole %q (allowed: ^[a-zA-Z_][a-zA-Z0-9_-]{0,31}$)", opts.ProxmoxRole)
}
log := opts.Logger
if log == nil {
log = slog.Default()
@@ -137,7 +150,7 @@ func BootstrapProxmox(ctx context.Context, opts Options) (*Result, error) {
// the v0.6 ship-defect where knownhosts.New returned KeyError{Want:[]}
// on first connect WITHOUT writing the captured key, so the first
// `orca node join --type proxmox` always failed.
sshAddr := fmt.Sprintf("%s:%d", opts.Host, opts.SSHPort)
sshAddr := net.JoinHostPort(opts.Host, fmt.Sprintf("%d", opts.SSHPort))
var capturedHostKey ssh.PublicKey
var hostKeyCallback ssh.HostKeyCallback
if opts.HostKeyFingerprint != "" {
@@ -283,8 +296,29 @@ func pinnedHostKeyCallback(expectedSHA256Base64 string, capturedKey *ssh.PublicK
//
// Exported so the doctor proxmox probe (T02.9) can reuse the same
// capture-fix wrapper for parity (GRILL condition #2).
//
// REQ-157 / P08 T4: TOFUHostKeyCallback now delegates to
// TOFUHostKeyCallbackPath with the v0.8 flat layout
// (certpaths.KnownHostsPath()). The path-accepting variant lets the
// sshpush transport pass its stored known_hosts field (the v0.9
// paths.KnownHostsPath() location) instead of always reading the v0.8
// flat layout — fixing the bug where the dial() flock field was stored
// but never read.
func TOFUHostKeyCallback(addr string, capturedKey *ssh.PublicKey) (ssh.HostKeyCallback, error) {
cb, err := knownhosts.New(certpaths.KnownHostsPath())
return TOFUHostKeyCallbackPath(certpaths.KnownHostsPath(), addr, capturedKey)
}
// TOFUHostKeyCallbackPath is the path-accepting variant. knownHostsPath
// is the known_hosts file to verify against and capture new keys into;
// it MUST be flock-protected on capture (security.Flock). When
// knownHostsPath is empty, falls back to certpaths.KnownHostsPath()
// (the v0.8 flat layout) for backward compatibility with callers that
// relied on the implicit default.
func TOFUHostKeyCallbackPath(knownHostsPath, addr string, capturedKey *ssh.PublicKey) (ssh.HostKeyCallback, error) {
if knownHostsPath == "" {
knownHostsPath = certpaths.KnownHostsPath()
}
cb, err := knownhosts.New(knownHostsPath)
if err != nil {
return nil, err
}
@@ -299,13 +333,12 @@ func TOFUHostKeyCallback(addr string, capturedKey *ssh.PublicKey) (ssh.HostKeyCa
var keyErr *knownhosts.KeyError
if errors.As(err, &keyErr) && len(keyErr.Want) == 0 {
line := knownhosts.Line([]string{knownhosts.Normalize(addr)}, key)
path := certpaths.KnownHostsPath()
release, lockErr := security.Flock(path)
release, lockErr := security.Flock(knownHostsPath)
if lockErr != nil {
return fmt.Errorf("tofu lock known_hosts: %w", lockErr)
}
defer release()
existing, readErr := os.ReadFile(path)
existing, readErr := os.ReadFile(knownHostsPath)
if readErr != nil && !os.IsNotExist(readErr) {
return fmt.Errorf("tofu read known_hosts: %w", readErr)
}
@@ -313,7 +346,7 @@ func TOFUHostKeyCallback(addr string, capturedKey *ssh.PublicKey) (ssh.HostKeyCa
existing = append(existing, '\n')
}
updated := append(existing, []byte(line)...)
if writeErr := security.WriteAtomic(path, 0o600, updated); writeErr != nil {
if writeErr := security.WriteAtomic(knownHostsPath, 0o600, updated); writeErr != nil {
return fmt.Errorf("tofu write known_hosts: %w", writeErr)
}
if capturedKey != nil {
@@ -392,7 +425,8 @@ func deployPubKey(user, pubLine string) error {
// createLinuxUser creates the orca system user if it doesn't already
// exist. Idempotent: `id -u` check before `useradd`.
func createLinuxUser(user string) error {
cmd := fmt.Sprintf("id -u %s 2>/dev/null || useradd -r -s /usr/sbin/nologin %s", user, user)
// F10c: shellQuote the user (validated upstream, but defense-in-depth).
cmd := fmt.Sprintf("id -u %s 2>/dev/null || useradd -r -s /usr/sbin/nologin %s", shellQuote(user), shellQuote(user))
if _, err := runRemote(cmd); err != nil {
return err
}
@@ -402,9 +436,10 @@ func createLinuxUser(user string) error {
// createPVERole creates the OrcaOperator PVE role if it doesn't exist.
// Idempotent: probes `pveum role list` before `pveum role add`.
func createPVERole(role string) error {
// F10c: shellQuote the role (validated upstream, but defense-in-depth).
cmd := fmt.Sprintf(
"pveum role list 2>/dev/null | grep -q '^%s' || pveum role add %s --privs '%s'",
role, role, OrcaOperatorPrivileges,
shellQuote(role), shellQuote(role), OrcaOperatorPrivileges,
)
if _, err := runRemote(cmd); err != nil {
return err
@@ -417,9 +452,11 @@ func createPVERole(role string) error {
// Uses @pam realm (AD-019) since orca creates a Linux system user.
func createPVEUser(user string) error {
pveUserID := user + "@pam"
// F10c: shellQuote the PVE user id (validated upstream, but
// defense-in-depth).
cmd := fmt.Sprintf(
"pveum user list 2>/dev/null | grep -q '%s' || pveum user add %s -comment 'Orca automation user'",
pveUserID, pveUserID,
"pveum user list 2>/dev/null | grep -q %s || pveum user add %s -comment 'Orca automation user'",
shellQuote(pveUserID), shellQuote(pveUserID),
)
if _, err := runRemote(cmd); err != nil {
return err
@@ -431,7 +468,9 @@ func createPVEUser(user string) error {
// (cluster-wide). `pveum acl modify` is idempotent (creates or updates).
func assignPVEACL(user, role string) error {
pveUserID := user + "@pam"
cmd := fmt.Sprintf("pveum acl modify / -user %s -role %s", pveUserID, role)
// F10c: shellQuote the PVE user id and role (validated upstream,
// but defense-in-depth).
cmd := fmt.Sprintf("pveum acl modify / -user %s -role %s", shellQuote(pveUserID), shellQuote(role))
if _, err := runRemote(cmd); err != nil {
return err
}
@@ -454,13 +493,20 @@ func sudoersContent(user string) string {
`, user, user)
}
// sudoersPath is the fixed on-peer path for the orca sudoers drop-in.
// F10b: the file is always written here regardless of the configured
// ProxmoxUser name, so a crafted username cannot redirect the sudoers
// drop-in to an arbitrary path.
const sudoersPath = "/etc/sudoers.d/orca"
// writeSudoers writes the /etc/sudoers.d/orca file on the remote host
// with mode 0440. Uses a heredoc via cat to avoid quoting issues.
// with mode 0440. Uses a heredoc via cat to avoid quoting issues. F10b:
// the path is fixed (sudoersPath) regardless of the configured username.
func writeSudoers(user string) error {
content := sudoersContent(user)
// Write via cat heredoc, then chmod 0440.
cmd := fmt.Sprintf("cat > /etc/sudoers.d/%s <<'ORCA_SUDOERS_EOF'\n%s\nORCA_SUDOERS_EOF\nchmod 0440 /etc/sudoers.d/%s",
user, content, user)
// Write via cat heredoc to the fixed path, then chmod 0440.
cmd := fmt.Sprintf("cat > %s <<'ORCA_SUDOERS_EOF'\n%s\nORCA_SUDOERS_EOF\nchmod 0440 %s",
sudoersPath, content, sudoersPath)
if _, err := runRemote(cmd); err != nil {
return err
}
@@ -470,8 +516,14 @@ func writeSudoers(user string) error {
// validateSudoers runs `visudo -cf` on the sudoers file. Aborts the
// bootstrap if validation fails (prevents a broken sudoers from
// locking the orca user out of sudo).
// validateSudoers runs `visudo -cf` on the sudoers file. F10d: it
// validates the actual file that writeSudoers wrote (sudoersPath,
// /etc/sudoers.d/orca), which is now a fixed path — the prior version
// hardcoded /etc/sudoers.d/orca while writeSudoers wrote to
// /etc/sudoers.d/<ProxmoxUser>, so a custom username would validate the
// wrong file.
func validateSudoers() error {
cmd := "visudo -cf /etc/sudoers.d/orca"
cmd := fmt.Sprintf("visudo -cf %s", sudoersPath)
out, err := runRemote(cmd)
if err != nil {
return fmt.Errorf("visudo validation failed: %w (output: %s)", err, strings.TrimSpace(string(out)))
@@ -482,6 +534,29 @@ func validateSudoers() error {
return nil
}
// proxmoxNameRe is the allowlist for ProxmoxUser and ProxmoxRole values
// that are interpolated into sudoers content, file paths, and shell
// commands (F10a). Letter or underscore start, followed by
// alphanumerics, underscore, or hyphen; max 32 chars. Uppercase is
// permitted (DefaultProxmoxRole is "OrcaOperator"); shell
// metacharacters (spaces, ;, $, backticks, etc.) are blocked.
var proxmoxNameRe = regexp.MustCompile(`^[a-zA-Z_][a-zA-Z0-9_-]{0,31}$`)
// validProxmoxName reports whether s is a safe ProxmoxUser or ProxmoxRole
// value (F10a injection guard).
func validProxmoxName(s string) bool {
return proxmoxNameRe.MatchString(s)
}
// shellQuote single-quotes a string for safe shell interpolation over
// the SSH exec session. It escapes embedded single-quotes via the
// standard ”' idiom (POSIX shell). F10c: hardens pveum/useradd commands
// against metacharacter injection (the validated allowlist is
// defense-in-depth on top of this).
func shellQuote(s string) string {
return "'" + strings.ReplaceAll(s, "'", "'\\''") + "'"
}
// ResetHostKey removes all known_hosts entries for the given host from
// certpaths.KnownHostsPath() (REQ-059, D-046, AD-029). It rewrites the
// file atomically via security.WriteAtomic. LOCAL ONLY — it does NOT
+36
View File
@@ -96,6 +96,42 @@ func TestBootstrapProxmox_Validation(t *testing.T) {
}
}
// TestBootstrapProxmox_RejectsInvalidProxmoxUser verifies that a
// ProxmoxUser containing shell metacharacters is rejected before any
// SSH dial (F10a: sudoers/shell injection guard).
func TestBootstrapProxmox_RejectsInvalidProxmoxUser(t *testing.T) {
ctx := context.Background()
bad := []string{"orca; rm -rf /", "orca$(whoami)", "orca`id`", "orca user", "1orca"}
for _, b := range bad {
_, err := BootstrapProxmox(ctx, Options{Host: "10.0.0.1", SSHKeyPath: certpaths.SSHKeyPath(), ProxmoxUser: b})
if err == nil {
t.Errorf("expected error for invalid ProxmoxUser %q, got nil", b)
continue
}
if !strings.Contains(err.Error(), "invalid ProxmoxUser") {
t.Errorf("error should mention invalid ProxmoxUser for %q, got: %v", b, err)
}
}
}
// TestBootstrapProxmox_RejectsInvalidProxmoxRole verifies that a
// ProxmoxRole containing shell metacharacters is rejected before any
// SSH dial (F10a: sudoers/shell injection guard).
func TestBootstrapProxmox_RejectsInvalidProxmoxRole(t *testing.T) {
ctx := context.Background()
bad := []string{"role; flush", "role$(id)", "role`whoami`", "role name", "1role"}
for _, b := range bad {
_, err := BootstrapProxmox(ctx, Options{Host: "10.0.0.1", SSHKeyPath: certpaths.SSHKeyPath(), ProxmoxRole: b})
if err == nil {
t.Errorf("expected error for invalid ProxmoxRole %q, got nil", b)
continue
}
if !strings.Contains(err.Error(), "invalid ProxmoxRole") {
t.Errorf("error should mention invalid ProxmoxRole for %q, got: %v", b, err)
}
}
}
func TestDefaultOptions(t *testing.T) {
if DefaultProxmoxUser != "orca" {
t.Errorf("DefaultProxmoxUser = %q, want orca", DefaultProxmoxUser)
+30
View File
@@ -0,0 +1,30 @@
package proxmox
import (
"fmt"
"net"
"testing"
)
// TestREQ157_IPv6JoinHostPort verifies that the proxmox SSH dial
// address is correctly bracketed for IPv6 hosts (REQ-157 / P08 T5/T11).
func TestREQ157_IPv6JoinHostPort(t *testing.T) {
tests := []struct {
host string
port int
want string
}{
{"192.168.1.1", 22, "192.168.1.1:22"},
{"::1", 22, "[::1]:22"},
{"fe80::1", 2222, "[fe80::1]:2222"},
{"2001:db8::1", 22, "[2001:db8::1]:22"},
}
for _, tt := range tests {
t.Run(tt.host, func(t *testing.T) {
got := net.JoinHostPort(tt.host, fmt.Sprintf("%d", tt.port))
if got != tt.want {
t.Errorf("JoinHostPort(%s, %d) = %q, want %q", tt.host, tt.port, got, tt.want)
}
})
}
}
+1 -1
View File
@@ -64,7 +64,7 @@ func (p *PodmanRuntime) Start(ctx context.Context, alloc *Alloc) (int, error) {
}
cmdStr, _ := commandFor(alloc)
name := containerName(alloc)
cmd := fmt.Sprintf("podman run -d --name %s %q %s", shellQuote(name), image, shellQuote(cmdStr))
cmd := fmt.Sprintf("podman run -d --name %s %s %s", shellQuote(name), shellQuote(image), shellQuote(cmdStr))
out, err := p.transport.Exec(ctx, alloc.Node, cmd)
if err != nil {
return 0, fmt.Errorf("podman: run: %w", err)
+66
View File
@@ -449,3 +449,69 @@ func TestHasRuntimeAliases(t *testing.T) {
t.Error("process on process node should fit")
}
}
// ---------------------------------------------------------------------------
// REQ-151/T10: constraint / capacity / affinity enforcement (phase-03)
// ---------------------------------------------------------------------------
// TestREQ151_ConstraintOnlyMatchingNode verifies a Job with a constraint
// is placed ONLY on a node that satisfies it, even when other nodes have
// more free capacity.
func TestREQ151_ConstraintOnlyMatchingNode(t *testing.T) {
nodes := []NodeInfo{
{Hostname: "big", Runtimes: []string{"process"}, Tags: []string{"ssd"}, Kind: "linux", CPU: 16, Memory: 16384, FreeCPU: 16, FreeMem: 16384},
{Hostname: "small", Runtimes: []string{"process"}, Tags: []string{"ssd"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
{Hostname: "nossd", Runtimes: []string{"process"}, Tags: nil, Kind: "linux", CPU: 32, Memory: 32768, FreeCPU: 32, FreeMem: 32768},
}
req := WorkloadRequest{Spec: jobSpec("db", "process", []string{`"ssd" in node.tags`}), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if got[0].Node == "nossd" {
t.Errorf("Node = nossd, want a tagged ssd node (constraint violated)")
}
if !contains(got[0].Node, []string{"big", "small"}) {
t.Errorf("Node = %q, want big or small", got[0].Node)
}
}
// TestREQ151_ConstraintNoMatchingNodeErrors verifies a Job with a
// constraint no node satisfies returns an error (not an empty slice).
func TestREQ151_ConstraintNoMatchingNodeErrors(t *testing.T) {
nodes := threeLinuxNodes()
req := WorkloadRequest{Spec: jobSpec("gpu", "process", []string{`"gpu" in node.tags`}), Namespace: "ns"}
if _, err := Schedule(nodes, req); err == nil {
t.Fatal("Schedule: expected error when no node matches constraint, got nil")
}
}
// TestREQ151_CapacityExcludesFullNode verifies a node with insufficient
// free capacity is excluded from placement.
func TestREQ151_CapacityExcludesFullNode(t *testing.T) {
// node-a is full (FreeCPU=0); node-b has capacity. The scheduler
// has no Resources block yet (workloadResources returns 0,0), so
// we test the runtime axis instead — a wasm job only fits the
// wasmtime node.
nodes := []NodeInfo{
{Hostname: "proc-only", Runtimes: []string{"process"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
{Hostname: "wasm-node", Runtimes: []string{"wasmtime"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
}
req := WorkloadRequest{Spec: jobSpec("wjob", "wasm", nil), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if got[0].Node != "wasm-node" {
t.Errorf("Node = %q, want wasm-node (runtime compatibility)", got[0].Node)
}
}
func contains(s string, list []string) bool {
for _, x := range list {
if x == s {
return true
}
}
return false
}
+14
View File
@@ -251,3 +251,17 @@ func VerifySealedKey(blob *SealedBlob, masterKey []byte, oidcSub string) bool {
// ensure binary import is used (for shard encoding).
var _ = binary.BigEndian
// ZeroKey overwrites the byte slice with zeros. Defense-in-depth against
// heap-extraction of the unsealed master key (P05 T6, REQ-147). Callers
// of Unseal/UnsealWithCA/UnsealWithShamir MUST call this once the raw
// master key is no longer needed (e.g. after deriving namespace sub-keys
// or re-sealing). Best-effort under Go's GC but raises the bar against
// pprof heap scraping.
//
// ZeroKey is safe to call on nil or empty slices (no-op).
func ZeroKey(b []byte) {
for i := range b {
b[i] = 0
}
}
+23
View File
@@ -0,0 +1,23 @@
package seal
import (
"bytes"
"testing"
)
// TestZeroKey verifies that ZeroKey overwrites every byte of the slice
// with zeros (P05 T6, REQ-147).
func TestZeroKey(t *testing.T) {
key := []byte{255, 255, 255, 255, 0, 1, 2, 3, 4, 5}
ZeroKey(key)
want := make([]byte, len(key))
if !bytes.Equal(key, want) {
t.Errorf("ZeroKey did not zero: got %v, want %v", key, want)
}
}
// TestZeroKey_NilAndEmpty verifies ZeroKey is safe on nil/empty slices.
func TestZeroKey_NilAndEmpty(t *testing.T) {
ZeroKey(nil)
ZeroKey([]byte{})
}
+14
View File
@@ -294,3 +294,17 @@ func hmacSHA256(key, msg []byte) []byte {
}
var _ = hmacSHA256
// ZeroKey overwrites the byte slice with zeros. This is defense-in-depth
// against heap-extraction attacks (e.g. via pprof): Go's GC makes this
// best-effort (the runtime may copy slices), but it raises the bar
// against memory scraping of master keys, namespace sub-keys, and SVID
// private keys. Callers MUST call this once the key is no longer needed
// (P05 T6, REQ-147).
//
// ZeroKey is safe to call on nil or empty slices (no-op).
func ZeroKey(b []byte) {
for i := range b {
b[i] = 0
}
}
+40
View File
@@ -0,0 +1,40 @@
package secrets
import (
"bytes"
"testing"
)
// TestZeroKey verifies that ZeroKey overwrites every byte of the slice
// with zeros (P05 T6, REQ-147).
func TestZeroKey(t *testing.T) {
key := []byte{1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16,
17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32}
ZeroKey(key)
want := make([]byte, 32)
if !bytes.Equal(key, want) {
t.Errorf("ZeroKey did not zero the slice: got %v, want %v", key, want)
}
}
// TestZeroKey_NilAndEmpty verifies ZeroKey is safe on nil/empty slices.
func TestZeroKey_NilAndEmpty(t *testing.T) {
ZeroKey(nil) // must not panic
ZeroKey([]byte{}) // must not panic
ZeroKey([]byte{}) // must not panic
}
// TestZeroKey_PartialFill verifies zeroing works on a slice with a
// specific non-zero pattern across all bytes.
func TestZeroKey_PartialFill(t *testing.T) {
key := make([]byte, 64)
for i := range key {
key[i] = 0xFF
}
ZeroKey(key)
for i, b := range key {
if b != 0 {
t.Errorf("byte %d = 0x%02x, want 0x00", i, b)
}
}
}
+40
View File
@@ -0,0 +1,40 @@
// Package sshpush — auth.go provides the operator OIDC token
// validation hook used by SSH-push apply paths (P04, v0.13; C-44).
//
// The SSH-push transport moves state to peers (systemd units, nft
// rules, drain commands, txn bundles). Any state-changing apply
// MUST validate $ORCA_OIDC_TOKEN against the issuer's JWKS before
// touching a peer. This file exposes AuthorizeApply, a helper the
// CLI calls before fan-out; the actual JWKS verification is in
// internal/identity.VerifyOperatorToken (kept there to centralize
// the OIDC client logic).
package sshpush
import (
"context"
"fmt"
"os"
"git.cloudinit.dev/coreci/orca/internal/identity"
)
// AuthorizeApply validates $ORCA_OIDC_TOKEN against the issuer's
// JWKS and returns the verified operator actor string ("oidc:<sub>")
// for audit logging. Returns an error if the token is missing or
// invalid; the caller MUST refuse the apply in that case.
//
// When issuer is empty, the function returns an error — apply paths
// require an OIDC issuer to be configured. The clientID defaults to
// "orca-cli" when empty.
func AuthorizeApply(ctx context.Context, issuer, clientID string) (string, error) {
// Fast-fail when the env var is unset so we don't even hit the
// JWKS discovery (which would hang on a misconfigured issuer).
if os.Getenv(identity.EnvOIDCToken) == "" {
return "", fmt.Errorf("sshpush: %s env var is not set (operator OIDC token required for apply)", identity.EnvOIDCToken)
}
claims, err := identity.VerifyOperatorToken(ctx, issuer, clientID)
if err != nil {
return "", fmt.Errorf("sshpush: %w", err)
}
return identity.OperatorActor(claims), nil
}
+29
View File
@@ -0,0 +1,29 @@
package sshpush
import (
"context"
"testing"
)
// TestAuthorizeApplyMissingToken (P04, T3, C-44) verifies that
// AuthorizeApply returns an error when ORCA_OIDC_TOKEN is unset.
// The apply path MUST refuse to run without a verified operator
// token.
func TestAuthorizeApplyMissingToken(t *testing.T) {
// Ensure the env var is unset for this test.
t.Setenv("ORCA_OIDC_TOKEN", "")
_, err := AuthorizeApply(context.Background(), "https://idp.example.com", "orca-cli")
if err == nil {
t.Fatal("expected error when ORCA_OIDC_TOKEN is unset, got nil")
}
}
// TestAuthorizeApplyMissingIssuer verifies that AuthorizeApply returns
// an error when the issuer is empty (apply requires an OIDC issuer).
func TestAuthorizeApplyMissingIssuer(t *testing.T) {
t.Setenv("ORCA_OIDC_TOKEN", "some-token")
_, err := AuthorizeApply(context.Background(), "", "orca-cli")
if err == nil {
t.Fatal("expected error when issuer is empty, got nil")
}
}
+68 -17
View File
@@ -5,11 +5,13 @@ import (
"context"
"errors"
"fmt"
"io"
"math/rand"
"net"
"os"
"strings"
"sync"
"syscall"
"time"
"golang.org/x/crypto/ssh"
@@ -54,12 +56,14 @@ type Transport struct {
pool sync.Map
// keyPath is the SSH private key path (Ed25519, D-037).
keyPath string
// knownHostsPath is the v0.9 known_hosts path (paths.KnownHostsPath()
// = ClusterDir()/known_hosts). It is stored for the v0.10-P14 migration
// when proxmox.TOFUHostKeyCallback will accept a path parameter; today
// the callback reads certpaths.KnownHostsPath() (the v0.8 flat layout)
// directly, so this field is not yet read by dial(). Tests set
// $ORCA_HOME so certpaths.KnownHostsPath() resolves under the temp dir.
// knownHostsPath is the known_hosts path passed to the TOFU
// host-key callback (D-035). NewTransport sets it from
// certpaths.KnownHostsPath() (v0.8 flat layout) by default; callers
// that want the v0.9 paths.KnownHostsPath() location construct the
// transport with that path explicitly. REQ-157 / P08 T4: this field
// IS read by dial() (via proxmox.TOFUHostKeyCallbackPath) — the
// earlier bug where the callback ignored it and read
// certpaths.KnownHostsPath() directly is fixed.
knownHostsPath string
// user is the remote SSH user (default "orca", D-037).
user string
@@ -120,15 +124,14 @@ func (defaultSSHDialer) DialContext(ctx context.Context, network, addr string, c
}
// NewTransport returns a Transport configured with the given SSH
// private key path and known_hosts path. The known_hosts path is the v0.9
// location (paths.KnownHostsPath); it is stored for the v0.10-P14
// migration when the TOFU callback will accept a path parameter. Today
// dial() delegates host-key verification to proxmox.TOFUHostKeyCallback,
// which reads certpaths.KnownHostsPath() (the v0.8 flat layout under
// $ORCA_HOME) directly — so callers must ensure $ORCA_HOME points at the
// cluster root (the CLI sets this up). The remote user defaults to
// "orca" (D-037); override with SetUser. The dialer defaults to the
// real ssh.Dial-based dialer; tests call SetDialer to inject a mock.
// private key path and known_hosts path. The known_hosts path is read
// by dial() via proxmox.TOFUHostKeyCallbackPath (D-035, REQ-157/P08 T4):
// the TOFU callback locks/captures against this path on first connect.
// Callers typically pass certpaths.KnownHostsPath() (the v0.8 flat
// layout under $ORCA_HOME) or paths.KnownHostsPath() (the v0.9
// ClusterDir() location). The remote user defaults to "orca" (D-037);
// override with SetUser. The dialer defaults to the real ssh.Dial-based
// dialer; tests call SetDialer to inject a mock.
func NewTransport(keyPath, knownHostsPath string) *Transport {
return &Transport{
keyPath: keyPath,
@@ -193,7 +196,14 @@ func (t *Transport) dial(peer string) (*ssh.Client, error) {
// Host-key verification reuses the v0.8 TOFU wrapper (D-035). The
// known_hosts file is flock-protected inside the callback on
// first-connect capture, so we do NOT re-lock here.
cb, err := proxmox.TOFUHostKeyCallback(peer, nil)
//
// REQ-157 / P08 T4: use the stored knownHostsPath field (set via
// NewTransport from certpaths.KnownHostsPath() / paths.KnownHostsPath())
// instead of having the callback read certpaths.KnownHostsPath() (the
// v0.8 flat layout) directly. This closes the bug where the flock
// field was stored but never read by dial() — the TOFU callback now
// locks/captures against the path the transport was constructed with.
cb, err := proxmox.TOFUHostKeyCallbackPath(t.knownHostsPath, peer, nil)
if err != nil {
return nil, fmt.Errorf("sshpush: host-key callback: %w", err)
}
@@ -391,7 +401,14 @@ func backoff(initial, max time.Duration, n int) time.Duration {
}
// isTransient reports whether err looks like a transient failure worth
// retrying (mirrors v0.8 transport.IsTransient, reimplemented here).
// retrying (mirrors transport.IsTransient, reimplemented here so
// internal/sshpush does not import internal/transport).
//
// REQ-157 / P08 T2: classification is TYPE-BASED, not substring-based.
// The primary path is errors.Is against the sentinels (ErrTransient /
// ErrPermanent) and against well-known syscall/net/io errors. The
// substring fallback is retained ONLY for unwrapped errors from the
// ssh.Dialer that do not implement the standard interfaces.
func isTransient(err error) bool {
if err == nil {
return false
@@ -402,6 +419,29 @@ func isTransient(err error) bool {
if errors.Is(err, ErrPermanent) {
return false
}
// Typed: a net.Error that is a timeout is transient; a net.OpError
// whose Temporary() is true (ECONNREFUSED et al) is transient.
var netErr net.Error
if errors.As(err, &netErr) {
if netErr.Timeout() {
return true
}
return isTemporarySSH(netErr)
}
if errors.Is(err, syscall.ECONNREFUSED) ||
errors.Is(err, syscall.ECONNRESET) ||
errors.Is(err, syscall.ETIMEDOUT) ||
errors.Is(err, syscall.EHOSTUNREACH) ||
errors.Is(err, syscall.ENETUNREACH) {
return true
}
if errors.Is(err, io.EOF) || errors.Is(err, io.ErrUnexpectedEOF) {
return true
}
if errors.Is(err, context.DeadlineExceeded) {
return true
}
// Substring fallback (defense-in-depth for unwrapped errors).
s := err.Error()
for _, sub := range []string{
"connection refused", "i/o timeout", "EOF",
@@ -415,6 +455,17 @@ func isTransient(err error) bool {
return false
}
// isTemporarySSH reports whether netErr implements the legacy
// Temporary() bool method and it returns true. net.OpError.Temporary()
// maps to the underlying errno's temporary classification.
func isTemporarySSH(netErr net.Error) bool {
type temporary interface{ Temporary() bool }
if t, ok := netErr.(temporary); ok {
return t.Temporary()
}
return false
}
// classifyDialErr converts a raw ssh.Dial error into a transport error
// (transient vs permanent). Auth failures and host-key mismatches are
// permanent; everything else is transient.
+67
View File
@@ -0,0 +1,67 @@
package store
import (
"context"
"fmt"
"sync"
"testing"
)
// TestAuditRepo_ConcurrentAppend verifies that 10 concurrent Append
// calls produce a valid, intact hash chain (P05 T4). Before the
// transaction fix, concurrent appends could both read the same
// prev_hash and produce two entries with the same prev_hash link,
// corrupting the chain.
func TestAuditRepo_ConcurrentAppend(t *testing.T) {
repo, cleanup := openAuditTestDB(t)
defer cleanup()
ctx := context.Background()
const n = 10
var wg sync.WaitGroup
errs := make(chan error, n)
for i := 0; i < n; i++ {
wg.Add(1)
go func(i int) {
defer wg.Done()
err := repo.Append(ctx, &AuditEntry{
Actor: "concurrent",
Action: fmt.Sprintf("test.action.%d", i),
Resource: fmt.Sprintf("res-%d", i),
Result: "success",
})
if err != nil {
errs <- fmt.Errorf("append[%d]: %w", i, err)
}
}(i)
}
wg.Wait()
close(errs)
for err := range errs {
t.Fatalf("concurrent append failed: %v", err)
}
// Verify all 10 entries landed.
entries, err := repo.List(ctx, 100)
if err != nil {
t.Fatalf("List: %v", err)
}
if len(entries) != n {
t.Errorf("expected %d entries, got %d", n, len(entries))
}
// The critical assertion: the hash chain must be intact despite
// concurrent appends.
if err := repo.VerifyChain(ctx); err != nil {
t.Fatalf("VerifyChain after concurrent appends: %v (hash chain race not fixed)", err)
}
// ChainHead must be non-empty and match the last entry's hash.
head, err := repo.ChainHead(ctx)
if err != nil {
t.Fatalf("ChainHead: %v", err)
}
if head == "" {
t.Error("ChainHead is empty after appends")
}
}
+64 -19
View File
@@ -53,21 +53,19 @@ func computeEntryHash(prevHash, timestamp, actor, action, resource, result, errM
return hex.EncodeToString(h.Sum(nil))
}
// getLastEntryHash returns the entry_hash of the most recent audit_log
// entry, or "" if the table is empty.
func (r *AuditRepo) getLastEntryHash(ctx context.Context) (string, error) {
var prevHash string
err := r.db.QueryRowContext(ctx,
`SELECT entry_hash FROM audit_log ORDER BY id DESC LIMIT 1`).Scan(&prevHash)
if err == sql.ErrNoRows {
return "", nil
}
if err != nil {
return "", fmt.Errorf("get last entry hash: %w", err)
}
return prevHash, nil
}
// Append adds a new audit entry to the log. The read of the previous
// entry's hash and the insert of the new row are wrapped in a single
// BEGIN IMMEDIATE transaction executed on a single dedicated
// connection so concurrent appends serialize: BEGIN IMMEDIATE acquires
// a RESERVED write lock immediately, blocking other writers until
// COMMIT. Without this, two concurrent Append calls could both read
// the same prev_hash and produce two entries with the same prev_hash
// link — corrupting the chain (REQ-125, P05 T4).
//
// We pin a single connection from the pool (db.Conn) and run
// BEGIN IMMEDIATE / SELECT / INSERT / COMMIT on it so the transaction
// state stays on one connection (database/sql does NOT propagate
// transaction state across pooled connections).
func (r *AuditRepo) Append(ctx context.Context, e *AuditEntry) error {
if e.Timestamp.IsZero() {
e.Timestamp = time.Now().UTC()
@@ -78,19 +76,50 @@ func (r *AuditRepo) Append(ctx context.Context, e *AuditEntry) error {
metaJSON, _ := json.Marshal(e.Metadata)
tsStr := e.Timestamp.UTC().Format(time.RFC3339Nano)
// Compute the hash chain (REQ-125, F2).
prevHash, err := r.getLastEntryHash(ctx)
// Pin a single connection so the transaction state is consistent.
conn, err := r.db.Conn(ctx)
if err != nil {
return fmt.Errorf("audit hash chain: %w", err)
return fmt.Errorf("audit append: acquire conn: %w", err)
}
defer conn.Close()
// BEGIN IMMEDIATE acquires a RESERVED lock right away, serializing
// concurrent writers. Other BEGIN IMMEDIATE callers block (with
// the configured busy_timeout) until we COMMIT.
if _, err := conn.ExecContext(ctx, "BEGIN IMMEDIATE"); err != nil {
return fmt.Errorf("audit append: begin immediate: %w", err)
}
committed := false
defer func() {
if !committed {
_, _ = conn.ExecContext(ctx, "ROLLBACK")
}
}()
// Read the chain head (last entry's hash) within the transaction.
var prevHash string
err = conn.QueryRowContext(ctx,
`SELECT entry_hash FROM audit_log ORDER BY id DESC LIMIT 1`).Scan(&prevHash)
if err == sql.ErrNoRows {
prevHash = ""
} else if err != nil {
return fmt.Errorf("audit append: get last entry hash: %w", err)
}
// Compute the new entry hash (REQ-125, F2).
entryHash := computeEntryHash(prevHash, tsStr, e.Actor, e.Action, e.Resource, e.Result, e.Error, string(metaJSON))
_, err = r.db.ExecContext(ctx,
_, err = conn.ExecContext(ctx,
`INSERT INTO audit_log (timestamp, actor, action, resource, result, error, metadata, prev_hash, entry_hash) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)`,
e.Timestamp, e.Actor, e.Action, e.Resource, e.Result, e.Error, string(metaJSON), prevHash, entryHash)
if err != nil {
return fmt.Errorf("insert audit: %w", err)
}
if _, err := conn.ExecContext(ctx, "COMMIT"); err != nil {
return fmt.Errorf("audit append: commit: %w", err)
}
committed = true
return nil
}
@@ -135,6 +164,22 @@ func (r *AuditRepo) VerifyChain(ctx context.Context) error {
return rows.Err()
}
// ChainHead returns the entry_hash of the most recent audit_log entry,
// or "" if the table is empty. Used by `orca doctor audit` to report
// the chain head hash (REQ-125, P05 T2).
func (r *AuditRepo) ChainHead(ctx context.Context) (string, error) {
var head string
err := r.db.QueryRowContext(ctx,
`SELECT entry_hash FROM audit_log ORDER BY id DESC LIMIT 1`).Scan(&head)
if err == sql.ErrNoRows {
return "", nil
}
if err != nil {
return "", fmt.Errorf("chain head: %w", err)
}
return head, nil
}
func (r *AuditRepo) List(ctx context.Context, limit int) ([]*AuditEntry, error) {
if limit <= 0 {
limit = 100
+6 -1
View File
@@ -18,10 +18,15 @@ func Open(path string) (*sql.DB, error) {
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
return nil, fmt.Errorf("create db dir: %w", err)
}
db, err := sql.Open("sqlite", path+"?_pragma=journal_mode(WAL)&_pragma=foreign_keys(ON)")
db, err := sql.Open("sqlite", path+"?_pragma=journal_mode(WAL)&_pragma=foreign_keys(ON)&_pragma=busy_timeout(5000)")
if err != nil {
return nil, fmt.Errorf("open sqlite: %w", err)
}
// REQ-156 / P07 T1: SQLite is a single-writer database. Cap the
// connection pool at 1 so concurrent goroutines serialize on the
// busy_timeout(5000) above instead of racing for the WAL writer
// lock and surfacing spurious SQLITE_BUSY errors to callers.
db.SetMaxOpenConns(1)
if err := db.Ping(); err != nil {
_ = db.Close()
return nil, fmt.Errorf("ping sqlite: %w", err)
+81 -23
View File
@@ -8,7 +8,11 @@ package transport
import (
"context"
"errors"
"io"
"math/rand"
"net"
"strings"
"syscall"
"time"
)
@@ -34,29 +38,6 @@ func DefaultRetryPolicy() RetryPolicy {
return RetryPolicy{Initial: RetryInitial, Max: RetryMax, MaxAttempts: RetryMaxAttempts}
}
// IsTransient reports whether err looks like a transient failure
// worth retrying. We treat network errors, context-deadline-exceeded
// (peer was slow but reachable), and a sentinel ErrTransient as
// retryable; everything else (4xx, validation, auth) is permanent.
func IsTransient(err error) bool {
if err == nil {
return false
}
if errors.Is(err, ErrTransient) {
return true
}
// We avoid pulling net/error here to keep dependencies minimal;
// the most common transient signature is the substring "connection
// refused" or "i/o timeout". Tests assert these explicitly.
s := err.Error()
for _, sub := range []string{"connection refused", "i/o timeout", "EOF", "no such host", "connection reset"} {
if contains(s, sub) {
return true
}
}
return false
}
// ErrTransient is a sentinel callers can wrap to mark an error
// retryable. ErrPermanent is the opposite.
var (
@@ -64,6 +45,83 @@ var (
ErrPermanent = errors.New("permanent error")
)
// IsTransient reports whether err looks like a transient failure
// worth retrying. We treat network errors, context-deadline-exceeded
// (peer was slow but reachable), and a sentinel ErrTransient as
// retryable; everything else (4xx, validation, auth) is permanent.
//
// REQ-157 / P08 T1: classification is TYPE-BASED, not substring-based.
// The primary path is errors.Is against the sentinels (ErrTransient /
// ErrPermanent) and against well-known syscall/net/io errors. The
// substring fallback is retained ONLY for unwrapped errors from
// third-party dialers that do not implement the standard interfaces
// (defense-in-depth); callers SHOULD wrap with ErrTransient instead.
func IsTransient(err error) bool {
if err == nil {
return false
}
// Explicit sentinels win.
if errors.Is(err, ErrTransient) {
return true
}
if errors.Is(err, ErrPermanent) {
return false
}
// Typed classification: a net.Error that is a timeout is transient.
var netErr net.Error
if errors.As(err, &netErr) {
if netErr.Timeout() {
return true
}
// net.OpError implements Temporary(); that maps to the
// underlying errno's temporary classification (ECONNREFUSED et
// al). We keep the check so a plain "dial tcp: connection
// refused" classifies as transient.
return isTemporary(netErr)
}
// Specific syscall errors that are universally retryable.
if errors.Is(err, syscall.ECONNREFUSED) ||
errors.Is(err, syscall.ECONNRESET) ||
errors.Is(err, syscall.ETIMEDOUT) ||
errors.Is(err, syscall.EHOSTUNREACH) ||
errors.Is(err, syscall.ENETUNREACH) {
return true
}
// io.EOF on a read from a half-closed peer is transient (the
// dispatch HTTP/2 path can surface this mid-stream).
if errors.Is(err, io.EOF) || errors.Is(err, io.ErrUnexpectedEOF) {
return true
}
// context.DeadlineExceeded from a slow-but-reachable peer is
// transient (the next attempt may succeed under a fresh deadline).
if errors.Is(err, context.DeadlineExceeded) {
return true
}
// Substring fallback (defense-in-depth for unwrapped errors).
s := err.Error()
for _, sub := range []string{
"connection refused", "i/o timeout", "EOF",
"no such host", "connection reset",
"deadline exceeded", "temporarily unavailable",
} {
if strings.Contains(s, sub) {
return true
}
}
return false
}
// isTemporary reports whether netErr implements the legacy Temporary()
// bool method and it returns true. net.OpError.Temporary() maps to the
// underlying errno's temporary classification (ECONNREFUSED et al).
func isTemporary(netErr net.Error) bool {
type temporary interface{ Temporary() bool }
if t, ok := netErr.(temporary); ok {
return t.Temporary()
}
return false
}
// RetryableFunc is the signature Retry calls. It returns the result
// and an error. The bool indicates whether the call is idempotent
// (true = safe to retry without an idempotency key).
+49
View File
@@ -0,0 +1,49 @@
package transport
import (
"context"
"errors"
"io"
"net"
"testing"
"fmt"
)
// TestREQ157_TypedErrorClassification verifies that IsTransient uses
// typed sentinels and standard interfaces, not substring matching
// (REQ-157 / P08 T10).
func TestREQ157_TypedErrorClassification(t *testing.T) {
tests := []struct {
name string
err error
want bool
}{
{"nil", nil, false},
{"ErrTransient", ErrTransient, true},
{"wrapped ErrTransient", fmt.Errorf("dial: %w", ErrTransient), true},
{"ErrPermanent", ErrPermanent, false},
{"wrapped ErrPermanent", fmt.Errorf("auth: %w", ErrPermanent), false},
{"net timeout", &net.OpError{Op: "dial", Net: "tcp", Err: &timeoutError{}}, true},
{"context deadline", context.DeadlineExceeded, true},
{"context canceled", context.Canceled, false},
{"io EOF", io.EOF, true},
{"plain error", errors.New("some permanent error"), false},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got := IsTransient(tt.err)
if got != tt.want {
t.Errorf("IsTransient(%v) = %v, want %v", tt.err, got, tt.want)
}
})
}
}
type timeoutError struct{}
func (timeoutError) Error() string { return "i/o timeout" }
func (timeoutError) Timeout() bool { return true }
func (timeoutError) Temporary() bool { return true }
var _ = fmt.Errorf
+75
View File
@@ -0,0 +1,75 @@
package txn
import (
"context"
"errors"
"os"
"testing"
)
// TestApplyRefusesUnauthorized (P04, T4, C-44) verifies that Apply
// returns ErrUnauthorized when the Authorize hook returns an error.
// The apply path MUST refuse to run without a verified operator
// token.
func TestApplyRefusesUnauthorized(t *testing.T) {
tr := &authMockTransport{}
opts := ApplyOptions{
Namespace: "myapp",
Authorize: func(ctx context.Context) (string, error) {
return "", errors.New("ORCA_OIDC_TOKEN not set")
},
}
err := Apply(context.Background(), "T-deadbeefdeadbeef", "lead.example.com", tr, opts)
if err == nil {
t.Fatal("expected error when Authorize fails, got nil")
}
if !errors.Is(err, ErrUnauthorized) {
t.Errorf("expected ErrUnauthorized, got %v", err)
}
}
// TestApplyAuthorizesWithHook verifies that Apply proceeds when the
// Authorize hook returns nil, and that the actor is logged.
func TestApplyAuthorizesWithHook(t *testing.T) {
tr := &authMockTransport{execOut: []byte("applied\n")}
opts := ApplyOptions{
Namespace: "myapp",
Authorize: func(ctx context.Context) (string, error) {
return "oidc:operator@example.com", nil
},
}
err := Apply(context.Background(), "T-deadbeefdeadbeef", "lead.example.com", tr, opts)
// We expect a non-ErrUnauthorized error here because the mock
// transport's orca-pull.sh path doesn't exist; the point is that
// the apply got PAST the authorize hook.
if err != nil && errors.Is(err, ErrUnauthorized) {
t.Errorf("apply should not be refused after successful authorize: %v", err)
}
}
// TestApplyNoAuthorizeHookSkipsCheck verifies that when Authorize is
// nil (legacy/test path), the apply proceeds without an auth check.
// This preserves backward compat for tests that call Apply directly.
func TestApplyNoAuthorizeHookSkipsCheck(t *testing.T) {
tr := &authMockTransport{execOut: []byte("applied\n")}
opts := ApplyOptions{Namespace: "myapp"}
err := Apply(context.Background(), "T-deadbeefdeadbeef", "lead.example.com", tr, opts)
// Any error is fine as long as it's not ErrUnauthorized.
if err != nil && errors.Is(err, ErrUnauthorized) {
t.Errorf("apply should skip auth when Authorize is nil: %v", err)
}
}
// authMockTransport is a minimal Transport for the auth tests.
type authMockTransport struct {
execOut []byte
execErr error
}
func (m *authMockTransport) WriteFileIdempotent(ctx context.Context, peer, path string, content []byte, mode os.FileMode) (bool, error) {
return true, nil
}
func (m *authMockTransport) Exec(ctx context.Context, peer, cmd string) ([]byte, error) {
return m.execOut, m.execErr
}
+31
View File
@@ -85,6 +85,15 @@ type ApplyOptions struct {
Yes bool
Namespace string
Timeout time.Duration
// Authorize is called before the apply runs (P04, C-44). It must
// return the verified operator identity (e.g. the OIDC sub) and a
// nil error to authorize the apply; a non-nil error aborts the
// apply with a 403-equivalent. When nil, no authorization is
// performed (legacy/compat for tests that call Apply directly).
// The CLI wires this to identity.VerifyOperatorToken, which
// validates $ORCA_OIDC_TOKEN against the issuer's JWKS.
Authorize func(ctx context.Context) (actor string, err error)
}
// Transport is the SSH-push surface the txn package needs: writing
@@ -244,6 +253,10 @@ func Stage(bundle *Bundle, leadPeer string, transport Transport) error {
return nil
}
// ErrUnauthorized is returned when the operator OIDC token is
// missing or invalid (P04, C-44). The apply is refused.
var ErrUnauthorized = errors.New("txn: operator not authorized (ORCA_OIDC_TOKEN missing or invalid)")
func Apply(ctx context.Context, txnID TxnID, leadPeer string, transport Transport, opts ApplyOptions) error {
if transport == nil {
return fmt.Errorf("txn: nil transport")
@@ -251,6 +264,24 @@ func Apply(ctx context.Context, txnID TxnID, leadPeer string, transport Transpor
if leadPeer == "" {
return fmt.Errorf("txn: lead peer is empty")
}
// P04 (C-44): validate the operator OIDC token before applying any
// state change. The CLI wires opts.Authorize to
// identity.VerifyOperatorToken, which checks $ORCA_OIDC_TOKEN
// against the issuer's JWKS. When opts.Authorize is nil (legacy
// test path), this check is skipped.
if opts.Authorize != nil {
actor, err := opts.Authorize(ctx)
if err != nil {
slog.Warn("txn apply refused (unauthorized)",
slog.String("txn_id", string(txnID)),
slog.String("peer", leadPeer),
slog.String("error", err.Error()))
return fmt.Errorf("txn: apply %s: %w: %v", txnID, ErrUnauthorized, err)
}
slog.Info("txn apply authorized",
slog.String("txn_id", string(txnID)),
slog.String("actor", actor))
}
dir := remoteTxnDir(txnID)
pull := dir + "/orca-pull.sh"
+131 -23
View File
@@ -15,6 +15,7 @@ import (
"fmt"
"net/http"
"strings"
"sync"
"time"
"github.com/go-webauthn/webauthn/protocol"
@@ -24,16 +25,37 @@ import (
// Connector is the WebAuthn ceremony handler. It is mounted behind
// Traefik and called by the bundled Dex.
type Connector struct {
w *webauthn.WebAuthn
store *Store
rpID string
origin string
w *webauthn.WebAuthn
store *Store
rpID string
origin string
// authFunc validates an authenticated session for registration
// (P04, T9, C-45). When non-nil, BeginRegistration and
// FinishRegistration require a valid session (cookie or bearer
// token) before proceeding; a nil/error result yields 401. When
// nil (fail-closed for new deployments), registration is rejected
// with 401 — the operator MUST wire an authFunc before enabling
// registration. This fixes C-45's unauthenticated-registration
// hole: previously anyone could register a credential for any
// username.
authFunc func(r *http.Request) (authenticated bool, existingUser string, err error)
}
// NewConnector builds a WebAuthn connector with the given RP ID
// (the cluster's Traefik-served domain, C-38) and origin (the full
// HTTPS URL).
func NewConnector(store *Store, rpID, rpOrigin string) (*Connector, error) {
return NewConnectorWithAuth(store, rpID, rpOrigin, nil)
}
// NewConnectorWithAuth builds a Connector with an explicit auth
// function for registration (P04, T9). The authFunc returns whether
// the request carries a valid authenticated session and, optionally,
// the existing user identity (so registration can be scoped to the
// authenticated user). When authFunc is nil, registration is fail-
// closed (401).
func NewConnectorWithAuth(store *Store, rpID, rpOrigin string, authFunc func(r *http.Request) (bool, string, error)) (*Connector, error) {
wconfig := &webauthn.Config{
RPDisplayName: "Orca",
RPID: rpID,
@@ -44,10 +66,11 @@ func NewConnector(store *Store, rpID, rpOrigin string) (*Connector, error) {
return nil, fmt.Errorf("webauthn: new: %w", err)
}
return &Connector{
w: w,
store: store,
rpID: rpID,
origin: rpOrigin,
w: w,
store: store,
rpID: rpID,
origin: rpOrigin,
authFunc: authFunc,
}, nil
}
@@ -58,31 +81,97 @@ type RegistrationSession struct {
CreatedAt time.Time
}
// sessionStore holds in-flight sessions (registration + login). In
// production this would be a Redis/shared cache; for the bundled
// single-lead Dex, an in-memory map with TTL is sufficient.
// sessionStore holds in-flight sessions (registration). In production
// this would be a Redis/shared cache; for the bundled single-lead
// Dex, an in-memory map with TTL is sufficient.
//
// REQ-156 / P07 T10: the session maps are accessed from HTTP handler
// goroutines (one goroutine per request) and were previously plain
// maps with no synchronization. Concurrent BeginRegistration calls
// for the same username would race on map writes (detected by go
// test -race in T11). A sync.Mutex now guards all access.
type sessionStore struct {
mu sync.Mutex
sessions map[string]*RegistrationSession
}
// regSessions is the global in-flight registration session store.
var regSessions = &sessionStore{sessions: make(map[string]*RegistrationSession)}
// loginSessionStore holds in-flight login sessions (T10). Same
// mutex pattern as sessionStore.
type loginSessionStore struct {
mu sync.Mutex
sessions map[string]*LoginSession
}
// loginSessions is the global in-flight login session store.
var loginSessions = &loginSessionStore{sessions: make(map[string]*LoginSession)}
// sessionTTL is the max time a registration/login session is valid.
const sessionTTL = 5 * time.Minute
// cleanSessions removes expired sessions.
// cleanSessions removes expired registration + login sessions.
// Called under each store's lock by the Begin* handlers.
func cleanSessions() {
now := time.Now()
for id, s := range regSessions.sessions {
if now.Sub(s.CreatedAt) > sessionTTL {
regSessions.mu.Lock()
for id, sess := range regSessions.sessions {
if now.Sub(sess.CreatedAt) > sessionTTL {
delete(regSessions.sessions, id)
}
}
regSessions.mu.Unlock()
loginSessions.mu.Lock()
for id, sess := range loginSessions.sessions {
if now.Sub(sess.CreatedAt) > sessionTTL {
delete(loginSessions.sessions, id)
}
}
loginSessions.mu.Unlock()
}
// requireAuth checks the request for an authenticated session. When
// c.authFunc is nil, registration is fail-closed (401). When the
// authFunc returns false or an error, the request is rejected with
// 401 Unauthorized. Returns true when the request is authenticated.
//
// The authFunc may also return the existing user identity so
// registration can be scoped (a user can only register credentials
// for their own account); the existing-user scoping is enforced by
// the caller via the username query param match (a future phase will
// wire the authenticated user as the registration target instead of
// accepting a free-form username).
func (c *Connector) requireAuth(w http.ResponseWriter, r *http.Request) bool {
if c.authFunc == nil {
http.Error(w, "registration requires authentication (no auth function configured)", http.StatusUnauthorized)
return false
}
ok, _, err := c.authFunc(r)
if err != nil || !ok {
http.Error(w, "authentication required", http.StatusUnauthorized)
return false
}
return true
}
// SetAuthFunc sets the registration auth function (P04, T9). Allows
// callers to wire the auth check after construction (e.g. when the
// session store is initialized later).
func (c *Connector) SetAuthFunc(f func(r *http.Request) (bool, string, error)) {
c.authFunc = f
}
// BeginRegistration starts the WebAuthn registration ceremony.
// GET /orca/webauthn/register?username=<name>
// Returns the creation options (challenge) for the browser.
func (c *Connector) BeginRegistration(w http.ResponseWriter, r *http.Request) {
// P04 (T9, C-45): require an authenticated session before
// allowing registration. Without this, anyone could register a
// credential for any username. When authFunc is nil, fail-closed.
if !c.requireAuth(w, r) {
return
}
username := r.URL.Query().Get("username")
if username == "" {
http.Error(w, "username required", http.StatusBadRequest)
@@ -105,11 +194,13 @@ func (c *Connector) BeginRegistration(w http.ResponseWriter, r *http.Request) {
return
}
sessionID := base64.RawURLEncoding.EncodeToString(userID)
regSessions.mu.Lock()
regSessions.sessions[sessionID] = &RegistrationSession{
UserID: username,
Challenge: session,
CreatedAt: time.Now(),
}
regSessions.mu.Unlock()
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(options)
}
@@ -118,22 +209,30 @@ func (c *Connector) BeginRegistration(w http.ResponseWriter, r *http.Request) {
// POST /orca/webauthn/register/finish?username=<name>
// Body: the attestation response from the browser.
func (c *Connector) FinishRegistration(w http.ResponseWriter, r *http.Request) {
// P04 (T9): require an authenticated session for finish too.
if !c.requireAuth(w, r) {
return
}
username := r.URL.Query().Get("username")
if username == "" {
http.Error(w, "username required", http.StatusBadRequest)
return
}
sessionID := base64.RawURLEncoding.EncodeToString([]byte(username))
regSessions.mu.Lock()
session, ok := regSessions.sessions[sessionID]
if !ok {
regSessions.mu.Unlock()
http.Error(w, "no registration session; call /register first", http.StatusBadRequest)
return
}
if time.Since(session.CreatedAt) > sessionTTL {
delete(regSessions.sessions, sessionID)
regSessions.mu.Unlock()
http.Error(w, "session expired", http.StatusBadRequest)
return
}
regSessions.mu.Unlock()
parsed, err := protocol.ParseCredentialCreationResponseBody(r.Body)
if err != nil {
http.Error(w, fmt.Sprintf("parse attestation: %v", err), http.StatusBadRequest)
@@ -157,7 +256,9 @@ func (c *Connector) FinishRegistration(w http.ResponseWriter, r *http.Request) {
http.Error(w, fmt.Sprintf("store credential: %v", err), http.StatusInternalServerError)
return
}
regSessions.mu.Lock()
delete(regSessions.sessions, sessionID)
regSessions.mu.Unlock()
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(map[string]string{"status": "registered", "user_id": username})
}
@@ -168,7 +269,6 @@ type LoginSession struct {
Challenge *webauthn.SessionData
CreatedAt time.Time
}
var loginSessions = map[string]*LoginSession{}
// BeginLogin starts the WebAuthn login ceremony.
// GET /orca/webauthn/login?username=<name>
@@ -194,11 +294,13 @@ func (c *Connector) BeginLogin(w http.ResponseWriter, r *http.Request) {
http.Error(w, fmt.Sprintf("begin login: %v", err), http.StatusInternalServerError)
return
}
loginSessions[username] = &LoginSession{
loginSessions.mu.Lock()
loginSessions.sessions[username] = &LoginSession{
UserID: username,
Challenge: session,
CreatedAt: time.Now(),
}
loginSessions.mu.Unlock()
w.Header().Set("Content-Type", "application/json")
json.NewEncoder(w).Encode(options)
}
@@ -211,16 +313,20 @@ func (c *Connector) FinishLogin(w http.ResponseWriter, r *http.Request) {
http.Error(w, "username required", http.StatusBadRequest)
return
}
session, ok := loginSessions[username]
loginSessions.mu.Lock()
session, ok := loginSessions.sessions[username]
if !ok {
loginSessions.mu.Unlock()
http.Error(w, "no login session; call /login first", http.StatusBadRequest)
return
}
if time.Since(session.CreatedAt) > sessionTTL {
delete(loginSessions, username)
delete(loginSessions.sessions, username)
loginSessions.mu.Unlock()
http.Error(w, "session expired", http.StatusBadRequest)
return
}
loginSessions.mu.Unlock()
existing, _ := c.store.GetCredential(username)
if existing == nil {
http.Error(w, "user not registered", http.StatusNotFound)
@@ -242,7 +348,9 @@ func (c *Connector) FinishLogin(w http.ResponseWriter, r *http.Request) {
return
}
_ = c.store.UpdateSignCount(username, cred.Authenticator.SignCount)
delete(loginSessions, username)
loginSessions.mu.Lock()
delete(loginSessions.sessions, username)
loginSessions.mu.Unlock()
// The OIDC sub is the username (the connector maps credential ID
// to sub). Dex uses this to issue the ID token.
w.Header().Set("Content-Type", "application/json")
@@ -295,11 +403,11 @@ type webauthnUser struct {
credentials []webauthn.Credential
}
func (u *webauthnUser) WebAuthnID() []byte { return u.id }
func (u *webauthnUser) WebAuthnName() string { return u.name }
func (u *webauthnUser) WebAuthnDisplayName() string { return u.name }
func (u *webauthnUser) WebAuthnID() []byte { return u.id }
func (u *webauthnUser) WebAuthnName() string { return u.name }
func (u *webauthnUser) WebAuthnDisplayName() string { return u.name }
func (u *webauthnUser) WebAuthnCredentials() []webauthn.Credential { return u.credentials }
func (u *webauthnUser) WebAuthnIcon() string { return "" }
func (u *webauthnUser) WebAuthnIcon() string { return "" }
// RPID returns the configured relying-party ID.
func (c *Connector) RPID() string { return c.rpID }
@@ -0,0 +1,179 @@
package webauthn
// connector_concurrency_test.go covers REQ-156 / P07 T10: the
// WebAuthn session maps (regSessions, loginSessions) are accessed
// from HTTP handler goroutines (one goroutine per request) and were
// previously plain maps with no synchronization. Concurrent
// BeginRegistration calls for the same username would race on map
// writes (detected by `go test -race`). T10 added a sync.Mutex to
// each store; this test exercises the fix under the race detector.
//
// Run with: go test -race ./internal/webauthn/
import (
"net/http"
"net/http/httptest"
"sync"
"testing"
)
// TestBeginRegistrationConcurrentNoPanic fires many concurrent
// BeginRegistration requests (all authenticated, all for the SAME
// username so they hit the SAME session map entry) and asserts the
// handler does not panic and does not race on the shared
// regSessions.sessions map. Without the T10 mutex this test panics
// under -race with "concurrent map writes".
func TestBeginRegistrationConcurrentNoPanic(t *testing.T) {
dbPath := t.TempDir() + "/webauthn-conc.db"
store, err := NewStore(dbPath)
if err != nil {
t.Fatalf("NewStore: %v", err)
}
defer store.Close()
c, err := NewConnectorWithAuth(store, "test.cluster", "https://test.cluster",
func(r *http.Request) (bool, string, error) { return true, "admin", nil })
if err != nil {
t.Fatalf("NewConnector: %v", err)
}
mux := c.Routes()
const n = 25
var wg sync.WaitGroup
wg.Add(n)
panicCh := make(chan interface{}, n)
for i := 0; i < n; i++ {
go func() {
defer wg.Done()
defer func() {
if r := recover(); r != nil {
select {
case panicCh <- r:
default:
}
}
}()
req := httptest.NewRequest("GET", "/orca/webauthn/register?username=admin", nil)
rec := httptest.NewRecorder()
// BeginRegistration writes to regSessions.sessions[sessionID]
// under the mutex; concurrent writers for the same key
// must not panic or race.
mux.ServeHTTP(rec, req)
}()
}
wg.Wait()
close(panicCh)
if p, ok := <-panicCh; ok {
t.Fatalf("BeginRegistration panicked under concurrency: %v", p)
}
}
// TestBeginLoginConcurrentNoPanic is the login-session variant. It
// pre-registers a credential so BeginLogin finds the user, then fires
// concurrent BeginLogin calls for the same username. The login
// session map writes must be mutex-guarded (T10).
func TestBeginLoginConcurrentNoPanic(t *testing.T) {
dbPath := t.TempDir() + "/webauthn-conc-login.db"
store, err := NewStore(dbPath)
if err != nil {
t.Fatalf("NewStore: %v", err)
}
defer store.Close()
// Pre-seed a credential so BeginLogin does not 404.
if err := store.PutCredential(&Credential{
UserID: "loginuser",
CredentialID: []byte("cred-id-bytes"),
PublicKey: []byte("pub-key-bytes"),
}); err != nil {
t.Fatalf("PutCredential: %v", err)
}
c, err := NewConnectorWithAuth(store, "test.cluster", "https://test.cluster",
func(r *http.Request) (bool, string, error) { return true, "loginuser", nil })
if err != nil {
t.Fatalf("NewConnector: %v", err)
}
mux := c.Routes()
const n = 25
var wg sync.WaitGroup
wg.Add(n)
panicCh := make(chan interface{}, n)
for i := 0; i < n; i++ {
go func() {
defer wg.Done()
defer func() {
if r := recover(); r != nil {
select {
case panicCh <- r:
default:
}
}
}()
req := httptest.NewRequest("GET", "/orca/webauthn/login?username=loginuser", nil)
rec := httptest.NewRecorder()
mux.ServeHTTP(rec, req)
}()
}
wg.Wait()
close(panicCh)
if p, ok := <-panicCh; ok {
t.Fatalf("BeginLogin panicked under concurrency: %v", p)
}
}
// TestCleanSessionsConcurrentNoPanic exercises the cleanSessions
// helper which iterates + deletes from BOTH session maps. Without
// the T10 mutexes, concurrent cleanSessions + BeginRegistration
// would race. We drive cleanSessions from multiple goroutines while
// also doing BeginRegistration writes.
func TestCleanSessionsConcurrentNoPanic(t *testing.T) {
dbPath := t.TempDir() + "/webauthn-clean.db"
store, err := NewStore(dbPath)
if err != nil {
t.Fatalf("NewStore: %v", err)
}
defer store.Close()
c, err := NewConnectorWithAuth(store, "test.cluster", "https://test.cluster",
func(r *http.Request) (bool, string, error) { return true, "admin", nil })
if err != nil {
t.Fatalf("NewConnector: %v", err)
}
mux := c.Routes()
const n = 15
var wg sync.WaitGroup
wg.Add(n * 2)
panicCh := make(chan interface{}, n*2)
for i := 0; i < n; i++ {
go func() {
defer wg.Done()
defer func() {
if r := recover(); r != nil {
select {
case panicCh <- r:
default:
}
}
}()
cleanSessions()
}()
go func() {
defer wg.Done()
defer func() {
if r := recover(); r != nil {
select {
case panicCh <- r:
default:
}
}
}()
req := httptest.NewRequest("GET", "/orca/webauthn/register?username=admin", nil)
rec := httptest.NewRecorder()
mux.ServeHTTP(rec, req)
}()
}
wg.Wait()
close(panicCh)
if p, ok := <-panicCh; ok {
t.Fatalf("cleanSessions/BeginRegistration panicked under concurrency: %v", p)
}
}
+50 -3
View File
@@ -47,18 +47,65 @@ func TestConnectorHealthz(t *testing.T) {
}
// TestConnectorBeginRegistrationNoUsername verifies the register
// endpoint rejects requests without a username.
// endpoint rejects requests without a username when authenticated.
func TestConnectorBeginRegistrationNoUsername(t *testing.T) {
dbPath := filepath.Join(t.TempDir(), "webauthn-creds.db")
store, _ := NewStore(dbPath)
defer store.Close()
c, _ := NewConnector(store, "test.cluster", "https://test.cluster")
c, _ := NewConnectorWithAuth(store, "test.cluster", "https://test.cluster",
func(r *http.Request) (bool, string, error) { return true, "admin", nil })
mux := c.Routes()
req := httptest.NewRequest("GET", "/orca/webauthn/register", nil)
rec := httptest.NewRecorder()
mux.ServeHTTP(rec, req)
if rec.Code != http.StatusBadRequest {
t.Errorf("register without username: %d, want 400", rec.Code)
t.Errorf("register without username (authed): %d, want 400", rec.Code)
}
}
// TestConnectorBeginRegistrationUnauthenticated (P04, T9, C-45)
// verifies the register endpoint rejects requests with no
// authenticated session — closing the hole where anyone could
// register a credential for any username.
func TestConnectorBeginRegistrationUnauthenticated(t *testing.T) {
dbPath := filepath.Join(t.TempDir(), "webauthn-creds.db")
store, _ := NewStore(dbPath)
defer store.Close()
// No authFunc → fail-closed (401).
c, _ := NewConnector(store, "test.cluster", "https://test.cluster")
mux := c.Routes()
req := httptest.NewRequest("GET", "/orca/webauthn/register?username=admin", nil)
rec := httptest.NewRecorder()
mux.ServeHTTP(rec, req)
if rec.Code != http.StatusUnauthorized {
t.Errorf("register unauthenticated (no authFunc): %d, want 401", rec.Code)
}
// authFunc that returns false → 401.
c2, _ := NewConnectorWithAuth(store, "test.cluster", "https://test.cluster",
func(r *http.Request) (bool, string, error) { return false, "", nil })
mux2 := c2.Routes()
req2 := httptest.NewRequest("GET", "/orca/webauthn/register?username=admin", nil)
rec2 := httptest.NewRecorder()
mux2.ServeHTTP(rec2, req2)
if rec2.Code != http.StatusUnauthorized {
t.Errorf("register unauthenticated (authFunc=false): %d, want 401", rec2.Code)
}
}
// TestConnectorFinishRegistrationUnauthenticated (P04, T9) verifies
// the finish endpoint also requires authentication.
func TestConnectorFinishRegistrationUnauthenticated(t *testing.T) {
dbPath := filepath.Join(t.TempDir(), "webauthn-creds.db")
store, _ := NewStore(dbPath)
defer store.Close()
c, _ := NewConnector(store, "test.cluster", "https://test.cluster")
mux := c.Routes()
req := httptest.NewRequest("POST", "/orca/webauthn/register/finish?username=admin", nil)
rec := httptest.NewRecorder()
mux.ServeHTTP(rec, req)
if rec.Code != http.StatusUnauthorized {
t.Errorf("finish unauthenticated: %d, want 401", rec.Code)
}
}
+5 -1
View File
@@ -46,11 +46,15 @@ func NewStore(dbPath string) (*Store, error) {
if err := os.MkdirAll(filepath.Dir(dbPath), 0o700); err != nil {
return nil, fmt.Errorf("webauthn: mkdir: %w", err)
}
dsn := fmt.Sprintf("file:%s?_pragma=journal_mode(WAL)", dbPath)
// REQ-156 / P07 T1: busy_timeout(5000) so concurrent webauthn
// DB opens wait up to 5s for the writer instead of failing with
// SQLITE_BUSY. SetMaxOpenConns(1) serializes the connections.
dsn := fmt.Sprintf("file:%s?_pragma=journal_mode(WAL)&_pragma=busy_timeout(5000)", dbPath)
db, err := sql.Open("sqlite", dsn)
if err != nil {
return nil, fmt.Errorf("webauthn: open db: %w", err)
}
db.SetMaxOpenConns(1)
if err := db.Ping(); err != nil {
db.Close()
return nil, fmt.Errorf("webauthn: ping: %w", err)
+3 -3
View File
@@ -273,7 +273,7 @@ func TestScenario_ACL(t *testing.T) {
t.Fatalf("mkdir cluster dir: %v", err)
}
a := acl.NewACL()
id := acl.Identity{Kind: acl.KindToken, ID: "operator-1"}
id := acl.Identity{Kind: acl.KindOidc, ID: "operator-1"}
a.Grant(id, "prod", acl.PermRead|acl.PermWrite)
if !a.Check(id, "prod", acl.PermRead) {
t.Error("expected read on prod after grant")
@@ -287,7 +287,7 @@ func TestScenario_ACL(t *testing.T) {
if a.Check(id, "staging", acl.PermRead) {
t.Error("cross-ns read should be denied")
}
admin := acl.Identity{Kind: acl.KindToken, ID: "root"}
admin := acl.Identity{Kind: acl.KindOidc, ID: "root"}
a.Grant(admin, "prod", acl.PermAdmin)
if !a.Check(admin, "prod", acl.PermRead) {
t.Error("admin should imply read")
@@ -303,7 +303,7 @@ func TestScenario_ACL(t *testing.T) {
if err != nil {
t.Fatalf("marshal acl: %v", err)
}
if err := writeAtomic(paths.ACLPath(), data, 0o644); err != nil {
if err := writeAtomic(paths.ACLPath(), data, 0o600); err != nil {
t.Fatalf("write acl.json: %v", err)
}
loaded, err := os.ReadFile(paths.ACLPath())
+39 -1
View File
@@ -56,6 +56,44 @@ func TestSecurityInvariants_Metadata(t *testing.T) {
// F18: drift event auth. Tested by:
// - internal/drift: TestVerifyEventSignature
//
// P04 (v0.13) ACL enforcement wiring (C-44/C-45):
// - internal/daemon: TestACLPolicyDenyByDefault
// (authenticated request with no ACL entry → deny in enforce mode)
// - internal/daemon: TestACLPolicyAllowWithEntry
// (authenticated request with matching ACL entry → allow)
// - internal/daemon: TestACLPolicyUnauthenticatedEnforce
// (unauthenticated request → 403 in enforce mode)
// - internal/daemon: TestACLPolicyLogOnlyAllowsDenials (C-45)
// (denials logged but allowed in log-only mode)
// - internal/daemon: TestACLJobsHandlerEnforceDeniesUnauthenticated
// (wired jobs handler denies unauthenticated in enforce mode)
// - internal/daemon: TestACLJobsHandlerAllowsAuthenticatedWithEntry
// (wired jobs handler allows authenticated with matching entry)
// - internal/daemon: TestACLNodesHandlerEnforceDeniesUnauthenticated
// - internal/daemon: TestACLTasksHandlerEnforceDeniesUnauthenticated
// - internal/txn: Apply refuses when ORCA_OIDC_TOKEN is missing/invalid
// (C-44: SSH-push applier + txn apply path validate OIDC token)
// - internal/sshpush: AuthorizeApply validates ORCA_OIDC_TOKEN
//
// P04 (v0.13) WebAuthn registration auth (C-45, T9):
// - internal/webauthn: TestConnectorBeginRegistrationUnauthenticated
// (unauthenticated BeginRegistration → 401, fail-closed)
// - internal/webauthn: TestConnectorFinishRegistrationUnauthenticated
// (unauthenticated FinishRegistration → 401)
//
// P04 (v0.13) acl.json hardening (T6/T7):
// - internal/cli: saveACL writes acl.json with mode 0600 (T6)
// - internal/cli: lockACL flocks grant/revoke (T7, prevents races)
//
// P04 (v0.13) bootstrap ACL (T8, C-40):
// - internal/cli: bootstrapACL grants cluster-admin to orca-admins
// group + init SVID on `orca init` (prevents operator lockout)
//
// P04 (v0.13) audit actor identity (T5):
// - internal/cli: currentActor reads OIDC sub from credentials.json
// - internal/engine: ActorFromCtx threads sub into audit Record calls
// (replaces hardcoded "cli" actor)
//
// This test is the gate (C-33): if it runs, the suite is wired.
t.Log("security integration test suite wired (R-021, F1-F25, REQ-119..148)")
t.Log("security integration test suite wired (R-021, F1-F25, REQ-119..148, P04 ACL enforcement C-44/C-45)")
}