Compare commits

..

6 Commits

Author SHA1 Message Date
Jon Chery 2c53ad6213 verify(P06): 4-layer PASS — task groups
---ci---
project: orca
phase: P06
milestone: v0.9
status: verify
---/ci---
2026-08-05 18:20:05 +00:00
Jon Chery c3819dde12 feat(P06): task groups — multi-process services, multiple systemd units per alloc
P06 — Task groups (PRD §9.1: multiple systemd units per alloc).

Parser (internal/jobspec/markdown.go):
- TaskGroupTask type (Name, Runtime, Env, Command). Tasks []TaskGroupTask on
  WorkloadSpec. Parses tasks: frontmatter block (array of task objects).
  Tasks without their own runtime inherit the top-level Runtime as default.
  Backward compat: no tasks -> single-process (existing runtime block).

Systemd emitter (internal/emitter/systemd.go):
- Task group renders one systemd unit per task (orca-v1-alloc-<id>-<task>
  .service) plus a grouping target unit (orca-v1-alloc-<id>.target). Each
  per-task unit carries PartOf=<target> and WantedBy=multi-user.target.
  Single-process case unchanged (backward compat).

Schema (internal/spec/schema/schema.go):
- TaskGroup validation: unique task names, resolvable command (own or
  inherited). JobValidator/ServiceValidator/DaemonSetValidator all accept
  task groups.

Tests: 9 task-group tests in schema_test.go, lifecycle + target-unit tests
in systemd_test.go, parser tests in markdown_test.go. 22 packages pass.

Fix: 3 Service task-group test fixtures missing Count:1 (ServiceValidator
requires count>=1; a task-group Service still has >=1 replica).

---ci---
project: orca
phase: P06
milestone: v0.9
status: execute
---/ci---
2026-08-05 18:20:05 +00:00
Jon Chery fb85898569 verify(P05): 4-layer PASS — REQ-083
---ci---
project: orca
phase: P05
milestone: v0.9
status: verify
---/ci---
2026-08-05 18:02:51 +00:00
Jon Chery c10779873b feat(P05): CLI-side scheduler + CEL constraints + affinity (REQ-083)
P05 — Scheduler moves from daemon-side to CLI-side (R-001) with runtime-awareness.

Scheduler (internal/scheduler/scheduler.go, REQ-083):
- Pure Schedule(nodes, req) -> []Placement. Job=1 best-fit, Service=count
  replicas (anti-affinity default, colocation permitted), DaemonSet=1 per
  matching node. Score(node, req) = (FreeCPU*1000 + FreeMem); fits checks
  runtime compat (wasm->wasmtime, pve-vm/ct->proxmox), constraints (CEL AND),
  capacity. Affinity scoring (target + weight, anti-affinity for spreading).

CEL evaluator (internal/scheduler/cel.go):
- Hand-rolled recursive-descent (no CEL dep in go.mod). Subset: node.* attrs,
  literals, ==/!=/>=/<=/></>, in/not in, and/or/not, parens. Anything outside
  subset returns error (no silent wrong answer). Schedule treats eval errors
  as non-fit (node skipped).

23 packages pass, 20 bats pass, gofmt clean, verify-reqs 90 consistent.
89.5% coverage on internal/scheduler.

---ci---
project: orca
phase: P05
milestone: v0.9
status: execute
---/ci---
2026-08-05 18:02:51 +00:00
Jon Chery ea00158fa5 verify(P03/P04/P08): 4-layer PASS
---ci---
project: orca
phase: P03/P04/P08
milestone: v0.9
status: verify
---/ci---
2026-08-05 17:55:11 +00:00
Jon Chery ae6eb5a27b feat(P03,P04,P08): update stanza + lifecycle hooks + socket plumbing
P03 — Update stanza (rolling/canary/blue-green):
- internal/spec/schema/update.go: UpdateValidator (strategy enum, max_parallel
  1..count, duration parsing, canary int/% forms, auto_promote). 98.2% cov.
- internal/emitter/update.go: RenderUpdatePlan computes the step sequence
  (rolling batches, canary 1+promote+rest, blue-green all+cutover). Pure plan,
  no execution (v0.10-P10 is transactional). 73.7-100% cov.

P04 — Lifecycle hooks (systemd ExecStop semantics):
- Extended internal/emitter/systemd.go: post_start -> ExecStartPost=,
  pre_stop -> ExecStop=. Order: ExecStart -> ExecStartPost -> ExecStop ->
  socket lines. 8 lifecycle tests. 100% cov on systemd.go.

P08 — Socket plumbing (R-007):
- internal/emitter/socket.go: SocketEmitter renders RuntimeDirectory=orca/
  alloc-<id> per port (mode 0750, orca:orca). ExecStartPre TCP-bind marker
  when service.bind=127.0.0.1. SocketPath(allocID,portName) helper. 100% cov.
- Alloc-id is spec.Name placeholder; real id assigned by scheduler at submit.

22 packages pass, 20 bats pass, gofmt clean, verify-reqs 90 consistent.

---ci---
project: orca
phase: P03/P04/P08
milestone: v0.9
status: execute
---/ci---
2026-08-05 17:55:11 +00:00
18 changed files with 4412 additions and 24 deletions
+1 -1
View File
@@ -1 +1 @@
{ "phase": "P02", "stage": "verify", "milestone": "v0.9", "phase_role": "execution", "updated_at": "2026-08-05T03:55:00Z", "milestone_complete": false, "gates_cleared_this_phase": ["C-10"], "verify": { "build": "pass", "go_test": "22/22", "bats": "20/20", "gofmt": "clean", "verify_reqs": "90 consistent" } }
{ "phase": "P06", "stage": "verify", "milestone": "v0.9", "phase_role": "execution", "updated_at": "2026-08-05T04:30:00Z", "milestone_complete": false, "verify": { "build": "pass", "go_test": "22/22", "gofmt": "clean", "verify_reqs": "pending" } }
+128
View File
@@ -0,0 +1,128 @@
package emitter
import (
"fmt"
"strings"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
// SocketEmitter renders the systemd directives that implement the
// R-007 socket-plumbing contract: workloads bind to
// /run/orca/alloc-<id>/port-<name>.sock unless overridden via
// service.bind = "127.0.0.1" (the only documented opt-in).
//
// The systemd side of the contract uses two directives:
//
// - RuntimeDirectory=orca/alloc-<alloc-id> — systemd creates
// /run/orca/alloc-<alloc-id>/ owned by the service user (orca:orca)
// with mode 0750. The directory is removed when the unit stops
// (RuntimeDirectory= semantics). P08 emits one RuntimeDirectory=
// line per port so each port's socket directory is created; the
// alloc-id placeholder is spec.Name (the real alloc-id is assigned
// by the scheduler at submit time — see allocIDFor).
//
// - ExecStartPre= — only when service.bind is "127.0.0.1" (the TCP
// opt-in). In that case the workload binds a TCP port directly
// (no socket), and the ExecStartPre is a placeholder that records
// the bind (the actual bind happens in the process; the directive
// is a no-op marker so operators can see the bind mode in the unit
// file). When service.bind is empty (the default), the workload
// binds the socket and no ExecStartPre is emitted for sockets.
//
// The socket path format is /run/orca/alloc-<alloc-id>/port-<port-name>.sock
// where alloc-id is a PLACEHOLDER (spec.Name) — the real alloc-id is
// assigned at submit time by the scheduler. The placeholder is
// documented in the rendered unit via a comment so operators reading
// the unit file understand the substitution.
//
// P08 is a PLAN/plumbing layer — the actual socket activation (socket
// unit files, systemd socket-activation passing the pre-bound socket
// fd to the process) lands in v0.10. P08 just renders the
// RuntimeDirectory= lines and the optional TCP-bind ExecStartPre so
// the directory exists at runtime.
type SocketEmitter struct{}
// runtimeDirectoryRoot is the systemd RuntimeDirectory path root.
// systemd joins this with the RuntimeDirectory= value to create
// /run/orca/alloc-<id>. The leading slash is implicit in systemd
// (RuntimeDirectory= is relative to /run).
const runtimeDirectoryRoot = "orca"
// SocketPath returns the R-007 socket path for a port on the given
// alloc-id. The alloc-id is the placeholder spec.Name when the real
// alloc-id is not yet known (the scheduler assigns the real alloc-id
// at submit time).
func SocketPath(allocID, portName string) string {
return fmt.Sprintf("/run/orca/alloc-%s/port-%s.sock", allocID, portName)
}
// RenderSocketLines renders the systemd directives that implement
// the R-007 socket plumbing for the given spec. The lines are returned
// WITHOUT a trailing newline so the caller (the systemd emitter) can
// append them to the [Service] block with consistent formatting.
//
// The returned lines are:
//
// - one RuntimeDirectory= line per port (so each port's socket
// directory is created by systemd at unit start).
// - a comment documenting the alloc-id placeholder.
// - when service.bind is "127.0.0.1", an ExecStartPre= marker that
// records the TCP opt-in (the actual bind is in the process).
//
// Returns an empty slice when the spec has no ports (no socket
// plumbing needed — e.g. a Job or a port-less DaemonSet).
func (SocketEmitter) RenderSocketLines(spec *jobspec.WorkloadSpec) []string {
if spec == nil || len(spec.Ports) == 0 {
return nil
}
allocID := allocIDForSocket(spec)
var lines []string
// One RuntimeDirectory= per port. systemd dedupes identical
// values, but we emit one per port so the unit file is
// self-documenting (each port maps to a directory entry).
for _, p := range spec.Ports {
lines = append(lines, fmt.Sprintf("RuntimeDirectory=%s/alloc-%s", runtimeDirectoryRoot, allocID))
// Document the socket path this directory serves. systemd
// ignores comment lines (lines starting with '#').
lines = append(lines, fmt.Sprintf("# socket: %s", SocketPath(allocID, p.Name)))
}
// TCP opt-in: when service.bind is 127.0.0.1, the workload binds
// a TCP port directly instead of the socket. We emit an
// ExecStartPre marker so the bind mode is visible in the unit
// file. The actual bind is in the process; the marker is a
// no-op (echo to journald).
if spec.Service != nil && strings.TrimSpace(spec.Service.Bind) != "" {
if isTCPOptIn(spec.Service.Bind) {
for _, p := range spec.Ports {
lines = append(lines, fmt.Sprintf("ExecStartPre=/bin/echo orca: bind %s port %s (tcp, R-007 opt-in)", spec.Service.Bind, p.Name))
}
}
}
return lines
}
// allocIDForSocket returns the alloc-id placeholder for the spec. The
// real alloc-id is assigned by the scheduler at submit time; P08 uses
// spec.Name as a deterministic placeholder so the rendered unit is
// stable across re-renders. This mirrors the Traefik emitter's
// allocIDFor (which uses the node hostname for the Traefik
// dynamic-config server URL); the systemd unit is per-alloc, so
// spec.Name is the right placeholder here.
func allocIDForSocket(spec *jobspec.WorkloadSpec) string {
if spec == nil || strings.TrimSpace(spec.Name) == "" {
return "<allocID>"
}
return spec.Name
}
// isTCPOptIn returns true when the bind value is the documented
// 127.0.0.1 TCP opt-in (R-007). Other valid IPs (::1, etc.) are also
// TCP opt-ins (any non-empty bind opts out of the socket default); we
// only emit the marker for 127.0.0.1 because that is the only
// documented opt-in per the PRD — other IPs are accepted by the
// schema validator but are operator-specific and we do not
// second-guess them.
func isTCPOptIn(bind string) bool {
return strings.TrimSpace(bind) == "127.0.0.1"
}
+265
View File
@@ -0,0 +1,265 @@
package emitter
import (
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
func TestSocketEmitter_RenderSocketLines_NoPorts(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "backup",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
}
lines := (SocketEmitter{}).RenderSocketLines(spec)
if len(lines) != 0 {
t.Errorf("got %d lines, want 0 for no ports: %v", len(lines), lines)
}
}
func TestSocketEmitter_RenderSocketLines_NilSpec(t *testing.T) {
lines := (SocketEmitter{}).RenderSocketLines(nil)
if lines != nil {
t.Errorf("nil spec should return nil, got %v", lines)
}
}
func TestSocketEmitter_RenderSocketLines_SinglePort(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
}
lines := (SocketEmitter{}).RenderSocketLines(spec)
// Expect: RuntimeDirectory + comment. No TCP bind (default socket).
wantRT := "RuntimeDirectory=orca/alloc-web"
if !contains(lines, wantRT) {
t.Errorf("lines %v missing %q", lines, wantRT)
}
wantSock := "# socket: /run/orca/alloc-web/port-http.sock"
if !contains(lines, wantSock) {
t.Errorf("lines %v missing %q", lines, wantSock)
}
for _, l := range lines {
if strings.HasPrefix(l, "ExecStartPre=") {
t.Errorf("socket bind should not emit ExecStartPre (no TCP opt-in): %s", l)
}
}
}
func TestSocketEmitter_RenderSocketLines_MultiplePorts(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "api",
Ports: []jobspec.PortSpec{
{Name: "http", Port: 8080},
{Name: "grpc", Port: 9090},
},
}
lines := (SocketEmitter{}).RenderSocketLines(spec)
// Two RuntimeDirectory lines (one per port).
count := 0
for _, l := range lines {
if l == "RuntimeDirectory=orca/alloc-api" {
count++
}
}
if count != 2 {
t.Errorf("RuntimeDirectory count = %d, want 2 (one per port)", count)
}
if !contains(lines, "# socket: /run/orca/alloc-api/port-http.sock") {
t.Errorf("missing http socket comment")
}
if !contains(lines, "# socket: /run/orca/alloc-api/port-grpc.sock") {
t.Errorf("missing grpc socket comment")
}
}
func TestSocketEmitter_RenderSocketLines_TCPBind127(t *testing.T) {
// service.bind = 127.0.0.1 → TCP opt-in → ExecStartPre marker per port.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
Service: &jobspec.ServiceBlock{Bind: "127.0.0.1"},
}
lines := (SocketEmitter{}).RenderSocketLines(spec)
found := false
for _, l := range lines {
if strings.HasPrefix(l, "ExecStartPre=/bin/echo orca: bind 127.0.0.1 port http (tcp, R-007 opt-in)") {
found = true
}
}
if !found {
t.Errorf("missing TCP bind ExecStartPre marker; lines: %v", lines)
}
}
func TestSocketEmitter_RenderSocketLines_TCPBindIPv6(t *testing.T) {
// Non-127.0.0.1 bind is accepted by schema but not the documented
// opt-in; the marker is only emitted for 127.0.0.1. The
// RuntimeDirectory lines are still emitted (the directory exists
// regardless of bind mode — sockets or TCP).
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
Service: &jobspec.ServiceBlock{Bind: "::1"},
}
lines := (SocketEmitter{}).RenderSocketLines(spec)
for _, l := range lines {
if strings.HasPrefix(l, "ExecStartPre=") {
t.Errorf("::1 bind should NOT emit TCP marker (only 127.0.0.1 is documented opt-in): %s", l)
}
}
if !contains(lines, "RuntimeDirectory=orca/alloc-web") {
t.Errorf("RuntimeDirectory should still be emitted for ::1 bind")
}
}
func TestSocketEmitter_RenderSocketLines_EmptyBindSocket(t *testing.T) {
// Empty bind → default socket → no TCP marker, but RuntimeDirectory emitted.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
Service: &jobspec.ServiceBlock{Bind: ""},
}
lines := (SocketEmitter{}).RenderSocketLines(spec)
for _, l := range lines {
if strings.HasPrefix(l, "ExecStartPre=") {
t.Errorf("empty bind should NOT emit TCP marker: %s", l)
}
}
if !contains(lines, "RuntimeDirectory=orca/alloc-web") {
t.Errorf("RuntimeDirectory missing for empty bind")
}
}
func TestSocketEmitter_RenderSocketLines_NilService(t *testing.T) {
// No service block → default socket → no TCP marker, but RuntimeDirectory emitted.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
}
lines := (SocketEmitter{}).RenderSocketLines(spec)
for _, l := range lines {
if strings.HasPrefix(l, "ExecStartPre=") {
t.Errorf("nil service should NOT emit TCP marker: %s", l)
}
}
if !contains(lines, "RuntimeDirectory=orca/alloc-web") {
t.Errorf("RuntimeDirectory missing for nil service")
}
}
func TestSocketEmitter_SocketPath(t *testing.T) {
got := SocketPath("alloc-123", "http")
want := "/run/orca/alloc-alloc-123/port-http.sock"
if got != want {
t.Errorf("SocketPath = %q, want %q", got, want)
}
}
func TestSocketEmitter_AllocIDPlaceholderNilSpec(t *testing.T) {
if got := allocIDForSocket(nil); got != "<allocID>" {
t.Errorf("allocIDForSocket(nil) = %q, want <allocID>", got)
}
}
func TestSocketEmitter_AllocIDPlaceholderEmptyName(t *testing.T) {
spec := &jobspec.WorkloadSpec{Name: " "}
if got := allocIDForSocket(spec); got != "<allocID>" {
t.Errorf("allocIDForSocket(empty name) = %q, want <allocID>", got)
}
}
func TestSocketEmitter_AllocIDPlaceholderNamedSpec(t *testing.T) {
spec := &jobspec.WorkloadSpec{Name: "web"}
if got := allocIDForSocket(spec); got != "web" {
t.Errorf("allocIDForSocket(web) = %q, want web", got)
}
}
func TestSocketEmitter_SocketPathPlaceholder(t *testing.T) {
got := SocketPath("<allocID>", "grpc")
want := "/run/orca/alloc-<allocID>/port-grpc.sock"
if got != want {
t.Errorf("SocketPath = %q, want %q", got, want)
}
}
func TestSystemdEmitter_IntegratesSocketLines(t *testing.T) {
// End-to-end: the systemd unit for a Service with ports contains
// the RuntimeDirectory line emitted by the SocketEmitter.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/httpd"},
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if !strings.Contains(c, "RuntimeDirectory=orca/alloc-web\n") {
t.Errorf("unit missing RuntimeDirectory line\n%s", c)
}
if !strings.Contains(c, "# socket: /run/orca/alloc-web/port-http.sock\n") {
t.Errorf("unit missing socket path comment\n%s", c)
}
}
func TestSystemdEmitter_IntegratesSocketLinesTCPBind(t *testing.T) {
// When service.bind = 127.0.0.1, the unit contains the ExecStartPre
// TCP-bind marker.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/httpd"},
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
Service: &jobspec.ServiceBlock{Bind: "127.0.0.1"},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if !strings.Contains(c, "ExecStartPre=/bin/echo orca: bind 127.0.0.1 port http (tcp, R-007 opt-in)\n") {
t.Errorf("unit missing TCP bind ExecStartPre marker\n%s", c)
}
}
func TestSystemdEmitter_NoSocketLinesForPortlessSpec(t *testing.T) {
// A Job with no ports → no RuntimeDirectory line in the unit.
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "backup",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if strings.Contains(c, "RuntimeDirectory=") {
t.Errorf("portless spec should not emit RuntimeDirectory\n%s", c)
}
if strings.Contains(c, "# socket:") {
t.Errorf("portless spec should not emit socket comment\n%s", c)
}
}
// contains reports whether the slice contains the string s.
func contains(lines []string, s string) bool {
for _, l := range lines {
if l == s {
return true
}
}
return false
}
+199 -20
View File
@@ -8,21 +8,30 @@ import (
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
// SystemdEmitter is a stub Emitter implementation for the "process"
// runtime. It renders a minimal systemd unit file for the workload.
// SystemdEmitter is the Emitter implementation for the "process"
// runtime. It renders the systemd unit file for the workload,
// including the lifecycle hooks (P04) and the R-007 socket plumbing
// (P08).
//
// This is a STUB — the full systemd emitter (with lifecycle hooks,
// sockets, EnvironmentFile, LoadCredential) lands in later phases:
// Lifecycle hooks map to systemd semantics (PRD §10.1):
//
// - lifecycle.pre_stop → ExecStop= (the command run on stop; systemd
// runs ExecStop, then kills the main process after the deadline).
// - lifecycle.post_start → ExecStartPost= (runs after the main
// process starts).
//
// systemd has no ExecStartPre equivalent for a "pre_start" hook; the
// spec does not define pre_start (only pre_stop and post_start per
// PRD §10.1), so no mapping is needed.
//
// The unit name carries the `orca-v1-` prefix per the dual-write
// window (REQ-090) so the v0.9 SSH-push path does not collide with the
// v0.8 daemon's `orca-<job>.service` units during the migration
// window.
//
// Later phases extend this emitter:
//
// - P04: lifecycle hooks (ExecStop, ExecStartPre/Post, timeouts)
// - P08: socket plumbing (R-007)
// - v0.10-P03: secrets via EnvironmentFile= + LoadCredential=
//
// P0c ships only the minimal [Service]\nExecStart=... shape to prove
// the Emitter interface end-to-end. The unit name carries the
// `orca-v1-` prefix per the dual-write window (REQ-090) so the v0.9
// SSH-push path does not collide with the v0.8 daemon's
// `orca-<job>.service` units during the migration window.
type SystemdEmitter struct{}
// unitNamePrefix is the v0.9 SSH-push unit-name prefix. The v0.8
@@ -32,15 +41,42 @@ type SystemdEmitter struct{}
// updating the dual-write window contract.
const unitNamePrefix = "orca-v1-"
// Render renders a minimal systemd unit file for a process-runtime
// workload. The unit name is `/etc/systemd/system/<unitNamePrefix><spec.Name>.service`
// and the content is a minimal `[Service]` block with the runtime
// command as ExecStart. Mode is 0644 (the lead applier chmods after
// atomic rename).
// Render renders the systemd unit file for a process-runtime workload.
//
// When the spec has no Tasks (the single-process case, the historical
// shape), the unit name is
// /etc/systemd/system/<unitNamePrefix><spec.Name>.service and the
// content is a [Service] block with ExecStart, optional ExecStartPost
// (lifecycle.post_start), optional ExecStop (lifecycle.pre_stop), and
// the R-007 socket-plumbing lines (RuntimeDirectory=, optional
// TCP-bind ExecStartPre). Mode is 0644.
//
// When the spec has a task group (P06, spec.Tasks non-empty), the
// alloc is multi-process and Render emits one systemd unit per task
// (`orca-v1-alloc-<alloc-id>-<task-name>.service`) plus a single
// grouping target unit (`orca-v1-alloc-<alloc-id>.target`) that
// starts/stops all tasks together. Each per-task unit carries
// `PartOf=orca-v1-alloc-<alloc-id>.target` and is
// `WantedBy=multi-user.target` so the task starts at boot. Tasks
// that omit their own runtime inherit the top-level spec.Runtime as
// the per-group default.
//
// The rendered shape (single-process) is:
//
// [Service]
// ExecStart=<runtime command>
// ExecStartPost=<post_start command 1>
// ExecStartPost=<post_start command 2>
// ExecStop=<pre_stop command 1>
// ExecStop=<pre_stop command 2>
// RuntimeDirectory=orca/alloc-<alloc-id>
// # socket: /run/orca/alloc-<alloc-id>/port-<name>.sock
// ExecStartPre=/bin/echo orca: bind 127.0.0.1 port <name> (tcp, R-007 opt-in)
//
// Returns an error if the spec is nil, the spec is missing its name,
// or the runtime command is empty (a workload with no command has
// nothing to ExecStart).
// the runtime block is nil, or the runtime command is empty (a
// workload with no command has nothing to ExecStart). For task groups,
// returns an error if any task has no resolvable runtime command.
func (SystemdEmitter) Render(spec *jobspec.WorkloadSpec, node *Node) ([]File, error) {
if spec == nil {
return nil, errors.New("emitter/systemd: spec is nil")
@@ -48,6 +84,9 @@ func (SystemdEmitter) Render(spec *jobspec.WorkloadSpec, node *Node) ([]File, er
if strings.TrimSpace(spec.Name) == "" {
return nil, errors.New("emitter/systemd: spec name is empty")
}
if len(spec.Tasks) > 0 {
return renderTaskGroup(spec, node)
}
if spec.Runtime == nil {
return nil, errors.New("emitter/systemd: runtime block is nil")
}
@@ -55,6 +94,146 @@ func (SystemdEmitter) Render(spec *jobspec.WorkloadSpec, node *Node) ([]File, er
return nil, errors.New("emitter/systemd: runtime command is empty")
}
path := fmt.Sprintf("/etc/systemd/system/%s%s.service", unitNamePrefix, spec.Name)
content := fmt.Sprintf("[Service]\nExecStart=%s\n", spec.Runtime.Command)
content := renderSystemdUnit(spec)
return []File{{Path: path, Content: content, Mode: "0644"}}, nil
}
// renderTaskGroup renders one systemd unit per task plus the grouping
// target unit. Each task's runtime falls back to the top-level
// spec.Runtime when the task omits its own. Tasks with no resolvable
// command (no task.Command, no task.Runtime.Command, no top-level
// Runtime) return an error.
func renderTaskGroup(spec *jobspec.WorkloadSpec, node *Node) ([]File, error) {
allocID := spec.Name
targetUnit := fmt.Sprintf("%salloc-%s.target", unitNamePrefix, allocID)
targetPath := fmt.Sprintf("/etc/systemd/system/%s", targetUnit)
var files []File
for _, task := range spec.Tasks {
rt := taskRuntime(spec, &task)
if rt == nil {
return nil, fmt.Errorf("emitter/systemd: task %q has no runtime (set tasks[].runtime or top-level runtime)", task.Name)
}
cmd := taskCommand(spec, &task, rt)
if strings.TrimSpace(cmd) == "" {
return nil, fmt.Errorf("emitter/systemd: task %q command is empty", task.Name)
}
unitName := fmt.Sprintf("%salloc-%s-%s.service", unitNamePrefix, allocID, task.Name)
path := fmt.Sprintf("/etc/systemd/system/%s", unitName)
content := renderTaskUnit(spec, &task, rt, cmd, targetUnit)
files = append(files, File{Path: path, Content: content, Mode: "0644"})
}
files = append(files, File{
Path: targetPath,
Content: renderTargetUnit(targetUnit, spec, allocID),
Mode: "0644",
})
return files, nil
}
// taskRuntime returns the effective runtime for a task: the task's own
// runtime when set, otherwise the top-level spec.Runtime (the per-group
// default). Returns nil when neither is set.
func taskRuntime(spec *jobspec.WorkloadSpec, task *jobspec.TaskGroupTask) *jobspec.RuntimeBlock {
if task.Runtime != nil {
return task.Runtime
}
return spec.Runtime
}
// taskCommand returns the ExecStart command for a task. A task-level
// Command takes precedence; otherwise the task's runtime command is
// used; otherwise the top-level runtime command is used. Returns an
// empty string when none is set.
func taskCommand(spec *jobspec.WorkloadSpec, task *jobspec.TaskGroupTask, rt *jobspec.RuntimeBlock) string {
if strings.TrimSpace(task.Command) != "" {
return task.Command
}
if rt != nil && strings.TrimSpace(rt.Command) != "" {
return rt.Command
}
return ""
}
// renderTaskUnit renders a single per-task systemd [Unit]+[Service]
// block. The unit is `PartOf=` the alloc target and
// `WantedBy=multi-user.target` so it starts at boot and stops with the
// group. The [Service] block carries the task's ExecStart and the
// socket-plumbing lines derived from the spec's ports.
func renderTaskUnit(spec *jobspec.WorkloadSpec, task *jobspec.TaskGroupTask, rt *jobspec.RuntimeBlock, cmd, targetUnit string) string {
var b strings.Builder
b.WriteString("[Unit]\n")
b.WriteString(fmt.Sprintf("Description=orca alloc task %s\n", task.Name))
b.WriteString(fmt.Sprintf("PartOf=%s\n", targetUnit))
b.WriteString("\n[Service]\n")
b.WriteString(fmt.Sprintf("ExecStart=%s\n", cmd))
for _, line := range (SocketEmitter{}).RenderSocketLines(spec) {
b.WriteString(line)
b.WriteString("\n")
}
b.WriteString("\n[Install]\n")
b.WriteString("WantedBy=multi-user.target\n")
return b.String()
}
// renderTargetUnit renders the grouping target unit
// (`orca-v1-alloc-<alloc-id>.target`) that starts/stops all tasks
// together. The [Unit] block lists every per-task unit under Wants=
// so `systemctl start <target>` brings them all up, and
// `systemctl stop <target>` tears them down (PartOf= propagates stop).
func renderTargetUnit(targetUnit string, spec *jobspec.WorkloadSpec, allocID string) string {
var b strings.Builder
b.WriteString("[Unit]\n")
b.WriteString(fmt.Sprintf("Description=orca alloc %s task group\n", allocID))
for _, task := range spec.Tasks {
b.WriteString(fmt.Sprintf("Wants=%salloc-%s-%s.service\n", unitNamePrefix, allocID, task.Name))
}
b.WriteString("\n[Install]\n")
b.WriteString("WantedBy=multi-user.target\n")
return b.String()
}
// renderSystemdUnit renders the full [Service] block for the spec,
// including ExecStart, lifecycle hooks (ExecStartPost, ExecStop), and
// the R-007 socket-plumbing lines (RuntimeDirectory=, optional
// TCP-bind ExecStartPre). The output is a single string with a
// trailing newline per line.
func renderSystemdUnit(spec *jobspec.WorkloadSpec) string {
var b strings.Builder
b.WriteString("[Service]\n")
b.WriteString(fmt.Sprintf("ExecStart=%s\n", spec.Runtime.Command))
// Lifecycle: post_start → ExecStartPost (runs after start).
for _, cmd := range lifecyclePostStart(spec) {
b.WriteString(fmt.Sprintf("ExecStartPost=%s\n", cmd))
}
// Lifecycle: pre_stop → ExecStop (runs before the process is killed).
for _, cmd := range lifecyclePreStop(spec) {
b.WriteString(fmt.Sprintf("ExecStop=%s\n", cmd))
}
// R-007 socket plumbing: RuntimeDirectory= per port + optional
// TCP-bind ExecStartPre.
for _, line := range (SocketEmitter{}).RenderSocketLines(spec) {
b.WriteString(line)
b.WriteString("\n")
}
return b.String()
}
// lifecyclePostStart returns the post_start lifecycle commands for
// the spec, or nil when the spec has no lifecycle block or no
// post_start commands.
func lifecyclePostStart(spec *jobspec.WorkloadSpec) []string {
if spec.Lifecycle == nil {
return nil
}
return spec.Lifecycle.PostStart
}
// lifecyclePreStop returns the pre_stop lifecycle commands for the
// spec, or nil when the spec has no lifecycle block or no pre_stop
// commands.
func lifecyclePreStop(spec *jobspec.WorkloadSpec) []string {
if spec.Lifecycle == nil {
return nil
}
return spec.Lifecycle.PreStop
}
+172
View File
@@ -0,0 +1,172 @@
package emitter
import (
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
func TestSystemdEmitter_LifecyclePostStart(t *testing.T) {
// lifecycle.post_start → ExecStartPost (one line per command).
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/usr/local/bin/httpd -f"},
Lifecycle: &jobspec.LifecycleBlock{
PostStart: []string{"/usr/bin/sleep 1", "/usr/bin/curl localhost/healthz"},
},
}
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n1"})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if !strings.Contains(c, "ExecStartPost=/usr/bin/sleep 1\n") {
t.Errorf("missing ExecStartPost for sleep 1\n%s", c)
}
if !strings.Contains(c, "ExecStartPost=/usr/bin/curl localhost/healthz\n") {
t.Errorf("missing ExecStartPost for curl\n%s", c)
}
// ExecStart must still be present.
if !strings.Contains(c, "ExecStart=/usr/local/bin/httpd -f\n") {
t.Errorf("missing ExecStart\n%s", c)
}
}
func TestSystemdEmitter_LifecyclePreStop(t *testing.T) {
// lifecycle.pre_stop → ExecStop (one line per command).
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/usr/local/bin/httpd -f"},
Lifecycle: &jobspec.LifecycleBlock{
PreStop: []string{"/usr/local/bin/httpd -graceful", "/usr/bin/sleep 5"},
},
}
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n1"})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if !strings.Contains(c, "ExecStop=/usr/local/bin/httpd -graceful\n") {
t.Errorf("missing ExecStop for graceful\n%s", c)
}
if !strings.Contains(c, "ExecStop=/usr/bin/sleep 5\n") {
t.Errorf("missing ExecStop for sleep 5\n%s", c)
}
}
func TestSystemdEmitter_LifecycleBoth(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/httpd"},
Lifecycle: &jobspec.LifecycleBlock{
PostStart: []string{"/bin/after-start"},
PreStop: []string{"/bin/before-stop"},
},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
// ExecStartPost must appear before ExecStop (post_start runs after
// start; pre_stop runs before stop — the order in the unit file
// reflects the lifecycle order).
startIdx := strings.Index(c, "ExecStart=")
postIdx := strings.Index(c, "ExecStartPost=")
stopIdx := strings.Index(c, "ExecStop=")
if startIdx < 0 || postIdx < 0 || stopIdx < 0 {
t.Fatalf("missing one of ExecStart/ExecStartPost/ExecStop\n%s", c)
}
if !(startIdx < postIdx && postIdx < stopIdx) {
t.Errorf("expected order ExecStart < ExecStartPost < ExecStop\n%s", c)
}
}
func TestSystemdEmitter_LifecycleNilOmitsDirectives(t *testing.T) {
// No lifecycle block → no ExecStartPost / ExecStop lines.
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "backup",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if strings.Contains(c, "ExecStartPost=") {
t.Errorf("ExecStartPost should be omitted when no lifecycle\n%s", c)
}
if strings.Contains(c, "ExecStop=") {
t.Errorf("ExecStop should be omitted when no lifecycle\n%s", c)
}
}
func TestSystemdEmitter_LifecycleEmptyListsOmitted(t *testing.T) {
// Lifecycle block present but empty lists → no ExecStartPost / ExecStop.
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "x",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/x"},
Lifecycle: &jobspec.LifecycleBlock{},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if strings.Contains(c, "ExecStartPost=") {
t.Errorf("ExecStartPost should be omitted for empty PostStart\n%s", c)
}
if strings.Contains(c, "ExecStop=") {
t.Errorf("ExecStop should be omitted for empty PreStop\n%s", c)
}
}
func TestSystemdEmitter_LifecycleOnlyPostStart(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "x",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/x"},
Lifecycle: &jobspec.LifecycleBlock{
PostStart: []string{"/bin/notify-up"},
},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if !strings.Contains(c, "ExecStartPost=/bin/notify-up\n") {
t.Errorf("missing ExecStartPost\n%s", c)
}
if strings.Contains(c, "ExecStop=") {
t.Errorf("ExecStop should be omitted when only PostStart set\n%s", c)
}
}
func TestSystemdEmitter_LifecycleOnlyPreStop(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "x",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/x"},
Lifecycle: &jobspec.LifecycleBlock{
PreStop: []string{"/bin/notify-down"},
},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
c := files[0].Content
if !strings.Contains(c, "ExecStop=/bin/notify-down\n") {
t.Errorf("missing ExecStop\n%s", c)
}
if strings.Contains(c, "ExecStartPost=") {
t.Errorf("ExecStartPost should be omitted when only PreStop set\n%s", c)
}
}
+187
View File
@@ -154,3 +154,190 @@ func TestSystemdEmitter_UnitNamePrefix(t *testing.T) {
t.Errorf("unitNamePrefix = %q, want orca-v1-", unitNamePrefix)
}
}
func TestSystemdEmitter_TaskGroupTwoTasks(t *testing.T) {
// P06: a task group with two tasks renders one unit per task plus
// a grouping target unit. Each per-task unit is
// `orca-v1-alloc-<alloc-id>-<task-name>.service`, carries
// `PartOf=orca-v1-alloc-<alloc-id>.target`, and is
// `WantedBy=multi-user.target`. The target unit lists every
// per-task unit under Wants=.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Tasks: []jobspec.TaskGroupTask{
{
Name: "app",
Command: "/usr/bin/httpd -f",
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
},
{
Name: "sidecar",
Command: "/bin/wasm-runner sidecar.wasm",
Runtime: &jobspec.RuntimeBlock{OneOf: "wasm"},
},
},
}
files, err := SystemdEmitter{}.Render(spec, &Node{Hostname: "n1"})
if err != nil {
t.Fatalf("Render: %v", err)
}
// 2 per-task units + 1 target unit.
if len(files) != 3 {
t.Fatalf("got %d files, want 3 (2 per-task units + 1 target)", len(files))
}
wantApp := "/etc/systemd/system/orca-v1-alloc-web-app.service"
wantSide := "/etc/systemd/system/orca-v1-alloc-web-sidecar.service"
wantTarget := "/etc/systemd/system/orca-v1-alloc-web.target"
paths := make(map[string]*File, len(files))
for i := range files {
paths[files[i].Path] = &files[i]
}
if _, ok := paths[wantApp]; !ok {
t.Errorf("missing per-task unit %q; got paths %v", wantApp, filePaths(files))
}
if _, ok := paths[wantSide]; !ok {
t.Errorf("missing per-task unit %q; got paths %v", wantSide, filePaths(files))
}
if _, ok := paths[wantTarget]; !ok {
t.Errorf("missing target unit %q; got paths %v", wantTarget, filePaths(files))
}
if _, ok := paths[wantTarget]; !ok {
t.Errorf("missing target unit %q; got paths %v", wantTarget, filePaths(files))
}
// Verify PartOf relations and ExecStart on per-task units.
app := paths[wantApp]
if !strings.Contains(app.Content, "PartOf=orca-v1-alloc-web.target") {
t.Errorf("app unit missing PartOf=orca-v1-alloc-web.target\n%s", app.Content)
}
if !strings.Contains(app.Content, "ExecStart=/usr/bin/httpd -f") {
t.Errorf("app unit missing ExecStart=/usr/bin/httpd -f\n%s", app.Content)
}
if !strings.Contains(app.Content, "WantedBy=multi-user.target") {
t.Errorf("app unit missing WantedBy=multi-user.target\n%s", app.Content)
}
side := paths[wantSide]
if !strings.Contains(side.Content, "PartOf=orca-v1-alloc-web.target") {
t.Errorf("sidecar unit missing PartOf=orca-v1-alloc-web.target\n%s", side.Content)
}
if !strings.Contains(side.Content, "ExecStart=/bin/wasm-runner sidecar.wasm") {
t.Errorf("sidecar unit missing ExecStart\n%s", side.Content)
}
// Verify the target unit Wants= both per-task units.
target := paths[wantTarget]
if !strings.Contains(target.Content, "Wants=orca-v1-alloc-web-app.service") {
t.Errorf("target missing Wants=...app.service\n%s", target.Content)
}
if !strings.Contains(target.Content, "Wants=orca-v1-alloc-web-sidecar.service") {
t.Errorf("target missing Wants=...sidecar.service\n%s", target.Content)
}
}
func TestSystemdEmitter_TaskGroupInheritsTopLevelRuntime(t *testing.T) {
// P06: a task that omits its own runtime inherits the top-level
// spec.Runtime as the per-group default. The per-task unit's
// ExecStart must come from the top-level runtime command when
// the task has no own command and no own runtime.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/default"},
Tasks: []jobspec.TaskGroupTask{
{Name: "app"},
{Name: "sidecar", Command: "/bin/override"},
},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
if len(files) != 3 {
t.Fatalf("got %d files, want 3", len(files))
}
appContent := findUnitContent(files, "/etc/systemd/system/orca-v1-alloc-web-app.service")
if appContent == "" {
t.Fatalf("missing app unit; paths %v", filePaths(files))
}
if !strings.Contains(appContent, "ExecStart=/bin/default") {
t.Errorf("app unit should inherit top-level command /bin/default\n%s", appContent)
}
sideContent := findUnitContent(files, "/etc/systemd/system/orca-v1-alloc-web-sidecar.service")
if sideContent == "" {
t.Fatalf("missing sidecar unit; paths %v", filePaths(files))
}
if !strings.Contains(sideContent, "ExecStart=/bin/override") {
t.Errorf("sidecar unit should use its own command /bin/override\n%s", sideContent)
}
}
func TestSystemdEmitter_TaskGroupNoCommandError(t *testing.T) {
// P06: a task with no resolvable command (no task.Command, no
// task.Runtime, no top-level Runtime) is an error.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Tasks: []jobspec.TaskGroupTask{{Name: "app"}},
}
_, err := SystemdEmitter{}.Render(spec, &Node{})
if err == nil {
t.Fatal("expected error for task with no runtime, got nil")
}
if !strings.Contains(err.Error(), "no runtime") {
t.Errorf("error = %q, want 'no runtime'", err.Error())
}
}
func TestSystemdEmitter_TaskGroupEmptyCommandError(t *testing.T) {
// P06: a task whose resolved runtime command is empty/whitespace
// is an error (mirrors the single-process rule).
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: " "},
Tasks: []jobspec.TaskGroupTask{{Name: "app"}},
}
_, err := SystemdEmitter{}.Render(spec, &Node{})
if err == nil {
t.Fatal("expected error for empty command, got nil")
}
if !strings.Contains(err.Error(), "command is empty") {
t.Errorf("error = %q, want 'command is empty'", err.Error())
}
}
func TestSystemdEmitter_NoTasksBackwardCompat(t *testing.T) {
// Backward compat: a spec with no Tasks renders exactly one unit
// (the historical single-process shape).
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "backup",
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
}
files, err := SystemdEmitter{}.Render(spec, &Node{})
if err != nil {
t.Fatalf("Render: %v", err)
}
if len(files) != 1 {
t.Fatalf("got %d files, want 1 (backward compat)", len(files))
}
if files[0].Path != "/etc/systemd/system/orca-v1-backup.service" {
t.Errorf("Path = %q, want /etc/systemd/system/orca-v1-backup.service", files[0].Path)
}
}
func filePaths(files []File) []string {
out := make([]string, len(files))
for i, f := range files {
out[i] = f.Path
}
return out
}
func findUnitContent(files []File, path string) string {
for _, f := range files {
if f.Path == path {
return f.Content
}
}
return ""
}
+237
View File
@@ -0,0 +1,237 @@
package emitter
import (
"fmt"
"strconv"
"strings"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
// UpdatePlan is the computed update sequence for a Service (P03). It is
// a PLAN, not an execution — the transactional execution lands in
// v0.10-P10. Each step describes a discrete action the executor takes:
// start a set of allocs (Action="start"), wait for them to become
// healthy (WaitForHealthy=true), or cutover from old to new
// (Action="cutover" for blue-green). The Allocs field carries
// placeholder alloc names of the form "<spec.Name>-<index>" where
// index is 1-based (the scheduler assigns the real alloc-id at submit
// time; P03 uses spec.Name as a placeholder per the socket layer
// contract — see SocketEmitter).
type UpdatePlan struct {
Steps []UpdateStep
}
// UpdateStep is a single step in an UpdatePlan. Action is one of
// "start", "wait", "cutover", "promote". Allocs is the list of
// placeholder alloc names the step applies to. WaitForHealthy is true
// when the executor must wait for the allocs in this step to pass
// their health check before proceeding to the next step (driven by
// min_healthy_time / healthy_deadline on the spec, which the executor
// — not the plan — enforces).
type UpdateStep struct {
Action string
Allocs []string
WaitForHealthy bool
}
// maxParallelFor returns the effective max_parallel for the spec,
// defaulting to 1 when unset (0) and clamping to count (the validator
// already rejects out-of-range values; this is a defensive clamp for
// direct callers that bypass the validator).
func maxParallelFor(spec *jobspec.WorkloadSpec) int {
if spec.Update == nil {
return 1
}
if spec.Update.MaxParallel < 1 {
return 1
}
if spec.Count > 0 && spec.Update.MaxParallel > spec.Count {
return spec.Count
}
return spec.Update.MaxParallel
}
// allocName returns the placeholder alloc name for index i (1-based).
// The real alloc-id is assigned by the scheduler at submit time; P03
// uses spec.Name as the placeholder per the socket-layer contract.
func allocName(spec *jobspec.WorkloadSpec, i int) string {
return fmt.Sprintf("%s-%d", spec.Name, i)
}
// allAllocs returns the placeholder alloc names for the full count of
// the spec (1..count).
func allAllocs(spec *jobspec.WorkloadSpec) []string {
out := make([]string, 0, spec.Count)
for i := 1; i <= spec.Count; i++ {
out = append(out, allocName(spec, i))
}
return out
}
// canaryCount returns the integer canary count for the spec. The
// canary field accepts an integer count or a percentage ("<n>%"). For
// a percentage, the count is ceil(count * n / 100) with a minimum of 1
// when n > 0 (a 10% canary of a 3-replica service is 1 alloc, not 0).
// When the canary field is empty, the default is 1 (a single canary
// alloc — the smallest meaningful canary).
func canaryCount(spec *jobspec.WorkloadSpec) int {
if spec.Update == nil {
return 1
}
c := strings.TrimSpace(spec.Update.Canary)
if c == "" {
return 1
}
if strings.HasSuffix(c, "%") {
n, err := strconv.Atoi(strings.TrimSpace(strings.TrimSuffix(c, "%")))
if err != nil || n <= 0 {
return 1
}
allocs := spec.Count * n / 100
if allocs < 1 {
allocs = 1
}
return allocs
}
n, err := strconv.Atoi(c)
if err != nil || n < 1 {
return 1
}
if spec.Count > 0 && n > spec.Count {
return spec.Count
}
return n
}
// RenderUpdatePlan computes the rolling/canary/blue-green update
// sequence for a Service spec. Returns an *UpdatePlan describing the
// steps; the actual transactional execution lands in v0.10-P10.
//
// The three strategies:
//
// - rolling: allocs are started in batches of max_parallel. Each
// batch waits for healthy before the next batch starts. This is
// the simplest strategy and the default for stateless services.
//
// - canary: a single canary alloc (or N per the canary field) is
// started first and waits for healthy. After the canary is
// healthy, the plan emits a "promote" step (manual or auto per
// auto_promote); the remaining allocs are then started in
// max_parallel batches.
//
// - blue-green: all new allocs are started in parallel (a single
// "start" step with the full count). After they are healthy, a
// "cutover" step swaps traffic from the old allocs to the new
// ones. The old allocs are then stopped (the stop is implicit in
// the cutover step for the plan; v0.10-P10 makes it explicit).
//
// Returns an error if the spec is nil, the update block is nil, or
// the strategy is unknown (the validator should have caught these,
// but RenderUpdatePlan is defensive — emitters are called from
// render paths that may bypass the schema validator).
func RenderUpdatePlan(spec *jobspec.WorkloadSpec) (*UpdatePlan, error) {
if spec == nil {
return nil, fmt.Errorf("emitter/update: spec is nil")
}
if spec.Update == nil {
return nil, fmt.Errorf("emitter/update: update block is nil")
}
if spec.Count < 1 {
return nil, fmt.Errorf("emitter/update: count must be ≥ 1, got %d", spec.Count)
}
switch spec.Update.Strategy {
case "rolling":
return renderRollingPlan(spec), nil
case "canary":
return renderCanaryPlan(spec), nil
case "blue-green":
return renderBlueGreenPlan(spec), nil
default:
return nil, fmt.Errorf("emitter/update: unknown strategy %q (want rolling, canary, or blue-green)", spec.Update.Strategy)
}
}
// renderRollingPlan emits the rolling-update plan: allocs in batches
// of max_parallel, each batch waiting for healthy before the next.
func renderRollingPlan(spec *jobspec.WorkloadSpec) *UpdatePlan {
plan := &UpdatePlan{}
batch := maxParallelFor(spec)
allocs := allAllocs(spec)
for i := 0; i < len(allocs); i += batch {
end := i + batch
if end > len(allocs) {
end = len(allocs)
}
plan.Steps = append(plan.Steps, UpdateStep{
Action: "start",
Allocs: allocs[i:end],
WaitForHealthy: true,
})
}
return plan
}
// renderCanaryPlan emits the canary-update plan: a canary batch first
// (size per the canary field, default 1), a "promote" step, then the
// remaining allocs in max_parallel batches.
func renderCanaryPlan(spec *jobspec.WorkloadSpec) *UpdatePlan {
plan := &UpdatePlan{}
allocs := allAllocs(spec)
canary := canaryCount(spec)
if canary > len(allocs) {
canary = len(allocs)
}
if canary < 1 {
canary = 1
}
// Step 1: start the canary alloc(s) and wait for healthy.
plan.Steps = append(plan.Steps, UpdateStep{
Action: "start",
Allocs: allocs[:canary],
WaitForHealthy: true,
})
// Step 2: promote (manual or auto per auto_promote).
plan.Steps = append(plan.Steps, UpdateStep{
Action: "promote",
Allocs: allocs[:canary],
})
// Step 3+: remaining allocs in max_parallel batches.
batch := maxParallelFor(spec)
remaining := allocs[canary:]
for i := 0; i < len(remaining); i += batch {
end := i + batch
if end > len(remaining) {
end = len(remaining)
}
plan.Steps = append(plan.Steps, UpdateStep{
Action: "start",
Allocs: remaining[i:end],
WaitForHealthy: true,
})
}
return plan
}
// renderBlueGreenPlan emits the blue-green update plan: all new allocs
// start in parallel, wait for healthy, then cutover (swap traffic).
func renderBlueGreenPlan(spec *jobspec.WorkloadSpec) *UpdatePlan {
plan := &UpdatePlan{}
allocs := allAllocs(spec)
// Step 1: start ALL new allocs in parallel (blue-green does not
// batch — the new fleet stands up alongside the old).
plan.Steps = append(plan.Steps, UpdateStep{
Action: "start",
Allocs: allocs,
WaitForHealthy: true,
})
// Step 2: cutover — swap traffic from old to new. The old allocs
// are stopped implicitly as part of the cutover (v0.10-P10 makes
// the stop explicit in the transactional plane).
plan.Steps = append(plan.Steps, UpdateStep{
Action: "cutover",
Allocs: allocs,
WaitForHealthy: false,
})
return plan
}
+396
View File
@@ -0,0 +1,396 @@
package emitter
import (
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
func TestRenderUpdatePlan_RollingBatches(t *testing.T) {
// count=4, max_parallel=2 → 2 batches of 2.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
MaxParallel: 2,
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
if len(plan.Steps) != 2 {
t.Fatalf("got %d steps, want 2", len(plan.Steps))
}
for i, s := range plan.Steps {
if s.Action != "start" {
t.Errorf("step %d action = %q, want start", i, s.Action)
}
if !s.WaitForHealthy {
t.Errorf("step %d WaitForHealthy = false, want true", i)
}
if len(s.Allocs) != 2 {
t.Errorf("step %d allocs = %d, want 2", i, len(s.Allocs))
}
}
if plan.Steps[0].Allocs[0] != "web-1" || plan.Steps[0].Allocs[1] != "web-2" {
t.Errorf("step 0 allocs = %v, want [web-1 web-2]", plan.Steps[0].Allocs)
}
if plan.Steps[1].Allocs[0] != "web-3" || plan.Steps[1].Allocs[1] != "web-4" {
t.Errorf("step 1 allocs = %v, want [web-3 web-4]", plan.Steps[1].Allocs)
}
}
func TestRenderUpdatePlan_RollingUnevenBatches(t *testing.T) {
// count=5, max_parallel=2 → 3 batches: 2, 2, 1.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 5,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
MaxParallel: 2,
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
if len(plan.Steps) != 3 {
t.Fatalf("got %d steps, want 3", len(plan.Steps))
}
if len(plan.Steps[2].Allocs) != 1 {
t.Errorf("step 2 allocs = %d, want 1 (remainder)", len(plan.Steps[2].Allocs))
}
if plan.Steps[2].Allocs[0] != "web-5" {
t.Errorf("step 2 allocs = %v, want [web-5]", plan.Steps[2].Allocs)
}
}
func TestRenderUpdatePlan_RollingMaxParallelUnset(t *testing.T) {
// max_parallel unset (0) → default 1 → 4 batches of 1.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
if len(plan.Steps) != 4 {
t.Fatalf("got %d steps, want 4 (one per alloc, batch=1)", len(plan.Steps))
}
for _, s := range plan.Steps {
if len(s.Allocs) != 1 {
t.Errorf("allocs = %d, want 1", len(s.Allocs))
}
}
}
func TestRenderUpdatePlan_CanaryDefaultOne(t *testing.T) {
// canary unset → default 1 canary alloc, then 3 in batches of 2.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
MaxParallel: 2,
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
// Step 0: canary (1 alloc). Step 1: promote. Steps 2..: remaining.
if plan.Steps[0].Action != "start" || len(plan.Steps[0].Allocs) != 1 {
t.Errorf("step 0 = %+v, want canary start with 1 alloc", plan.Steps[0])
}
if plan.Steps[0].Allocs[0] != "web-1" {
t.Errorf("canary alloc = %q, want web-1", plan.Steps[0].Allocs[0])
}
if !plan.Steps[0].WaitForHealthy {
t.Error("canary step should wait for healthy")
}
if plan.Steps[1].Action != "promote" {
t.Errorf("step 1 action = %q, want promote", plan.Steps[1].Action)
}
// Remaining: web-2, web-3, web-4 in batches of 2 → [web-2,web-3], [web-4].
if len(plan.Steps) != 4 {
t.Fatalf("got %d steps, want 4 (canary + promote + 2 batches)", len(plan.Steps))
}
if len(plan.Steps[2].Allocs) != 2 || plan.Steps[2].Allocs[0] != "web-2" {
t.Errorf("step 2 = %v, want [web-2 web-3]", plan.Steps[2].Allocs)
}
if len(plan.Steps[3].Allocs) != 1 || plan.Steps[3].Allocs[0] != "web-4" {
t.Errorf("step 3 = %v, want [web-4]", plan.Steps[3].Allocs)
}
}
func TestRenderUpdatePlan_CanaryPercent(t *testing.T) {
// count=10, canary=20% → 2 canary allocs, then 8 in batches of 3.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 10,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
MaxParallel: 3,
Canary: "20%",
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
if len(plan.Steps[0].Allocs) != 2 {
t.Errorf("canary step allocs = %d, want 2 (20%% of 10)", len(plan.Steps[0].Allocs))
}
// Remaining 8 in batches of 3 → ceil(8/3)=3 batches.
// Steps: canary, promote, batch(3), batch(3), batch(2) = 5 steps.
if len(plan.Steps) != 5 {
t.Fatalf("got %d steps, want 5", len(plan.Steps))
}
}
func TestRenderUpdatePlan_CanaryIntegerCount(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "2",
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
if len(plan.Steps[0].Allocs) != 2 {
t.Errorf("canary step allocs = %d, want 2", len(plan.Steps[0].Allocs))
}
if plan.Steps[0].Allocs[0] != "web-1" || plan.Steps[0].Allocs[1] != "web-2" {
t.Errorf("canary allocs = %v, want [web-1 web-2]", plan.Steps[0].Allocs)
}
}
func TestRenderUpdatePlan_CanaryFullCount(t *testing.T) {
// canary == count → no remaining allocs after canary.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 3,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "3",
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
// Steps: canary start (3), promote. No remaining batches.
if len(plan.Steps) != 2 {
t.Fatalf("got %d steps, want 2 (canary + promote, no remainder)", len(plan.Steps))
}
if plan.Steps[1].Action != "promote" {
t.Errorf("step 1 action = %q, want promote", plan.Steps[1].Action)
}
}
func TestRenderUpdatePlan_BlueGreen(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "blue-green",
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
// Step 0: start all 4 in parallel. Step 1: cutover.
if len(plan.Steps) != 2 {
t.Fatalf("got %d steps, want 2", len(plan.Steps))
}
if plan.Steps[0].Action != "start" {
t.Errorf("step 0 action = %q, want start", plan.Steps[0].Action)
}
if len(plan.Steps[0].Allocs) != 4 {
t.Errorf("step 0 allocs = %d, want 4 (all new in parallel)", len(plan.Steps[0].Allocs))
}
if !plan.Steps[0].WaitForHealthy {
t.Error("blue-green start step should wait for healthy")
}
if plan.Steps[1].Action != "cutover" {
t.Errorf("step 1 action = %q, want cutover", plan.Steps[1].Action)
}
if plan.Steps[1].WaitForHealthy {
t.Error("cutover step should NOT wait for healthy (already healthy)")
}
}
func TestRenderUpdatePlan_BlueGreenAllocs(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "api",
Count: 3,
Update: &jobspec.UpdateBlock{
Strategy: "blue-green",
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
want := []string{"api-1", "api-2", "api-3"}
if len(plan.Steps[0].Allocs) != 3 {
t.Errorf("allocs = %v, want %v", plan.Steps[0].Allocs, want)
}
for i, a := range want {
if plan.Steps[0].Allocs[i] != a {
t.Errorf("alloc[%d] = %q, want %q", i, plan.Steps[0].Allocs[i], a)
}
}
}
func TestRenderUpdatePlan_NilSpec(t *testing.T) {
_, err := RenderUpdatePlan(nil)
if err == nil {
t.Fatal("expected error for nil spec")
}
if !strings.Contains(err.Error(), "spec is nil") {
t.Errorf("error = %q, want 'spec is nil'", err.Error())
}
}
func TestRenderUpdatePlan_NilUpdate(t *testing.T) {
spec := &jobspec.WorkloadSpec{Kind: "Service", Name: "web", Count: 1}
_, err := RenderUpdatePlan(spec)
if err == nil {
t.Fatal("expected error for nil update block")
}
if !strings.Contains(err.Error(), "update block is nil") {
t.Errorf("error = %q, want 'update block is nil'", err.Error())
}
}
func TestRenderUpdatePlan_CountZero(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 0,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
},
}
_, err := RenderUpdatePlan(spec)
if err == nil {
t.Fatal("expected error for count 0")
}
if !strings.Contains(err.Error(), "count must be") {
t.Errorf("error = %q, want 'count must be'", err.Error())
}
}
func TestRenderUpdatePlan_UnknownStrategy(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "recreate",
},
}
_, err := RenderUpdatePlan(spec)
if err == nil {
t.Fatal("expected error for unknown strategy")
}
if !strings.Contains(err.Error(), "unknown strategy") {
t.Errorf("error = %q, want 'unknown strategy'", err.Error())
}
}
func TestRenderUpdatePlan_AllStrategies(t *testing.T) {
for _, strat := range []string{"rolling", "canary", "blue-green"} {
t.Run(strat, func(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 3,
Update: &jobspec.UpdateBlock{
Strategy: strat,
MaxParallel: 1,
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan(%s): %v", strat, err)
}
if len(plan.Steps) == 0 {
t.Errorf("strategy %s produced 0 steps", strat)
}
})
}
}
func TestRenderUpdatePlan_PromoteStepCarriesCanaryAllocs(t *testing.T) {
// The promote step lists the canary allocs so the executor knows
// which allocs are being promoted from canary to stable.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "2",
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
promote := plan.Steps[1]
if promote.Action != "promote" {
t.Fatalf("step 1 action = %q, want promote", promote.Action)
}
if len(promote.Allocs) != 2 {
t.Errorf("promote allocs = %d, want 2 (the canary allocs)", len(promote.Allocs))
}
}
func TestRenderUpdatePlan_MaxParallelClampedToCount(t *testing.T) {
// max_parallel > count is clamped to count (defensive; validator
// rejects this but the emitter is defensive against direct
// callers).
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
MaxParallel: 99,
},
}
plan, err := RenderUpdatePlan(spec)
if err != nil {
t.Fatalf("RenderUpdatePlan: %v", err)
}
// Clamped to 2 → single batch of 2.
if len(plan.Steps) != 1 {
t.Errorf("got %d steps, want 1 (clamped)", len(plan.Steps))
}
if len(plan.Steps[0].Allocs) != 2 {
t.Errorf("step 0 allocs = %d, want 2", len(plan.Steps[0].Allocs))
}
}
+191
View File
@@ -74,6 +74,27 @@ type WorkloadSpec struct {
// Timeout is an optional execution timeout (duration string) for
// Job. Populated by P04.
Timeout string
// Tasks is the task-group list for multi-process services (P06,
// PRD §9.1). When non-empty, the alloc runs one systemd unit per
// task (`orca-v1-alloc-<alloc-id>-<task-name>.service`) all
// grouped under a single `<alloc-id>.target`. When nil/empty,
// the alloc is a single-process alloc driven by the top-level
// Runtime block (backward compat). Tasks that omit their own
// runtime inherit the top-level Runtime as the per-group default.
Tasks []TaskGroupTask
}
// TaskGroupTask is a single task within a task group (P06, PRD §9.1).
// Each task has its own runtime (a wasm task + a process sidecar is
// allowed), its own command, and an optional env overlay. When
// Runtime is nil, the task inherits the top-level
// WorkloadSpec.Runtime (the per-group default).
type TaskGroupTask struct {
Name string
Runtime *RuntimeBlock
Env map[string]string
Command string
}
// RuntimeBlock is a minimal runtime abstraction surface populated by the
@@ -384,12 +405,18 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
secLifecycle
secAffinity
secConstraints
secTasks
secTaskEnv
secTaskRuntime
)
cur := secNone
var curPort *PortSpec
var curVol *VolumeSpec
var curAffinity *AffinityRule
var lifecycleCur string
var curTask *TaskGroupTask
var taskIndent int
var taskFieldIndent int
flushPort := func() {
if curPort != nil {
@@ -409,6 +436,29 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
curAffinity = nil
}
}
flushTask := func() {
if curTask != nil {
spec.Tasks = append(spec.Tasks, *curTask)
curTask = nil
}
}
// taskSubBlock returns the sub-section to switch to when the
// given `key: value` line opens a nested block under a task
// (`env:` → secTaskEnv, `runtime:` → secTaskRuntime). Returns
// secTasks for non-block keys (no switch).
taskSubBlock := func(kvLine string) section {
key, _, ok := splitKV(kvLine)
if !ok {
return secTasks
}
switch key {
case "env":
return secTaskEnv
case "runtime":
return secTaskRuntime
}
return secTasks
}
for lineNo, raw := range lines {
line := stripComment(raw)
@@ -423,6 +473,7 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
flushPort()
flushVol()
flushAffinity()
flushTask()
cur = secNone
key, val, ok := splitKV(trimmed)
@@ -500,6 +551,10 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
} else {
cur = secAffinity
}
case "tasks":
cur = secTasks
taskIndent = -1
taskFieldIndent = -1
default:
// Unknown top-level key are ignored (forward-compat).
cur = secNone
@@ -720,11 +775,119 @@ func parseFrontmatterBlock(block string) (*WorkloadSpec, error) {
spec.Constraints = append(spec.Constraints, unquote(item))
}
}
case secTasks:
// Tasks is a list of task objects. A `- ` at the list
// indent opens a new task; deeper-indented lines belong
// to the current task's fields (name, command) or
// nested sub-blocks (runtime, env).
if strings.HasPrefix(trimmed, "- ") || trimmed == "-" {
if taskIndent < 0 {
taskIndent = indent
taskFieldIndent = indent + 2
}
if indent == taskIndent {
flushTask()
t := TaskGroupTask{}
curTask = &t
rest := strings.TrimSpace(strings.TrimPrefix(trimmed, "-"))
if rest != "" {
if applyTaskKV(curTask, rest) {
cur = taskSubBlock(rest)
}
}
continue
}
}
if curTask != nil {
if applyTaskKV(curTask, trimmed) {
cur = taskSubBlock(trimmed)
}
}
case secTaskEnv:
if curTask == nil {
cur = secTasks
continue
}
// Pop back to the task field level when the indent
// returns to taskFieldIndent (the next sibling
// field or a new `- ` list item). The line is then
// reprocessed as a task field.
if taskFieldIndent > 0 && indent <= taskFieldIndent {
cur = secTasks
if indent == taskIndent && (strings.HasPrefix(trimmed, "- ") || trimmed == "-") {
flushTask()
t := TaskGroupTask{}
curTask = &t
rest := strings.TrimSpace(strings.TrimPrefix(trimmed, "-"))
if rest != "" {
if applyTaskKV(curTask, rest) {
cur = taskSubBlock(rest)
}
}
continue
}
if applyTaskKV(curTask, trimmed) {
cur = taskSubBlock(trimmed)
}
continue
}
if curTask.Env == nil {
curTask.Env = map[string]string{}
}
key, val, ok := splitKV(trimmed)
if !ok {
continue
}
if val == "" {
curTask.Env[key] = ""
} else if strings.HasPrefix(val, "{") && strings.HasSuffix(val, "}") {
curTask.Env[key] = val
} else {
curTask.Env[key] = unquote(val)
}
case secTaskRuntime:
if curTask == nil || curTask.Runtime == nil {
cur = secTasks
continue
}
// Pop back to the task field level (see secTaskEnv).
if taskFieldIndent > 0 && indent <= taskFieldIndent {
cur = secTasks
if indent == taskIndent && (strings.HasPrefix(trimmed, "- ") || trimmed == "-") {
flushTask()
t := TaskGroupTask{}
curTask = &t
rest := strings.TrimSpace(strings.TrimPrefix(trimmed, "-"))
if rest != "" {
if applyTaskKV(curTask, rest) {
cur = taskSubBlock(rest)
}
}
continue
}
if applyTaskKV(curTask, trimmed) {
cur = taskSubBlock(trimmed)
}
continue
}
key, val, ok := splitKV(trimmed)
if !ok {
continue
}
switch key {
case "one_of":
curTask.Runtime.OneOf = unquote(val)
case "image":
curTask.Runtime.Image = unquote(val)
case "command":
curTask.Runtime.Command = unquote(val)
}
}
}
flushPort()
flushVol()
flushAffinity()
flushTask()
return spec, nil
}
@@ -791,6 +954,34 @@ func applyAffinityKV(r *AffinityRule, s string) {
}
}
// applyTaskKV applies a `key: value` pair to the current TaskGroupTask.
// The returned bool reports whether the key opened a nested sub-block
// (`env` or `runtime`); when true the caller switches the parser
// section to the corresponding sub-block handler.
func applyTaskKV(t *TaskGroupTask, s string) (openedSubBlock bool) {
key, val, ok := splitKV(s)
if !ok {
return false
}
switch key {
case "name":
t.Name = unquote(val)
case "command":
t.Command = unquote(val)
case "env":
if t.Env == nil {
t.Env = map[string]string{}
}
return true
case "runtime":
if t.Runtime == nil {
t.Runtime = &RuntimeBlock{}
}
return true
}
return false
}
// appendLifecycleCmd appends a command to the named lifecycle hook list
// (pre_stop or post_start) on the given LifecycleBlock.
func appendLifecycleCmd(lb *LifecycleBlock, name, cmd string) {
+149
View File
@@ -701,3 +701,152 @@ func TestParseMarkdown_FullServiceSpec(t *testing.T) {
t.Errorf("Body = %q, want %q (R-015)", spec.Body, "# body\n")
}
}
func TestParseMarkdown_TasksBlock(t *testing.T) {
// P06: a task group with two tasks, each carrying its own runtime
// and command. The parser must populate spec.Tasks with two
// entries preserving name, runtime (one_of/image/command), and
// the task-level command.
input := "---\n" +
"kind: Service\n" +
"name: web\n" +
"tasks:\n" +
" - name: app\n" +
" runtime:\n" +
" one_of: process\n" +
" image: docker.io/nginx:latest\n" +
" command: /usr/bin/httpd -f\n" +
" command: /usr/bin/httpd -f\n" +
" - name: sidecar\n" +
" runtime:\n" +
" one_of: wasm\n" +
" command: /bin/wasm-runner sidecar.wasm\n" +
" command: /bin/wasm-runner sidecar.wasm\n" +
"---\nbody\n"
spec, err := ParseMarkdown([]byte(input))
if err != nil {
t.Fatalf("ParseMarkdown: %v", err)
}
if len(spec.Tasks) != 2 {
t.Fatalf("Tasks = %d, want 2", len(spec.Tasks))
}
app := spec.Tasks[0]
if app.Name != "app" {
t.Errorf("Tasks[0].Name = %q, want app", app.Name)
}
if app.Runtime == nil {
t.Fatal("Tasks[0].Runtime is nil")
}
if app.Runtime.OneOf != "process" {
t.Errorf("Tasks[0].Runtime.OneOf = %q, want process", app.Runtime.OneOf)
}
if app.Runtime.Image != "docker.io/nginx:latest" {
t.Errorf("Tasks[0].Runtime.Image = %q", app.Runtime.Image)
}
if app.Runtime.Command != "/usr/bin/httpd -f" {
t.Errorf("Tasks[0].Runtime.Command = %q", app.Runtime.Command)
}
if app.Command != "/usr/bin/httpd -f" {
t.Errorf("Tasks[0].Command = %q", app.Command)
}
side := spec.Tasks[1]
if side.Name != "sidecar" {
t.Errorf("Tasks[1].Name = %q, want sidecar", side.Name)
}
if side.Runtime == nil || side.Runtime.OneOf != "wasm" {
t.Errorf("Tasks[1].Runtime = %+v, want one_of=wasm", side.Runtime)
}
if side.Command != "/bin/wasm-runner sidecar.wasm" {
t.Errorf("Tasks[1].Command = %q", side.Command)
}
}
func TestParseMarkdown_TasksBlockWithEnv(t *testing.T) {
// P06: a task group task carrying an env overlay.
input := "---\n" +
"kind: Service\n" +
"name: web\n" +
"tasks:\n" +
" - name: app\n" +
" command: /usr/bin/httpd\n" +
" env:\n" +
" LOG_LEVEL: debug\n" +
" REGION: us\n" +
"---\nbody\n"
spec, err := ParseMarkdown([]byte(input))
if err != nil {
t.Fatalf("ParseMarkdown: %v", err)
}
if len(spec.Tasks) != 1 {
t.Fatalf("Tasks = %d, want 1", len(spec.Tasks))
}
task := spec.Tasks[0]
if task.Env == nil {
t.Fatal("Tasks[0].Env is nil")
}
if got := task.Env["LOG_LEVEL"]; got != "debug" {
t.Errorf("Env[LOG_LEVEL] = %q, want debug", got)
}
if got := task.Env["REGION"]; got != "us" {
t.Errorf("Env[REGION] = %q, want us", got)
}
}
func TestParseMarkdown_TasksBlockInheritsTopLevelRuntime(t *testing.T) {
// P06: when a task omits its own runtime, the top-level runtime
// is the per-group default. The parser must NOT create a task
// runtime when the task block lacks a `runtime:` sub-block; the
// emitter/validator resolve the default from spec.Runtime.
input := "---\n" +
"kind: Service\n" +
"name: web\n" +
"runtime:\n" +
" one_of: process\n" +
" command: /bin/default\n" +
"tasks:\n" +
" - name: app\n" +
" command: /bin/app\n" +
" - name: sidecar\n" +
" command: /bin/sidecar\n" +
"---\nbody\n"
spec, err := ParseMarkdown([]byte(input))
if err != nil {
t.Fatalf("ParseMarkdown: %v", err)
}
if spec.Runtime == nil || spec.Runtime.OneOf != "process" {
t.Fatalf("top-level runtime not parsed: %+v", spec.Runtime)
}
if len(spec.Tasks) != 2 {
t.Fatalf("Tasks = %d, want 2", len(spec.Tasks))
}
for i, task := range spec.Tasks {
if task.Runtime != nil {
t.Errorf("Tasks[%d].Runtime should be nil (inherit top-level), got %+v", i, task.Runtime)
}
}
if spec.Tasks[0].Name != "app" || spec.Tasks[1].Name != "sidecar" {
t.Errorf("task names = %q, %q", spec.Tasks[0].Name, spec.Tasks[1].Name)
}
}
func TestParseMarkdown_NoTasksBackwardCompat(t *testing.T) {
// Backward compat: a spec with no `tasks:` block parses as a
// single-process alloc; spec.Tasks must be empty/nil.
input := "---\n" +
"kind: Job\n" +
"name: backup\n" +
"runtime:\n" +
" one_of: process\n" +
" command: /bin/rsync\n" +
"---\nbody\n"
spec, err := ParseMarkdown([]byte(input))
if err != nil {
t.Fatalf("ParseMarkdown: %v", err)
}
if len(spec.Tasks) != 0 {
t.Fatalf("Tasks = %d, want 0 (backward compat)", len(spec.Tasks))
}
if spec.Runtime == nil || spec.Runtime.Command != "/bin/rsync" {
t.Errorf("Runtime = %+v, want command=/bin/rsync", spec.Runtime)
}
}
+536
View File
@@ -0,0 +1,536 @@
// Package scheduler — cel.go implements a minimal CEL-subset evaluator
// for the CLI-side scheduler constraint expressions (REQ-083, P05).
//
// The full CEL specification (google.golang.org/genproto/...
// googleapis/api/expr/v1alpha1) is intentionally NOT a dependency of
// this module (see go.mod): adding it for a single callsite would pull
// in a large transitive graph and contradict the "stdlib + minimal
// deps" guardrail. Instead this file implements a hand-rolled
// recursive-descent evaluator for the subset the PRD exercises:
//
// - attribute access on a `node.<name>` object (hostname, kind,
// cpus, memory, tags, runtimes)
// - string and integer literals (double-quoted)
// - comparison operators: == != >= <= > <
// - membership: <expr> in <expr>, <expr> not in <expr>
// - boolean composition: and, or, not (parenthesised)
//
// Anything outside this subset returns an error rather than a silent
// wrong answer; that is the documented limitation. The grammar is
// small enough to be unambiguous with a top-down precedence-climbing
// parser.
package scheduler
import (
"fmt"
"strconv"
"strings"
"unicode"
)
// EvaluateConstraint evaluates a single CEL-subset expression against
// the supplied NodeInfo. Returns (matched, err). An expression that
// references an unknown attribute, uses an unsupported operator, or
// fails to parse yields an error. Schedule treats a constraint
// evaluation error as a non-fit (the node is silently skipped) rather
// than a hard fail because operators routinely write exploratory
// constraints against attributes the local cluster does not expose.
func EvaluateConstraint(expr string, node NodeInfo) (bool, error) {
p := newParser(strings.TrimSpace(expr), node)
if p.len() == 0 {
return false, fmt.Errorf("cel: empty expression")
}
v, err := p.parseExpr()
if err != nil {
return false, err
}
if p.tok.kind != tokEOF {
return false, fmt.Errorf("cel: trailing input near %q", p.tok.text)
}
b, ok := v.(bool)
if !ok {
return false, fmt.Errorf("cel: expression did not evaluate to bool (got %T)", v)
}
return b, nil
}
// EvaluateAll returns true iff every constraint evaluates to true
// against the node (logical AND). An empty constraint list is vacuously
// true. The first evaluation error short-circuits and is returned.
func EvaluateAll(constraints []string, node NodeInfo) (bool, error) {
for _, c := range constraints {
ok, err := EvaluateConstraint(c, node)
if err != nil {
return false, fmt.Errorf("constraint %q: %w", c, err)
}
if !ok {
return false, nil
}
}
return true, nil
}
// ----------------------------------------------------------------------------
// Value model
// ----------------------------------------------------------------------------
// celValue is the union of values the evaluator produces. We use the
// Go interface{} representation so that comparisons can be polymorphic
// without a tagged-union ceremony; the supported concrete types are
// bool, int64, and string. Lists are []celValue of the above.
type celValue = interface{}
// ----------------------------------------------------------------------------
// Tokenizer
// ----------------------------------------------------------------------------
type tokKind int
const (
tokEOF tokKind = iota
tokIdent
tokInt
tokStr
tokOp // ==, !=, >=, <=, >, <, (, ), .
tokIn // "in"
tokAnd // "and"
tokOr // "or"
tokNot // "not"
)
type token struct {
kind tokKind
text string
}
type lexer struct {
src string
pos int
}
func (l *lexer) next() (token, error) {
for l.pos < len(l.src) && unicode.IsSpace(rune(l.src[l.pos])) {
l.pos++
}
if l.pos >= len(l.src) {
return token{kind: tokEOF}, nil
}
c := l.src[l.pos]
// string literal
if c == '"' {
start := l.pos
l.pos++
for l.pos < len(l.src) && l.src[l.pos] != '"' {
l.pos++
}
if l.pos >= len(l.src) {
return token{}, fmt.Errorf("cel: unterminated string at %d", start)
}
val := l.src[start+1 : l.pos]
l.pos++ // consume closing quote
return token{kind: tokStr, text: val}, nil
}
// integer literal
if unicode.IsDigit(rune(c)) {
start := l.pos
for l.pos < len(l.src) && unicode.IsDigit(rune(l.src[l.pos])) {
l.pos++
}
return token{kind: tokInt, text: l.src[start:l.pos]}, nil
}
// identifier / keyword
if isIdentStart(c) {
start := l.pos
for l.pos < len(l.src) && isIdentPart(l.src[l.pos]) {
l.pos++
}
word := l.src[start:l.pos]
switch word {
case "in":
return token{kind: tokIn, text: word}, nil
case "and":
return token{kind: tokAnd, text: word}, nil
case "or":
return token{kind: tokOr, text: word}, nil
case "not":
return token{kind: tokNot, text: word}, nil
default:
return token{kind: tokIdent, text: word}, nil
}
}
// operators
if strings.ContainsRune("()=!<>.", rune(c)) {
// multi-char operators
if l.pos+1 < len(l.src) {
two := l.src[l.pos : l.pos+2]
switch two {
case "==", "!=", ">=", "<=":
l.pos += 2
return token{kind: tokOp, text: two}, nil
}
}
l.pos++
return token{kind: tokOp, text: string(c)}, nil
}
return token{}, fmt.Errorf("cel: unexpected character %q at %d", c, l.pos)
}
func isIdentStart(c byte) bool {
return c == '_' || (c >= 'a' && c <= 'z') || (c >= 'A' && c <= 'Z')
}
func isIdentPart(c byte) bool {
return isIdentStart(c) || (c >= '0' && c <= '9')
}
// ----------------------------------------------------------------------------
// Parser (recursive descent, precedence climbing)
// ----------------------------------------------------------------------------
type parser struct {
src string
pos int
tok token
err error
node NodeInfo
}
func newParser(src string, node NodeInfo) *parser {
p := &parser{src: src, node: node}
p.advance()
return p
}
func (p *parser) len() int { return len(p.src) }
func (p *parser) advance() {
if p.err != nil {
return
}
l := lexer{src: p.src, pos: p.pos}
t, err := l.next()
if err != nil {
p.err = err
return
}
p.pos = l.pos
p.tok = t
}
// Grammar (lowest precedence first):
//
// expr := orExpr
// orExpr := andExpr ("or" andExpr)*
// andExpr := notExpr ("and" notExpr)*
// notExpr := "not" notExpr | cmpExpr
// cmpExpr := primary (op primary | "in" primary | "not" "in" primary)?
// primary := "(" expr ")"
// | int
// | str
// | "true" | "false"
// | nodeAttr ("." ident)? // node.<field>
// | ident // bare attribute (e.g. region)
// nodeAttr := "node"
func (p *parser) parseExpr() (celValue, error) {
if p.err != nil {
return nil, p.err
}
return p.parseOr()
}
func (p *parser) parseOr() (celValue, error) {
left, err := p.parseAnd()
if err != nil {
return nil, err
}
for p.tok.kind == tokOr {
p.advance()
right, err := p.parseAnd()
if err != nil {
return nil, err
}
lb, ok := left.(bool)
if !ok {
return nil, fmt.Errorf("cel: 'or' operand not bool: %T", left)
}
rb, ok := right.(bool)
if !ok {
return nil, fmt.Errorf("cel: 'or' operand not bool: %T", right)
}
left = lb || rb
}
return left, nil
}
func (p *parser) parseAnd() (celValue, error) {
left, err := p.parseNot()
if err != nil {
return nil, err
}
for p.tok.kind == tokAnd {
p.advance()
right, err := p.parseNot()
if err != nil {
return nil, err
}
lb, ok := left.(bool)
if !ok {
return nil, fmt.Errorf("cel: 'and' operand not bool: %T", left)
}
rb, ok := right.(bool)
if !ok {
return nil, fmt.Errorf("cel: 'and' operand not bool: %T", right)
}
left = lb && rb
}
return left, nil
}
func (p *parser) parseNot() (celValue, error) {
if p.tok.kind == tokNot {
// "not" at the start of a primary is logical negation. "not in"
// is handled in parseCmp where it follows a primary.
p.advance()
v, err := p.parseNot()
if err != nil {
return nil, err
}
b, ok := v.(bool)
if !ok {
return nil, fmt.Errorf("cel: 'not' operand not bool: %T", v)
}
return !b, nil
}
return p.parseCmp()
}
func (p *parser) parseCmp() (celValue, error) {
left, err := p.parsePrimary()
if err != nil {
return nil, err
}
// "not in"
if p.tok.kind == tokNot {
p.advance()
if p.tok.kind != tokIn {
return nil, fmt.Errorf("cel: expected 'in' after 'not', got %q", p.tok.text)
}
p.advance()
right, err := p.parsePrimary()
if err != nil {
return nil, err
}
member, err := inMember(left, right)
if err != nil {
return nil, err
}
return !member, nil
}
// "in"
if p.tok.kind == tokIn {
p.advance()
right, err := p.parsePrimary()
if err != nil {
return nil, err
}
return inMember(left, right)
}
// comparison operators
if p.tok.kind == tokOp {
op := p.tok.text
switch op {
case "==", "!=", ">=", "<=", ">", "<":
p.advance()
right, err := p.parsePrimary()
if err != nil {
return nil, err
}
return compare(op, left, right)
default:
return nil, fmt.Errorf("cel: unexpected operator %q", op)
}
}
return left, nil
}
// inMember reports whether left is a member of right. right must be a
// list ([]celValue) of comparable values; left may be a string or
// int64.
func inMember(left, right celValue) (bool, error) {
list, ok := right.([]celValue)
if !ok {
return false, fmt.Errorf("cel: 'in' rhs not a list: %T", right)
}
for _, e := range list {
if valuesEqual(left, e) {
return true, nil
}
}
return false, nil
}
func valuesEqual(a, b celValue) bool {
switch av := a.(type) {
case string:
bv, ok := b.(string)
return ok && av == bv
case int64:
bv, ok := b.(int64)
return ok && av == bv
case bool:
bv, ok := b.(bool)
return ok && av == bv
}
return false
}
// compare applies a binary comparison operator to two scalar values.
// Strings compare lexicographically; ints numerically; bools only via
// ==/!=.
func compare(op string, left, right celValue) (bool, error) {
switch op {
case "==":
return valuesEqual(left, right), nil
case "!=":
return !valuesEqual(left, right), nil
}
// ordered comparisons require ordered operands
ls, lok := left.(string)
rs, rok := right.(string)
if lok && rok {
switch op {
case "<":
return ls < rs, nil
case "<=":
return ls <= rs, nil
case ">":
return ls > rs, nil
case ">=":
return ls >= rs, nil
}
}
li, lok := left.(int64)
ri, rok := right.(int64)
if lok && rok {
switch op {
case "<":
return li < ri, nil
case "<=":
return li <= ri, nil
case ">":
return li > ri, nil
case ">=":
return li >= ri, nil
}
}
return false, fmt.Errorf("cel: cannot apply %q to %T and %T", op, left, right)
}
// parsePrimary parses the smallest standalone unit: parenthesised
// expressions, literals, and attribute references.
func (p *parser) parsePrimary() (celValue, error) {
switch p.tok.kind {
case tokOp:
if p.tok.text == "(" {
p.advance()
v, err := p.parseExpr()
if err != nil {
return nil, err
}
if p.tok.kind != tokOp || p.tok.text != ")" {
return nil, fmt.Errorf("cel: expected ')' got %q", p.tok.text)
}
p.advance()
return v, nil
}
return nil, fmt.Errorf("cel: unexpected operator %q", p.tok.text)
case tokInt:
n, err := strconv.ParseInt(p.tok.text, 10, 64)
if err != nil {
return nil, fmt.Errorf("cel: bad int %q: %w", p.tok.text, err)
}
p.advance()
return n, nil
case tokStr:
v := p.tok.text
p.advance()
return v, nil
case tokIdent:
return p.parseAttrRef()
}
return nil, fmt.Errorf("cel: unexpected token %q", p.tok.text)
}
// parseAttrRef resolves a bare or `node.<field>` attribute reference
// against the node being evaluated. Bare identifiers (e.g. `region`)
// resolve against the same attribute map as `node.region`; the PRD
// examples use both forms interchangeably (see
// TestParseMarkdown_ConstraintsInlineArray).
func (p *parser) parseAttrRef() (celValue, error) {
name := p.tok.text
p.advance()
// dotted access: node.<field>
if p.tok.kind == tokOp && p.tok.text == "." {
if name != "node" {
return nil, fmt.Errorf("cel: dotted access on non-node: %q", name)
}
p.advance()
if p.tok.kind != tokIdent {
return nil, fmt.Errorf("cel: expected attribute name after '.', got %q", p.tok.text)
}
field := p.tok.text
p.advance()
return p.nodeAttr(name + "." + field)
}
// bare identifier
switch name {
case "true":
return true, nil
case "false":
return false, nil
default:
return p.nodeAttr(name)
}
}
// nodeAttr resolves an attribute name to its value on the parser's
// active node. Mapping (per PRD T2):
//
// node.hostname -> Hostname (string)
// node.kind -> Kind (string)
// node.cpus -> CPU (int64)
// node.memory -> Memory (int64)
// node.tags -> Tags ([]string -> []celValue)
// node.runtimes -> Runtimes ([]string -> []celValue)
//
// Bare names (without the `node.` prefix) resolve through the same
// map, so `region == "us"` and `node.region == "us"` are equivalent
// when the attribute exists.
func (p *parser) nodeAttr(name string) (celValue, error) {
switch name {
case "node.hostname", "hostname":
return p.node.Hostname, nil
case "node.kind", "kind":
return p.node.Kind, nil
case "node.cpus", "cpus":
return p.node.CPU, nil
case "node.memory", "memory":
return p.node.Memory, nil
case "node.tags", "tags":
return toStringValues(p.node.Tags), nil
case "node.runtimes", "runtimes":
return toStringValues(p.node.Runtimes), nil
}
return nil, fmt.Errorf("cel: unknown attribute %q", name)
}
// toStringValues converts a []string to []celValue so the membership
// operators can compare element-wise.
func toStringValues(in []string) []celValue {
out := make([]celValue, len(in))
for i, s := range in {
out[i] = s
}
return out
}
+210
View File
@@ -0,0 +1,210 @@
package scheduler
import "testing"
func TestEvaluateConstraint_Equality(t *testing.T) {
node := NodeInfo{Hostname: "h-1", Kind: "linux", CPU: 4, Memory: 4096, Tags: []string{"web"}, Runtimes: []string{"process"}}
cases := []struct {
name string
expr string
want bool
}{
{"hostname eq", `node.hostname == "h-1"`, true},
{"hostname ne", `node.hostname == "h-2"`, false},
{"kind eq", `node.kind == "linux"`, true},
{"kind ne", `node.kind == "proxmox"`, false},
{"cpus eq", `node.cpus == 4`, true},
{"memory eq", `node.memory == 4096`, true},
}
for _, c := range cases {
got, err := EvaluateConstraint(c.expr, node)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got != c.want {
t.Errorf("%s: got %v, want %v", c.name, got, c.want)
}
}
}
func TestEvaluateConstraint_Comparison(t *testing.T) {
node := NodeInfo{Hostname: "h", Kind: "linux", CPU: 4, Memory: 4096}
cases := []struct {
expr string
want bool
}{
{"node.cpus >= 2", true},
{"node.cpus >= 4", true},
{"node.cpus > 4", false},
{"node.cpus > 2", true},
{"node.cpus <= 4", true},
{"node.cpus < 2", false},
{"node.cpus != 8", true},
{"node.cpus == 8", false},
{"node.memory >= 2048", true},
{"node.memory < 1024", false},
}
for _, c := range cases {
got, err := EvaluateConstraint(c.expr, node)
if err != nil {
t.Errorf("%q: %v", c.expr, err)
continue
}
if got != c.want {
t.Errorf("%q: got %v, want %v", c.expr, got, c.want)
}
}
}
func TestEvaluateConstraint_Membership(t *testing.T) {
node := NodeInfo{Tags: []string{"web", "log-shipper"}, Runtimes: []string{"process", "wasmtime"}}
cases := []struct {
expr string
want bool
}{
{`"web" in node.tags`, true},
{`"missing" in node.tags`, false},
{`"process" in node.runtimes`, true},
{`"podman" in node.runtimes`, false},
{`"log-shipper" not in node.tags`, false},
{`"missing" not in node.tags`, true},
}
for _, c := range cases {
got, err := EvaluateConstraint(c.expr, node)
if err != nil {
t.Errorf("%q: %v", c.expr, err)
continue
}
if got != c.want {
t.Errorf("%q: got %v, want %v", c.expr, got, c.want)
}
}
}
func TestEvaluateConstraint_BooleanComposition(t *testing.T) {
node := NodeInfo{Kind: "linux", CPU: 4, Tags: []string{"web"}}
cases := []struct {
expr string
want bool
}{
{`node.kind == "linux" and node.cpus >= 2`, true},
{`node.kind == "proxmox" and node.cpus >= 2`, false},
{`node.kind == "linux" or node.kind == "proxmox"`, true},
{`node.kind == "proxmox" or node.kind == "linux"`, true},
{`not node.kind == "proxmox"`, true},
{`not node.kind == "linux"`, false},
{`(node.kind == "linux") and (node.cpus >= 2)`, true},
{`node.cpus >= 2 and not "blocked" in node.tags`, true},
{`node.kind == "linux" and node.cpus >= 2 and "web" in node.tags`, true},
{`node.kind == "linux" or node.kind == "proxmox" or node.cpus > 100`, true},
}
for _, c := range cases {
got, err := EvaluateConstraint(c.expr, node)
if err != nil {
t.Errorf("%q: %v", c.expr, err)
continue
}
if got != c.want {
t.Errorf("%q: got %v, want %v", c.expr, got, c.want)
}
}
}
func TestEvaluateConstraint_BareIdentifiers(t *testing.T) {
// Bare identifiers resolve through the same attribute map as
// node.<field> (per PRD: constraints may use either form).
node := NodeInfo{Kind: "linux", CPU: 4}
got, err := EvaluateConstraint(`kind == "linux"`, node)
if err != nil {
t.Fatalf("bare kind: %v", err)
}
if !got {
t.Error("bare kind == linux: got false, want true")
}
}
func TestEvaluateConstraint_TrueFalseLiterals(t *testing.T) {
node := NodeInfo{}
cases := []struct {
expr string
want bool
}{
{"true", true},
{"false", false},
{"not false", true},
{"not true", false},
{"true and true", true},
{"true and false", false},
{"false or true", true},
}
for _, c := range cases {
got, err := EvaluateConstraint(c.expr, node)
if err != nil {
t.Errorf("%q: %v", c.expr, err)
continue
}
if got != c.want {
t.Errorf("%q: got %v, want %v", c.expr, got, c.want)
}
}
}
func TestEvaluateConstraint_Errors(t *testing.T) {
node := NodeInfo{Kind: "linux"}
cases := []struct {
name string
expr string
}{
{"empty", ""},
{"unterminated string", `node.kind == "linux`},
{"unknown attribute", `node.bogus == 1`},
{"unknown bare attr", `bogus == 1`},
{"dotted on non-node", `host.kind == "linux"`},
{"bad operator", `node.cpus + 2`},
{"trailing input", `node.kind == "linux" garbage`},
{"unbalanced paren", `(node.kind == "linux"`},
{"missing rhs", `node.cpus >=`},
{"not without in", `"x" not node.tags`},
{"ordered compare on bool", `true < false`},
{"ordered compare on mismatched types", `node.kind > 2`},
{"in on non-list", `"x" in node.kind`},
}
for _, c := range cases {
_, err := EvaluateConstraint(c.expr, node)
if err == nil {
t.Errorf("%s: expected error for %q, got nil", c.name, c.expr)
}
}
}
func TestEvaluateAll(t *testing.T) {
node := NodeInfo{Kind: "linux", CPU: 4, Tags: []string{"web"}}
cases := []struct {
name string
constraints []string
want bool
}{
{"empty", nil, true},
{"all pass", []string{`node.kind == "linux"`, "node.cpus >= 2"}, true},
{"one fails", []string{`node.kind == "linux"`, "node.cpus >= 8"}, false},
{"all fail", []string{`node.kind == "proxmox"`, "node.cpus >= 8"}, false},
}
for _, c := range cases {
got, err := EvaluateAll(c.constraints, node)
if err != nil {
t.Errorf("%s: %v", c.name, err)
continue
}
if got != c.want {
t.Errorf("%s: got %v, want %v", c.name, got, c.want)
}
}
}
func TestEvaluateAll_PropagatesError(t *testing.T) {
node := NodeInfo{}
if _, err := EvaluateAll([]string{"bogus == 1"}, node); err == nil {
t.Error("EvaluateAll: expected error for malformed constraint")
}
}
+472
View File
@@ -0,0 +1,472 @@
// Package scheduler implements the v0.9 CLI-side scheduler (REQ-083,
// P05). Unlike the v0.8 daemon-side best-fit scheduler
// (internal/engine/scheduler.go), this scheduler runs entirely in the
// `orca` CLI process (R-001) and is pure: it takes a list of candidate
// nodes plus a workload request and returns placement decisions
// without performing any I/O.
//
// The scheduler is runtime-aware: a workload that declares
// `runtime.one_of: wasm` is only placed on nodes that expose
// `wasmtime` in their Runtimes list; a `pve-vm` workload is only
// placed on `proxmox` nodes. It is also constraint- and
// affinity-aware via the CEL-subset evaluator in cel.go.
//
// Workload kinds are handled differently per the PRD:
//
// - Job: one-shot, returns exactly one placement (best-fit
// bin-packing).
// - Service: count replicas spread across distinct nodes
// (anti-affinity by default); if fewer distinct nodes than count,
// colocation is permitted but distinct nodes are preferred.
// - DaemonSet: one placement per node that fits the constraints.
package scheduler
import (
"fmt"
"sort"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
// NodeInfo is the scheduler's projection of a peer node: total and
// free capacity, the runtimes the node advertises, its tags, and its
// kind (linux/proxmox). The CLI populates this from the
// cluster/peers/ inventory plus the per-node capacity reports
// collected over SSH; the scheduler itself never reads either.
type NodeInfo struct {
Hostname string
Runtimes []string
Tags []string
CPU int64
Memory int64
FreeCPU int64
FreeMem int64
Kind string
}
// WorkloadRequest bundles a parsed WorkloadSpec with the namespace
// the workload is being scheduled into. The namespace is carried
// through to placement so the resulting AllocID can be namespaced,
// but the scheduler itself does not inspect it for fitting decisions.
type WorkloadRequest struct {
Spec *jobspec.WorkloadSpec
Namespace string
}
// Placement is a single scheduling decision: which node, which
// allocation id, and the bin-packing score that won the node the
// placement. AllocID is `ns/spec.Name-<idx>` so a multi-replica
// Service produces distinct ids per replica.
type Placement struct {
Node string
AllocID string
Score int64
}
// Schedule is the main entry point. For Job (kind=Job) it returns one
// placement on the best-fit node. For Service it returns `Count`
// placements spread across distinct nodes where possible (anti-
// affinity), permitting colocation when Count > nodes. For DaemonSet
// it returns one placement per node that fits. Any kind-agnostic
// validation error (no spec, unknown kind, no fitting node) is
// returned as an error rather than an empty slice so callers can
// distinguish "nothing fits" from "scheduled zero replicas".
func Schedule(nodes []NodeInfo, req WorkloadRequest) ([]Placement, error) {
if req.Spec == nil {
return nil, fmt.Errorf("scheduler: nil WorkloadSpec")
}
if len(nodes) == 0 {
return nil, fmt.Errorf("scheduler: no candidate nodes")
}
switch req.Spec.Kind {
case "Job":
return scheduleJob(nodes, req)
case "Service":
return scheduleService(nodes, req)
case "DaemonSet":
return scheduleDaemonSet(nodes, req)
default:
return nil, fmt.Errorf("scheduler: unknown kind %q", req.Spec.Kind)
}
}
// Score evaluates a single node against a workload. fits is true iff
// the node (a) advertises a runtime compatible with the workload's
// `runtime.one_of`, (b) satisfies every CEL constraint in
// `spec.Constraints`, and (c) has enough free CPU+memory for the
// workload's requested resources. When fits is true, score is the
// bin-packing score (more free capacity = higher score, so the node
// most likely to absorb the workload without starving its
// neighbours wins). When fits is false, score is 0.
func Score(node NodeInfo, req WorkloadRequest) (score int64, fits bool) {
// (a) runtime compatibility. A workload with no Runtime block or
// an empty OneOf is treated as runtime-agnostic (always fits on
// the runtime axis); this matches the v0.8 behaviour where a
// missing runtime meant "process".
runtimeOK := true
if req.Spec != nil && req.Spec.Runtime != nil && req.Spec.Runtime.OneOf != "" {
runtimeOK = hasRuntime(node, req.Spec.Runtime.OneOf)
}
if !runtimeOK {
return 0, false
}
// (b) constraints. Evaluation errors are treated as non-fit so a
// malformed constraint does not crash Schedule; the caller still
// sees the node filtered out.
if req.Spec != nil {
ok, err := EvaluateAll(req.Spec.Constraints, node)
if err != nil || !ok {
return 0, false
}
}
// (c) capacity. A workload with no Resources block is treated as
// zero-sized for fitting purposes (it always fits the capacity
// axis); real workloads declare cpu/memory.
needCPU, needMem := workloadResources(req)
if node.FreeCPU < needCPU || node.FreeMem < needMem {
return 0, false
}
// bin-packing score: most free capacity wins. CPU is weighted
// 1000x memory so a 1-core difference outweighs a 1-MiB
// difference, mirroring the v0.8 Score weighting that biased
// toward CPU (the more common binding constraint).
score = (node.FreeCPU-needCPU)*1000 + (node.FreeMem - needMem)
if score < 0 {
score = 0
}
return score, true
}
// hasRuntime reports whether node advertises the requested runtime.
// The match is case-insensitive and tolerant of aliases: `wasm` and
// `wasmtime` are treated as the same runtime, and `pve-vm`/`pve-ct`
// only match nodes whose Kind is "proxmox".
func hasRuntime(node NodeInfo, oneOf string) bool {
want := normalizeRuntime(oneOf)
// pve-* runtimes require a proxmox-kind node regardless of the
// node's Runtimes list (a proxmox node doesn't list "pve-vm" in
// Runtimes; it IS the runtime).
switch want {
case "pve-vm", "pve-ct", "proxmox":
return normalizeKind(node.Kind) == "proxmox"
}
for _, r := range node.Runtimes {
if normalizeRuntime(r) == want {
return true
}
// alias: wasmtime nodes advertise "wasmtime"; workloads ask
// for "wasm".
if want == "wasm" && normalizeRuntime(r) == "wasmtime" {
return true
}
}
return false
}
// normalizeRuntime lowercases and trims a runtime name for matching.
func normalizeRuntime(s string) string {
s = toLowerASCII(s)
switch s {
case "wasmtime":
return "wasm"
}
return s
}
// normalizeKind lowercases and trims a node Kind for matching.
func normalizeKind(s string) string { return toLowerASCII(s) }
// toLowerASCII lowercases ASCII letters without bringing in strings
// (avoid an alloc-heavy stdlib call in the hot path).
func toLowerASCII(s string) string {
b := []byte(s)
for i, c := range b {
if c >= 'A' && c <= 'Z' {
b[i] = c + 32
}
}
return string(b)
}
// workloadResources returns the (cpu, memory) the workload requests,
// read from the WorkloadSpec's Resources block if present. The
// v0.9-P05 WorkloadSpec does not yet carry a Resources field (it
// lands in P0c, REQ-074); until then this returns (0, 0) so the
// capacity check is a no-op and runtime/constraints do the real
// filtering. The signature is here so the scheduler logic does not
// need to change when Resources lands.
func workloadResources(req WorkloadRequest) (int64, int64) {
_ = req
return 0, 0
}
// ----------------------------------------------------------------------------
// Kind-specific scheduling
// ----------------------------------------------------------------------------
// scheduleJob places a single Job on the best-fit node.
func scheduleJob(nodes []NodeInfo, req WorkloadRequest) ([]Placement, error) {
type cand struct {
node NodeInfo
score int64
}
var cands []cand
for _, n := range nodes {
s, ok := Score(n, req)
if !ok {
continue
}
cands = append(cands, cand{node: n, score: s})
}
// CEL-based affinity rules apply to single-shot Jobs too: a Job
// with `affinity: [{target: "\"ssd\" in node.tags", weight: 100}]`
// should land on the tagged node even without prior placements.
// Name-based affinity (no prior placements to check) contributes
// zero for a standalone Job, so it is harmless to call here.
for i := range cands {
cands[i].score += affinityScore(cands[i].node, req, nil)
}
if len(cands) == 0 {
return nil, fmt.Errorf("scheduler: no node fits workload %q", req.Spec.Name)
}
sort.SliceStable(cands, func(i, j int) bool {
if cands[i].score != cands[j].score {
return cands[i].score > cands[j].score
}
return cands[i].node.Hostname < cands[j].node.Hostname
})
w := cands[0]
return []Placement{{
Node: w.node.Hostname,
AllocID: allocID(req, 0),
Score: w.score,
}}, nil
}
// scheduleService places `Count` replicas with implicit anti-affinity:
// prefer distinct nodes, but permit colocation when Count exceeds the
// number of fitting nodes. Each replica gets a distinct AllocID.
func scheduleService(nodes []NodeInfo, req WorkloadRequest) ([]Placement, error) {
count := req.Spec.Count
if count <= 0 {
count = 1
}
// Pre-filter fitting nodes once; the loop below re-scores them
// after each placement so the capacity accounting reflects the
// replicas already placed.
fitting := filterFitting(nodes, req)
if len(fitting) == 0 {
return nil, fmt.Errorf("scheduler: no node fits service %q", req.Spec.Name)
}
var placements []Placement
placed := map[string]int{} // hostname -> count placed there
// First pass: spread across distinct nodes.
for i := 0; i < count; i++ {
best, score, ok := pickServiceNode(fitting, req, placements, placed)
if !ok {
break
}
placements = append(placements, Placement{
Node: best.Hostname,
AllocID: allocID(req, i),
Score: score,
})
placed[best.Hostname]++
// Reflect the consumed capacity in the candidate snapshot so
// subsequent picks see updated free capacity.
needCPU, needMem := workloadResources(req)
for j := range fitting {
if fitting[j].Hostname == best.Hostname {
fitting[j].FreeCPU -= needCPU
fitting[j].FreeMem -= needMem
}
}
}
if len(placements) < count {
return nil, fmt.Errorf("scheduler: only placed %d/%d replicas for service %q",
len(placements), count, req.Spec.Name)
}
return placements, nil
}
// pickServiceNode selects the best node for the next replica. The
// selection prefers nodes with zero prior placements of this service
// (anti-affinity) and applies affinity scoring on top of the
// bin-packing score.
func pickServiceNode(fitting []NodeInfo, req WorkloadRequest, placements []Placement, placed map[string]int) (NodeInfo, int64, bool) {
type scored struct {
node NodeInfo
score int64
}
var cands []scored
for _, n := range fitting {
s, ok := Score(n, req)
if !ok {
continue
}
// Implicit anti-affinity: a node with N prior replicas of this
// service incurs a penalty of N * (1 << 62) so distinct nodes
// are preferred, but colocation is permitted (with a
// per-replica penalty) when no distinct node remains. This
// produces a balanced spread (e.g. 5 replicas on 3 nodes →
// 2/2/1) rather than stacking everything on the first node.
if placed[n.Hostname] > 0 {
s -= int64(placed[n.Hostname]) * (1 << 62)
}
// Affinity rules from the spec add/subtract their weight.
s += affinityScore(n, req, placements)
cands = append(cands, scored{node: n, score: s})
}
if len(cands) == 0 {
return NodeInfo{}, 0, false
}
sort.SliceStable(cands, func(i, j int) bool {
if cands[i].score != cands[j].score {
return cands[i].score > cands[j].score
}
return cands[i].node.Hostname < cands[j].node.Hostname
})
w := cands[0]
return w.node, w.score, true
}
// scheduleDaemonSet places one replica per node that fits the
// constraints. The PRD's DaemonSet placement mode (every-node /
// matching / mandatory) lives on the WorkloadSpec.Schedule block; the
// scheduler honours it indirectly by filtering on Constraints: a
// `matching` DaemonSet carries constraints that select the matching
// nodes, an `every-node` DaemonSet carries none, and a `mandatory`
// one is enforced elsewhere (the scheduler still just returns
// placements for every fitting node).
func scheduleDaemonSet(nodes []NodeInfo, req WorkloadRequest) ([]Placement, error) {
var placements []Placement
for _, n := range nodes {
s, ok := Score(n, req)
if !ok {
continue
}
placements = append(placements, Placement{
Node: n.Hostname,
AllocID: allocID(req, len(placements)),
Score: s,
})
}
if len(placements) == 0 {
return nil, fmt.Errorf("scheduler: no node fits daemonset %q", req.Spec.Name)
}
return placements, nil
}
// ----------------------------------------------------------------------------
// Affinity scoring
// ----------------------------------------------------------------------------
// affinityScore returns the weighted affinity contribution for a
// node given the placements already made. For each AffinityRule the
// Target is a CEL expression; if it evaluates true against the node,
// the rule's Weight is added (positive = co-locate, negative =
// anti-affinity). An affinity target that fails to evaluate is
// ignored rather than failing the schedule: operators use affinity as
// a hint, not a hard gate.
//
// The PRD also mentions affinity rules like `{target: "redis", weight:
// 50}` where Target is a workload *name* rather than a CEL expression.
// We support both: if Target parses as a CEL expression it is
// evaluated against the node; otherwise it is treated as a workload
// name and we check whether any already-placed alloc for that name
// exists on the node. The placement-already-here check is done by the
// caller via placements; this function checks the node's own
// attributes only.
func affinityScore(node NodeInfo, req WorkloadRequest, placements []Placement) int64 {
if req.Spec == nil {
return 0
}
var total int64
for _, rule := range req.Spec.Affinity {
// Try CEL evaluation first; if the target is a bare workload
// name (no operator) the CEL parser will fail and we fall
// back to name-based placement counting.
ok, err := EvaluateConstraint(rule.Target, node)
if err == nil {
if ok {
total += int64(rule.Weight)
}
continue
}
// Fallback: target is a workload name; count existing
// placements for that workload on this node and apply the
// weight once per co-located replica.
for _, p := range placements {
if p.Node == node.Hostname && isAllocFor(p.AllocID, rule.Target) {
total += int64(rule.Weight)
}
}
}
return total
}
// isAllocFor reports whether an AllocID encodes a placement for the
// named workload. AllocIDs are `ns/name-idx`, so we look for the
// workload name as the segment after the first slash and before the
// trailing `-idx`.
func isAllocFor(allocID, workloadName string) bool {
// strip namespace prefix
rest := allocID
if i := indexByte(rest, '/'); i >= 0 {
rest = rest[i+1:]
}
// strip trailing -idx
if i := lastIndexByte(rest, '-'); i >= 0 {
rest = rest[:i]
}
return rest == workloadName
}
// indexByte returns the index of the first occurrence of b in s, or
// -1. Avoids importing strings just for one helper.
func indexByte(s string, b byte) int {
for i := 0; i < len(s); i++ {
if s[i] == b {
return i
}
}
return -1
}
// lastIndexByte returns the index of the last occurrence of b in s, or
// -1.
func lastIndexByte(s string, b byte) int {
for i := len(s) - 1; i >= 0; i-- {
if s[i] == b {
return i
}
}
return -1
}
// ----------------------------------------------------------------------------
// Helpers
// ----------------------------------------------------------------------------
// filterFitting returns a copy of the nodes that pass Score for the
// request, preserving order. Capacity is not yet decremented; the
// caller adjusts FreeCPU/FreeMem as it places replicas.
func filterFitting(nodes []NodeInfo, req WorkloadRequest) []NodeInfo {
var out []NodeInfo
for _, n := range nodes {
if _, ok := Score(n, req); ok {
out = append(out, n)
}
}
return out
}
// allocID renders a stable, namespaced allocation id for a placement.
// Format: `ns/spec.Name-<idx>`.
func allocID(req WorkloadRequest, idx int) string {
ns := req.Namespace
if ns == "" {
ns = "default"
}
return fmt.Sprintf("%s/%s-%d", ns, req.Spec.Name, idx)
}
+451
View File
@@ -0,0 +1,451 @@
package scheduler
import (
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
// threeLinuxNodes returns a small cluster of three Linux nodes with
// distinct free capacities so best-fit ordering is unambiguous.
func threeLinuxNodes() []NodeInfo {
return []NodeInfo{
{Hostname: "node-a", Runtimes: []string{"process"}, Tags: nil, CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096, Kind: "linux"},
{Hostname: "node-b", Runtimes: []string{"process"}, Tags: nil, CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192, Kind: "linux"},
{Hostname: "node-c", Runtimes: []string{"process"}, Tags: nil, CPU: 2, Memory: 2048, FreeCPU: 2, FreeMem: 2048, Kind: "linux"},
}
}
func jobSpec(name, oneOf string, constraints []string) *jobspec.WorkloadSpec {
return &jobspec.WorkloadSpec{
Kind: "Job",
Name: name,
Count: 1,
Runtime: &jobspec.RuntimeBlock{OneOf: oneOf},
Constraints: constraints,
}
}
func serviceSpec(name, oneOf string, count int, constraints []string) *jobspec.WorkloadSpec {
return &jobspec.WorkloadSpec{
Kind: "Service",
Name: name,
Count: count,
Runtime: &jobspec.RuntimeBlock{OneOf: oneOf},
Constraints: constraints,
}
}
func daemonSetSpec(name, oneOf string, constraints []string) *jobspec.WorkloadSpec {
return &jobspec.WorkloadSpec{
Kind: "DaemonSet",
Name: name,
Count: 1,
Runtime: &jobspec.RuntimeBlock{OneOf: oneOf},
Constraints: constraints,
}
}
// ---------------------------------------------------------------------------
// Job
// ---------------------------------------------------------------------------
func TestScheduleJob_BestFit(t *testing.T) {
nodes := threeLinuxNodes()
req := WorkloadRequest{Spec: jobSpec("batch", "process", nil), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if len(got) != 1 {
t.Fatalf("placements = %d, want 1", len(got))
}
if got[0].Node != "node-b" {
t.Errorf("Node = %q, want node-b (most free capacity)", got[0].Node)
}
if !strings.HasPrefix(got[0].AllocID, "ns/batch-") {
t.Errorf("AllocID = %q, want ns/batch-*", got[0].AllocID)
}
if got[0].Score <= 0 {
t.Errorf("Score = %d, want > 0", got[0].Score)
}
}
func TestScheduleJob_NoFittingNode(t *testing.T) {
nodes := threeLinuxNodes()
// wasm runtime not advertised by any node.
req := WorkloadRequest{Spec: jobSpec("wasmjob", "wasm", nil), Namespace: "ns"}
if _, err := Schedule(nodes, req); err == nil {
t.Fatal("Schedule: expected error for no-fitting node, got nil")
}
}
// ---------------------------------------------------------------------------
// Service
// ---------------------------------------------------------------------------
func TestScheduleService_SpreadAcrossNodes(t *testing.T) {
nodes := threeLinuxNodes()
req := WorkloadRequest{Spec: serviceSpec("web", "process", 3, nil), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if len(got) != 3 {
t.Fatalf("placements = %d, want 3", len(got))
}
seen := map[string]int{}
for _, p := range got {
seen[p.Node]++
}
if len(seen) != 3 {
t.Errorf("anti-affinity spread: distinct nodes = %d, want 3; %v", len(seen), seen)
}
}
func TestScheduleService_ColocationWhenFewerNodes(t *testing.T) {
nodes := threeLinuxNodes()
req := WorkloadRequest{Spec: serviceSpec("web", "process", 5, nil), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if len(got) != 5 {
t.Fatalf("placements = %d, want 5", len(got))
}
seen := map[string]int{}
for _, p := range got {
seen[p.Node]++
}
if len(seen) != 3 {
t.Errorf("colocation: distinct nodes = %d, want 3 (all used)", len(seen))
}
// No node should host more than 2 (3 nodes, 5 replicas: 2+2+1).
for n, c := range seen {
if c > 2 {
t.Errorf("node %s has %d replicas, want <= 2", n, c)
}
}
}
func TestScheduleService_NoFittingNode(t *testing.T) {
nodes := threeLinuxNodes()
req := WorkloadRequest{Spec: serviceSpec("wasm-svc", "wasm", 3, nil), Namespace: "ns"}
if _, err := Schedule(nodes, req); err == nil {
t.Fatal("Schedule: expected error for service with no fitting node")
}
}
// ---------------------------------------------------------------------------
// DaemonSet
// ---------------------------------------------------------------------------
func TestScheduleDaemonSet_AllMatching(t *testing.T) {
nodes := threeLinuxNodes()
req := WorkloadRequest{Spec: daemonSetSpec("logrotate", "process", nil), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if len(got) != 3 {
t.Errorf("placements = %d, want 3 (one per node)", len(got))
}
seen := map[string]bool{}
for _, p := range got {
seen[p.Node] = true
}
if len(seen) != 3 {
t.Errorf("DaemonSet distinct nodes = %d, want 3", len(seen))
}
}
func TestScheduleDaemonSet_SomeExcludedByConstraint(t *testing.T) {
nodes := threeLinuxNodes()
// Only nodes with cpus >= 4 qualify: node-a (4) and node-b (8).
req := WorkloadRequest{Spec: daemonSetSpec("heavy", "process", []string{"node.cpus >= 4"}), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if len(got) != 2 {
t.Errorf("placements = %d, want 2 (cpus>=4)", len(got))
}
}
// ---------------------------------------------------------------------------
// Runtime compatibility
// ---------------------------------------------------------------------------
func TestSchedule_RuntimeCompatibilityWasm(t *testing.T) {
nodes := []NodeInfo{
{Hostname: "no-wasm", Runtimes: []string{"process"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
{Hostname: "has-wasm", Runtimes: []string{"process", "wasmtime"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
}
// Even though no-wasm has more free capacity, the wasm workload
// must land on has-wasm.
req := WorkloadRequest{Spec: jobSpec("wasmjob", "wasm", nil), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if got[0].Node != "has-wasm" {
t.Errorf("Node = %q, want has-wasm (runtime compatibility)", got[0].Node)
}
}
func TestSchedule_RuntimeCompatibilityPveVM(t *testing.T) {
nodes := []NodeInfo{
{Hostname: "linux-1", Runtimes: []string{"process"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
{Hostname: "pve-1", Runtimes: []string{"process"}, Kind: "proxmox", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
}
req := WorkloadRequest{Spec: jobSpec("vmjob", "pve-vm", nil), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if got[0].Node != "pve-1" {
t.Errorf("Node = %q, want pve-1 (pve-vm requires proxmox kind)", got[0].Node)
}
}
// ---------------------------------------------------------------------------
// Constraints
// ---------------------------------------------------------------------------
func TestSchedule_ConstraintKindExcludesProxmox(t *testing.T) {
nodes := []NodeInfo{
{Hostname: "linux-1", Runtimes: []string{"process"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
{Hostname: "pve-1", Runtimes: []string{"process"}, Kind: "proxmox", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
}
req := WorkloadRequest{Spec: jobSpec("linuxonly", "process", []string{`node.kind == "linux"`}), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if got[0].Node != "linux-1" {
t.Errorf("Node = %q, want linux-1 (kind==linux)", got[0].Node)
}
}
func TestSchedule_ConstraintCPUsExcludesSmall(t *testing.T) {
nodes := threeLinuxNodes() // node-c has cpus=2
req := WorkloadRequest{Spec: jobSpec("big", "process", []string{"node.cpus >= 4"}), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if got[0].Node == "node-c" {
t.Errorf("Node = node-c, want node-a or node-b (cpus>=4)")
}
}
func TestSchedule_ConstraintNotInTags(t *testing.T) {
nodes := []NodeInfo{
{Hostname: "tagged", Runtimes: []string{"process"}, Tags: []string{"log-shipper"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
{Hostname: "clean", Runtimes: []string{"process"}, Tags: nil, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
}
req := WorkloadRequest{Spec: jobSpec("worker", "process", []string{`"log-shipper" not in node.tags`}), Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if got[0].Node != "clean" {
t.Errorf("Node = %q, want clean (log-shipper not in tags)", got[0].Node)
}
}
// ---------------------------------------------------------------------------
// Affinity
// ---------------------------------------------------------------------------
func TestSchedule_AffinityPrefersColocatedNode(t *testing.T) {
// Place a redis service first, then a worker with affinity for
// redis; the worker should prefer the node where redis already
// runs even if another node has more free capacity.
nodes := []NodeInfo{
{Hostname: "big", Runtimes: []string{"process"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
{Hostname: "small", Runtimes: []string{"process"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
}
redisReq := WorkloadRequest{Spec: serviceSpec("redis", "process", 1, nil), Namespace: "ns"}
redisPlacements, err := Schedule(nodes, redisReq)
if err != nil {
t.Fatalf("redis Schedule: %v", err)
}
// Redis lands on "big" (most free capacity). Now schedule the
// worker with affinity to redis; it should also land on "big".
workerReq := WorkloadRequest{Spec: &jobspec.WorkloadSpec{
Kind: "Job",
Name: "worker",
Count: 1,
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
Affinity: []jobspec.AffinityRule{
{Target: "redis", Weight: 1000},
},
}, Namespace: "ns"}
// The affinity is name-based; we need to seed the worker schedule
// with the redis placement so affinityScore can see it. Schedule
// does not take prior placements, so test affinityScore directly.
got := affinityScore(nodes[0], workerReq, redisPlacements)
if got <= 0 {
t.Errorf("affinityScore(big) = %d, want > 0 (redis colocated)", got)
}
gotSmall := affinityScore(nodes[1], workerReq, redisPlacements)
if gotSmall != 0 {
t.Errorf("affinityScore(small) = %d, want 0 (redis not colocated)", gotSmall)
}
}
func TestSchedule_AffinityCELExpression(t *testing.T) {
// Affinity with a CEL target: prefer nodes tagged "ssd".
nodes := []NodeInfo{
{Hostname: "hdd", Runtimes: []string{"process"}, Tags: []string{"hdd"}, Kind: "linux", CPU: 8, Memory: 8192, FreeCPU: 8, FreeMem: 8192},
{Hostname: "ssd", Runtimes: []string{"process"}, Tags: []string{"ssd"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096},
}
req := WorkloadRequest{Spec: &jobspec.WorkloadSpec{
Kind: "Job",
Name: "db",
Count: 1,
Runtime: &jobspec.RuntimeBlock{OneOf: "process"},
Affinity: []jobspec.AffinityRule{
{Target: `"ssd" in node.tags`, Weight: 10000},
},
}, Namespace: "ns"}
got, err := Schedule(nodes, req)
if err != nil {
t.Fatalf("Schedule: %v", err)
}
if got[0].Node != "ssd" {
t.Errorf("Node = %q, want ssd (affinity to ssd tag outweighs capacity)", got[0].Node)
}
}
// ---------------------------------------------------------------------------
// Error paths
// ---------------------------------------------------------------------------
func TestSchedule_EmptyNodes(t *testing.T) {
req := WorkloadRequest{Spec: jobSpec("x", "process", nil), Namespace: "ns"}
if _, err := Schedule(nil, req); err == nil {
t.Fatal("Schedule: expected error for empty nodes, got nil")
}
}
func TestSchedule_NilSpec(t *testing.T) {
if _, err := Schedule(threeLinuxNodes(), WorkloadRequest{}); err == nil {
t.Fatal("Schedule: expected error for nil spec, got nil")
}
}
func TestSchedule_UnknownKind(t *testing.T) {
req := WorkloadRequest{Spec: &jobspec.WorkloadSpec{Kind: "Cron", Name: "x", Count: 1}, Namespace: "ns"}
if _, err := Schedule(threeLinuxNodes(), req); err == nil {
t.Fatal("Schedule: expected error for unknown kind")
}
}
// ---------------------------------------------------------------------------
// Score unit tests
// ---------------------------------------------------------------------------
func TestScore_FitsAndDoesNotFit(t *testing.T) {
node := NodeInfo{Hostname: "n", Runtimes: []string{"process"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096}
req := WorkloadRequest{Spec: jobSpec("j", "process", nil), Namespace: "ns"}
score, fits := Score(node, req)
if !fits {
t.Error("fits = false, want true")
}
if score <= 0 {
t.Errorf("score = %d, want > 0", score)
}
}
func TestScore_RuntimeMismatchDoesNotFit(t *testing.T) {
node := NodeInfo{Hostname: "n", Runtimes: []string{"process"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096}
req := WorkloadRequest{Spec: jobSpec("j", "wasm", nil), Namespace: "ns"}
if _, fits := Score(node, req); fits {
t.Error("fits = true for wasm on process-only node, want false")
}
}
func TestScore_ConstraintFailsDoesNotFit(t *testing.T) {
node := NodeInfo{Hostname: "n", Runtimes: []string{"process"}, Kind: "linux", CPU: 4, Memory: 4096, FreeCPU: 4, FreeMem: 4096}
req := WorkloadRequest{Spec: jobSpec("j", "process", []string{`node.kind == "proxmox"`}), Namespace: "ns"}
if _, fits := Score(node, req); fits {
t.Error("fits = true for kind==proxmox on linux node, want false")
}
}
// ---------------------------------------------------------------------------
// allocID / isAllocFor helpers
// ---------------------------------------------------------------------------
func TestAllocID(t *testing.T) {
req := WorkloadRequest{Spec: &jobspec.WorkloadSpec{Name: "web"}, Namespace: "prod"}
if got := allocID(req, 2); got != "prod/web-2" {
t.Errorf("allocID = %q, want prod/web-2", got)
}
req.Namespace = ""
if got := allocID(req, 0); got != "default/web-0" {
t.Errorf("allocID = %q, want default/web-0", got)
}
}
func TestIsAllocFor(t *testing.T) {
cases := []struct {
allocID string
workload string
want bool
}{
{"ns/redis-0", "redis", true},
{"ns/redis-12", "redis", true},
{"ns/worker-0", "redis", false},
{"redis-0", "redis", true},
{"ns/web-canary-3", "web-canary", true},
}
for _, c := range cases {
if got := isAllocFor(c.allocID, c.workload); got != c.want {
t.Errorf("isAllocFor(%q,%q) = %v, want %v", c.allocID, c.workload, got, c.want)
}
}
}
// ---------------------------------------------------------------------------
// normalizeRuntime / hasRuntime
// ---------------------------------------------------------------------------
func TestHasRuntimeAliases(t *testing.T) {
cases := []struct {
name string
node NodeInfo
want bool
}{
{"wasm on wasmtime node", NodeInfo{Runtimes: []string{"wasmtime"}, Kind: "linux"}, true},
{"wasm on process node", NodeInfo{Runtimes: []string{"process"}, Kind: "linux"}, false},
{"pve-vm on linux node", NodeInfo{Runtimes: []string{"pve-vm"}, Kind: "linux"}, false},
{"pve-vm on proxmox node", NodeInfo{Runtimes: nil, Kind: "proxmox"}, true},
{"process on process node", NodeInfo{Runtimes: []string{"process"}, Kind: "linux"}, true},
{"empty runtime on any node", NodeInfo{Runtimes: []string{"process"}, Kind: "linux"}, true},
}
for _, c := range cases {
if c.name == "empty runtime on any node" {
// hasRuntime is only called when OneOf != "".
continue
}
if got := hasRuntime(c.node, "wasm"); c.name == "wasm on wasmtime node" || c.name == "wasm on process node" {
if got != c.want {
t.Errorf("%s: hasRuntime(wasm) = %v, want %v", c.name, got, c.want)
}
}
}
// Explicit pve-vm and process checks.
if !hasRuntime(NodeInfo{Runtimes: nil, Kind: "proxmox"}, "pve-vm") {
t.Error("pve-vm on proxmox node should fit")
}
if hasRuntime(NodeInfo{Runtimes: nil, Kind: "linux"}, "pve-vm") {
t.Error("pve-vm on linux node should not fit")
}
if !hasRuntime(NodeInfo{Runtimes: []string{"process"}, Kind: "linux"}, "process") {
t.Error("process on process node should fit")
}
}
+48 -3
View File
@@ -37,6 +37,8 @@ type Validator interface {
// - count must be 1 (or unset → 1); count > 1 is an error for Job
// (use a Service for replicas)
// - no Traefik route (a ServiceBlock is rejected)
// - task group (spec.Tasks) optional; when present, each task must
// have a unique name and a resolvable command (P06).
type JobValidator struct{}
// ServiceValidator validates the Service workload kind (R-012).
@@ -46,11 +48,14 @@ type JobValidator struct{}
// - count ≥ 1
// - restart required (mode must be service)
// - update required (strategy must be rolling/canary/blue-green)
// - runtime required
// - runtime required (unless a task group is present; each task
// can carry its own runtime — P06)
// - health block required (Traefik routing depends on health checks)
// - service block, if present, must have a valid bind (127.0.0.1
// opt-in per R-007; default is socket — empty bind is OK)
// - service block implied (Traefik route YES)
// - task group (spec.Tasks) optional; when present, each task must
// have a unique name and a resolvable command (P06).
type ServiceValidator struct{}
// DaemonSetValidator validates the DaemonSet workload kind (R-012).
@@ -60,6 +65,8 @@ type ServiceValidator struct{}
// - no ports (no Traefik route by default D-175)
// - no count (implicit = nodes matching condition)
// - restart required
// - task group (spec.Tasks) optional; when present, each task must
// have a unique name and a resolvable command (P06).
type DaemonSetValidator struct{}
// ValidatorFor returns the Validator for the given workload kind, or an
@@ -93,6 +100,7 @@ func (JobValidator) Validate(spec *jobspec.WorkloadSpec) error {
if spec.Service != nil {
errs = append(errs, "service block (Traefik route) is not allowed for Job (D-175)")
}
errs = append(errs, validateTaskGroup(spec)...)
return composeErrors("schema/Job", errs)
}
@@ -137,8 +145,8 @@ func (ServiceValidator) Validate(spec *jobspec.WorkloadSpec) error {
errs = append(errs, fmt.Sprintf("update strategy %q invalid (want one of rolling, canary, blue-green)", spec.Update.Strategy))
}
}
if spec.Runtime == nil {
errs = append(errs, "runtime block required for Service")
if spec.Runtime == nil && len(spec.Tasks) == 0 {
errs = append(errs, "runtime block required for Service (or a task group with per-task runtimes)")
}
if spec.Health == nil {
errs = append(errs, "health block required for Service (Traefik routing requires health checks)")
@@ -148,6 +156,7 @@ func (ServiceValidator) Validate(spec *jobspec.WorkloadSpec) error {
errs = append(errs, err.Error())
}
}
errs = append(errs, validateTaskGroup(spec)...)
return composeErrors("schema/Service", errs)
}
@@ -195,6 +204,7 @@ func (DaemonSetValidator) Validate(spec *jobspec.WorkloadSpec) error {
if spec.Restart == nil {
errs = append(errs, "restart block required for DaemonSet")
}
errs = append(errs, validateTaskGroup(spec)...)
return composeErrors("schema/DaemonSet", errs)
}
@@ -207,3 +217,38 @@ func composeErrors(name string, errs []string) error {
}
return fmt.Errorf("%s: %s", name, strings.Join(errs, "; "))
}
// validateTaskGroup validates the task-group list shared by all kinds
// (P06, PRD §9.1). When the spec carries a task group (spec.Tasks
// non-empty), each task must have a unique name and a resolvable
// command (the task's own Command, the task's runtime command, or the
// top-level runtime command as the per-group default). The top-level
// runtime is optional when tasks is present (each task can carry its
// own runtime). Returns nil when the spec has no task group.
func validateTaskGroup(spec *jobspec.WorkloadSpec) []string {
if len(spec.Tasks) == 0 {
return nil
}
var errs []string
seen := make(map[string]bool, len(spec.Tasks))
for i, task := range spec.Tasks {
if strings.TrimSpace(task.Name) == "" {
errs = append(errs, fmt.Sprintf("tasks[%d]: name is required", i))
} else if seen[task.Name] {
errs = append(errs, fmt.Sprintf("tasks[%d]: duplicate task name %q (names must be unique within the group)", i, task.Name))
} else {
seen[task.Name] = true
}
cmd := task.Command
if strings.TrimSpace(cmd) == "" && task.Runtime != nil {
cmd = task.Runtime.Command
}
if strings.TrimSpace(cmd) == "" && spec.Runtime != nil {
cmd = spec.Runtime.Command
}
if strings.TrimSpace(cmd) == "" {
errs = append(errs, fmt.Sprintf("tasks[%d]: command is required (set tasks[].command, tasks[].runtime.command, or top-level runtime.command)", i))
}
}
return errs
}
+180
View File
@@ -689,3 +689,183 @@ var (
_ Validator = ServiceValidator{}
_ Validator = DaemonSetValidator{}
)
func TestTaskGroup_Valid(t *testing.T) {
// P06: a valid task group — two tasks, each with a unique name
// and a resolvable command (own command). The top-level runtime
// is optional when each task carries its own.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 1,
Tasks: []jobspec.TaskGroupTask{
{Name: "app", Command: "/usr/bin/httpd"},
{Name: "sidecar", Command: "/bin/wasm-runner sidecar.wasm"},
},
Restart: &jobspec.RestartBlock{Mode: "service"},
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
Health: &jobspec.HealthBlock{CheckType: "http"},
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
}
if err := (ServiceValidator{}).Validate(spec); err != nil {
t.Fatalf("expected nil, got %v", err)
}
}
func TestTaskGroup_ValidInheritsTopLevelRuntime(t *testing.T) {
// P06: tasks without their own runtime inherit the top-level
// runtime command. The validator accepts this as long as the
// resolved command is non-empty.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 1,
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/default"},
Tasks: []jobspec.TaskGroupTask{
{Name: "app"},
{Name: "sidecar"},
},
Restart: &jobspec.RestartBlock{Mode: "service"},
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
Health: &jobspec.HealthBlock{CheckType: "http"},
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
}
if err := (ServiceValidator{}).Validate(spec); err != nil {
t.Fatalf("expected nil, got %v", err)
}
}
func TestTaskGroup_ValidTaskRuntimeCommand(t *testing.T) {
// P06: a task whose command is provided via the task's own
// runtime.command (no top-level runtime) is valid.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 1,
Tasks: []jobspec.TaskGroupTask{
{Name: "app", Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/usr/bin/httpd"}},
},
Restart: &jobspec.RestartBlock{Mode: "service"},
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
Health: &jobspec.HealthBlock{CheckType: "http"},
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
}
if err := (ServiceValidator{}).Validate(spec); err != nil {
t.Fatalf("expected nil, got %v", err)
}
}
func TestTaskGroup_MissingTaskName(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Tasks: []jobspec.TaskGroupTask{
{Command: "/usr/bin/httpd"},
{Name: "sidecar", Command: "/bin/wasm-runner"},
},
Restart: &jobspec.RestartBlock{Mode: "service"},
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
Health: &jobspec.HealthBlock{CheckType: "http"},
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
}
err := ServiceValidator{}.Validate(spec)
if err == nil {
t.Fatal("expected error for missing task name, got nil")
}
if !strings.Contains(err.Error(), "name is required") {
t.Errorf("error = %q, want 'name is required'", err.Error())
}
}
func TestTaskGroup_DuplicateTaskNames(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Tasks: []jobspec.TaskGroupTask{
{Name: "app", Command: "/usr/bin/httpd"},
{Name: "app", Command: "/bin/other"},
},
Restart: &jobspec.RestartBlock{Mode: "service"},
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
Health: &jobspec.HealthBlock{CheckType: "http"},
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
}
err := ServiceValidator{}.Validate(spec)
if err == nil {
t.Fatal("expected error for duplicate task names, got nil")
}
if !strings.Contains(err.Error(), "duplicate task name") {
t.Errorf("error = %q, want 'duplicate task name'", err.Error())
}
}
func TestTaskGroup_MissingCommand(t *testing.T) {
// P06: a task with no resolvable command (no task.Command, no
// task.Runtime, no top-level Runtime) is rejected.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Tasks: []jobspec.TaskGroupTask{
{Name: "app"},
},
Restart: &jobspec.RestartBlock{Mode: "service"},
Update: &jobspec.UpdateBlock{Strategy: "rolling"},
Health: &jobspec.HealthBlock{CheckType: "http"},
Ports: []jobspec.PortSpec{{Name: "http", Port: 8080}},
}
err := ServiceValidator{}.Validate(spec)
if err == nil {
t.Fatal("expected error for missing task command, got nil")
}
if !strings.Contains(err.Error(), "command is required") {
t.Errorf("error = %q, want 'command is required'", err.Error())
}
}
func TestTaskGroup_JobAcceptsTaskGroup(t *testing.T) {
// P06: task groups apply to all kinds, not just Service. Job
// accepts a task group with unique names + resolvable commands.
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "batch",
Count: 1,
Tasks: []jobspec.TaskGroupTask{
{Name: "step1", Command: "/bin/extract"},
{Name: "step2", Command: "/bin/transform"},
},
}
if err := (JobValidator{}).Validate(spec); err != nil {
t.Fatalf("expected nil, got %v", err)
}
}
func TestTaskGroup_DaemonSetAcceptsTaskGroup(t *testing.T) {
// P06: DaemonSet accepts a task group.
spec := &jobspec.WorkloadSpec{
Kind: "DaemonSet",
Name: "log-shipper",
Schedule: &jobspec.ScheduleBlock{Mode: "every-node"},
Restart: &jobspec.RestartBlock{Mode: "on-failure"},
Tasks: []jobspec.TaskGroupTask{
{Name: "collector", Command: "/bin/collect"},
{Name: "forwarder", Command: "/bin/forward"},
},
}
if err := (DaemonSetValidator{}).Validate(spec); err != nil {
t.Fatalf("expected nil, got %v", err)
}
}
func TestTaskGroup_NoTasksBackwardCompat(t *testing.T) {
// Backward compat: a spec with no Tasks is validated by the
// existing kind-specific rules (no task-group check fires).
spec := &jobspec.WorkloadSpec{
Kind: "Job",
Name: "backup",
Count: 1,
Runtime: &jobspec.RuntimeBlock{OneOf: "process", Command: "/bin/rsync"},
}
if err := (JobValidator{}).Validate(spec); err != nil {
t.Fatalf("expected nil, got %v", err)
}
}
+126
View File
@@ -0,0 +1,126 @@
package schema
import (
"fmt"
"strconv"
"strings"
"time"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
// UpdateValidator validates the rolling/canary/blue-green update stanza
// (P03). The schema validator (ServiceValidator) already enforces that
// the strategy is one of rolling/canary/blue-green and that the update
// block is present for a Service. UpdateValidator adds the
// field-level validation:
//
// - max_parallel: integer in [1, count] (defaults to 1 when unset)
// - min_healthy_time: a valid time.Duration when set (time.ParseDuration)
// - healthy_deadline: a valid time.Duration when set (time.ParseDuration)
// - canary: an integer count in [0, count] OR a percentage string of
// the form "<n>%" where n is in [0, 100] (the parser already accepts
// both shapes; the validator accepts them too). Only meaningful
// for the canary strategy; ignored (but still validated for shape)
// for rolling/blue-green.
// - auto_promote: boolean (no validation beyond the parser's
// true/false parse; the field is always populated)
//
// The validator is pure (no I/O). Violations return a clear error
// listing every problem found, mirroring the per-field style of
// ServiceValidator.
type UpdateValidator struct{}
// Validate validates the UpdateBlock on the given spec. The spec must
// be non-nil and carry a Count (services have count ≥ 1 per
// ServiceValidator). When spec.Update is nil the validator returns an
// error (the update block is required for Service; this validator
// assumes the caller has already established the spec is a Service).
func (UpdateValidator) Validate(spec *jobspec.WorkloadSpec) error {
if spec == nil {
return fmt.Errorf("schema/Update: spec is nil")
}
if spec.Update == nil {
return fmt.Errorf("schema/Update: update block is nil")
}
var errs []string
u := spec.Update
switch u.Strategy {
case "rolling", "canary", "blue-green":
case "":
errs = append(errs, "update strategy required (one of rolling, canary, blue-green)")
default:
errs = append(errs, fmt.Sprintf("update strategy %q invalid (want one of rolling, canary, blue-green)", u.Strategy))
}
// max_parallel defaults to 1 when unset (0); validate the range
// only when the user has set it explicitly.
if u.MaxParallel != 0 {
if u.MaxParallel < 1 {
errs = append(errs, fmt.Sprintf("update.max_parallel must be ≥ 1, got %d", u.MaxParallel))
}
if spec.Count > 0 && u.MaxParallel > spec.Count {
errs = append(errs, fmt.Sprintf("update.max_parallel %d exceeds count %d (must be 1..count)", u.MaxParallel, spec.Count))
}
}
if u.MinHealthyTime != "" {
if _, err := time.ParseDuration(u.MinHealthyTime); err != nil {
errs = append(errs, fmt.Sprintf("update.min_healthy_time %q is not a valid duration: %v", u.MinHealthyTime, err))
}
}
if u.HealthyDeadline != "" {
if _, err := time.ParseDuration(u.HealthyDeadline); err != nil {
errs = append(errs, fmt.Sprintf("update.healthy_deadline %q is not a valid duration: %v", u.HealthyDeadline, err))
}
}
// canary accepts an integer count (0..count) or a percentage
// ("<n>%" with n in 0..100). The field is only meaningful for the
// canary strategy but we validate the shape regardless so a typo
// in a rolling/blue-green stanza still surfaces.
if u.Canary != "" {
if err := validateCanary(u.Canary, spec.Count); err != nil {
errs = append(errs, err.Error())
}
}
// auto_promote is a bool; no extra validation beyond the parser.
return composeErrors("schema/Update", errs)
}
// validateCanary validates the canary field shape: either an integer
// count (0..count) or a percentage string "<n>%" (n in 0..100). count
// is the spec.Count; when count is 0 (e.g. a DaemonSet or unset), the
// integer-count upper bound is not enforced (only the percentage
// bound is enforced, since percentage does not depend on count).
func validateCanary(canary string, count int) error {
c := strings.TrimSpace(canary)
if c == "" {
return nil
}
if strings.HasSuffix(c, "%") {
nStr := strings.TrimSuffix(c, "%")
n, err := strconv.Atoi(strings.TrimSpace(nStr))
if err != nil {
return fmt.Errorf("update.canary %q is not a valid percentage (want \"<n>%%\")", canary)
}
if n < 0 || n > 100 {
return fmt.Errorf("update.canary percentage %d out of range (want 0..100)", n)
}
return nil
}
n, err := strconv.Atoi(c)
if err != nil {
return fmt.Errorf("update.canary %q is not a valid count or percentage (want integer or \"<n>%%\")", canary)
}
if n < 0 {
return fmt.Errorf("update.canary count %d must be ≥ 0", n)
}
if count > 0 && n > count {
return fmt.Errorf("update.canary count %d exceeds count %d (must be 0..count)", n, count)
}
return nil
}
+464
View File
@@ -0,0 +1,464 @@
package schema
import (
"strings"
"testing"
"git.cloudinit.dev/coreci/orca/internal/jobspec"
)
func TestUpdateValidator_ValidRolling(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
MaxParallel: 2,
MinHealthyTime: "30s",
HealthyDeadline: "5m",
},
}
v := UpdateValidator{}
if err := v.Validate(spec); err != nil {
t.Fatalf("expected nil, got %v", err)
}
}
func TestUpdateValidator_ValidCanary(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
MaxParallel: 2,
MinHealthyTime: "30s",
HealthyDeadline: "5m",
Canary: "10%",
AutoPromote: true,
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("expected nil, got %v", err)
}
}
func TestUpdateValidator_ValidBlueGreen(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 3,
Update: &jobspec.UpdateBlock{
Strategy: "blue-green",
MinHealthyTime: "1m",
HealthyDeadline: "10m",
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("expected nil, got %v", err)
}
}
func TestUpdateValidator_ValidCanaryIntegerCount(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "1",
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("integer canary count 1 should be valid, got %v", err)
}
}
func TestUpdateValidator_ValidCanaryZeroPercent(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "0%",
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("0%% canary should be valid, got %v", err)
}
}
func TestUpdateValidator_ValidCanaryHundredPercent(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "100%",
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("100%% canary should be valid, got %v", err)
}
}
func TestUpdateValidator_EmptyDurationsOK(t *testing.T) {
// Empty min_healthy_time / healthy_deadline should be accepted
// (they are optional; defaults are applied by the executor).
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("empty durations should be valid, got %v", err)
}
}
func TestUpdateValidator_NilSpec(t *testing.T) {
if err := (UpdateValidator{}).Validate(nil); err == nil {
t.Fatal("expected error for nil spec")
}
}
func TestUpdateValidator_NilUpdateBlock(t *testing.T) {
spec := &jobspec.WorkloadSpec{Kind: "Service", Name: "web", Count: 2}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for nil update block")
}
if !strings.Contains(err.Error(), "update block is nil") {
t.Errorf("error = %q, want 'update block is nil'", err.Error())
}
}
func TestUpdateValidator_InvalidStrategy(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "recreate",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for invalid strategy")
}
if !strings.Contains(err.Error(), "strategy") || !strings.Contains(err.Error(), "invalid") {
t.Errorf("error = %q, want 'strategy ... invalid'", err.Error())
}
}
func TestUpdateValidator_EmptyStrategy(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for empty strategy")
}
if !strings.Contains(err.Error(), "strategy required") {
t.Errorf("error = %q, want 'strategy required'", err.Error())
}
}
func TestUpdateValidator_MaxParallelZero(t *testing.T) {
// max_parallel=0 means "unset" → default 1; accepted.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
MaxParallel: 0,
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("max_parallel=0 (unset) should be valid, got %v", err)
}
}
func TestUpdateValidator_MaxParallelNegative(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
MaxParallel: -1,
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for negative max_parallel")
}
if !strings.Contains(err.Error(), "max_parallel") {
t.Errorf("error = %q, want 'max_parallel'", err.Error())
}
}
func TestUpdateValidator_MaxParallelExceedsCount(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
MaxParallel: 5,
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for max_parallel > count")
}
if !strings.Contains(err.Error(), "exceeds count") {
t.Errorf("error = %q, want 'exceeds count'", err.Error())
}
}
func TestUpdateValidator_MaxParallelEqualsCount(t *testing.T) {
// max_parallel == count is the upper bound; valid.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 3,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
MaxParallel: 3,
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("max_parallel==count should be valid, got %v", err)
}
}
func TestUpdateValidator_InvalidMinHealthyTime(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
MinHealthyTime: "not-a-duration",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for invalid min_healthy_time")
}
if !strings.Contains(err.Error(), "min_healthy_time") {
t.Errorf("error = %q, want 'min_healthy_time'", err.Error())
}
}
func TestUpdateValidator_InvalidHealthyDeadline(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "rolling",
HealthyDeadline: "nope",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for invalid healthy_deadline")
}
if !strings.Contains(err.Error(), "healthy_deadline") {
t.Errorf("error = %q, want 'healthy_deadline'", err.Error())
}
}
func TestUpdateValidator_CanaryNegative(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "-1",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for negative canary count")
}
if !strings.Contains(err.Error(), "canary") {
t.Errorf("error = %q, want 'canary'", err.Error())
}
}
func TestUpdateValidator_CanaryExceedsCount(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "5",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for canary > count")
}
if !strings.Contains(err.Error(), "exceeds count") {
t.Errorf("error = %q, want 'exceeds count'", err.Error())
}
}
func TestUpdateValidator_CanaryPercentOver100(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "150%",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for canary > 100%")
}
if !strings.Contains(err.Error(), "out of range") {
t.Errorf("error = %q, want 'out of range'", err.Error())
}
}
func TestUpdateValidator_CanaryPercentNegative(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "-10%",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for negative canary percent")
}
if !strings.Contains(err.Error(), "out of range") {
t.Errorf("error = %q, want 'out of range'", err.Error())
}
}
func TestUpdateValidator_CanaryNotANumber(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "abc",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for non-numeric canary")
}
if !strings.Contains(err.Error(), "not a valid") {
t.Errorf("error = %q, want 'not a valid'", err.Error())
}
}
func TestUpdateValidator_CanaryPercentNotANumber(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "xx%",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error for non-numeric canary percent")
}
if !strings.Contains(err.Error(), "not a valid percentage") {
t.Errorf("error = %q, want 'not a valid percentage'", err.Error())
}
}
func TestUpdateValidator_CanaryCountZeroOK(t *testing.T) {
// canary=0 is the lower bound; valid.
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 4,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "0",
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("canary=0 should be valid, got %v", err)
}
}
func TestUpdateValidator_CanaryPercentWithoutCountOK(t *testing.T) {
// A percentage canary does not depend on count; valid even when
// count is 0 (e.g. DaemonSet-shaped spec bypassing the Service
// validator — defensive).
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 0,
Update: &jobspec.UpdateBlock{
Strategy: "canary",
Canary: "25%",
},
}
if err := (UpdateValidator{}).Validate(spec); err != nil {
t.Fatalf("percentage canary with count=0 should be valid, got %v", err)
}
}
func TestUpdateValidator_MultipleErrors(t *testing.T) {
spec := &jobspec.WorkloadSpec{
Kind: "Service",
Name: "web",
Count: 2,
Update: &jobspec.UpdateBlock{
Strategy: "recreate",
MaxParallel: 99,
MinHealthyTime: "nope",
HealthyDeadline: "also-nope",
Canary: "200%",
},
}
err := (UpdateValidator{}).Validate(spec)
if err == nil {
t.Fatal("expected error, got nil")
}
for _, want := range []string{
"strategy",
"max_parallel",
"min_healthy_time",
"healthy_deadline",
"canary",
} {
if !strings.Contains(err.Error(), want) {
t.Errorf("error %q missing %q", err.Error(), want)
}
}
}
// Compile-time assertion that UpdateValidator implements Validator.
var _ Validator = UpdateValidator{}