2 Commits

Author SHA1 Message Date
Jon Chery 47674969a1 docs(P1): complete gitops-operators domain phase
---ci---
project: atelier
phase: 1
milestone: v0.3
status: complete
requirements:
  covered: [ATELIER-60, ATELIER-61, ATELIER-62, ATELIER-63, ATELIER-64]
  partial: []
---/ci---
2026-08-05 03:20:48 +00:00
Jon Chery ce36db0579 docs(ship): P0 complete — checkpoint + roadmap status
---ci---
project: atelier
phase: 0
milestone: v0.3
status: complete
phase_tag: v0.2.0
release_id: 468
---/ci---
2026-08-05 03:17:10 +00:00
7 changed files with 788 additions and 3 deletions
+4 -2
View File
@@ -1,9 +1,11 @@
{
"phase": 0,
"stage": "clarify",
"stage": "complete",
"milestone": "v0.3",
"phase_role": "pre_execution",
"project": "atelier",
"attempts": 0,
"updated_at": "2026-08-05T03:08:00Z"
"updated_at": "2026-08-05T03:17:00Z",
"phase_tag": "v0.2.0",
"release_id": 468
}
+1 -1
View File
@@ -84,7 +84,7 @@ NFR milestone: no separate minor tag. The final patch (v0.1.5) IS the v0.2 deliv
| Phase | Name | Type | Status | Key Deliverables |
|-------|------|------|--------|------------------|
| 0 | Pre-Execution | docs | active | Spec, clarify, research, ideate, plan, PERSONAS.md (extends platform-engineer, adds ml-engineer) |
| 0 | Pre-Execution | docs | complete | Spec, clarify, research, ideate, plan, PERSONAS.md (extends platform-engineer, adds ml-engineer) — shipped v0.2.0 |
| 1 | GitOps + Operators Domain | docs | pending | domains/gitops-operators/{first-principles, argocd, flux, operators, progressive-delivery}.md |
| 2 | AI/ML Domain | docs | pending | domains/ai-ml/{first-principles, data-versioning, model-evaluation, serving, monitoring-drift}.md |
| 3 | i18n + Compliance Domains | docs | pending | domains/i18n/{first-principles, locale-resources, formatting, rtl-bidi, testing-i18n}.md, domains/compliance/{first-principles, audit-logs, data-retention, policy-as-code, evidence}.md |
+176
View File
@@ -0,0 +1,176 @@
# ArgoCD — Derived Rules
> Derives from `domains/gitops-operators/first-principles.md`.
> Applies P1P10 to ArgoCD specifically. For the ArgoCD-vs-Flux
> decision, see the decision matrix at the end of this doc and in
> `flux.md`.
## What ArgoCD Is (P1 Git is the Source of Truth, P3 Pull, Don't Push)
- ArgoCD is a pull-based GitOps controller for Kubernetes. It runs
inside the target cluster, pulls desired state from git, and
reconciles the cluster to match. CI never holds `kubectl` rights
against the cluster (P3).
- An Application is a declarative binding of "this git path" to
"this cluster destination." The Application CRD is the unit of
reconciliation. The cluster is a derivative of git, never the
authority (P1).
- ArgoCD supports Helm charts, Kustomize overlays, ksonnet, and raw
manifests as source formats — see `domains/kubernetes/helm.md`
and `domains/kubernetes/kustomize.md`.
## Application CRD (P2 Declarative Over Imperative, P4 Continuous Reconciliation)
- An Application declares `source` (repo, path, revision, chart),
`destination` (server, namespace), and `syncPolicy`. The
reconciler loops continuously; drift is corrected automatically,
not on-demand (P4).
```yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: payments-api
namespace: argocd
spec:
source:
repoURL: https://git.example.com/platform/payments
targetRevision: 1.2.3
path: manifests/prod
destination:
server: https://kubernetes.default.svc
namespace: payments
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=false
```
- `automated.prune: true` deletes resources removed from git.
`selfHeal: true` corrects hand-edited drift back to git (P8).
Disable both for workloads that need manual approval gates.
## App-of-Apps (P6 Operators Encode Domain Knowledge, C6 Composability)
- The App-of-Apps pattern: one root Application points at a git
directory of child Application manifests. The root app reconciles
the children; the children reconcile the workloads. This is the
ArgoCD expression of composition — a fleet of apps as a tree of
Applications.
- Use App-of-Apps for cluster bootstrapping (one repo, many
clusters, many apps). Do not use it as a substitute for a package
manager; if you are templating hundreds of near-identical
Applications, use a generator (ApplicationSet) instead.
## Sync Waves and Hooks (P4 Continuous Reconciliation, P7 Reversibility)
- Sync waves order resources within a sync: `PreSync``Sync`
`PostSync`. Use waves to run a job before a Deployment, or a
migration before the app that depends on it.
- Sync hooks (`PreSync`, `Sync`, `PostSync`, `SyncFail`) are
Resources annotated to execute at a wave boundary. A `SyncFail`
hook runs on sync failure — the abort path (P7).
- Wave ordering is a correctness mechanism, not a performance one.
Mis-ordered waves (e.g., app starts before its migration job)
are a correctness bug.
## Health and Status (P9 Failure is Observable and Surfaced)
- ArgoCD assesses every resource's health (`Healthy`, `Progressing`,
`Degraded`, `Missing`, `Suspended`) and surfaces the aggregate as
Application status. Sync status (`Synced`, `OutOfSync`) reports
drift against git.
- Health checks are pluggable via Lua scripts for custom CRDs. An
Operator-managed CRD without a health check reads as `Progressing`
forever — write one (see `operators.md`).
- Out-of-sync or degraded status must emit a notification (Slack,
PagerDuty, webhook). Silent drift is the bug (P9). Wire status to
`domains/observability/metrics.md`.
## Diff and Drift (P8 Reconcile, Don't Mutate by Hand, P4)
- `argocd app diff` shows the diff between git and live cluster.
A non-empty diff on a synced app is hand-edit drift — the
recovery is `selfHeal`, not a manual `kubectl apply` (P8).
- Drift detection runs continuously (P4). The gap between "git
changed" and "cluster matches git" is observable, not assumed.
## RBAC and SSO (P10 Least Privilege Reconciliation)
- ArgoCD's own RBAC governs who can view, sync, and admin
Applications. Bind to SSO (OIDC, SAML) for human identity; bind
the controller's service account to a Role scoped to the
namespaces it reconciles.
- The controller's credentials must not be `cluster-admin` (P10).
Use namespace-scoped Roles via `ApplicationSet` namespaces or
cluster-wide AppProject restrictions. See
`domains/kubernetes/rbac.md` and `domains/security/authorization.md`.
- AppProjects bound the blast radius of what an Application can
deploy (allowed repos, destinations, roles). One AppProject per
team or environment; the default project is for nothing in
production.
## Multi-Cluster (P4 Locality, P10)
- ArgoCD registers external clusters by secret. The controller
pulls from git and pushes to the registered cluster's API server.
The "pull, don't push" boundary (P3) is between the target
cluster's reconciler and CI — the controller-to-apiserver hop is
internal to the platform.
- Scope each registered cluster's credentials to the namespaces
ArgoCD manages there. Do not register a cluster with cluster-admin
and call it done (P10).
## Sync Windows (P5 Reversibility, P7)
- Sync windows restrict when automated sync runs (e.g., no syncs
during business hours, or syncs only in a maintenance window).
They are a reversibility mechanism: a bad commit lands in git,
but the sync window holds it until review.
- Sync windows do not replace health monitoring (P9). A degraded
app inside a window is still an incident.
## Secrets (P10, cross-link security/secrets)
- Do not store raw Secrets in the GitOps repo. Use a sealed-secret
controller (Bitnami Sealed Secrets, SOPS, External Secrets
Operator) so the git store holds encrypted material only. See
`domains/security/secrets.md` for the general secret-hygiene
principles.
## ArgoCD vs Flux — Decision Matrix (IDEATE-21, D-039)
| Axis | ArgoCD | Flux |
|------|--------|------|
| Architecture | Monolithic controller + Application CRD | Composable GitOps Toolkit controllers (source, kustomize, helm, notification) |
| Reconciliation unit | Application (one CRD per app) | Kustomization / HelmRelease (one per deploy unit) |
| UI | Web UI + CLI (full dashboard, tree view, diff viewer) | CLI-first; UI via Weave GitOps or FluxUI (add-on) |
| Sync model | Periodic poll or webhook; sync waves + hooks | Poll + webhook; runs continuously, no explicit sync waves |
| Multi-cluster | One ArgoCD manages many clusters (hub-and-spoke) | One Flux per cluster (per-cluster autonomy) |
| Templating in repo | Helm, Kustomize, ksonnet, raw manifests, Jsonnet | Helm, Kustomize, raw manifests |
| RBAC | Built-in RBAC + SSO + AppProjects | Kubernetes RBAC (no built-in RBAC layer) |
| Progressive delivery | Argo Rollouts (sister project, tight integration) | Flagger (sister project, tight integration) |
| Best for | Teams wanting a UI, multi-cluster from one pane, App-of-Apps bootstrapping | Teams wanting composable controllers, per-cluster autonomy, minimal footprint |
| Watch out for | Monolithic controller scaling, UI as ops crutch, AppProject sprawl | No native UI, steeper learning curve, manual multi-cluster orchestration |
- Use ArgoCD when you want a UI, central multi-cluster management,
and sync-wave ordering. Use Flux when you want composable
controllers, per-cluster autonomy, and a minimal footprint.
- Both are CNCF graduated and both implement the OpenGitOps
principles. The choice is architectural fit, not correctness. See
`flux.md` for the Flux-side perspective.
## What Violates ArgoCD Discipline
| Violation | Principle |
|-----------|-----------|
| CI pipeline with `kubectl` rights pushing to the cluster | P3 Pull, Don't Push |
| `argocd app set` used as the steady state instead of git | P1 Git is the Source of Truth |
| `selfHeal: false` on a prod app with no manual gate | P8 Reconcile, Don't Mutate by Hand |
| Controller ServiceAccount bound to `cluster-admin` | P10 Least Privilege Reconciliation |
| Sync failure with no notification wired | P9 Failure is Observable and Surfaced |
| AppProject with no destination restrictions in prod | P10 Least Privilege Reconciliation |
| Raw Secret in the GitOps repo | P10, `domains/security/secrets.md` |
| Manual `kubectl edit` on an ArgoCD-managed resource | P8 Reconcile, Don't Mutate by Hand |
@@ -0,0 +1,131 @@
# GitOps + Operators — First Principles
## 1. The Principles
### P1. Git is the Source of Truth
Desired state lives in a versioned, immutable git store. The
cluster is a derivative of git, never the authority. If a state
exists only in the cluster and not in git, it is drift, not truth.
The commit history is the audit trail and the rollback path.
### P2. Declarative Over Imperative
Express the desired cluster state, not the commands to reach it.
A manifest says what should exist; the reconciler makes it so.
Imperative `kubectl` is for inspection and incident response, not
for the steady state. This is the GitOps expression of
`domains/kubernetes/P1 Declarative Desired State` and
`domains/infrastructure-as-code/P1 Declarative Intent`.
### P3. Pull, Don't Push
Agents running inside the target pull desired state from git; the
target never accepts outside push credentials. No CI pipeline holds
`kubectl` rights against the production cluster. The cluster reaches
out to git, not the other way around. This is the security primitive
of GitOps: the blast radius of a compromised CI is bounded by what CI
can push, and a pull model gives CI nothing to push.
### P4. Continuous Reconciliation
The reconciliation loop is the primitive. Drift is detected and
corrected automatically, not on-demand. A manual `apply` is an
exception, not the workflow. The loop runs continuously; the gap
between "git changed" and "cluster matches git" is measured in
seconds, not tickets.
### P5. State is Immutable and Versioned
Every change to desired state is a commit. History is the audit
trail and the rollback path. A revert is a rollback; a force-push is
history deletion. The git store is treated like
`domains/infrastructure-as-code/P3 State is Truth` — lose it or
tamper with it, and you lose the ability to reason about the system.
### P6. Operators Encode Domain Knowledge
Operational expertise lives as CRDs plus controllers, not as
runbooks that humans must remember. An operator is a control loop
that encodes how to reconcile a specific domain (a database, a
message queue, a certificate). The operator is the deepest
expression of `domains/kubernetes/P1 Declarative Desired State`
the domain knowledge is the desired state.
### P7. Progressive Delivery is Reversible by Construction
Canary and blue-green are staged, metric-gated, and one-command
abortable. Promotion without a rollback path is a violation. A
rollout that cannot be aborted is a deploy, not a progressive
delivery. This is the GitOps extension of
`domains/devops/P5 Progressive Delivery` and
`domains/kubernetes/P10 Roll Forward, Roll Back`.
### P8. Reconcile, Don't Mutate by Hand
Manual `kubectl apply` or `kubectl edit` on a GitOps-managed
resource is an incident. The reconciler will overwrite the hand
edit on the next loop; the hand edit was never truth. Drift back to
git is the recovery, not the failure. This is the GitOps angle on
`domains/infrastructure-as-code/P9 Drift is Recoverable`.
### P9. Failure is Observable and Surfaced
Sync failures, health degradation, and rollout-stall events emit
status and notifications. Silent drift is the bug. A GitOps
controller that fails to sync without surfacing the failure has
violated the contract — you cannot fix what you cannot see
(`domains/observability/metrics.md`).
### P10. Least Privilege Reconciliation
The controller's credentials are scoped to the namespaces and
resources it reconciles. No `cluster-admin` GitOps robots. One
credential set per boundary; the reconciler sees only what it
reconciles. This is the GitOps angle on
`domains/kubernetes/P7 RBAC by Intent, Not Identity` and
`domains/security/authorization.md`.
## 2. Core Principle Trace
Each GitOps + Operators P-rule derives from one or more core
C-rules (C1C8). The matrix extension lands in P4 of the v0.3
plan; the traces below are authoritative.
| P-rule | Core | Why |
|--------|------|-----|
| P1 Git is the Source of Truth | C1, C5 | Correctness of state; reversibility via history |
| P2 Declarative Over Imperative | C2, C3 | Clarity of intent; simplicity of mental model |
| P3 Pull, Don't Push | C1, C4 | Correctness via security; locality of credentials |
| P4 Continuous Reconciliation | C7, C1 | Observability of drift; correctness of convergence |
| P5 State is Immutable and Versioned | C5 | Reversibility via version history |
| P6 Operators Encode Domain Knowledge | C6, C2 | Composability of expertise; clarity of operational intent |
| P7 Progressive Delivery is Reversible | C5, C1 | Reversibility of promotion; correctness of abort |
| P8 Reconcile, Don't Mutate by Hand | C1, C7 | Correctness of single source; observability of drift |
| P9 Failure is Observable and Surfaced | C7 | Observability of reconciliation |
| P10 Least Privilege Reconciliation | C1, C8 | Correctness via security; economy of trust |
## 3. What Violates These Principles
| Violation | Principle Breached |
|-----------|-------------------|
| CI pipeline pushes manifests to the cluster | P3 Pull, Don't Push |
| A resource exists in the cluster but not in git | P1 Git is the Source of Truth |
| `kubectl edit` on a GitOps-managed resource | P8 Reconcile, Don't Mutate by Hand |
| Reconciler with `cluster-admin` ClusterRoleBinding | P10 Least Privilege Reconciliation |
| Sync failure with no status or notification | P9 Failure is Observable and Surfaced |
| Canary with no abort/rollback path | P7 Progressive Delivery is Reversible |
| Operator runbook that exists only in a wiki | P6 Operators Encode Domain Knowledge |
| Reconciler that applies on a cron, not continuously | P4 Continuous Reconciliation |
| Force-push rewrites GitOps repo history | P5 State is Immutable and Versioned |
| Imperative deploy script as the steady state | P2 Declarative Over Imperative |
## 4. Relationship to Other Domains
GitOps + Operators is the deployment-automation layer above
`domains/kubernetes/` and `domains/infrastructure-as-code/`. It
borrows their declarative-reconciliation model and adds the
git-as-source-of-truth and pull-based credential boundaries. Cross
links are one-directional (per D-026 extended):
- `domains/kubernetes/P1 Declarative Desired State` ← P2
- `domains/kubernetes/P10 Roll Forward, Roll Back` ← P7
- `domains/infrastructure-as-code/P1 Declarative Intent` ← P2
- `domains/infrastructure-as-code/P3 State is Truth` ← P1, P5
- `domains/infrastructure-as-code/P9 Drift is Recoverable` ← P4, P8
- `domains/devops/P4 Rollback First` ← P5, P7
- `domains/devops/P5 Progressive Delivery` ← P7
- `domains/devops/P6 Configuration as Code` ← P1, P2
- `domains/security/secrets.md` ← P3, P10 (reconciliation credentials)
- `domains/security/supply-chain.md` ← P5 (signed, immutable provenance)
- `domains/observability/metrics.md` ← P4, P9 (reconciliation + rollout metrics)
+159
View File
@@ -0,0 +1,159 @@
# Flux — Derived Rules
> Derives from `domains/gitops-operators/first-principles.md`.
> Applies P1P10 to Flux specifically. For the ArgoCD-vs-Flux
> decision, see the decision matrix at the end of this doc and in
> `argocd.md`.
## What Flux Is (P1 Git is the Source of Truth, P3 Pull, Don't Push)
- Flux is a set of composable controllers — the GitOps Toolkit —
that run inside the target cluster, pull desired state from git
or OCI registries, and reconcile the cluster to match. CI never
holds `kubectl` rights against the cluster (P3).
- The composable-controller architecture is a C6 (Composability)
exemplar: each controller does one thing (source, kustomize, helm,
notification) and the controllers compose into a full GitOps
system.
- Flux supports Helm releases, Kustomize overlays, and raw
manifests — see `domains/kubernetes/helm.md` and
`domains/kubernetes/kustomize.md`.
## GitOps Toolkit Controllers (P6 Composability, P4 Continuous Reconciliation)
- **source-controller** — pulls git, Helm, OCI, and bucket sources;
emits artifacts (tarballs) with a digest. The source is the
pinned input to reconciliation (P5 versioning by digest).
- **kustomize-controller** — reconciles Kustomization CRDs against
the artifacts from source-controller. Runs continuously (P4).
- **helm-controller** — reconciles HelmRelease CRDs against Helm
charts from source-controller.
- **notification-controller** — emits events and notifications for
sync, health, and source-readiness events (P9).
- **image-automation-controller** (optional) — updates git with new
image tags when a policy matches, closing the "latest image"
loop declaratively.
## Kustomization CRD (P2 Declarative Over Imperative, P4)
- A Kustomization binds "this source" to "this target namespace"
with a reconciliation interval. The reconciler loops
continuously; drift is corrected automatically (P4).
```yaml
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: payments-api
namespace: flux-system
spec:
sourceRef:
kind: GitRepository
name: platform
namespace: flux-system
path: ./manifests/prod
targetNamespace: payments
interval: 1m
prune: true
wait: true
healthChecks:
- apiVersion: apps/v1
kind: Deployment
name: payments-api
namespace: payments
```
- `prune: true` deletes resources removed from git. `wait: true`
waits for health checks before declaring the Kustomization ready.
Disable prune for workloads that need manual removal gates.
## HelmRelease CRD (P6 Composability, cross-link helm.md)
- A HelmRelease binds a Helm chart (from a HelmRepository or OCI
source) to target values and a target namespace. helm-controller
renders and applies it. See `domains/kubernetes/helm.md` for the
chart model.
- Pin the chart version in the HelmRepository or the HelmRelease.
Never float `latest` — unversioned charts drift (P5).
## OCI Sources (P5 State is Immutable and Versioned)
- source-controller can pull from OCI registries (Helm charts as
OCI artifacts, or generic OCI repositories). The digest is the
version — immutable by construction (P5).
- OCI sources close the supply-chain loop: the manifest is signed
and immutable in the registry, and Flux pulls it by digest. Cross-
link `domains/security/supply-chain.md` for signed-provenance
principles.
## Reconciliation and Drift (P4 Continuous Reconciliation, P8)
- Flux reconciles on `interval` (default 1m) and on webhook event.
Drift between git and cluster is detected each interval and
corrected (with `prune` + `selfHeal` semantics).
- Hand-edited drift on a Flux-managed resource is overwritten on the
next loop — the hand edit was never truth (P8). The recovery is
to fix git, not to `kubectl apply`.
## Notifications and Events (P9 Failure is Observable and Surfaced)
- notification-controller emits events for source readiness, sync
success/failure, and health transitions. Wire them to Slack,
PagerDuty, or a webhook. Silent drift is the bug (P9).
- Events flow to `domains/observability/metrics.md` via the
notification controller's provider model — sync and health as
first-class signals.
## RBAC and Multi-Cluster (P10 Least Privilege Reconciliation, P4)
- Flux's controllers run with a ServiceAccount in `flux-system`.
Scope that account to the namespaces Flux reconciles. Do not bind
it to `cluster-admin` (P10). See `domains/kubernetes/rbac.md` and
`domains/security/authorization.md`.
- Flux is per-cluster by design (one Flux install per cluster). For
multi-cluster, use one repo with per-cluster paths, or a fleet
tool that bootstraps Flux per cluster. Per-cluster autonomy is a
feature, not a limitation — it bounds the blast radius of a
compromised controller (P4 locality, P10).
## Secrets (P10, cross-link security/secrets)
- Do not store raw Secrets in the GitOps repo. Use the
SOPS-compatible decryption in kustomize-controller, or External
Secrets Operator, so the git store holds encrypted material only.
See `domains/security/secrets.md`.
## ArgoCD vs Flux — Decision Matrix (IDEATE-21, D-039)
| Axis | ArgoCD | Flux |
|------|--------|------|
| Architecture | Monolithic controller + Application CRD | Composable GitOps Toolkit controllers (source, kustomize, helm, notification) |
| Reconciliation unit | Application (one CRD per app) | Kustomization / HelmRelease (one per deploy unit) |
| UI | Web UI + CLI (full dashboard, tree view, diff viewer) | CLI-first; UI via Weave GitOps or FluxUI (add-on) |
| Sync model | Periodic poll or webhook; sync waves + hooks | Poll + webhook; runs continuously, no explicit sync waves |
| Multi-cluster | One ArgoCD manages many clusters (hub-and-spoke) | One Flux per cluster (per-cluster autonomy) |
| Templating in repo | Helm, Kustomize, ksonnet, raw manifests, Jsonnet | Helm, Kustomize, raw manifests |
| RBAC | Built-in RBAC + SSO + AppProjects | Kubernetes RBAC (no built-in RBAC layer) |
| Progressive delivery | Argo Rollouts (sister project, tight integration) | Flagger (sister project, tight integration) |
| Best for | Teams wanting a UI, multi-cluster from one pane, App-of-Apps bootstrapping | Teams wanting composable controllers, per-cluster autonomy, minimal footprint |
| Watch out for | Monolithic controller scaling, UI as ops crutch, AppProject sprawl | No native UI, steeper learning curve, manual multi-cluster orchestration |
- Use Flux when you want composable controllers, per-cluster
autonomy, and a minimal footprint. Use ArgoCD when you want a UI,
central multi-cluster management, and sync-wave ordering.
- Both are CNCF graduated and both implement the OpenGitOps
principles. The choice is architectural fit, not correctness. See
`argocd.md` for the ArgoCD-side perspective.
## What Violates Flux Discipline
| Violation | Principle |
|-----------|-----------|
| CI pipeline with `kubectl` rights pushing to the cluster | P3 Pull, Don't Push |
| HelmRelease with no pinned chart version | P5 State is Immutable and Versioned |
| Flux ServiceAccount bound to `cluster-admin` | P10 Least Privilege Reconciliation |
| Kustomization with no `healthChecks` on a prod app | P9 Failure is Observable and Surfaced |
| No notification provider wired for sync failures | P9 Failure is Observable and Surfaced |
| Raw Secret in the GitOps repo | P10, `domains/security/secrets.md` |
| Manual `kubectl edit` on a Flux-managed resource | P8 Reconcile, Don't Mutate by Hand |
| `interval: 24h` on a prod Kustomization (drift window too wide) | P4 Continuous Reconciliation |
+140
View File
@@ -0,0 +1,140 @@
# Operators — Derived Rules
> Derives from `domains/gitops-operators/first-principles.md`.
> Applies P6 (Operators Encode Domain Knowledge) primarily, with
> P1, P4, P8, P9, P10. Cross-links `domains/kubernetes/workloads.md`
> and `domains/kubernetes/rbac.md` for the underlying controller
> model, and `domains/infrastructure-as-code/modules.md` for the
> module-vs-operator boundary.
## What an Operator Is (P6 Operators Encode Domain Knowledge)
- An Operator is a Kubernetes controller that encodes human
operational knowledge as CRDs plus a control loop. The operator
reconciles a domain-specific resource (a database, a message
queue, a certificate, a ML model) to a desired state.
- The operator is the deepest expression of
`domains/kubernetes/P1 Declarative Desired State`: the domain
knowledge itself is the desired state. A runbook that lives only
in a wiki is operational knowledge that has not been encoded —
the operator is the encoding (P6).
- An operator runs inside the cluster, observes its CRDs, and acts.
It is a pull-based reconciler by construction — see
`domains/gitops-operators/first-principles.md` P3.
## CRDs and Controllers (P2 Declarative Over Imperative, P4 Continuous Reconciliation)
- A CustomResourceDefinition (CRD) defines the schema of the
domain resource. The controller watches instances of that CRD
and reconciles current → desired (P4).
- The CRD is the public contract of the operator. Version it
(`v1alpha1``v1beta1``v1`) and preserve backward
compatibility — see `domains/api/versioning.md` for the general
API-evolution principles. A CRD is an API surface, not an
internal type.
```yaml
apiVersion: postgres.example.com/v1
kind: PostgresCluster
metadata:
name: payments-db
namespace: payments
spec:
replicas: 3
version: "16"
storage:
size: 100Gi
storageClass: fast-ssd
backup:
schedule: "0 2 * * *"
retention: 7d
```
- The controller reconciles this spec: creates StatefulSets, PVCs,
Services, backup CronJobs. The user declares intent; the operator
makes it so (P2, P6).
## The Control Loop (P4 Continuous Reconciliation, P8)
- The loop watches CRD instances, compares current vs desired, and
acts to converge. Drift (a hand-deleted pod, a failed backup) is
detected and corrected each loop (P4).
- An operator-managed resource should not be hand-edited (P8). The
operator owns the subordinate resources (StatefulSets, PVCs); a
manual `kubectl edit` on a subordinate is drift the operator will
overwrite.
## Operator SDK and OLM (P6 Composability, C6)
- The Operator SDK scaffolds a controller from a CRD (Go, Ansible,
Helm). Use it to avoid re-implementing the controller boilerplate.
- Operator Lifecycle Manager (OLM) installs, updates, and manages
operators as first-class cluster components. OLM is the package
manager for operators — the operator analogue of
`domains/kubernetes/helm.md` for workloads.
- An operator published via OLM is a versioned, catalog-tracked
artifact. Pin the operator version; do not float `latest` (P5
applies to operators as much as to manifests).
## When to Write an Operator vs a Helm Chart (P6, C6 Composability)
| Axis | Helm chart | Operator |
|------|-----------|----------|
| Day-2 operations | None — chart installs, you operate | Encoded — operator reconciles lifecycle (backup, resize, failover, upgrade) |
| State | Static manifests | Live control loop watching CRDs |
| Day-1 install | Strong fit — package and install | Overkill if install is all you need |
| Day-2 reconcile | None — drift is manual | Continuous — drift corrected each loop |
| Domain knowledge | Lives in runbooks + on-call | Lives in the controller code |
| Best for | Off-the-shelf apps, stateless services, one-shot deploys | Stateful apps, complex lifecycles, day-2 automation (backup, scale, failover, version upgrades) |
| Watch out for | Templating complexity, no day-2 reconcile | Controller complexity, multi-team maintenance burden, scope creep |
- Write an operator when the day-2 operations (backup, failover,
resize, version upgrade) are non-trivial and repeated. Write a
Helm chart when install is all you need and day-2 is run by a
human or a separate tool.
- Do not write an operator to wrap a Helm chart and call it day-2
automation — that is a Helm chart with extra steps. See
`domains/infrastructure-as-code/modules.md` for the
module-vs-copy boundary (the operator-vs-chart boundary is its
analogue).
## Scope and Responsibility Boundaries (P10 Least Privilege, C6)
- An operator owns one domain. An operator that manages databases
and message queues and certificates is doing three jobs — split
it. Scope creep is the most common operator failure mode (P6
violation: the encoded knowledge is no longer coherent).
- The operator's ServiceAccount must be scoped to the resources it
manages (P10). A database operator that needs `cluster-admin` to
create a StatefulSet has the wrong RBAC — see
`domains/kubernetes/rbac.md` and `domains/security/authorization.md`.
- One operator per CRD family; one ServiceAccount per operator; one
namespace per operator (or a shared `operators` namespace with
strict RoleBindings). Default namespace is for nothing in
production.
## Failure and Observability (P9 Failure is Observable and Surfaced)
- An operator must surface its reconcile status on the CRD
(`status.conditions`, `status.observedGeneration`). A CRD with no
status is an operator that fails silently (P9).
- Wire operator events to notifications and metrics. A failed
backup, a stuck failover, a version-upgrade stall must emit a
signal — see `domains/observability/metrics.md`.
- An operator that reconciles but does not report health is a
black box. The GitOps controller (ArgoCD/Flux) will read it as
`Progressing` forever — write the health check (see `argocd.md`
"Health and Status").
## What Violates Operator Discipline
| Violation | Principle |
|-----------|-----------|
| Operator that manages databases + queues + certs | P6 Operators Encode Domain Knowledge (scope creep) |
| Operator ServiceAccount bound to `cluster-admin` | P10 Least Privilege Reconciliation |
| CRD with no `status.conditions` | P9 Failure is Observable and Surfaced |
| Operator with no health check wired to GitOps | P9, `argocd.md` Health and Status |
| Unversioned CRD (`v1` shipped without alpha/beta) | P5, `domains/api/versioning.md` |
| Manual `kubectl edit` on an operator-managed subordinate | P8 Reconcile, Don't Mutate by Hand |
| Operator that wraps a Helm chart and adds no day-2 logic | P6 (no knowledge encoded) |
| Operator runbook that exists only in a wiki | P6 Operators Encode Domain Knowledge |
@@ -0,0 +1,177 @@
# Progressive Delivery — Derived Rules
> Derives from `domains/gitops-operators/first-principles.md`.
> Applies P7 (Progressive Delivery is Reversible by Construction)
> primarily, with P4, P9. Cross-links `domains/devops/first-principles.md`
> P4 Rollback First and P5 Progressive Delivery, and
> `domains/observability/metrics.md` for the analysis signals.
## What Progressive Delivery Is (P7 Reversible by Construction)
- Progressive delivery shifts traffic in stages (canary, blue-green)
gated by analysis (metrics, counters, error rates). Each stage is
metric-checked; a failed gate aborts the rollout and reverts to
the prior stable version. Promotion without a rollback path is a
violation (P7).
- Progressive delivery is the GitOps extension of
`domains/devops/P5 Progressive Delivery` and
`domains/kubernetes/P10 Roll Forward, Roll Back`. The k8s rolling
update is the floor; progressive delivery adds metric-gated
promotion and one-command abort.
- Two sister projects dominate: **Argo Rollouts** (Argo ecosystem)
and **Flagger** (Flux ecosystem). Both implement the same pattern
— a Rollout CRD replaces a Deployment, an analysis drives the
gates, an abort reverts traffic.
## The Rollout CRD (P2 Declarative Over Imperative, P7)
- A Rollout (Argo Rollouts) or Canary/Flag (Flagger) is a CRD that
replaces the Deployment as the reconciled resource. It declares
the strategy (canary, blue-green), the traffic split, and the
analysis gates. The controller reconciles traffic and pods to
match.
```yaml
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: payments-api
namespace: payments
spec:
replicas: 10
selector:
matchLabels:
app: payments-api
template:
metadata:
labels:
app: payments-api
spec:
containers:
- name: api
image: registry.example.com/payments-api:1.2.3
strategy:
canary:
trafficRouting:
istio:
virtualService:
name: payments-vs
routes: [primary]
steps:
- setWeight: 5
- pause: { duration: 2m }
- analysis:
templates:
- templateName: success-rate
- setWeight: 25
- pause: { duration: 5m }
- analysis:
templates:
- templateName: success-rate
- setWeight: 50
- pause: { duration: 5m }
- setWeight: 100
```
- Each `setWeight` shifts traffic; each `pause` holds for
observation; each `analysis` runs a metric gate. A failed
analysis aborts the rollout and reverts traffic to the stable
ReplicaSet (P7).
## Canary vs Blue-Green (P7, C3 Simplicity)
| Strategy | Mechanism | Cost | Best for |
|----------|-----------|------|----------|
| Canary | Shift a small % of traffic to the new version; increase on gate success | Low (few new pods) | Most production rollouts; metric-gated, gradual |
| Blue-Green | Run two full environments; switch traffic all-at-once | High (2× capacity) | Schema-breaking changes, instant rollback, low-frequency deploys |
- Canary is the default — it is reversible by construction (P7)
and economical (C8). Blue-green is for changes that cannot be
partial (a breaking schema migration, a full cutover).
- A canary with no analysis gate is a slow blue-green — it is not
progressive delivery. The gate is what makes it progressive (P7).
## Analysis Templates (P9 Failure is Observable and Surfaced, P7)
- An AnalysisTemplate declares the metric query, the success
threshold, and the count of samples. The rollout controller runs
the analysis at each gate; a failed analysis aborts the rollout.
```yaml
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: success-rate
namespace: payments
spec:
metrics:
- name: success-rate
interval: 1m
successCondition: result[0] >= 0.99
failureLimit: 2
provider:
prometheus:
address: http://prometheus.observability:9090
query: |
sum(rate(http_requests_total{job="payments-api",code!~"5.."}[2m]))
/
sum(rate(http_requests_total{job="payments-api"}[2m]))
```
- `successCondition` is the gate; `failureLimit` is the tolerance
for transient blips. A single failed sample aborts immediately if
`failureLimit: 0`; tolerate noise with `failureLimit: 2`.
- The metric is the abort signal — see `domains/observability/metrics.md`
for the SLI/SLO discipline that makes the gate meaningful. A gate
on an undefined SLO is a gate on noise.
## Argo Rollouts vs Flagger (P6 Composability, P7)
| Axis | Argo Rollouts | Flagger |
|------|---------------|---------|
| Ecosystem | Argo (ArgoCD sister project) | Flux (Flux sister project) |
| CRD | `Rollout` (replaces `Deployment`) | `Canary` / `Flag` (wraps a `Deployment`) |
| Traffic providers | Istio, NGINX, ALB, SMI, Traefik, Ambassador | Istio, NGINX, Linkerd, SMI, App Mesh, Gloo, Contour |
| Analysis sources | Prometheus, Datadog, Wavefront, NewRelic, CloudWatch, Graphite, Kayenta | Prometheus, Datadog, CloudWatch, Stackdriver, Elasticsearch, Graphite |
| Integration | Tight with ArgoCD (UI shows rollout) | Tight with Flux (events via notification-controller) |
| Learning curve | Rollout CRD replaces Deployment (migration cost) | Wraps existing Deployment (lower migration cost) |
| Best for | ArgoCD shops wanting rollout in the Argo UI | Flux shops wanting progressive delivery with minimal migration |
- Both implement the same pattern. The choice follows your GitOps
controller — Argo Rollouts with ArgoCD, Flagger with Flux. Mixing
is possible but not idiomatic.
## Abort and Rollback (P7 Reversible by Construction, P5)
- An abort reverts traffic to the stable ReplicaSet immediately. A
rollout without a tested abort is a prototype (P7).
- The abort must be one-command (or one-gate-failure). A
progressive delivery that requires manual rollback steps has
lost the "reversible by construction" property — it is a deploy
with extra steps.
- Test the abort path in staging. An abort that has never been
exercised will fail when you need it most — see
`domains/devops/first-principles.md` P4 Rollback First.
## Observability (P9 Failure is Observable and Surfaced)
- Progressive delivery is only as good as its metrics. A rollout
gated on a metric that is not tracked is ungated — the gate is
theater (P9).
- Wire rollout status (phase, weight, analysis result) to
notifications and dashboards. A stalled rollout with no signal is
silent drift (P9). See `domains/observability/metrics.md`.
- Cross-link `domains/kubernetes/workloads.md` for the underlying
Deployment/ReplicaSet model that progressive delivery replaces.
## What Violates Progressive Delivery Discipline
| Violation | Principle |
|-----------|-----------|
| Canary with no analysis gate | P7 Progressive Delivery is Reversible by Construction |
| Rollout with no tested abort path | P7, `domains/devops/P4 Rollback First` |
| Analysis gate on an undefined SLO | P9 Failure is Observable and Surfaced |
| Blue-green with no 2× capacity budget | C8 Economy (blue-green is a cost decision) |
| Rollout stalled with no notification | P9 Failure is Observable and Surfaced |
| Manual `kubectl` traffic shift on a Rollout-managed service | P8 Reconcile, Don't Mutate by Hand |
| `failureLimit: 0` on a noisy metric (constant false aborts) | P4 Continuous Reconciliation (gate noise tolerance) |