docs(P02): complete kubernetes phase
---ci--- project: atelier phase: 2 milestone: v0.2 status: complete requirements: covered: [ATELIER-41, ATELIER-42, ATELIER-43, ATELIER-44, ATELIER-45, ATELIER-46, ATELIER-47] partial: [] ---/ci---
This commit is contained in:
@@ -1,12 +1,10 @@
|
||||
{
|
||||
"phase": 1,
|
||||
"stage": "complete",
|
||||
"phase": 2,
|
||||
"stage": "execute",
|
||||
"milestone": "v0.2",
|
||||
"phase_role": "execution",
|
||||
"project": "atelier",
|
||||
"attempts": 0,
|
||||
"updated_at": "2026-08-05T02:05:00Z",
|
||||
"milestone_complete": false,
|
||||
"phase_tag": "v0.1.1",
|
||||
"release_id": 463
|
||||
"updated_at": "2026-08-05T02:10:00Z",
|
||||
"milestone_complete": false
|
||||
}
|
||||
@@ -0,0 +1,67 @@
|
||||
# Kubernetes — First Principles
|
||||
|
||||
## 1. The Principles
|
||||
|
||||
### P1. Declarative Desired State
|
||||
You declare the desired state; controllers reconcile current →
|
||||
desired. Imperative `kubectl` is for inspection and incident
|
||||
response, not for the steady state. The cluster's job is to make
|
||||
reality match the manifest.
|
||||
|
||||
### P2. Pods are Mortal
|
||||
A pod is born, runs, and dies. Never assume its identity, its IP,
|
||||
or its lifetime. Use controllers (Deployment, StatefulSet,
|
||||
DaemonSet), not bare pods. A bare pod has no recovery, no
|
||||
scaling, no rollback.
|
||||
|
||||
### P3. Labels Select
|
||||
Labels and selectors are the join mechanism of the platform —
|
||||
workloads to services, policies to workloads, workloads to nodes.
|
||||
Label by intent (`app`, `tier`, `env`), not by infrastructure
|
||||
(`node-3`, `ip-10.0.0.5`). Selectors compose; ad-hoc naming does
|
||||
not.
|
||||
|
||||
### P4. Requests and Limits are Contracts
|
||||
Resource requests drive scheduling; limits drive quality of
|
||||
service. A workload with no requests is `BestEffort` — first
|
||||
evicted under pressure. A workload with no limits is unbounded.
|
||||
Specifying requests is not optional in production.
|
||||
|
||||
### P5. Probes Drive Health
|
||||
Liveness, readiness, and startup probes are how the platform
|
||||
sees your workload. Without a readiness probe, traffic routes to
|
||||
a pod that is not ready. Without a liveness probe, a wedged
|
||||
container runs forever. The platform cannot heal what it cannot
|
||||
see.
|
||||
|
||||
### P6. Namespaces Bound Blast Radius
|
||||
Namespaces are the unit of quota, RBAC, network policy, and
|
||||
cleanup. A namespace is the boundary of "this thing and all its
|
||||
parts." Default namespace is for nothing in production; every
|
||||
workload gets a named namespace sized to its blast radius.
|
||||
|
||||
### P7. RBAC by Intent, Not Identity
|
||||
Bind roles to service accounts by the workload's purpose, not to
|
||||
user identities. Least privilege: the role grants the minimum
|
||||
the workload needs. `cluster-admin` is a smell, not a shortcut.
|
||||
Cross-link `domains/security/authorization.md`.
|
||||
|
||||
### P8. Storage is Explicit
|
||||
Storage is ephemeral by default. Persistence requires a
|
||||
PVC, a StorageClass, and a reclaim policy decision. `emptyDir`
|
||||
for state that must survive is a bug. The choice of
|
||||
reclaim policy (`Retain`, `Delete`) is a data-safety decision,
|
||||
not a default.
|
||||
|
||||
### P9. Config and Secrets are Separate
|
||||
ConfigMaps are non-sensitive configuration; Secrets are
|
||||
sensitive configuration. Both are injected at runtime, never
|
||||
baked into the image. A configuration change should not require
|
||||
a rebuild; a secret rotation should not require a redeploy of the
|
||||
image. Cross-link `domains/security/secrets.md`.
|
||||
|
||||
### P10. Roll Forward, Roll Back
|
||||
Every Deployment has a rolling update strategy and a rollout
|
||||
history. A deploy is reversible: `kubectl rollout undo`. A deploy
|
||||
without a tested rollback is a prototype. Canary and blue-green
|
||||
are the k8s expression of `domains/devops/P5 Progressive Delivery`.
|
||||
@@ -0,0 +1,68 @@
|
||||
# Helm — Derived Rules
|
||||
|
||||
> Derives from `domains/kubernetes/first-principles.md`. Applies P1, P3, P6. For the Helm-vs-Kustomize decision, see the decision matrix at the end of this doc and in `kustomize.md`.
|
||||
|
||||
## What Helm Is (P6 Modules Compose)
|
||||
|
||||
- Helm is a package manager for Kubernetes. A chart is a versioned package of templated manifests. `helm install` renders the templates against `values.yaml` and applies the result.
|
||||
- A chart encapsulates a reusable deployment (an application, a database, a full stack). It is the k8s analogue of an IaC module — see `domains/infrastructure-as-code/modules.md`.
|
||||
- Charts live in registries (Helm registry via OCI, or the classic chart repos) and are versioned per SemVer.
|
||||
|
||||
## Chart Structure (P1 Declarative Desired State, C2 Clarity)
|
||||
|
||||
- `Chart.yaml` — metadata (name, version, appVersion, dependencies).
|
||||
- `values.yaml` — default inputs; the chart's public interface.
|
||||
- `templates/` — Go-templated manifests. `templates/_helpers.tpl` holds reusable template partials.
|
||||
- `values.schema.json` — optional schema for values, giving type checking on inputs. Use it for published charts.
|
||||
- A chart should have one logical purpose. A chart that deploys an app and a database and an ingress and an observability stack has too many jobs — split it.
|
||||
|
||||
## Values (P5 Version Everything, C2 Clarity)
|
||||
|
||||
- `values.yaml` holds defaults. Override per release: `helm install --set key=value` or `helm install -f my-values.yaml`.
|
||||
- Pin values files in git per environment. A release is reproducible from the chart version + the values file.
|
||||
- Sensitive values do not belong in `values.yaml`. Inject via Secrets (see `rbac.md` P9 and `domains/security/secrets.md`). Some charts accept `existingSecret` to reference a pre-created Secret.
|
||||
|
||||
## Release Management (P5 Version Everything, P10 Roll Forward Roll Back)
|
||||
|
||||
- A release is a named instantiation of a chart. `helm upgrade` applies a new chart version or new values. `helm rollback` reverts to the previous release revision.
|
||||
- `helm history <release>` lists revisions; `helm rollback <release> <revision>` is the rollback. The rollback must be tested like any deploy (P10).
|
||||
- Pin the chart version: `helm install --version 1.2.3`. Never `--version latest` in production — unversioned charts drift (same anti-pattern as unpinned IaC modules).
|
||||
|
||||
## Templating Discipline (P1 Declarative Desired State, C2 Clarity)
|
||||
|
||||
- Templates render to valid manifests. The chart author's job is that the rendered output is correct k8s, not that the template is clever.
|
||||
- Keep `templates/` readable. Heavy logic belongs in `_helpers.tpl` or in a values structure that the template merely projects.
|
||||
- `helm template` renders to stdout without applying — use it to review what a release will create before installing it.
|
||||
|
||||
## Registries (P5 Version Everything)
|
||||
|
||||
- OCI registries are the modern chart distribution (same registry as container images, charts as OCI artifacts). Classic chart repos are legacy.
|
||||
- Pull from a pinned registry reference: `oci://registry/chart:1.2.3`. The digest + tag is the version.
|
||||
|
||||
## Helm vs Kustomize — Decision Matrix (IDEATE-10)
|
||||
|
||||
| Axis | Helm | Kustomize |
|
||||
|------|------|----------|
|
||||
| Mechanism | Templating (Go templates) | Overlays (base + patches) |
|
||||
| Reuse unit | Chart (versioned package) | Base directory (kustomization.yaml) |
|
||||
| Distribution | Registry (OCI, chart repo) | Git (base dir in a repo) |
|
||||
| Values | `values.yaml` + overrides | `kustomization.yaml` + patches |
|
||||
| Release mgmt | `helm` tracks releases, history, rollback | None native — apply with `kubectl apply -k` |
|
||||
| Learning curve | Template language to learn | YAML patching, no DSL |
|
||||
| Blast radius | One chart, many resources, templated | One base, many overlays, patched |
|
||||
| Best for | Off-the-shelf apps, packaged stacks, multi-env via values | Internal apps, patching upstream manifests, env-specific deltas |
|
||||
| Watch out for | Template complexity, `latest` chart drift, secrets in values | No release tracking, manual rollback, patch sprawl |
|
||||
|
||||
- Use Helm when you distribute a reusable app or consume third-party charts. Use Kustomize when you patch existing manifests or keep env deltas in one repo.
|
||||
- Mixing both is fine and common: Helm for the packaged parts, Kustomize for the last-mile per-env patching. Do not fight the tool that fits the job.
|
||||
|
||||
## What Violates Helm Discipline
|
||||
|
||||
| Violation | Principle |
|
||||
|-----------|-----------|
|
||||
| `helm install --version latest` in prod | P5 Version Everything |
|
||||
| Secrets in `values.yaml` | P9 Config and Secrets are Separate, security |
|
||||
| Chart with 15 subcharts doing unrelated things | P6 Modules Compose (split it) |
|
||||
| No `values.schema.json` on a published chart | C2 Clarity |
|
||||
| `helm upgrade` without reviewing `helm template` output | P1 Declarative Desired State, P4 Plan Before Apply |
|
||||
| Untested `helm rollback` | P10 Roll Forward Roll Back |
|
||||
@@ -0,0 +1,62 @@
|
||||
# Kustomize — Derived Rules
|
||||
|
||||
> Derives from `domains/kubernetes/first-principles.md`. Applies P1, P3, P6. For the Helm-vs-Kustomize decision, see the decision matrix at the end of this doc and in `helm.md`.
|
||||
|
||||
## What Kustomize Is (P1 Declarative Desired State)
|
||||
|
||||
- Kustomize customizes manifests without templating. A base directory holds the canonical manifests; overlays hold the deltas. The result is plain YAML applied with `kubectl apply -k`.
|
||||
- No DSL, no template language, no rendering step hidden from `kubectl`. The patch is a YAML file; the result is inspectable.
|
||||
- Kustomize is built into `kubectl` (`kubectl apply -k`, `kubectl diff -k`). No separate runtime is required to apply.
|
||||
|
||||
## Base and Overlays (P6 Namespaces Bound Blast Radius, C4 Locality)
|
||||
|
||||
- A `kustomization.yaml` in a base directory lists the resources (Deployment, Service, etc.) the application needs. It is the canonical manifest.
|
||||
- An overlay is a directory with its own `kustomization.yaml` that references the base (`resources: - ../../base`) and applies patches or additional resources.
|
||||
- Typical structure: `base/`, `overlays/dev/`, `overlays/staging/`, `overlays/prod/`. The overlay is the environment axis; the base is the shared truth.
|
||||
|
||||
## Patches (P1 Declarative Desired State, C2 Clarity)
|
||||
|
||||
- Strategic merge patches — a YAML document that overrides matching fields. Simple for single-resource changes.
|
||||
- JSON patches (RFC 6902) — precise operations (`add`, `replace`, `remove`) on a path. Use when a strategic merge is ambiguous (e.g., list operations).
|
||||
- `patches` field (modern) takes a list of patch files with targets, replacing the older `patchesStrategicMerge` and `patchesJson6902`. Prefer it.
|
||||
- A patch is a delta. It is reviewed as "what changes from base," which is exactly the diff a reviewer wants to see.
|
||||
|
||||
## Generators and Transformers (P3 Labels Select)
|
||||
|
||||
- `configMapGenerator` and `secretGenerator` create ConfigMaps and Secrets from files or literals, with content hashes in the names. A change to the source file changes the hash, which changes the name, which rolls the workload. This is the kustomize pattern for "config change = redeploy."
|
||||
- `namePrefix`, `nameSuffix`, and `namespace` transformers rewrite names across the base. Use for namespace isolation (P6) or to run the same base multiple times in one cluster without collisions.
|
||||
- `commonLabels` and `commonAnnotations` stamp labels onto everything in the base — the kustomize-native way to enforce the labelling discipline of P3.
|
||||
|
||||
## No Release Tracking (P10 Roll Forward Roll Back)
|
||||
|
||||
- Kustomize has no release object, no history, no built-in rollback. `kubectl apply -k` is a one-shot apply; the previous state is in git, not in a Helm-style release record.
|
||||
- Rollback is `git revert` + `kubectl apply -k`. The git history IS the release history. This is fine — and arguably cleaner — but it means rollback is a git operation, not a `helm rollback` command.
|
||||
- Use a GitOps tool (ArgoCD, Flux) on top of Kustomize for automated reconciliation and rollback tracking. The tool watches the git ref; rollback is a git revert.
|
||||
|
||||
## Helm vs Kustomize — Decision Matrix (IDEATE-10)
|
||||
|
||||
| Axis | Kustomize | Helm |
|
||||
|------|----------|------|
|
||||
| Mechanism | Overlays (base + patches) | Templating (Go templates) |
|
||||
| Reuse unit | Base directory (kustomization.yaml) | Chart (versioned package) |
|
||||
| Distribution | Git (base dir in a repo) | Registry (OCI, chart repo) |
|
||||
| Values | `kustomization.yaml` + patches | `values.yaml` + overrides |
|
||||
| Release mgmt | None native — `kubectl apply -k` | `helm` tracks releases, history, rollback |
|
||||
| Learning curve | YAML patching, no DSL | Template language to learn |
|
||||
| Blast radius | One base, many overlays, patched | One chart, many resources, templated |
|
||||
| Best for | Internal apps, patching upstream manifests, env-specific deltas | Off-the-shelf apps, packaged stacks, multi-env via values |
|
||||
| Watch out for | No release tracking, manual rollback, patch sprawl | Template complexity, `latest` chart drift, secrets in values |
|
||||
|
||||
- Use Kustomize when you patch existing manifests or keep env deltas in one repo. Use Helm when you distribute a reusable app or consume third-party charts.
|
||||
- Mixing both is fine and common: Kustomize for the internal apps, Helm for the packaged parts. The decision is per-workload, not per-cluster.
|
||||
|
||||
## What Violates Kustomize Discipline
|
||||
|
||||
| Violation | Principle |
|
||||
|-----------|-----------|
|
||||
| Duplicated base instead of an overlay | P6 Modules Compose (use an overlay) |
|
||||
| Patch that overrides most of the base | C3 Simplicity (the base is wrong — fix the base) |
|
||||
| No `commonLabels` on a multi-team base | P3 Labels Select |
|
||||
| No git-based rollback strategy | P10 Roll Forward Roll Back |
|
||||
| Hand-edited rendered output instead of `apply -k` | P1 Declarative Desired State |
|
||||
| Patch sprawl (10 overlays each patching 15 fields) | C3 Simplicity (refactor the base) |
|
||||
@@ -0,0 +1,48 @@
|
||||
# Networking — Derived Rules
|
||||
|
||||
> Derives from `domains/kubernetes/first-principles.md`. Covers Service, Ingress, Gateway API, EndpointSlices, NetworkPolicy, and DNS. Applies P1, P3, P6.
|
||||
|
||||
## The Service (P3 Labels Select)
|
||||
|
||||
- A Service routes traffic to pods selected by a label selector. The selector is the join between the network abstraction and the workloads.
|
||||
- Service types: `ClusterIP` (in-cluster only, default), `NodePort` (exposed on every node's IP at a fixed port), `LoadBalancer` (cloud-managed LB points to the Service). Default to `ClusterIP`; expose only what must be exposed.
|
||||
- A Service fronts a Deployment (or other controller), never a bare pod. The controller keeps pods available; the Service routes to whichever are ready (per the readiness probe — see `workloads.md`).
|
||||
|
||||
## EndpointSlices (P3 Labels Select, P5 Probes Drive Health)
|
||||
|
||||
- An EndpointSlice lists the pod IPs currently backing a Service. Only pods passing their readiness probe appear.
|
||||
- The Service routes by EndpointSlice, not by selector directly. A pod with the right labels but a failed readiness probe is not in the Service.
|
||||
|
||||
## Ingress and Gateway API (P6 Namespaces Bound Blast Radius)
|
||||
|
||||
- Ingress routes HTTP/HTTPS traffic from outside the cluster to Services. It is L7 routing by host and path.
|
||||
- Gateway API is the successor to Ingress: more expressive (TCP, UDP, TLS passthrough), role-oriented (GatewayClass → Gateway → Route), and cross-platform. Prefer Gateway API for new L7 needs.
|
||||
- Both Ingress and Gateway API are implemented by a controller (nginx-ingress, Traefik, Istio, Envoy Gateway). Pick one; mixing ingress controllers in a cluster is operational debt.
|
||||
|
||||
## NetworkPolicy (P6 Namespaces Bound Blast Radius, P7 RBAC by Intent)
|
||||
|
||||
- A NetworkPolicy is a firewall rule for pods. Default-deny ingress; allow by namespace and pod selector.
|
||||
- Without a default-deny NetworkPolicy, every pod can reach every other pod. In production, default-deny is the baseline; allows are the exceptions.
|
||||
- NetworkPolicy is enforced by the CNI plugin (Calico, Cilium, etc.). A NetworkPolicy with no supporting CNI is a no-op. Verify the CNI enforces before relying on it.
|
||||
|
||||
## DNS (P3 Labels Select)
|
||||
|
||||
- Every Service gets a DNS record: `<service>.<namespace>.svc.cluster.local`. Pods get `pod-ip-address.<namespace>.pod.cluster.local` (with dots replaced).
|
||||
- Headless Services (`clusterIP: None`) resolve directly to pod IPs — use for StatefulSet peer discovery (`<statefulset>-0.<service>`).
|
||||
- DNS is how workloads find each other without hardcoded IPs. Use the DNS name, not the ClusterIP.
|
||||
|
||||
## Dual-Stack (P4 Locality)
|
||||
|
||||
- IPv4/IPv6 dual-stack is opt-in per cluster. Services can be single-stack or dual-stack per Service.
|
||||
- Decide at cluster creation. Migrating a single-stack cluster to dual-stack is disruptive and rarely worth it.
|
||||
|
||||
## What Violates Networking Discipline
|
||||
|
||||
| Violation | Principle |
|
||||
|-----------|-----------|
|
||||
| `LoadBalancer` on an internal-only Service | P6 Namespaces Bound Blast Radius |
|
||||
| No default-deny NetworkPolicy | P6 Namespaces Bound Blast Radius, P7 RBAC by Intent |
|
||||
| Hardcoded pod IP in config | P3 Labels Select (use DNS) |
|
||||
| Service pointing at a bare pod | P3 Labels Select (point at a controller) |
|
||||
| Multiple ingress controllers in one cluster | C3 Simplicity (operational debt) |
|
||||
| No readiness probe on a Service-backed workload | P5 Probes Drive Health (empty EndpointSlices) |
|
||||
@@ -0,0 +1,45 @@
|
||||
# RBAC and Pod Security — Derived Rules
|
||||
|
||||
> Derives from `domains/kubernetes/first-principles.md`. P7 (RBAC by Intent, Not Identity) lives here. Cross-link `domains/security/authorization.md` for the general authorization principles and `domains/security/secrets.md` for secret handling.
|
||||
|
||||
## RBAC Objects (P7 RBAC by Intent, Not Identity)
|
||||
|
||||
- **Role** — permissions within a namespace (verb on resource). **ClusterRole** — permissions cluster-wide or usable across namespaces.
|
||||
- **RoleBinding** — binds a Role to a subject (ServiceAccount, User, Group) within a namespace. **ClusterRoleBinding** — binds a ClusterRole cluster-wide.
|
||||
- Prefer Role + RoleBinding per namespace over ClusterRole + ClusterRoleBinding. Cluster-level is the broad axe; namespace-level is the scalpel.
|
||||
|
||||
## Bind to Service Accounts, Not Users (P7 RBAC by Intent, Not Identity)
|
||||
|
||||
- A workload authenticates as a ServiceAccount. Bind the Role to the ServiceAccount, scoped to the workload's namespace.
|
||||
- The Role encodes the workload's intent: "this workload reads ConfigMaps in this namespace." Not "this user is an admin."
|
||||
- One ServiceAccount per workload (or workload family). Do not reuse the `default` ServiceAccount for production workloads; it is a shared identity.
|
||||
|
||||
## Least Privilege (P7 RBAC by Intent, C1 Correctness via security)
|
||||
|
||||
- Grant the minimum verbs on the minimum resources. `get, list, watch` on `pods` is fine for a monitoring sidecar; `*` on `*` is not.
|
||||
- `cluster-admin` is a smell. If a workload "needs" `cluster-admin`, the workload is either doing something it should not, or it is a cluster operator that should be reviewed as such.
|
||||
- Audit `ClusterRoleBindings` regularly. They are the broadest grant in the system and the easiest to leave behind.
|
||||
|
||||
## Pod Security Standards and Admission (P7 RBAC by Intent, security)
|
||||
|
||||
- Pod Security Standards (PSS) define three profiles: `privileged` (unrestricted), `baseline` (some restrictions), `restricted` (hardened).
|
||||
- Pod Security Admission (built-in) enforces a PSS profile per namespace via labels: `pod-security.kubernetes.io/enforce: restricted`. It replaces the deprecated PodSecurityPolicy.
|
||||
- Map namespaces to profiles: `restricted` for prod workloads, `baseline` for most, `privileged` only for system add-ons (CNI, CSI, node agents) that need it. A workload in `privileged` is a security event, not a default.
|
||||
|
||||
## Service Accounts and Token Automation (P9 Config and Secrets are Separate)
|
||||
|
||||
- ServiceAccount tokens are auto-mounted into pods unless `automountServiceAccountToken: false`. For workloads that do not call the API, disable auto-mount.
|
||||
- Long-lived ServiceAccount tokens are deprecated. Use projected tokens (bound to the pod, time-limited) via `TokenRequest`.
|
||||
- A workload that does not need API access should not have a token. A workload that needs API access should have a Role scoped to its intent.
|
||||
|
||||
## What Violates RBAC Discipline
|
||||
|
||||
| Violation | Principle |
|
||||
|-----------|-----------|
|
||||
| `cluster-admin` bound to a workload | P7 RBAC by Intent, Not Identity |
|
||||
| Reused `default` ServiceAccount for prod | P7 RBAC by Intent, Not Identity |
|
||||
| `automountServiceAccountToken: true` on a non-API workload | P9 Config and Secrets are Separate |
|
||||
| `privileged` PSS on an application namespace | P7 RBAC by Intent, security |
|
||||
| ClusterRoleBinding where a RoleBinding would suffice | P6 Namespaces Bound Blast Radius, P7 |
|
||||
| Long-lived static token instead of projected | P9 Config and Secrets are Separate |
|
||||
| Leftover ClusterRoleBindings after a workload is removed | P7 RBAC by Intent (audit) |
|
||||
@@ -0,0 +1,59 @@
|
||||
# Storage — Derived Rules
|
||||
|
||||
> Derives from `domains/kubernetes/first-principles.md`. P8 (Storage is Explicit) lives here. Covers Volumes, PV/PVC, StorageClass, CSI, snapshots, and reclaim policies. Cross-link `domains/data/` for the data-model angle.
|
||||
|
||||
## Ephemeral by Default (P8 Storage is Explicit)
|
||||
|
||||
- A container's filesystem is ephemeral. When the pod dies, the filesystem dies with it. This is the design, not a flaw.
|
||||
- `emptyDir` is an ephemeral volume scoped to the pod's lifetime (survives container restarts within the pod, dies with the pod). It is scratch space, never durable storage.
|
||||
- Any data that must survive a pod restart requires a PersistentVolumeClaim (PVC). The choice of "must survive" is the data-safety decision at the heart of P8.
|
||||
|
||||
## PersistentVolume and PersistentVolumeClaim (P8 Storage is Explicit)
|
||||
|
||||
- A PersistentVolume (PV) is a piece of storage in the cluster. A PersistentVolumeClaim (PVC) is a request for that storage by a workload.
|
||||
- The PV is the resource; the PVC is the consumer. A workload mounts the PVC, not the PV directly.
|
||||
- For StatefulSets, use `volumeClaimTemplates` so each replica gets its own PVC with a stable name (`data-<statefulset>-0`). Do not share one PVC across replicas of a stateful workload.
|
||||
|
||||
## StorageClass and Dynamic Provisioning (P8 Storage is Explicit, P5 Version Everything)
|
||||
|
||||
- A StorageClass describes the "flavour" of storage (e.g., `fast-ssd`, `cold-hdd`, `encrypted`). A PVC names a StorageClass or gets the cluster default.
|
||||
- Dynamic provisioning creates the PV on demand when the PVC is created, via the CSI driver. Manual PV creation is for specific cases (a pre-existing disk, a static NFS export).
|
||||
- Mark a default StorageClass only if the default is safe for all workloads. A fast-but-expensive default can cause cost surprises; a slow default can cause performance surprises.
|
||||
|
||||
## CSI (P5 Version Everything)
|
||||
|
||||
- The Container Storage Interface (CSI) is the standard driver interface. Each storage backend ships a CSI driver. Pin the CSI driver version in the cluster; treat it as infrastructure.
|
||||
- CSI enables features beyond mount/unmount: snapshots, cloning, volume expansion, and topology-aware provisioning. Not all drivers implement all features; verify before relying.
|
||||
|
||||
## Volume Snapshots (P5 Reversibility, P8 Storage is Explicit)
|
||||
|
||||
- A VolumeSnapshot is a point-in-time copy of a PVC, taken by the CSI driver. Restore creates a new PVC from the snapshot.
|
||||
- Snapshots are not backups. They are local to the storage backend and may share blocks with the source. An off-cluster backup is still required for disaster recovery.
|
||||
- Snapshot scheduling is a workload concern (use a CronJob or a tool like Velero), not a k8s-native feature.
|
||||
|
||||
## Reclaim Policies (P8 Storage is Explicit, P5 Reversibility)
|
||||
|
||||
| Policy | On PVC delete | When |
|
||||
|--------|---------------|------|
|
||||
| `Retain` | PV and its data persist; PV must be manually reclaimed | Production, data-safety default |
|
||||
| `Delete` | PV and the underlying storage are deleted | Ephemeral, dev, scratch |
|
||||
| `Recycle` (deprecated) | PV scrubbed and made available again | Do not use — use dynamic provisioning |
|
||||
|
||||
- The reclaim policy is a data-safety decision. `Delete` on a production PVC is a footgun: deleting the PVC destroys the data. Default to `Retain` for prod, `Delete` for dev.
|
||||
- For StatefulSet PVCs, the reclaim policy on the StorageClass governs what happens when the PVC is deleted (which happens when the StatefulSet is scaled down or deleted, depending on the policy).
|
||||
|
||||
## Ephemeral Volumes (P8 Storage is Explicit)
|
||||
|
||||
- `configMap`, `secret`, `downwardAPI` volumes are read-only (by default) projections injected at pod start. They are configuration, not storage.
|
||||
- `emptyDir` with `medium: Memory` is a tmpfs — fast, ephemeral, memory-charged. Use for scratch that must be fast and never persist.
|
||||
|
||||
## What Violates Storage Discipline
|
||||
|
||||
| Violation | Principle |
|
||||
|-----------|-----------|
|
||||
| `emptyDir` for data that must survive pod restart | P8 Storage is Explicit |
|
||||
| Shared PVC across StatefulSet replicas | P8 Storage is Explicit (use `volumeClaimTemplates`) |
|
||||
| `Delete` reclaim policy on production storage | P8 Storage is Explicit, P5 Reversibility |
|
||||
| Snapshot treated as a backup | P5 Reversibility (snapshots are local, not DR) |
|
||||
| No default StorageClass decision (accidental default) | P8 Storage is Explicit |
|
||||
| Manual PV creation when dynamic provisioning exists | C3 Simplicity |
|
||||
@@ -0,0 +1,53 @@
|
||||
# Workloads — Derived Rules
|
||||
|
||||
> Derives from `domains/kubernetes/first-principles.md`. Covers Pod, ReplicaSet, Deployment, StatefulSet, DaemonSet, Job, and CronJob. Applies P1–P10.
|
||||
|
||||
## The Pod (P2 Pods are Mortal)
|
||||
|
||||
- A pod is the smallest deployable unit: one or more containers sharing network and storage namespaces.
|
||||
- Never deploy a bare pod. A bare pod has no controller to restart, scale, or replace it. Use a controller.
|
||||
- Pods are replaceable by design. Do not store state in a pod's filesystem (`emptyDir` is scratch, not storage — see `storage.md`).
|
||||
|
||||
## Controllers (P1 Declarative Desired State)
|
||||
|
||||
| Controller | When | Identity | Ordering |
|
||||
|------------|------|----------|----------|
|
||||
| Deployment | Stateless workloads | None (pods interchangeable) | No ordering |
|
||||
| StatefulSet | Stateful workloads (databases, queues) | Stable name (`pod-0`, `pod-1`) + stable PVC | Ordered, sequential |
|
||||
| DaemonSet | One pod per node (logging, monitoring, CNI) | Per-node | — |
|
||||
| Job | Run to completion (batch) | — | — |
|
||||
| CronJob | Scheduled batch | — | — |
|
||||
|
||||
- A Deployment manages a ReplicaSet, which manages pods. You interact with the Deployment; the ReplicaSet is an implementation detail except during rollouts.
|
||||
- StatefulSet gives stable network identity and stable persistent storage per replica. Use it when the workload needs a stable name or per-replica data (databases, distributed systems). Do not use StatefulSet for stateless workloads — the ordering is overhead.
|
||||
|
||||
## Probes (P5 Probes Drive Health)
|
||||
|
||||
- **Readiness probe** — is the pod ready to serve traffic? Failing readiness removes the pod from the Service's endpoints but does not restart it. Use for "warm-up" and transient unavailability.
|
||||
- **Liveness probe** — is the pod alive? Failing liveness restarts the container. Use for "wedged but running." Do not use liveness to check dependencies (it will cascade-restart on a dependency blip).
|
||||
- **Startup probe** — has the pod finished starting? Disables liveness/readiness until it succeeds. Use for slow-starting workloads so liveness does not kill them before they are ready.
|
||||
- Probes must check the workload's own health, not the health of its dependencies. A readiness probe that fails on a downstream outage causes the Service to drain all pods simultaneously.
|
||||
|
||||
## Lifecycle and Disruption (P2 Pods are Mortal, P10 Roll Forward Roll Back)
|
||||
|
||||
- `kubectl rollout status` watches a Deployment's rollout to completion. `kubectl rollout undo` reverts to the previous ReplicaSet.
|
||||
- PodDisruptionBudgets (PDBs) protect voluntary disruptions (node drain, cluster autoscaler). An involuntary disruption (node failure) ignores the PDB. Set a PDB on every workload that must keep a minimum available.
|
||||
- Rolling update strategy: `maxUnavailable` and `maxSurge` control the speed of rollout. Slow rollouts (low `maxSurge`) are safer; fast rollouts (high `maxUnavailable`) risk availability.
|
||||
|
||||
## Resource Contracts (P4 Requests and Limits are Contracts)
|
||||
|
||||
- Every container in production has a CPU request, a memory request, and a memory limit. CPU limits are optional but recommended to bound noisy neighbours.
|
||||
- QoS classes: `Guaranteed` (requests == limits), `Burstable` (requests < limits), `BestEffort` (no requests). `BestEffort` is first evicted under node pressure — never for prod.
|
||||
- A workload without requests is an unbounded gamble on the scheduler. Set them.
|
||||
|
||||
## What Violates Workload Discipline
|
||||
|
||||
| Violation | Principle |
|
||||
|-----------|-----------|
|
||||
| Bare pod (no controller) | P2 Pods are Mortal |
|
||||
| StatefulSet for a stateless workload | P1 Declarative Desired State (overhead) |
|
||||
| No probes | P5 Probes Drive Health |
|
||||
| Liveness probe checks a dependency | P5 Probes Drive Health |
|
||||
| No PDB on a critical workload | P10 Roll Forward Roll Back |
|
||||
| No resource requests in prod | P4 Requests and Limits are Contracts |
|
||||
| `emptyDir` for data that must persist | P8 Storage is Explicit |
|
||||
Reference in New Issue
Block a user