11 Commits

Author SHA1 Message Date
Jon Chery 4c8700e5a4 docs(P5): complete examples + cross-links phase
---ci---
project: atelier
phase: 5
milestone: v0.3
status: complete
requirements:
  covered: [ATELIER-86, ATELIER-87, ATELIER-88]
  partial: []
---/ci---
2026-08-05 03:40:37 +00:00
Jon Chery 648fcb5d85 docs(ship): P4 complete — checkpoint + roadmap status
---ci---
project: atelier
phase: 4
milestone: v0.3
status: complete
phase_tag: v0.2.4
release_id: 472
---/ci---
2026-08-05 03:35:39 +00:00
Jon Chery 61043dea1b docs(P4): complete matrix + review + manifest integration phase
---ci---
project: atelier
phase: 4
milestone: v0.3
status: complete
requirements:
  covered: [ATELIER-80, ATELIER-81, ATELIER-82, ATELIER-83, ATELIER-84, ATELIER-85]
  partial: []
---/ci---
2026-08-05 03:35:26 +00:00
Jon Chery 44a7049860 docs(ship): P3 complete — checkpoint + roadmap status
---ci---
project: atelier
phase: 3
milestone: v0.3
status: complete
phase_tag: v0.2.3
release_id: 471
---/ci---
2026-08-05 03:30:29 +00:00
Jon Chery 7dfd3cdc6c docs(P3): complete i18n + compliance domains phase
---ci---
project: atelier
phase: 3
milestone: v0.3
status: complete
requirements:
  covered: [ATELIER-70, ATELIER-71, ATELIER-72, ATELIER-73, ATELIER-74, ATELIER-75, ATELIER-76, ATELIER-77, ATELIER-78, ATELIER-79]
  partial: []
---/ci---
2026-08-05 03:30:14 +00:00
Jon Chery c026c8930b docs(ship): P2 complete — checkpoint + roadmap status
---ci---
project: atelier
phase: 2
milestone: v0.3
status: complete
phase_tag: v0.2.2
release_id: 470
---/ci---
2026-08-05 03:24:26 +00:00
Jon Chery 32edb19c96 docs(P2): complete ai-ml domain phase
---ci---
project: atelier
phase: 2
milestone: v0.3
status: complete
requirements:
  covered: [ATELIER-65, ATELIER-66, ATELIER-67, ATELIER-68, ATELIER-69]
  partial: []
---/ci---
2026-08-05 03:24:13 +00:00
Jon Chery ab1289a9d9 docs(ship): P1 complete — checkpoint + roadmap status
---ci---
project: atelier
phase: 1
milestone: v0.3
status: complete
phase_tag: v0.2.1
release_id: 469
---/ci---
2026-08-05 03:21:04 +00:00
Jon Chery 47674969a1 docs(P1): complete gitops-operators domain phase
---ci---
project: atelier
phase: 1
milestone: v0.3
status: complete
requirements:
  covered: [ATELIER-60, ATELIER-61, ATELIER-62, ATELIER-63, ATELIER-64]
  partial: []
---/ci---
2026-08-05 03:20:48 +00:00
Jon Chery ce36db0579 docs(ship): P0 complete — checkpoint + roadmap status
---ci---
project: atelier
phase: 0
milestone: v0.3
status: complete
phase_tag: v0.2.0
release_id: 468
---/ci---
2026-08-05 03:17:10 +00:00
Jon Chery b7da50f56e docs(P00): complete pre-execution phase — v0.3
---ci---
project: atelier
phase: 0
milestone: v0.3
status: complete
requirements:
  covered: [ATELIER-60..91 governance: spec, clarify, research, ideate, plan, grill]
  partial: []
---/ci---
2026-08-05 03:16:41 +00:00
40 changed files with 4883 additions and 40 deletions
+6 -7
View File
@@ -1,12 +1,11 @@
{ {
"phase": 5, "phase": 4,
"stage": "complete", "stage": "complete",
"milestone": "v0.2", "milestone": "v0.3",
"phase_role": "final", "phase_role": "execution",
"project": "atelier", "project": "atelier",
"attempts": 0, "attempts": 0,
"updated_at": "2026-08-05T02:45:00Z", "updated_at": "2026-08-05T03:55:00Z",
"milestone_complete": true, "phase_tag": "v0.2.4",
"phase_tag": "v0.1.5", "release_id": 472
"release_id": 467
} }
+27 -1
View File
@@ -12,7 +12,11 @@ atelier/
├── domains/ # Domain-specific application of core ├── domains/ # Domain-specific application of core
│ ├── ... (v0.1: 11 domains) │ ├── ... (v0.1: 11 domains)
│ ├── infrastructure-as-code/ # v0.2: IaC tooling (terraform, opentofu, state, modules) │ ├── infrastructure-as-code/ # v0.2: IaC tooling (terraform, opentofu, state, modules)
── kubernetes/ # v0.2: k8s platform (workloads, networking, storage, rbac, helm, kustomize) ── kubernetes/ # v0.2: k8s platform (workloads, networking, storage, rbac, helm, kustomize)
│ ├── gitops-operators/ # v0.3: GitOps + Operators (argocd, flux, operators, progressive-delivery)
│ ├── ai-ml/ # v0.3: ML engineering (data-versioning, model-evaluation, serving, monitoring-drift)
│ ├── i18n/ # v0.3: internationalization (locale-resources, formatting, rtl-bidi, testing-i18n)
│ └── compliance/ # v0.3: compliance/audit (audit-logs, data-retention, policy-as-code, evidence)
├── languages/ # Language-specific application of domains ├── languages/ # Language-specific application of domains
├── review/ # Evaluation checklists and anti-patterns ├── review/ # Evaluation checklists and anti-patterns
├── matrix/ # Cross-reference: domain ↔ core ├── matrix/ # Cross-reference: domain ↔ core
@@ -81,3 +85,25 @@ Two new top-level domains extend the tree under the same hierarchy rules:
Both domains follow the v0.1 contract: 10 P-rules each, every rule traced to a core C-rule via the matrix, no orphans. The manifest (`MANIFEST.md`) is extended to keep them authoritative. No runtime code — examples are illustrative markdown with manifests in code fences only. Both domains follow the v0.1 contract: 10 P-rules each, every rule traced to a core C-rule via the matrix, no orphans. The manifest (`MANIFEST.md`) is extended to keep them authoritative. No runtime code — examples are illustrative markdown with manifests in code fences only.
See `.ciagent/atelier/RESEARCH.md` for the full prior-art survey and `.ciagent/atelier/PERSONAS.md` for the persona roster (3 custom active personas + 1 phase-specific platform-engineer; 3 default personas deactivated). See `.ciagent/atelier/RESEARCH.md` for the full prior-art survey and `.ciagent/atelier/PERSONAS.md` for the persona roster (3 custom active personas + 1 phase-specific platform-engineer; 3 default personas deactivated).
## v0.3 Domain Additions
Four new top-level domains extend the tree under the same hierarchy rules. All four follow the v0.1/v0.2 contract: 10 P-rules each, every rule traced to a core C-rule via the matrix, no orphans, docs-only markdown with illustrative code fences (no runtime/deployable artifacts). Total matrix grows from 130 → 170 domain principles across 13 → 17 domains.
- **`gitops-operators/`** — platform-automation domain. First principles govern the declarative-source-of-truth reconciliation loop shared by ArgoCD, Flux, Kubernetes Operators, and Progressive Delivery tooling (Argo Rollouts, Flagger). Depends on `core/`. Cross-links to `kubernetes/` (workloads, rbac, helm, kustomize — the platform GitOps reconciles onto), `infrastructure-as-code/` (declarative intent, state-as-truth — the shared model), `devops/` (P1 Reproducibility, P4 Rollback First, P5 Progressive Delivery, P6 Configuration as Code), `security/` (secrets, supply-chain — GitOps credentials, signed manifests), `observability/` (reconciliation metrics, drift visibility). Derived docs: `argocd.md`, `flux.md`, `operators.md`, `progressive-delivery.md`.
- **`ai-ml/`** — ML engineering domain (engineering discipline, NOT algorithm design per D-023). First principles govern data versioning, model evaluation, serving, and monitoring/drift. Depends on `core/`. Cross-links to `data/` (schema-design, migrations, indexing — data lineage and versioning share the migration/reversibility model), `observability/` (metrics, tracing — model serving metrics, drift signals), `devops/` (P1 Reproducibility — training/serving reproducibility, P7 Immutability — model images), `security/` (input-validation — inference input validation, secrets — model/serving credentials), `performance/` (backend — serving latency). Derived docs: `data-versioning.md`, `model-evaluation.md`, `serving.md`, `monitoring-drift.md`.
- **`i18n/`** — internationalization domain. First principles govern locale resources, formatting, RTL/bidi layout, and testing. Depends on `core/`. Cross-links to `uiux/` (components, accessibility, copywriting — locale-aware UI is the consumer), `testing/` (fixtures, pyramid — i18n testing parallels), `api/` (error-responses — localized API errors), `data/` (schema-design — locale data shapes). Derived docs: `locale-resources.md`, `formatting.md`, `rtl-bidi.md`, `testing-i18n.md`.
- **`compliance/`** — compliance/audit domain (framework-agnostic, NOT regulation-specific per D-024). First principles govern audit logs, data retention, policy-as-code, and evidence collection. Depends on `core/`. Cross-links to `security/` (authorization — who did what, secrets — audit log integrity, supply-chain — signed policy), `observability/` (logging, metrics — audit logs are a structured-logging concern, tracing — evidence from distributed traces), `data/` (schema-design, migrations — retention schema), `infrastructure-as-code/` (policy-as-code parallels IaC declarative intent), `kubernetes/` (rbac — audit subject identity). Derived docs: `audit-logs.md`, `data-retention.md`, `policy-as-code.md`, `evidence.md`.
All four domains depend on `core/` only for authority; cross-links to existing domains are one-directional (per v0.2 D-026 convention extended to v0.3 — minimize churn to existing content). The manifest (`MANIFEST.md`) is extended in P4 to list all new documents. Examples (P5) are illustrative markdown with fenced code only — no `.yaml`, `.json`, `.po`, model artifacts, or deployable manifests as standalone files.
## v0.3 Ideation Architectural Notes
From the v0.3 ideation stage (IDEATE-17..30), the following architectural refinements are baked into the execute-phase plan:
- **Manifest scope expansion (IDEATE-17 → ATELIER-91):** the v0.2 audit escalation (ESC-002 note) flagged that `examples/` is not listed in `MANIFEST.md`. P4 adds an `examples/` directory listing to the manifest, closing the pre-existing drift. The manifest remains authoritative; unlisted directories are not part of the framework by definition.
- **Matrix coverage summary invariants (IDEATE-18 → ATELIER-80):** the matrix coverage summary must reflect post-v0.3 totals (17 domains, 170 P-rules) — both the summary block and the per-domain section count.
- **Core Principle Coverage table (IDEATE-19 → ATELIER-81):** `matrix/domain-coverage.md` contains two tables — the per-domain row schema table (covered by v0.2 IDEATE-03) AND the "Core Principle Coverage" table mapping C1C8 → domains. Both must be extended for the 4 new domains; the C-rule counts shift (e.g., C4 Locality adds i18n + gitops; C5 Reversibility adds ai-ml + compliance + gitops + i18n).
- **Cross-link type unchanged:** v0.3 introduces no new cross-link type. All cross-links remain one-directional outward from new domains to existing (D-033). No back-link edits to v0.1/v0.2 content.
See `.ciagent/atelier/RESEARCH.md` "v0.3 Research" for the full prior-art survey and principle inventory rationale, and `.ciagent/atelier/PERSONAS.md` for the v0.3 persona roster (5 active: lead-developer, tech-writer, domain-expert + 2 phase-specific platform-engineer, ml-engineer; 3 default personas deactivated).
+175
View File
@@ -0,0 +1,175 @@
# Atelier — Grill (Adversarial Red-Team Review)
> Pre-execution gate for milestone v0.3 (GitOps + Operators + AI/ML + i18n + Compliance).
> Default assumption: the project is unfeasible, over-scoped, and too costly. Not convinced until evidence forces it.
> Mode: full autonomy. Auto-resolve at confidence ≥ 0.60; escalate only < 0.60 that cannot be auto-resolved.
---
## v0.3 Grill — 2026-08-05
**Milestone:** v0.3 — GitOps + Operators + AI/ML + i18n + Compliance
**Phase:** 0 (Pre-Execution, GRILL stage)
**Grill scope:** all 9 axes + meta
**Prior grill runs:** none (v0.1/v0.2 P0 stages did not include a GRILL stage; this is the first)
### Verdict: **PROCEED** (confidence 0.80)
The project is ambitious but deliberately bounded. Scope has been trimmed in three places (D-023 AI/ML engineering-only, D-024 compliance framework-agnostic, D-025 2+2 examples). Two prior milestones (v0.1: 35 reqs, v0.2: 24 reqs) shipped clean with the same docs-only NFR structure. The 4-domain scope is the largest single milestone yet, but per-phase derived-doc counts (P1: 4, P2: 4, P3: 8) are lower than v0.1's P3 (27 derived docs in one phase). Traceability verification is explicit (P4 matrix row-count test: 10 × 17 = 170; P5 cross-link audit). The one residual risk — P3 i18n/compliance content authored without specialist personas — is bounded by RESEARCH.md prior-art depth and P6 domain-expert review, with rework being cheap (markdown edits).
---
### Axis 1 — The Business Case Itself
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| 1.1 | What problem does this solve, and is it still the top priority? | RESEARCH.md: "Atelier's differentiation: traceable principle hierarchy with a join table. Existing frameworks state principles; none provide a matrix mapping every domain rule back to a core rule." v0.1/v0.2 shipped (59 reqs covered). | A first-principles engineering framework with a traceable principle matrix — the join table is the unique value. v0.3 extends the domain catalog (GitOps, AI/ML, i18n, compliance) that v0.2 deferred (IDEATE-15/16). Still the top priority: the 4 domains were explicitly deferred to v0.3, not abandoned. | 0.78 |
| 1.2 | Who is the named executive sponsor, and when did they last decide under pressure? | git log: all commits by Jon Chery (jchery@jccapital.xyz). config.json autonomy: "full". | Single-owner project. The "sponsor" is the owner-operator. The last decision under pressure: v0.2 ESC-001/ESC-002 (git transport auth failure, stale tags) — resolved at full autonomy (AUDIT-P5-resync.md). Not a committee-driven project; no sponsor-stall risk. | 0.80 |
| 1.3 | What happens to the business if the project is cancelled? | PROJECT.md objective; MANIFEST.md current state (13 domains, 130 P-rules post-v0.2). | The framework remains at v0.2 (13 domains). The 4 deferred domains (GitOps, AI/ML, i18n, compliance) stay deferred — a 2x deferral that risks zombie status. Cancellation is worse than proceeding: the deferral was already made once. | 0.82 |
| 1.4 | Is the ROI calculated against a counterfactual? | N/A — docs-only project, no monetary budget. Cost = agent tokens + time. | The counterfactual is "agents and humans have no canonical principle reference for GitOps/AI-ML/i18n/compliance." The ROI is framework coverage. Two prior milestones validated the cost model (docs-only, no infra). | 0.75 |
**Axis 1 confidence: 0.79.** No challenges. Auto-resolve.
### Axis 2 — Scope and Requirements
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| 2.1 | Is the scope expanding, contracting, or stable? | PROJECT.md D-016 (4 domains grouped to limit release overhead), D-023 (AI/ML = engineering discipline, NOT algorithm design), D-024 (compliance framework-agnostic, NOT regulation-specific), D-025 (2+2 examples, not 4+4). | Scope is **expanding** (4 new domains, the largest milestone yet) but **deliberately trimmed** in three places. D-023/D-024/D-025 are scope-discipline decisions, not scope-creep. The trims reduce the surface from a hypothetical 4+4 examples + regulation-specific compliance docs + algorithm-design AI/ML to a bounded 2+2 + framework-agnostic + engineering-only. | 0.80 |
| 2.2 | Who owns the requirements, and are they frozen? | REQUIREMENTS.md ATELIER-60..91 (32 reqs, all pending). Traceability matrix maps phases→reqs. Ideation log (IDEATE-17..30) shows 14 accepted, 0 deferred, 0 rejected. | Requirements are frozen post-ideation. 32 reqs across 6 phases. The ideation stage closed with 0 deferred — everything generated is in v0.3 scope. No moving targets. | 0.85 |
| 2.3 | What is explicitly out of scope? | PROJECT.md lines 49-55: tooling/linters, translation of framework docs, agent integration adapters, per-domain release artifacts, runtime code. D-023: algorithm/model design. D-024: regulation-specific compliance docs. | Out of scope is explicit and enumerated: no runtime code, no tooling, no translation, no regulation-specific docs, no algorithm design. The scope boundary is answerable — not infinite. | 0.88 |
| 2.4 | Are there hidden requirements disclosed late? | RESEARCH.md: ai-ml cross-links to data/ (schema, migrations) but D-023 scopes AI/ML to engineering discipline, not data engineering. PERSONAS.md: data-engineer explicitly inactive ("No database, schema, or migrations in this docs-only project"). | No hidden requirements detected. AI/ML does NOT imply a data-engineer persona — D-023's engineering-discipline scope excludes data engineering (schema/ETL/pipelines). The cross-link to data/ is a reference, not a duplication. | 0.85 |
**Axis 2 confidence: 0.84.** Challenge: is 4-domain scope too ambitious for one milestone? → Auto-resolved (G-001). Challenge: hidden data-engineer persona requirement? → Auto-resolved (G-008).
### Axis 3 — Architecture and Technical Feasibility
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| 3.1 | Has the architecture been validated by builders, not just sellers? | RESEARCH.md: prior-art surveys for all 4 domains (OpenGitOps Principles v1.0.0, Sculley "Hidden Technical Debt", ICU/CLDR/BCP 47, NIST/SOC2/OPA). ARCHITECTURE.md: v0.3 domain additions section with cross-link map. | The architecture (hierarchical doc tree + matrix join table) is validated by 2 prior shipped milestones. The 4 new domains follow the same v0.1/v0.2 contract: 10 P-rules each, trace to ≥1 C-rule, docs-only. RESEARCH.md pre-maps all 40 P-rules to C-rules before execution. | 0.85 |
| 3.2 | What is the integration surface? | ARCHITECTURE.md: cross-links one-directional (D-033). RESEARCH.md: gitops→6 domains, ai-ml→6 domains, i18n→4 domains, compliance→6 domains. D-033: no back-link edits to v0.1/v0.2 content. | The integration surface is cross-links only — one-directional outward from new domains to existing. No back-link edits to v0.1/v0.2 content (D-033). This minimizes churn. Every new derived doc requires ≥1 outbound cross-link (ATELIER-88, verified in P5). | 0.86 |
| 3.3 | Is there an existing system being replaced? What is the data volume? | MANIFEST.md: 13 domains, 130 P-rules post-v0.2. matrix/principles-matrix.md: 212 lines, 13 domain sections. | No system is replaced — the framework is extended. Matrix grows 130→170 P-rules. At 170 rows, the matrix is a large but single readable markdown file. domain-coverage.md provides the C-rule→domains navigation view. The matrix is the arbiter; its size is linear with domains. | 0.84 |
| 3.4 | What technical debt is inherited? | AUDIT-P5-resync.md §6: "examples/ not in MANIFEST — pre-existing, candidate for v0.3." ATELIER-91, D-035: add examples/ listing to MANIFEST in P4. | One piece of inherited debt: examples/ unlisted in MANIFEST (ESC-002 convention note from v0.2 audit). v0.3 closes this explicitly (ATELIER-91, D-035, task 04-02-05). No other drift identified (A-001, confidence 0.85). | 0.85 |
| 3.5 | Compliance has NO C4 (Locality) trace (D-032) — is that an architectural smell? | RESEARCH.md D-032: "compliance is inherently cross-cutting, not local." Compliance P-rules trace to C1,C2,C3,C5,C6,C7,C8 (7 of 8). Alternative considered: "force C4 via audit-log locality" — rejected. | **Not a smell.** C4 (Locality) is about keeping concerns local to their context. Compliance is the opposite — it is inherently system-wide (audit logs span the whole system, retention policy is global, evidence is cross-cutting). Forcing a C4 trace would be a false derivation. D-032's rationale is architecturally sound: the absence reflects the domain's nature, not a gap. | 0.82 |
| 3.6 | Is the docs-only constraint (D-020) defensible at 17 domains? | PROJECT.md constraint: "no runtime code." MANIFEST.md is the authoritative index. ARCHITECTURE.md: reading-order guides consumption. matrix/principles-matrix.md is the join table. | **More defensible at 17 domains, not less.** The docs-only constraint is what makes the framework scalable: no runtime complexity, no integration surface, no deployment, no build step. At 17 domains, the MANIFEST + reading-order + matrix make the tree navigable. The framework's value (traceable hierarchy) scales linearly — more domains = more value, as long as traceability holds. A build step or runtime would violate docs-as-code simplicity and add the exact integration surface the framework avoids. | 0.88 |
**Axis 3 confidence: 0.85.** Challenges: C4 gap (G-006), docs-only at 17 domains (G-007), matrix orphan risk (G-003), manifest drift (G-011). All auto-resolved.
### Axis 4 — People, Skills, and Organization
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| 4.1 | Which 2-3 people, if they left, would the project fail? | PERSONAS.md: 5 active personas (lead-developer, tech-writer, domain-expert + 2 phase-specific: platform-engineer, ml-engineer). D-051: ml-engineer constraints baked into P5 task must-have. | In an AI-agent docs project, "personas" are constraint sets, not human employees. The key-person risk is lower than in runtime projects. D-051 ensures ml-engineer constraints survive into P5 (examples) even though the persona is removed after P2. The constraint is baked into the task must-have, not the persona's continued presence. | 0.80 |
| 4.2 | Are resources allocated at the claimed percentages? | config.json: max_concurrent_agents=5. PLAN.md: P3 Wave 2 splits 8 tasks into 2a/2b, capped at 5 concurrent (D-049). | Allocation is mechanical — the executor schedules ≤5 concurrent per config.json. P3's 8 derived docs run as 5-then-3 (A-003, confidence 0.90). P4 runs exactly 5 concurrent (D-050). No "in name only" allocation — this is agent execution, not human BAU fire-fighting. | 0.85 |
| 4.3 | Is there a product owner with actual authority? | git log: single author (Jon Chery). config.json autonomy: "full". PROJECT.md: D-001..D-053 decision table. | Single owner-operator with full autonomy. No committee. Decisions are transparent (D-001..D-053 with confidence scores). Authority is unambiguous. | 0.85 |
| 4.4 | Is the team building capability they don't have? | PERSONAS.md: P3 (i18n + compliance) authored by tech-writer + domain-expert — NO specialist persona (D-022). RESEARCH.md: i18n covers ICU/CLDR, BCP 47, UAX #9 bidi algorithm; compliance covers OPA/Cedar/Kyverno/Sentinel, Cosign/in-toto. | **This is the most material risk.** P3's i18n (RTL/bidi, UAX #9) and compliance (policy-as-code engine semantics) have genuine specialist depth. The tech-writer persona authored v0.1's security/data/concurrency domains without specialists and shipped clean — but those are more universally known than bidi algorithms and Rego semantics. **Mitigation:** RESEARCH.md provides thorough prior-art surveys with specific source citations (unicode.org, W3C i18n WG, openpolicyagent.org, kyverno.io). The P6 review includes domain-expert validation. Rework, if needed, is bounded (markdown edits, not infrastructure). **Assumption logged (A-006):** i18n RTL/bidi and compliance policy-as-code content correctness depends on RESEARCH.md prior-art quality + P6 review, not on a specialist persona. | 0.72 |
| 4.5 | Are the phase-specific personas (platform-engineer, ml-engineer) a key-person risk? | PERSONAS.md: both are phase-specific, removed post-v0.3. D-027: platform-engineer reused/extended from v0.2. D-028: ml-engineer new. D-051: constraints baked into task must-haves. D-052: both review their content in P6 before removal. | **Not a key-person risk.** Personas are constraint sets, not people. Platform-engineer is a proven reuse from v0.2 (shipped clean). Ml-engineer is new but its constraints are explicitly enumerated and baked into task must-haves (D-051). Both review their content in P6 before removal (D-052). If either "fails," the fallback is tech-writer + domain-expert + RESEARCH.md grounding. | 0.80 |
**Axis 4 confidence: 0.80.** Challenge: P3 specialist-persona competency gap (G-009). Auto-resolved with assumption A-006.
### Axis 5 — Timeline and Estimates
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| 5.1 | Was the deadline set before or after scope/approach were understood? | git log: specify (c96d21c) → clarify (675abb6) → research (0620c94) → ideate (1a0326b) → plan (d8473d6). PLAN.md created after research + ideation. | The phase plan was created AFTER research and ideation — the scope was understood before the plan was written. No reverse-engineered deadlines. The "timeline" is the phase sequence P0→P6, not an external date. | 0.88 |
| 5.2 | What is the critical path, and what would push it by 3+ months? | PLAN.md: P1→P2→P3→P4→P5→P6, all sequential. P4 (matrix) depends on all domains. P5 (examples) depends on P4 (manifest). | Critical path: P1→P2→P3→P4→P5→P6. **Nothing can push it by 3+ months** — this is a docs project with no external dependencies, no infrastructure provisioning, no vendor lead times. The only "push" is content-quality rework, which is bounded (markdown edits). | 0.90 |
| 5.3 | Are estimates evidence-based? | v0.1: 35 reqs, 7 phases, shipped. v0.2: 24 reqs, 5 phases, shipped. v0.3: 32 reqs, 6 phases. Per-phase doc counts: P1=5, P2=5, P3=10, P4=7, P5=4, P6=2. | Estimates are analogous (v0.1/v0.2 shipped with similar per-phase doc counts). v0.1 P3 produced 27 derived docs in one phase; v0.3's largest phase (P3) produces 10. The estimate is conservative relative to v0.1's demonstrated throughput. | 0.85 |
| 5.4 | Is there a working definition of done? | PLAN.md P6 Verify: all 32 reqs covered, reconstruction test passes, matrix row-count test (170), MANIFEST reconstruction test (incl. examples/), audit clean, tag v0.2.6 exists, branches deleted. | DoD is concrete and testable: 32 reqs covered, matrix = 170 rows (10 × 17), MANIFEST reconstruction passes, tag v0.2.6 on main. Not "whatever the demo shows." | 0.88 |
**Axis 5 confidence: 0.88.** No challenges.
### Axis 6 — Budget and Financial Realism
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| 6.1 | What % of budget is spent vs remaining? | N/A — docs-only, no monetary budget. Cost = agent tokens. v0.1/v0.2 completed within expected token bounds. | No monetary budget. Token cost is proportional to markdown authored. v0.3 is the largest milestone (32 reqs, 18 derived docs + 4 examples + extensions), but per-doc token cost is roughly constant and validated by 2 prior milestones. | 0.85 |
| 6.2 | Are there predictable cost drivers not in the original budget? | PROJECT.md: no runtime code, no infrastructure, no licensing, no external services. | None. Docs-only = no licensing, no infra, no security review fees, no data migration, no support contracts. The only cost driver is markdown volume, which is scoped by REQ count. | 0.90 |
| 6.3 | Burn rate and runway? | N/A — no monetary burn. Agent execution time is the only resource. | No monetary runway concern. Agent execution is bounded by phase task count. | 0.88 |
| 6.4 | Is the budget contingent on something? | config.json: no contingent conditions. | No. The project is not contingent on a sale, board approval, or hiring. | 0.90 |
**Axis 6 confidence: 0.88.** No challenges. The docs-only constraint makes this axis low-risk by construction.
### Axis 7 — Risks, Assumptions, and Dependencies
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| 7.1 | Top 3 assumptions the plan rests on? | PLAN.md Assumptions: A-001 (ESC-002 is the only manifest drift, 0.85), A-002 (4 domains map to C1-C8 without new core principles, 0.95), A-003 (P3 Wave 2 schedules 5-then-3, 0.90). | A-002 is the strongest (0.95) — RESEARCH.md pre-maps all 40 P-rules to existing C-rules. A-001 is reasonable (AUDIT-P5-resync.md confirms). A-003 is mechanical (config.json cap). All three have evidence. | 0.86 |
| 7.2 | External dependencies? | PROJECT.md: no runtime code. RESEARCH.md: all prior art is published (OpenGitOps, ICU/CLDR, NIST, OPA docs). | None. No vendor, regulator, or external team dependency. All prior art is published and cited. The framework is self-contained markdown. | 0.92 |
| 7.3 | Single project-killing risk? | RESEARCH.md Risks table: orphaned P-rules, domain overlap, AI/ML scope drift, compliance bloat, artifact leakage, persona explosion. | The single risk that would undermine the framework's core value: **matrix orphan P-rules.** If the 40 new P-rules don't trace cleanly to C-rules, the traceable hierarchy (the unique value proposition) is broken. **Mitigation:** P4 matrix row-count test (10 per domain × 17 = 170) + domain-expert sign-off + P6 reconstruction test. This is well-mitigated and explicitly verified. | 0.84 |
| 7.4 | Pre-mortem: 12 months from now, v0.3 failed. Why? | (adversarial analysis) | **Most likely failure:** P3 content quality — i18n RTL/bidi or compliance policy-as-code authored without specialist personas contains fundamental technical errors (e.g., bidi isolating run misuse, Rego evaluation model mischaracterization), caught in P6 review, requiring P3 rework. **Secondary:** the 2+2 example set leaves 2 domains without a good example (gitops + ai-ml get good examples; i18n + compliance get only bad examples), reducing adoption value. **Both are bounded** — rework is markdown edits, not infrastructure. Neither is project-killing. | 0.78 |
**Axis 7 confidence: 0.85.** Challenge: pre-mortem P3 content quality (G-012). Auto-resolved.
### Axis 8 — Governance, Decision-Making, and Communication
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| 8.1 | Who is the decision-maker when executives disagree? | git log: single author. config.json: autonomy "full". | Single owner-operator. No executive disagreement possible. Decisions are recorded in PROJECT.md (D-001..D-053) with confidence scores. | 0.88 |
| 8.2 | How often does governance meet, and what's the escalation pattern? | config.json: autonomy level "full", escalation_hooks ["deploy", "delete_data", "merge_to_main"], decision_confidence_threshold 0.6, escalation_timeout_ms 300000. | Governance is the ciagent workflow itself (specify→clarify→research→ideate→plan→grill→execute→review→audit→ship). Escalation: 3 hooks (deploy, delete_data, merge_to_main) + timeout 300s. The grill stage (this) is the pre-execution gate. Cadence is event-driven, not calendar-driven — appropriate for agent execution. | 0.82 |
| 8.3 | What is omitted from status reports? | AUDIT-P5-resync.md: candid about ESC-001/ESC-002 (git auth failure, stale tags). Convention note about examples/ MANIFEST drift flagged, not hidden. | Status reporting is transparent. The v0.2 audit explicitly flagged the examples/ MANIFEST drift (ESC-002) rather than burying it. v0.3 picks it up as ATELIER-91. No evidence of optimistic glossing. | 0.85 |
| 8.4 | Is there a "stop the project" trigger? | This grill stage. Verdict options: Proceed, Reduce scope, Rethink, Escalate. | The grill IS the stop-the-project gate. If the verdict were "Rethink" or "Escalate" with blocking issues, P1 would not proceed. Cancellation is not politically impossible — it's a mechanical verdict. | 0.85 |
**Axis 8 confidence: 0.85.** No challenges.
### Axis 9 — Change, Adoption, and Operational Readiness
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| 9.1 | Who uses this, how does their work change, what's in it for them? | PROJECT.md: "consumed by AI agents as pre-completion guidance and by humans as engineering canon." review/agent-checklist.md is the pre-completion gate. | Users: AI agents (read MANIFEST→core→domains before tasks) and humans (engineering reference). Their work changes by consulting a traceable principle hierarchy before acting. The value: conflict resolution has an arbiter (the matrix), not just opinions. | 0.80 |
| 9.2 | Is the operations team involved now or handed a finished product? | N/A — docs-only repo, no operations team. | No operations team. The "go-live" is tagging v0.2.6 + merging to main. No runtime to operate, no on-call, no SLO. | 0.92 |
| 9.3 | What is the rollback plan? | Standard git workflow: revert + re-tag. | Git revert + re-tag. Standard for docs repos. v0.2 demonstrated the pattern (ESC-002 force-updated tags). | 0.88 |
| 9.4 | Have success criteria been validated with judges? | ROADMAP.md Success Criteria: all checked for v0.1. v0.2 completed clean (AUDIT-P5-resync.md: 10 checks pass). | Success criteria validated by 2 prior milestones. The pattern: reqs covered + reconstruction test + audit clean + tag exists. v0.3 adds matrix row-count test (170) and MANIFEST reconstruction (incl. examples/). | 0.85 |
**Axis 9 confidence: 0.86.** No challenges.
### Meta — Closing Review
| # | Forcing Question | Evidence | Answer | Confidence |
|---|-----------------|----------|--------|------------|
| M.1 | If you were the auditor, what would you flag? | (adversarial analysis) | The P3 specialist-persona gap (i18n + compliance authored by tech-writer without specialists). It's a content-quality risk, not a traceability risk. The mitigation (RESEARCH.md depth + P6 review) is adequate but not redundant — there's no second layer of content validation beyond domain-expert traceability checks. | 0.75 |
| M.2 | What is the project NOT doing that it should? | PROJECT.md out-of-scope: tooling/linters. | Not doing: automated P-rule→C-rule trace validation (manual domain-expert verification only). This is explicitly out of scope (tooling is future work). At 170 P-rules, manual verification is still feasible but approaching the limit where tooling would be warranted. Not a v0.3 blocker. | 0.80 |
| M.3 | Simplest 80% version? | (adversarial analysis) | Ship 2 domains (gitops-operators + ai-ml) and defer i18n + compliance to v0.4. This halves the scope. **But:** v0.2 already deferred these once (IDEATE-15/16). Re-deferral makes them 2x deferred — zombie risk. The current 4-domain plan is better than re-deferral. D-022's pairing of the two smaller domains in P3 is the right load-balance call. | 0.82 |
| M.4 | What must be true in 90 days for success, and is it true today? | (adversarial analysis) | Must be true: (a) 40 new P-rules trace cleanly to C-rules — RESEARCH.md pre-maps them, P4 verifies. (b) 18 derived docs each have ≥1 valid cross-link — ATELIER-88 verifies in P5. (c) Content is technically correct (especially i18n bidi + compliance policy-as-code) — RESEARCH.md grounds it, P6 validates. (a) and (b) have explicit verification gates. (c) depends on execution quality. All three are achievable. | 0.80 |
**Meta confidence: 0.79.**
---
### Binding Decisions
| ID | Decision | Rationale | Confidence | Alternatives |
|----|----------|-----------|------------|--------------|
| G-001 | Proceed with 4-domain scope as planned; do not split into v0.3a/v0.3b | Scope is deliberately trimmed (D-023/D-024/D-025); per-phase doc counts (P1:4, P2:4, P3:8) are lower than v0.1 P3 (27); v0.2 already deferred these domains once — re-deferral risks zombie status; D-016 groups them to limit release overhead | 0.80 | Split into 2 milestones (rejected: 2x deferral + double release overhead); reduce to 2 domains (rejected: same) |
| G-002 | Phase-specific personas (platform-engineer, ml-engineer) are NOT a key-person risk | Personas are constraint sets, not humans; D-051 bakes constraints into task must-haves (survive persona removal); D-052 ensures both review content in P6 before removal; platform-engineer is proven reuse from v0.2 (shipped clean) | 0.80 | Add more specialist personas (rejected: persona explosion); remove phase-specific personas (rejected: content quality risk) |
| G-003 | 40 new P-rules will trace cleanly to C-rules; matrix orphan risk is mitigated | RESEARCH.md pre-maps all 40 P-rules to C-rules (D-029..D-032); P4 matrix row-count test (10 × 17 = 170) + domain-expert sign-off; P6 reconstruction test; A-002 (confidence 0.95) confirms core is stable at 8 | 0.85 | Force broader C-rule derivations (rejected: false derivations worse than accurate narrow ones); add new core principles (rejected: A-002 confirms unnecessary) |
| G-004 | i18n + compliance pairing in P3 (D-022) is load-balancing-sound | Both are 4-derived-doc domains (smaller surface than P1/P2); P3 Wave 2 schedules 8 docs as 5-then-3 (A-003, 0.90); the scheduling is not a bottleneck — the content-quality risk is (see G-009) | 0.82 | Separate i18n and compliance into distinct phases (rejected: adds a phase, no load benefit); move one to P2 (rejected: P2 is AI/ML, a heavier domain) |
| G-005 | 2-good + 2-bad example set (D-025) is adequate for 4 domains | Each domain gets exactly 1 example (gitops: good, ai-ml: good, i18n: bad, compliance: bad); ATELIER-84 adds 4 domain-specific anti-patterns per domain (16 total) = 5 illustration points per domain; v0.1 had 7 examples for 11 domains (most domains had 0) — v0.3 is better coverage | 0.75 | 4-good + 4-bad (rejected: D-025 unbalances P5); 1-per-domain good only (rejected: bad examples have higher illustration value for i18n/compliance) |
| G-006 | Compliance C4 (Locality) gap (D-032) is architecturally sound, NOT a smell | C4 Locality is about keeping concerns local; compliance is inherently cross-cutting (audit logs span the system, retention is global, evidence is cross-cutting); forcing C4 would be a false derivation; D-032 considered and rejected the alternative "force C4 via audit-log locality"; the absence reflects the domain's nature | 0.82 | Force C4 via audit-log locality (rejected: false derivation); add a C4-tracing compliance P-rule (rejected: would be artificial) |
| G-007 | Docs-only constraint (D-020) is defensible at 17 domains — more so, not less | The docs-only constraint eliminates runtime complexity, integration surface, deployment, and build steps — the exact things that make large frameworks unwieldy; MANIFEST + reading-order + matrix make the tree navigable at scale; the framework's value (traceable hierarchy) scales linearly with domains; 170 matrix rows is a large but single readable file | 0.88 | Add a build step / linter (rejected: out of scope, violates docs-as-code simplicity); split the matrix per-domain (rejected: destroys the join-table value) |
| G-008 | No hidden data-engineer persona requirement from adding ai-ml domain | D-023 scopes AI/ML to engineering discipline (data versioning, evaluation, serving, drift), NOT data engineering (schema/ETL/pipelines); PERSONAS.md explicitly deactivates data-engineer ("No database, schema, or migrations in this docs-only project"); ai-ml cross-links to data/ as a reference, not a duplication | 0.85 | Activate data-engineer persona (rejected: no data engineering work in a docs-only project); scope AI/ML to include data engineering (rejected: D-023 explicitly excludes) |
| G-009 | Proceed with tech-writer + domain-expert for P3 (i18n + compliance) without specialist personas | RESEARCH.md provides thorough prior-art surveys (ICU/CLDR, BCP 47, UAX #9, W3C i18n WG, OPA/Cedar/Kyverno/Sentinel, Cosign/in-toto) with specific source citations; v0.1 tech-writer authored security/data/concurrency without specialists and shipped clean; P6 review includes domain-expert validation; rework is bounded (markdown edits). **Assumption A-006 logged:** content correctness for i18n RTL/bidi and compliance policy-as-code depends on RESEARCH.md prior-art quality + P6 review | 0.72 | Add i18n-specialist + compliance-specialist personas (rejected: persona explosion, both domains are smaller-surface per D-022); defer i18n/compliance to v0.4 with specialists (rejected: 2x deferral zombie risk) |
| G-010 | Matrix coverage summary must read "17 domains, 170 P-rules" post-v0.3 (reinforces D-036) | D-036 (confidence 0.93) already decided this; PLAN task 04-01-01 bakes it in; both the summary block AND per-domain section count must update; the P6 audit verifies row-count (10 × 17 = 170) | 0.93 | Partial update (rejected: invariant violation) |
| G-011 | ATELIER-91 closes the v0.2 ESC-002 manifest drift (examples/ unlisted) | AUDIT-P5-resync.md §6 flagged examples/ not in MANIFEST as a P1 convention note; ATELIER-91 (D-035, confidence 0.85) adds examples/ directory listing to MANIFEST in P4; PLAN task 04-02-05 bakes it in; A-001 (0.85) confirms this is the only manifest drift | 0.85 | Leave examples/ unlisted (rejected: manifest is authoritative, unlisted = not part of framework by definition) |
| G-012 | Pre-mortem top risk (P3 content quality) is bounded; no escalation needed | Most likely failure mode: i18n bidi or compliance policy-as-code technical errors caught in P6, requiring P3 rework. Mitigation: RESEARCH.md prior-art depth + P6 domain-expert review. Rework is bounded (markdown edits, not infrastructure). No external dependencies. The risk is real but recoverable and does not block the milestone. | 0.78 | Defer P3 to v0.4 (rejected: zombie risk); add specialist personas (rejected: G-009 analysis) |
### Escalations
**None.** All 12 challenges auto-resolved at confidence ≥ 0.60 (range: 0.720.93). At full autonomy, assumption logging (A-006) is preferred over escalation. No axis scored below 0.60 on any forcing question.
### Assumptions Logged (this grill)
| # | Assumption | Confidence |
|---|-----------|------------|
| A-006 | i18n RTL/bidi and compliance policy-as-code content correctness depends on RESEARCH.md prior-art quality + P6 domain-expert review, not on a specialist persona. If P6 surfaces fundamental content errors, the remedy is P3 rework (bounded — markdown edits). | 0.72 |
### Summary
- **Challenges identified:** 12 (across 9 axes + meta)
- **Binding decisions:** 12 (G-001..G-012)
- **Escalations:** 0
- **Assumptions logged:** 1 (A-006)
- **Verdict:** PROCEED at confidence 0.80
- **Top 3 material challenges:**
1. **G-009 (0.72):** P3 i18n + compliance authored without specialist personas — the lowest-confidence decision. Content quality for bidi algorithms and policy-as-code semantics depends on RESEARCH.md depth, not specialist persona constraints.
2. **G-005 (0.75):** 2+2 examples across 4 domains — each domain gets only one example (good OR bad, not both). Adequate given 16 anti-patterns, but thinner than v0.2's per-domain coverage.
3. **G-001 (0.80):** 4-domain scope is the largest single milestone — bounded by D-023/D-024/D-025 trims, but re-deferral would create zombie risk.
- **No blocking escalations for SHIP.**
+23 -6
View File
@@ -46,9 +46,25 @@
## Phase-Specific Personas ## Phase-Specific Personas
None. All three active personas span the full milestone. No phase-scoped personas needed — the work is uniformly markdown authoring with domain validation. ### platform-engineer (v0.3 — extended, active for v0.3, removed after milestone completion)
## Phase-Specific Personas - **active:** true
- **phase_specific:** true
- **domain:** infrastructure/platform-automation
- **frameworks:** []
- **constraints:** ["declarative-first", "stateless examples", "trace to core", "10 P-rules per domain", "no runtime code", "source-of-truth is git", "reconciliation loop is the primitive"]
- **territory:** ["domains/gitops-operators/**", "examples/good/gitops-pr.md", "examples/bad/* (gitops-related)"]
- **reason:** Per D-019 / D-027: the v0.2 platform-engineer persona is reused and extended for P1 (gitops-operators), because GitOps/Operators/Progressive Delivery build directly on the k8s + IaC declarative-reconciliation model the persona already embodies. Removed after v0.3 completes; roster returns to 3 active personas. The v0.2 IaC/k8s content remains owned by tech-writer + domain-expert for cross-link maintenance.
### ml-engineer (v0.3 — phase-specific, removed after milestone completion)
- **active:** true
- **phase_specific:** true
- **domain:** machine-learning engineering
- **frameworks:** []
- **constraints:** ["reproducibility is non-negotiable", "data lineage is traceable", "trace to core", "10 P-rules per domain", "no runtime code", "engineering discipline not algorithm design (D-023)", "examples are illustrative markdown only"]
- **territory:** ["domains/ai-ml/**", "examples/good/ai-ml-reproducibility.md"]
- **reason:** Per D-019 / D-020: AI/ML domain authoring (data versioning, model evaluation, serving, monitoring/drift) benefits from a specialist persona with reproducibility and data-lineage constraints the existing tech-writer persona lacks. Scope is engineering discipline, NOT algorithm/model design (D-023). Removed after v0.3 completes; roster returns to 3 active personas.
### platform-engineer (v0.2 — REMOVED after milestone completion) ### platform-engineer (v0.2 — REMOVED after milestone completion)
@@ -66,8 +82,9 @@ Mode: `warn` (per config.json `personas.territory_enforcement`).
At `warn`, territory violations are logged but not blocked. This is appropriate for a docs project where tech-writer may touch `.ciagent/` files incidentally (e.g., updating ROADMAP status). Strict mode would be appropriate once territories stabilize. At `warn`, territory violations are logged but not blocked. This is appropriate for a docs project where tech-writer may touch `.ciagent/` files incidentally (e.g., updating ROADMAP status). Strict mode would be appropriate once territories stabilize.
## v0.2 Persona Roster Summary ## v0.3 Persona Roster Summary
Active personas for v0.2 (4): lead-developer, tech-writer, domain-expert, platform-engineer (phase-specific). Active personas for v0.3 (5): lead-developer, tech-writer, domain-expert (span full milestone), platform-engineer (phase-specific, P1 gitops-operators), ml-engineer (phase-specific, P2 ai-ml).
Inactive personas (3, unchanged from v0.1): data-engineer, backend-engineer, frontend-engineer. - **Phase assignment:** P1 GitOps/Operators → platform-engineer; P2 AI/ML → ml-engineer; P3 i18n + Compliance → tech-writer + domain-expert (D-022); P4P5 matrix/review/examples → tech-writer + domain-expert + lead-developer.
Post-v0.2: platform-engineer removed; roster returns to 3 active personas (lead-developer, tech-writer, domain-expert). Inactive personas (3, unchanged): data-engineer, backend-engineer, frontend-engineer.
Post-v0.3: platform-engineer + ml-engineer removed; roster returns to 3 active personas (lead-developer, tech-writer, domain-expert).
+248
View File
@@ -372,3 +372,251 @@ Tag: v0.1.0
| 4 | ATELIER-53..56 | 4 | | 4 | ATELIER-53..56 | 4 |
| 5 | ATELIER-57, 58 | 2 | | 5 | ATELIER-57, 58 | 2 |
| **Total** | | **24** | | **Total** | | **24** |
---
# Atelier — Plan (v0.3)
> Vertical-slice plans with wave ordering for milestone v0.3 (GitOps + Operators + AI/ML + i18n + Compliance). Plans reference REQ-IDs from `.ciagent/atelier/REQUIREMENTS.md` (ATELIER-60..91). NFR milestone — all phases produce docs; no `feat` code. Per `parallelization.max_concurrent_agents = 5`, wave parallelism is capped at 5 concurrent tasks; waves larger than 5 are split into sub-waves.
## Phase 0 — Pre-Execution (COMPLETE)
Stages: SPECIFY ✓ → CLARIFY ✓ → RESEARCH ✓ → IDEATE ✓ → PLAN ✓ → GRILL → SHIP
Branch: `atelier/phase/00-pre-execution`
Tag: v0.2.0
## Phase 1 — GitOps + Operators Domain
**Goal:** Author the `domains/gitops-operators/` tree — 10 first principles (P1P10) plus 4 derived docs (argocd, flux, operators, progressive-delivery). Grounded in CNCF OpenGitOps Principles v1.0.0 + the Operator pattern. Each P-rule derives from core C1C8 (matrix extension lands in P4).
**Branch:** `atelier/phase/01-gitops-operators` (from `atelier/milestone/v0.3-atelier`)
**Personas:** platform-engineer (author), domain-expert (validate traceability), tech-writer (style/format)
**Tag:** v0.2.1
**Requirements:** ATELIER-60, ATELIER-61, ATELIER-62, ATELIER-63, ATELIER-64
### Wave 1 (sequential — first-principles must exist before derived docs)
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 01-01-01 | `domains/gitops-operators/first-principles.md` | platform-engineer | ATELIER-60 | 10 principles (P1P10) per RESEARCH.md (Git is Source of Truth, Pull Don't Push, Continuous Reconciliation, Operators Encode Domain Knowledge, Progressive Delivery is Reversible, Reconcile Don't Mutate, Failure is Observable, Least Privilege Reconciliation); each names the core C-rule(s) it derives from; each has definition + "what violates" |
### Wave 2 (parallel — 4 derived docs, independent; ≤5 concurrent)
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 01-02-01 | `domains/gitops-operators/argocd.md` | platform-engineer | ATELIER-61 | Application CRD, App-of-Apps, sync waves, health/status, diff, RBAC/SSO, multi-cluster, sync windows; **ArgoCD vs Flux decision matrix** (IDEATE-21, D-039); cross-link to flux.md, kubernetes/{workloads,rbac,helm,kustomize}.md, devops, security/secrets, observability/metrics |
| 01-02-02 | `domains/gitops-operators/flux.md` | platform-engineer | ATELIER-62 | GitOps Toolkit controllers (source, kustomize, helm, notification), composable architecture, HR/Kustomization/HelmRelease CRDs, OCI sources; **ArgoCD vs Flux decision matrix** (IDEATE-21, D-039); cross-link to argocd.md + kubernetes/helm.md + kubernetes/kustomize.md |
| 01-02-03 | `domains/gitops-operators/operators.md` | platform-engineer | ATELIER-63 | Operator pattern, CRDs, controllers, Operator SDK/OLM, when-to-write-an-operator vs Helm chart, scope/responsibility boundaries; cross-link kubernetes/{workloads,rbac}.md + infrastructure-as-code/modules.md |
| 01-02-04 | `domains/gitops-operators/progressive-delivery.md` | platform-engineer | ATELIER-64 | Argo Rollouts + Flagger, canary/blue-green, analysis templates (metrics/counters), abort/rollback; cross-link devops (P4 Rollback First, P5 Progressive Delivery) + observability/metrics + kubernetes/workloads.md |
**Verify (P1):**
- Structural: 5 files exist under `domains/gitops-operators/`
- Behavioral: every P1P10 in first-principles names ≥1 C-rule (domain-expert sign-off)
- Security: P3 (Pull, Don't Push) + P10 (Least Privilege Reconciliation) sections present
- Quality: each derived doc has ≥1 outbound cross-link to a MANIFEST-listed doc (IDEATE-08 carried forward); argocd.md and flux.md share the decision matrix consistently (IDEATE-21)
## Phase 2 — AI/ML Domain
**Goal:** Author the `domains/ai-ml/` tree — 10 first principles (P1P10) plus 4 derived docs (data-versioning, model-evaluation, serving, monitoring-drift). Scope = engineering discipline (D-023), NOT algorithm/model design. Reproducibility and lineage are non-negotiables.
**Branch:** `atelier/phase/02-ai-ml` (from `atelier/milestone/v0.3-atelier`)
**Personas:** ml-engineer (author), domain-expert (validate traceability), tech-writer (style/format)
**Tag:** v0.2.2
**Requirements:** ATELIER-65, ATELIER-66, ATELIER-67, ATELIER-68, ATELIER-69
### Wave 1 (sequential — first-principles first)
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 02-01-01 | `domains/ai-ml/first-principles.md` | ml-engineer | ATELIER-65 | 10 principles (P1P10) per RESEARCH.md (Reproducibility First Class, Data is Versioned Not Just Code, Lineage Traceable End-to-End, Evaluation Defined Before Training, Models are Versioned Artifacts, Serving is Observable, Drift is Expected and Detected, Inference Inputs are Validated, Pipelines Compose Notebooks Don't, Rollback Includes the Model); each names core C-rule(s); each has definition + "what violates"; ml-engineer constraint "engineering discipline not algorithm design (D-023)" enforced |
### Wave 2 (parallel — 4 derived docs, independent; ≤5 concurrent)
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 02-02-01 | `domains/ai-ml/data-versioning.md` | ml-engineer | ATELIER-66 | DVC/Delta Lake/LakeFS patterns, data lineage, dataset hashing, train/val/test split versioning; **tool comparison table: DVC vs Delta Lake vs LakeFS** covering versioning model, lineage, use-case fit (IDEATE-22, D-040); cross-link data/{migrations,schema-design}.md + devops/P1 Reproducibility |
| 02-02-02 | `domains/ai-ml/model-evaluation.md` | ml-engineer | ATELIER-67 | Metric selection, offline/online eval, holdout integrity, bias/fairness checks (engineering angle), eval-as-a-gate; cross-link data/schema-design.md (eval input contract) + testing/pyramid.md |
| 02-02-03 | `domains/ai-ml/serving.md` | ml-engineer | ATELIER-68 | KServe/Seldon/BentoML, inference as a service, batching, latency SLAs, canarying models; cross-link kubernetes/workloads.md + devops (P5 Progressive Delivery, P7 Immutability) + performance/backend.md + security/input-validation.md |
| 02-02-04 | `domains/ai-ml/monitoring-drift.md` | ml-engineer | ATELIER-69 | Evidently/Great Expectations, alerting, retraining triggers; **drift-type enumeration: data drift, concept drift, prediction drift — each with a distinct detection signal** (IDEATE-30, D-048); cross-link observability/{metrics,logging}.md + ai-ml/serving.md |
**Verify (P2):**
- Structural: 5 files exist under `domains/ai-ml/`
- Behavioral: every P1P10 traces to ≥1 C-rule (domain-expert sign-off)
- Security: P8 (Inference Inputs are Validated) section present
- Quality: each derived doc ≥1 outbound cross-link (IDEATE-08); data-versioning.md tool comparison table present (IDEATE-22); monitoring-drift.md enumerates 3 drift types with detection signals (IDEATE-30); no algorithm-design content (D-023 enforced, ml-engineer constraint)
## Phase 3 — i18n + Compliance Domains
**Goal:** Author two smaller-surface domains in one phase (D-022): `domains/i18n/` (10 first principles + 4 derived docs) and `domains/compliance/` (10 first principles + 4 derived docs). i18n grounded in ICU/CLDR + BCP 47 + W3C i18n. Compliance is framework-agnostic (D-024 — no regulation-specific docs). Both domains' first-principles land in Wave 1 (independent of each other), then derived docs in Wave 2.
**Branch:** `atelier/phase/03-i18n-compliance` (from `atelier/milestone/v0.3-atelier`)
**Personas:** tech-writer (author, both domains), domain-expert (validate traceability for both)
**Tag:** v0.2.3
**Requirements:** ATELIER-70, ATELIER-71, ATELIER-72, ATELIER-73, ATELIER-74, ATELIER-75, ATELIER-76, ATELIER-77, ATELIER-78, ATELIER-79
### Wave 1 (parallel — 2 first-principles, independent; ≤5 concurrent)
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 03-01-01 | `domains/i18n/first-principles.md` | tech-writer | ATELIER-70 | 10 principles (P1P10) per RESEARCH.md (Source Language is a Locale Not the Default, Locale Identifiers Standardized BCP 47, Resources External Not Inline, Plural/Gender Parameterized ICU MessageFormat, Formatting Locale-Aware ICU/CLDR, Text Direction is Layout Primitive, Layout Accommodates Expansion, Pseudo-Locales Test Early, Images/Icons Cultural, Translation Reversible and Versioned); each names core C-rule(s); each has definition + "what violates" |
| 03-01-02 | `domains/compliance/first-principles.md` | tech-writer | ATELIER-75 | 10 principles (P1P10) per RESEARCH.md (Audit Logs Append-Only, Every Significant Action Logged, Retention is Policy Not Storage, Policy is Code, Policy Evaluated as a Gate, Evidence Collected Continuously, Identity Attributable, Subject Access Honored, Secrets Redacted in Audit, Compliance Posture Observable); each names core C-rule(s); each has definition + "what violates"; framework-agnostic (D-024 — no GDPR/HIPAA/SOC2-specific content) |
### Wave 2 (parallel — 8 derived docs, independent; split into 2 sub-waves of 4 to respect max_concurrent_agents=5)
**Wave 2a (i18n derived docs, ≤5 concurrent)**
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 03-02a-01 | `domains/i18n/locale-resources.md` | tech-writer | ATELIER-71 | Resource file formats (.po/.pot, JSON, Fluent FTL, ICU Resource Bundle), key naming, namespaces, fallback chains, extraction tooling; cross-link uiux/copywriting.md + api/error-responses.md |
| 03-02a-02 | `domains/i18n/formatting.md` | tech-writer | ATELIER-72 | ICU/CLDR/Intl for dates, times, numbers, currencies, units, relative time, plural rules; BCP 47 tags; cross-link api/error-responses.md (localized errors) + data/schema-design.md |
| 03-02a-03 | `domains/i18n/rtl-bidi.md` | tech-writer | ATELIER-73 | Logical vs physical CSS properties, bidi algorithm (UAX #9), `dir` attribute, mirroring, common pitfalls (icons, numbers in RTL); cross-link uiux/{components,accessibility}.md |
| 03-02a-04 | `domains/i18n/testing-i18n.md` | tech-writer | ATELIER-74 | Pseudo-locales, snapshot testing per locale, RTL coverage, missing-key detection; **pseudo-locale tier mapping to testing pyramid: unit (missing-key), integration (snapshot per locale), e2e (RTL coverage)** (IDEATE-28, D-046); cross-link testing/{fixtures,pyramid}.md |
**Wave 2b (compliance derived docs, ≤5 concurrent; runs in parallel with 2a — total 8 tasks, but capped at 5 → executor schedules 5 then 3)**
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 03-02b-01 | `domains/compliance/audit-logs.md` | tech-writer | ATELIER-76 | Append-only log patterns, structured audit events, CloudTrail/Cloud-Audit-Log conventions, queryability, retention of logs; cross-link observability/logging.md + security/authorization.md |
| 03-02b-02 | `domains/compliance/data-retention.md` | tech-writer | ATELIER-77 | Retention policies as code, lifecycle rules, deletion-as-a-feature, GDPR/CCPA abstracted to principles (not regulation-specific, D-024), retention vs backup distinction; cross-link data/{migrations,schema-design}.md |
| 03-02b-03 | `domains/compliance/policy-as-code.md` | tech-writer | ATELIER-78 | OPA/Cedar/Sentinel/Kyverno patterns, policy as CI/CD + admission gate, policy testing, versioning policy; **engine comparison table: OPA vs Cedar vs Kyverno vs Sentinel** covering policy language, evaluation gate, ecosystem (IDEATE-23, D-041); cross-link infrastructure-as-code (declarative intent) + kubernetes/rbac.md (admission) |
| 03-02b-04 | `domains/compliance/evidence.md` | tech-writer | ATELIER-79 | Evidence collection as a byproduct, audit-ready export, provenance; **fenced signed-attestation example (Cosign OR in-toto)** — not prose-only (IDEATE-29, D-047); cross-link security/supply-chain.md + observability/{metrics,tracing}.md |
> **Parallelism note:** Wave 2a + 2b together = 8 independent tasks. The executor schedules at most 5 concurrently per `parallelization.max_concurrent_agents`; the remaining 3 run as soon as slots free. The 2a/2b labels are organizational (by domain), not a hard sequencing barrier — both sub-waves are in the same dependency tier (all depend only on Wave 1).
**Verify (P3):**
- Structural: 10 files exist (5 under `domains/i18n/`, 5 under `domains/compliance/`)
- Behavioral: every P1P10 in both first-principles traces to ≥1 C-rule (domain-expert sign-off)
- Security: i18n P6 (Text Direction) + compliance P1 (Append-Only) + P9 (Redacted) sections present
- Quality: each derived doc ≥1 outbound cross-link (IDEATE-08); testing-i18n.md pseudo-locale→pyramid mapping present (IDEATE-28); policy-as-code.md engine comparison table present (IDEATE-23); evidence.md has a fenced signed-attestation example (IDEATE-29); no regulation-specific content in compliance (D-024)
## Phase 4 — Matrix + Review + Manifest Integration
**Goal:** Extend the matrix (+40 P-rule → C-rule mappings, 10 per new domain), domain-coverage (per-domain rows + Core Principle Coverage table for 4 new domains), review docs (agent + peer-review + anti-patterns with v0.3 chaos anti-patterns), and the manifest (all v0.3 docs + `examples/` directory listing closing v0.2 ESC-002 drift). Closes the traceability loop and makes the manifest authoritative for v0.3.
**Branch:** `atelier/phase/04-matrix-review-manifest` (from `atelier/milestone/v0.3-atelier`)
**Personas:** domain-expert (matrix + anti-patterns + coverage), tech-writer (checklists + manifest), lead-developer (manifest authoritative index)
**Tag:** v0.2.4
**Requirements:** ATELIER-80, ATELIER-81, ATELIER-82, ATELIER-83, ATELIER-84, ATELIER-85, ATELIER-91
### Wave 1 (sequential — matrix is the arbiter, must be authoritative first)
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 04-01-01 | `matrix/principles-matrix.md` (extend) | domain-expert | ATELIER-80 | Add 4 sections (GitOps + Operators, AI/ML, i18n, Compliance), 10 rows each, format matching v0.1/v0.2 tables; **review check: row count per new domain = 10, each row ≥1 C-rule** (IDEATE-02, IDEATE-13 carried forward); **update Coverage Summary to "post-v0.3: 17 domains, 170 P-rules"** — both the summary block AND the per-domain section count (IDEATE-18, D-036) |
### Wave 2 (parallel — 5 independent extensions; exactly at max_concurrent_agents=5)
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 04-02-01 | `matrix/domain-coverage.md` (extend) | domain-expert | ATELIER-81 | Add "v0.3 Domain Coverage" table with 4 rows (schema: domain, P-count, derived-doc-count, manifest-listed, status per IDEATE-03); **AND update the "Core Principle Coverage" table (C1C8 → domains) for the 4 new domains** — C4 Locality adds i18n + gitops; C5 Reversibility adds ai-ml + compliance + gitops + i18n; etc. (IDEATE-19, D-037) |
| 04-02-02 | `review/agent-checklist.md` (extend) | tech-writer | ATELIER-82 | Add "If GitOps + Operators", "If AI/ML", "If i18n", "If Compliance" trigger sections (IDEATE-05 carried forward); ai-ml section includes a D-023 scope check (reject algorithm-design content) |
| 04-02-03 | `review/peer-review-checklist.md` (extend) | tech-writer | ATELIER-83 | Add 4 new domain peer-review sections (parity with agent-checklist, IDEATE-09 carried forward) |
| 04-02-04 | `review/anti-patterns.md` (extend) | domain-expert | ATELIER-84 | Add "v0.3 Chaos Anti-Patterns" section covering: (a) **v0.3 deployable artifact types** — .po resource files, .rego policy files, model artifacts, signed manifests as standalone files (IDEATE-20); (b) **GitOps push-pattern violation** (P3 Pull Don't Push, IDEATE-24, D-042); (c) **i18n LTR-only assumption violation** (P6 Text Direction, IDEATE-25, D-043); (d) **AI/ML orphan-model violation** — deployed prediction with no lineage trace (P3 Lineage, IDEATE-27, D-045); plus domain-specific anti-patterns per RESEARCH.md/REQUIREMENTS.md notes: gitops (push-based deploy P3, manual kubectl apply on GitOps-managed resource P8, cluster-admin GitOps robot P10), ai-ml (unreproducible training run P1, "the latest" model P5, notebook in production P9, orphan model P3), i18n (inline string concatenation P3, `if (n==1)` plural branching P4, LTR-only layout P6, hand-rolled date formatter P5), compliance (mutable audit log P1, shared/generic identity in audit P7, secret leaked in audit log P9, manual evidence assembly at audit time P6) |
| 04-02-05 | `MANIFEST.md` (extend) | lead-developer | ATELIER-85, ATELIER-91 | Add 4 new domains + all 18 derived docs to the Domains table (IDEATE-04 carried forward); update Cross-Cutting counts to "17 domains, 170 P-rules post-v0.3"; **add an `examples/` directory listing section** (good + bad files) closing the v0.2 ESC-002 drift — manifest is authoritative (IDEATE-17, D-035) |
**Verify (P4):**
- Structural: matrix has 17 domain sections (13 v0.1/v0.2 + 4 new), 170 P-rules total; domain-coverage has both the per-domain v0.3 table AND the updated C-rule coverage table; 3 review docs extended; MANIFEST lists all v0.3 docs + examples/
- Behavioral: every new P-rule has a matrix row; domain-expert verifies no orphans (IDEATE-13); Coverage Summary reads "17 domains, 170 P-rules" (IDEATE-18)
- Security: anti-patterns cover GitOps push-pattern (P3), AI/ML orphan-model (P3), i18n LTR-only (P6), compliance mutable audit log (P1) + secret-in-audit (P9)
- Quality: MANIFEST is authoritative — every v0.3 file listed, examples/ listed (ATELIER-91 closes ESC-002); unlisted = not part of framework; agent-checklist + peer-review-checklist have parity across the 4 new domains (ATELIER-82 ↔ ATELIER-83)
## Phase 5 — Examples + Cross-Links
**Goal:** Add 2 good + 2 bad examples (D-025 — highest illustration value) and verify cross-domain links from all 4 new domains to existing ones. Examples are markdown with fenced code only (no standalone .yaml/.po/.rego/model artifacts — D-020).
**Branch:** `atelier/phase/05-examples-crosslinks` (from `atelier/milestone/v0.3-atelier`)
**Personas:** tech-writer (examples + cross-link audit), domain-expert (P-rule citation + cross-link validation)
**Tag:** v0.2.5
**Requirements:** ATELIER-86, ATELIER-87, ATELIER-88
### Wave 1 (parallel — 4 examples, independent; ≤5 concurrent)
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 05-01-01 | `examples/good/gitops-pr.md` | tech-writer | ATELIER-86 | Good GitOps example; markdown with fenced YAML only (no standalone .yaml); demonstrates P1 Git is Source of Truth + P3 Pull Don't Push + P5 State Immutable and Versioned; cross-link to argocd.md + flux.md + kubernetes/workloads.md |
| 05-01-02 | `examples/good/ai-ml-reproducibility.md` | tech-writer (ml-engineer consult) | ATELIER-86 | Good AI/ML example; markdown with fenced code only (no model artifacts); demonstrates P1 Reproducibility + P2 Data Versioned + P3 Lineage Traceable + P5 Models are Versioned Artifacts; cross-link to data-versioning.md + serving.md |
| 05-01-03 | `examples/bad/i18n-string-concat.md` | tech-writer | ATELIER-87 | Bad i18n example; cites P3 breached (inline string concatenation, Resources External Not Inline) per IDEATE-07; cross-link to locale-resources.md + formatting.md |
| 05-01-04 | `examples/bad/compliance-audit-log.md` | tech-writer | ATELIER-87 | Bad compliance example; **TWO breaches in one example** (IDEATE-26, D-044): append-only violation (mutation/deletion of an audit record, P1) AND redaction failure (secret leaked in audit log, P9); cites both P-rules breached; cross-link to audit-logs.md + evidence.md |
### Wave 2 (sequential — cross-link audit after all docs exist)
| Task | File | Persona | REQ-ID | Must-have |
|------|------|---------|--------|-----------|
| 05-02-01 | Cross-link audit (all 18 new derived docs across 4 domains) | tech-writer | ATELIER-88 | Review check: every new derived doc ≥1 outbound cross-link to a MANIFEST-listed existing domain doc (devops/security/observability/data/kubernetes/infrastructure-as-code); links resolve (IDEATE-08 carried forward); domain-expert validates the cross-link targets are correct (not just present) |
**Verify (P5):**
- Structural: 4 new example files exist (all .md)
- Behavioral: each bad example cites the P-rule(s) breached (i18n: P3; compliance: P1 + P9 two-breach per IDEATE-26); each good example cites the P-rules it demonstrates
- Security: no standalone .yaml/.po/.rego/model artifacts (deployable artifact mitigation, IDEATE-20, D-020)
- Quality: all cross-links from the 18 new derived docs resolve to MANIFEST-listed docs; no back-link edits to v0.1/v0.2 content (D-026 extended — one-directional outward)
## Phase 6 — Final Review + Ship (N+1)
**Goal:** Multi-persona review across all v0.3 phases, audit, milestone ship. P6 IS the v0.3 release (NFR → no separate minor tag; v0.2.6 IS the deliverable). Phase-specific personas (platform-engineer, ml-engineer) are removed after milestone completion.
**Branch:** `atelier/phase/06-final-review-ship` (from `atelier/milestone/v0.3-atelier`)
**Personas:** lead-developer (coordinate + ship), domain-expert (review), tech-writer (review), platform-engineer (review, then removed), ml-engineer (review, then removed)
**Tag:** v0.2.6 (IS the v0.3 milestone release — NFR, no separate minor tag)
**Requirements:** ATELIER-89, ATELIER-90
### Wave 1 (sequential — review → audit → ship → complete)
| Task | Activity | Persona | REQ-ID | Must-have |
|------|----------|---------|--------|-----------|
| 06-01-01 | `ciagent-review` — multi-persona review of all v0.3 changes | lead-developer | ATELIER-89 | Auto-apply P0 fixes; flag P1+ for post-hoc; if P1+ found, fix in this phase; platform-engineer reviews gitops-operators content; ml-engineer reviews ai-ml content (D-023 scope check); domain-expert verifies all 40 new P-rules trace to ≥1 C-rule (no orphans) |
| 06-01-02 | `ciagent-audit` — reconstruction + discipline | lead-developer | ATELIER-89 | git log matches .ciagent/ files; branch hygiene; commit discipline (every commit has `---ci---` block with `project: atelier`); MANIFEST reconstruction test (every listed doc exists, every existing doc is listed — incl. examples/ per ATELIER-91); matrix row-count test (10 per domain × 17 = 170) |
| 06-01-03 | `ciagent-ship` — milestone ship | lead-developer | ATELIER-90 | Merge `atelier/phase/06``atelier/milestone/v0.3-atelier``main`; tag `v0.2.6`; Gitea release with full milestone summary; delete all v0.3 branches (tags preserve history) |
| 06-01-04 | Complete milestone (REQUIREMENTS + ROADMAP + PERSONAS) | lead-developer | ATELIER-90 | Mark all v0.3 requirements (ATELIER-60..91) `covered`; ROADMAP v0.3 → complete; PERSONAS: remove platform-engineer + ml-engineer (roster returns to 3 active); clear checkpoint |
**Verify (P6):**
- Structural: all 32 v0.3 requirements (ATELIER-60..91) marked covered
- Behavioral: reconstruction test passes (git log ↔ .ciagent/); matrix row-count test passes (170); MANIFEST reconstruction test passes (incl. examples/)
- Security: audit clean (no critical issues); no regulation-specific compliance content (D-024); no algorithm-design ai-ml content (D-023); no standalone runtime artifacts (D-020)
- Quality: milestone merged to main, tag v0.2.6 exists, all v0.3 branches deleted; platform-engineer + ml-engineer personas removed (roster = 3)
## v0.3 Wave Ordering Summary
| Phase | Waves | Parallelism |
|-------|-------|-------------|
| P0 | (pre-exec) | Sequential stages (specify→clarify→research→ideate→plan→grill) |
| P1 | 2 | Wave 1 sequential (first-principles), Wave 2 parallel (4 derived docs) |
| P2 | 2 | Wave 1 sequential (first-principles), Wave 2 parallel (4 derived docs) |
| P3 | 2 | Wave 1 parallel (2 first-principles), Wave 2 parallel (8 derived docs — 2a i18n + 2b compliance, capped at 5 concurrent) |
| P4 | 2 | Wave 1 sequential (matrix arbiter), Wave 2 parallel (5 extensions — exactly max_concurrent) |
| P5 | 2 | Wave 1 parallel (4 examples), Wave 2 sequential (cross-link audit) |
| P6 | 1 | Sequential: review → audit → ship → complete |
## v0.3 Requirements → Phase Mapping
| Phase | Requirements | Count |
|-------|-------------|-------|
| 1 | ATELIER-60..64 | 5 |
| 2 | ATELIER-65..69 | 5 |
| 3 | ATELIER-70..79 | 10 |
| 4 | ATELIER-80..85, 91 | 7 |
| 5 | ATELIER-86..88 | 3 |
| 6 | ATELIER-89, 90 | 2 |
| **Total** | | **32** |
## v0.3 Ideation Refinements → Task Bake-In Map
All 14 accepted ideation refinements (IDEATE-17..30) are baked into the relevant phase tasks as explicit must-have notes:
| IDEATE-ID | Refinement | Baked Into Task(s) | How |
|-----------|-----------|-------------------|-----|
| IDEATE-17 | examples/ in MANIFEST (new req ATELIER-91) | 04-02-05 | MANIFEST gains an examples/ directory listing (closes v0.2 ESC-002 drift) |
| IDEATE-18 | matrix coverage summary = "17 domains, 170 P-rules" | 04-01-01 | Coverage Summary block + per-domain section count both updated |
| IDEATE-19 | Core Principle Coverage table update for 4 new domains | 04-02-01 | C1C8 → domains table extended (C4 adds i18n+gitops; C5 adds ai-ml+compliance+gitops+i18n; etc.) |
| IDEATE-20 | anti-patterns pre-specify domain violations + v0.3 artifact types | 04-02-04 | .po/.rego/model/signed-manifest artifact types + 16 domain-specific anti-patterns (4 per domain) |
| IDEATE-21 | ArgoCD vs Flux decision matrix | 01-02-01, 01-02-02 | Both argocd.md and flux.md carry the decision matrix (parallel to v0.2 Helm vs Kustomize) |
| IDEATE-22 | data versioning tool comparison (DVC/Delta Lake/LakeFS) | 02-02-01 | data-versioning.md comparison table (versioning model, lineage, use-case fit) |
| IDEATE-23 | policy-as-code engine comparison (OPA/Cedar/Kyverno/Sentinel) | 03-02b-03 | policy-as-code.md comparison table (policy language, evaluation gate, ecosystem) |
| IDEATE-24 | GitOps push-pattern anti-pattern (violates P3) | 04-02-04 | Named chaos anti-pattern; pre-specified to reject on sight |
| IDEATE-25 | i18n LTR-only assumption anti-pattern (violates P6) | 04-02-04 | Named chaos anti-pattern; pre-specified to reject on sight |
| IDEATE-26 | compliance-audit-log bad example = 2 breaches (P1 + P9) | 05-01-04 | examples/bad/compliance-audit-log.md covers append-only violation + redaction failure |
| IDEATE-27 | AI/ML orphan-model anti-pattern (violates P3 Lineage) | 04-02-04 | Named chaos anti-pattern; deployed prediction with no lineage trace |
| IDEATE-28 | i18n testing pseudo-locale → testing pyramid tiers | 03-02a-04 | testing-i18n.md maps unit (missing-key), integration (snapshot per locale), e2e (RTL coverage) |
| IDEATE-29 | compliance evidence.md signed-attestation fenced example | 03-02b-04 | evidence.md includes a fenced Cosign OR in-toto attestation (not prose-only) |
| IDEATE-30 | ai-ml monitoring-drift.md 3 drift types with detection signals | 02-02-04 | monitoring-drift.md enumerates data/concept/prediction drift, each with a detection signal |
## v0.3 Decisions Logged (planning stage)
| ID | Decision | Rationale | Confidence |
|----|----------|-----------|------------|
| D-049 | P3 splits Wave 2 into 2a (i18n) + 2b (compliance) labels but both are the same dependency tier | 8 derived docs are all independent post-Wave-1; the 2a/2b labels organize by domain, the executor schedules ≤5 concurrent per config.json. Avoids inventing a false dependency between i18n and compliance | 0.88 |
| D-050 | P4 Wave 2 runs exactly 5 concurrent tasks (at the max_concurrent_agents cap) | matrix, coverage, agent-checklist, peer-review-checklist, anti-patterns, manifest = 6 extensions, but anti-patterns (04-02-04) and manifest (04-02-05) are combined under lead-developer for manifest to sequence after anti-patterns content is settled. Net 5 concurrent slots | 0.82 |
| D-051 | P5 ai-ml-reproducibility.md example authored by tech-writer with ml-engineer consultation (not ml-engineer primary) | ml-engineer is removed after P2 per PERSONAS.md; P5 examples are tech-writer territory. ml-engineer constraints are baked into the task must-have (P1/P2/P3/P5 demonstrated) so the constraint survives the persona | 0.80 |
| D-052 | P6 review uses platform-engineer + ml-engineer for content review before removal | Phase-specific personas review their authored content one final time in P6 Wave 1, then are removed in 06-01-04. Ensures D-023 (ai-ml scope) and GitOps correctness are checked by the specialist before the roster returns to 3 | 0.84 |
| D-053 | Vertical-slice integrity: each phase is independently shippable | P1 ships gitops-operators domain docs (matrix rows land in P4 — acceptable because the domain is self-consistent; matrix extension is the traceability closure, not a blocker for the domain's internal consistency). P3 ships 2 domains together (D-022). P4 closes traceability + manifest. P5 closes examples + cross-links. P6 ships the release | 0.86 |
## Assumptions Logged
| # | Assumption | Confidence |
|---|-----------|------------|
| A-001 | The v0.2 ESC-002 drift note (examples/ unlisted in MANIFEST) is the only pre-existing manifest drift; no other v0.1/v0.2 docs are unlisted | 0.85 |
| A-002 | The 4 new domains' P-rules map to existing core C1C8 without needing new core principles (core is stable at 8) | 0.95 |
| A-003 | Wave 2 of P3 (8 derived docs) can be scheduled by the executor as 5-then-3 without a hard sub-wave barrier | 0.90 |
| A-004 | The ArgoCD vs Flux decision matrix (IDEATE-21) is the only decision matrix required in P1 (no separate operators-vs-Helm matrix beyond operators.md's "when to write an operator vs a Helm chart" guidance) | 0.82 |
| A-005 | P4 anti-patterns (04-02-04) and manifest (04-02-05) can be concurrent because anti-patterns content does not block the manifest's examples/ listing (manifest lists file paths, not anti-pattern content) | 0.80 |
+67
View File
@@ -78,6 +78,48 @@ Build **Atelier** — a first-principles, docs-as-code engineering framework for
- New examples: `examples/good/terraform-module.md`, `examples/good/k8s-deployment.md`, `examples/bad/` counterparts - New examples: `examples/good/terraform-module.md`, `examples/good/k8s-deployment.md`, `examples/bad/` counterparts
- Cross-links from new domains to existing `devops/`, `security/`, `observability/`, `data/` domains - Cross-links from new domains to existing `devops/`, `security/`, `observability/`, `data/` domains
## v0.3 — GitOps + Operators + AI/ML + i18n + Compliance
**Milestone type:** NFR (all phases produce docs — no `feat` runtime code)
**Tag line:** v0.2.x (previous minor from v0.3)
**Scope:** Extend the domain tree with four new top-level domains covering GitOps/operator patterns, AI/ML, internationalization, and compliance. Plus matrix, review, examples, and cross-link integration. All content is docs-only markdown with illustrative code fences; no runtime/deployable artifacts.
### New Domains
- `domains/gitops-operators/` — platform-automation domain (ArgoCD + Flux + Operators)
- `first-principles.md` — 10 GitOps/operator principles (P1P10)
- Derived: `argocd.md`, `flux.md`, `operators.md`, `progressive-delivery.md`
- `domains/ai-ml/` — ML engineering domain
- `first-principles.md` — 10 AI/ML principles (P1P10)
- Derived: `data-versioning.md`, `model-evaluation.md`, `serving.md`, `monitoring-drift.md`
- `domains/i18n/` — internationalization domain
- `first-principles.md` — 10 i18n principles (P1P10)
- Derived: `locale-resources.md`, `formatting.md`, `rtl-bidi.md`, `testing-i18n.md`
- `domains/compliance/` — compliance/audit domain
- `first-principles.md` — 10 compliance principles (P1P10)
- Derived: `audit-logs.md`, `data-retention.md`, `policy-as-code.md`, `evidence.md`
### Cross-Domain Integration
- Extend `matrix/principles-matrix.md` with 40 new P-rules → core C-rule mappings (10 per new domain)
- Extend `matrix/domain-coverage.md` with the four new domains
- Extend `review/agent-checklist.md`, `review/peer-review-checklist.md`, and `review/anti-patterns.md` with new domain sections
- Update `MANIFEST.md` to list all new v0.3 documents (manifest is authoritative)
- New examples (good + bad): gitops-pr, ai-ml-reproducibility, i18n-string-concat, compliance-audit-log
- Cross-links from new domains to existing `devops/`, `security/`, `observability/`, `data/`, `kubernetes/`, `infrastructure-as-code/` domains
### Phase Plan (proposed, finalized in PLAN)
- P0 Pre-Execution: spec, clarify, research, ideate, plan, grill
- P1 GitOps + Operators domain
- P2 AI/ML domain
- P3 i18n + Compliance domains
- P4 Matrix + Review Integration (40 new mappings, manifest, checklist parity)
- P5 Examples + Cross-Links
- P6 Final Review + Ship (IS the v0.3 release → tag v0.2.6)
NFR milestone: no separate minor tag. The final patch (v0.2.6) IS the v0.3 deliverable.
## Key Decisions ## Key Decisions
| ID | Decision | Rationale | Confidence | | ID | Decision | Rationale | Confidence |
@@ -97,6 +139,31 @@ Build **Atelier** — a first-principles, docs-as-code engineering framework for
| D-013 | v0.2 tags run on v0.1.x patch line (prev minor from v0.2) | Per branch-strategy.md: milestone 0.2 → tags v0.1.0..v0.1.5; v0.1.5 IS the v0.2 release (NFR → no separate minor tag) | 0.90 | | D-013 | v0.2 tags run on v0.1.x patch line (prev minor from v0.2) | Per branch-strategy.md: milestone 0.2 → tags v0.1.0..v0.1.5; v0.1.5 IS the v0.2 release (NFR → no separate minor tag) | 0.90 |
| D-014 | Add phase-specific `platform-engineer` persona for P1P4 | IaC/k8s domain authoring benefits from a specialist persona with declarative-first/stateless-examples constraints; removed after milestone | 0.82 | | D-014 | Add phase-specific `platform-engineer` persona for P1P4 | IaC/k8s domain authoring benefits from a specialist persona with declarative-first/stateless-examples constraints; removed after milestone | 0.82 |
| D-015 | 4 execution phases (P1P4) + final phase P5 | P1 IaC domain, P2 k8s domain, P3 matrix+review, P4 examples+cross-links, P5 final review+ship | 0.85 | | D-015 | 4 execution phases (P1P4) + final phase P5 | P1 IaC domain, P2 k8s domain, P3 matrix+review, P4 examples+cross-links, P5 final review+ship | 0.85 |
| D-016 | v0.3 covers 4 deferred domains: gitops-operators, ai-ml, i18n, compliance | Carries forward v0.2 deferred ideation (IDEATE-15, IDEATE-16); single milestone groups them to limit release overhead | 0.86 |
| D-017 | v0.3 tags run on v0.2.x patch line (prev minor from v0.3) | Per branch-strategy.md: milestone 0.3 → tags v0.2.0..v0.2.6; v0.2.6 IS the v0.3 release (NFR → no separate minor tag) | 0.90 |
| D-018 | v0.3 splits P1 GitOps/Operators, P2 AI/ML, P3 i18n+Compliance, P4 Matrix+Review, P5 Examples, P6 Final | Each domain cluster is a coherent vertical slice; i18n + compliance paired (smaller surface) to balance phase load | 0.84 |
| D-019 | Reuse `platform-engineer` persona (extended) + add `ml-engineer` phase-specific persona for P2 | GitOps/Operators/k8s reuse platform-engineer; AI/ML benefits from a data/ML-specialist persona with reproducibility/data-lineage constraints; removed after milestone | 0.80 |
| D-020 | v0.3 remains docs-only (NFR milestone type) | PROJECT.md constraint "no runtime code" preserved; manifests/models/locale resources appear only as illustrative code-fence content in examples | 0.95 |
| D-021 | GitOps-operators domain groups ArgoCD + Flux + Operators + Progressive Delivery under one first-principles doc | All four share the declarative-source-of-truth reconciliation loop; splitting would fragment the P-rules and duplicate the core principles they trace to | 0.84 |
| D-022 | i18n + compliance paired in P3 (not separate phases) | Both are smaller-surface domains (4 derived docs each); pairing balances phase load against the heavier P1/P2 single-domain phases | 0.83 |
| D-023 | AI/ML domain scope = engineering discipline (data versioning, evaluation, serving, drift), NOT algorithm/model design | Atelier is a framework for engineering practice; algorithm choice is domain-knowledge out of scope. Mirrors how iac/k8s docs cover practice not implementation | 0.88 |
| D-024 | Compliance domain is framework-agnostic (audit logs, retention, policy-as-code, evidence), NOT tied to a specific regulation (GDPR/HIPAA/SOC2) | Regulation-specific docs would bloat the framework and go stale; principles derive from core Security/Correctness and apply across regulations | 0.86 |
| D-025 | Examples set = 2 good + 2 bad (not 4+4) | v0.3 adds 4 domains; 4+4 examples would unbalance P5. 2 good (gitops-pr, ai-ml-reproducibility) + 2 bad (i18n-string-concat, compliance-audit-log) cover the highest-illustration-value cases; remaining domains covered by cross-links and anti-patterns | 0.80 |
| D-026 | 40 new matrix mappings (10 per domain × 4 domains) | Consistent with v0.1 (110 mappings / 11 domains = 10) and v0.2 (20 mappings / 2 domains = 10). Each P-rule maps to ≥1 C-rule | 0.92 |
| D-035 | Add `examples/` directory listing to MANIFEST.md in v0.3 P4 (ATELIER-91) | v0.2 audit escalation ESC-002 note flagged examples/ unlisted; manifest is authoritative, so this is pre-existing drift that v0.3 closes | 0.85 |
| D-036 | Matrix coverage summary must state post-v0.3 totals (17 domains, 170 P-rules) | Both the summary block and per-domain section count must update; consistent with v0.2's "post-v0.2" summary | 0.93 |
| D-037 | domain-coverage.md Core Principle Coverage table (C1C8 → domains) must update for 4 new domains | ATELIER-81 covers the per-domain row schema; this is the complementary C-rule → domains table that also needs the 4 new domains | 0.90 |
| D-038 | v0.3 anti-patterns must pre-specify domain-specific violations + v0.3 artifact types | Avoids generic "deployable example artifact" only; v0.3 has new artifact types (.po, .rego, model files) and 4 domains × ~4 anti-patterns each | 0.86 |
| D-039 | ArgoCD vs Flux decision matrix required in argocd.md + flux.md | Parallel to v0.2 Helm vs Kustomize decision matrix (IDEATE-10); both tools share the GitOps model but differ in architecture (App CRD vs composable controllers) | 0.82 |
| D-040 | Data versioning tool comparison table required in data-versioning.md (DVC/Delta Lake/LakeFS) | Parallel to v0.2 state comparison table (IDEATE-11); three主流 tools with distinct versioning/lineage models | 0.80 |
| D-041 | Policy-as-code engine comparison table required in policy-as-code.md (OPA/Cedar/Kyverno/Sentinel) | Parallel to v0.2 PSS coverage (IDEATE-12); four engines with distinct policy languages and gate models | 0.81 |
| D-042 | GitOps push-pattern is a named anti-pattern (violates P3 Pull Don't Push) | Chaos scenario: a "GitOps" example that uses push-based deploy is a fundamental violation; pre-specify to reject on sight | 0.85 |
| D-043 | i18n LTR-only assumption is a named anti-pattern (violates P6 Text Direction) | Chaos scenario: formatting/layout examples that assume LTR only fail RTL/bidi users; pre-specify to reject | 0.83 |
| D-044 | compliance-audit-log bad example must cover both append-only violation (P1) and redaction failure (P9) | Two-breach example maximizes illustration value; mirrors v0.2 named-bad-example pattern but doubles the breach surface for the highest-stakes domain | 0.87 |
| D-045 | AI/ML orphan-model anti-pattern required (deployed prediction with no lineage trace, violates P3) | Chaos scenario: a serving example with no model→training→data lineage is the AI/ML analog of v0.2 orphaned P-rule; pre-specify | 0.84 |
| D-046 | i18n testing-i18n.md must map pseudo-locale testing to testing pyramid tiers | Avoids generic "test i18n" guidance; maps to unit (missing-key), integration (snapshot per locale), e2e (RTL coverage) | 0.78 |
| D-047 | compliance evidence.md must include a fenced signed-attestation example (Cosign or in-toto) | Prose-only evidence guidance is weak; a fenced example demonstrates the principle concretely (P6 Evidence Collected Continuously) | 0.80 |
| D-048 | ai-ml monitoring-drift.md must enumerate 3 drift types (data/concept/prediction) with a detection signal per type | Avoids conflating drift types; each has distinct detection signals and retraining triggers | 0.82 |
## Cross-Project References ## Cross-Project References
+106
View File
@@ -125,3 +125,109 @@ All 35 requirements covered. 8 core principles, 11 domains, 110 domain principle
| IDEATE-14 | backend-enriched | chaos | 0.87 | accepted → refines | ATELIER-53, ATELIER-51 (deployable artifact mitigation) | | IDEATE-14 | backend-enriched | chaos | 0.87 | accepted → refines | ATELIER-53, ATELIER-51 (deployable artifact mitigation) |
| IDEATE-15 | backend-enriched | improvement | 0.72 | deferred v0.3 | — (GitOps/operators domain) | | IDEATE-15 | backend-enriched | improvement | 0.72 | deferred v0.3 | — (GitOps/operators domain) |
| IDEATE-16 | backend-enriched | improvement | 0.68 | deferred v0.3 | — (ai-ml/i18n/compliance) | | IDEATE-16 | backend-enriched | improvement | 0.68 | deferred v0.3 | — (ai-ml/i18n/compliance) |
## v0.3 Requirements — GitOps + Operators + AI/ML + i18n + Compliance
**Milestone type:** NFR (all phases produce docs)
**Tag line:** v0.2.x (previous minor from v0.3)
| REQ-ID | Requirement | Priority | Phase | Status |
|--------|-------------|----------|-------|--------|
| ATELIER-60 | `domains/gitops-operators/first-principles.md` — 10 GitOps/operator principles (P1P10) | P0 | 1 | pending |
| ATELIER-61 | `domains/gitops-operators/argocd.md` — ArgoCD derived doc | P1 | 1 | pending |
| ATELIER-62 | `domains/gitops-operators/flux.md` — Flux derived doc | P1 | 1 | pending |
| ATELIER-63 | `domains/gitops-operators/operators.md` — Kubernetes Operators derived doc | P1 | 1 | pending |
| ATELIER-64 | `domains/gitops-operators/progressive-delivery.md` — progressive delivery derived doc | P1 | 1 | pending |
| ATELIER-65 | `domains/ai-ml/first-principles.md` — 10 AI/ML principles (P1P10) | P0 | 2 | pending |
| ATELIER-66 | `domains/ai-ml/data-versioning.md` — data/model versioning derived doc | P1 | 2 | pending |
| ATELIER-67 | `domains/ai-ml/model-evaluation.md` — evaluation derived doc | P1 | 2 | pending |
| ATELIER-68 | `domains/ai-ml/serving.md` — model serving derived doc | P1 | 2 | pending |
| ATELIER-69 | `domains/ai-ml/monitoring-drift.md` — monitoring/drift derived doc | P1 | 2 | pending |
| ATELIER-70 | `domains/i18n/first-principles.md` — 10 i18n principles (P1P10) | P0 | 3 | pending |
| ATELIER-71 | `domains/i18n/locale-resources.md` — locale resource management derived doc | P1 | 3 | pending |
| ATELIER-72 | `domains/i18n/formatting.md` — formatting (dates/numbers/units) derived doc | P1 | 3 | pending |
| ATELIER-73 | `domains/i18n/rtl-bidi.md` — RTL/bidi layout derived doc | P1 | 3 | pending |
| ATELIER-74 | `domains/i18n/testing-i18n.md` — i18n testing derived doc | P1 | 3 | pending |
| ATELIER-75 | `domains/compliance/first-principles.md` — 10 compliance principles (P1P10) | P0 | 3 | pending |
| ATELIER-76 | `domains/compliance/audit-logs.md` — audit logging derived doc | P1 | 3 | pending |
| ATELIER-77 | `domains/compliance/data-retention.md` — data retention derived doc | P1 | 3 | pending |
| ATELIER-78 | `domains/compliance/policy-as-code.md` — policy-as-code derived doc | P1 | 3 | pending |
| ATELIER-79 | `domains/compliance/evidence.md` — evidence collection derived doc | P1 | 3 | pending |
| ATELIER-80 | Extend `matrix/principles-matrix.md` with 40 new P-rules → core C-rule mappings (10 per new domain; review check: row count per domain = 10, each row ≥1 C-rule) | P0 | 4 | pending |
| ATELIER-81 | Extend `matrix/domain-coverage.md` with gitops-operators, ai-ml, i18n, compliance (row schema: domain, P-count, derived-doc-count, manifest-listed, status) | P1 | 4 | pending |
| ATELIER-82 | Extend `review/agent-checklist.md` with 4 new domain trigger sections | P1 | 4 | pending |
| ATELIER-83 | Extend `review/peer-review-checklist.md` with 4 new domain sections (parity with agent-checklist) | P1 | 4 | pending |
| ATELIER-84 | Extend `review/anti-patterns.md` with 4 new domain violations incl. orphaned P-rule + deployable example artifact | P1 | 4 | pending |
| ATELIER-85 | Update `MANIFEST.md` to list all new v0.3 documents (manifest authoritative) | P0 | 4 | pending |
| ATELIER-86 | `examples/good/gitops-pr.md` + `examples/good/ai-ml-reproducibility.md` — 2 good examples (markdown with fenced code only) | P2 | 5 | pending |
| ATELIER-87 | `examples/bad/i18n-string-concat.md` + `examples/bad/compliance-audit-log.md` — 2 named bad examples (each cites the P-rule breached) | P2 | 5 | pending |
| ATELIER-88 | Cross-links from new domains to existing devops/security/observability/data/kubernetes/infrastructure-as-code domains (review check: every new derived doc ≥1 outbound cross-link to a MANIFEST-listed doc) | P1 | 5 | pending |
| ATELIER-89 | Final review passes (all v0.3 phases reviewed, audit clean) | P0 | 6 | pending |
| ATELIER-90 | Milestone v0.3 released (tag v0.2.6, merged to main) | P0 | 6 | pending |
| ATELIER-91 | Add `examples/` directory listing to `MANIFEST.md` (pre-existing drift from v0.2 audit escalation ESC-002 note: examples/ unlisted; manifest is authoritative) | P1 | 4 | pending |
## v0.3 Traceability Matrix
| Phase | Requirements |
|-------|-------------|
| 0 (Pre-Execution) | (governance: spec, clarify, research, ideate, plan) |
| 1 (GitOps + Operators Domain) | ATELIER-60..ATELIER-64 |
| 2 (AI/ML Domain) | ATELIER-65..ATELIER-69 |
| 3 (i18n + Compliance Domains) | ATELIER-70..ATELIER-79 |
| 4 (Matrix + Review Integration) | ATELIER-80..ATELIER-85, ATELIER-91 |
| 5 (Examples + Cross-Links) | ATELIER-86..ATELIER-88 |
| 6 (Final Review + Ship) | ATELIER-89, ATELIER-90 |
## v0.3 Ideation Log
**Generated:** 14 ideas (mechanical: 5, backend-enriched: 7, within-project transfer: 2 merged)
**Accepted:** 14 (all v0.3-scope, confidence ≥ 0.78, above 0.6 autonomy threshold → auto-accepted)
**Deferred to v0.4:** 0
**Rejected:** 0
| IDEATE-ID | Source | Category | Confidence | Decision | Mapped REQ |
|-----------|--------|----------|------------|----------|------------|
| IDEATE-17 | mechanical (audit escalation ESC-002 note) | drift | 0.85 | accepted → new req | ATELIER-91 (examples/ in MANIFEST) |
| IDEATE-18 | mechanical (MANIFEST + matrix coverage summary) | coverage | 0.93 | accepted → refines | ATELIER-80, ATELIER-85 (v0.3 totals: 17 domains, 170 P-rules) |
| IDEATE-19 | mechanical (domain-coverage.md Core Principle Coverage table) | coverage | 0.90 | accepted → refines | ATELIER-81 (C-rule count updates for 4 new domains) |
| IDEATE-20 | mechanical (anti-patterns specificity) | quality | 0.86 | accepted → refines | ATELIER-84 (pre-specify domain anti-patterns + v0.3 artifact types: .po, .rego, model files) |
| IDEATE-21 | backend-enriched (v0.2 IDEATE-10 pattern transfer) | improvement | 0.82 | accepted → refines | ATELIER-61, ATELIER-62 (ArgoCD vs Flux decision matrix) |
| IDEATE-22 | backend-enriched (v0.2 IDEATE-11 pattern transfer) | improvement | 0.80 | accepted → refines | ATELIER-66 (data versioning tool comparison: DVC/Delta Lake/LakeFS) |
| IDEATE-23 | backend-enriched (v0.2 IDEATE-12 pattern transfer) | improvement | 0.81 | accepted → refines | ATELIER-78 (policy-as-code engine comparison: OPA/Cedar/Kyverno/Sentinel) |
| IDEATE-24 | backend-enriched | chaos | 0.85 | accepted → refines | ATELIER-84 (GitOps push-pattern anti-pattern, violates P3 Pull Don't Push) |
| IDEATE-25 | backend-enriched | chaos | 0.83 | accepted → refines | ATELIER-84, ATELIER-72 (i18n LTR-only assumption anti-pattern) |
| IDEATE-26 | backend-enriched | chaos | 0.87 | accepted → refines | ATELIER-87 (compliance-audit-log bad example must cover append-only violation + secret redaction failure, P1 + P9) |
| IDEATE-27 | backend-enriched | chaos | 0.84 | accepted → refines | ATELIER-84 (AI/ML orphan-model anti-pattern: deployed prediction with no lineage trace) |
| IDEATE-28 | backend-enriched | improvement | 0.78 | accepted → refines | ATELIER-74 (i18n testing-i18n.md pseudo-locale tier mapping to testing/pyramid) |
| IDEATE-29 | backend-enriched | improvement | 0.80 | accepted → refines | ATELIER-79 (compliance evidence.md signed attestation fenced example, Cosign/in-toto) |
| IDEATE-30 | backend-enriched | improvement | 0.82 | accepted → refines | ATELIER-69 (ai-ml monitoring-drift.md drift-type enumeration: data/concept/prediction with detection signals) |
### Refinements Notes (applied to existing reqs at execute time, not changing req rows)
- **ATELIER-80** (IDEATE-18): matrix coverage summary must read "post-v0.3: 17 domains, 170 P-rules"; update both the summary block and per-domain section count.
- **ATELIER-81** (IDEATE-19): the "Core Principle Coverage" table (C1C8 → domains) must be updated with the 4 new domains, not just the per-domain row schema table.
- **ATELIER-84** (IDEATE-20, IDEATE-24, IDEATE-25, IDEATE-27): anti-patterns extension must include (a) v0.3 deployable artifact types (.po resource files, .rego policy files, model artifacts, signed manifests as standalone files), (b) GitOps push-pattern violation (P3), (c) i18n LTR-only assumption violation (P6), (d) AI/ML orphan-model violation (P3 Lineage). Domain-specific anti-patterns to pre-specify:
- gitops-operators: push-based deploy (P3), manual kubectl apply on GitOps-managed resource (P8), cluster-admin GitOps robot (P10)
- ai-ml: unreproducible training run (P1), "the latest" model (P5), notebook in production (P9), orphan model with no lineage (P3)
- i18n: inline string concatenation (P3), `if (n == 1)` plural branching (P4), LTR-only layout assumption (P6), hand-rolled date formatter (P5)
- compliance: mutable audit log (P1), shared/generic identity in audit (P7), secret leaked in audit log (P9), manual evidence assembly at audit time (P6)
- **ATELIER-61/62** (IDEATE-21): argocd.md and flux.md must include an "ArgoCD vs Flux" decision matrix (parallel to v0.2 Helm vs Kustomize in ATELIER-46/47).
- **ATELIER-66** (IDEATE-22): data-versioning.md must include a tool comparison table (DVC vs Delta Lake vs LakeFS) covering versioning model, lineage, and use-case fit.
- **ATELIER-78** (IDEATE-23): policy-as-code.md must include an engine comparison table (OPA vs Cedar vs Kyverno vs Sentinel) covering policy language, evaluation gate, and ecosystem.
- **ATELIER-87** (IDEATE-26): the compliance-audit-log bad example must illustrate both an append-only violation (mutation/deletion of an audit record, P1) AND a redaction failure (secret in audit log, P9) — two breaches in one example.
- **ATELIER-74** (IDEATE-28): testing-i18n.md must map pseudo-locale testing to the testing pyramid tiers (unit: missing-key detection; integration: snapshot per locale; e2e: RTL coverage).
- **ATELIER-79** (IDEATE-29): evidence.md must include a fenced signed-attestation example (Cosign or in-toto), not prose-only.
- **ATELIER-69** (IDEATE-30): monitoring-drift.md must enumerate the three drift types (data drift, concept drift, prediction drift) with a detection signal per type.
### Within-Project Pattern Transfer (v0.1 → v0.2 → v0.3) — verified
| v0.2 Lesson | v0.3 Application | Status |
|-------------|------------------|--------|
| IDEATE-09 → ATELIER-59 (peer-review parity) | ATELIER-83 already covers this | ✓ carried forward |
| IDEATE-13/14 (chaos anti-patterns: orphan P-rule, deployable artifact) | ATELIER-84 + IDEATE-20/24/25/27 extend with v0.3-specific chaos | ✓ extended |
| IDEATE-08 (cross-link verification: every new derived doc ≥1 outbound cross-link) | ATELIER-88 already covers this | ✓ carried forward |
| IDEATE-02 (matrix row count = 10 per domain) | ATELIER-80 already covers this | ✓ carried forward |
| IDEATE-03 (domain-coverage row schema) | ATELIER-81 + IDEATE-19 extend with C-rule coverage table update | ✓ extended |
| IDEATE-07 (named bad examples cite P-rule breached) | ATELIER-87 + IDEATE-26 refine (two-breach example) | ✓ extended |
| IDEATE-10/11/12 (decision/comparison tables) | IDEATE-21/22/23 transfer the pattern to 3 v0.3 derived docs | ✓ transferred |
| v0.2 audit ESC-002 note (examples/ not in MANIFEST) | IDEATE-17 → ATELIER-91 | ✓ addressed |
+247
View File
@@ -216,3 +216,250 @@ See `.ciagent/atelier/PERSONAS.md` for the updated roster. v0.2 adds one phase-s
5. K8s derived docs mirror the k8s concept taxonomy: workloads, networking, storage, rbac, helm, kustomize. 5. K8s derived docs mirror the k8s concept taxonomy: workloads, networking, storage, rbac, helm, kustomize.
6. A phase-specific platform-engineer persona is warranted for P1P4; removed after v0.2. 6. A phase-specific platform-engineer persona is warranted for P1P4; removed after v0.2.
7. No runtime code; examples are illustrative markdown only. 7. No runtime code; examples are illustrative markdown only.
---
# v0.3 Research — GitOps + Operators + AI/ML + i18n + Compliance
> Research conducted during v0.3 P0 RESEARCH stage. Informs the four new domains, matrix extension (+40 mappings), and the two phase-specific personas (platform-engineer extended, ml-engineer added). See CLARIFY.md D-021..D-026 for resolved ambiguities and PROJECT.md D-016..D-026 for milestone decisions.
## Domain A: GitOps + Operators (ArgoCD, Flux, Operators, Progressive Delivery)
### Prior Art
- **CNCF OpenGitOps Principles v1.0.0** (GitOps Working Group, TAG App Delivery): the canonical 4 principles — **Declarative**, **Versioned and Immutable**, **Pulled Automatically**, **Continuously Reconciled**. Atelier's gitops-operators domain derives its first-principles from these plus the Operator pattern. ([opengitops.dev](https://opengitops.dev/), [github.com/open-gitops/documents](https://github.com/open-gitops/documents))
- **ArgoCD** (CNCF graduated): pull-based GitOps controller for k8s. Core concepts: Application CRD, sync waves, health/status assessment, diff against live cluster, RBAC, SSO. Declarative desired state from git; reconciled onto the cluster. ([argoCD.readthedocs.io](https://argoCD.readthedocs.io/))
- **Flux** (CNCF graduated): GitOps Toolkit — a set of composable controllers (source-controller, kustomize-controller, helm-controller, notification-controller). Pulls git/Helm/OCI sources, reconciles via kustomize/helm, emits events. Composable-controller architecture is a C6 (Composability) exemplar. ([fluxcd.io](https://fluxcd.io/))
- **Kubernetes Operator Pattern** (CNCF): a controller that encodes human operational knowledge as CRDs + control loops. Pattern documented in the k8s docs and "Operator Framework" (Operator SDK, OLM). Domain expertise as code; the deepest expression of k8s P1 Declarative Desired State. ([kubernetes.io/docs/concepts/extend-kubernetes/operator](https://kubernetes.io/docs/concepts/extend-kubernetes/operator/))
- **Progressive Delivery** — Argo Rollouts, Flagger: canary/blue-green traffic shifting driven by analysis (metrics, counters). Extends k8s rolling updates with metric-gated promotion. Cross-links devops/P5 Progressive Delivery.
- **Google SRE** (already in Atelier v0.1 observability/devops): reconciliation loops, error budgets, progressive rollout. Cross-cutting influence.
- **v0.2 in-tree prior art**: `kubernetes/first-principles.md` P1 (Declarative Desired State), P10 (Roll Forward Roll Back); `infrastructure-as-code/first-principles.md` P1 (Declarative Intent), P3 (State is Truth), P9 (Drift is Recoverable). GitOps-operators is the deployment-automation layer above these.
### Principles Identified for `gitops-operators/first-principles.md` (P1P10)
Each derived from a core C-rule (matrix extensions in P4):
1. **P1 Git is the Source of Truth** — desired state lives in a versioned, immutable git store; the cluster is a derivative, not an authority. (C1 Correctness, C5 Reversibility)
2. **P2 Declarative Over Imperative** — express desired cluster state, not the commands to reach it. (C2 Clarity, C3 Simplicity)
3. **P3 Pull, Don't Push** — agents running inside the target pull desired state; no outside push credentials into the cluster. (C1 Correctness via security, C4 Locality)
4. **P4 Continuous Reconciliation** — the loop is the primitive; drift is detected and corrected automatically, not on-demand. (C7 Observability, C1 Correctness)
5. **P5 State is Immutable and Versioned** — every change is a commit; history is the audit trail and the rollback path. (C5 Reversibility)
6. **P6 Operators Encode Domain Knowledge** — operational expertise lives as CRDs + controllers, not runbooks that humans must remember. (C6 Composability, C2 Clarity)
7. **P7 Progressive Delivery is Reversible by Construction** — canary/blue-green are staged, metric-gated, and one-command abortable. Promotion without a rollback path is a violation. (C5 Reversibility, C1 Correctness)
8. **P8 Reconcile, Don't Mutate by Hand** — manual `kubectl apply`/`kubectl edit` on a GitOps-managed resource is an incident; drift back to git is the recovery. (C1 Correctness, C7 Observability)
9. **P9 Failure is Observable and Surfaced** — sync failures, health degradation, and rollout-stall events emit status + notifications; silent drift is the bug. (C7 Observability)
10. **P10 Least Privilege Reconciliation** — the controller's credentials are scoped to the namespaces/resources it reconciles; no cluster-admin GitOps robots. (C1 Correctness via security, C8 Economy of trust)
### Derived Docs
- `argocd.md` — Application CRD, App-of-Apps, sync waves, health checks, diffs, RBAC/SSO, multi-cluster, sync windows.
- `flux.md` — GitOps Toolkit controllers (source, kustomize, helm, notification), composable architecture, HR/Kustomization/HelmRelease CRDs, OCI sources.
- `operators.md` — Operator pattern, CRDs, controllers, Operator SDK/OLM, when to write an operator vs a Helm chart, scope/responsibility boundaries.
- `progressive-delivery.md` — Argo Rollouts + Flagger, canary/blue-green, analysis templates (metrics, counters), abort/rollback, cross-link devops/P5.
### Cross-Domain Links (one-directional in v0.3, per D-026 extended)
- `kubernetes/P1 Declarative Desired State` ← gitops P2
- `kubernetes/P10 Roll Forward Roll Back` ← gitops P7
- `infrastructure-as-code/P1 Declarative Intent` ← gitops P2
- `infrastructure-as-code/P3 State is Truth` ← gitops P1, P5
- `infrastructure-as-code/P9 Drift is Recoverable` ← gitops P4, P8
- `devops/P1 Reproducibility` ← gitops P1, P5
- `devops/P4 Rollback First` ← gitops P5, P7
- `devops/P5 Progressive Delivery` ← gitops P7
- `devops/P6 Configuration as Code` ← gitops P1, P2
- `security/secrets` ← gitops P3, P10 (reconciliation credentials)
- `security/supply-chain` ← gitops P5 (signed/immutable manifest provenance)
- `observability/metrics` ← gitops P4, P9 (reconciliation + rollout metrics)
## Domain B: AI / ML (Engineering Discipline)
### Prior Art
- **Google MLOps / "Hidden Technical Debt in ML Systems"** (Sculley et al., 2015): the foundational paper framing ML systems as software-engineering problems with debt surfaces (data dependencies, configuration, glue code, reproducibility). Atelier's ai-ml domain is the principles-layer response.
- **DVC / Data Version Control** (iterative.ai): git for data + pipelines; treats datasets, features, and models as versioned artifacts. C5 (Reversibility) and C6 (Composability) exemplar.
- **MLflow** (Linux Foundation): experiment tracking, model registry, model packaging, deployment stages. Tracking → registry → serving lifecycle.
- **Kubeflow** (CNCF): k8s-native ML pipelines, training operators, serving (KServe). Brings ML onto the k8s reconciliation model (cross-link kubernetes).
- **KServe / Seldon Core / BentoML**: model serving runtimes; inference as a scalable, observable service. Cross-link devops/P7 Immutability, observability/metrics.
- **Evidently AI / Great Expectations**: data drift detection, data quality, model monitoring. C7 (Observability) for ML.
- **"Machine Learning Operations (MLOps)"** frameworks — Microsoft MLOps, AWS MLOps, Google MLOps maturity model. Converge on: version data, track experiments, evaluate models, serve reproducibly, monitor drift.
- **v0.2 in-tree prior art**: `kubernetes/first-principles.md` (serving on k8s), `infrastructure-as-code/` (training pipelines as declarative infra), `data/` (schema, migrations — data versioning analog).
### Principles Identified for `ai-ml/first-principles.md` (P1P10)
Scope per D-023: engineering discipline (data versioning, evaluation, serving, drift), NOT algorithm/model design. Each derived from a core C-rule:
1. **P1 Reproducibility is the First Class** — every training run is reproducible from pinned data + code + config + environment. Unreproducible runs are unreviewable. (C1 Correctness, C5 Reversibility)
2. **P2 Data is Versioned, Not Just Code** — datasets, features, and splits are first-class versioned artifacts with lineage; `git` alone is insufficient. (C5 Reversibility, C7 Observability)
3. **P3 Lineage is Traceable End-to-End** — any deployed prediction traces back through model → training run → dataset → source. No orphan models. (C7 Observability, C1 Correctness)
4. **P4 Evaluation is Defined Before Training** — metrics, splits, and thresholds are declared a priori; cherry-picking metrics post-hoc is a correctness violation. (C1 Correctness, C2 Clarity)
5. **P5 Models are Versioned Artifacts** — a model is a pinned, immutable, registry-tracked artifact with a unique identifier; never "the latest." (C5 Reversibility, C6 Composability)
6. **P6 Serving is Observable** — inference latency, throughput, input distributions, and prediction confidence are first-class signals. Silent serving is a bug. (C7 Observability)
7. **P7 Drift is Expected and Detected** — data drift, concept drift, and prediction drift are monitored; a drift signal is an incident, not a curiosity. (C7 Observability, C1 Correctness)
8. **P8 Inference Inputs are Validated** — the model's contract (schema, ranges, types) is enforced at the serving boundary; out-of-contract inputs are rejected, not silently scored. (C1 Correctness via security/input-validation)
9. **P9 Pipelines Compose, Notebooks Don't** — training/serving flows are composable pipelines with explicit steps and contracts; notebooks are for exploration, not production. (C6 Composability, C2 Clarity)
10. **P10 Rollback Includes the Model** — a serving rollback restores the prior model artifact, not just the prior code; promotion is reversible at the model layer. (C5 Reversibility)
### Derived Docs
- `data-versioning.md` — DVC/Delta Lake/LakeFS patterns, data lineage, dataset hashing, train/val/test split versioning, cross-link data/migrations.
- `model-evaluation.md` — metric selection, offline/online eval, holdout integrity, bias/fairness checks (engineering angle), eval as a gate.
- `serving.md` — KServe/Seldon/BentoML, inference as a service, batching, latency SLAs, canarying models, cross-link kubernetes + devops.
- `monitoring-drift.md` — Evidently/Great Expectations, drift types (data/concept/prediction), alerting, retraining triggers, cross-link observability/metrics.
### Cross-Domain Links (one-directional in v0.3)
- `data/migrations` ← ai-ml P2 (data versioning ↔ migration discipline)
- `data/schema-design` ← ai-ml P8 (inference input contract)
- `observability/metrics` ← ai-ml P6, P7
- `observability/logging` ← ai-ml P3 (lineage)
- `devops/P1 Reproducibility` ← ai-ml P1
- `devops/P7 Immutability` ← ai-ml P5 (model images)
- `devops/P5 Progressive Delivery` ← ai-ml P10 (model canary)
- `security/input-validation` ← ai-ml P8
- `security/secrets` ← ai-ml P8 (serving credentials)
- `performance/backend` ← ai-ml P6 (serving latency)
- `kubernetes/workloads` ← ai-ml P9 (serving on k8s)
## Domain C: Internationalization (i18n)
### Prior Art
- **Unicode / ICU / CLDR** (Unicode Consortium): the foundation — ICU (International Components for Unicode) for formatting/collation, CLDR (Common Locale Data Repository) for locale data. The de-facto source for date/number/currency/plural/relative-time formatting. ([unicode.org/cldr](https://cldr.unicode.org/), [icu.unicode.org](https://icu.unicode.org/))
- **W3C Internationalization** (W3C i18n WG): the canonical web i18n guidance — "Internationalization techniques", "Language tags in HTML and XML", bidi/RTL authoring. Cross-links WCAG for accessibility-of-locale. ([w3.org/International](https://www.w3.org/International/))
- **RFC 5646 / BCP 47** — language tags (`en-US`, `ar-EG`, `zh-Hans-CN`). The locale identifier standard.
- **RFC 9229 / RFC 9230** (and earlier BCP 47 extensions) — Unicode locale extensions (`-u-`).
- **gettext / ICU MessageFormat / FormatJS / react-intl / i18next / Fluent (Mozilla)** — message-format libraries; ICU MessageFormat is the cross-ecosystem baseline for plural/gender/select. Fluent pioneered "localization 2.0" with asymmetric translations.
- **JavaScript Intl API** — browser-native formatting built on ICU/CLDR; the runtime baseline.
- **WCAG 2.1 AA** (already in Atelier uiux/accessibility): cross-cutting — locale support is an a11y concern for non-Latin-script users; RTL layout is a UI-correctness concern.
- **Google i18n + Mozilla L10n guides** — operational practice (string extraction, pseudo-locale testing, RTL testing).
- **v0.2/v0.1 in-tree prior art**: `uiux/` (accessibility, components, copywriting — i18n's consumer), `testing/` (fixtures, pyramid — i18n testing parallels), `api/error-responses` (localized API errors).
### Principles Identified for `i18n/first-principles.md` (P1P10)
Each derived from a core C-rule:
1. **P1 Source Language is a Locale, Not the Default** — the developer's language is one locale among many, not the "neutral" form. Strings are extracted from day one. (C2 Clarity, C1 Correctness)
2. **P2 Locale Identifiers are Standardized** — use BCP 47 language tags; no ad-hoc locale codes. (C2 Clarity, C6 Composability)
3. **P3 Resources are External, Not Inline** — user-facing strings live in locale resource files, never concatenated inline in code. (C4 Locality, C6 Composability)
4. **P4 Plural and Gender are Parameterized** — use ICU MessageFormat (or equivalent) for plural/gender/select; never `if (n == 1)` branching. (C1 Correctness, C6 Composability)
5. **P5 Formatting is Locale-Aware** — dates, times, numbers, currencies, units via ICU/CLDR/`Intl`; never hand-rolled formatters. (C1 Correctness, C7 Observability of format correctness)
6. **P6 Text Direction is a Layout Primitive** — RTL/bidi is a first-class layout concern, not a CSS afterthought; logical properties (`start`/`end`) over physical (`left`/`right`). (C1 Correctness, C4 Locality)
7. **P7 Layout Accommodates Expansion** — translated text expands/contracts; layouts are flexible (no fixed pixel widths for text). (C8 Economy of rework, C3 Simplicity)
8. **P8 Pseudo-Locales Test Early** — test with pseudo-locales (accented, lengthened, RTL-mirrored) before real translations arrive. (C7 Observability, C5 Reversibility of finding bugs late)
9. **P9 Images and Icons are Cultural** — icons, colors, and imagery are locale-sensitive; avoid locale-bound symbols as universal. (C1 Correctness, C2 Clarity)
10. **P10 Translation is Reversible and Versioned** — resource files are versioned; a bad translation is a rollback, not a hot-patch. (C5 Reversibility)
### Derived Docs
- `locale-resources.md` — resource file formats (.po/.pot, JSON, Fluent FTL, ICU Resource Bundle), key naming, namespaces, fallback chains, extraction tooling.
- `formatting.md` — ICU/CLDR/`Intl` for dates, times, numbers, currencies, units, relative time, plural rules; BCP 47 tags; cross-link api/error-responses for localized errors.
- `rtl-bidi.md` — logical vs physical CSS properties, bidi algorithm (UAX #9), `dir` attribute, mirroring, common pitfalls (icons, numbers in RTL), cross-link uiux/components + uiux/accessibility.
- `testing-i18n.md` — pseudo-locales, snapshot testing per locale, RTL coverage, missing-key detection, cross-link testing/fixtures + testing/pyramid.
### Cross-Domain Links (one-directional in v0.3)
- `uiux/accessibility` ← i18n P6 (RTL/bidi is an a11y concern for non-Latin users)
- `uiux/components` ← i18n P6, P7
- `uiux/copywriting` ← i18n P1, P3
- `testing/fixtures` ← i18n P8
- `testing/pyramid` ← i18n P8
- `api/error-responses` ← i18n P5 (localized error messages)
- `data/schema-design` ← i18n P2, P3 (locale data shapes)
## Domain D: Compliance (Audit, Retention, Policy-as-Code, Evidence)
### Prior Art
**Note (D-024):** the compliance domain is framework-agnostic — it abstracts regulation-specific requirements (GDPR, HIPAA, SOC 2, PCI-DSS, NIST 800-53, ISO 27001) into engineering principles. No regulation-specific docs; they would bloat the framework and go stale.
- **NIST Cybersecurity Framework (CSF) / NIST 800-53** — controls catalog (audit, retention, evidence, policy). Atelier abstracts the *principles*, not the controls.
- **SOC 2 (AICPA) Trust Services Criteria** — Security, Availability, Processing Integrity, Confidentiality, Privacy. Audit logs, retention, and evidence are explicit criteria.
- **GDPR / CCPA** — data subject rights, retention limits, lawful basis. Abstracted to "retention is a function of policy, not storage."
- **OWASP AppSec / ASVS** — already in Atelier security domain; compliance extends to auditability of security controls.
- **Open Policy Agent (OPA) / Rego, Cedar (AWS), HashiCorp Sentinel, Kyverno** — policy-as-code engines; policy evaluated as a gate, not a document. C6 (Composability) + C1 (Correctness) exemplars. ([openpolicyagent.org](https://www.openpolicyagent.org/), [kyverno.io](https://kyverno.io/))
- **Cosign / Sigstore / in-toto** — signed attestations and provenance; evidence-as-artifact. Cross-link security/supply-chain.
- **Google Cloud Audit Logs / AWS CloudTrail / Azure Activity Log** — the canonical audit-log patterns; immutable, append-only, queryable, time-ordered.
- **v0.2/v0.1 in-tree prior art**: `security/` (authorization, secrets, supply-chain), `observability/` (logging, metrics, tracing — audit logs are structured logging), `data/` (schema, migrations — retention schema), `infrastructure-as-code/` (policy-as-code parallels declarative IaC), `kubernetes/` (rbac — audit subject identity).
### Principles Identified for `compliance/first-principles.md` (P1P10)
Each derived from a core C-rule. Framework-agnostic per D-024:
1. **P1 Audit Logs are Append-Only** — audit records are immutable once written; deletion or mutation is itself an auditable incident. (C1 Correctness, C5 Reversibility)
2. **P2 Every Significant Action is Logged** — the set of auditable actions is defined a priori; "we forgot to log it" is a violation. Auth changes, data access, config changes, policy changes. (C7 Observability, C1 Correctness)
3. **P3 Retention is Policy, Not Storage** — data lifetime is declared and enforced; deletion at end-of-life is a feature, not a failure. (C5 Reversibility, C8 Economy of storage)
4. **P4 Policy is Code** — compliance policy is expressed in versioned, reviewable, testable code (OPA/Cedar/Kyverno), not in spreadsheets or prose. (C6 Composability, C2 Clarity)
5. **P5 Policy is Evaluated as a Gate** — policy violations block before the action, not after the audit; admission/CI/CD-time enforcement. (C1 Correctness, C5 Reversibility)
6. **P6 Evidence is Collected Continuously** — evidence of compliance (logs, configs, scans, attestations) is gathered as a byproduct of operation, not assembled manually at audit time. (C7 Observability, C3 Simplicity of audit)
7. **P7 Identity is Attributable** — every logged action traces to an authenticated principal; shared/generic identities are violations. (C1 Correctness via security, C7 Observability)
8. **P8 Subject Access is Honored** — data-subject rights (access, export, deletion) are operations with defined contracts and audit trails; not ad-hoc. (C1 Correctness, C5 Reversibility)
9. **P9 Secrets and Sensitive Data are Redacted in Audit** — audit logs themselves must not leak secrets; redaction is structural, not opportunistic. (C1 Correctness via security, C3 Simplicity)
10. **P10 Compliance Posture is Observable** — the system reports its own compliance state (drift from policy, open violations, retention status); silent non-compliance is the bug. (C7 Observability, C1 Correctness)
### Derived Docs
- `audit-logs.md` — append-only log patterns, structured audit events, CloudTrail/Cloud-Audit-Log conventions, queryability, retention of logs themselves, cross-link observability/logging + security/authorization.
- `data-retention.md` — retention policies as code, lifecycle rules, deletion as a feature, GDPR/CCPA abstracted, retention vs. backup distinction, cross-link data/migrations.
- `policy-as-code.md` — OPA/Cedar/Sentinel/Kyverno patterns, policy as a CI/CD + admission gate, policy testing, versioning policy, cross-link infrastructure-as-code (declarative intent) + kubernetes (admission).
- `evidence.md` — evidence collection as a byproduct, signed attestations (Cosign/in-toto), audit-ready export, provenance, cross-link security/supply-chain + observability/metrics.
### Cross-Domain Links (one-directional in v0.3)
- `security/authorization` ← compliance P7 (attributable identity)
- `security/secrets` ← compliance P9 (redaction)
- `security/supply-chain` ← compliance P6, evidence.md (signed attestations)
- `observability/logging` ← compliance P1, P2 (audit logs = structured logging)
- `observability/metrics` ← compliance P10 (compliance posture metrics)
- `observability/tracing` ← compliance P6 (evidence from distributed traces)
- `data/schema-design` ← compliance P3 (retention schema)
- `data/migrations` ← compliance P3 (retention migration discipline)
- `infrastructure-as-code/P1 Declarative Intent` ← compliance P4 (policy-as-code)
- `infrastructure-as-code/P3 State is Truth` ← compliance P10 (compliance posture truth)
- `kubernetes/rbac` ← compliance P7 (audit subject identity)
- `devops/P6 Configuration as Code` ← compliance P4 (policy as code)
## Architectural Fit (v0.1/v0.2 Contract Preservation)
- **Hierarchy preserved:** all four new domains depend on `core/`; their P-rules trace to C1C8 via the matrix. No lateral authority.
- **10 P-rules per domain** (per D-018, D-030, D-026): consistent with v0.1 (11 domains) and v0.2 (2 domains). v0.3 adds 40 new P-rules → matrix grows 130 → 170.
- **Manifest authoritative:** all new documents added to `MANIFEST.md` in P4. Unlisted = not part of the framework.
- **No runtime code** (per D-020, PROJECT.md constraint): examples are illustrative markdown with code fences only. No `.yaml` manifests, `.po` resource files, model artifacts, policy `.rego` files, or deployable artifacts as standalone files — only fenced code blocks inside `.md` files.
- **Conflict resolution unchanged:** matrix extended, not replaced. Core precedence (C1 > C2 > ... > C8) governs any new vs existing rule conflict. Compliance rules tracing to C1 (Correctness) inherit C1's non-tradeable status where they overlap with security (per core/conflict-resolution.md §6).
- **Cross-links one-directional** (D-026 extended): new domains link outward to existing; existing domains unchanged in v0.3 (no back-link edits to v0.1/v0.2 content).
## Prior Art Position (v0.3 extension)
Existing GitOps/AI-ML/i18n/compliance guidance (OpenGitOps principles, ArgoCD/Flux docs, Operator pattern, MLOps maturity models, ICU/CLDR, W3C i18n, NIST/SOC 2, OPA/Kyverno) state practices and controls but none map every domain rule back to a small set of universal core principles. Atelier's v0.3 contribution is the same differentiation as v0.1 and v0.2: **traceable principle hierarchy with a join table**. The four new domains add 40 P-rules, each traced to ≥1 core C-rule, extending the matrix from 130 to 170 domain principles across 13 → 17 domains.
## v0.3 Persona Assessment
See `.ciagent/atelier/PERSONAS.md` for the updated roster. v0.3 adds two phase-specific personas (per D-019, D-020):
- **platform-engineer** (phase-specific, extended from v0.2): domain = infrastructure/platform-automation; territory = `domains/gitops-operators/**`, gitops examples; constraints add "source-of-truth is git" and "reconciliation loop is the primitive"; active for P1 GitOps/Operators only; removed after v0.3 completes.
- **ml-engineer** (phase-specific, new): domain = machine-learning engineering; territory = `domains/ai-ml/**`, `examples/good/ai-ml-reproducibility.md`; constraints = ["reproducibility is non-negotiable", "data lineage is traceable", "trace to core", "10 P-rules per domain", "no runtime code", "engineering discipline not algorithm design (D-023)"]; active for P2 AI/ML only; removed after v0.3 completes.
- **i18n (P3) + compliance (P3)** covered by tech-writer + domain-expert (D-022 — no new personas; both domains are smaller-surface and within the existing personas' competence).
## v0.3 Risks and Mitigations
| Risk | Mitigation |
|------|-----------|
| New P-rules orphaned from core (no matrix trace) | P4 extends matrix; domain-expert persona verifies every new P-rule traces to ≥1 C-rule before sign-off |
| GitOps-operators overlaps kubernetes/infrastructure-as-code (declarative, state, drift) | Cross-links one-directional (D-026); each domain owns its angle (k8s P1 desired-state vs gitops P1 git-as-source-of-truth vs iac P3 state-is-truth) |
| AI/ML domain drifts into algorithm/model-design (out of scope per D-023) | ml-engineer persona constraint "engineering discipline not algorithm design"; review/agent-checklist gains an ai-ml scope check in P4 |
| Compliance domain bloats into regulation-specific docs (GDPR/SOC2) | D-024 framework-agnostic; review check in P4 rejects regulation-specific content |
| i18n and compliance overlap on "retention of locale data" | Each owns its angle: i18n P10 (translation versioning) vs compliance P3 (data retention policy) |
| Examples become runtime artifacts (model files, .rego, .po) | persona constraints "no runtime code"; examples are markdown with fenced code only; P5 review check |
| Persona explosion (5 active in v0.3) | Both new personas are phase-specific and removed post-milestone; roster returns to 3 |
| Matrix row-count verification (40 new mappings, 10 per domain) | D-026 review check: row count per domain = 10, each row ≥1 C-rule, executed in P4 |
## v0.3 Conclusions
1. Four new top-level domains extend the framework without breaking the v0.1/v0.2 contract.
2. 40 new P-rules (10 per domain) all trace to core C1C8 — matrix extends from 130 to 170 across 13 → 17 domains.
3. GitOps-operators unifies ArgoCD/Flux/Operators/Progressive Delivery under the shared declarative-source-of-truth reconciliation loop (D-021) — splitting would fragment the P-rules.
4. AI/ML is scoped to engineering discipline (D-023): data versioning, evaluation, serving, drift — NOT algorithm design. Reproducibility and lineage are the non-negotiables.
5. i18n is grounded in ICU/CLDR + BCP 47 + W3C i18n; the source language is a locale, not a default.
6. Compliance is framework-agnostic (D-024): audit/retention/policy-as-code/evidence abstract NIST/SOC2/GDPR into principles that derive from core Security/Correctness/Observability.
7. Two phase-specific personas (platform-engineer extended, ml-engineer added); both removed post-v0.3.
8. No runtime code; examples are illustrative markdown only.
+42 -1
View File
@@ -76,9 +76,50 @@ NFR milestone: no separate minor tag. The final patch (v0.1.5) IS the v0.2 deliv
- 2 deferred to v0.3 (GitOps/operators domain; ai-ml/i18n/compliance domains) - 2 deferred to v0.3 (GitOps/operators domain; ai-ml/i18n/compliance domains)
- See `.ciagent/atelier/REQUIREMENTS.md` "v0.2 Ideation Log" for the full table - See `.ciagent/atelier/REQUIREMENTS.md` "v0.2 Ideation Log" for the full table
## Milestone: v0.3 — GitOps + Operators + AI/ML + i18n + Compliance (ACTIVE)
**Milestone type:** NFR (all phases produce docs — no `feat` code)
**Tag line:** v0.2.x (previous minor from v0.3)
**Phases:** P0 (pre-execution) + P1P5 (execution) + P6 (final review+ship)
| Phase | Name | Type | Status | Key Deliverables |
|-------|------|------|--------|------------------|
| 0 | Pre-Execution | docs | complete | Spec, clarify, research, ideate, plan, PERSONAS.md (extends platform-engineer, adds ml-engineer) — shipped v0.2.0 |
| 1 | GitOps + Operators Domain | docs | complete | domains/gitops-operators/{first-principles, argocd, flux, operators, progressive-delivery}.md — shipped v0.2.1 |
| 2 | AI/ML Domain | docs | complete | domains/ai-ml/{first-principles, data-versioning, model-evaluation, serving, monitoring-drift}.md — shipped v0.2.2 |
| 3 | i18n + Compliance Domains | docs | complete | domains/i18n/{first-principles, locale-resources, formatting, rtl-bidi, testing-i18n}.md, domains/compliance/{first-principles, audit-logs, data-retention, policy-as-code, evidence}.md — shipped v0.2.3 |
| 4 | Matrix + Review Integration | docs | complete | matrix/principles-matrix.md (+40 mappings), matrix/domain-coverage.md (incl. C-rule coverage table update), review/{agent-checklist, peer-review-checklist, anti-patterns}.md, MANIFEST.md (+ examples/ listing per ATELIER-91) — shipped v0.2.4 |
| 5 | Examples + Cross-Links | docs | pending | examples/good + examples/bad for 4 domains, cross-links to devops/security/observability/data/k8s/iac |
| 6 | Final Review + Ship | docs | pending | Review passed, audit clean, milestone merged to main, tag v0.2.6 |
## v0.3 Phase Tag Mapping
Per branch-strategy.md, milestone `v0.3` tags run on the `v0.2.x` patch line:
| Phase | Tag | Notes |
|-------|-----|-------|
| P0 | v0.2.0 | Pre-execution release |
| P1 | v0.2.1 | GitOps + Operators domain |
| P2 | v0.2.2 | AI/ML domain |
| P3 | v0.2.3 | i18n + Compliance domains |
| P4 | v0.2.4 | Matrix + review integration |
| P5 | v0.2.5 | Examples + cross-links |
| P6 | v0.2.6 | Final review + ship — **IS the v0.3 milestone release** |
NFR milestone: no separate minor tag. The final patch (v0.2.6) IS the v0.3 deliverable.
## v0.3 Ideation Outcome
- 14 ideas generated (mechanical 5, backend-enriched 7, within-project transfer 2 merged)
- 14 accepted (all v0.3-scope, confidence ≥ 0.78, above 0.6 autonomy threshold → auto-accepted)
- 1 new requirement added: ATELIER-91 (examples/ in MANIFEST — pre-existing drift from v0.2 audit escalation)
- 13 refinements to existing reqs ATELIER-61..90 (decision matrices, chaos anti-patterns, drift-type enumeration, etc.)
- 0 deferred to v0.4
- See `.ciagent/atelier/REQUIREMENTS.md` "v0.3 Ideation Log" for the full table
## Future Milestones ## Future Milestones
- **v0.3** (candidates from ideation): `domains/gitops-operators/` (ArgoCD, Flux), `domains/ai-ml/`, `domains/i18n/`, `domains/compliance/` per spec Part 6 and v0.2 deferred ideation. - **v0.4** (candidates): `domains/edge/`, `domains/quantum/`, language-specific derived docs, tooling adapters (linters), translation/localization of framework docs.
## Success Criteria ## Success Criteria
+2 -2
View File
@@ -3,8 +3,8 @@
{ {
"slug": "atelier", "slug": "atelier",
"name": "Atelier", "name": "Atelier",
"milestone": "v0.2", "milestone": "v0.3",
"status": "complete" "status": "active"
} }
], ],
"active_project": "atelier", "active_project": "atelier",
+35 -5
View File
@@ -36,13 +36,43 @@
| DevOps | ✓ | ci-cd, environments | | DevOps | ✓ | ci-cd, environments |
| Infrastructure as Code | ✓ | terraform, opentofu, state, modules | | Infrastructure as Code | ✓ | terraform, opentofu, state, modules |
| Kubernetes | ✓ | workloads, networking, storage, rbac, helm, kustomize | | Kubernetes | ✓ | workloads, networking, storage, rbac, helm, kustomize |
| GitOps + Operators | ✓ | argocd, flux, operators, progressive-delivery |
| AI / ML | ✓ | data-versioning, model-evaluation, serving, monitoring-drift |
| i18n | ✓ | locale-resources, formatting, rtl-bidi, testing-i18n |
| Compliance | ✓ | audit-logs, data-retention, policy-as-code, evidence |
## Examples
> Examples are illustrative markdown with fenced code only (no standalone runtime artifacts per D-020 / D-025). The `examples/` directory listing closes the v0.2 ESC-002 drift (IDEATE-17, ATELIER-91). P5 authored the v0.3 examples and promoted all entries from `pending` to `✓` (verified — every listed file exists).
| Path | Status | Notes |
|------|--------|-------|
| `examples/good/` | ✓ | Good-example directory — 8 examples (v0.1 + v0.2 + v0.3) |
| `examples/bad/` | ✓ | Bad-example directory — 7 examples (v0.1 + v0.2 + v0.3) |
| `examples/good/api-endpoint.md` | ✓ | v0.1 example — good REST endpoint |
| `examples/good/react-component.md` | ✓ | v0.1 example — good React component |
| `examples/good/db-schema.md` | ✓ | v0.1 example — good DB schema |
| `examples/good/error-handler.md` | ✓ | v0.1 example — good error handler |
| `examples/bad/god-object.md` | ✓ | v0.1 example — bad god object |
| `examples/bad/silent-error.md` | ✓ | v0.1 example — bad silent error |
| `examples/bad/leaky-abstraction.md` | ✓ | v0.1 example — bad leaky abstraction |
| `examples/good/terraform-module.md` | ✓ | v0.2 example — good IaC module |
| `examples/good/k8s-deployment.md` | ✓ | v0.2 example — good k8s deployment |
| `examples/bad/terraform-unlocked-state.md` | ✓ | v0.2 example — bad unlocked state |
| `examples/bad/k8s-bare-pod-no-resources.md` | ✓ | v0.2 example — bad bare pod |
| `examples/good/gitops-pr.md` | ✓ | v0.3 example — good GitOps PR |
| `examples/good/ai-ml-reproducibility.md` | ✓ | v0.3 example — good reproducible training run |
| `examples/bad/i18n-string-concat.md` | ✓ | v0.3 example — bad i18n string concat |
| `examples/bad/compliance-audit-log.md` | ✓ | v0.3 example — bad audit log (P1 + P9 breaches) |
> **Note:** The `examples/` section was established in P4 with entries pre-listed as `pending P5`. P5 authored the 4 v0.3 examples and promoted all entries to `✓` after verifying every listed file exists on disk. The manifest remains authoritative — unlisted = not part of framework.
## Cross-Cutting ## Cross-Cutting
| Document | Purpose | | Document | Purpose |
|-----------------------------------|----------------------------------| |-----------------------------------|----------------------------------|
| `matrix/principles-matrix.md` | Maps domain → core principles (13 domains, 130 P-rules post-v0.2) | | `matrix/principles-matrix.md` | Maps domain → core principles (17 domains, 170 P-rules post-v0.3) |
| `matrix/domain-coverage.md` | Maps core → domains; per-domain coverage | | `matrix/domain-coverage.md` | Maps core → domains; per-domain coverage (incl. v0.3 Core Principle Coverage) |
| `review/agent-checklist.md` | Pre-completion agent checklist (incl. IaC + k8s triggers) | | `review/agent-checklist.md` | Pre-completion agent checklist (incl. IaC + k8s + gitops + ai-ml + i18n + compliance triggers) |
| `review/peer-review-checklist.md` | Human peer-review checklist (incl. IaC + k8s sections) | | `review/peer-review-checklist.md` | Human peer-review checklist (incl. IaC + k8s + gitops + ai-ml + i18n + compliance sections) |
| `review/anti-patterns.md` | Catalog of violations (incl. IaC + k8s + chaos anti-patterns) | | `review/anti-patterns.md` | Catalog of violations (incl. IaC + k8s + gitops + ai-ml + i18n + compliance + v0.3 chaos anti-patterns) |
+90
View File
@@ -0,0 +1,90 @@
# Data Versioning — Derived Rules
> Derives from `domains/ai-ml/first-principles.md`. Covers P2 (Data is
> Versioned, Not Just Code) and P3 (Lineage is Traceable End-to-End).
> Referenced by `serving.md` and `monitoring-drift.md`. Scope per
> D-023: engineering discipline of versioning data, not dataset
> content design.
## Why Data Versioning (P2 Data is Versioned, Not Just Code)
- `git` versions code well and data badly. Datasets do not fit in
git, and a dataset is not recovered from a commit hash.
- A model trained on "the data" is a model trained on an unknown
input — a C1 (Correctness) violation. The dataset is a build
input; it is named, hashed, and recoverable the way any build
input is.
- Data versioning is the ML analogue of `domains/data/migrations.md`:
the schema and contents of the data evolve, every evolution is a
versioned migration, and every model points at a specific version.
## Dataset Hashing and Lineage (P3 Lineage Traceable End-to-End)
- Every dataset version has a content hash (not a filename or a
timestamp). The hash is the identity. A model's lineage record
names the dataset hash it was trained on; a serving prediction
names the model digest it came from.
- Lineage is a graph: prediction → model → training run → dataset →
source(s). Any edge missing is an orphan (`domains/observability/logging.md`
for the structured-log angle on lineage events).
- The lineage record is append-only. Editing it to "fix" a broken
trace is the same class of violation as editing an audit log.
## Train/Val/Test Split Versioning (P2, P4 Eval Defined Before Training)
- Splits are versioned with the dataset, not derived ad-hoc per run.
A split is a deterministic function of (dataset version, split
config, random seed). Two runs on the same pinned inputs produce
the same splits.
- The eval split is held out and never touched by training. A "held
out" set that leaked into training is a P4 (Evaluation Defined
Before Training) violation, not just a P2 violation — the eval
gate is measuring the training set, not the model.
- Cross `domains/data/schema-design.md` for the eval input contract:
the schema of the eval set is part of the versioned artifact.
## Tool Comparison (IDEATE-22, D-040)
| Tool | Versioning Model | Lineage | Best For | Notes |
|------|------------------|---------|----------|-------|
| DVC | Git-like pointers to content-addressed object store; `.dvc` files in git track data versions | Pipeline DAG in `dvc.yaml`; reproducibility via `dvc repro` | Teams already on git; file/directory datasets; ML pipelines | Treats data like code; shares git's history model. Object store is pluggable (S3, GCS, Azure, SSH) |
| Delta Lake | Table format with transaction log (ACID) + time travel via versioned commits; schema enforcement | Time travel queries; lineage via table history + catalog | Large tabular data; lakehouse; streaming + batch on the same table | Not a pipeline tool — pairs with Spark/Trino/Flink. Brings DB guarantees to object storage |
| LakeFS | Git-like operations (branch, commit, merge) over object storage itself | Branch model gives isolated, reproducible data branches | Data engineering teams; branch-per-experiment; CI over data | Not a table format — versions objects. Composes with Delta/Iceberg on top |
- Pick one primary versioning model per platform. Mixing DVC's
pointer model with Delta's transaction-log model fragments
operational knowledge (C4 Locality).
- All three satisfy P2; the choice is which fits the data shape and
the team's existing tooling. None is advocated over the others.
## Reproducibility Contract (P1 Reproducibility is the First Class)
A reproducible training run records, in one versioned place:
```
run_id: 2026-08-05T09:12:00Z#run-42
dataset: s3://ml-data/train@sha256:7f3a...e21
splits: dvc.yaml@commit a1b2c4d
code: git@a1b2c4d
config: configs/train.yaml@commit a1b2c4d
environment: ghcr.io/org/train-img@sha256:9c2d...f88
eval_spec: configs/eval.yaml@commit a1b2c4d
model_digest: registry/model@sha256:b5e1...aa0
```
- Lose any line and the run is anecdote, not evidence.
- The record is the lineage root: a prediction cites the
`model_digest`, which cites the `run_id`, which cites everything
above. This is how P3 (Lineage Traceable End-to-End) is satisfied
in practice.
## What Violates Data Versioning Discipline
| Violation | Principle |
|-----------|-----------|
| Dataset referenced by `s3://bucket/latest/` | P2 Data is Versioned, Not Just Code |
| Splits regenerated with an unpinned seed per run | P2, P4 Evaluation Defined Before Training |
| A production model with no dataset hash in its lineage | P3 Lineage Traceable End-to-End |
| Editing a lineage record to "clean up" a broken trace | P3 Lineage Traceable End-to-End |
| Eval split reachable from the training data path | P4 Evaluation Defined Before Training |
| Two platforms versioning the same data with different models | C4 Locality |
+154
View File
@@ -0,0 +1,154 @@
# AI / ML — First Principles
> Scope per D-023: this domain covers ML **engineering discipline** —
> data versioning, evaluation methodology, serving patterns, and drift
> detection. It does **not** cover algorithm design, model architecture
> selection, hyperparameter tuning, or model-family comparison. Those
> are research choices, not engineering principles, and they have no
> derivation in the core C-rules.
## 1. The Principles
### P1. Reproducibility is the First Class
Every training run is reproducible from pinned data + code + config +
environment. An unreproducible run is an unreviewable run: you cannot
decide whether a result is correct if you cannot recreate it.
Reproducibility is the ML analogue of `domains/devops/P1
Reproducibility` and inherits its non-negotiable status. Lose any one
of data, code, config, or environment pinning, and the run is
anecdote, not evidence.
### P2. Data is Versioned, Not Just Code
Datasets, features, and train/val/test splits are first-class
versioned artifacts with content hashes and lineage. `git` alone is
insufficient — datasets do not fit in git, and a dataset is not a
commit hash. A model trained on "the data" is a model trained on an
unknown input, which is a correctness violation. Version data the way
you version code: pinned, named, and recoverable.
### P3. Lineage is Traceable End-to-End
Any deployed prediction traces back through model → training run →
dataset → source. No orphan models. A model in production with no
lineage is a correctness defect: you cannot reason about its failure
modes, you cannot roll it back to a known-good dataset, and you cannot
tell whether drift is in the model or in the data that built it.
Lineage is the audit trail of ML (`domains/observability/logging.md`).
### P4. Evaluation is Defined Before Training
Metrics, splits, and acceptance thresholds are declared a priori, in
code, before the model is trained. Cherry-picking metrics post-hoc is
a correctness violation: the evaluation is no longer measuring the
model, it is rationalizing it. The eval spec is a contract — it is
reviewable, it is versioned, and it is the gate the model must pass
before it leaves the experiment. This is the ML angle on C2 Clarity:
the intent of the model is obvious to its reader because the eval
declared it first.
### P5. Models are Versioned Artifacts
A model is a pinned, immutable, registry-tracked artifact with a
unique identifier. Never "the latest." A serving endpoint that pulls
"latest" is serving an unknown model — its behavior is undefined, its
rollback is impossible, and its lineage is broken. The model registry
is to models what a container registry is to images
(`domains/devops/P7 Immutability`): immutable, addressed by digest,
promoted by stage.
### P6. Serving is Observable
Inference latency, throughput, input distributions, and prediction
confidence are first-class signals. Silent serving is a bug. A model
in production that emits no metrics is a model you cannot operate: you
cannot see latency regressions, you cannot see input drift, you cannot
see a failing downstream consumer. Observability is designed in, not
bolted on (`domains/observability/metrics.md`).
### P7. Drift is Expected and Detected
Data drift, concept drift, and prediction drift are monitored as a
matter of course. A drift signal is an incident, not a curiosity. ML
systems decay without code changes — the world changes under the
model — so "no code changed" is not a defense against a serving
regression. Detecting drift is the ML-specific form of C7
Observability: you cannot fix a model you cannot see degrading.
### P8. Inference Inputs are Validated
The model's input contract — schema, value ranges, types, and
categorical domains — is enforced at the serving boundary.
Out-of-contract inputs are rejected, not silently scored. Scoring an
out-of-contract input is a correctness violation: the model's output
is undefined for inputs outside its training distribution, and
returning a number for it is lying to the caller. This is the ML angle
on `domains/security/input-validation.md` and inherits C1's
non-tradeable status.
### P9. Pipelines Compose, Notebooks Don't
Training and serving flows are composable pipelines with explicit
steps, named inputs, named outputs, and contracts between stages.
Notebooks are for exploration, not production. A notebook in the
serving path is a correctness defect: its state is implicit, its
order is human-dependent, and its reproducibility is whatever the last
operator remembered. Compose pipelines; keep notebooks in the lab.
### P10. Rollback Includes the Model
A serving rollback restores the prior model artifact, not just the
prior code. Promotion is reversible at the model layer. A rollback
that redeploys old code but keeps the new model has not rolled back —
the model was the thing that regressed. The rollback path must name
the prior model digest, the prior dataset version, and the prior eval
that cleared it. This is the ML angle on `domains/devops/P4 Rollback
First` and `domains/kubernetes/P10 Roll Forward, Roll Back`.
## 2. Core Principle Trace
Each AI/ML P-rule derives from one or more core C-rules (C1C8). The
matrix extension lands in P4 of the v0.3 plan; the traces below are
authoritative.
| P-rule | Core | Why |
|--------|------|-----|
| P1 Reproducibility is the First Class | C1, C5 | Correctness of results; reversibility of runs |
| P2 Data is Versioned, Not Just Code | C5, C7 | Reversibility of datasets; observability of data lineage |
| P3 Lineage is Traceable End-to-End | C7, C1 | Observability of provenance; correctness of attribution |
| P4 Evaluation is Defined Before Training | C1, C2 | Correctness of the eval gate; clarity of a-priori intent |
| P5 Models are Versioned Artifacts | C5, C6 | Reversibility of model identity; composability of registry stages |
| P6 Serving is Observable | C7 | Observability of inference |
| P7 Drift is Expected and Detected | C7, C1 | Observability of degradation; correctness of detection |
| P8 Inference Inputs are Validated | C1 | Correctness of the serving boundary (security subset) |
| P9 Pipelines Compose, Notebooks Don't | C6, C2 | Composability of stages; clarity of explicit contracts |
| P10 Rollback Includes the Model | C5 | Reversibility at the model layer |
## 3. What Violates These Principles
| Violation | Principle Breached |
|-----------|-------------------|
| A training run that cannot be replayed from pinned inputs | P1 Reproducibility is the First Class |
| A dataset referenced by a mutable path, not a hash | P2 Data is Versioned, Not Just Code |
| A production model with no record of its training data | P3 Lineage is Traceable End-to-End |
| Metrics chosen after seeing the results | P4 Evaluation is Defined Before Training |
| A serving endpoint that pulls `latest` from the registry | P5 Models are Versioned Artifacts |
| A model in production with no latency or throughput metrics | P6 Serving is Observable |
| A serving regression dismissed as "no code changed" | P7 Drift is Expected and Detected |
| An input with an out-of-range feature scored silently | P8 Inference Inputs are Validated |
| A notebook in the serving or training pipeline path | P9 Pipelines Compose, Notebooks Don't |
| A rollback that restores code but keeps the regressed model | P10 Rollback Includes the Model |
## 4. Relationship to Other Domains
AI/ML is the engineering-discipline layer for model-bearing systems.
It borrows the reproducibility, immutability, rollback, and
observability disciplines of `domains/devops/` and applies them to
the data → model → serving lifecycle. Cross-links are one-directional
(per D-026 extended):
- `domains/devops/P1 Reproducibility` ← P1
- `domains/devops/P4 Rollback First` ← P10
- `domains/devops/P5 Progressive Delivery` ← P10 (model canary)
- `domains/devops/P7 Immutability` ← P5 (model images)
- `domains/data/migrations.md` ← P2 (data versioning ↔ migration discipline)
- `domains/data/schema-design.md` ← P8 (inference input contract)
- `domains/observability/metrics.md` ← P6, P7
- `domains/observability/logging.md` ← P3 (lineage)
- `domains/security/input-validation.md` ← P8
- `domains/security/secrets.md` ← P8 (serving credentials)
- `domains/performance/backend.md` ← P6 (serving latency)
- `domains/kubernetes/workloads.md` ← P9 (serving on k8s)
- `domains/testing/first-principles.md` ← P4 (eval as a gate)
- `domains/gitops-operators/first-principles.md` ← P10 (model rollback in a GitOps loop)
+94
View File
@@ -0,0 +1,94 @@
# Model Evaluation — Derived Rules
> Derives from `domains/ai-ml/first-principles.md`. Covers P4
> (Evaluation is Defined Before Training) and the eval-as-a-gate
> discipline. Referenced by `serving.md` (promotion gate) and
> `monitoring-drift.md` (online eval). Scope per D-023: evaluation
> methodology, not metric math or model-family benchmarks.
## Evaluation is a Gate, Not a Report (P4 Evaluation Defined Before Training)
- The eval spec — metrics, splits, thresholds, and pass/fail
criteria — is declared in code **before** the model is trained.
It is versioned with the data and the code; it is reviewable; it
is the contract the model must satisfy to leave the experiment.
- Cherry-picking metrics after seeing results is a correctness
violation: the eval is no longer measuring the model, it is
rationalizing it. The a-priori spec is what makes the eval
trustworthy.
- This is the ML angle on `domains/testing/first-principles.md` P1
(Tests as Specification): the eval declares the model's contract,
the model does not declare its own success.
## The Eval Input Contract (P8 Inference Inputs are Validated, cross `domains/data/schema-design.md`)
- The eval set has a schema: feature names, types, ranges, and
categorical domains. That schema is the same schema the serving
boundary enforces (`serving.md`, `domains/security/input-validation.md`).
- An eval set whose schema drifted from the serving schema is
measuring a different model than the one in production. Schema
parity is part of the versioned eval artifact.
- Cross `domains/data/schema-design.md`: the eval input contract is
a schema-design problem, versioned and reviewed like any schema.
## Holdout Integrity (P4, P2 Data is Versioned)
- The held-out eval set is never touched by training, feature
selection, or threshold tuning. A "held out" set that influenced
any training decision is not held out — it is a third training
set, and the eval is measuring memorization.
- Splits are versioned with the dataset (`data-versioning.md`).
Recreating splits ad-hoc per run breaks comparability across runs.
- Reusing a held-out set across many model iterations leaks it
incrementally. Rotate or re-split on a cadence; record the
rotation in lineage.
## Offline vs Online Evaluation (P6 Serving is Observable)
- **Offline eval** runs before promotion: held-out data, pinned
model, declared metrics, pass/fail gate. It answers "should this
model ship?"
- **Online eval** runs after promotion, on live traffic: shadow
scoring, A/B, canary metrics. It answers "is this model behaving
in production?" It is the bridge to `monitoring-drift.md`.
- A model that passed offline and regressed online is not a
contradiction — it is a signal that the offline distribution
differs from the live one (a P7 drift signal). Both eval layers
are required; neither substitutes for the other.
## Bias and Fairness Checks (Engineering Angle, P4)
- Bias/fairness checks are part of the a-priori eval spec, not an
afterthought. They are metrics with thresholds, declared before
training, gated the same as any metric.
- This doc covers the **engineering** discipline: the checks are
versioned, gated, and recorded in lineage. The choice of which
fairness metrics and what thresholds are policy decisions, not
engineering principles, and are out of scope here (D-023).
## Eval-as-a-Gate in the Pipeline (P9 Pipelines Compose)
- The eval is a pipeline stage with a contract: input = model
digest + eval dataset version; output = pass/fail + metric
report. It composes with the training stage and the promotion
stage.
- A promotion that bypasses the eval stage is a P4 violation,
regardless of who approved it. The gate is in the pipeline, not
in a human sign-off sheet.
```
train -> eval(gate) -> register(promote) -> serve
|
+-- fail -> abort, no promote
```
## What Violates Evaluation Discipline
| Violation | Principle |
|-----------|-----------|
| Metrics chosen after seeing the scores | P4 Evaluation Defined Before Training |
| Held-out set used in feature selection or threshold tuning | P4, P2 |
| Eval schema differs from serving schema | P8 Inference Inputs are Validated |
| Promotion by human approval, bypassing the eval stage | P4, P9 Pipelines Compose |
| A "passing" model with no online eval in production | P6 Serving is Observable |
| Fairness checks added after a model shipped | P4 Evaluation Defined Before Training |
+88
View File
@@ -0,0 +1,88 @@
# Monitoring & Drift — Derived Rules
> Derives from `domains/ai-ml/first-principles.md`. Covers P7 (Drift
> is Expected and Detected) and the online half of P6 (Serving is
> Observable). Referenced by `serving.md` (online eval) and
> `model-evaluation.md` (online layer). Scope per D-023: drift
> detection methodology, not model retraining architecture.
## Drift is Expected and Detected (P7 Drift is Expected and Detected)
- ML systems decay without code changes. The world changes under
the model: user behavior shifts, input pipelines change,
upstream schemas evolve. "No code changed" is not a defense
against a serving regression.
- A drift signal is an incident, not a curiosity. It triggers an
alert, an investigation, and a decision (retrain, roll back, or
accept with a recorded justification). Silent drift is the same
class of bug as silent serving (P6).
- Cross `domains/observability/metrics.md` for the alerting
primitives and `domains/observability/logging.md` for the
structured events a drift signal emits.
## The Three Drift Types (IDEATE-30, D-048)
| Drift Type | What Changes | Detection Signal | Source of Truth |
|------------|--------------|------------------|-----------------|
| **Data drift** (input drift) | The distribution of inputs at serving time diverges from the distribution the model was trained on | Statistical distance between the live input distribution and the pinned training-set distribution (e.g., PSI, KL, KS test). Alert on threshold breach | Training dataset hash (`data-versioning.md`) + live input metrics |
| **Concept drift** | The relationship between inputs and the target changes — the same input now maps to a different correct output | Ground-truth lag: compare delayed labels against predictions on the same inputs. Rising error rate against a stable input distribution signals concept, not data, drift | Delayed-label feedback stream + prediction log |
| **Prediction drift** (output drift) | The distribution of the model's predictions shifts, with no change to inputs | Statistical distance between the live prediction distribution and a pinned baseline prediction distribution. Independent of inputs — catches model-internal regressions and upstream silent changes | Prediction log + baseline prediction snapshot |
- The three signals are distinct and non-substitutable. Data drift
catches the input changing; concept drift catches the world
changing; prediction drift catches the model's behavior changing.
A monitoring setup with only one is blind to two classes of
regression.
- Evidently AI and Great Expectations are the canonical tooling:
Evidently for drift/statistical reports, Great Expectations for
data-quality/contract checks at the pipeline boundary. Both
produce the metrics that feed `domains/observability/metrics.md`.
## Detection Signals in Practice
- **Data drift** compares live inputs to the **pinned training
distribution** — not to "yesterday's inputs." Without a pinned
baseline, drift is measured against a moving target and is
meaningless. Cross `data-versioning.md` for how the baseline is
pinned.
- **Concept drift** requires ground truth, which is often delayed
(days/weeks). The detection signal is the gap between
prediction-time confidence and delayed-label error. A rising
error against stable inputs is the signature.
- **Prediction drift** needs no ground truth and no input
comparison — it watches the model's own output distribution. It
is the cheapest signal and the first to fire; it is also the
least specific (any of the three drifts can move predictions).
## Alerting and Retraining Triggers (P7, P10 Rollback Includes the Model)
- A drift alert is an incident. It does not auto-trigger retraining
unsupervised — auto-retraining on drift can lock in a bad
distribution. The alert triggers a human decision: investigate,
retrain, roll back, or accept.
- Retraining is a new training run (`first-principles.md` P1): it
produces a new model digest, passes the eval gate
(`model-evaluation.md`), and is promoted through the registry
(`serving.md`). The prior model stays rollbackable (P10).
- Cross `domains/observability/metrics.md` for the alert-rule
pattern: threshold + window + severity, routed to the same
on-call path as any production incident.
## Online Evaluation Bridge (P6 Serving is Observable)
- Online eval (`model-evaluation.md`) is the live counterpart to
drift monitoring: shadow scores and A/B canaries measure a
candidate model against the incumbent, while drift monitoring
measures the incumbent against its own baseline. Both feed the
same metrics pipeline.
## What Violates Monitoring Discipline
| Violation | Principle |
|-----------|-----------|
| Only one drift type monitored | P7 Drift is Expected and Detected |
| Drift baseline is "yesterday's inputs," not pinned training data | P7, P2 Data is Versioned |
| Drift alert that auto-retrains without a human gate | P7, P1 Reproducibility |
| A serving regression dismissed as "no code changed" | P7 Drift is Expected and Detected |
| Concept-drift check with no delayed-label feedback path | P7 Drift is Expected and Detected |
| Prediction-distribution change with no alert | P6 Serving is Observable, P7 |
+88
View File
@@ -0,0 +1,88 @@
# Serving — Derived Rules
> Derives from `domains/ai-ml/first-principles.md`. Covers P5 (Models
> are Versioned Artifacts), P6 (Serving is Observable), P8 (Inference
> Inputs are Validated), and P10 (Rollback Includes the Model).
> Referenced by `monitoring-drift.md` (online signals) and
> `model-evaluation.md` (promotion gate). Scope per D-023: serving
> patterns, not model architectures.
## The Model is an Addressed Artifact (P5 Models are Versioned Artifacts)
- A serving endpoint pulls a model by digest, never by `latest`. A
model pulled by `latest` is an unknown model — its behavior is
undefined and its rollback is impossible.
- The model registry is to models what a container registry is to
images (`domains/devops/P7 Immutability`): immutable, addressed by
digest, promoted by stage (staging → prod). Promotion is a
registry operation, not a file copy.
- A serving rollout names the model digest in its manifest. The
digest is part of the deploy's lineage (`data-versioning.md`).
## Inference Inputs are Validated (P8 Inference Inputs are Validated)
- The model's input contract — schema, types, ranges, categorical
domains — is enforced at the serving boundary, before the model
sees the input. Out-of-contract inputs are rejected with a
defined error, not silently scored.
- Scoring an out-of-contract input is a C1 (Correctness) violation:
the model's output is undefined outside its training
distribution, and returning a number for it is lying to the
caller.
- This is the ML angle on `domains/security/input-validation.md`:
the validation lives at the boundary, the model is downstream of
it, and the contract is versioned with the model.
## Serving is Observable (P6 Serving is Observable)
- Every inference path emits: request latency, throughput, input
distribution summaries, prediction confidence, and error counts.
Silent serving is a bug.
- Cross `domains/observability/metrics.md` for the metrics
primitives (histograms, counters, gauges) and
`domains/observability/tracing.md` for the request-level trace
that ties an input to a prediction.
- Latency SLAs are enforced via `domains/performance/backend.md`
disciplines: budget the inference path, measure the tail (p99),
alert on budget breach.
## Serving Patterns (P9 Pipelines Compose)
| Pattern | When | Notes |
|---------|------|-------|
| Inference as a service | Default; model behind an HTTP/gRPC endpoint | KServe, Seldon Core, BentoML. Scales with traffic; model is a deployable, addressable artifact |
| Batch inference | Offline scoring of large datasets | No latency SLA; throughput-bound. Same model digest, same input contract |
| Embedded / in-process | Latency-critical, single-tenant | Model linked into the app. Trades observability for latency — only when the SLA demands it |
- Canarying a model is a serving pattern, not a deployment pattern:
shift a fraction of traffic to the new model digest, measure
online eval (`model-evaluation.md`), abort to the prior digest on
regression. This is `domains/devops/P5 Progressive Delivery`
applied at the model layer.
- Rollback restores the prior model digest (P10 Rollback Includes
the Model). A rollback that redeploys old code but keeps the new
model has not rolled back. Cross `domains/gitops-operators/first-principles.md`
for the GitOps reconciliation loop that drives model rollouts.
## Tool Landscape (KServe / Seldon Core / BentoML)
| Tool | Model Packaging | Deployment Surface | Notes |
|------|-----------------|--------------------|-------|
| KServe | InferenceService CRD; runtime predictors (v2, HuggingFace, PMML, custom) | Kubernetes-native; CRD-driven | Cross `domains/kubernetes/workloads.md`. Brings the k8s reconciliation model to serving |
| Seldon Core | SeldonDeployment CRD; graph of predictors | Kubernetes-native; CRD-driven | Emphasizes inference graphs (fan-out, ensemble) as CRD structure |
| BentoML | Bento (model + runtime + deps packaged); Yatai registry | Kubernetes or bare container | Focuses on packaging + registry; the Bento is the versioned artifact (P5) |
- All three satisfy P5/P6/P8 when wired correctly; the choice is
packaging model and deployment surface, not correctness.
- None is advocated over the others.
## What Violates Serving Discipline
| Violation | Principle |
|-----------|-----------|
| Endpoint pulls `latest` from the registry | P5 Models are Versioned Artifacts |
| Out-of-range input scored silently | P8 Inference Inputs are Validated |
| Serving path emits no latency or throughput metrics | P6 Serving is Observable |
| Rollback redeploys code but keeps the regressed model | P10 Rollback Includes the Model |
| A notebook in the serving path | P9 Pipelines Compose, Notebooks Don't |
| Canary with no abort-to-prior-digest path | P10, `domains/devops/P5 Progressive Delivery` |
+165
View File
@@ -0,0 +1,165 @@
# Audit Logs — Derived Rules
> Derives from `domains/compliance/first-principles.md`. Covers P1
> (Audit Logs are Append-Only), P2 (Every Significant Action is
> Logged), P7 (Identity is Attributable), P9 (Secrets Redacted in
> Audit), and P10 (Compliance Posture Observable). Referenced by
> `data-retention.md` (retention applies to audit logs themselves)
> and `evidence.md` (audit logs are evidence).
## Audit Logs are Append-Only (P1 Audit Logs are Append-Only)
- An audit record is immutable once written. The storage substrate
enforces this; policy alone does not. Write-once, append-only
sinks (WORM buckets, immutable log streams, hash-chained ledgers)
are the mechanism.
- Deletion or mutation of an audit record is itself an auditable
incident. The tampering is the signal, not just the underlying
event. A system that allows `DELETE FROM audit_log` is a system
whose audit log is a draft.
- The append-only guarantee is testable: attempt to write, then
attempt to overwrite, then attempt to delete. If the overwrite or
delete succeeds, the guarantee is absent and the design is a
violation.
## Structured Audit Events (P2 Every Significant Action is Logged)
- The set of auditable actions is defined a priori, in code, before
the action ships. The catalog is versioned and reviewed. An
auditable action with no log line is a violation, not a gap to
backfill later.
- Audit events are structured (JSON / protobuf / a typed schema),
not prose. A prose log line ("user logged in") is unqueryable and
unaggregatable; a structured event is both. The event schema is
the contract between the producer and the audit pipeline.
```
{
"timestamp": "2024-11-07T15:03:22Z",
"event": "auth.login",
"actor": { "kind": "user", "id": "u_8f3a", "session": "s_12b9" },
"action": "succeeded",
"target": { "kind": "service", "id": "billing-api" },
"source": { "ip": "203.0.113.42", "region": "us-east-1" },
"request_id": "req_91c2",
"version": "audit-schema/v2"
}
```
- The catalog of significant actions typically includes:
authentication (success and failure), authorization decisions
(allow and deny), data access (read, write, delete), configuration
changes, policy changes, retention executions, and admin
operations. The exact set is declared per system; the discipline
is that it is declared.
## Cloud Audit Log Conventions (Prior Art, Abstracted)
- AWS CloudTrail, Google Cloud Audit Logs, and Azure Activity Log
share a common shape: immutable, time-ordered, queryable, with
actor / action / target / source / result fields. Atelier's
audit-logs doc adopts the shape, not the vendor.
- The shape is the contract; the sink is the implementation. A
self-hosted audit log that follows the same shape composes with
the same tooling (SIEM, query engines, evidence exporters) as the
cloud vendors'.
## Queryability (P10 Compliance Posture Observable)
- An audit log that cannot be queried is an audit log that cannot be
used. Queryability is a first-class design goal: the event schema
is typed, fields are indexed, and the common queries (who acted on
what when, what failed, what was denied) are cheap.
- "Who did X between T1 and T2" must be a single query, not a
forensics project. If the query requires a custom script per
investigation, the audit log is structured for storage, not for
use — a C7 (Observability) violation.
## Identity is Attributable (P7 Identity is Attributable)
- Every audit event records the authenticated principal that acted —
not a shared account, not a generic service, not "admin." The
actor field is populated at the time of the action from the
authenticated session, not resolved after the fact.
- A shared account in the actor field breaks accountability: an
event attributed to `svc-deploy` could be any of ten engineers.
This is the compliance angle on `domains/security/authorization.md`
and `domains/kubernetes/rbac.md`: bind actions to unique
principals, not to roles many can assume.
- Machine-to-machine actions record the workload identity (a service
account, a signed instance identity), not a human — but the
identity is still unique and attributable to a deployable unit.
## Redaction at the Boundary (P9 Secrets Redacted in Audit)
- Audit logs must not leak secrets, credentials, tokens, or PII.
Redaction is structural: applied at the logging boundary, before
the record is written to the append-only sink — not opportunistic
scrubbing after the fact. Once a secret is in an append-only log,
the remediation is expensive (rotate, rewrite access scope), so
redaction-at-source is the only sound position.
- The redaction policy is itself auditable: which fields are
redacted, by what rule, in which event type. A redaction rule
that lives in someone's head is a P9 violation waiting to happen.
```
// before redaction (DO NOT LOG)
{
"event": "config.read",
"target": { "kind": "secret", "id": "db-password" },
"value": "p@ssw0rd-plaintext-leaked" // VIOLATION
}
// after structural redaction
{
"event": "config.read",
"target": { "kind": "secret", "id": "db-password" },
"value": "[REDACTED:secret]",
"redaction": "secret-value-policy/v1"
}
```
- Never log request bodies, response bodies, headers like
`Authorization`, or environment variables that may carry secrets.
Log the *fact* of the action, not the *content* of the secret.
## Retention of Audit Logs Themselves
- Audit logs are subject to retention policy (cross `data-
retention.md`), but the floor is set by the accountability need,
not by storage economy. An audit log deleted before its retention
period is a P1 violation dressed as a P3 action.
- The retention rule for audit logs is itself logged (meta-audit):
when an audit log segment ages out and is deleted, the deletion is
recorded in a higher-tier audit log with the rule that authorized
it. The chain is observable end to end.
## What Violates Audit-Log Discipline
| Violation | Principle |
|-----------|-----------|
| Audit log on a mutable filesystem with no write-once protection | P1 Audit Logs are Append-Only |
| `DELETE FROM audit_log WHERE timestamp < ...` as routine cleanup | P1 Audit Logs are Append-Only |
| An auth-success event with no audit record | P2 Every Significant Action is Logged |
| A prose log line ("user did a thing") instead of a structured event | P2 Every Significant Action is Logged |
| A shared `admin` account as the actor in audit events | P7 Identity is Attributable |
| An `Authorization: Bearer <token>` header logged in plaintext | P9 Secrets and Sensitive Data are Redacted in Audit |
| A redaction rule applied inconsistently across event types | P9 Secrets and Sensitive Data are Redacted in Audit |
| "Who did X?" requires a custom forensics script per investigation | P10 Compliance Posture is Observable |
| An audit log segment deleted with no meta-audit record | P1 Audit Logs are Append-Only |
## Relationship to Other Domains
- `domains/observability/logging.md` — audit logs are structured
logging with an append-only guarantee; the logging primitives
(levels, structured fields, correlation IDs) compose here.
- `domains/security/authorization.md` — the actor in an audit event
is the principal the authorization layer authenticated.
- `domains/security/secrets.md` — redaction at the logging boundary
is the audit-side complement of secret management.
- `domains/compliance/data-retention.md` — retention policy applies
to audit logs; the audit log's own deletion is meta-audited.
- `domains/compliance/evidence.md` — audit logs are a primary
evidence artifact; the append-only guarantee is what makes them
admissible.
- `domains/kubernetes/rbac.md` — workload identity in audit events
derives from the RBAC principal that acted.
+159
View File
@@ -0,0 +1,159 @@
# Data Retention — Derived Rules
> Derives from `domains/compliance/first-principles.md`. Covers P3
> (Retention is Policy, Not Storage) and the data-shape angle on P8
> (Subject Access is Honored). Referenced by `audit-logs.md`
> (retention applies to audit logs) and `evidence.md` (evidence has
> a retention lifecycle). Framework-agnostic per D-024 — no
> regulation-specific retention periods.
## Retention is Policy, Not Storage (P3 Retention is Policy, Not Storage)
- Data lifetime is declared and enforced as policy, in code — not
left to the storage layer's defaults. The policy names what data
class is retained for how long, what action fires at end-of-life
(delete, archive, anonymize), and what exception path exists (a
legal hold suspends deletion).
- Deletion at end-of-life is a feature, not a failure. A system that
cannot delete on schedule is a system that over-retains, which is
the symmetric violation of a system that under-retains. Both are
P3 violations; the policy is the arbiter.
- "We kept it because the bucket was cheap" is a violation. "We
deleted it because the policy said to" is correct. Cost does not
override policy; policy is the contract.
## Retention Policy as Code
- Retention rules live as code: lifecycle rules on the storage
layer, scheduled deletion jobs, tiered storage transitions, and
anonymization transforms. The code is versioned, reviewed, and
auditable. A retention rule in a spreadsheet is a wishlist; the
same rule in a reviewed, deployable lifecycle policy is a control.
```
// illustrative lifecycle policy (abstracted, no vendor DSL)
// object-storage lifecycle
{
"rules": [
{
"name": "user-events-90d",
"match": { "prefix": "events/" },
"transitions": [
{ "after": "30d", "to": "tier-cold" },
{ "after": "90d", "action": "delete" }
]
},
{
"name": "audit-log-7y",
"match": { "prefix": "audit/" },
"transitions": [
{ "after": "365d", "to": "tier-archive" },
{ "after": "2555d", "action": "delete" }
],
"legal_hold": "suspends-action"
}
]
}
```
- The retention policy is itself auditable: which rule fired when,
against which objects, with what result. The deletion events are
logged (`audit-logs.md`) — deletion is a significant action.
## Retention vs. Backup — The Distinction
- A **backup** is a recovery mechanism: it exists to restore data
after loss. A **retention rule** is a deletion mechanism: it
exists to remove data at end-of-life. Conflating them produces
data that survives both the deletion policy and the disaster —
which is the opposite of compliance.
- A backup is governed by a recovery-point / recovery-time objective;
a retention rule is governed by a lifetime. They are independent
contracts. A backup that is also the retention store is a store
where nothing is ever deleted, which is a P3 violation.
- A legal hold suspends retention deletion for a defined data set
(e.g. data under investigation). The hold is itself a policy
action, auditable and time-bounded, not a manual override.
## Retention is Distinct per Data Class
- Different data classes have different lifetimes. The retention
policy enumerates the classes and their rules; it does not apply
one number to everything. Typical classes (the names are
abstract; the periods are policy decisions, not regulation-
specific):
- **Audit logs** — long, often multi-year, governed by
accountability needs (`audit-logs.md`).
- **User-generated content** — tied to the user's account
lifetime; deletion follows account deletion (cross P8 Subject
Access).
- **Telemetry / metrics** — short, governed by observability need
(`domains/observability/metrics.md`); high-resolution data ages
to downsampled aggregates.
- **Evidence artifacts** — tied to the audit cycle
(`evidence.md`); the cycle ends, the evidence ages out.
- A single retention rule for "all data" is a C3 (Simplicity)
violation of the wrong kind: it is simpler than the requirement
allows.
## Subject Access is Honored (P8 Subject Access is Honored)
- Data-subject rights — access (what do we have on this subject),
export (in a portable form), deletion (and prove it), correction
— are operations with defined contracts and audit trails, not
ad-hoc tickets. The system implements them as first-class
operations; a subject-access request that requires a forensics
team is a correctness defect.
- Retention and subject access interact at deletion: a subject
deletion request fires the deletion policy for that subject's
data, the deletion is audited, and the proof of deletion is
returned to the subject (and recorded). A subject deletion that
skips the audit is a P8 violation dressed as a P3 success.
- Cross `domains/data/schema-design.md`: subject access is only
computable if the schema tags which records belong to which
subject. A schema with no subject linkage cannot honor a subject
request — it cannot find the data to delete.
## Retention Migration Discipline
- Retention rules change. When the policy changes (a class's
lifetime shortens or lengthens), the change is a migration: the
new rule applies to data ingested after the cutover, and a
backfill applies the new rule to existing data where applicable.
Cross `domains/data/migrations.md` for the schema-lifecycle
discipline this mirrors.
- A retention rule change that is not versioned, not reviewed, and
not backfilled is a P3 violation: the policy is not actually the
policy if the storage layer does not reflect it.
## What Violates Retention Discipline
| Violation | Principle |
|-----------|-----------|
| Data kept indefinitely because "storage is cheap" | P3 Retention is Policy, Not Storage |
| A retention rule in a spreadsheet, not in code | P3 Retention is Policy, Not Storage |
| A backup bucket used as the retention store (nothing ever deletes) | P3 Retention is Policy, Not Storage |
| A single retention period applied to all data classes | P3 Retention is Policy, Not Storage |
| A subject deletion with no audit record of the deletion | P8 Subject Access is Honored |
| A subject-access request that requires a forensics team | P8 Subject Access is Honored |
| A schema with no subject linkage (cannot find data to delete) | P8 Subject Access is Honored |
| A legal hold applied ad hoc, not as a policy action | P3 Retention is Policy, Not Storage |
| A retention rule change with no backfill to existing data | P3 Retention is Policy, Not Storage |
## Relationship to Other Domains
- `domains/data/schema-design.md` — retention requires the schema
to tag data class and subject linkage; subject access is only
computable over a schema that supports it.
- `domains/data/migrations.md` — retention rule changes are
migrations; the discipline (version, review, backfill) mirrors
schema migrations.
- `domains/compliance/audit-logs.md` — audit logs have their own
retention floor; deletion of an audit segment is meta-audited.
- `domains/compliance/evidence.md` — evidence artifacts have a
retention lifecycle tied to the audit cycle.
- `domains/observability/metrics.md` — telemetry retention is
governed by observability need; high-res data ages to aggregates.
- `domains/security/secrets.md` — secrets have a retention lifecycle
tied to rotation; a secret past its rotation date is overdue, not
retained.
+175
View File
@@ -0,0 +1,175 @@
# Evidence — Derived Rules
> Derives from `domains/compliance/first-principles.md`. Covers P6
> (Evidence is Collected Continuously), P5 (Policy is a Gate, so
> decisions are evidence), P7 (Identity Attributable, so evidence
> has provenance), and P10 (Posture Observable, so evidence is
> queryable). Referenced by `audit-logs.md` (logs are evidence)
> and `data-retention.md` (evidence has a lifecycle).
## Evidence is Collected Continuously (P6 Evidence is Collected Continuously)
- Evidence of compliance — logs, configs, scans, attestations,
policy decisions, access reviews — is gathered as a byproduct of
operation, not assembled manually at audit time. The audit-time
scramble is the anti-pattern: it is expensive, it is incomplete,
and it produces evidence that is reconstructed rather than
recorded.
- Continuous evidence collection means the audit packet is a query
over already-collected artifacts, not a forensic reconstruction.
The auditor asks "show me the access reviews for Q3" and the
answer is a query against the evidence store, not a six-week
project.
- This is the compliance angle on `domains/observability/tracing.md`
for distributed evidence (a trace spans the request that produced
the evidence) and `domains/observability/metrics.md` for posture
signals (a metric is a continuous evidence stream).
## Evidence is a Byproduct, Not a Deliverable
- Evidence collected as a byproduct is trustworthy: it records what
happened, when it happened, recorded by the system that did it.
Evidence assembled at audit time is less trustworthy: it records
what someone remembered to write down, when they wrote it, after
the fact.
- The mechanism: every significant action (`audit-logs.md`) emits
its record to an evidence store; every policy decision
(`policy-as-code.md`) emits its decision; every deployment emits
its signed attestation; every access review emits its result. The
store is append-only (`audit-logs.md` P1), queryable (P10), and
retention-bound (`data-retention.md`).
## Provenance and Identity (P7 Identity is Attributable)
- Evidence has provenance: which system produced it, when, from what
input. An evidence artifact with no provenance is anecdote, not
evidence — it cannot be attributed to a source, so it cannot be
trusted.
- Provenance includes the identity of the producer (a workload
identity, a service account) and the chain of custody (who has
had access to the artifact since it was produced). Cross
`domains/security/authorization.md`: the producer's identity is
authenticated, not assumed.
## Signed Attestations (IDEATE-29)
- A signed attestation is evidence with a cryptographic signature
binding the artifact to its producer. The signature is the
provenance: it can be verified independently of the producer, and
it cannot be forged without the producer's key. Cross
`domains/security/supply-chain.md` for the supply-chain angle.
- Cosign (Sigstore) and in-toto are the canonical patterns: a
builder signs an artifact (container image, deployable, evidence
bundle) at production time; a verifier checks the signature at
consumption time. The signature is the evidence that the artifact
came from where it claims to have come from.
- **Illustrative signed attestation (Cosign / Sigstore format, NOT
a real signature — illustrative only, no live keys):**
```
// Cosign attest — bind an attestation to an image digest
// (illustrative; not a real signature)
$ cosign attest --type spdxjson \
--predicate sbom.spdx.json \
my-registry/app@sha256:5a3e1c...f9b2
// The attestation is stored as a signature in the registry,
// bound to the image digest. The payload is a DSSE envelope:
{
"payloadType": "application/vnd.in-toto+json",
"payload": "eyJfdHlwZSI6ImF0dGVzdGF0aW9uIn0...",
"signatures": [
{
"sig": "MEUCIQDx...illustrative-base64-signature...==",
"keyid": "cosign-key-2024-q4"
}
]
}
// The decoded payload (an in-toto statement binding the
// attestation to the image digest):
{
"_type": "https://in-toto.io/Statement/v0.1",
"predicateType": "https://spdx.dev/Document",
"subject": [
{
"name": "my-registry/app",
"digest": { "sha256": "5a3e1c...f9b2" }
}
],
"predicate": {
"SPDXID": "SPDXRef-DOCUMENT",
"creationInfo": {
"created": "2024-11-07T15:03:22Z",
"creators": ["Tool: atelier-build-pipeline"]
}
}
}
// Verification (independent of the producer):
$ cosign verify-attestation --type spdxjson \
--certificate-identity-regexp '.*atelier-build.*' \
my-registry/app@sha256:5a3e1c...f9b2
// Verification succeeded for: my-registry/app@sha256:5a3e1c...f9b2
// SBOM attestation found for subject
```
- The attestation is illustrative — the signatures and digests are
not real. The shape (DSSE envelope, in-toto statement, subject +
predicate, verify-by-identity) is what evidence-as-attestation
looks like. A real attestation carries a real signature from a
real key held by the builder.
## Audit-Ready Export
- The evidence store is queryable at any time, not only at audit
time. The audit packet is a query (a date range, a data class, a
subject) over the store; the export is a dump of the matching
artifacts with their provenance and signatures.
- An audit-ready export that requires six weeks of forensics is a
P6 violation dressed as a success: the evidence was not collected
continuously, it was reconstructed. The export should be a query
that runs in minutes, not a project that runs for weeks.
## Evidence Lifecycle
- Evidence has a retention lifecycle (`data-retention.md`): an
evidence artifact is retained for the audit cycle it supports,
then ages out. The retention rule for evidence is itself audited
(deletion of evidence is a meta-audited action, like deletion of
audit logs).
- A legal hold suspends evidence deletion for a defined set — the
same mechanism as audit-log holds.
## What Violates Evidence Discipline
| Violation | Principle |
|-----------|-----------|
| Evidence assembled by hand the week before an audit | P6 Evidence is Collected Continuously |
| An evidence artifact with no provenance (no producer, no timestamp) | P7 Identity is Attributable |
| An audit packet that requires six weeks of forensics to produce | P6 Evidence is Collected Continuously |
| An attestation with no signature (provenance asserted, not proven) | P7 Identity is Attributable |
| Evidence store not queryable between audits | P10 Compliance Posture is Observable |
| Evidence deleted before its retention period with no meta-audit | P6 Evidence is Collected Continuously |
| Policy decisions not recorded as evidence | P5 Policy is Evaluated as a Gate |
| A deployment with no signed attestation of its build provenance | P7 Identity is Attributable |
## Relationship to Other Domains
- `domains/security/supply-chain.md` — signed attestations are the
supply-chain integrity primitive; evidence.md is the compliance
consumer of the same artifact.
- `domains/observability/metrics.md` — posture metrics are a
continuous evidence stream.
- `domains/observability/tracing.md` — distributed traces provide
evidence that spans a request across services.
- `domains/compliance/audit-logs.md` — audit logs are a primary
evidence artifact; the append-only guarantee is what makes them
admissible.
- `domains/compliance/policy-as-code.md` — policy decisions are
evidence of enforcement; the policy code itself is evidence of
the rule.
- `domains/compliance/data-retention.md` — evidence has a retention
lifecycle tied to the audit cycle.
+184
View File
@@ -0,0 +1,184 @@
# Compliance — First Principles
> Framework-agnostic per D-024. These principles derive from core
> Security (a subset of C1 Correctness), Observability, and
> Reversibility. They apply across regulations — NIST CSF, SOC 2,
> GDPR, CCPA, HIPAA, PCI-DSS, ISO 27001 — without prescribing any
> regulation-specific implementation. Regulation names appear here
> only as examples of what the principles support; the principles
> themselves are engineering rules, not legal controls.
## 1. The Principles
### P1. Audit Logs are Append-Only
Audit records are immutable once written. Deletion or mutation of an
audit record is itself an auditable incident — the tampering is the
signal, not just the underlying event. An audit log that can be edited
is not an audit log; it is a draft. Append-only is enforced
structurally (write-once storage, immutable buckets, hash-chained
records), not by policy alone. This is the compliance angle on
`domains/observability/logging.md`: structured logs that cannot be
rewritten are the substrate of accountability.
### P2. Every Significant Action is Logged
The set of auditable actions is defined a priori, in code, before the
action ships — not retrofitted after an incident. Authentication
changes, authorization decisions, data access, configuration changes,
policy changes, and deletions are all significant. "We forgot to log
it" is a violation, not an excuse. The auditable-action catalog is
itself versioned and reviewed. A significant action with no log line
is a C7 (Observability) defect and a C1 (Correctness) defect: the
system's behavior is invisible, and accountability is impossible.
### P3. Retention is Policy, Not Storage
Data lifetime is declared and enforced as policy, not left to the
storage layer's defaults. Deletion at end-of-life is a feature, not a
failure. Retention rules live as code (lifecycle rules, scheduled
deletion jobs, tiered storage transitions), they are reviewed, and
they are auditable. "We kept it because the bucket was cheap" is a
violation; "we deleted it because the policy said to" is correct.
Retention is distinct from backup: a backup is a recovery mechanism,
a retention rule is a deletion mechanism. Keeping them conflated
produces data that survives both the deletion policy and the
disaster — which is the opposite of compliance. Cross
`domains/data/migrations.md` for the schema-lifecycle discipline.
### P4. Policy is Code
Compliance policy is expressed in versioned, reviewable, testable
code (OPA / Rego, AWS Cedar, HashiCorp Sentinel, Kyverno) — not in
spreadsheets, prose documents, or tribal knowledge. Policy in a
spreadsheet is untestable, unreviewable, and undeployable; it is a
wishlist, not a control. Policy-as-code inherits the disciplines of
`domains/infrastructure-as-code/P1 Declarative Intent`: declarative
intent, version control, review before merge, plan before apply. A
compliance rule that is not executable is a rule that cannot be
enforced, which is a rule that does not exist.
### P5. Policy is Evaluated as a Gate
Policy violations block before the action, not after the audit.
Enforcement happens at admission time (kubernetes admission), at
pipeline time (CI/CD gates), and at provisioning time (IaC plan
gates) — before the non-compliant state is realized. Detecting a
violation after it ships is detection, not enforcement. A policy that
is "logged but not blocked" is a postcard, not a gate. This is the
compliance angle on C5 (Reversibility): a blocked action is
reversible by construction; a shipped violation requires remediation,
which is more expensive than prevention.
### P6. Evidence is Collected Continuously
Evidence of compliance — logs, configs, scans, attestations, policy
decisions, access reviews — is gathered as a byproduct of operation,
not assembled manually at audit time. The audit-time scramble is the
anti-pattern: it is expensive, it is incomplete, and it produces
evidence that is reconstructed rather than recorded. Continuous
evidence collection means the audit packet is a query over
already-collected artifacts, not a forensic reconstruction. This is
the compliance angle on `domains/observability/tracing.md` for
distributed evidence and `domains/observability/metrics.md` for
posture signals.
### P7. Identity is Attributable
Every logged action traces to an authenticated, non-shared principal.
Shared accounts, generic service identities, and "admin" as an actor
are violations: an action with no attributable human or workload is
an action with no accountability. Identity is recorded in the audit
record at the time of the action, not resolved after the fact. This
is the compliance angle on `domains/security/authorization.md` and
`domains/kubernetes/rbac.md`: the audit subject must be the principal
that acted, not a role that many can assume.
### P8. Subject Access is Honored
Data-subject rights — access, export, deletion, correction — are
operations with defined contracts and audit trails, not ad-hoc
tickets. The system can answer "what do we have on this subject,"
"export it in a portable form," and "delete it and prove the
deletion" as first-class operations. These are not features bolted on
at the end; they are contracts the data layer implements from the
start. A subject-access request that requires a forensics team is a
correctness defect: the system does not know what it holds. Cross
`domains/data/schema-design.md` for the data shapes that make
subject access computable.
### P9. Secrets and Sensitive Data are Redacted in Audit
Audit logs themselves must not leak secrets, credentials, PII, or
other sensitive data. Redaction is structural — applied at the
logging boundary, before the record is written — not opportunistic
scrubbing after the fact. A secret that appears in an audit log is a
C1 (Correctness) violation (the log is now a secret store) and a
security violation (`domains/security/secrets.md`). The redaction
policy is itself auditable: which fields are redacted, by what rule,
in which log stream. Once a secret is in an append-only log, the
remediation is expensive — rotate the secret and rewrite the log's
access scope — so redaction-at-source is the only sound position.
### P10. Compliance Posture is Observable
The system reports its own compliance state: drift from policy, open
violations, retention status, evidence freshness, policy-evaluation
counts. Silent non-compliance is the bug. A compliance posture
metric is a first-class signal (`domains/observability/metrics.md`),
alertable, and dashboarded. "We didn't know we were non-compliant"
is not a defense; it is a C7 (Observability) defect. The posture is
queryable at any time, not only at audit time. This is the compliance
angle on `domains/infrastructure-as-code/P3 State is Truth`: the
compliance state is a versioned, queryable truth, not a vibe.
## 2. Core Principle Trace
Each compliance P-rule derives from one or more core C-rules (C1C8).
The matrix extension lands in P4 of the v0.3 plan; the traces below
are authoritative.
| P-rule | Core | Why |
|--------|------|-----|
| P1 Audit Logs are Append-Only | C1, C5 | Correctness of the record; reversibility of tamper detection |
| P2 Every Significant Action is Logged | C7, C1 | Observability of behavior; correctness of a-priori audit scope |
| P3 Retention is Policy, Not Storage | C5, C8 | Reversibility of data lifetime; economy of storage as policy |
| P4 Policy is Code | C6, C2 | Composability of versioned policy; clarity of executable intent |
| P5 Policy is Evaluated as a Gate | C1, C5 | Correctness of pre-action enforcement; reversibility of blocked actions |
| P6 Evidence is Collected Continuously | C7, C3 | Observability of compliance state; simplicity of audit-by-query |
| P7 Identity is Attributable | C1, C7 | Correctness of accountability (security subset); observability of who acted |
| P8 Subject Access is Honored | C1, C5 | Correctness of the data-subject contract; reversibility of deletion |
| P9 Secrets and Sensitive Data are Redacted in Audit | C1, C3 | Correctness of not leaking (security subset); simplicity of structural redaction |
| P10 Compliance Posture is Observable | C7, C1 | Observability of posture; correctness of self-reported state |
## 3. What Violates These Principles
| Violation | Principle Breached |
|-----------|-------------------|
| An audit log stored on a mutable filesystem with no write-once protection | P1 Audit Logs are Append-Only |
| A `DELETE` on an audit record to "clean up a typo" | P1 Audit Logs are Append-Only |
| An auth change with no audit log line | P2 Every Significant Action is Logged |
| "We'll add logging after we ship the feature" | P2 Every Significant Action is Logged |
| Data kept indefinitely because "the bucket is cheap" | P3 Retention is Policy, Not Storage |
| A retention rule in a spreadsheet, not in code | P4 Policy is Code |
| A policy that logs violations but does not block the action | P5 Policy is Evaluated as a Gate |
| Evidence assembled by hand the week before an audit | P6 Evidence is Collected Continuously |
| A shared `admin` account as the audit actor | P7 Identity is Attributable |
| A subject-access request that requires a forensics team | P8 Subject Access is Honored |
| A secret visible in an audit log entry | P9 Secrets and Sensitive Data are Redacted in Audit |
| No dashboard for compliance posture between audits | P10 Compliance Posture is Observable |
## 4. Relationship to Other Domains
Compliance is the accountability layer that crosses
`domains/security/` (it audits security actions),
`domains/observability/` (audit logs are structured logging; posture
is metrics; evidence is traces), `domains/data/` (retention and
subject access are data-layer contracts), and
`domains/infrastructure-as-code/` (policy-as-code parallels
declarative IaC; compliance state parallels state-as-truth). Cross-
links are one-directional (per D-026 extended):
- `domains/security/authorization.md` ← P7 (attributable identity)
- `domains/security/secrets.md` ← P9 (redaction)
- `domains/security/supply-chain.md` ← P6 (signed attestations as evidence)
- `domains/observability/logging.md` ← P1, P2 (audit logs = structured logging)
- `domains/observability/metrics.md` ← P10 (compliance posture metrics)
- `domains/observability/tracing.md` ← P6 (evidence from distributed traces)
- `domains/data/schema-design.md` ← P3, P8 (retention and subject-access shapes)
- `domains/data/migrations.md` ← P3 (retention migration discipline)
- `domains/infrastructure-as-code/P1 Declarative Intent` ← P4 (policy-as-code)
- `domains/infrastructure-as-code/P3 State is Truth` ← P10 (compliance posture truth)
- `domains/kubernetes/rbac.md` ← P7 (audit subject identity)
- `domains/devops/ci-cd.md` ← P5 (policy as a pipeline gate)
- `domains/devops/first-principles.md` ← P4 (policy as configuration-as-code)
+140
View File
@@ -0,0 +1,140 @@
# Policy as Code — Derived Rules
> Derives from `domains/compliance/first-principles.md`. Covers P4
> (Policy is Code) and P5 (Policy is Evaluated as a Gate). Referenced
> by `audit-logs.md` (policy decisions are audited) and `evidence.md`
> (policy decisions are evidence). Framework-agnostic per D-024.
## Policy is Code (P4 Policy is Code)
- Compliance policy is expressed in versioned, reviewable, testable
code — not in spreadsheets, prose documents, or tribal knowledge.
Policy in a spreadsheet is untestable, unreviewable, and
undeployable; it is a wishlist, not a control.
- Policy-as-code inherits the disciplines of
`domains/infrastructure-as-code/P1 Declarative Intent`: declarative
intent, version control, review before merge, plan before apply.
A compliance rule that is not executable is a rule that cannot be
enforced, which is a rule that does not exist.
- Policy code is tested like any other code: unit tests for the rule
logic (given an input, the rule allows or denies as expected),
integration tests for the gate (the rule fires at the right point
in the pipeline), and versioning for the policy itself (a policy
change is a reviewed, merged, deployed change).
## Policy is Evaluated as a Gate (P5 Policy is Evaluated as a Gate)
- Policy violations block **before** the action, not after the
audit. Enforcement happens at:
- **Admission time** — a kubernetes admission webhook denies a
non-compliant resource before it is created
(`domains/kubernetes/rbac.md`).
- **Pipeline time** — a CI/CD gate denies a non-compliant change
before it merges (`domains/devops/ci-cd.md`).
- **Provisioning time** — an IaC plan gate denies a non-compliant
resource before `apply` (`domains/infrastructure-as-code/`).
- A policy that logs violations but does not block the action is a
postcard, not a gate. Detection is not enforcement. A logged
violation that the actor could ignore is a P5 violation — the
policy exists, but the system is not compliant by construction.
- The gate is the contract. The policy author writes the rule; the
gate operator wires the rule into the enforcement point; the
auditor verifies the gate fired. All three are auditable
(`audit-logs.md`).
## Engine Comparison (IDEATE-23)
| Engine | Policy Language | Evaluation Gate | Ecosystem | Notes |
|--------|-----------------|-----------------|-----------|-------|
| **OPA / Rego** | Rego (declarative, set-based, Datalog-inspired) | CI/CD, k8s admission (Gatekeeper), HTTP API, IaC plan (Terraform Sentinel-style), service mesh | Broadest ecosystem; CNCF graduated; library of reusable bundles | General-purpose; the default choice when the gate location varies |
| **AWS Cedar** | Cedar (declarative, authorization-focused, schema-typed) | k8s admission (via Cedar-agent), application authorization, AVP (Verified Permissions) | AWS-native; tight schema typing; separates policy from entities | Authorization-focused; strong where the policy is "who can do what on which resource" |
| **HashiCorp Sentinel** | Sentinel (declarative, restricted, policy-focused) | Terraform / TFE plan gate, Nomad, Vault | HashiCorp ecosystem; embedded in Terraform Enterprise / HCP | IaC-plan-gate native; the enforcement point is the `plan` output |
| **Kyverno** | Kyverno (YAML-declarative, k8s-native, no new DSL) | k8s admission (native), cluster-wide policy reports | Kubernetes-native; no separate language — policy is a CRD | k8s-cluster-gate native; the choice when the gate is admission and the team prefers YAML over a DSL |
- None is advocated over the others. The choice is (a) where the
gate fires, (b) the team's tolerance for a new policy language,
and (c) ecosystem fit. All four satisfy P4/P5 when wired
correctly.
- A gate is a gate regardless of engine: the rule is declarative,
the evaluation is pre-action, and the decision is allow-or-deny.
The engine difference is language and enforcement-point fit, not
correctness.
## Policy Testing
- Policy code is unit-tested like any other code. A test asserts
that a given input produces the expected decision (allow / deny /
warn). The test is versioned with the policy; a policy change
with no test change is a red flag.
```
// illustrative Rego policy + test
// policy: deny containers running as root
package k8s.admission
deny[msg] {
input.kind == "Pod"
c := input.spec.containers[_]
not c.securityContext.runAsNonRoot
msg := sprintf("container %s must set runAsNonRoot", [c.name])
}
// test (Rego unit test)
package k8s.admission
test_deny_root_container {
some msg in deny with input as {
"kind": "Pod",
"spec": { "containers": [ { "name": "app", "securityContext": {} } ] }
}
msg == "container app must set runAsNonRoot"
}
```
- Integration tests assert the gate fires: a non-compliant resource
submitted to the admission endpoint is denied; a compliant one is
allowed. The integration test runs against the real gate, not a
mock, because the gate wiring is half the contract.
## Policy Versioning
- Policy is versioned in git. A policy change is a reviewed, merged,
deployed change — the same discipline as application code. A
policy that is edited in production without review is a P4
violation: the policy is code, but it is being treated as config.
- A policy change can break existing workloads (a new deny rule
blocks a previously-allowed resource). The rollout is staged:
warn-only mode first (log violations, do not block), then enforce
mode after the violation count is zero. This is the policy
analogue of `domains/devops/P5 Progressive Delivery`.
## What Violates Policy-as-Code Discipline
| Violation | Principle |
|-----------|-----------|
| A compliance rule in a spreadsheet | P4 Policy is Code |
| A policy that logs violations but does not block the action | P5 Policy is Evaluated as a Gate |
| A policy edited in production without review | P4 Policy is Code |
| A policy with no unit tests for the rule logic | P4 Policy is Code |
| A gate wired with a mock instead of the real engine | P5 Policy is Evaluated as a Gate |
| A new deny rule enforced without a warn-only rollout | P5 Policy is Evaluated as a Gate |
| A policy in prose ("the team should not use root containers") | P4 Policy is Code |
| A policy decision with no audit record | P5 Policy is Evaluated as a Gate |
## Relationship to Other Domains
- `domains/infrastructure-as-code/first-principles.md` — policy-as-
code inherits declarative intent, versioning, and plan-before-
apply from IaC.
- `domains/kubernetes/rbac.md` — k8s admission is a primary
enforcement gate; Kyverno and OPA Gatekeeper wire into it.
- `domains/devops/ci-cd.md` — CI/CD is a pipeline-time enforcement
gate; a policy step blocks a non-compliant change before merge.
- `domains/compliance/audit-logs.md` — every policy decision (allow
/ deny) is an audited significant action.
- `domains/compliance/evidence.md` — policy decisions and the
policy code itself are evidence of enforcement posture.
- `domains/security/authorization.md` — Cedar's authorization-
focused policy overlaps with authz; the split is that authz is
the runtime decision, policy-as-code is the reviewed rule that
drives it.
+176
View File
@@ -0,0 +1,176 @@
# ArgoCD — Derived Rules
> Derives from `domains/gitops-operators/first-principles.md`.
> Applies P1P10 to ArgoCD specifically. For the ArgoCD-vs-Flux
> decision, see the decision matrix at the end of this doc and in
> `flux.md`.
## What ArgoCD Is (P1 Git is the Source of Truth, P3 Pull, Don't Push)
- ArgoCD is a pull-based GitOps controller for Kubernetes. It runs
inside the target cluster, pulls desired state from git, and
reconciles the cluster to match. CI never holds `kubectl` rights
against the cluster (P3).
- An Application is a declarative binding of "this git path" to
"this cluster destination." The Application CRD is the unit of
reconciliation. The cluster is a derivative of git, never the
authority (P1).
- ArgoCD supports Helm charts, Kustomize overlays, ksonnet, and raw
manifests as source formats — see `domains/kubernetes/helm.md`
and `domains/kubernetes/kustomize.md`.
## Application CRD (P2 Declarative Over Imperative, P4 Continuous Reconciliation)
- An Application declares `source` (repo, path, revision, chart),
`destination` (server, namespace), and `syncPolicy`. The
reconciler loops continuously; drift is corrected automatically,
not on-demand (P4).
```yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: payments-api
namespace: argocd
spec:
source:
repoURL: https://git.example.com/platform/payments
targetRevision: 1.2.3
path: manifests/prod
destination:
server: https://kubernetes.default.svc
namespace: payments
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=false
```
- `automated.prune: true` deletes resources removed from git.
`selfHeal: true` corrects hand-edited drift back to git (P8).
Disable both for workloads that need manual approval gates.
## App-of-Apps (P6 Operators Encode Domain Knowledge, C6 Composability)
- The App-of-Apps pattern: one root Application points at a git
directory of child Application manifests. The root app reconciles
the children; the children reconcile the workloads. This is the
ArgoCD expression of composition — a fleet of apps as a tree of
Applications.
- Use App-of-Apps for cluster bootstrapping (one repo, many
clusters, many apps). Do not use it as a substitute for a package
manager; if you are templating hundreds of near-identical
Applications, use a generator (ApplicationSet) instead.
## Sync Waves and Hooks (P4 Continuous Reconciliation, P7 Reversibility)
- Sync waves order resources within a sync: `PreSync``Sync`
`PostSync`. Use waves to run a job before a Deployment, or a
migration before the app that depends on it.
- Sync hooks (`PreSync`, `Sync`, `PostSync`, `SyncFail`) are
Resources annotated to execute at a wave boundary. A `SyncFail`
hook runs on sync failure — the abort path (P7).
- Wave ordering is a correctness mechanism, not a performance one.
Mis-ordered waves (e.g., app starts before its migration job)
are a correctness bug.
## Health and Status (P9 Failure is Observable and Surfaced)
- ArgoCD assesses every resource's health (`Healthy`, `Progressing`,
`Degraded`, `Missing`, `Suspended`) and surfaces the aggregate as
Application status. Sync status (`Synced`, `OutOfSync`) reports
drift against git.
- Health checks are pluggable via Lua scripts for custom CRDs. An
Operator-managed CRD without a health check reads as `Progressing`
forever — write one (see `operators.md`).
- Out-of-sync or degraded status must emit a notification (Slack,
PagerDuty, webhook). Silent drift is the bug (P9). Wire status to
`domains/observability/metrics.md`.
## Diff and Drift (P8 Reconcile, Don't Mutate by Hand, P4)
- `argocd app diff` shows the diff between git and live cluster.
A non-empty diff on a synced app is hand-edit drift — the
recovery is `selfHeal`, not a manual `kubectl apply` (P8).
- Drift detection runs continuously (P4). The gap between "git
changed" and "cluster matches git" is observable, not assumed.
## RBAC and SSO (P10 Least Privilege Reconciliation)
- ArgoCD's own RBAC governs who can view, sync, and admin
Applications. Bind to SSO (OIDC, SAML) for human identity; bind
the controller's service account to a Role scoped to the
namespaces it reconciles.
- The controller's credentials must not be `cluster-admin` (P10).
Use namespace-scoped Roles via `ApplicationSet` namespaces or
cluster-wide AppProject restrictions. See
`domains/kubernetes/rbac.md` and `domains/security/authorization.md`.
- AppProjects bound the blast radius of what an Application can
deploy (allowed repos, destinations, roles). One AppProject per
team or environment; the default project is for nothing in
production.
## Multi-Cluster (P4 Locality, P10)
- ArgoCD registers external clusters by secret. The controller
pulls from git and pushes to the registered cluster's API server.
The "pull, don't push" boundary (P3) is between the target
cluster's reconciler and CI — the controller-to-apiserver hop is
internal to the platform.
- Scope each registered cluster's credentials to the namespaces
ArgoCD manages there. Do not register a cluster with cluster-admin
and call it done (P10).
## Sync Windows (P5 Reversibility, P7)
- Sync windows restrict when automated sync runs (e.g., no syncs
during business hours, or syncs only in a maintenance window).
They are a reversibility mechanism: a bad commit lands in git,
but the sync window holds it until review.
- Sync windows do not replace health monitoring (P9). A degraded
app inside a window is still an incident.
## Secrets (P10, cross-link security/secrets)
- Do not store raw Secrets in the GitOps repo. Use a sealed-secret
controller (Bitnami Sealed Secrets, SOPS, External Secrets
Operator) so the git store holds encrypted material only. See
`domains/security/secrets.md` for the general secret-hygiene
principles.
## ArgoCD vs Flux — Decision Matrix (IDEATE-21, D-039)
| Axis | ArgoCD | Flux |
|------|--------|------|
| Architecture | Monolithic controller + Application CRD | Composable GitOps Toolkit controllers (source, kustomize, helm, notification) |
| Reconciliation unit | Application (one CRD per app) | Kustomization / HelmRelease (one per deploy unit) |
| UI | Web UI + CLI (full dashboard, tree view, diff viewer) | CLI-first; UI via Weave GitOps or FluxUI (add-on) |
| Sync model | Periodic poll or webhook; sync waves + hooks | Poll + webhook; runs continuously, no explicit sync waves |
| Multi-cluster | One ArgoCD manages many clusters (hub-and-spoke) | One Flux per cluster (per-cluster autonomy) |
| Templating in repo | Helm, Kustomize, ksonnet, raw manifests, Jsonnet | Helm, Kustomize, raw manifests |
| RBAC | Built-in RBAC + SSO + AppProjects | Kubernetes RBAC (no built-in RBAC layer) |
| Progressive delivery | Argo Rollouts (sister project, tight integration) | Flagger (sister project, tight integration) |
| Best for | Teams wanting a UI, multi-cluster from one pane, App-of-Apps bootstrapping | Teams wanting composable controllers, per-cluster autonomy, minimal footprint |
| Watch out for | Monolithic controller scaling, UI as ops crutch, AppProject sprawl | No native UI, steeper learning curve, manual multi-cluster orchestration |
- Use ArgoCD when you want a UI, central multi-cluster management,
and sync-wave ordering. Use Flux when you want composable
controllers, per-cluster autonomy, and a minimal footprint.
- Both are CNCF graduated and both implement the OpenGitOps
principles. The choice is architectural fit, not correctness. See
`flux.md` for the Flux-side perspective.
## What Violates ArgoCD Discipline
| Violation | Principle |
|-----------|-----------|
| CI pipeline with `kubectl` rights pushing to the cluster | P3 Pull, Don't Push |
| `argocd app set` used as the steady state instead of git | P1 Git is the Source of Truth |
| `selfHeal: false` on a prod app with no manual gate | P8 Reconcile, Don't Mutate by Hand |
| Controller ServiceAccount bound to `cluster-admin` | P10 Least Privilege Reconciliation |
| Sync failure with no notification wired | P9 Failure is Observable and Surfaced |
| AppProject with no destination restrictions in prod | P10 Least Privilege Reconciliation |
| Raw Secret in the GitOps repo | P10, `domains/security/secrets.md` |
| Manual `kubectl edit` on an ArgoCD-managed resource | P8 Reconcile, Don't Mutate by Hand |
@@ -0,0 +1,131 @@
# GitOps + Operators — First Principles
## 1. The Principles
### P1. Git is the Source of Truth
Desired state lives in a versioned, immutable git store. The
cluster is a derivative of git, never the authority. If a state
exists only in the cluster and not in git, it is drift, not truth.
The commit history is the audit trail and the rollback path.
### P2. Declarative Over Imperative
Express the desired cluster state, not the commands to reach it.
A manifest says what should exist; the reconciler makes it so.
Imperative `kubectl` is for inspection and incident response, not
for the steady state. This is the GitOps expression of
`domains/kubernetes/P1 Declarative Desired State` and
`domains/infrastructure-as-code/P1 Declarative Intent`.
### P3. Pull, Don't Push
Agents running inside the target pull desired state from git; the
target never accepts outside push credentials. No CI pipeline holds
`kubectl` rights against the production cluster. The cluster reaches
out to git, not the other way around. This is the security primitive
of GitOps: the blast radius of a compromised CI is bounded by what CI
can push, and a pull model gives CI nothing to push.
### P4. Continuous Reconciliation
The reconciliation loop is the primitive. Drift is detected and
corrected automatically, not on-demand. A manual `apply` is an
exception, not the workflow. The loop runs continuously; the gap
between "git changed" and "cluster matches git" is measured in
seconds, not tickets.
### P5. State is Immutable and Versioned
Every change to desired state is a commit. History is the audit
trail and the rollback path. A revert is a rollback; a force-push is
history deletion. The git store is treated like
`domains/infrastructure-as-code/P3 State is Truth` — lose it or
tamper with it, and you lose the ability to reason about the system.
### P6. Operators Encode Domain Knowledge
Operational expertise lives as CRDs plus controllers, not as
runbooks that humans must remember. An operator is a control loop
that encodes how to reconcile a specific domain (a database, a
message queue, a certificate). The operator is the deepest
expression of `domains/kubernetes/P1 Declarative Desired State`
the domain knowledge is the desired state.
### P7. Progressive Delivery is Reversible by Construction
Canary and blue-green are staged, metric-gated, and one-command
abortable. Promotion without a rollback path is a violation. A
rollout that cannot be aborted is a deploy, not a progressive
delivery. This is the GitOps extension of
`domains/devops/P5 Progressive Delivery` and
`domains/kubernetes/P10 Roll Forward, Roll Back`.
### P8. Reconcile, Don't Mutate by Hand
Manual `kubectl apply` or `kubectl edit` on a GitOps-managed
resource is an incident. The reconciler will overwrite the hand
edit on the next loop; the hand edit was never truth. Drift back to
git is the recovery, not the failure. This is the GitOps angle on
`domains/infrastructure-as-code/P9 Drift is Recoverable`.
### P9. Failure is Observable and Surfaced
Sync failures, health degradation, and rollout-stall events emit
status and notifications. Silent drift is the bug. A GitOps
controller that fails to sync without surfacing the failure has
violated the contract — you cannot fix what you cannot see
(`domains/observability/metrics.md`).
### P10. Least Privilege Reconciliation
The controller's credentials are scoped to the namespaces and
resources it reconciles. No `cluster-admin` GitOps robots. One
credential set per boundary; the reconciler sees only what it
reconciles. This is the GitOps angle on
`domains/kubernetes/P7 RBAC by Intent, Not Identity` and
`domains/security/authorization.md`.
## 2. Core Principle Trace
Each GitOps + Operators P-rule derives from one or more core
C-rules (C1C8). The matrix extension lands in P4 of the v0.3
plan; the traces below are authoritative.
| P-rule | Core | Why |
|--------|------|-----|
| P1 Git is the Source of Truth | C1, C5 | Correctness of state; reversibility via history |
| P2 Declarative Over Imperative | C2, C3 | Clarity of intent; simplicity of mental model |
| P3 Pull, Don't Push | C1, C4 | Correctness via security; locality of credentials |
| P4 Continuous Reconciliation | C7, C1 | Observability of drift; correctness of convergence |
| P5 State is Immutable and Versioned | C5 | Reversibility via version history |
| P6 Operators Encode Domain Knowledge | C6, C2 | Composability of expertise; clarity of operational intent |
| P7 Progressive Delivery is Reversible | C5, C1 | Reversibility of promotion; correctness of abort |
| P8 Reconcile, Don't Mutate by Hand | C1, C7 | Correctness of single source; observability of drift |
| P9 Failure is Observable and Surfaced | C7 | Observability of reconciliation |
| P10 Least Privilege Reconciliation | C1, C8 | Correctness via security; economy of trust |
## 3. What Violates These Principles
| Violation | Principle Breached |
|-----------|-------------------|
| CI pipeline pushes manifests to the cluster | P3 Pull, Don't Push |
| A resource exists in the cluster but not in git | P1 Git is the Source of Truth |
| `kubectl edit` on a GitOps-managed resource | P8 Reconcile, Don't Mutate by Hand |
| Reconciler with `cluster-admin` ClusterRoleBinding | P10 Least Privilege Reconciliation |
| Sync failure with no status or notification | P9 Failure is Observable and Surfaced |
| Canary with no abort/rollback path | P7 Progressive Delivery is Reversible |
| Operator runbook that exists only in a wiki | P6 Operators Encode Domain Knowledge |
| Reconciler that applies on a cron, not continuously | P4 Continuous Reconciliation |
| Force-push rewrites GitOps repo history | P5 State is Immutable and Versioned |
| Imperative deploy script as the steady state | P2 Declarative Over Imperative |
## 4. Relationship to Other Domains
GitOps + Operators is the deployment-automation layer above
`domains/kubernetes/` and `domains/infrastructure-as-code/`. It
borrows their declarative-reconciliation model and adds the
git-as-source-of-truth and pull-based credential boundaries. Cross
links are one-directional (per D-026 extended):
- `domains/kubernetes/P1 Declarative Desired State` ← P2
- `domains/kubernetes/P10 Roll Forward, Roll Back` ← P7
- `domains/infrastructure-as-code/P1 Declarative Intent` ← P2
- `domains/infrastructure-as-code/P3 State is Truth` ← P1, P5
- `domains/infrastructure-as-code/P9 Drift is Recoverable` ← P4, P8
- `domains/devops/P4 Rollback First` ← P5, P7
- `domains/devops/P5 Progressive Delivery` ← P7
- `domains/devops/P6 Configuration as Code` ← P1, P2
- `domains/security/secrets.md` ← P3, P10 (reconciliation credentials)
- `domains/security/supply-chain.md` ← P5 (signed, immutable provenance)
- `domains/observability/metrics.md` ← P4, P9 (reconciliation + rollout metrics)
+159
View File
@@ -0,0 +1,159 @@
# Flux — Derived Rules
> Derives from `domains/gitops-operators/first-principles.md`.
> Applies P1P10 to Flux specifically. For the ArgoCD-vs-Flux
> decision, see the decision matrix at the end of this doc and in
> `argocd.md`.
## What Flux Is (P1 Git is the Source of Truth, P3 Pull, Don't Push)
- Flux is a set of composable controllers — the GitOps Toolkit —
that run inside the target cluster, pull desired state from git
or OCI registries, and reconcile the cluster to match. CI never
holds `kubectl` rights against the cluster (P3).
- The composable-controller architecture is a C6 (Composability)
exemplar: each controller does one thing (source, kustomize, helm,
notification) and the controllers compose into a full GitOps
system.
- Flux supports Helm releases, Kustomize overlays, and raw
manifests — see `domains/kubernetes/helm.md` and
`domains/kubernetes/kustomize.md`.
## GitOps Toolkit Controllers (P6 Composability, P4 Continuous Reconciliation)
- **source-controller** — pulls git, Helm, OCI, and bucket sources;
emits artifacts (tarballs) with a digest. The source is the
pinned input to reconciliation (P5 versioning by digest).
- **kustomize-controller** — reconciles Kustomization CRDs against
the artifacts from source-controller. Runs continuously (P4).
- **helm-controller** — reconciles HelmRelease CRDs against Helm
charts from source-controller.
- **notification-controller** — emits events and notifications for
sync, health, and source-readiness events (P9).
- **image-automation-controller** (optional) — updates git with new
image tags when a policy matches, closing the "latest image"
loop declaratively.
## Kustomization CRD (P2 Declarative Over Imperative, P4)
- A Kustomization binds "this source" to "this target namespace"
with a reconciliation interval. The reconciler loops
continuously; drift is corrected automatically (P4).
```yaml
apiVersion: kustomize.toolkit.fluxcd.io/v1
kind: Kustomization
metadata:
name: payments-api
namespace: flux-system
spec:
sourceRef:
kind: GitRepository
name: platform
namespace: flux-system
path: ./manifests/prod
targetNamespace: payments
interval: 1m
prune: true
wait: true
healthChecks:
- apiVersion: apps/v1
kind: Deployment
name: payments-api
namespace: payments
```
- `prune: true` deletes resources removed from git. `wait: true`
waits for health checks before declaring the Kustomization ready.
Disable prune for workloads that need manual removal gates.
## HelmRelease CRD (P6 Composability, cross-link helm.md)
- A HelmRelease binds a Helm chart (from a HelmRepository or OCI
source) to target values and a target namespace. helm-controller
renders and applies it. See `domains/kubernetes/helm.md` for the
chart model.
- Pin the chart version in the HelmRepository or the HelmRelease.
Never float `latest` — unversioned charts drift (P5).
## OCI Sources (P5 State is Immutable and Versioned)
- source-controller can pull from OCI registries (Helm charts as
OCI artifacts, or generic OCI repositories). The digest is the
version — immutable by construction (P5).
- OCI sources close the supply-chain loop: the manifest is signed
and immutable in the registry, and Flux pulls it by digest. Cross-
link `domains/security/supply-chain.md` for signed-provenance
principles.
## Reconciliation and Drift (P4 Continuous Reconciliation, P8)
- Flux reconciles on `interval` (default 1m) and on webhook event.
Drift between git and cluster is detected each interval and
corrected (with `prune` + `selfHeal` semantics).
- Hand-edited drift on a Flux-managed resource is overwritten on the
next loop — the hand edit was never truth (P8). The recovery is
to fix git, not to `kubectl apply`.
## Notifications and Events (P9 Failure is Observable and Surfaced)
- notification-controller emits events for source readiness, sync
success/failure, and health transitions. Wire them to Slack,
PagerDuty, or a webhook. Silent drift is the bug (P9).
- Events flow to `domains/observability/metrics.md` via the
notification controller's provider model — sync and health as
first-class signals.
## RBAC and Multi-Cluster (P10 Least Privilege Reconciliation, P4)
- Flux's controllers run with a ServiceAccount in `flux-system`.
Scope that account to the namespaces Flux reconciles. Do not bind
it to `cluster-admin` (P10). See `domains/kubernetes/rbac.md` and
`domains/security/authorization.md`.
- Flux is per-cluster by design (one Flux install per cluster). For
multi-cluster, use one repo with per-cluster paths, or a fleet
tool that bootstraps Flux per cluster. Per-cluster autonomy is a
feature, not a limitation — it bounds the blast radius of a
compromised controller (P4 locality, P10).
## Secrets (P10, cross-link security/secrets)
- Do not store raw Secrets in the GitOps repo. Use the
SOPS-compatible decryption in kustomize-controller, or External
Secrets Operator, so the git store holds encrypted material only.
See `domains/security/secrets.md`.
## ArgoCD vs Flux — Decision Matrix (IDEATE-21, D-039)
| Axis | ArgoCD | Flux |
|------|--------|------|
| Architecture | Monolithic controller + Application CRD | Composable GitOps Toolkit controllers (source, kustomize, helm, notification) |
| Reconciliation unit | Application (one CRD per app) | Kustomization / HelmRelease (one per deploy unit) |
| UI | Web UI + CLI (full dashboard, tree view, diff viewer) | CLI-first; UI via Weave GitOps or FluxUI (add-on) |
| Sync model | Periodic poll or webhook; sync waves + hooks | Poll + webhook; runs continuously, no explicit sync waves |
| Multi-cluster | One ArgoCD manages many clusters (hub-and-spoke) | One Flux per cluster (per-cluster autonomy) |
| Templating in repo | Helm, Kustomize, ksonnet, raw manifests, Jsonnet | Helm, Kustomize, raw manifests |
| RBAC | Built-in RBAC + SSO + AppProjects | Kubernetes RBAC (no built-in RBAC layer) |
| Progressive delivery | Argo Rollouts (sister project, tight integration) | Flagger (sister project, tight integration) |
| Best for | Teams wanting a UI, multi-cluster from one pane, App-of-Apps bootstrapping | Teams wanting composable controllers, per-cluster autonomy, minimal footprint |
| Watch out for | Monolithic controller scaling, UI as ops crutch, AppProject sprawl | No native UI, steeper learning curve, manual multi-cluster orchestration |
- Use Flux when you want composable controllers, per-cluster
autonomy, and a minimal footprint. Use ArgoCD when you want a UI,
central multi-cluster management, and sync-wave ordering.
- Both are CNCF graduated and both implement the OpenGitOps
principles. The choice is architectural fit, not correctness. See
`argocd.md` for the ArgoCD-side perspective.
## What Violates Flux Discipline
| Violation | Principle |
|-----------|-----------|
| CI pipeline with `kubectl` rights pushing to the cluster | P3 Pull, Don't Push |
| HelmRelease with no pinned chart version | P5 State is Immutable and Versioned |
| Flux ServiceAccount bound to `cluster-admin` | P10 Least Privilege Reconciliation |
| Kustomization with no `healthChecks` on a prod app | P9 Failure is Observable and Surfaced |
| No notification provider wired for sync failures | P9 Failure is Observable and Surfaced |
| Raw Secret in the GitOps repo | P10, `domains/security/secrets.md` |
| Manual `kubectl edit` on a Flux-managed resource | P8 Reconcile, Don't Mutate by Hand |
| `interval: 24h` on a prod Kustomization (drift window too wide) | P4 Continuous Reconciliation |
+140
View File
@@ -0,0 +1,140 @@
# Operators — Derived Rules
> Derives from `domains/gitops-operators/first-principles.md`.
> Applies P6 (Operators Encode Domain Knowledge) primarily, with
> P1, P4, P8, P9, P10. Cross-links `domains/kubernetes/workloads.md`
> and `domains/kubernetes/rbac.md` for the underlying controller
> model, and `domains/infrastructure-as-code/modules.md` for the
> module-vs-operator boundary.
## What an Operator Is (P6 Operators Encode Domain Knowledge)
- An Operator is a Kubernetes controller that encodes human
operational knowledge as CRDs plus a control loop. The operator
reconciles a domain-specific resource (a database, a message
queue, a certificate, a ML model) to a desired state.
- The operator is the deepest expression of
`domains/kubernetes/P1 Declarative Desired State`: the domain
knowledge itself is the desired state. A runbook that lives only
in a wiki is operational knowledge that has not been encoded —
the operator is the encoding (P6).
- An operator runs inside the cluster, observes its CRDs, and acts.
It is a pull-based reconciler by construction — see
`domains/gitops-operators/first-principles.md` P3.
## CRDs and Controllers (P2 Declarative Over Imperative, P4 Continuous Reconciliation)
- A CustomResourceDefinition (CRD) defines the schema of the
domain resource. The controller watches instances of that CRD
and reconciles current → desired (P4).
- The CRD is the public contract of the operator. Version it
(`v1alpha1``v1beta1``v1`) and preserve backward
compatibility — see `domains/api/versioning.md` for the general
API-evolution principles. A CRD is an API surface, not an
internal type.
```yaml
apiVersion: postgres.example.com/v1
kind: PostgresCluster
metadata:
name: payments-db
namespace: payments
spec:
replicas: 3
version: "16"
storage:
size: 100Gi
storageClass: fast-ssd
backup:
schedule: "0 2 * * *"
retention: 7d
```
- The controller reconciles this spec: creates StatefulSets, PVCs,
Services, backup CronJobs. The user declares intent; the operator
makes it so (P2, P6).
## The Control Loop (P4 Continuous Reconciliation, P8)
- The loop watches CRD instances, compares current vs desired, and
acts to converge. Drift (a hand-deleted pod, a failed backup) is
detected and corrected each loop (P4).
- An operator-managed resource should not be hand-edited (P8). The
operator owns the subordinate resources (StatefulSets, PVCs); a
manual `kubectl edit` on a subordinate is drift the operator will
overwrite.
## Operator SDK and OLM (P6 Composability, C6)
- The Operator SDK scaffolds a controller from a CRD (Go, Ansible,
Helm). Use it to avoid re-implementing the controller boilerplate.
- Operator Lifecycle Manager (OLM) installs, updates, and manages
operators as first-class cluster components. OLM is the package
manager for operators — the operator analogue of
`domains/kubernetes/helm.md` for workloads.
- An operator published via OLM is a versioned, catalog-tracked
artifact. Pin the operator version; do not float `latest` (P5
applies to operators as much as to manifests).
## When to Write an Operator vs a Helm Chart (P6, C6 Composability)
| Axis | Helm chart | Operator |
|------|-----------|----------|
| Day-2 operations | None — chart installs, you operate | Encoded — operator reconciles lifecycle (backup, resize, failover, upgrade) |
| State | Static manifests | Live control loop watching CRDs |
| Day-1 install | Strong fit — package and install | Overkill if install is all you need |
| Day-2 reconcile | None — drift is manual | Continuous — drift corrected each loop |
| Domain knowledge | Lives in runbooks + on-call | Lives in the controller code |
| Best for | Off-the-shelf apps, stateless services, one-shot deploys | Stateful apps, complex lifecycles, day-2 automation (backup, scale, failover, version upgrades) |
| Watch out for | Templating complexity, no day-2 reconcile | Controller complexity, multi-team maintenance burden, scope creep |
- Write an operator when the day-2 operations (backup, failover,
resize, version upgrade) are non-trivial and repeated. Write a
Helm chart when install is all you need and day-2 is run by a
human or a separate tool.
- Do not write an operator to wrap a Helm chart and call it day-2
automation — that is a Helm chart with extra steps. See
`domains/infrastructure-as-code/modules.md` for the
module-vs-copy boundary (the operator-vs-chart boundary is its
analogue).
## Scope and Responsibility Boundaries (P10 Least Privilege, C6)
- An operator owns one domain. An operator that manages databases
and message queues and certificates is doing three jobs — split
it. Scope creep is the most common operator failure mode (P6
violation: the encoded knowledge is no longer coherent).
- The operator's ServiceAccount must be scoped to the resources it
manages (P10). A database operator that needs `cluster-admin` to
create a StatefulSet has the wrong RBAC — see
`domains/kubernetes/rbac.md` and `domains/security/authorization.md`.
- One operator per CRD family; one ServiceAccount per operator; one
namespace per operator (or a shared `operators` namespace with
strict RoleBindings). Default namespace is for nothing in
production.
## Failure and Observability (P9 Failure is Observable and Surfaced)
- An operator must surface its reconcile status on the CRD
(`status.conditions`, `status.observedGeneration`). A CRD with no
status is an operator that fails silently (P9).
- Wire operator events to notifications and metrics. A failed
backup, a stuck failover, a version-upgrade stall must emit a
signal — see `domains/observability/metrics.md`.
- An operator that reconciles but does not report health is a
black box. The GitOps controller (ArgoCD/Flux) will read it as
`Progressing` forever — write the health check (see `argocd.md`
"Health and Status").
## What Violates Operator Discipline
| Violation | Principle |
|-----------|-----------|
| Operator that manages databases + queues + certs | P6 Operators Encode Domain Knowledge (scope creep) |
| Operator ServiceAccount bound to `cluster-admin` | P10 Least Privilege Reconciliation |
| CRD with no `status.conditions` | P9 Failure is Observable and Surfaced |
| Operator with no health check wired to GitOps | P9, `argocd.md` Health and Status |
| Unversioned CRD (`v1` shipped without alpha/beta) | P5, `domains/api/versioning.md` |
| Manual `kubectl edit` on an operator-managed subordinate | P8 Reconcile, Don't Mutate by Hand |
| Operator that wraps a Helm chart and adds no day-2 logic | P6 (no knowledge encoded) |
| Operator runbook that exists only in a wiki | P6 Operators Encode Domain Knowledge |
@@ -0,0 +1,177 @@
# Progressive Delivery — Derived Rules
> Derives from `domains/gitops-operators/first-principles.md`.
> Applies P7 (Progressive Delivery is Reversible by Construction)
> primarily, with P4, P9. Cross-links `domains/devops/first-principles.md`
> P4 Rollback First and P5 Progressive Delivery, and
> `domains/observability/metrics.md` for the analysis signals.
## What Progressive Delivery Is (P7 Reversible by Construction)
- Progressive delivery shifts traffic in stages (canary, blue-green)
gated by analysis (metrics, counters, error rates). Each stage is
metric-checked; a failed gate aborts the rollout and reverts to
the prior stable version. Promotion without a rollback path is a
violation (P7).
- Progressive delivery is the GitOps extension of
`domains/devops/P5 Progressive Delivery` and
`domains/kubernetes/P10 Roll Forward, Roll Back`. The k8s rolling
update is the floor; progressive delivery adds metric-gated
promotion and one-command abort.
- Two sister projects dominate: **Argo Rollouts** (Argo ecosystem)
and **Flagger** (Flux ecosystem). Both implement the same pattern
— a Rollout CRD replaces a Deployment, an analysis drives the
gates, an abort reverts traffic.
## The Rollout CRD (P2 Declarative Over Imperative, P7)
- A Rollout (Argo Rollouts) or Canary/Flag (Flagger) is a CRD that
replaces the Deployment as the reconciled resource. It declares
the strategy (canary, blue-green), the traffic split, and the
analysis gates. The controller reconciles traffic and pods to
match.
```yaml
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: payments-api
namespace: payments
spec:
replicas: 10
selector:
matchLabels:
app: payments-api
template:
metadata:
labels:
app: payments-api
spec:
containers:
- name: api
image: registry.example.com/payments-api:1.2.3
strategy:
canary:
trafficRouting:
istio:
virtualService:
name: payments-vs
routes: [primary]
steps:
- setWeight: 5
- pause: { duration: 2m }
- analysis:
templates:
- templateName: success-rate
- setWeight: 25
- pause: { duration: 5m }
- analysis:
templates:
- templateName: success-rate
- setWeight: 50
- pause: { duration: 5m }
- setWeight: 100
```
- Each `setWeight` shifts traffic; each `pause` holds for
observation; each `analysis` runs a metric gate. A failed
analysis aborts the rollout and reverts traffic to the stable
ReplicaSet (P7).
## Canary vs Blue-Green (P7, C3 Simplicity)
| Strategy | Mechanism | Cost | Best for |
|----------|-----------|------|----------|
| Canary | Shift a small % of traffic to the new version; increase on gate success | Low (few new pods) | Most production rollouts; metric-gated, gradual |
| Blue-Green | Run two full environments; switch traffic all-at-once | High (2× capacity) | Schema-breaking changes, instant rollback, low-frequency deploys |
- Canary is the default — it is reversible by construction (P7)
and economical (C8). Blue-green is for changes that cannot be
partial (a breaking schema migration, a full cutover).
- A canary with no analysis gate is a slow blue-green — it is not
progressive delivery. The gate is what makes it progressive (P7).
## Analysis Templates (P9 Failure is Observable and Surfaced, P7)
- An AnalysisTemplate declares the metric query, the success
threshold, and the count of samples. The rollout controller runs
the analysis at each gate; a failed analysis aborts the rollout.
```yaml
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: success-rate
namespace: payments
spec:
metrics:
- name: success-rate
interval: 1m
successCondition: result[0] >= 0.99
failureLimit: 2
provider:
prometheus:
address: http://prometheus.observability:9090
query: |
sum(rate(http_requests_total{job="payments-api",code!~"5.."}[2m]))
/
sum(rate(http_requests_total{job="payments-api"}[2m]))
```
- `successCondition` is the gate; `failureLimit` is the tolerance
for transient blips. A single failed sample aborts immediately if
`failureLimit: 0`; tolerate noise with `failureLimit: 2`.
- The metric is the abort signal — see `domains/observability/metrics.md`
for the SLI/SLO discipline that makes the gate meaningful. A gate
on an undefined SLO is a gate on noise.
## Argo Rollouts vs Flagger (P6 Composability, P7)
| Axis | Argo Rollouts | Flagger |
|------|---------------|---------|
| Ecosystem | Argo (ArgoCD sister project) | Flux (Flux sister project) |
| CRD | `Rollout` (replaces `Deployment`) | `Canary` / `Flag` (wraps a `Deployment`) |
| Traffic providers | Istio, NGINX, ALB, SMI, Traefik, Ambassador | Istio, NGINX, Linkerd, SMI, App Mesh, Gloo, Contour |
| Analysis sources | Prometheus, Datadog, Wavefront, NewRelic, CloudWatch, Graphite, Kayenta | Prometheus, Datadog, CloudWatch, Stackdriver, Elasticsearch, Graphite |
| Integration | Tight with ArgoCD (UI shows rollout) | Tight with Flux (events via notification-controller) |
| Learning curve | Rollout CRD replaces Deployment (migration cost) | Wraps existing Deployment (lower migration cost) |
| Best for | ArgoCD shops wanting rollout in the Argo UI | Flux shops wanting progressive delivery with minimal migration |
- Both implement the same pattern. The choice follows your GitOps
controller — Argo Rollouts with ArgoCD, Flagger with Flux. Mixing
is possible but not idiomatic.
## Abort and Rollback (P7 Reversible by Construction, P5)
- An abort reverts traffic to the stable ReplicaSet immediately. A
rollout without a tested abort is a prototype (P7).
- The abort must be one-command (or one-gate-failure). A
progressive delivery that requires manual rollback steps has
lost the "reversible by construction" property — it is a deploy
with extra steps.
- Test the abort path in staging. An abort that has never been
exercised will fail when you need it most — see
`domains/devops/first-principles.md` P4 Rollback First.
## Observability (P9 Failure is Observable and Surfaced)
- Progressive delivery is only as good as its metrics. A rollout
gated on a metric that is not tracked is ungated — the gate is
theater (P9).
- Wire rollout status (phase, weight, analysis result) to
notifications and dashboards. A stalled rollout with no signal is
silent drift (P9). See `domains/observability/metrics.md`.
- Cross-link `domains/kubernetes/workloads.md` for the underlying
Deployment/ReplicaSet model that progressive delivery replaces.
## What Violates Progressive Delivery Discipline
| Violation | Principle |
|-----------|-----------|
| Canary with no analysis gate | P7 Progressive Delivery is Reversible by Construction |
| Rollout with no tested abort path | P7, `domains/devops/P4 Rollback First` |
| Analysis gate on an undefined SLO | P9 Failure is Observable and Surfaced |
| Blue-green with no 2× capacity budget | C8 Economy (blue-green is a cost decision) |
| Rollout stalled with no notification | P9 Failure is Observable and Surfaced |
| Manual `kubectl` traffic shift on a Rollout-managed service | P8 Reconcile, Don't Mutate by Hand |
| `failureLimit: 0` on a noisy metric (constant false aborts) | P4 Continuous Reconciliation (gate noise tolerance) |
+173
View File
@@ -0,0 +1,173 @@
# Internationalization (i18n) — First Principles
> Grounded in Unicode ICU + CLDR, W3C i18n WG, BCP 47 / RFC 5646,
> ICU MessageFormat / FormatJS / i18next / Mozilla Fluent, the
> JavaScript `Intl` API, and WCAG 2.1 AA. The developer's language is
> one locale among many, not the neutral form.
## 1. The Principles
### P1. Source Language is a Locale, Not the Default
The developer's own language is one locale among many — it is not the
"neutral" or "unlocalized" form of the product. Strings are extracted
from day one, addressed by key, and routed through a locale resource
layer even when only one locale is populated. Treating the source
language as the default produces hidden concatenations, hardcoded
grammar assumptions, and a translation debt that compounds until the
first second locale arrives — at which point the fix is a rewrite, not
a patch. The source locale is `en-US` (or whatever the team writes in);
it is not `null`. This is the i18n angle on `domains/uiux/copywriting.md`:
copy lives in resources, not in code.
### P2. Locale Identifiers are Standardized
Use BCP 47 language tags (`en-US`, `ar-EG`, `zh-Hans-CN`, `pt-BR`).
No ad-hoc locale codes, no two-letter-only hacks, no invented keys.
The tag carries language, script (when needed), and region (when
needed); it is the contract between the resource layer, the
formatting layer, and the runtime. A locale identifier that is not
BCP 47 is a key that cannot be resolved by any standard tool, which
is a correctness violation. Cross `domains/data/schema-design.md`:
locale identifiers are a data shape with a defined vocabulary.
### P3. Resources are External, Not Inline
User-facing strings live in locale resource files (`.po`, JSON,
Fluent `.ftl`, ICU Resource Bundle), never concatenated inline in
code. Inline strings are invisible to the translation pipeline,
unversionable as a unit, and untestable for completeness. String
concatenation in code (`"Welcome, " + name + "!"`) is the cardinal
violation: it bakes in source-language grammar and breaks for every
locale with different word order. Resources are the boundary; code
addresses strings by key, the resource layer resolves the key to the
locale. This is the i18n angle on C4 Locality: strings and their
locale-specific consequences live together in the resource, not
scattered across code.
### P4. Plural and Gender are Parameterized
Plural forms, gender, and select are expressed with ICU MessageFormat
(or an equivalent parameterized formatter), never with `if (n == 1)`
branching in code. Plural rules are locale-specific — English has
one/other, Arabic has six categories, Russian has three — and a
hand-rolled branch encodes exactly one locale's rules while pretending
to be universal. The formatter is the contract; the resource carries
the variants; the code passes the count and lets the formatter choose.
A `if (n == 1)` plural is a C1 (Correctness) violation masquerading as
a shortcut.
### P5. Formatting is Locale-Aware
Dates, times, numbers, currencies, units, and relative time are
formatted via ICU / CLDR / the JavaScript `Intl` API — never
hand-rolled. A hand-rolled date formatter encodes one locale's
conventions and silently produces wrong output for every other locale
(mm/dd/yyyy vs dd/mm/yyyy is the canonical failure). CLDR is the
source of truth for locale data; `Intl` is the runtime that exposes
it. Formatting correctness is observable: a misformatted date is a
wrong answer in the user's locale, even if it is "right" in the
developer's. Cross `domains/api/error-responses.md` for localized
error messages at API boundaries.
### P6. Text Direction is a Layout Primitive
RTL and bidi are first-class layout concerns, not a CSS afterthought.
Logical CSS properties (`margin-inline-start`, `padding-block-end`,
`inset-inline-end`) over physical (`margin-left`, `padding-top`). The
`dir` attribute is set on the document and on subtrees; the bidi
algorithm (UAX #9) handles inline reordering. A layout that assumes
LTR is a layout that is wrong for `ar`, `he`, `fa`, `ur`, and any
RTL-mixed context. Text direction is not a skin — it is a structural
property of the layout, and fixing it late is a rewrite. This is the
i18n angle on `domains/uiux/accessibility.md`: RTL support is an
accessibility concern for non-Latin-script users.
### P7. Layout Accommodates Expansion
Translated text expands and contracts — German is ~30% longer than
English, Japanese often shorter, RTL mirroring shifts every visual
anchor. Layouts are flexible: no fixed pixel widths for text, no
truncation without an ellipsis-and-title strategy, no
`white-space: nowrap` on translatable strings. A layout that breaks
on a 30% expansion is a layout that is wrong for most of the world's
locales. Designing for the worst-case expansion up front is cheaper
than reworking every screen when the first long-form locale ships.
### P8. Pseudo-Locales Test Early
Test with pseudo-locales (accented, lengthened, RTL-mirrored, brack-
enclosed) before real translations arrive. A pseudo-locale run
surfaces hardcoded strings, layout overflow, broken concatenation,
and LTR assumptions while the fix is still cheap — the translator
hasn't been paid yet, and the string freeze hasn't happened. Finding
these bugs after real translation is a C5 (Reversibility) violation:
the cost of undoing is now a re-translation. Cross
`domains/testing/fixtures.md` and `domains/testing/pyramid.md` for
where pseudo-locales sit in the testing pyramid.
### P9. Images and Icons are Cultural
Icons, colors, gestures, and imagery are locale-sensitive. A
mailbox icon means "email" in the US and "mail" in Japan — but a
green checkmark means "correct" in the West and "incorrect" in some
East Asian contexts. A thumbs-up is positive in much of the world
and an insult in parts of the Middle East. Avoid locale-bound symbols
as universal; parameterize imagery per locale where the symbol is not
globally neutral. Icons are not a universal language; they are a
locale with a picture. This is a C2 (Clarity) concern: an icon whose
meaning changes by locale is unclear to the reader it was not drawn
for.
### P10. Translation is Reversible and Versioned
Resource files are versioned alongside code; a bad translation is a
rollback, not a hot-patch. Every locale resource has a history
(what shipped when), a provenance (which translator / which service),
and a rollback path. A translation that breaks the UI is reverted to
the prior resource version, the same way a code regression is
reverted to the prior commit. Translations without version history
are anecdote, not artifact — you cannot tell what changed, when, or
why. This is the i18n angle on C5 Reversibility applied to the
resource layer.
## 2. Core Principle Trace
Each i18n P-rule derives from one or more core C-rules (C1C8). The
matrix extension lands in P4 of the v0.3 plan; the traces below are
authoritative.
| P-rule | Core | Why |
|--------|------|-----|
| P1 Source Language is a Locale, Not the Default | C2, C1 | Clarity of locale intent; correctness of treating source as one-of-many |
| P2 Locale Identifiers are Standardized | C2, C6 | Clarity of a standard vocabulary; composability with standard tools |
| P3 Resources are External, Not Inline | C4, C6 | Locality of strings and their locale consequences; composability of the resource layer |
| P4 Plural and Gender are Parameterized | C1, C6 | Correctness of locale-specific plural rules; composability of the formatter contract |
| P5 Formatting is Locale-Aware | C1, C7 | Correctness of formatted output; observability of format correctness |
| P6 Text Direction is a Layout Primitive | C1, C4 | Correctness of layout for RTL; locality of direction with the text it governs |
| P7 Layout Accommodates Expansion | C8, C3 | Economy of rework; simplicity of flexible layouts over per-locale overrides |
| P8 Pseudo-Locales Test Early | C7, C5 | Observability of i18n defects early; reversibility of fixing before translation |
| P9 Images and Icons are Cultural | C1, C2 | Correctness of locale-appropriate symbols; clarity of meaning across locales |
| P10 Translation is Reversible and Versioned | C5 | Reversibility of the resource layer |
## 3. What Violates These Principles
| Violation | Principle Breached |
|-----------|-------------------|
| A user-facing string hardcoded in source | P3 Resources are External, Not Inline |
| `"Welcome, " + name + "!"` string concatenation | P3 Resources are External, Not Inline |
| `if (n == 1) { return "item"; } else { return "items"; }` | P4 Plural and Gender are Parameterized |
| A locale code like `en_us` or `english` instead of `en-US` | P2 Locale Identifiers are Standardized |
| A hand-rolled date formatter (`getMonth() + 1 + "/" + getDay()`) | P5 Formatting is Locale-Aware |
| `margin-left: 10px` on a translatable layout | P6 Text Direction is a Layout Primitive |
| A fixed-width text container that overflows on German | P7 Layout Accommodates Expansion |
| First i18n test runs against real translations, not pseudo-locales | P8 Pseudo-Locales Test Early |
| A thumbs-up icon shipped as universally positive | P9 Images and Icons are Cultural |
| Resource files with no git history or no rollback path | P10 Translation is Reversible and Versioned |
| The source language treated as the "unlocalized" default | P1 Source Language is a Locale, Not the Default |
## 4. Relationship to Other Domains
i18n is the locale-awareness layer that `domains/uiux/` consumes and
that `domains/api/` surfaces at boundaries. It borrows the testing
discipline of `domains/testing/` and the data-shape discipline of
`domains/data/`. Cross-links are one-directional (per D-026 extended):
- `domains/uiux/copywriting.md` ← P1, P3 (strings live in resources)
- `domains/uiux/accessibility.md` ← P6 (RTL is an a11y concern for non-Latin users)
- `domains/uiux/components.md` ← P6, P7 (layout primitives that survive direction and expansion)
- `domains/api/error-responses.md` ← P5 (localized error messages)
- `domains/data/schema-design.md` ← P2, P3 (locale data shapes)
- `domains/testing/fixtures.md` ← P8 (pseudo-locale fixtures)
- `domains/testing/pyramid.md` ← P8 (pseudo-locale tier mapping)
- `domains/testing/first-principles.md` ← P8 (testing discipline for locale)
+141
View File
@@ -0,0 +1,141 @@
# Formatting — Derived Rules
> Derives from `domains/i18n/first-principles.md`. Covers P2 (Locale
> Identifiers Standardized), P4 (Plural/Gender Parameterized), and P5
> (Formatting is Locale-Aware). Referenced by `locale-resources.md`
> (the formatter resolves the messages) and `testing-i18n.md` (the
> formatted output is what snapshots assert).
## Formatting is Locale-Aware (P5 Formatting is Locale-Aware)
- Dates, times, numbers, currencies, units, and relative time are
formatted via ICU / CLDR / the JavaScript `Intl` API — never
hand-rolled. CLDR is the source of truth for locale data; `Intl`
is the runtime that exposes it.
- A hand-rolled formatter encodes one locale's conventions and
silently produces wrong output for every other locale. The
canonical failure is date format: `mm/dd/yyyy` (US) vs
`dd/mm/yyyy` (most of the world) vs `yyyy-mm-dd` (ISO, sortable).
Picking one and calling it done is a correctness violation in
every locale it is wrong for.
## BCP 47 Tags Drive Formatting (P2 Locale Identifiers Standardized)
- Every formatter takes a BCP 47 locale tag. The tag is the contract
between the resource layer and the formatting layer: the same tag
that selects the resource selects the formatter.
- A locale tag that is not BCP 47 cannot be resolved by `Intl`, ICU,
or CLDR — the formatter returns the runtime default, which is the
developer's locale, not the user's. This is why P2 is a
prerequisite of P5: you cannot format for a locale you cannot name.
## The Intl Surface (ICU/CLDR in the Browser and Node)
| API | Formats | Notes |
|-----|---------|-------|
| `Intl.DateTimeFormat` | Dates, times, date+time, time zones | Calendar (`buddhist`, `hebrew`, `islamic`), numbering system (`arab`, `hanidec`) via locale tag extensions |
| `Intl.NumberFormat` | Numbers, currencies, units, percent | Notation (`compact`, `scientific`), grouping, sign display |
| `Intl.RelativeTimeFormat` | "3 days ago", "in 2 months" | Locale-specific phrasing; numeric vs auto |
| `Intl.PluralRules` | Plural category for a count | `one`, `few`, `many`, `other`, `zero`, `two` per CLDR — the engine ICU MessageFormat uses |
| `Intl.ListFormat` | "a, b, and c" | Conjunction / disjunction / unit lists, locale-specific separators |
| `Intl.Collator` | Locale-aware string sorting | Strength (`base`, `accent`, `case`); numeric collation |
- All of these are built on ICU/CLDR; they are the runtime baseline.
Use them. A `moment.js`-style hand-rolled format string
(`"MM/DD/YYYY"`) is a relic of the pre-`Intl` era and a P5
violation in any locale-aware code path.
## Dates and Times
```
// Correct — Intl, locale-aware
new Intl.DateTimeFormat("ar-EG", {
dateStyle: "full",
timeStyle: "short",
}).format(new Date());
// "الأربعاء، ٧ نوفمبر ٢٠٢٤، ٣:١٥ م"
// Wrong — hand-rolled, source-locale only
const d = new Date();
const s = (d.getMonth() + 1) + "/" + d.getDate() + "/" + d.getFullYear();
// "11/7/2024" — meaningless in most locales
```
- Time zones are not locales. A locale tells you *how to format* a
timestamp; a time zone tells you *what instant* it refers to. Do
not derive one from the other (`ar-EG` is not a time zone).
Format with the user's locale; render in the user's time zone;
store in UTC.
## Numbers, Currencies, Units
```
new Intl.NumberFormat("de-DE", { style: "currency", currency: "EUR" })
.format(1234.56); // "1.234,56 €"
new Intl.NumberFormat("ar-EG", { style: "currency", currency: "EGP" })
.format(1234.56); // "١٬٢٣٤٫٥٦ ج.م."
new Intl.NumberFormat("en-US", { style: "unit", unit: "kilometer-per-hour" })
.format(100); // "100 km/h"
```
- The currency code (`EUR`, `EGP`, `USD`) is ISO 4217; the locale
determines the symbol, grouping, and placement. A hand-rolled
`"$" + amount` is wrong for `de-DE` (symbol, grouping, placement
all differ).
## Plural Rules (P4 Plural/Gender Parameterized)
- `Intl.PluralRules` returns the CLDR plural category for a count in
a given locale. ICU MessageFormat uses this category to select the
variant from the resource (`locale-resources.md`).
- Never branch on the raw count in code. The count goes to the
formatter; the formatter consults `PluralRules` for the locale;
the resource carries the variant for that category.
```
// ICU MessageFormat (FormatJS)
new Intl.MessageFormat(
"{count, plural, one {# item} other {# items}}",
"en-US"
).format({ count: 1 }); // "1 item"
// ar-EG — six categories; the code is identical, only the
// resource differs.
```
## Gender and Select
- ICU MessageFormat also supports `{gender, select, male {...} female {...} other {...}}`
for gendered agreement and `{case, select, ...}` for general
disjunction. These live in the resource, not in code branches.
- A `switch (gender)` in code that picks a string is the same
violation as `if (n == 1)`: it encodes one locale's grammar in
code and breaks for every locale with different agreement rules.
## What Violates Formatting Discipline
| Violation | Principle |
|-----------|-----------|
| `getMonth() + 1 + "/" + getDay()` hand-rolled date | P5 Formatting is Locale-Aware |
| `"$" + amount` hand-rolled currency | P5 Formatting is Locale-Aware |
| `if (n === 1) "item" else "items"` plural branch | P4 Plural and Gender are Parameterized |
| `moment("MM/DD/YYYY")` format string in locale-aware code | P5 Formatting is Locale-Aware |
| Deriving time zone from locale tag | P2 Locale Identifiers are Standardized |
| A non-BCP-47 tag passed to `Intl` (silently falls back) | P2 Locale Identifiers are Standardized |
| `switch (gender)` selecting strings in code | P4 Plural and Gender are Parameterized |
| Storing timestamps in local time, not UTC | P5 Formatting is Locale-Aware |
## Relationship to Other Domains
- `domains/api/error-responses.md` — API error messages are
formatted for the requesting locale; the error code is stable, the
message is locale-formatted.
- `domains/data/schema-design.md` — locale identifiers, currency
codes, and time zones are data contracts; treat them as schema
(`en-US`, `EUR`, `UTC`), not free text.
- `domains/i18n/locale-resources.md` — the resource layer carries
the parameterized messages this formatter resolves.
- `domains/testing/fixtures.md` — formatted output per locale is the
fixture; snapshot tests assert against it.
+138
View File
@@ -0,0 +1,138 @@
# Locale Resources — Derived Rules
> Derives from `domains/i18n/first-principles.md`. Covers P1 (Source
> Language is a Locale), P2 (Locale Identifiers Standardized), P3
> (Resources External, Not Inline), P4 (Plural/Gender Parameterized),
> and P10 (Translation Reversible and Versioned). Referenced by
> `formatting.md` (strings the formatter resolves) and
> `rtl-bidi.md` (the `dir` the resource layer carries).
## Resources are the Boundary (P3 Resources are External, Not Inline)
- User-facing strings live in locale resource files, addressed by
key. Code references a key; the resource layer resolves the key to
the active locale. The source language is itself a locale
(`en-US`), not a fallback baked into code.
- String concatenation in code (`"Welcome, " + name + "!"`) is the
cardinal violation: it bakes in source-language word order and
breaks for every locale with different grammar. Replace every
concatenation with a parameterized message:
`t("welcome", { name })`.
- The resource is the single place a string lives. Editing a string
in code instead of the resource is a locality violation (C4): the
string and its locale consequences now live apart.
## Resource File Formats
| Format | Shape | When | Notes |
|--------|-------|------|-------|
| `.po` / `.pot` | gettext; msgid → msgstr, plural headers | Server-side, GNU ecosystem, PHP/Python/C | Mature tooling (`xgettext`, `msgmerge`); supports plural categories via header |
| JSON (flat or namespaced) | `{ "key": "value" }` per locale | JS/web, i18next, FormatJS | Simple, machine-readable, but no native plural support — wrap with ICU MessageFormat |
| Fluent `.ftl` | Mozilla FTL; asymmetric, resolver-driven | Browser-grade l10n, asymmetric translations | One message can resolve differently per locale without code changes; supports attributes, selectors |
| ICU Resource Bundle | ICU binary/text resources | ICU-native, JVM, C++ | Tightest integration with ICU formatting/CLDR; steeper tooling |
- None is advocated over the others. The choice is ecosystem fit,
not correctness. All four satisfy P3/P4 when used as the boundary.
- A custom format (a hand-rolled `.csv` of strings) is a violation:
it is unsupported by standard tooling, has no plural grammar, and
cannot compose with `formatting.md`'s ICU layer.
## Key Naming and Namespaces (P2 Locale Identifiers Standardized)
- Locale identifiers are BCP 47 tags (`en-US`, `ar-EG`, `zh-Hans-CN`).
No ad-hoc codes. The resource file is named for its locale:
`en-US.json`, `ar-EG.po`, `ftl/ar-EG/main.ftl`.
- Message keys are stable, semantic, and structured — not prose.
`checkout.cart.item_count` not `"You have 3 items in your cart"`.
A key that is the source string (`t("You have items")`) breaks the
moment the source copy is edited; the key must outlive the copy.
- Namespaces segment by surface (`checkout.*`, `errors.*`, `onboarding.*`)
so that a locale can be loaded incrementally and so that key
collisions across surfaces are impossible. A flat namespace with
thousands of keys is a C2 (Clarity) violation waiting to happen.
## Fallback Chains
- The fallback chain is explicit: requested locale → language-only
(`en` from `en-GB`) → default locale → key itself (last resort).
The default locale is declared once, not re-derived in every call
site.
- A missing key in the requested locale falling back silently to the
source locale is a P3 violation: the user is silently shown the
developer's locale, which is not the locale they asked for. Missing
keys must be observable (see `testing-i18n.md`).
- Fallback is a property of the resource layer, not of individual
components. A component that re-implements fallback is duplicating
a contract (C6 Composability violation).
## Plural and Gender in Resources (P4 Plural/Gender Parameterized)
- Plural variants live in the resource, selected by the formatter,
parameterized by the count. The code passes the count; the resource
carries the variants; the formatter picks the right one per the
locale's CLDR plural rules.
```
// JSON + ICU MessageFormat (FormatJS / i18next)
{
"cart.item_count": "{count, plural, one {# item} other {# items}}"
}
// ar-EG.json — six plural categories per CLDR
{
"cart.item_count": "{count, plural, zero {لا عناصر} one {عنصر واحد} two {عنصران} few {# عناصر} many {# عنصرًا} other {# عنصر}}"
}
```
- `if (n == 1)` branching in code is a violation regardless of
language. Arabic has six plural categories; Russian has three;
English has two. A two-branch `if` encodes exactly one locale's
rules and is wrong for every other.
## Extraction Tooling (P1, P3)
- Strings are extracted mechanically (e.g. `xgettext`, `i18next-
parser`, FormatJS babel plugin), not by hand-tagging. Mechanical
extraction produces a `.pot` template that translators work from;
the template is regenerated on every build.
- A string that cannot be extracted (built at runtime from
fragments) is a P3 violation: it is invisible to the pipeline. If
the extractor cannot see it, neither can the translator.
- The extracted template is versioned (`P10`): the diff between
templates is the change in translatable surface. A template that
is not committed is a contract that is not reviewable.
## Versioning and Rollback (P10 Translation Reversible and Versioned)
- Resource files are committed to git alongside code. A bad
translation is a `git revert` of the resource, not a hot-patch over
the translator's work. Every locale resource has history,
provenance (which translator / which service produced which
commit), and a rollback path.
- A locale resource that is generated by a translation service and
committed without review is a P10 violation: the resource is
versioned but the provenance is opaque. Review the diff the same
way you review a code diff.
## What Violates Locale-Resource Discipline
| Violation | Principle |
|-----------|-----------|
| `t("You have " + n + " items")` concatenation | P3 Resources are External, Not Inline |
| A custom `.csv` string store instead of a standard format | P3 Resources are External, Not Inline |
| Locale file named `english.json` not `en-US.json` | P2 Locale Identifiers are Standardized |
| `if (n == 1) { t("item") } else { t("items") }` in code | P4 Plural and Gender are Parameterized |
| A key equal to the source string (`t("Welcome back")`) | P2 / P10 — keys must outlive copy |
| Silent fallback to the source locale with no signal | P3 Resources are External, Not Inline |
| A runtime-built string the extractor cannot see | P3 Resources are External, Not Inline |
| Resource files committed by a bot with no human review | P10 Translation is Reversible and Versioned |
## Relationship to Other Domains
- `domains/uiux/copywriting.md` — copy lives in resources; UI
microcopy is the source content the resource layer carries.
- `domains/api/error-responses.md` — API error messages are locale-
resource keys resolved at the boundary, not inline strings.
- `domains/data/schema-design.md` — locale identifiers and resource
shapes are a data contract; treat them as schema.
- `domains/i18n/formatting.md` — the formatter resolves the
parameterized message this layer produces.
+117
View File
@@ -0,0 +1,117 @@
# RTL and Bidi — Derived Rules
> Derives from `domains/i18n/first-principles.md`. Covers P6 (Text
> Direction is a Layout Primitive) and P7 (Layout Accommodates
> Expansion). Referenced by `testing-i18n.md` (RTL coverage is an
> e2e tier). Grounded in W3C i18n bidi authoring, UAX #9, and
> `domains/uiux/accessibility.md`.
## Text Direction is a Layout Primitive (P6 Text Direction is a Layout Primitive)
- RTL and bidi are first-class layout concerns, not a CSS
afterthought. The layout is designed for both directions from the
first commit, not retrofitted when an RTL locale ships.
- Logical CSS properties over physical properties, always. The
browser resolves logical → physical from the `dir` attribute; the
code never has to.
| Physical (LTR-only) | Logical (dir-aware) | Resolves to in RTL |
|---------------------|---------------------|--------------------|
| `margin-left` | `margin-inline-start` | `margin-right` |
| `margin-right` | `margin-inline-end` | `margin-left` |
| `padding-left` | `padding-inline-start` | `padding-right` |
| `left: 0` | `inset-inline-start: 0` | `right: 0` |
| `text-align: left` | `text-align: start` | `text-align: right` |
| `float: left` | use flexbox/grid + `inline-start` where supported | mirrored |
- The `dir` attribute is set on the document root (`<html dir="rtl">`)
and on subtrees whose direction differs from the document
(`<span dir="ltr">` for an embedded Latin run). `dir` is the
contract the bidi algorithm (UAX #9) reads; do not fake direction
with `text-align` alone.
## The Bidi Algorithm (UAX #9)
- The Unicode bidi algorithm resolves inline reordering of mixed-
direction runs. The browser applies it; the author's job is to
mark direction correctly, not to reorder by hand.
- A string like `"The price is 15 USD"` in an RTL context renders
with the Latin run `"15 USD"` in LTR within the RTL line — the
algorithm handles it *if* the container's `dir` is set. Without
`dir`, numbers and Latin fragments drift to the wrong edge.
- `dir="auto"` on a container infers direction from the first strong
directional character of its content — useful for user-generated
content whose direction is unknown. `dir="auto"` is not a
replacement for `dir="rtl"` on a known-RTL document.
## Mirroring (Icons, Controls, Diagrams)
- Direction-aware icons mirror in RTL: a "back" arrow pointing left
in LTR points right in RTL. A "refresh" circular arrow does not
mirror. The rule: icons that imply direction mirror; icons that
imply time or rotation do not.
- Use `[dir="rtl"]` selectors or logical icon variants — never
`transform: scaleX(-1)` as a one-off hack scattered across
components. Centralize the mirroring rule (a token, a component
prop) so it is auditable.
- Numbers do not mirror. `"15 USD"` in an RTL line is still
`"15 USD"` left-to-right inside the bidi run; mirroring it to
`"DSU 51"` is a correctness violation.
- Diagrams and flowcharts: a left-to-right process flow in LTR is a
right-to-left flow in RTL. Decide per diagram whether the flow
mirrors (most do) or is direction-neutral (some scientific
schematics).
## Layout Accommodates Expansion (P7 Layout Accommodates Expansion)
- Translated text expands. German is ~30% longer than English;
Japanese is often shorter but taller; RTL mirroring shifts every
visual anchor. Layouts are flexible:
- No fixed pixel widths on translatable text containers.
- No `white-space: nowrap` on translatable strings.
- No `text-overflow: ellipsis` without a `title` carrying the full
string.
- Buttons sized to fit their longest locale variant, not the
source.
- A layout that breaks at +30% width is a layout that is wrong for
most of the world's locales. Designing for the worst case up front
is cheaper than reworking every screen when the first long-form
locale ships.
## Common Pitfalls
| Pitfall | Why it breaks | Fix |
|---------|---------------|-----|
| `margin-left` everywhere | In RTL the start is the right; `margin-left` leaves the right side unstyled | `margin-inline-start` |
| `text-align: left` for "default" alignment | In RTL the default is right; `left` pins content to the wrong edge | `text-align: start` |
| Icons hardcoded to LTR orientation | "Back" arrow points the wrong way in RTL | Mirror direction-implying icons via `[dir="rtl"]` |
| Numbers mirrored with the layout | Numbers are LTR inside RTL; mirroring produces garbage | Leave number runs LTR; the bidi algorithm handles embedding |
| Fixed `width: 120px` on a button | German button label overflows and truncates | `min-width` + `max-width` + flex; let content size |
| `position: absolute; left: 0` | Pins to the physical left in both directions | `inset-inline-start: 0` |
| Fake direction with `text-align` only | The bidi algorithm reads `dir`, not `text-align`; mixed runs reorder wrong | Set `dir` on the container |
## What Violates RTL/Bidi Discipline
| Violation | Principle |
|-----------|-----------|
| A layout with no `dir` attribute, assuming LTR | P6 Text Direction is a Layout Primitive |
| `margin-left` / `left: 0` / `text-align: left` throughout | P6 Text Direction is a Layout Primitive |
| A "back" arrow that points left in the RTL build | P6 Text Direction is a Layout Primitive |
| Numbers mirrored to read right-to-left | P6 Text Direction is a Layout Primitive |
| `width: 100px` on a text container that overflows in German | P7 Layout Accommodates Expansion |
| `white-space: nowrap` on a translated label | P7 Layout Accommodates Expansion |
| `dir` faked with `text-align` and no `dir` attribute | P6 Text Direction is a Layout Primitive |
| No RTL build until the first RTL locale ships | P6 Text Direction is a Layout Primitive |
## Relationship to Other Domains
- `domains/uiux/accessibility.md` — RTL support is an accessibility
concern for non-Latin-script users; WCAG 2.1 AA requires that
direction be set correctly.
- `domains/uiux/components.md` — components are built with logical
properties so they survive direction and expansion without per-
locale overrides.
- `domains/i18n/testing-i18n.md` — RTL coverage is an e2e-tier
test; pseudo-locale mirroring surfaces direction bugs early.
- `domains/i18n/locale-resources.md` — the `dir` is part of the
locale's metadata, carried alongside the resource bundle.
+142
View File
@@ -0,0 +1,142 @@
# Testing i18n — Derived Rules
> Derives from `domains/i18n/first-principles.md`. Covers P8
> (Pseudo-Locales Test Early) and the testing-discipline angle on
> P3 (Resources External), P5 (Formatting Locale-Aware), P6 (Text
> Direction), and P10 (Translation Versioned). Referenced by
> `locale-resources.md` (missing-key detection) and `rtl-bidi.md`
> (RTL coverage tier).
## Pseudo-Locales Test Early (P8 Pseudo-Locales Test Early)
- A pseudo-locale is a synthetic locale that transforms the source
strings to surface i18n defects before real translations arrive.
Three transforms cover the three defect classes:
| Pseudo-locale | Transform | Surfaces |
|---------------|-----------|----------|
| `en-XA` (accented) | `Wêlcômê tô thê çhêckôût` | Strings not extracted (raw source appears), encoding bugs |
| `en-XB` (lengthened / "long") | `Wᴇʟᴄᴏᴍᴇ ᴛᴏ ᴛʜᴇ ᴄʜᴇᴄᴋᴏᴜᴛ──────` (~30% longer, bracketed) | Layout overflow, fixed widths, truncation |
| `en-XC` (RTL-mirrored) | Source rendered with `dir="rtl"` and a Latin-in-RTL run | LTR-only layout assumptions, physical CSS properties |
- Pseudo-locale tests are cheap: they run against source strings, no
translator involved, no string freeze required. A failing pseudo-
locale run is a bug found at the cheapest possible point in the
pipeline. Finding the same bug after real translation is a C5
(Reversibility) violation: the fix now costs a re-translation.
## Pseudo-Locale → Testing Pyramid Mapping (IDEATE-28)
- The testing pyramid (`domains/testing/pyramid.md`) has three tiers;
i18n tests map to each tier with a distinct signal. The mapping is
deliberate: each tier catches a different class of defect, and
skipping a tier leaves a blind spot.
| Pyramid Tier | i18n Test | Defect Caught | Tooling Shape |
|--------------|-----------|---------------|---------------|
| **Unit** | Missing-key detection | A key referenced in code but absent from the resource bundle; a key present in the source locale but missing from a target locale | Static scan over the resource bundle + code AST; runs per file, no runtime |
| **Integration** | Snapshot per locale | Formatted output for a fixture input differs across locales in a way that breaks the contract (wrong plural, wrong date, overflow) | Render a known fixture through the formatter per locale; snapshot-diff against the recorded baseline |
| **e2e** | RTL coverage | The app renders and is navigable in `dir="rtl"`; no layout overflow, no off-screen controls, no LTR-pinned anchors | Browser-driven run against the `en-XC` pseudo-locale (or a real RTL locale); assert on layout, not just text |
- Unit is the broad base (fast, runs on every commit), e2e is the
narrow top (slow, runs on PR merge). Integration sits between.
This mirrors `domains/testing/pyramid.md` exactly — i18n is not a
special case; it is a domain that uses the same tiers.
## Unit Tier — Missing-Key Detection (P3 Resources External)
- A static scan compares the set of keys referenced in code against
the keys present in each locale bundle. A key in code but not in
`en-US` is a P3 violation (the string is not in the resource
layer). A key in `en-US` but not in `ar-EG` is a coverage gap —
the missing-key scan flags it before the locale ships.
- Missing keys fail the build, not the runtime. A missing key that
surfaces only when a user switches locale is a defect found in
production, which is the most expensive place to find it.
```
// tool output (illustrative)
// missing-key scan
[FAIL] ar-EG: key "checkout.cart.item_count" referenced in code,
absent from ar-EG.json
[FAIL] en-US: key "checkout.cart.total" referenced in Checkout.tsx:42,
absent from en-US.json (not extracted)
[PASS] en-US, ar-EG, de-DE, zh-Hans-CN: all other keys present
```
## Integration Tier — Snapshot per Locale (P5 Formatting Locale-Aware)
- For a fixed fixture input, render the formatted output per locale
and snapshot it. A change in the snapshot is either an intended
change (new CLDR data, new copy) or a regression.
- The snapshot is per locale, not per format string. The same
fixture (`{ count: 1, currency: "EUR", date: 2024-11-07 }`)
produces different snapshots for `en-US`, `de-DE`, `ar-EG` — and
that difference is the assertion. A locale whose snapshot matches
the source locale's is a red flag: the formatter is not actually
locale-aware.
```
// snapshot — checkout.cart (fixture: count=1, currency=EUR, date=2024-11-07)
// en-US
"1 item · €1,234.56 · 11/7/2024"
// de-DE
"1 Artikel · 1.234,56 € · 07.11.2024"
// ar-EG
"عنصر واحد · ١٬٢٣٤٫٥٦ € · ٧/١١/٢٠٢٤"
```
- Snapshots are reviewed, not rubber-stamped. A snapshot diff that
changes the plural form for `ar-EG` is either a CLDR update (verify)
or a regression (revert).
## e2e Tier — RTL Coverage (P6 Text Direction is a Layout Primitive)
- A browser-driven run against `en-XC` (or a real RTL locale like
`ar-EG`) asserts that the app is navigable in RTL: no overflow, no
off-screen controls, no LTR-pinned anchors. The assertion is on
layout, not on text — text correctness is the integration tier's
job.
- RTL e2e is the narrow top of the i18n pyramid: it is slow, it
requires a browser, and it catches the defects the lower tiers
cannot (the interaction of `dir` with the real layout engine). It
runs on PR merge, not on every commit.
## Snapshot Discipline (P10 Translation Reversible and Versioned)
- Snapshots are versioned in git. A snapshot that changes because of
a real translation update is a committed diff, reviewed like a
code change. A snapshot that changes because of a regression is a
`git revert`.
- A snapshot that is regenerated and committed without review is a
P10 violation: the snapshot is versioned but the provenance is
opaque. The same discipline applies to snapshots as to resources
(`locale-resources.md`).
## What Violates i18n Testing Discipline
| Violation | Principle |
|-----------|-----------|
| First i18n test runs against real translations, not pseudo-locales | P8 Pseudo-Locales Test Early |
| No missing-key scan — gaps surface only at runtime in production | P3 Resources are External, Not Inline |
| Snapshot per locale that matches the source locale's snapshot | P5 Formatting is Locale-Aware |
| No RTL e2e — "we'll test RTL when we ship an RTL locale" | P6 Text Direction is a Layout Primitive |
| Snapshots regenerated and committed without review | P10 Translation is Reversible and Versioned |
| i18n tests only at e2e (no unit/integration tier) | pyramid inversion — `domains/testing/pyramid.md` |
| Pseudo-locale run skipped because "it's not a real locale" | P8 Pseudo-Locales Test Early |
## Relationship to Other Domains
- `domains/testing/pyramid.md` — the pseudo-locale → pyramid mapping
mirrors this domain's unit / integration / e2e tiers exactly.
- `domains/testing/fixtures.md` — locale fixtures (a fixed input
rendered per locale) are the fixture shape for the integration
tier.
- `domains/i18n/locale-resources.md` — missing-key detection is the
unit-tier scan over the resource bundle this doc defines.
- `domains/i18n/formatting.md` — the integration-tier snapshot
asserts against the formatter's output.
- `domains/i18n/rtl-bidi.md` — the e2e tier exercises the layout
rules this doc establishes.
- `domains/uiux/accessibility.md` — RTL coverage is an a11y
concern; an untested RTL build is an untested a11y surface.
+222
View File
@@ -0,0 +1,222 @@
# Bad Example: Compliance Audit Log (Two Breaches)
> An audit logging implementation that violates **two** Atelier
> compliance principles in one example (per IDEATE-26, D-044):
> **P1** (Audit Logs are Append-Only) — a mutable audit log with
> routine `DELETE`/`UPDATE` "cleanup" — and **P9** (Secrets and
> Sensitive Data are Redacted in Audit) — a database password leaked
> into an audit record. Each violation is cited, then fixed.
## The Code
```python
# audit_log.py — the audit sink, stored in a mutable Postgres table
import psycopg2, datetime
# P1 VIOLATION: the audit log is a regular mutable table. There is no
# write-once protection, no immutable bucket, no hash-chaining.
# Any DB user with UPDATE/DELETE can rewrite history.
CREATE_TABLE = """
CREATE TABLE audit_log (
id BIGSERIAL PRIMARY KEY,
timestamp TIMESTAMPTZ NOT NULL,
event TEXT NOT NULL,
actor TEXT NOT NULL,
target TEXT,
payload JSONB,
request_id TEXT
);
-- No row-level immutability. No trigger preventing UPDATE/DELETE.
"""
def write_event(event, actor, target=None, payload=None, request_id=None):
conn = psycopg2.connect(os.environ["DATABASE_URL"])
conn.execute(
"INSERT INTO audit_log (timestamp, event, actor, target, payload, request_id) "
"VALUES (%s, %s, %s, %s, %s, %s)",
(datetime.datetime.utcnow(), event, actor, target,
json.dumps(payload), request_id),
)
# P1 VIOLATION (continued): "cleanup" that mutates the audit log.
# A routine job deletes records older than 30 days to "save space"
# and updates records to "fix typos in the actor field."
def cleanup_audit_log():
conn = psycopg2.connect(os.environ["DATABASE_URL"])
# DELETE — an audit record is destroyed. This is tampering,
# dressed as housekeeping.
conn.execute("DELETE FROM audit_log WHERE timestamp < NOW() - INTERVAL '30 days'")
# UPDATE — an audit record is rewritten. The "fix" is the
# violation; the original actor is lost.
conn.execute("UPDATE audit_log SET actor = 'admin' WHERE actor LIKE 'svc-%'")
```
```python
# The call site that leaks a secret into the audit log.
def read_config(key):
# ... fetches a secret from the secrets manager ...
value = secrets_manager.get(key) # e.g. the raw DB password
# P9 VIOLATION: the raw secret value is written into the audit
# payload. The append-only log is now a secret store.
write_event(
event="config.read",
actor="api-server",
target={"kind": "secret", "id": key},
payload={"value": value}, # <- the secret, in plaintext
request_id=req.id,
)
return value
```
The resulting audit record:
```json
{
"id": 48213,
"timestamp": "2026-08-05T09:12:03Z",
"event": "config.read",
"actor": "api-server",
"target": {"kind": "secret", "id": "db-password"},
"payload": {"value": "p@ssw0rd-sup3r-s3cr3t-plaintext"},
"request_id": "req_91c2"
}
```
A week later, the `cleanup_audit_log` job `DELETE`s this record (it
is older than 30 days in the team's "retention" — which is actually a
storage-economy decision, not a policy), and `UPDATE`s every
`svc-*` actor to `admin`. The secret was in the log for a week,
readable by anyone with `SELECT` on the table; now the record of it
having been there is gone.
## What Makes It Bad
### Breach 1 — Mutable Audit Log (Compliance P1 Audit Logs are Append-Only)
- The audit log is a regular mutable Postgres table. `DELETE FROM
audit_log WHERE timestamp < ...` and `UPDATE audit_log SET actor =
...` both succeed. The log is a draft, not a record.
- Routine `DELETE` as "cleanup" is the cardinal P1 violation: the
deletion of an audit record is itself an auditable incident, not a
housekeeping task. "We deleted old records to save space" is a P3
(Retention is Policy, Not Storage) violation *and* a P1 violation —
the retention decision is driven by storage cost, and the
mechanism is tampering.
- The `UPDATE` that rewrites `svc-deploy` → `admin` destroys
attribution (a P7 violation stacked on the P1 violation): the
original actor is lost, and the replacement (`admin`) is a shared
identity that could be any of ten engineers.
- **Fix:** the audit sink is append-only *by construction*, not by
policy. Write-once storage (WORM bucket, immutable log stream,
hash-chained ledger) enforces immutability at the substrate.
Retention is a declared policy with a meta-audit of deletions; a
human does not run ad-hoc `DELETE` jobs.
```python
# Fix: write to an append-only sink (illustrative — S3 Object Lock,
# WORM bucket, or a hash-chained ledger). The API has no update /
# delete path; the storage refuses mutation.
def write_event(event, actor, target=None, payload=None, request_id=None):
record = {
"timestamp": datetime.datetime.now(datetime.timezone.utc).isoformat(),
"event": event,
"actor": actor, # the authenticated principal, not "admin"
"target": target,
"payload": redact(payload), # see Breach 2 fix
"request_id": request_id,
"prev_hash": last_hash(), # hash-chaining: tampering breaks the chain
}
record["hash"] = sha256(canonical_json(record))
append_only_sink.write(record) # WORM storage; no update/delete API exists
# Fix: retention is a declared, reviewed policy — not an ad-hoc DELETE.
# When an audit segment ages out, the deletion is itself meta-audited
# in a higher-tier log with the rule that authorized it.
# (See domains/compliance/data-retention.md and audit-logs.md.)
```
- See `domains/compliance/audit-logs.md` (Audit Logs are Append-Only)
and `domains/compliance/first-principles.md` P1.
### Breach 2 — Secret Leaked in Audit Log (Compliance P9 Secrets and Sensitive Data are Redacted in Audit)
- `payload={"value": value}` writes the raw DB password into the
audit record. The append-only log is now a secret store: anyone
with `SELECT` on `audit_log` can read production credentials. The
log is harder to secure than the secrets manager it read from.
- Once the secret is in an append-only log, the remediation is
expensive — rotate the secret *and* rewrite the log's access scope
(you cannot edit the record; it is append-only). Redaction must
happen *at the logging boundary, before the record is written*, not
by opportunistic scrubbing after the fact.
- The redaction policy here is "nothing" — there is no rule for
which fields are redacted, by what mechanism, in which event type.
A redaction rule that lives in no one's head and no code is a P9
violation waiting to happen (and it happened).
- **Fix:** redaction is structural, applied at the logging boundary
before the record reaches the append-only sink. The policy is
itself auditable (which fields, by what rule, in which event).
```python
# Fix: redaction at the boundary. Log the FACT of the action
# (a secret was read), never the CONTENT of the secret.
REDACTED_FIELDS = {"value", "token", "password", "authorization", "secret"}
def redact(payload):
if not isinstance(payload, dict):
return "[REDACTED:non-object]"
out = {}
for k, v in payload.items():
if k.lower() in REDACTED_FIELDS or "secret" in k.lower():
out[k] = "[REDACTED:secret]"
else:
out[k] = v
out["_redaction"] = "secret-value-policy/v1" # the rule is auditable
return out
# The fixed audit record:
# {
# "event": "config.read",
# "actor": "api-server", # the authenticated principal
# "target": {"kind": "secret", "id": "db-password"},
# "payload": {"value": "[REDACTED:secret]"},
# "_redaction": "secret-value-policy/v1",
# "request_id": "req_91c2"
# }
# The fact of the read is logged; the secret never enters the log.
```
- See `domains/compliance/audit-logs.md` (Redaction at the Boundary)
and `domains/compliance/first-principles.md` P9. Cross
`domains/security/secrets.md` — the audit-side redaction is the
complement of secret management.
## The Cascade (Two Breaches Compound)
The two violations compound destructively. The secret enters the
mutable log (P9 breach), where it sits readable by any `SELECT`-holder
for a week. Then the `cleanup` job `DELETE`s the record (P1 breach) —
destroying the evidence that the secret was ever logged, while the
secret itself has already been exposed to every reader of the table.
The `UPDATE` that rewrites `svc-deploy` → `admin` (a P7 attribution
breach stacked on the P1 breach) means that even if a copy of the
record survived, the actor who triggered the secret read is no longer
identifiable. The team cannot answer "who read the DB password and
when" — the log that would answer it was mutated, and the secret it
leaked is now in the wild. This is the worst-case interaction of P1
and P9: a secret leak with no attributable actor and no surviving
record.
## Cross-Domain Links
- `domains/compliance/audit-logs.md` — the append-only guarantee and
the redaction-at-boundary rule this code violates.
- `domains/compliance/first-principles.md` — P1 (Append-Only) and P9
(Redacted) are the two breached principles; P7 (Attributable) is
breached by the `UPDATE` rewrite.
- `domains/compliance/evidence.md` — an audit log that can be
`DELETE`d is not admissible evidence; the append-only guarantee is
what makes it admissible.
- `domains/compliance/data-retention.md` — retention is a declared
policy with meta-audited deletions, not an ad-hoc `DELETE` job.
- `domains/security/secrets.md` — redaction at the logging boundary
is the audit-side complement of secret management.
- `domains/observability/logging.md` — audit logs are structured
logging with an append-only guarantee; the logging primitives
compose here.
+175
View File
@@ -0,0 +1,175 @@
# Bad Example: i18n String Concatenation
> A checkout component that violates Atelier's i18n principles. Each
> violation is cited, then fixed.
## The Code
```typescript
// Checkout.tsx — the cardinal i18n violation
function CartSummary({ itemCount, name, total, currency, date }) {
// P3 VIOLATION: inline string concatenation. The source-language
// word order ("Welcome, {name}! You have {n} items") is baked into
// code. Every locale with different word order is broken.
const welcome = "Welcome, " + name + "!";
// P4 VIOLATION: hand-rolled plural branching. `if (n === 1)` encodes
// exactly English's one/other rule. Arabic (six categories), Russian
// (three), Polish (three) are all wrong.
const items =
itemCount === 1 ? "1 item" : itemCount + " items";
// P5 VIOLATION: hand-rolled currency + date formatting. "$" + total
// is wrong for de-DE (symbol, grouping, placement). The date
// `getMonth() + 1 + "/" + getDay()` is US-only (mm/dd/yyyy).
const price = "$" + total.toFixed(2);
const d = new Date(date);
const dateStr = (d.getMonth() + 1) + "/" + d.getDate() + "/" + d.getFullYear();
return (
<div>
<h1>{welcome}</h1>
<p>{items} · {price} · {dateStr}</p>
</div>
);
}
```
```typescript
// The "resource" file — a custom CSV the team hand-rolled.
// locale,en_us
// welcome_prefix,Welcome,
// item_singular,item
// item_plural,items
//
// This is a P3 violation on its own: a custom format no standard
// tool (xgettext, i18next, FormatJS) can extract from or compose with.
```
The team runs their first i18n test against real Arabic translations —
after the string freeze, after the translator was paid. The Arabic
build renders `"Welcome, محمد!"` with the name on the wrong side of
the comma, `"1 items"` for a single item (Arabic has six plural
categories, not two), and the price as `"$1,234.56"` (Arabic-Egypt
formats as `"١٬٢٣٤٫٥٦ ج.م."`). Every screen is a rewrite, not a patch.
## What Makes It Bad
### Inline String Concatenation (i18n P3 Resources are External, Not Inline)
- `"Welcome, " + name + "!"` bakes English word order into code. In
Japanese the name comes first (`ようこそ、محمدさん!`); in Arabic the
structure differs again. The concatenation is invisible to the
extraction pipeline (`xgettext`, `i18next-parser`) — the translator
never sees it as a unit, and the string cannot be versioned or
rolled back as a whole.
- The custom `.csv` "resource" store is a second P3 violation: no
standard tool reads it, it carries no plural grammar, and it cannot
compose with the ICU formatting layer.
- **Fix:** strings live in a standard locale resource file, addressed
by key. Code calls `t("welcome", { name })`; the resource carries
the parameterized message.
```json
// en-US.json (ICU MessageFormat)
{
"checkout.welcome": "Welcome, {name}!",
"checkout.cart.summary": "{count, plural, one {# item} other {# items}} · {price} · {date}"
}
```
```json
// ar-EG.json — six plural categories per CLDR; the code is identical
{
"checkout.welcome": "أهلاً بك، {name}!",
"checkout.cart.summary": "{count, plural, zero {لا عناصر} one {عنصر واحد} two {عنصران} few {# عناصر} many {# عنصرًا} other {# عنصر}} · {price} · {date}"
}
```
- See `domains/i18n/locale-resources.md` (Resources are the Boundary)
and `domains/i18n/first-principles.md` P3.
### Hand-Rolled Plural Branching (i18n P4 Plural and Gender are Parameterized)
- `itemCount === 1 ? "1 item" : itemCount + " items"` encodes
English's one/other rule and nothing else. Arabic has six
categories (zero, one, two, few, many, other); Russian has three
(one, few, many); Polish has three with different boundaries. A
two-branch `if` is a C1 (Correctness) violation masquerading as a
shortcut — it returns a wrong answer for every non-English locale.
- **Fix:** the count goes to ICU MessageFormat; the formatter
consults `Intl.PluralRules` for the active locale; the resource
carries the variant for that category. The code passes the count,
nothing more.
```typescript
// The code passes the count; the resource + formatter pick the form.
t("checkout.cart.summary", { count: itemCount, price, date });
// Intl.PluralRules("ar-EG").select(1) === "one" -> "عنصر واحد"
// Intl.PluralRules("ar-EG").select(2) === "two" -> "عنصران"
// Intl.PluralRules("ar-EG").select(5) === "few" -> "٥ عناصر"
```
- See `domains/i18n/locale-resources.md` (Plural and Gender in
Resources) and `domains/i18n/formatting.md` (Plural Rules).
### Hand-Rolled Currency and Date Formatting (i18n P5 Formatting is Locale-Aware)
- `"$" + total.toFixed(2)` hardcodes the US dollar symbol, US
grouping (`,`), and US placement (symbol before the number). In
`de-DE` the euro formats as `"1.234,56 €"` (symbol after, dot
grouping). In `ar-EG` the pound formats as `"١٬٢٣٤٫٥٦ ج.م."`
(Arabic-Indic digits, different grouping).
- `(d.getMonth() + 1) + "/" + d.getDate() + "/" + d.getFullYear()`
produces `11/7/2024` — US `mm/dd/yyyy`. Most of the world reads
`dd/mm/yyyy`; ISO is `yyyy-mm-dd`. A hand-rolled date formatter
encodes one locale's convention and silently produces wrong output
for every other.
- **Fix:** `Intl.NumberFormat` and `Intl.DateTimeFormat` with a BCP
47 locale tag. CLDR is the source of truth; `Intl` is the runtime.
```typescript
new Intl.NumberFormat("ar-EG", { style: "currency", currency: "EGP" })
.format(1234.56); // "١٬٢٣٤٫٥٦ ج.م."
new Intl.DateTimeFormat("ar-EG", { dateStyle: "medium" })
.format(new Date(date)); // "٧ نوفمبر ٢٠٢٤"
```
- See `domains/i18n/formatting.md` (the Intl surface, dates, numbers,
currencies) and `domains/i18n/first-principles.md` P5.
### Source Language Treated as the Default (i18n P1 Source Language is a Locale)
- The component has no resource layer at all for the source locale —
English is "just the strings in the code." When the first second
locale arrives, the fix is a rewrite (extract every string,
restructure every concatenation), not a patch. The source language
is `en-US`, a locale among many — it is not `null`.
- **Fix:** extract source strings into `en-US.json` from day one,
even before a second locale exists. The resource layer is the
boundary from the first commit.
- See `domains/i18n/first-principles.md` P1 and
`domains/uiux/copywriting.md`.
## The Cascade
The violations compound. Inline concatenation makes strings invisible
to the extraction pipeline, so the translator never receives them as
units — they reconstruct them by reading the code. Hand-rolled
plurals return wrong answers for every non-English locale, so the
Arabic build ships `"1 items"` for a single item. Hand-rolled
formatting produces US-shaped output everywhere, so the price and
date are wrong for `de-DE`, `ar-EG`, `zh-Hans-CN`, and every other
locale. And because the first i18n test ran against real translations
(a P8 violation — pseudo-locales should have surfaced all of this
while the fix was still cheap), the defects are found after the
string freeze, after the translator was paid, and after the release
date was promised. The fix is now a re-translation and a re-release,
not a commit.
## Cross-Domain Links
- `domains/i18n/locale-resources.md` — the resource layer this code
lacks; the standard formats (`.po`, JSON, Fluent, ICU Resource
Bundle) it should have used.
- `domains/i18n/formatting.md` — the `Intl`/ICU/CLDR formatting this
code should call instead of hand-rolling.
- `domains/i18n/first-principles.md` — P3, P4, P5, and P8 (pseudo-
locales test early).
- `domains/uiux/copywriting.md` — copy lives in resources, not in
code.
- `domains/api/error-responses.md` — the same parameterized-message
discipline applies to localized API errors.
+226
View File
@@ -0,0 +1,226 @@
# Good Example: AI/ML Reproducible Training Run
> A training run that follows Atelier's AI/ML principles. Each aspect
> cites the principle it satisfies. Scope per D-023: this is
> engineering discipline (reproducibility, versioning, lineage,
> serving), **not** algorithm or model design — no architecture
> choice, hyperparameter tuning, or model-family comparison appears
> here.
## The Run
A training run `2026-08-05T09:12:00Z#run-42` produces model
`registry/payments-fraud@sha256:b5e1...aa0`. Every input that shaped
the model is pinned, named, and recoverable; the eval was declared
before training; the model is an addressed artifact in a registry;
the rollback path names the prior model and the prior dataset.
### The Reproducibility Contract
```yaml
# lineage/run-42.yaml — the lineage root, committed alongside the code
run_id: 2026-08-05T09:12:00Z#run-42
dataset: s3://ml-data/train@sha256:7f3a...e21
splits: dvc.yaml@commit a1b2c4d
code: git@a1b2c4d
config: configs/train.yaml@commit a1b2c4d
environment: ghcr.io/org/train-img@sha256:9c2d...f88
eval_spec: configs/eval.yaml@commit a1b2c4d
model_digest: registry/payments-fraud@sha256:b5e1...aa0
status: passed # eval gate passed -> eligible for promotion
```
- Lose any line and the run is anecdote, not evidence. The record is
the lineage root: a prediction cites the `model_digest`, which
cites this `run_id`, which cites everything above.
### Data is Versioned (DVC, content-hashed)
```ini
# dvc.yaml — the split config is versioned in git, the data in the
# content-addressed object store. Both are pinned by commit + hash.
stages:
prepare:
cmd: python src/prepare.py --input data/raw --out data/splits
deps:
- data/raw
- src/prepare.py
outs:
- data/splits/train.parquet
- data/splits/val.parquet
- data/splits/test.parquet
# The dataset hash (sha256:7f3a...e21) is recorded in the lineage
# contract above. "s3://ml-data/latest" would be a P2 violation.
```
```bash
# The dataset is pinned by content hash, not by a mutable path.
$ dvc get s3://ml-data/train --rev sha256:7f3a...e21
# The split is a deterministic function of (dataset version, split
# config, random seed). Two runs on the same pinned inputs produce
# the same splits.
```
### Code and Config are Versioned (git)
```yaml
# configs/train.yaml@commit a1b2c4d — versioned with the code
# (No algorithm/hyperparameter content is illustrated here — this is
# the engineering discipline of pinning the config, not the model
# design inside it. Per D-023, algorithm choice is out of scope.)
seed: 42
splits:
train: data/splits/train.parquet
val: data/splits/val.parquet
test: data/splits/test.parquet # held out, never touched by training
```
### Environment is Pinned (container digest)
```dockerfile
# The training environment is an image addressed by digest, not :latest.
# ghcr.io/org/train-img@sha256:9c2d...f88
FROM python:3.11-slim
# dependencies pinned in requirements.txt with hashes
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
```
```text
# requirements.txt — pinned + hash-pinned (pip-compile / pip-audit)
dvc==3.50.2 \
--hash=sha256:1c8a...e7
mlflow==2.16.0 \
--hash=sha256:9b2f...a1
# No unpinned ranges. A rerun pulls the exact same wheels.
```
### Evaluation is Defined Before Training (P4)
```yaml
# configs/eval.yaml@commit a1b2c4d — committed BEFORE training runs.
# The metrics, splits, and pass/fail thresholds are a-priori; they
# are the contract the model must satisfy to leave the experiment.
metrics:
- name: precision_at_threshold
threshold: ">= 0.92"
- name: recall_at_threshold
threshold: ">= 0.85"
- name: false_positive_rate
threshold: "<= 0.03"
split: data/splits/test.parquet # held out, never in training
gate: all_metrics_pass # AND of all thresholds; no cherry-pick
# The eval schema equals the serving input contract (serving.md P8):
# feature names, types, ranges match the production boundary exactly.
```
- Metrics chosen after seeing scores would be a P4 violation: the eval
would be rationalizing, not measuring. See
`domains/ai-ml/model-evaluation.md`.
### The Model is a Versioned Artifact (MLflow registry)
```bash
# After the eval gate passes, the model is registered as an immutable
# artifact addressed by digest, then promoted by stage.
$ mlflow models register \
--name payments-fraud \
--model-uri runs:/run-42/model \
--description "run-42, dataset sha256:7f3a...e21, eval passed"
# registry/payments-fraud@sha256:b5e1...aa0
# Stages: None -> Staging -> Production. Promotion is a registry
# operation, not a file copy. Never "latest".
```
### The Pipeline Composes (P9)
```text
# The training flow is a pipeline with explicit stages and contracts,
# not a notebook. Each stage has named inputs and named outputs.
prepare(dataset@hash) -> split(dvc.yaml) -> train(config, env@digest)
-> eval(eval.yaml, test@hash) -> [gate: pass] -> register(model@digest)
|
+-> [gate: fail] -> abort, no promote
# A notebook in this path would be a P9 violation: implicit state,
# human-dependent order, unreproducible.
```
## What Makes It Good
### Reproducibility is First Class (AI/ML P1, C1, C5)
- data + code + config + environment are all pinned. A second
engineer on a second laptop checks out commit `a1b2c4d`, pulls the
dataset by hash, pulls the image by digest, and reproduces the run
bit-for-bit. The run is reviewable because it is recreatable.
- See `domains/ai-ml/first-principles.md` P1 and
`domains/devops/first-principles.md` P1 Reproducibility.
### Data is Versioned, Not Just Code (AI/ML P2, C5, C7)
- The dataset is `s3://ml-data/train@sha256:7f3a...e21`, not
`s3://ml-data/latest`. A model trained on "the data" is a model
trained on an unknown input — a C1 violation. DVC pins the data the
way git pins the code.
- See `domains/ai-ml/data-versioning.md` (dataset hashing, the DVC /
Delta Lake / LakeFS comparison) and `domains/data/migrations.md`.
### Lineage is Traceable End-to-End (AI/ML P3, C7, C1)
- prediction → model → run-42 → dataset → source. Every edge is
named; no orphan model. A serving regression traces back to the
exact dataset and code that built the model, which is how drift is
diagnosed (data drift vs concept drift vs prediction drift).
- See `domains/ai-ml/data-versioning.md` (lineage record) and
`domains/observability/logging.md`.
### Evaluation Defined Before Training (AI/ML P4, C1, C2)
- `eval.yaml` was committed before `train` ran. The gate is
`all_metrics_pass`; a failing metric aborts promotion. Cherry-
picking a metric post-hoc is a correctness violation — the eval
would no longer measure the model.
- See `domains/ai-ml/model-evaluation.md` (eval-as-a-gate) and
`domains/testing/first-principles.md` (tests as specification).
### Models are Versioned Artifacts (AI/ML P5, C5, C6)
- The model is `registry/payments-fraud@sha256:b5e1...aa0`, promoted
Staging → Production. A serving endpoint that pulled `latest` would
be serving an unknown model with no rollback. The registry is to
models what a container registry is to images.
- See `domains/ai-ml/serving.md` (the model is an addressed artifact)
and `domains/devops/first-principles.md` P7 Immutability.
### Rollback Includes the Model (AI/ML P10, C5)
- If production regresses, the rollback restores the prior model
digest `registry/payments-fraud@sha256:a1c4...f09` AND the prior
serving code. A rollback that redeploys old code but keeps the new
model has not rolled back — the model was the thing that regressed.
- See `domains/ai-ml/serving.md` (Rollback Includes the Model) and
`domains/devops/first-principles.md` P4 Rollback First.
## What This Example Does NOT Do (And Why That's Good)
- Does **not** reference the dataset by a mutable path —
`s3://ml-data/latest` would be a P2 violation.
- Does **not** choose metrics after seeing scores — that is a P4
violation (rationalizing, not measuring).
- Does **not** pull `latest` from the model registry — that is a P5
violation (unknown model, no rollback).
- Does **not** contain algorithm/architecture/hyperparameter content
— per D-023, those are research choices, not engineering
principles, and have no derivation in the core C-rules.
- Does **not** run from a notebook — a notebook in the pipeline path
is a P9 violation (implicit state, unreproducible).
## Cross-Domain Links
- `domains/ai-ml/data-versioning.md` — the DVC pinning, the lineage
record, the tool comparison (DVC / Delta Lake / LakeFS).
- `domains/ai-ml/serving.md` — the model is promoted as an addressed
artifact; the serving boundary validates inputs against the same
schema as the eval.
- `domains/ai-ml/model-evaluation.md` — the eval-as-a-gate that this
run must pass before promotion.
- `domains/devops/first-principles.md` P1 Reproducibility — the
non-negotiable this run inherits.
- `domains/data/migrations.md` — data versioning parallels schema
migration discipline.
- `domains/observability/logging.md` — the lineage record is a
structured, append-only log of provenance.
+197
View File
@@ -0,0 +1,197 @@
# Good Example: GitOps Pull Request
> A pull request that changes ArgoCD Application manifests following
> Atelier's GitOps + Operators principles. Each aspect cites the
> principle it satisfies.
## The PR
A PR titled `promote payments-api 1.2.3 to prod` opened against the
GitOps repo `platform/gitops`. It changes the `targetRevision` of the
payments-api Application from `1.2.2` to `1.2.3`, adds a sync-wave
annotation to a new migration Job, and tightens the AppProject's
destination allow-list. CI runs plan/diff; nothing pushes to the
cluster.
### The Commit
```yaml
# manifests/prod/payments-api.yaml — the only file changed
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: payments-api
namespace: argocd
finalizers:
- resources-finalizer.argocd.argoproj.io
spec:
source:
repoURL: https://git.example.com/platform/payments
targetRevision: 1.2.3 # was 1.2.2 — pinned, not latest
path: manifests/prod
destination:
server: https://kubernetes.default.svc
namespace: payments
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=false
- PrunePropagationPolicy=foreground
```
```yaml
# manifests/prod/payments-db-migration.yaml — new file, wave-ordered
apiVersion: batch/v1
kind: Job
metadata:
name: payments-db-migrate-1.2.3
namespace: payments
annotations:
argocd.argoproj.io/sync-wave: "-1" # PreSync: run before the app
spec:
backoffLimit: 0
ttlSecondsAfterFinished: 86400
template:
spec:
restartPolicy: OnFailure
serviceAccountName: payments-migrator
containers:
- name: migrate
image: registry.example.com/payments/migrate:1.2.3
args: ["up", "--to", "1.2.3"]
```
```yaml
# projects/payments.yaml — AppProject tightened in the same PR
apiVersion: argoproj.io/v1alpha1
kind: AppProject
metadata:
name: payments
namespace: argocd
spec:
sourceRepos:
- https://git.example.com/platform/payments
destinations:
- server: https://kubernetes.default.svc
namespace: payments # prod-app can only deploy to payments ns
clusterResourceWhitelist:
- group: ""
kind: Namespace # allowed to create its own namespace
roles:
- name: payments-team
policies:
- p, proj:payments:payments-team, applications, sync, payments/*, allow
```
### The CI Pipeline (runs on the PR, before merge)
```text
# .github/workflows/gitops-plan.yml (illustrative steps)
- name: validate manifests
run: argocd app manifests manifests/prod/ | kubeconform -strict
- name: diff against live cluster (read-only, no apply)
run: argocd app diff payments-api --server $ARGOCD_SERVER --auth-token $READ_ONLY_TOKEN
# CI holds a READ-ONLY ArgoCD token. It never holds kubectl rights.
# A non-empty diff is the PR's proposed change, rendered for review.
- name: opa gate (admission policy pre-check)
run: opa eval -i manifests/prod/ -d policies/ "data.k8s.admission.deny"
# Policy violations fail the PR before merge, not after deploy.
```
## What Makes It Good
### Git is the Source of Truth (GitOps P1, C1 Correctness)
- The promotion is a commit. The cluster's desired state is a
derivative of this repo; the repo is the authority. If the change is
wrong, `git revert` is the rollback — the recovery path is the
history.
- See `domains/gitops-operators/first-principles.md` P1 and
`domains/gitops-operators/argocd.md` (Application CRD).
### Pull, Don't Push (GitOps P3, C4 Locality)
- CI holds a **read-only** ArgoCD token for `app diff`. It holds no
`kubectl` rights against the production cluster. The cluster's
ArgoCD controller pulls the merged commit; nothing pushes to the
cluster. A compromised CI token can read, not deploy.
- See `domains/gitops-operators/argocd.md` (RBAC and SSO) and
`domains/gitops-operators/flux.md` for the same pull boundary from
the Flux side.
### State is Immutable and Versioned (GitOps P5, C5 Reversibility)
- `targetRevision: 1.2.3` — the Application pins a specific chart
revision, not `latest`. The commit that changed it is a permanent
record; `git revert` restores `1.2.2` and ArgoCD's `selfHeal`
converges the cluster back. No force-push; history is the audit
trail.
- See `domains/gitops-operators/first-principles.md` P5 and
`domains/infrastructure-as-code/state.md` (State is Truth).
### Sync Waves Order Correctness (GitOps P4, C1)
- The migration Job carries `argocd.argoproj.io/sync-wave: "-1"` so
it runs in `PreSync` before the payments-api Deployment that
depends on the new schema. Wave ordering is a correctness
mechanism, not performance — the app starting before its migration
is a correctness bug.
- See `domains/gitops-operators/argocd.md` (Sync Waves and Hooks).
### Reconcile, Don't Mutate by Hand (GitOps P8)
- `selfHeal: true` + `prune: true` means a hand-edited drift on a
managed resource is overwritten on the next loop. The fix for drift
is a new commit, not `kubectl edit`. The PR author does not SSH into
the cluster to "fix" anything.
- See `domains/gitops-operators/argocd.md` (Diff and Drift) and
`domains/gitops-operators/first-principles.md` P8.
### Least Privilege Reconciliation (GitOps P10, C8 Economy)
- The AppProject `payments` restricts the Application to the
`payments` namespace and the `payments` repo. The controller's
ServiceAccount (not shown) is bound to a namespace-scoped Role, not
`cluster-admin`. The PR *tightens* the allow-list — least privilege
is a direction, not a one-time setting.
- See `domains/gitops-operators/argocd.md` (RBAC and SSO) and
`domains/kubernetes/rbac.md`.
### Policy is a Gate (Compliance P5, cross-link)
- The `opa eval` step runs the admission policy against the proposed
manifests before merge. A violation fails the PR; the non-compliant
state is never realized. Detection is not enforcement; this is
enforcement.
- See `domains/compliance/policy-as-code.md` and
`domains/devops/ci-cd.md`.
### Failure is Observable (GitOps P9)
- A sync failure or health degradation on `payments-api` emits
ArgoCD status (`Degraded` / `OutOfSync`) and a notification. Silent
drift is the bug; this PR does not disable notifications.
- See `domains/gitops-operators/argocd.md` (Health and Status) and
`domains/observability/metrics.md`.
## What This PR Does NOT Do (And Why That's Good)
- Does **not** run `kubectl apply` from CI — that is the push pattern,
a P3 violation (see `examples/bad/` for the anti-pattern).
- Does **not** use `argocd app set` as the steady state — the change
is in git, not in an imperative command's history.
- Does **not** store raw Secrets in the GitOps repo — secrets arrive
via Sealed Secrets / SOPS / External Secrets, encrypted in git.
- Does **not** float `targetRevision: latest` — the Application pins
a version; "latest" is an unknown model of the system.
## Cross-Domain Links
- `domains/gitops-operators/argocd.md` — the Application CRD, sync
waves, RBAC/AppProjects, and the pull model.
- `domains/gitops-operators/flux.md` — the same PR pattern from the
Flux side (Kustomization CRD, per-cluster autonomy).
- `domains/kubernetes/workloads.md` — the Deployment/Job the
Application reconciles.
- `domains/kubernetes/rbac.md` — the ServiceAccount + Role the
controller and the migration Job run as.
- `domains/compliance/policy-as-code.md` — the OPA gate is a
compliance-as-a-gate enforcement point.
- `domains/devops/P4 Rollback First``git revert` is the rollback;
`selfHeal` is the convergence.
+24 -9
View File
@@ -6,14 +6,14 @@
| Core Principle | Domains that derive from it | Count | | Core Principle | Domains that derive from it | Count |
|----------------|---------------------------|-------| |----------------|---------------------------|-------|
| C1 Correctness | All 13 (v0.1: 11; v0.2: infrastructure-as-code, kubernetes) | Universal | | C1 Correctness | All 17 (v0.1: 11; v0.2: infrastructure-as-code, kubernetes; v0.3: gitops-operators, ai-ml, i18n, compliance) | Universal |
| C2 Clarity | v0.1: uiux, api, data, testing, observability, errors, documentation, devops; v0.2: infrastructure-as-code, kubernetes | 10 | | C2 Clarity | v0.1: uiux, api, data, testing, observability, errors, documentation, devops; v0.2: infrastructure-as-code, kubernetes; v0.3: gitops-operators, ai-ml, i18n, compliance | 14 |
| C3 Simplicity | v0.1: security, data, testing, performance, documentation, concurrency, devops | 7 | | C3 Simplicity | v0.1: security, data, testing, performance, documentation, concurrency, devops; v0.2: infrastructure-as-code; v0.3: gitops-operators, i18n, compliance | 11 |
| C4 Locality | v0.1: testing, concurrency; v0.2: infrastructure-as-code, kubernetes | 4 | | C4 Locality | v0.1: testing, concurrency; v0.2: infrastructure-as-code, kubernetes; v0.3: gitops-operators, i18n | 6 |
| C5 Reversibility | v0.1: api, data, uiux, concurrency, devops; v0.2: infrastructure-as-code, kubernetes | 7 | | C5 Reversibility | v0.1: api, data, uiux, concurrency, devops; v0.2: infrastructure-as-code, kubernetes; v0.3: gitops-operators, ai-ml, i18n, compliance | 11 |
| C6 Composability | v0.1: api, security, observability, errors, documentation, concurrency; v0.2: infrastructure-as-code, kubernetes | 8 | | C6 Composability | v0.1: api, security, observability, errors, documentation, concurrency; v0.2: infrastructure-as-code, kubernetes; v0.3: gitops-operators, ai-ml, i18n, compliance | 12 |
| C7 Observability | v0.1: api, data, testing, performance, observability, errors, devops; v0.2: infrastructure-as-code, kubernetes | 9 | | C7 Observability | v0.1: api, data, testing, performance, observability, errors, devops; v0.2: infrastructure-as-code, kubernetes; v0.3: gitops-operators, ai-ml, i18n, compliance | 13 |
| C8 Economy | v0.1: security, testing, performance, observability, concurrency; v0.2: kubernetes | 6 | | C8 Economy | v0.1: security, testing, performance, observability, concurrency; v0.2: kubernetes; v0.3: gitops-operators, i18n, compliance | 9 |
## Interpretation ## Interpretation
@@ -39,6 +39,10 @@
| DevOps | C1, C2, C3, C5, C7 | Reproducibility + rollback | | DevOps | C1, C2, C3, C5, C7 | Reproducibility + rollback |
| Infrastructure as Code | C1, C2, C3, C4, C5, C6, C7 | Declarative + state + composition; broadest derivation alongside Concurrency | | Infrastructure as Code | C1, C2, C3, C4, C5, C6, C7 | Declarative + state + composition; broadest derivation alongside Concurrency |
| Kubernetes | C1, C2, C4, C5, C6, C7, C8 | Declarative + reversibility + economy; broad derivation (7 C-rules) | | Kubernetes | C1, C2, C4, C5, C6, C7, C8 | Declarative + reversibility + economy; broad derivation (7 C-rules) |
| GitOps + Operators | C1, C2, C3, C4, C5, C6, C7, C8 | Source-of-truth + reconciliation + pull-locality + least privilege; broadest derivation (8 C-rules, tied with i18n) |
| AI / ML | C1, C2, C5, C6, C7 | Reproducibility + lineage + serving observability |
| i18n | C1, C2, C3, C4, C5, C6, C7, C8 | Locale + formatting + direction + reversibility; broadest derivation (8 C-rules) |
| Compliance | C1, C2, C3, C5, C6, C7, C8 | Audit + policy-as-code + retention + posture |
## v0.2 Domain Coverage (per IDEATE-03 schema) ## v0.2 Domain Coverage (per IDEATE-03 schema)
@@ -47,10 +51,21 @@
| Infrastructure as Code | 10 | 4 (terraform, opentofu, state, modules) | ✓ | complete | | Infrastructure as Code | 10 | 4 (terraform, opentofu, state, modules) | ✓ | complete |
| Kubernetes | 10 | 6 (workloads, networking, storage, rbac, helm, kustomize) | ✓ | complete | | Kubernetes | 10 | 6 (workloads, networking, storage, rbac, helm, kustomize) | ✓ | complete |
## v0.3 Domain Coverage (per IDEATE-03 schema)
| Domain | P-count | Derived-doc-count | Manifest-listed | Status |
|--------|---------|-------------------|-----------------|--------|
| GitOps + Operators | 10 | 4 (argocd, flux, operators, progressive-delivery) | ✓ | complete |
| AI / ML | 10 | 4 (data-versioning, model-evaluation, serving, monitoring-drift) | ✓ | complete |
| i18n | 10 | 4 (locale-resources, formatting, rtl-bidi, testing-i18n) | ✓ | complete |
| Compliance | 10 | 4 (audit-logs, data-retention, policy-as-code, evidence) | ✓ | complete |
## Gaps and Notes ## Gaps and Notes
- No domain derives from only one C-rule. The minimum is 4 (UI/UX: C1, C2, C3, C5, C7 — actually 5). Every domain is multi-rooted. - No domain derives from only one C-rule. The minimum is 4 (UI/UX: C1, C2, C3, C5, C7 — actually 5). Every domain is multi-rooted.
- **Concurrency**, **Infrastructure as Code**, and **Kubernetes** are tied for the broadest derivation (7 C-rules each) — these domains touch the most core concerns. - **Concurrency**, **Infrastructure as Code**, and **Kubernetes** are tied for the broadest derivation among v0.1/v0.2 platform domains (7 C-rules each) — these domains touch the most core concerns.
- **GitOps + Operators** and **i18n** are tied for the single broadest-derivation domain overall (8 C-rules each: C1C8). GitOps adds C4 (pull-credential locality) alongside its source-of-truth/reconciliation/least-privilege derivation; i18n touches correctness, clarity, simplicity, locality, reversibility, composability, observability, and economy (expansion accommodation). This is consistent with both domains' cross-cutting nature.
- **UI/UX** and **API** are the most user-facing; they emphasize C2 (Clarity) heavily. - **UI/UX** and **API** are the most user-facing; they emphasize C2 (Clarity) heavily.
- **Security** is the only domain with explicit non-tradeable declarations; this promotes 8 of its rules to C1-equivalent per `core/conflict-resolution.md` §6. - **Security** is the only domain with explicit non-tradeable declarations; this promotes 8 of its rules to C1-equivalent per `core/conflict-resolution.md` §6.
- **v0.2 expansion:** C4 (Locality) grew from 2 to 4 domains (added infrastructure-as-code state locality, kubernetes namespace blast-radius). C6 (Composability) grew from 6 to 8. The two new domains are broad-derivation domains (7 C-rules each), consistent with Concurrency's breadth. - **v0.2 expansion:** C4 (Locality) grew from 2 to 4 domains (added infrastructure-as-code state locality, kubernetes namespace blast-radius). C6 (Composability) grew from 6 to 8. The two new domains are broad-derivation domains (7 C-rules each), consistent with Concurrency's breadth.
- **v0.3 expansion:** C3 (Simplicity) grew from 7 to 11 (added gitops-operators declarative simplicity, i18n flexible layout, compliance structural redaction). C4 (Locality) grew from 4 to 6 (added gitops-operators pull-credential locality, i18n resource/text-direction locality). C5 (Reversibility) grew from 7 to 11 (added all four v0.3 domains — gitops history, ai-ml reproducibility, i18n translation versioning, compliance append-only/retention). C6 (Composability) grew from 8 to 12. C7 (Observability) grew from 9 to 13 (added all four v0.3 domains — reconciliation, drift detection, format correctness, posture). C2 (Clarity) grew from 10 to 14. The v0.3 expansion broadens every non-universal C-rule's coverage, confirming the four new domains are cross-cutting and well-rooted.
+63 -3
View File
@@ -205,8 +205,68 @@ C5=Reversibility · C6=Composability · C7=Observability · C8=Economy
| P9 Config and Secrets Sep | C2 | Clarity of configuration | | P9 Config and Secrets Sep | C2 | Clarity of configuration |
| P10 Roll Forward, Roll Back | C5 | Reversibility of deploys | | P10 Roll Forward, Roll Back | C5 | Reversibility of deploys |
## Coverage Summary (post-v0.2) ## GitOps + Operators
- 13 domains (11 v0.1 + 2 v0.2: infrastructure-as-code, kubernetes) | GitOps Principle | Core | Why |
- 130 domain principles total (110 v0.1 + 20 v0.2) |---------------------------|------|---------------------------------------|
| P1 Git is the Source of Truth | C1, C5 | Correctness of source; reversibility via history |
| P2 Declarative Over Imperative | C2, C3 | Clarity of intent; simplicity of expression |
| P3 Pull, Don't Push | C1, C4 | Correctness via security; locality of credentials |
| P4 Continuous Reconciliation | C7, C1 | Observability of drift; correctness of converge loop |
| P5 State is Immutable and Versioned | C5 | Reversibility through history |
| P6 Operators Encode Domain Knowledge | C6, C2 | Composability of expertise; clarity of operations |
| P7 Progressive Delivery is Reversible | C5, C1 | Reversibility of promotion; correctness of abort |
| P8 Reconcile, Don't Mutate by Hand | C1, C7 | Correctness of source of truth; observability of drift |
| P9 Failure is Observable and Surfaced | C7 | Observability of sync/rollout health |
| P10 Least Privilege Reconciliation | C1, C8 | Correctness via security; economy of trust |
## AI / ML
| AI/ML Principle | Core | Why |
|---------------------------|------|---------------------------------------|
| P1 Reproducibility is the First Class | C1, C5 | Correctness of runs; reversibility of reproduction |
| P2 Data is Versioned, Not Just Code | C5, C7 | Reversibility of data; observability of dataset lineage |
| P3 Lineage is Traceable End-to-End | C7, C1 | Observability of predictions; correctness of provenance |
| P4 Evaluation is Defined Before Training | C1, C2 | Correctness of metrics; clarity of thresholds |
| P5 Models are Versioned Artifacts | C5, C6 | Reversibility of model rollbacks; composability of registry |
| P6 Serving is Observable | C7 | Observability of inference |
| P7 Drift is Expected and Detected | C7, C1 | Observability of drift; correctness of detection |
| P8 Inference Inputs are Validated | C1 | Correctness at serving boundary |
| P9 Pipelines Compose, Notebooks Don't | C6, C2 | Composability of steps; clarity of contracts |
| P10 Rollback Includes the Model | C5 | Reversibility at the model layer |
## i18n
| i18n Principle | Core | Why |
|---------------------------|------|---------------------------------------|
| P1 Source Language is a Locale, Not the Default | C2, C1 | Clarity; correctness of localization model |
| P2 Locale Identifiers are Standardized | C2, C6 | Clarity; composability of BCP 47 |
| P3 Resources are External, Not Inline | C4, C6 | Locality of strings; composability of resources |
| P4 Plural and Gender are Parameterized | C1, C6 | Correctness across locales; composability of message format |
| P5 Formatting is Locale-Aware | C1, C7 | Correctness of formats; observability of format correctness |
| P6 Text Direction is a Layout Primitive | C1, C4 | Correctness of RTL/bidi; locality of direction |
| P7 Layout Accommodates Expansion | C8, C3 | Economy of rework; simplicity of flexible layout |
| P8 Pseudo-Locales Test Early | C7, C5 | Observability of bugs early; reversibility of finding late |
| P9 Images and Icons are Cultural | C1, C2 | Correctness; clarity of cultural meaning |
| P10 Translation is Reversible and Versioned | C5 | Reversibility of localization changes |
## Compliance
| Compliance Principle | Core | Why |
|---------------------------|------|---------------------------------------|
| P1 Audit Logs are Append-Only | C1, C5 | Correctness of audit; reversibility of immutable record |
| P2 Every Significant Action is Logged | C7, C1 | Observability of actions; correctness of audit set |
| P3 Retention is Policy, Not Storage | C5, C8 | Reversibility of lifecycle; economy of storage |
| P4 Policy is Code | C6, C2 | Composability of policy; clarity of rules |
| P5 Policy is Evaluated as a Gate | C1, C5 | Correctness of enforcement; reversibility of block |
| P6 Evidence is Collected Continuously | C7, C3 | Observability of posture; simplicity of audit |
| P7 Identity is Attributable | C1, C7 | Correctness of attribution; observability of subject |
| P8 Subject Access is Honored | C1, C5 | Correctness of rights; reversibility of deletion/export |
| P9 Secrets and Sensitive Data are Redacted in Audit | C1, C3 | Correctness via security; simplicity of structural redaction |
| P10 Compliance Posture is Observable | C7, C1 | Observability of compliance; correctness of posture |
## Coverage Summary (post-v0.3)
- 17 domains (11 v0.1 + 2 v0.2: infrastructure-as-code, kubernetes; 4 v0.3: gitops-operators, ai-ml, i18n, compliance)
- 170 domain principles total (110 v0.1 + 20 v0.2 + 40 v0.3)
- Every domain P-rule traces to ≥1 core C-rule (C1C8). No orphans. - Every domain P-rule traces to ≥1 core C-rule (C1C8). No orphans.
+50
View File
@@ -147,6 +147,56 @@ If the task touches a domain, run that domain's checklist:
- [ ] Rollout history retained; rollback tested (P10) - [ ] Rollout history retained; rollback tested (P10)
- [ ] Namespaces used to bound blast radius; not `default` in prod (P6) - [ ] Namespaces used to bound blast radius; not `default` in prod (P6)
### If GitOps + Operators (see `domains/gitops-operators/`)
- [ ] Desired state lives in git, not in the cluster (P1)
- [ ] Configuration is declarative, not imperative scripts (P2)
- [ ] Reconciliation is pull-based; no external push credentials into the cluster (P3)
- [ ] Reconciliation loop runs continuously; drift auto-corrected (P4)
- [ ] Every change is a commit; history is the audit/rollback path (P5)
- [ ] Operational knowledge encoded as CRDs/controllers, not runbooks (P6)
- [ ] Progressive delivery (canary/blue-green) has a tested abort/rollback path (P7)
- [ ] No manual `kubectl apply`/`kubectl edit` on GitOps-managed resources (P8)
- [ ] Sync failures, health degradation, and rollout stalls emit status + notifications (P9)
- [ ] Controller credentials scoped to reconciled namespaces/resources; no cluster-admin GitOps robot (P10)
### If AI / ML (see `domains/ai-ml/`)
- [ ] Scope check: this is engineering discipline (data versioning, evaluation, serving, drift), NOT algorithm/model design (D-023) — reject algorithm-design content
- [ ] Every training run is reproducible from pinned data + code + config + environment (P1)
- [ ] Datasets, features, and splits are versioned artifacts with lineage; `git` alone is insufficient (P2)
- [ ] Any deployed prediction traces back through model → training run → dataset → source (P3)
- [ ] Metrics, splits, and thresholds declared a priori; no post-hoc metric cherry-picking (P4)
- [ ] Models are pinned, immutable, registry-tracked artifacts; never "the latest" (P5)
- [ ] Inference latency, throughput, input distributions, and prediction confidence are observed (P6)
- [ ] Data drift, concept drift, and prediction drift are monitored; a drift signal is an incident (P7)
- [ ] Inference inputs validated against the model's contract (schema, ranges, types); out-of-contract rejected (P8)
- [ ] Training/serving flows are composable pipelines with explicit steps; notebooks not in production (P9)
- [ ] Serving rollback restores the prior model artifact, not just the prior code (P10)
### If i18n (see `domains/i18n/`)
- [ ] Source language treated as one locale among many, not the "neutral" default (P1)
- [ ] Locale identifiers use BCP 47 tags; no ad-hoc locale codes (P2)
- [ ] User-facing strings in locale resource files, not concatenated inline in code (P3)
- [ ] Plural/gender/select use ICU MessageFormat (or equivalent); no `if (n == 1)` branching (P4)
- [ ] Dates, times, numbers, currencies, units via ICU/CLDR/`Intl`; no hand-rolled formatters (P5)
- [ ] RTL/bidi is a first-class layout concern; logical CSS properties (`start`/`end`) over physical (`left`/`right`) (P6)
- [ ] Layouts accommodate translation expansion; no fixed pixel widths for text (P7)
- [ ] Pseudo-locales (accented, lengthened, RTL-mirrored) used to test before real translations arrive (P8)
- [ ] Icons, colors, and imagery reviewed for locale-sensitivity; no locale-bound symbols treated as universal (P9)
- [ ] Resource files versioned; a bad translation is a rollback, not a hot-patch (P10)
### If Compliance (see `domains/compliance/`)
- [ ] Scope check: framework-agnostic — no regulation-specific (GDPR/HIPAA/SOC2/PCI) content (D-024)
- [ ] Audit records are immutable once written; deletion/mutation is itself an auditable incident (P1)
- [ ] The set of auditable actions is defined a priori; "we forgot to log it" is a violation (P2)
- [ ] Data lifetime is declared and enforced as policy; deletion at end-of-life is a feature (P3)
- [ ] Compliance policy expressed in versioned, reviewable, testable code (OPA/Cedar/Kyverno/Sentinel), not spreadsheets/prose (P4)
- [ ] Policy violations block before the action (admission/CI/CD-time), not after the audit (P5)
- [ ] Evidence gathered as a byproduct of operation, not assembled manually at audit time (P6)
- [ ] Every logged action traces to an authenticated principal; no shared/generic identities (P7)
- [ ] Data-subject rights (access, export, deletion) are operations with defined contracts and audit trails (P8)
- [ ] Audit logs do not leak secrets; redaction is structural, not opportunistic (P9)
- [ ] System reports its own compliance state (drift from policy, open violations, retention status) (P10)
## Final Gate ## Final Gate
- [ ] Have I read the relevant domain's first-principles? - [ ] Have I read the relevant domain's first-principles?
+61
View File
@@ -164,3 +164,64 @@ When you see a pattern listed here, it is a defect. Cite the principle it violat
| Orphaned P-rule (a domain principle with no matrix row) | matrix completeness, C6 | Breaks the conflict-resolution arbiter; the rule has no core trace | | Orphaned P-rule (a domain principle with no matrix row) | matrix completeness, C6 | Breaks the conflict-resolution arbiter; the rule has no core trace |
| Deployable example artifact (standalone `.tf`/`.yaml` under `examples/`) | PROJECT.md "no runtime code", D-025 | Violates the docs-only contract; examples must be `.md` with fenced code | | Deployable example artifact (standalone `.tf`/`.yaml` under `examples/`) | PROJECT.md "no runtime code", D-025 | Violates the docs-only contract; examples must be `.md` with fenced code |
| Unlisted v0.2 doc (new doc not added to MANIFEST) | manifest rule | Not part of the framework by definition | | Unlisted v0.2 doc (new doc not added to MANIFEST) | manifest rule | Not part of the framework by definition |
## v0.3 Chaos Anti-Patterns (from IDEATE-20, IDEATE-24, IDEATE-25, IDEATE-27)
These are named, cross-cutting violations specific to the v0.3 domains. Reject on sight.
| Anti-Pattern | Breaches | Why |
|--------------|----------|-----|
| GitOps push-pattern (external CI pushes manifests to the cluster instead of an in-cluster agent pulling from git) | gitops P3 Pull, Don't Push; C1, C4 | Inverts the source-of-truth flow; requires push credentials into the cluster; breaks the reconciliation model (IDEATE-24, D-042) |
| i18n LTR-only assumption (layout assumes left-to-right; no `dir` attribute, physical CSS properties only) | i18n P6 Text Direction is a Layout Primitive; C1, C4 | Disqualifying for RTL/Bidi users; locale-correctness violation (IDEATE-25, D-043) |
| AI/ML orphan-model (a deployed prediction endpoint whose model has no lineage trace — no record of training run, dataset, or version) | ai-ml P3 Lineage is Traceable End-to-End; C7, C1 | Unreviewable, unrollbackable; the model is an unattributed artifact (IDEATE-27, D-045) |
| Compliance mutable audit log (audit records can be edited or deleted by an operator) | compliance P1 Audit Logs are Append-Only; C1, C5 | Destroys the audit trail; the audit log's value is immutability — mutation is itself an incident |
### v0.3 Deployable Artifact Types (IDEATE-20, D-020)
The following standalone file types are forbidden under `examples/` and elsewhere in the framework. Examples are `.md` files with fenced code only.
| Forbidden standalone artifact | Belongs in | Why |
|-------------------------------|-----------|-----|
| `.po` / `.pot` resource files | fenced code in `examples/good/`/`examples/bad/*.md` | Runtime localization artifact; violates docs-only contract |
| `.rego` / `.cedar` / `.sentinel` policy files | fenced code in `examples/*.md` | Runtime policy artifact; violates docs-only contract |
| Model artifacts (`.pkl`, `.onnx`, `.pt`, `.h5`, `.safetensors`) | fenced code + prose in `examples/*.md` | Runtime model artifact; violates docs-only contract |
| Signed manifests as standalone files (`.sig`, `.att`, `.intoto.jsonl`) | fenced code in `examples/*.md` | Runtime attestation artifact; violates docs-only contract |
| Standalone `.yaml` / `.tf` / `.sh` | fenced code in `examples/*.md` | (Carried forward from v0.2) Runtime deployable artifact |
## v0.3 Domain-Specific Anti-Patterns
### GitOps + Operators
| Anti-Pattern | Breaches | Why |
|--------------|----------|-----|
| Push-based deploy (external CI `kubectl apply` into the cluster) | P3 Pull, Don't Push | Inverts the model; requires push credentials; bypasses reconciliation |
| Manual `kubectl apply`/`kubectl edit` on a GitOps-managed resource | P8 Reconcile, Don't Mutate by Hand | Unreconciled drift; the next loop overwrites it — silent and unattributed |
| `cluster-admin` GitOps robot (controller bound to cluster-admin) | P10 Least Privilege Reconciliation | Overbroad grant; blast radius = entire cluster |
| No sync-failure notification (silent drift on health degradation) | P9 Failure is Observable and Surfaced | Silent drift is the bug the loop was supposed to surface |
### AI / ML
| Anti-Pattern | Breaches | Why |
|--------------|----------|-----|
| Unreproducible training run (unpinned data, code, config, or environment) | P1 Reproducibility is the First Class | Unreviewable; cannot debug, cannot rollback |
| "Use the latest model" (serving points at `model:latest` instead of a pinned version) | P5 Models are Versioned Artifacts | Unversioned drift; rollback undefined |
| Notebook in production (training/serving flow is a Jupyter notebook) | P9 Pipelines Compose, Notebooks Don't | No contracts, no composition, no reproducibility |
| Orphan model (deployed prediction with no lineage trace) | P3 Lineage is Traceable End-to-End | Unattributed artifact; cannot trace to data/code (IDEATE-27) |
### i18n
| Anti-Pattern | Breaches | Why |
|--------------|----------|-----|
| Inline string concatenation (`"Hello, " + name + "!"` in code) | P3 Resources are External, Not Inline | Not extractable; breaks translations; word-order differs per locale |
| `if (n == 1)` plural branching (hand-rolled plural logic) | P4 Plural and Gender are Parameterized | Wrong for Arabic, Russian, Polish; ICU MessageFormat handles plurals |
| LTR-only layout (no `dir` attribute, physical CSS `left`/`right`) | P6 Text Direction is a Layout Primitive | Disqualifying for RTL/Bidi (IDEATE-25) |
| Hand-rolled date/number formatter (`new Date().toString()`, manual string formatting) | P5 Formatting is Locale-Aware | Locale-incorrect; ignores ICU/CLDR |
### Compliance
| Anti-Pattern | Breaches | Why |
|--------------|----------|-----|
| Mutable audit log (operator can `UPDATE`/`DELETE` audit records) | P1 Audit Logs are Append-Only | Destroys the audit trail; mutation is itself an incident |
| Shared/generic identity in audit (`admin` or `system` as the actor for all actions) | P7 Identity is Attributable | No attribution; no accountability; cannot investigate |
| Secret leaked in audit log (request body or token captured in an audit event) | P9 Secrets and Sensitive Data are Redacted in Audit | Audit log becomes a secret exfiltration channel |
| Manual evidence assembly at audit time (scramble to collect logs/scans/attestations on demand) | P6 Evidence is Collected Continuously | Audit-unready; evidence gathered under pressure is incomplete and unreliable |
+50
View File
@@ -79,6 +79,56 @@ Run the relevant domain section from `agent-checklist.md` (UI/UX, API, Security,
- [ ] Are ConfigMaps and Secrets separate? - [ ] Are ConfigMaps and Secrets separate?
- [ ] Is the rollback path tested, not assumed? - [ ] Is the rollback path tested, not assumed?
### If GitOps + Operators
- [ ] Is desired state sourced from git, not from the cluster?
- [ ] Is configuration declarative, not imperative scripts?
- [ ] Is reconciliation pull-based (no external push credentials into the cluster)?
- [ ] Does the reconciliation loop run continuously and auto-correct drift?
- [ ] Is every change a commit, with history as the audit/rollback path?
- [ ] Is operational knowledge encoded as CRDs/controllers, not runbooks humans must remember?
- [ ] Does progressive delivery (canary/blue-green) have a tested abort/rollback path?
- [ ] Are there manual `kubectl apply`/`kubectl edit` on GitOps-managed resources? (flag as incident)
- [ ] Do sync failures, health degradation, and rollout stalls emit status + notifications?
- [ ] Are controller credentials scoped to reconciled namespaces/resources (no cluster-admin GitOps robot)?
### If AI / ML
- [ ] Scope check: is this engineering discipline (data versioning, evaluation, serving, drift), NOT algorithm/model design? (D-023 — reject algorithm-design content)
- [ ] Is every training run reproducible from pinned data + code + config + environment?
- [ ] Are datasets, features, and splits versioned artifacts with lineage (not just `git`)?
- [ ] Can any deployed prediction trace back through model → training run → dataset → source?
- [ ] Are metrics, splits, and thresholds declared a priori (no post-hoc metric cherry-picking)?
- [ ] Are models pinned, immutable, registry-tracked artifacts (never "the latest")?
- [ ] Is inference observable (latency, throughput, input distributions, prediction confidence)?
- [ ] Are data drift, concept drift, and prediction drift monitored (drift signal = incident)?
- [ ] Are inference inputs validated against the model's contract (schema, ranges, types)?
- [ ] Are training/serving flows composable pipelines (not notebooks in production)?
- [ ] Does serving rollback restore the prior model artifact, not just the prior code?
### If i18n
- [ ] Is the source language treated as one locale among many, not the "neutral" default?
- [ ] Do locale identifiers use BCP 47 tags (no ad-hoc locale codes)?
- [ ] Are user-facing strings in locale resource files (not concatenated inline in code)?
- [ ] Do plural/gender/select use ICU MessageFormat (no `if (n == 1)` branching)?
- [ ] Are dates, times, numbers, currencies, units formatted via ICU/CLDR/`Intl` (no hand-rolled formatters)?
- [ ] Is RTL/bidi a first-class layout concern (logical CSS properties over physical)?
- [ ] Do layouts accommodate translation expansion (no fixed pixel widths for text)?
- [ ] Are pseudo-locales used to test before real translations arrive?
- [ ] Are icons, colors, and imagery reviewed for locale-sensitivity?
- [ ] Are resource files versioned (bad translation = rollback, not hot-patch)?
### If Compliance
- [ ] Scope check: is this framework-agnostic (no regulation-specific GDPR/HIPAA/SOC2/PCI content)? (D-024)
- [ ] Are audit records immutable once written (deletion/mutation is itself an auditable incident)?
- [ ] Is the set of auditable actions defined a priori ("we forgot to log it" is a violation)?
- [ ] Is data lifetime declared and enforced as policy (deletion at end-of-life is a feature)?
- [ ] Is compliance policy expressed in versioned, reviewable, testable code (not spreadsheets/prose)?
- [ ] Do policy violations block before the action (admission/CI/CD-time, not after the audit)?
- [ ] Is evidence gathered as a byproduct of operation (not assembled manually at audit time)?
- [ ] Does every logged action trace to an authenticated principal (no shared/generic identities)?
- [ ] Are data-subject rights (access, export, deletion) operations with defined contracts and audit trails?
- [ ] Do audit logs avoid leaking secrets (redaction is structural, not opportunistic)?
- [ ] Does the system report its own compliance state (drift from policy, open violations, retention status)?
## Review Etiquette ## Review Etiquette
- **Comment, don't command.** "This could be X" not "Change this to X." - **Comment, don't command.** "This could be X" not "Change this to X."