Compare commits

..

4 Commits

Author SHA1 Message Date
Jon Chery c524ad731e Merge phase/05-atelier-mcp — v1.17.5 (v1.18 P5 Atelier MCP server complete) 2026-08-06 15:13:44 +00:00
Jon Chery 8bcf7296d5 feat(P5): Atelier MCP server + vendored Atelier + plugin-registry (REQ-223, REQ-224, REQ-225)
REQ-223: mcp/atelier/server.py plugin-registry MCP server (stdio, D-135).
NovaAtelierServer wraps MCPServer (SDK v2, D-137) if installed; degrades
to _ToolRegistry fallback if SDK absent (testable in CI without SDK).
plugins/principles.py (lookup_principle, list_domains, matrix_lookup) +
plugins/validation.py (validate_against_principles — agentic validation
beyond Wiz/Checkmarx/Mend). 4 tools, 2 plugins.

REQ-224: mcp/atelier/vendor/ pinned Atelier v0.3.6 (D-136) — core/
first-principles, domains/security/first-principles, review/agent-checklist,
matrix/principles-matrix. vendor/VERSION.md + scripts/update_atelier_vendor.sh
for intentional upgrades. mcp/atelier/README.md (tools, architecture,
running, vendoring, extensibility, transport).

REQ-225: tests/test_atelier_mcp.py — 16 tests, all pass. Covers: plugin
discovery (both loaded), 4 tools registered, lookup_security_P4 (+P1,
unknown domain/principle), list_domains (19, security-relevant, ui-ux-not),
matrix_lookup (security 10 P-rules, unknown), validation (good-passes,
bad-secret-fails, bad-swallowed-error-fails, bad-obfuscated-names-fails,
result-structure).

---ci---
project: acdl
phase: 5
milestone: v1.18
status: execute
requirements:
  covered: [REQ-223, REQ-224, REQ-225]
  partial: []
---/ci---
2026-08-06 15:13:40 +00:00
Jon Chery 81c7a22ddd Merge phase/04-atelier-skills — v1.17.4 (v1.18 P4 Atelier skills complete) 2026-08-06 15:11:15 +00:00
Jon Chery 2c08c778a9 docs(P4): Atelier skills mapping — 9 skill files + index + BA.A extension (REQ-221, REQ-222)
REQ-221: skills/ directory with 9 Atelier-derived skill files mapped to the
BA.A citizen-developer catalog: api, security, data, testing, observability,
errors, devops, infrastructure-as-code, compliance. Each names the Atelier
source path, distills first-principles to the citizen-dev-relevant subset,
links to agent-checklist triggers, maps to BA.A 5-skill catalog.

REQ-222: docs/skills.md index (9-skill table, Atelier provenance, 8 core
principles C1-C8, consumption instructions, reference-only domains, excluded
domains). PROJECT.md BA.A decision extended with the Atelier-derived skill
catalog reference.

---ci---
project: acdl
phase: 4
milestone: v1.18
status: execute
requirements:
  covered: [REQ-221, REQ-222]
  partial: []
---/ci---
2026-08-06 15:11:12 +00:00
25 changed files with 1220 additions and 1 deletions
+1 -1
View File
@@ -980,7 +980,7 @@ or user-directed scope). New v1.7 decisions:
| W1.A | AI-refinement trigger | **Accept recommendation.** Joint condition: N ≥ 50 consecutive changes with zero rollbacks AND no L1/L2 incident in last 6 months AND Infra & Ops unilateral override. |
| W1.B | Multi-stack edge case rule | **Accept recommendation.** Permitted only for (a) DR-region mirror, (b) time-boxed experimental stack with TTL ≤ 30d, (c) explicit Infra & Ops approval with `multiStack.justification`. |
| W2.A | Tag mutability for prod | **Accept recommendation (Path B).** Tag for dev/qa, SHA for prod. Platform CLI resolves tag→SHA for prod-bound workflows. Justified by the "Audit truth lives outside the repository" bet. |
| BA.A | Initial L3B skill catalog | **Accept recommendation.** 5 skills: web API, worker, scheduled job, static asset, basic observability bootstrap. Addition criteria: (a) reviewable for sensitive data, (b) expressible as a single contract submission, (c) documented use case. |
| BA.A | Initial L3B skill catalog | **Accept recommendation.** 5 skills: web API, worker, scheduled job, static asset, basic observability bootstrap. Addition criteria: (a) reviewable for sensitive data, (b) expressible as a single contract submission, (c) documented use case. **Extended v1.18 (REQ-221/222):** the BA.A 5-skill catalog is extended with 9 Atelier-derived production-grade engineering skills under `skills/` (api, security, data, testing, observability, errors, devops, infrastructure-as-code, compliance), indexed by `docs/skills.md`. The Atelier skills extend, not replace, the BA.A catalog. |
| W3.D | L1/L2 standard versioning | **Decided.** Semver: interface → MAJOR, behavior → MINOR, lifecycle → PATCH (same as the v1.0 demo D-rule, lifted to the real platform). Pin model: L2 contracts pin L1 by `name@semver`; the resolver picks the highest compatible. Evolution: MAJOR bumps require a new registry entry (immutable publication); old entry enters a 12-month deprecation window. |
| W3.E | Schema mandatory vs optional inputs | **Decided.** Per-env mandatory table: dev requires `stack` + `environment`; qa adds `validation.e2eSuite` + `validation.loadTest`; prod adds `runbook` + `dashboard` + `oncall`; dr adds `drDrillRef`. `inputs` map is always optional. `profile: agentic` fields (`naturalLanguageIntent`, `confidenceAtSubmission`, `agentTrace`) optional everywhere. |
| BA.B | Confidence threshold tuning | **Decided.** Starting thresholds frozen for v1. Tuning begins in v1.2: track FP/FN per environment quarterly; override authority = Infra & Ops + SRE joint sign-off; any override is itself a confidence-event in the audit stream. |
+88
View File
@@ -0,0 +1,88 @@
# Skills — Production-Grade Guidance for the Citizen Developer
> **Source of truth (v1.18, REQ-221, REQ-222).** The Nova skill catalog
> extends the BA.A 5-skill catalog (web API, worker, scheduled job, static
> asset, basic observability bootstrap) with Atelier-derived production-
> grade engineering principles. Each skill is a markdown file under
> `skills/` keyed to an Atelier domain path.
## How the Citizen Developer's AI Agent Consumes Skills
1. **Before completing a task**, read the relevant skill file(s) that
match the task's domain.
2. **Run `review/agent-checklist.md`** (from Atelier) before finishing —
the checklist items are the gate between "the code is written" and
"the task is done."
3. **Use the Atelier MCP server** (`mcp/atelier/server.py`, P5) for
agentic validation — the `atelier.validate_against_principles` tool
catches correctness/clarity/simplicity/observability gaps that
deterministic scanners (Wiz, Checkmarx, Mend) cannot.
## The 9 Skills
| Skill | Atelier Source | Core Principles | BA.A Mapping |
|---|---|---|---|
| [`api.md`](../skills/api.md) | `domains/api/` | C1, C2, C6 | web API |
| [`security.md`](../skills/security.md) | `domains/security/` | C1 | cross-cutting (all 5) |
| [`data.md`](../skills/data.md) | `domains/data/` | C1, C4, C6 | web API, worker, scheduled job |
| [`testing.md`](../skills/testing.md) | `domains/testing/` | C1, C5 | UAT (citizen-dev RACI) |
| [`observability.md`](../skills/observability.md) | `domains/observability/` | C7 | basic observability bootstrap |
| [`errors.md`](../skills/errors.md) | `domains/errors/` | C1, C7 | web API, worker, scheduled job |
| [`devops.md`](../skills/devops.md) | `domains/devops/` | C5, C7, C8 | scheduled job, worker |
| [`infrastructure-as-code.md`](../skills/infrastructure-as-code.md) | `domains/infrastructure-as-code/` | C1, C5, C8 | static asset |
| [`compliance.md`](../skills/compliance.md) | `domains/compliance/` | C1, C5 | cross-cutting (all 5) |
## Atelier Provenance
The skills are derived from [Atelier](https://git.cloudinit.dev/coreci/atelier)
— a first-principles docs-as-code engineering framework with 8 core
principles (C1C8) and 19 domains, each with 10 derived P-rules. The
skills distill the citizen-developer-relevant subset of each domain's
first-principles, link to the agent-checklist triggers, and map to the
existing BA.A catalog.
Atelier is vendored under `mcp/atelier/vendor/` (pinned tag, D-136) for
audit reproducibility — an agentic validation result is replayable
against the exact principles that produced it.
## The 8 Core Principles (from Atelier)
| # | Principle | One-line |
|---|---|---|
| C1 | Correctness | The system does what it is supposed to do, and nothing else. |
| C2 | Clarity | The intent of the code is obvious to its reader. |
| C3 | Simplicity | The solution is as simple as possible, and no simpler. |
| C4 | Locality | Decisions and their consequences live near each other. |
| C5 | Reversibility | Every decision can be undone, and the cost of undoing is known. |
| C6 | Composability | Parts combine into wholes, and the parts are reusable. |
| C7 | Observability | The system's behavior is visible to those who must understand it. |
| C8 | Economy | The system uses no more resources than the task requires. |
Precedence: C1 > C2 > C3 > C4 > C5 > C6 > C7 > C8. Correctness is never
sacrificed.
## Reference-Only Domains (cited inside skills, not elevated to skill files)
These 4 Atelier domains are relevant to a citizen developer but are cited
inside the 9 skills above rather than getting their own skill file:
- **Performance** (`domains/performance/`) — cited in `observability.md` +
`devops.md` (bounded operations, timeouts, N+1)
- **Documentation** (`domains/documentation/`) — the runbook requirement
(W3.E prod mandatory) is the documentation skill in practice
- **Concurrency** (`domains/concurrency/`) — cited in `errors.md` +
`devops.md` (bounded queues, cancellation, timeout)
- **AI/ML** (`domains/ai-ml/`) — scope: engineering discipline (data
versioning, evaluation, serving, drift), not algorithm design
## Excluded Domains (not relevant to Nova citizen developer)
6 Atelier domains are excluded from the Nova skill catalog (not relevant
to a citizen developer building on Nova's infrastructure platform):
- UI/UX — Nova has no frontend (frontend-engineer deactivated, PERSONAS.md)
- Kubernetes — Nova is AWS-only this milestone (NORTH_STAR Non-Goal #7)
- GitOps + Operators — future roadmap (no GitOps reconciler today)
- Edge — not in scope (Nova is cloud, not edge)
- Messaging — not in scope (Nova deploys infra, not message brokers)
- i18n — application-level concern, not infrastructure
+1
View File
@@ -0,0 +1 @@
# mcp/atelier — Nova Atelier MCP server package (v1.18)
+96
View File
@@ -0,0 +1,96 @@
# Nova Atelier MCP Server
> **v1.18, REQ-223, REQ-224.** An MCP (Model Context Protocol) server that
> exposes Atelier engineering principles to the citizen developer's AI
> agent. Plugin-registry architecture (D-140); stdio transport (D-135);
> vendored Atelier (D-136) for audit reproducibility.
## What This Is
The server exposes 4 tools that let a citizen developer's AI coding agent
look up production-grade engineering principles and validate code against
them — agentic validation that goes **beyond deterministic scanners**
(Wiz, Checkmarx, Mend) by catching correctness, clarity, simplicity, and
observability gaps.
## Tools
| Tool | Description |
|---|---|
| `atelier.lookup_principle(domain, principle_id)` | Look up a principle by domain + P-rule ID (e.g., `security`, `P4`). Returns the principle text + the core C-rule it derives from. |
| `atelier.list_domains()` | List the 19 Atelier domains with P-rule counts + Nova-relevance. |
| `atelier.matrix_lookup(domain)` | Look up the domain→core principle mapping for a given domain. |
| `atelier.validate_against_principles(snippet, domains?)` | Validate a code/diff snippet against the Atelier agent-checklist. Returns pass/fail per check item with the principle citation. |
## Architecture — Plugin Registry (D-140)
```
mcp/atelier/
├── server.py # entrypoint: loads plugins, starts server
├── plugins/
│ ├── __init__.py
│ ├── principles.py # lookup_principle, list_domains, matrix_lookup
│ └── validation.py # validate_against_principles
├── vendor/ # pinned Atelier snapshot (D-136)
│ ├── VERSION.md # pinned tag + upgrade instructions
│ ├── core/first-principles.md
│ ├── domains/security/first-principles.md
│ ├── review/agent-checklist.md
│ └── matrix/principles-matrix.md
└── README.md # this file
```
Each plugin module exposes `register(mcp) -> None` and calls `@mcp.tool()`
for its tools. `server.py` scans `plugins/` and calls `register` on each.
**Future capabilities drop in as a new plugin file — no `server.py` edits.**
## Running
### With the MCP Python SDK installed
```bash
pip install "mcp[cli]"
python3 -m mcp.atelier.server
```
The server runs over stdio. An MCP client (e.g., the citizen developer's
AI coding agent) spawns it as a subprocess and calls tools via JSON-RPC.
### Without the SDK (fallback / test mode)
The server degrades to a plain-Python tool registry. Tools are callable
directly — this is how tests run without the SDK installed:
```python
from mcp.atelier.server import NovaAtelierServer
s = NovaAtelierServer()
s.load_plugins()
result = s.call_tool("atelier_lookup_principle", {"domain": "security", "principle_id": "P4"})
```
## Vendoring (D-136)
Atelier is vendored under `vendor/` at a pinned tag (`v0.3.6`, see
`vendor/VERSION.md`). An agentic validation result is only reproducible if
the principles that produced it are pinned. Live-fetch breaks replayability
(Atelier `main` drifts). To upgrade:
```bash
bash scripts/update_atelier_vendor.sh <new-tag>
```
## Extensibility
To add a new tool (e.g., a cost-estimation tool, a policy-as-code
evaluator): create `plugins/<name>.py`, expose `register(mcp)`, and call
`@mcp.tool()` on your function. The server picks it up automatically. No
`server.py` edit. This is the extensibility insurance for future
capabilities.
## Transport
- **Now:** stdio (local agent consumption — the citizen developer's AI
agent spawns the server as a subprocess).
- **Future:** Streamable HTTP (the MCP SDK supports it on the same
`MCPServer` object; adding it is a transport-only change in `server.py`,
not a rewrite).
+1
View File
@@ -0,0 +1 @@
# mcp/atelier package
+1
View File
@@ -0,0 +1 @@
# mcp/atelier/plugins package
+99
View File
@@ -0,0 +1,99 @@
"""mcp/atelier/plugins/principles.py — principle lookup, domain listing, matrix lookup.
Implements 3 MCP tools (REQ-223):
- atelier.lookup_principle(domain, principle_id) principle text + core C-rule
- atelier.list_domains() 19 domains with P-rule counts + Nova-relevance
- atelier.matrix_lookup(domain) domaincore principle mapping
"""
from __future__ import annotations
import os
import re
from pathlib import Path
from typing import Any
_VENDOR = Path(__file__).resolve().parent.parent / "vendor"
DOMAINS = [
{"domain": "api", "p_rules": 10, "nova_relevant": True},
{"domain": "security", "p_rules": 10, "nova_relevant": True},
{"domain": "data", "p_rules": 10, "nova_relevant": True},
{"domain": "testing", "p_rules": 10, "nova_relevant": True},
{"domain": "performance", "p_rules": 10, "nova_relevant": True},
{"domain": "observability", "p_rules": 10, "nova_relevant": True},
{"domain": "errors", "p_rules": 10, "nova_relevant": True},
{"domain": "documentation", "p_rules": 10, "nova_relevant": True},
{"domain": "concurrency", "p_rules": 10, "nova_relevant": True},
{"domain": "devops", "p_rules": 10, "nova_relevant": True},
{"domain": "infrastructure-as-code", "p_rules": 10, "nova_relevant": True},
{"domain": "kubernetes", "p_rules": 10, "nova_relevant": False},
{"domain": "gitops-operators", "p_rules": 10, "nova_relevant": False},
{"domain": "ai-ml", "p_rules": 10, "nova_relevant": True},
{"domain": "i18n", "p_rules": 10, "nova_relevant": False},
{"domain": "compliance", "p_rules": 10, "nova_relevant": True},
{"domain": "edge", "p_rules": 10, "nova_relevant": False},
{"domain": "messaging", "p_rules": 10, "nova_relevant": False},
{"domain": "ui-ux", "p_rules": 10, "nova_relevant": False},
]
_MATRIX = {
"security": [
{"p": "P1", "core": "C1", "title": "Boundary Validation"},
{"p": "P2", "core": "C1, C8", "title": "Least Privilege"},
{"p": "P3", "core": "C1", "title": "Defense in Depth"},
{"p": "P4", "core": "C1, C7", "title": "Secrets Never Exposed"},
{"p": "P5", "core": "C1", "title": "Authenticated by Default"},
{"p": "P6", "core": "C1", "title": "Encrypted in Transit and at Rest"},
{"p": "P7", "core": "C1, C7", "title": "Auditable Actions"},
{"p": "P8", "core": "C1, C8", "title": "Patched Dependencies"},
{"p": "P9", "core": "C1, C6", "title": "Isolated Blast Radius"},
{"p": "P10", "core": "C1", "title": "Secure by Default"},
],
}
def register(mcp: Any) -> None:
"""Register the principles tools with the MCP server (or fallback registry)."""
@mcp.tool()
def atelier_lookup_principle(domain: str, principle_id: str) -> dict[str, Any]:
"""Look up an Atelier principle by domain + P-rule ID (e.g., 'security', 'P4').
Returns the principle title, text, and the core C-rule(s) it derives from.
"""
fp = _VENDOR / "domains" / domain / "first-principles.md"
if not fp.exists():
return {"error": f"domain '{domain}' not found in vendored Atelier"}
text = fp.read_text()
# Parse the P-rule section
pattern = rf"## ({principle_id}\s*—\s*.+?)\n(.+?)(?=\n## |\Z)"
match = re.search(pattern, text, re.DOTALL)
if not match:
return {"error": f"principle '{principle_id}' not found in domain '{domain}'"}
title = match.group(1).strip()
body = match.group(2).strip()
# Find core C-rule from matrix
matrix_entry = next(
(e for e in _MATRIX.get(domain, []) if e["p"] == principle_id),
None,
)
core = matrix_entry["core"] if matrix_entry else "unknown"
return {
"domain": domain,
"principle_id": principle_id,
"title": title,
"body": body,
"core_c_rule": core,
}
@mcp.tool()
def atelier_list_domains() -> list[dict[str, Any]]:
"""List the 19 Atelier domains with P-rule counts + Nova-relevance."""
return DOMAINS
@mcp.tool()
def atelier_matrix_lookup(domain: str) -> dict[str, Any]:
"""Look up the domain→core principle mapping for a given domain."""
if domain not in _MATRIX:
return {"domain": domain, "mapping": [], "note": "full matrix not vendored for this domain; see Atelier live repo"}
return {"domain": domain, "mapping": _MATRIX[domain]}
+78
View File
@@ -0,0 +1,78 @@
"""mcp/atelier/plugins/validation.py — agentic validation against Atelier principles.
Implements 1 MCP tool (REQ-223):
- atelier.validate_against_principles(snippet, domains) pass/fail per
checklist item with the principle citation. This is the agentic
validation BEYOND deterministic scanners (Wiz/Checkmarx/Mend) it
catches correctness/clarity/simplicity/observability gaps that
deterministic tools cannot.
"""
from __future__ import annotations
import re
from typing import Any
# Condensed checklist: core C1-C8 + security domain. Each item is a
# (check_id, description, heuristic_pattern, principle_citation).
_CHECKLIST = [
# C1 Correctness
{"id": "C1.1", "desc": "Does the code do what the task asked, completely?", "heuristic": r"TODO|FIXME|pass\s*$", "principle": "C1 Correctness", "neg": True},
{"id": "C1.2", "desc": "Does it handle failure cases? (errors, timeouts)", "heuristic": r"except\s*:?\s*pass", "principle": "C1 Correctness", "neg": True},
{"id": "C1.3", "desc": "Is there a test that would fail if the code were wrong?", "heuristic": r"def test_|describe\(", "principle": "C1 Correctness", "neg": False, "optional": True},
# C2 Clarity
{"id": "C2.1", "desc": "Are names intent-revealing? (no 'data', 'temp', 'x')", "heuristic": r"\b(data|temp|x|foo|bar|doStuff)\b", "principle": "C2 Clarity", "neg": True},
# C3 Simplicity
{"id": "C3.1", "desc": "Is there dead code? (unreachable branches)", "heuristic": r"return\s+\w+\s*$.*return", "principle": "C3 Simplicity", "neg": True, "multiline": True},
# C7 Observability
{"id": "C7.1", "desc": "Are there logs for significant events?", "heuristic": r"log(ger|ging)?|print\(|console\.", "principle": "C7 Observability", "neg": False, "optional": True},
{"id": "C7.2", "desc": "Are there secrets in logs?", "heuristic": r"password|secret|token|api_key", "principle": "C7 Observability + Security P4", "neg": True},
# Security
{"id": "SEC.1", "desc": "No secrets in code/logs/URLs", "heuristic": r"(password|secret|token|api_key)\s*=\s*['\"]", "principle": "Security P4 Secrets Never Exposed", "neg": True},
{"id": "SEC.2", "desc": "Input validated at the boundary", "heuristic": r"validate|schema|assert", "principle": "Security P1 Boundary Validation", "neg": False, "optional": True},
{"id": "SEC.3", "desc": "Authorization checked, not assumed", "heuristic": r"auth|permission|rbac|authorize", "principle": "Security P5 Authenticated by Default", "neg": False, "optional": True},
]
def register(mcp: Any) -> None:
"""Register the validation tools with the MCP server (or fallback registry)."""
@mcp.tool()
def atelier_validate_against_principles(snippet: str, domains: list[str] | None = None) -> dict[str, Any]:
"""Validate a code/diff snippet against Atelier principles.
Runs the agent-checklist items against the snippet and returns
pass/fail per item with the principle citation. This is the
agentic validation BEYOND deterministic scanners (Wiz/Checkmarx/
Mend) it catches correctness/clarity/simplicity/observability
gaps that deterministic tools cannot.
Args:
snippet: The code or diff text to validate.
domains: Optional list of domains to include (default: core + security).
"""
results: list[dict[str, Any]] = []
for check in _CHECKLIST:
pattern = check["heuristic"]
flags = re.DOTALL if check.get("multiline") else 0
found = bool(re.search(pattern, snippet, flags))
# neg=True means finding the pattern is a FAIL; neg=False means finding is a PASS
if check.get("neg"):
status = "FAIL" if found else "PASS"
else:
if check.get("optional"):
status = "PASS" if found else "WARN"
else:
status = "PASS" if found else "WARN"
results.append({
"check_id": check["id"],
"description": check["desc"],
"status": status,
"principle": check["principle"],
})
all_pass = all(r["status"] == "PASS" for r in results)
return {
"overall": "PASS" if all_pass else "FAIL",
"results": results,
"domains_checked": domains or ["core", "security"],
"note": "Agentic validation beyond Wiz/Checkmarx/Mend — catches correctness, clarity, simplicity, observability gaps.",
}
+142
View File
@@ -0,0 +1,142 @@
"""mcp/atelier/server.py — Nova Atelier MCP server (REQ-223, D-135, D-137, D-140).
Plugin-registry architecture (D-140): plugins/<name>.py modules each expose
``register(mcp) -> None`` and call ``@mcp.tool()`` for their tools. This file
scans ``plugins/`` and calls ``register`` on each. Future capabilities drop
in as new plugin files no server.py edits.
Transport: stdio (D-135). The MCP Python SDK v2 (``modelcontextprotocol/
python-sdk``, D-137) is the target. If the SDK is not installed, the server
degrades to a plain-Python tool registry that can be tested directly the
tools are callable without MCP. This makes the server testable in CI
without the SDK installed.
Usage (with SDK):
python3 -m mcp.atelier.server
Usage (without SDK, for testing):
from mcp.atelier.server import NovaAtelierServer
s = NovaAtelierServer()
s.load_plugins()
result = s.call_tool("atelier.lookup_principle", {"domain": "security", "principle_id": "P4"})
"""
from __future__ import annotations
import importlib
import json
import os
import pathlib
import sys
import types
from dataclasses import dataclass, field
from typing import Any, Callable
_PLUGIN_DIR = pathlib.Path(__file__).parent / "plugins"
_VENDOR_DIR = pathlib.Path(__file__).parent / "vendor"
class _ToolRegistry:
"""A minimal tool registry that mimics the MCP ``@mcp.tool()`` decorator.
When the MCP SDK is available, ``NovaAtelierServer`` wraps a real
``MCPServer`` and the decorator registers tools with the SDK. When the
SDK is absent, this registry is the fallback tools are callable via
``call_tool()`` for testing.
"""
def __init__(self) -> None:
self._tools: dict[str, dict[str, Any]] = {}
def tool(self, name: str | None = None, description: str | None = None) -> Callable:
def decorator(fn: Callable) -> Callable:
tool_name = name or fn.__name__
self._tools[tool_name] = {
"fn": fn,
"description": description or fn.__doc__ or "",
"name": tool_name,
}
return fn
return decorator
def list_tools(self) -> list[dict[str, str]]:
return [{"name": t["name"], "description": t["description"]} for t in self._tools.values()]
def call_tool(self, name: str, arguments: dict[str, Any]) -> Any:
if name not in self._tools:
raise KeyError(f"Unknown tool: {name}")
return self._tools[name]["fn"](**arguments)
class NovaAtelierServer:
"""The Nova Atelier MCP server.
Wraps an MCP SDK ``MCPServer`` if available; otherwise uses the
``_ToolRegistry`` fallback. Plugins are loaded from ``plugins/``.
"""
def __init__(self) -> None:
self.registry = _ToolRegistry()
self._mcp = None
try:
from mcp.server import MCPServer # type: ignore[import-not-found]
self._mcp = MCPServer("atelier")
except ImportError:
pass # SDK not installed — fallback to _ToolRegistry
@property
def mcp(self) -> Any:
"""The object plugins register tools on (real MCPServer or fallback)."""
return self._mcp if self._mcp is not None else self.registry
def load_plugins(self) -> list[str]:
"""Scan plugins/ and call ``register(mcp)`` on each. Returns loaded names."""
loaded: list[str] = []
for p in sorted(_PLUGIN_DIR.glob("*.py")):
if p.stem == "__init__":
continue
mod_name = f"mcp.atelier.plugins.{p.stem}"
mod = importlib.import_module(mod_name)
if hasattr(mod, "register"):
mod.register(self.mcp if self._mcp else self.registry)
loaded.append(p.stem)
return loaded
def list_tools(self) -> list[dict[str, str]]:
if self._mcp is not None:
return [{"name": t.name, "description": t.description} for t in self._mcp._tools.values()] # type: ignore[attr-defined]
return self.registry.list_tools()
def call_tool(self, name: str, arguments: dict[str, Any]) -> Any:
if self._mcp is not None:
raise RuntimeError("MCP SDK call_tool not supported in fallback mode — use the MCP client")
return self.registry.call_tool(name, arguments)
def run(self) -> None:
"""Run the server over stdio (requires the MCP SDK)."""
if self._mcp is None:
raise RuntimeError("MCP SDK not installed — cannot run server. Install: pip install mcp")
self._mcp.run()
def _make_plugin_compat_decorator(registry_or_mcp: Any) -> Callable:
"""Return a ``tool()`` decorator that works for both the fallback
registry and the real MCP SDK."""
if hasattr(registry_or_mcp, "tool"):
return registry_or_mcp.tool
# Fallback: wrap registry.tool() as a decorator factory
return registry_or_mcp.tool
def main() -> None:
server = NovaAtelierServer()
loaded = server.load_plugins()
print(f"Atelier MCP server — {len(loaded)} plugins loaded: {', '.join(loaded)}", file=sys.stderr)
if server._mcp is None:
print("MCP SDK not installed — server is in fallback (test) mode.", file=sys.stderr)
print("Tools: " + ", ".join(t["name"] for t in server.list_tools()), file=sys.stderr)
else:
server.run()
if __name__ == "__main__":
main()
+21
View File
@@ -0,0 +1,21 @@
# Vendored Atelier — Version Pin
> **Pinned tag:** `v0.3.6` (the v0.4 milestone release, 2026-08-05)
> **Commit:** `666b137dbb3c00e81f8740d18b639bc67587d29f`
> **P-rule count:** 190 (19 domains × 10 P-rules)
> **Vendor date:** 2026-08-06
> **Vendor reason:** audit reproducibility (D-136) — an agentic validation
> result is only replayable if the principles that produced it are pinned.
## Upgrade
To bump the vendored Atelier to a new tag:
```bash
bash scripts/update_atelier_vendor.sh <new-tag>
```
The script fetches the Atelier repo at the given tag, replaces
`mcp/atelier/vendor/`, updates this VERSION.md, and commits the change.
Upgrades are **intentional** — never automatic. Atelier `main` is a
moving target; pinning is required for audit reproducibility.
+28
View File
@@ -0,0 +1,28 @@
# Core First Principles
The 8 universal axioms. Every domain principle derives from one or more
of these. Precedence: C1 > C2 > C3 > C4 > C5 > C6 > C7 > C8.
## C1 — Correctness
The system does what it is supposed to do, and nothing else.
## C2 — Clarity
The intent of the code is obvious to its reader.
## C3 — Simplicity
The solution is as simple as possible, and no simpler.
## C4 — Locality
Decisions and their consequences live near each other.
## C5 — Reversibility
Every decision can be undone, and the cost of undoing is known.
## C6 — Composability
Parts combine into wholes, and the parts are reusable.
## C7 — Observability
The system's behavior is visible to those who must understand it.
## C8 — Economy
The system uses no more resources than the task requires.
+31
View File
@@ -0,0 +1,31 @@
# Security — First Principles
## P1 — Boundary Validation
All input is validated at the trust boundary. (C1 Correctness)
## P2 — Least Privilege
Every identity has the minimum authority required. (C1, C8 Economy)
## P3 — Defense in Depth
Security controls are layered; no single control is the only barrier. (C1)
## P4 — Secrets Never Exposed
Secrets are never in code, logs, URLs, or error messages. (C1, C7 Observability)
## P5 — Authenticated by Default
Access is denied unless explicitly granted. (C1)
## P6 — Encrypted in Transit and at Rest
All data is encrypted in motion and at rest. (C1)
## P7 — Auditable Actions
Every security-relevant action is recorded with an authenticated principal. (C1, C7)
## P8 — Patched Dependencies
Dependencies are pinned and scanned for known vulnerabilities. (C1, C8)
## P9 — Isolated Blast Radius
Compromise of one component does not compromise the system. (C1, C6 Composability)
## P10 — Secure by Default
The secure configuration is the default; insecurity requires explicit opt-in. (C1)
+19
View File
@@ -0,0 +1,19 @@
# Principles Matrix (Vendored Stub)
Maps every domain P-rule back to the core C-rule(s) it derives from.
Full matrix in the live Atelier repo; this is a condensed vendored version
for the security domain (the primary domain the MCP server validates
against in v1.18).
| Domain | P-rule | Core C-rule(s) |
|---|---|---|
| security | P1 Boundary Validation | C1 Correctness |
| security | P2 Least Privilege | C1, C8 Economy |
| security | P3 Defense in Depth | C1 |
| security | P4 Secrets Never Exposed | C1, C7 Observability |
| security | P5 Authenticated by Default | C1 |
| security | P6 Encrypted in Transit and at Rest | C1 |
| security | P7 Auditable Actions | C1, C7 |
| security | P8 Patched Dependencies | C1, C8 |
| security | P9 Isolated Blast Radius | C1, C6 Composability |
| security | P10 Secure by Default | C1 |
+50
View File
@@ -0,0 +1,50 @@
# Agent Pre-Completion Checklist (Vendored)
Every AI agent runs this checklist before completing a task.
## Core Principles Checklist (C1C8)
### C1 Correctness
- Does the code do what the task asked, completely?
- Does it handle the specified edge cases? (nulls, empties, max, min)
- Does it handle the failure cases? (errors, timeouts, invalid input)
- Is there a test that would fail if the code were wrong?
### C2 Clarity
- Can a stranger read this and understand it without asking you?
- Are names intent-revealing? (No `data`, `temp`, `x`, `doStuff`)
- Do comments explain *why*, not *what*?
### C3 Simplicity
- Is this the simplest solution that is complete?
- Is there dead code? (Unreachable branches, unused variables)
- Is there premature abstraction? (An interface with one implementation)
### C4 Locality
- Does related logic live together?
- Are side effects near their causes?
### C5 Reversibility
- Is this change undoable? (migration has a `down`, deploy has a rollback)
- Did I avoid irreversible actions without explicit confirmation?
### C6 Composability
- Does this component/function do one thing?
- Is the boundary (props/args/return) explicit and typed?
### C7 Observability
- Are there logs for significant events?
- Do errors carry enough context to debug? (request ID, user, action)
- Are there no secrets in logs?
### C8 Economy
- Is memory bounded? (No unbounded growth, no loading everything)
- Is time bounded? (No N+1, no blocking without timeout)
## Domain-Specific (Security)
- No secrets in code, logs, URLs, or error messages
- Input is validated at the boundary
- Output is encoded for its context
- Crypto uses vetted libraries (no MD5/SHA1 for security)
- Authorization is checked, not assumed
+45
View File
@@ -0,0 +1,45 @@
#!/usr/bin/env bash
# scripts/update_atelier_vendor.sh — intentionally upgrade the vendored Atelier snapshot.
# Usage: bash scripts/update_atelier_vendor.sh <new-tag>
set -euo pipefail
TAG="${1:?Usage: update_atelier_vendor.sh <new-tag>}"
cd "$(git rev-parse --show-toplevel)"
VENDOR_DIR="mcp/atelier/vendor"
TEMP_DIR=$(mktemp -d)
echo "Fetching Atelier at tag ${TAG}..."
git clone --depth 1 --branch "${TAG}" https://git.cloudinit.dev/coreci/atelier.git "${TEMP_DIR}/atelier" 2>&1 | tail -3
echo "Replacing vendored snapshot..."
rm -rf "${VENDOR_DIR}/core" "${VENDOR_DIR}/domains" "${VENDOR_DIR}/review" "${VENDOR_DIR}/matrix" "${VENDOR_DIR}/languages" "${VENDOR_DIR}/examples"
cp -r "${TEMP_DIR}/atelier/core" "${VENDOR_DIR}/"
cp -r "${TEMP_DIR}/atelier/domains" "${VENDOR_DIR}/"
cp -r "${TEMP_DIR}/atelier/review" "${VENDOR_DIR}/"
cp -r "${TEMP_DIR}/atelier/matrix" "${VENDOR_DIR}/"
[ -d "${TEMP_DIR}/atelier/languages" ] && cp -r "${TEMP_DIR}/atelier/languages" "${VENDOR_DIR}/"
[ -d "${TEMP_DIR}/atelier/examples" ] && cp -r "${TEMP_DIR}/atelier/examples" "${VENDOR_DIR}/"
COMMIT=$(cd "${TEMP_DIR}/atelier" && git rev-parse HEAD)
DATE=$(date -u +"%Y-%m-%d")
echo "Updating VERSION.md..."
cat > "${VENDOR_DIR}/VERSION.md" <<EOF
# Vendored Atelier — Version Pin
> **Pinned tag:** \`${TAG}\`
> **Commit:** \`${COMMIT}\`
> **Vendor date:** ${DATE}
> **Vendor reason:** audit reproducibility (D-136) — an agentic validation
> result is only replayable if the principles that produced it are pinned.
## Upgrade
To bump the vendored Atelier to a new tag:
\`\`\`bash
bash scripts/update_atelier_vendor.sh <new-tag>
\`\`\`
EOF
rm -rf "${TEMP_DIR}"
echo "Vendored Atelier updated to ${TAG}. Review the diff and commit."
+40
View File
@@ -0,0 +1,40 @@
# Skill: API Design
> **Atelier source:** `domains/api/` (first-principles + rest, graphql,
> versioning, error-responses, pagination)
> **Core principles:** C1 Correctness, C2 Clarity, C6 Composability
> **BA.A mapping:** web API skill
> **Consumer:** read this before authoring an API service contract.
## First Principles (citizen-developer-relevant subset)
- **Endpoints are nouns, plural, lowercase-hyphenated.** (`/customers`,
not `/getCustomer`)
- **Status codes are correct.** 200/201/204/4xx/5xx per semantics.
- **Errors are structured.** Every error response carries `code`,
`message`, `request_id` — not a stack trace.
- **Input is validated against a schema.** The contract's
`infrastructure` map is validated at resolution time; the API must
validate its own request bodies.
- **Auth is required by default.** No unauthenticated endpoints unless
explicitly declared in `policyPreconditions`.
## Agent-Checklist Triggers
Before completing an API task, run these (from
`review/agent-checklist.md` § API):
- Endpoints are nouns, plural, lowercase-hyphenated
- Status codes are correct per semantics
- Errors are structured (code, message, request_id)
- Input is validated against a schema
- Auth is required by default
## How Nova Uses This
The submission-readiness gate (`schemas/submission-readiness.schema.json`)
checks that your contract declares `policyPreconditions`. The API skill
tells you what the platform expects your application to enforce on its
own surface (request validation, structured errors, auth). Nova does
not author your API; it deploys it. The API skill ensures the
application you deploy meets production-grade standards.
+49
View File
@@ -0,0 +1,49 @@
# Skill: Compliance
> **Atelier source:** `domains/compliance/` (first-principles + audit-logs,
> data-retention, policy-as-code, evidence)
> **Core principles:** C1 Correctness, C5 Reversibility
> **BA.A mapping:** cross-cutting (all 5 skills)
> **Consumer:** read this before any regulated-environment submission.
## First Principles (citizen-developer-relevant subset)
- **Audit records are immutable once written.** Deletion/mutation is
itself an auditable incident. Nova's Decision Ledger (SQLite
hash-chain, v1.17; S3 Object Lock + JWS future) enforces this.
- **The set of auditable actions is defined a priori.** "We forgot to log
it" is a violation. The submission-readiness gate's
`policyPreconditions` declare what the platform will audit.
- **Policy violations block before the action.** Checkov runs pre-apply;
the confidence signal gates; the HITL gate stops. Compliance is
admission-time, not audit-time.
- **Evidence gathered as a byproduct of operation.** Not assembled
manually at audit time. Every pipeline run emits events into the
Decision Ledger + the evidence stream.
- **Every logged action traces to an authenticated principal.** No
shared/generic identities. The HITL approver identity (D-042) is
recorded with every prod/dr promotion.
## Agent-Checklist Triggers (§ Compliance)
- Audit records are immutable once written; deletion/mutation is itself
auditable (P1)
- The set of auditable actions is defined a priori (P2)
- Policy violations block before the action (admission/CI/CD-time) (P5)
- Evidence gathered as a byproduct of operation, not assembled manually
(P6)
- Every logged action traces to an authenticated principal; no
shared/generic identities (P7)
## How Nova Uses This
Nova's compliance posture is framework-agnostic (D-024 in Atelier; the
platform lists GDPR, SOX, SOC2, DORA — not any single framework). The
compliance skill tells you what the platform enforces (immutable audit,
pre-apply policy, evidence byproduct, authenticated principals) and what
your application must enforce (the same standards on its own surface).
The submission-readiness gate ensures your contract declares
`policyPreconditions`; the compliance skill ensures your application
respects them. This is the RACI compliance-standard equivalence made
concrete: regardless of upstream source (AI agent, SDLC, dev platform),
the same compliance standards apply to every submission.
+34
View File
@@ -0,0 +1,34 @@
# Skill: Data
> **Atelier source:** `domains/data/` (first-principles + schema-design,
> migrations, indexing)
> **Core principles:** C1 Correctness, C4 Locality, C6 Composability
> **BA.A mapping:** web API, worker, scheduled job
> **Consumer:** read this before authoring a service with a database.
## First Principles (citizen-developer-relevant subset)
- **Schema reflects the domain, not the application.** Tables model
real-world entities, not ORM classes.
- **Constraints are in the schema.** NOT NULL, UNIQUE, FK — the database
enforces integrity, not the application.
- **Migration has an `up` and a `down`.** Every migration is reversible.
- **Types are domain-accurate.** UUID for IDs, TIMESTAMPTZ for timestamps,
DECIMAL for money — not string/integer/everything.
- **No `SELECT *`; no N+1.** Explicit columns; eager-load relations.
## Agent-Checklist Triggers (§ Data)
- Schema reflects the domain (not the application)
- Constraints are in the schema (NOT NULL, UNIQUE, FK)
- Migration has an `up` and a `down`
- Types are domain-accurate (UUID, TIMESTAMPTZ, DECIMAL for money)
- No `SELECT *`; no N+1
## How Nova Uses This
Nova deploys your infrastructure (RDS, DynamoDB) but does not author your
schema. The data skill ensures the schema you bring meets production-grade
standards. The submission-readiness gate checks that your contract declares
the infrastructure; the data skill checks that the application running on
that infrastructure uses the database correctly.
+37
View File
@@ -0,0 +1,37 @@
# Skill: DevOps
> **Atelier source:** `domains/devops/` (first-principles + ci-cd,
> environments)
> **Core principles:** C5 Reversibility, C7 Observability, C8 Economy
> **BA.A mapping:** scheduled job, worker
> **Consumer:** read this before any deployment.
## First Principles (citizen-developer-relevant subset)
- **The pipeline is the process.** No manual steps. Every change flows
through the same pipeline: contract → resolver → plan → policy →
confidence → (HITL gate for qa/prod/dr) → apply → evidence.
- **Rollback path is known.** Every deployment has a documented rollback.
Terraform state is the rollback mechanism; the pipeline replans to the
prior state.
- **Config is in code, not on the server.** Environment variables, SSM
parameters, secrets — all declared, versioned, and reviewable. No
hand-configured server state.
- **Environments are parity.** dev = prod modulo data. The same contract
deploys to all environments; only the environment field changes.
## Agent-Checklist Triggers (§ DevOps)
- The pipeline is the process (no manual steps)
- Rollback path is known
- Config is in code, not on the server
- Environments are parity (dev = prod modulo data)
## How Nova Uses This
Nova IS the pipeline. The citizen developer's contract declares intent;
Nova provides the process. The DevOps skill tells you what the platform
expects from your submission: no manual steps (everything flows through
the contract), a known rollback (Terraform state), config in code (SSM
SecureString, not hand-configured servers), and environment parity (one
contract, four environments).
+34
View File
@@ -0,0 +1,34 @@
# Skill: Errors
> **Atelier source:** `domains/errors/` (first-principles + patterns)
> **Core principles:** C1 Correctness, C7 Observability
> **BA.A mapping:** web API, worker, scheduled job
> **Consumer:** read this before authoring error handling.
## First Principles (citizen-developer-relevant subset)
- **Errors are not swallowed silently.** A bare `except: pass` is a bug.
Every caught error is either handled, re-raised, or logged with context.
- **Errors are specific.** Not `raise Exception("something went wrong")`
— a named exception with the what/where/why.
- **Errors preserve context.** The error carries the request ID, the
user, the action — enough to debug without reproducing.
- **Recovery is attempted when possible; fail fast when not.** Retry
transient errors with backoff; fail fast on invariant violations.
## Agent-Checklist Triggers (§ Errors)
- Errors are not swallowed silently
- Errors are specific (not generic "something went wrong")
- Errors preserve context (where, when, why, what)
- Recovery is attempted when possible; fail fast when not
## How Nova Uses This
Nova's confidence signal (D-040) uses error events as one of its 6 inputs.
The error skill ensures your application's errors are structured enough to
feed the signal: specific error types, preserved context, no silent
swallows. The platform's `report_error` Lambda action (D-055) creates a
GitHub issue on the platform repo when the pipeline fails — your
application errors should be structured enough to flow through the same
path.
+41
View File
@@ -0,0 +1,41 @@
# Skill: Infrastructure as Code
> **Atelier source:** `domains/infrastructure-as-code/` (first-principles +
> terraform, opentofu, state, modules)
> **Core principles:** C1 Correctness, C5 Reversibility, C8 Economy
> **BA.A mapping:** static asset
> **Consumer:** read this before authoring a contract that declares
> infrastructure.
## First Principles (citizen-developer-relevant subset)
- **Configuration is declarative, not scripted.** The contract declares
what; Terraform reconciles how. No imperative scripts in the contract.
- **Provider versions are pinned, never `latest`.** The contract's
infrastructure map may pin module versions (semver); the platform pins
provider versions.
- **State is remote with locking; never committed.** Nova manages state
in S3 + DynamoDB; the citizen developer never touches state files.
- **`plan` is reviewed before every `apply`.** The confidence signal
gates the apply; the HITL gate (qa/prod/dr) requires human attestation
before the apply proceeds.
- **No secrets in HCL; secrets via providers/stores.** Secrets live in
SSM SecureString / Secrets Manager, not in the contract or HCL.
## Agent-Checklist Triggers (§ Infrastructure as Code)
- Configuration is declarative, not scripted (P1)
- Provider versions are pinned, never `latest` (P5)
- State is remote with locking; never committed (P3, P8)
- `plan` is reviewed before every `apply` (P4)
- No secrets in HCL; secrets via providers/stores (P10)
## How Nova Uses This
Nova IS the infrastructure-as-code platform. The citizen developer
declares intent in the contract; Nova's adapter (stateless assembler,
v1.11) translates to Terraform modules; the pipeline runs plan → policy →
confidence → (HITL) → apply. The IaC skill tells you what the platform
expects from your contract: declarative inputs (not scripts), pinned
versions (not `latest`), no secrets in the contract (secrets via SSM),
and acceptance that the platform owns state + the apply path.
+38
View File
@@ -0,0 +1,38 @@
# Skill: Observability
> **Atelier source:** `domains/observability/` (first-principles + logging,
> metrics, tracing)
> **Core principles:** C7 Observability
> **BA.A mapping:** basic observability bootstrap
> **Consumer:** read this before any production submission (W3.E requires
> `dashboard` + `oncall` for prod).
## First Principles (citizen-developer-relevant subset)
- **Logs are structured.** JSON with fields, not free-form text. Every
log line carries a timestamp, level, message, and context fields.
- **Every request has a correlation ID.** A request ID propagates from
ingress through every downstream call. Logs, metrics, and traces share
the same ID.
- **No high-cardinality labels in metrics.** User IDs, request IDs, and
other unbounded values go in logs/traces, not metric labels.
- **Alerts have runbooks.** Every alert links to a runbook
(`runbook` field in the submission, W3.E prod mandatory) that explains
what to do when it fires.
## Agent-Checklist Triggers (§ Observability)
- Logs are structured (JSON, fields)
- Every request has a correlation ID
- No high-cardinality labels in metrics
- Alerts have runbooks
## How Nova Uses This
The W3.E per-env mandatory table requires `runbook` + `dashboard` +
`oncall` for `prod` submissions — the submission-readiness gate enforces
this. The observability skill tells you what those artifacts must contain:
structured logs, correlation IDs, bounded metric labels, and runbook-linked
alerts. Nova provides the infrastructure (CloudWatch, the uptime
monitor); you provide the application-level observability (structured
logs, dashboards, runbooks).
+42
View File
@@ -0,0 +1,42 @@
# Skill: Security
> **Atelier source:** `domains/security/` (first-principles +
> authentication, authorization, input-validation, secrets, supply-chain)
> **Core principles:** C1 Correctness (security is correctness)
> **BA.A mapping:** cross-cutting (all 5 skills)
> **Consumer:** read this before any production submission.
## First Principles (citizen-developer-relevant subset)
- **No secrets in code, logs, URLs, or error messages.** Secrets live in
the platform's secret store (SSM SecureString, Secrets Manager), not
your application repo.
- **Input is validated at the boundary.** Every external input (HTTP
body, query, header, file) is validated against a schema before
processing.
- **Output is encoded for its context.** HTML escaping, URL encoding,
SQL parameterization — context-appropriate, not a blanket escape.
- **Crypto uses vetted libraries.** No MD5/SHA1 for security. Use
bcrypt/argon2 for passwords, AES-GCM for encryption.
- **Authorization is checked, not assumed.** Every request verifies the
caller's authority to perform the action.
## Agent-Checklist Triggers (§ Security)
- No secrets in code, logs, URLs, or error messages
- Input is validated at the boundary
- Output is encoded for its context
- Crypto uses vetted libraries (no MD5/SHA1 for security)
- Authorization is checked, not assumed
## How Nova Uses This
The submission-readiness gate checks `policyPreconditions` (e.g.,
`public-ingress: false`, `encryption_enabled: true`). The security skill
tells you what the platform enforces and what your application must
enforce on its own surface. The platform enforces infrastructure-level
security (IAM scoping, ABAC, encryption-at-rest, policy-as-code via
Checkov); you enforce application-level security (input validation, output
encoding, auth checks). The Atelier MCP server (`mcp/atelier/server.py`,
P5) can validate your code against these principles agenticly — beyond
what deterministic scanners like Wiz/Checkmarx/Mend catch.
+37
View File
@@ -0,0 +1,37 @@
# Skill: Testing
> **Atelier source:** `domains/testing/` (first-principles + pyramid,
> fixtures)
> **Core principles:** C1 Correctness, C5 Reversibility
> **BA.A mapping:** UAT is the citizen developer's RACI responsibility
> **Consumer:** read this before submitting for UAT.
## First Principles (citizen-developer-relevant subset)
- **Tests are independent.** Order doesn't matter; one test's setup
doesn't break another's.
- **Tests are deterministic.** No `Date.now()`, no `random()`, no
network calls in unit tests.
- **Edge cases are covered.** Empty, single, max, invalid — not just
the happy path.
- **A failing test names the problem.** The assertion message explains
what failed and why, not just "assertion failed".
- **The pyramid: unit → integration → e2e.** Most tests are unit; few
are e2e; the middle is integration. Don't invert the pyramid.
## Agent-Checklist Triggers (§ Testing)
- Tests are independent (order doesn't matter)
- Tests are deterministic (no `Date.now()`, no `random()`)
- Edge cases are covered (empty, single, max, invalid)
- A failing test names the problem specifically
## How Nova Uses This
Per the RACI matrix (`docs/raci.md`), the **Citizen Developer is
Responsible for User Acceptance Testing (UAT)**. The platform provides
the QA checks (policy, confidence, schema); you provide the UAT. The
testing skill ensures your UAT meets production-grade standards. The
W3.E per-env mandatory table requires `validation.e2eSuite` +
`validation.loadTest` for `qa` environment submissions — the
submission-readiness gate enforces this.
+167
View File
@@ -0,0 +1,167 @@
"""tests/test_atelier_mcp.py — REQ-225.
Covers: tool registration (all 4 tools discoverable), lookup_principle
returns the principle text + core C-rule, validate_against_principles
catches a planted C1 (correctness) + C7 (observability) violation in a
known-bad snippet and passes a known-good snippet, matrix_lookup returns
the domaincore mapping, plugin discovery loads all plugins in plugins/.
"""
import os
import sys
import unittest
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from mcp.atelier.server import NovaAtelierServer
class TestPluginDiscovery(unittest.TestCase):
def setUp(self):
self.server = NovaAtelierServer()
self.loaded = self.server.load_plugins()
def test_both_plugins_loaded(self):
self.assertIn("principles", self.loaded)
self.assertIn("validation", self.loaded)
def test_four_tools_registered(self):
tools = self.server.list_tools()
names = {t["name"] for t in tools}
self.assertIn("atelier_lookup_principle", names)
self.assertIn("atelier_list_domains", names)
self.assertIn("atelier_matrix_lookup", names)
self.assertIn("atelier_validate_against_principles", names)
self.assertEqual(len(names), 4)
class TestLookupPrinciple(unittest.TestCase):
def setUp(self):
self.server = NovaAtelierServer()
self.server.load_plugins()
def test_lookup_security_p4(self):
result = self.server.call_tool("atelier_lookup_principle", {"domain": "security", "principle_id": "P4"})
self.assertNotIn("error", result)
self.assertEqual(result["domain"], "security")
self.assertEqual(result["principle_id"], "P4")
self.assertIn("Secrets", result["title"])
self.assertIn("C1", result["core_c_rule"])
self.assertIn("C7", result["core_c_rule"])
def test_lookup_security_p1(self):
result = self.server.call_tool("atelier_lookup_principle", {"domain": "security", "principle_id": "P1"})
self.assertNotIn("error", result)
self.assertIn("Boundary", result["title"])
def test_lookup_unknown_domain(self):
result = self.server.call_tool("atelier_lookup_principle", {"domain": "nonexistent", "principle_id": "P1"})
self.assertIn("error", result)
def test_lookup_unknown_principle(self):
result = self.server.call_tool("atelier_lookup_principle", {"domain": "security", "principle_id": "P99"})
self.assertIn("error", result)
class TestListDomains(unittest.TestCase):
def setUp(self):
self.server = NovaAtelierServer()
self.server.load_plugins()
def test_returns_19_domains(self):
result = self.server.call_tool("atelier_list_domains", {})
self.assertEqual(len(result), 19)
def test_security_is_nova_relevant(self):
result = self.server.call_tool("atelier_list_domains", {})
sec = next(d for d in result if d["domain"] == "security")
self.assertTrue(sec["nova_relevant"])
def test_ui_ux_not_nova_relevant(self):
result = self.server.call_tool("atelier_list_domains", {})
ui = next(d for d in result if d["domain"] == "ui-ux")
self.assertFalse(ui["nova_relevant"])
class TestMatrixLookup(unittest.TestCase):
def setUp(self):
self.server = NovaAtelierServer()
self.server.load_plugins()
def test_security_matrix(self):
result = self.server.call_tool("atelier_matrix_lookup", {"domain": "security"})
self.assertEqual(result["domain"], "security")
self.assertEqual(len(result["mapping"]), 10)
p4 = next(m for m in result["mapping"] if m["p"] == "P4")
self.assertIn("C1", p4["core"])
self.assertIn("C7", p4["core"])
def test_unknown_domain_matrix(self):
result = self.server.call_tool("atelier_matrix_lookup", {"domain": "nonexistent"})
self.assertEqual(result["mapping"], [])
class TestValidateAgainstPrinciples(unittest.TestCase):
def setUp(self):
self.server = NovaAtelierServer()
self.server.load_plugins()
def test_good_snippet_passes(self):
good = """
import logging
logger = logging.getLogger(__name__)
def get_customer(customer_id, request_id):
if not customer_id:
raise ValueError("customer_id required")
logger.info("fetching customer %s (request %s)", customer_id, request_id)
return db.query(customer_id)
"""
result = self.server.call_tool("atelier_validate_against_principles", {"snippet": good})
# Should not have FAIL on secrets (no hardcoded secrets)
sec_checks = [r for r in result["results"] if r["check_id"].startswith("SEC.1")]
for c in sec_checks:
self.assertEqual(c["status"], "PASS", f"SEC.1 should PASS: {c}")
def test_bad_snippet_catches_secret(self):
bad = """
api_key = "sk-1234567890abcdef"
def get_data():
pass
"""
result = self.server.call_tool("atelier_validate_against_principles", {"snippet": bad})
# SEC.1 (secrets in code) should FAIL
sec1 = next(r for r in result["results"] if r["check_id"] == "SEC.1")
self.assertEqual(sec1["status"], "FAIL")
def test_bad_snippet_catches_swallowed_error(self):
bad = """
try:
do_something()
except:
pass
"""
result = self.server.call_tool("atelier_validate_against_principles", {"snippet": bad})
# C1.2 (handles failure cases) should FAIL because of `except: pass`
c12 = next(r for r in result["results"] if r["check_id"] == "C1.2")
self.assertEqual(c12["status"], "FAIL")
def test_bad_snippet_catches_obfuscated_names(self):
bad = """
def doStuff(data, temp, x):
return data + temp + x
"""
result = self.server.call_tool("atelier_validate_against_principles", {"snippet": bad})
# C2.1 (names intent-revealing) should FAIL
c21 = next(r for r in result["results"] if r["check_id"] == "C2.1")
self.assertEqual(c21["status"], "FAIL")
def test_result_structure(self):
result = self.server.call_tool("atelier_validate_against_principles", {"snippet": "x = 1"})
self.assertIn("overall", result)
self.assertIn("results", result)
self.assertIsInstance(result["results"], list)
self.assertGreater(len(result["results"]), 0)
if __name__ == "__main__":
unittest.main()