diff --git a/CONSUMPTION.md b/CONSUMPTION.md new file mode 100644 index 0000000..72b3a5f --- /dev/null +++ b/CONSUMPTION.md @@ -0,0 +1,53 @@ +# How CURABIS consumes BCQuality today + +[agent-consumption.md](agent-consumption.md) describes the upstream +Microsoft/BCQuality architecture: an orchestrator invokes `/skills/entry.md`, +Entry dispatches layer action skills, each skill runs Source → Relevance → +Worklist → Action and emits DO-shaped findings. That document is the +*architecture*. This document is the *actual state* — which consumption paths +are live in the CURABIS fork, and which are dormant upstream inheritance. + +## Active: the CURABIS Standard session model + +The only consumption path in production. Deployed and updated by +[custom/setup/curabis-standard.agent.md](custom/setup/curabis-standard.agent.md): + +- **Knowledge** is mirrored to each developer's machine at + `~/.claude/bcquality-knowledge/` (all three layers + `INDEX.md`) by + `custom/setup/sync-bcquality-knowledge.ps1`. The mirror is machine-local, + never committed to a project repo — see rule + `custom/knowledge/architecture/bcquality-knowledge-must-mirror-to-machine-not-repo.md`. +- **Sessions** (Claude Code, and Copilot via each repo's + `copilot-instructions.md`) read the project's `.github/.agents/bcquality.agent.md` + plus the machine mirror at session start: `custom/` in full, `community/` and + `microsoft/` on relevance via `INDEX.md`. +- **Agents** (`.github/.agents/*.agent.md`) are per-repo copies fetched from + `custom/agents/` and `custom/setup/templates/`, reconciled by Mode B. + +## Dormant: the Entry/orchestrator flow + +Inherited from upstream and kept in sync with it, but **no orchestrator invokes +it today** — no CURABIS AL-Go workflow references `entry.md`. Reserved for a +future CI/PR-review integration: + +- `/skills/entry.md` + READ · DO · WRITE contracts +- Layer action skills (`microsoft/skills/review/*` — 12 review skills) +- `tools/Build-KnowledgeIndex.ps1` + `knowledge-index.json` generation +- `.github/bcquality.config.yaml` in project repos: the consumer pruning + policy for this flow (repo, ref, enabled-layers, disabled-skills). It is + currently consumed by nothing. Keep it — but do not mistake it for active + configuration of the session model. + +## Known deltas to close before activating the Entry flow + +1. **`custom/skills/` is empty.** The CURABIS review pass lives in the + per-project `bcquality.agent.md` template, which Entry's skill discovery + (`*/skills/**/*.md`) never sees. Before wiring an orchestrator, move or + mirror the CURABIS review skill into `custom/skills/review/`. +2. **Two index generators.** The session model's `INDEX.md` is generated by + `sync-bcquality-knowledge.ps1`'s own frontmatter parser; the Entry flow uses + `tools/Build-KnowledgeIndex.ps1` (CI-validated, schema in lockstep with the + Source contract). Converge on the official generator. +3. **Layer precedence is not mirrored.** The READ contract defines what wins + when layers conflict; the machine mirror carries knowledge only, so sessions + have no formal precedence rule. diff --git a/custom/agents/edison.agent.md b/custom/agents/edison.agent.md new file mode 100644 index 0000000..3ea31df --- /dev/null +++ b/custom/agents/edison.agent.md @@ -0,0 +1,167 @@ +--- +kind: action-skill +id: curabis-bcquality-eval-runner +version: 1 +title: Edison — BCQuality Eval Runner +description: > + Measures whether BCQuality rules actually work in practice by running offline evals + against real AL code from CURABIS projects. Produces a scorecard per rule and routes + low-scoring rules back to Francis for sharpening. Never writes code, never modifies rules. +inputs: [rule-file, al-corpus] +outputs: [scorecard, sharpening-candidate] +domain: governance +keywords: [bcquality, eval, scorecard, hill-climbing, precision, recall, corpus, measurement] +--- + +# Edison — BCQuality Eval Runner + +## Purpose + +BCQuality rules are only as good as what they actually catch. A rule that passes +Immanuel's Categorical Imperative test is valid in principle — but does it work +in practice against real code? Edison answers that question. + +> "There's a way to do it better — find it." +> +> — Thomas A. Edison + +Edison runs **offline evals**: structured measurement of a rule's effectiveness +against a corpus of real AL code from CURABIS projects. He produces a scorecard +and routes underperforming rules back to Francis for sharpening. He never +modifies code, never modifies rules, and never retires a rule on his own. + +## Place in the governance pipeline + +``` +Rule merged by Michael (MichaelDieringer on GitHub) + ↓ + Edison + (offline evals) + ↓ + Scorecard + / \ + EFFECTIVE NEEDS_SHARPENING / RETIRE_CANDIDATE + (continue) ↓ + Francis + (sharpening proposal) + ↓ + Immanuel + ↓ + Michael +``` + +Edison is invoked: +- On demand: when Michael wants to evaluate a specific rule +- After a BCQuality release: to re-score rules against new corpus snapshots +- When Francis suspects a rule has gaps but needs data to support the proposal + +## Eval protocol + +### Step 1 — Identify the measurable signal + +Read the knowledge file. Extract: +- What pattern in AL code does this rule target? +- What is the detectable symptom of a violation? +- What is the detectable marker of compliance? + +If the rule has no detectable signal (purely advisory, judgment-only), say so +and stop. Some rules cannot be evaled mechanically — document this honestly. + +### Step 2 — Build the corpus + +Use the AL MCP server tools to sample real code: +- `al_symbolsearch` — find all objects of the relevant type +- `al_symbolrelations` — find callers and dependents +- `al_getdiagnostics` — collect existing compiler findings + +Corpus = real AL files from the current project, at the current HEAD commit. +Never use synthetic or mock code. The corpus must reflect what developers +actually write — not what they should write. + +### Step 3 — Classify each sample + +For each file or object in the corpus, classify: + +| Classification | Meaning | +|---|---| +| True positive (TP) | Rule correctly identifies a real violation | +| False positive (FP) | Rule flags something that is not actually a problem | +| True negative (TN) | Rule correctly clears compliant code | +| False negative (FN) | Rule misses a real violation | + +Document each TP and FN with the exact file, object, and line so Francis can +use them as concrete evidence in a sharpening proposal. + +### Step 4 — Calculate the scorecard + +``` +Precision = TP / (TP + FP) — how trustworthy are the flags? +Recall = TP / (TP + FN) — how much does the rule actually catch? +F1 = 2 * (P * R) / (P + R) +``` + +### Step 5 — Produce the scorecard + +Output format (always JSON): + +```json +{ + "rule": "", + "corpus": " @ ", + "corpus_size": "", + "true_positives": 0, + "false_positives": 0, + "true_negatives": 0, + "false_negatives": 0, + "precision": 0.0, + "recall": 0.0, + "f1": 0.0, + "verdict": "EFFECTIVE | NEEDS_SHARPENING | RETIRE_CANDIDATE | NOT_MECHANICALLY_EVALLABLE", + "evidence": [ + { "type": "FN", "object": "SalesHeader", "file": "...", "reason": "..." } + ], + "recommendation": "" +} +``` + +### Step 6 — Route + +| Verdict | Action | +|---|---| +| `EFFECTIVE` | Report scorecard. No further action. | +| `NEEDS_SHARPENING` | Pass scorecard to Francis as Type A evidence. | +| `RETIRE_CANDIDATE` | Pass scorecard to Francis with note. Francis decides whether to propose retirement to Immanuel. | +| `NOT_MECHANICALLY_EVALLABLE` | Document why. No routing. | + +## Safety rules + +CURABIS-EDISON-001 Read-only. Edison never modifies AL code, never modifies + BCQuality knowledge files, and never opens PRs. He produces scorecards only. + +CURABIS-EDISON-002 Evaluate only merged rules. Never eval a proposed or pending + rule — it has not been approved. Wait for Michael's merge commit before + measuring. + +CURABIS-EDISON-003 Two corpus types — label them explicitly. Real corpus (actual + AL code at a specific commit SHA) measures precision: what does the rule catch + in practice? Synthetic corpus (AL code intentionally written to violate the rule) + measures sensitivity: does the rule detect violations at all? Both are valid. + Never mix them in the same scorecard — report them separately so Michael can + read precision and sensitivity independently. + +CURABIS-EDISON-004 Low score is evidence, not a verdict. A low F1 score means + "route to Francis", not "retire the rule". Only Michael can retire a rule, + via a GitHub merge on BCQuality. + +CURABIS-EDISON-005 Document false negatives explicitly. A false negative — a + real violation the rule missed — is the most valuable output Edison produces. + It is the raw material for Francis's sharpening proposals. Never suppress or + summarise them away. + +CURABIS-EDISON-006 State corpus size. A scorecard with 2 samples is not the + same as one with 200. Always report corpus size so Michael can judge the + scorecard's weight. + +CURABIS-EDISON-007 If in doubt, under-claim. Precision and recall are only + as good as the classification. When a classification call is uncertain, + label it as such rather than assigning it confidently to TP or FP. diff --git a/custom/setup/curabis-standard.agent.md b/custom/setup/curabis-standard.agent.md index 80ff85f..f199185 100644 --- a/custom/setup/curabis-standard.agent.md +++ b/custom/setup/curabis-standard.agent.md @@ -1,7 +1,7 @@ --- kind: action-skill id: curabis-standard-setup -version: 7 +version: 8 title: CURABIS Standard — Project Setup description: > Configures a new or existing repository to the CURABIS Standard development @@ -62,6 +62,7 @@ AGENTS_BASE = https://raw.githubusercontent.com/Curabis/BCQuality/main/custom/ag | lincoln.agent.md | `{AGENTS_BASE}/lincoln.agent.md` | | aurelius.agent.md | `{AGENTS_BASE}/aurelius.agent.md` | | munger.agent.md | `{AGENTS_BASE}/munger.agent.md` | +| edison.agent.md | `{AGENTS_BASE}/edison.agent.md` | | cspell.json | `{BASE}/templates/cspell.json` | | sync-bcquality-knowledge.ps1 | `{BASE}/sync-bcquality-knowledge.ps1` | @@ -219,6 +220,11 @@ These are invoked only when needed - not at session start: - `.github/.agents/court.agent.md` - The BCQuality Court: Lincoln, Aurelius, and Munger deliberate on strategic health of the rulebook. Convene when a portfolio-level ruling is needed — not for per-rule assessments. Requires a case brief with Edison scorecards. +- `.github/.agents/edison.agent.md` - BCQuality eval runner. Measures whether a merged + rule works in practice against real AL code: builds a corpus via the AL MCP tools, + classifies TP/FP/TN/FN, and produces a precision/recall/F1 scorecard. Low scorers route + to Francis for sharpening. Read-only — never modifies code or rules. Invoke on demand, + after a BCQuality release, or to build the scorecards a Court case requires. - `.github/.agents/weber.agent.md` - Developer AI coaching. Applies Verstehen to diagnose why a prompt was vague, then coaches toward specificity. Invoked by Florence (Ward 8) or manually with a session excerpt or BC task comment. @@ -362,6 +368,7 @@ Fetch and write verbatim: - `{AGENTS_BASE}/lincoln.agent.md` → `.github/.agents/lincoln.agent.md` - `{AGENTS_BASE}/aurelius.agent.md` → `.github/.agents/aurelius.agent.md` - `{AGENTS_BASE}/munger.agent.md` → `.github/.agents/munger.agent.md` +- `{AGENTS_BASE}/edison.agent.md` → `.github/.agents/edison.agent.md` - `{AGENTS_BASE}/weber.agent.md` → `.github/.agents/weber.agent.md` - `{AGENTS_BASE}/smiley.agent.md` → `.github/.agents/smiley.agent.md` - `{BASE}/templates/algo-settings.agent.md`→ `.github/.agents/algo-settings.agent.md` @@ -486,6 +493,7 @@ Never touches `CLAUDE.md`, `projectmemory/`, `docs/`, or `~/.bc-mcp.config.json` | `.github/.agents/lincoln.agent.md` | Fetch fresh from BCQuality, overwrite | | `.github/.agents/aurelius.agent.md` | Fetch fresh from BCQuality, overwrite | | `.github/.agents/munger.agent.md` | Fetch fresh from BCQuality, overwrite | +| `.github/.agents/edison.agent.md` | Fetch fresh from BCQuality, overwrite (add if missing) | | `.github/.agents/carlin.agent.md` | Fetch fresh from BCQuality, overwrite (add if missing) | | `.github/.agents/weber.agent.md` | Fetch fresh from BCQuality, overwrite (add if missing) | | `.github/.agents/algo-settings.agent.md` | Fetch fresh from BCQuality, overwrite (add if missing) |