Narrow development guidance to provisional contract

Keep plan enrichment internal and read-only pending consumer agreement and runtime pilot evidence. Move knowledge to its independent PR and remove the consumer-owned forensic evaluator.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 638b66d2-9f06-4f60-8781-808709e1485c
This commit is contained in:
Jesper Schulz-Wedde 2026-09-18 15:46:52 +02:00
parent 8f025ac679
commit fa7eb750c6
40 changed files with 294 additions and 2053 deletions

View file

@ -1,4 +1,4 @@
# AL review and guidance evaluation
# AL review evaluation
The evaluation is convention-driven. The harness discovers every `<layer>/skills/review/al-<domain>-review.md` leaf across the enabled `microsoft`, `community`, and `custom` layers. Duplicate domains resolve with `custom > community > microsoft` precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both `.bad.al` and `.good.al` companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.
@ -54,117 +54,3 @@ This credential-free check proves every selected leaf maps to a same-named knowl
For a single combined stress-test result, use `-ResultsPath` instead.
The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.
## Read-only plan guidance
`development-guidance-fixtures.json` evaluates the planning interface used by
existing workflows. It supplies an existing plan and expects referenced
constraints without target-repository changes. The initial-plan fixture is
anonymized and synthetic: a full document with metadata and a markdown body
covering root cause, proposed fix, affected files, tests, and acceptance
criteria. It is integration-shaped input, not private consumer content, a
continuation checkpoint, or proof that any production consumer is integrated.
### Credential-free contract and scorer coverage
CI validates the manifest, prepares opaque model requests, and runs
deterministic scorer regressions with controlled reports and temporary Git
repositories. These checks cover report shape, outcomes, reference paths, and
the evaluator's pre/post read-only comparison. They do **not** run an agent,
compile AL, run Business Central tests, or establish better code authoring.
Prepare requests in a runner-owned artifact directory outside every target
workspace:
```powershell
$run = Join-Path ([IO.Path]::GetTempPath()) 'bcquality-guidance-run'
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . -PrepareDirectory $run
```
Preparation is not a model run. A scorer can validate a citation's path and
required fields, but only external agent traces and expert evaluation can
establish that the article was opened and its normative constraints faithfully
applied. Expected knowledge recall/precision is fixture-specific, not a corpus
coverage or authoring capability percentage.
`no-knowledge` means no additional applicable BCQuality constraints, not unsafe
work. It requires empty `knowledge`. Partial evaluation, failed retrieval,
unknown context, and materially unresolved applicability must remain visible
and distinct; they cannot be counted as successful enrichment merely because
the JSON is parseable.
Run the deterministic regression suite without an agent or AL environment:
```powershell
pwsh ./tools/Test-DevelopmentGuidanceEvaluator.ps1
```
### Runner-owned read-only evidence
For an external guidance run, first provision a representative, standalone Git
repository for each manifest case. The runner supplies a JSON workspace map
whose keys are the manifest IDs (not the hashed model IDs) and whose values are
absolute workspace roots. It may pass `-WorkspaceMapPath` during preparation
to bind the generated requests to those roots. The model must not select its
own workspace for scoring.
Capture evidence **before** invoking the agent, with the manifest, workspace
map, and source checkout already finalized:
```powershell
# Runner-selected paths, all outside the targets and BCQuality checkout.
$map = Join-Path $evidenceDirectory 'workspace-map.json'
$baseline = Join-Path $evidenceDirectory 'baseline.json'
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
-CaptureBaseline -WorkspaceMapPath $map -BaselinePath $baseline
# Retain the printed SHA256 in runner-only state BEFORE agent invocation.
# After the external agent writes result-case-<hash>.json files:
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
-ResultsDirectory $resultsDirectory -BaselinePath $baseline `
-BaselineSha256 $preRunDigest
```
`$evidenceDirectory`, `$resultsDirectory`, and `$preRunDigest` are supplied by
the runner; the digest must not be recomputed from potentially modified evidence
after the agent runs. Protect the baseline, digest, evaluator, and invocation
from agent changes. Capture refuses to overwrite an existing baseline. Results
contain only `caseId` and `guidanceReport`; a legacy `workspaceRoot`, if present,
must agree with the independently captured binding and never overrides it.
Missing baselines or digests, malformed reports, and escaped reference paths
fail scoring.
The comparison checks target identity, Git HEAD, refs and index, filesystem
content and stable metadata, including tracked, untracked, ignored files and
empty directories. Committing edits or making an empty commit does not evade
the check. It also compares the actual knowledge checkout and manifest identity.
Targets must have internal Git storage; linked target worktrees, submodules,
sparse checkouts, links/junctions/reparse points, hard links, and alternate data
streams are unsupported and rejected rather than silently excluded. A linked
**knowledge** checkout is supported with its Git storage identity recorded.
Use quiescent, isolated repositories; concurrent changes also fail the gate.
This is before/after evidence, not an OS sandbox or a complete write monitor.
It cannot prove that no transient write was reverted, that articles were opened,
or that constraints are semantically faithful. Reports and generated artifacts
must stay outside all target workspaces and the knowledge checkout. The
regression suite creates and removes its own uniquely named fixture directory;
it does not run against or clean a caller's target.
### External agent/runtime pilot (follow-up)
Consumer uptake and a real before-authoring pilot are not implemented by these
fixtures. The consumer must normalize its normal initial plan, persist guidance
after state initialization, inject it into existing phases, re-enrich on
material changes, and run an independent final review. See
[the integration boundary](../agent-consumption.md#repository-specific-development-orchestrators).
Before claiming improved repairs, run an independent pinned baseline without
enrichment and a matched enriched run. Hold starting code, task, model, tools,
runtime, and gates constant; record actual immutable BCQuality checkout and
policy identities rather than trusting a configured ref. Use that same
recorded checkout for enrichment and final review. Retain external logs,
article-read traces, resulting diffs, compile/test outcomes, and independent
review evidence, including failures, no-knowledge, partial, and unresolved
results. No compile/run or authoring-quality claim follows from the
credential-free checks above.

View file

@ -1,175 +0,0 @@
{
"version": 1,
"skill": "microsoft/skills/development/al-development-plan.md",
"minimumKnowledgeRecall": 1.0,
"minimumKnowledgePrecision": 0.67,
"cases": [
{
"id": "synthetic-normal-initial-plan",
"title": "Anonymized normal initial-plan consumer boundary",
"evidenceType": "integration-shaped-synthetic",
"boundary": "A synthetic consumer produces metadata plus a markdown plan body, serializes the full document, and passes that document as the generic existing development-plan. This is not an external real pilot and contains no consumer workflow-state schema.",
"expectedKind": "bug",
"expectedOutcome": "completed",
"expectedUnknown": [],
"requiresUnresolved": false,
"requiresMaterialUnresolved": false,
"development-plan": "{\"metadata\":{\"kind\":\"bug\",\"request\":\"Update every entry in the supplied filtered record set while preserving the caller's selection.\",\"origin\":\"anonymized synthetic initial plan\"},\"body\":\"## Root cause and design\\nThe routine reads and updates only the first record rather than iterating the supplied filtered set. Preserve the supplied filters and update each selected row.\\n\\n## Proposed fix\\nUse an update-safe FindSet/Next loop with an explicit update on each selected entry.\\n\\n## Affected files\\n- src/Batch/UpdateSelectedEntries.Codeunit.al\\n- test/Batch/UpdateSelectedEntriesTests.Codeunit.al\\n\\n## Test strategy\\nUse existing AL test library codeunits to arrange three selected rows and an excluded row. Assert all selected rows are updated and the excluded row is unchanged. This is a proposed test, not a reported result.\\n\\n## Acceptance criteria\\n- Every selected entry is updated exactly once.\\n- The supplied filters remain effective.\\n- No excluded entry changes.\\n- Empty selections cause no changes.\\n\"}",
"context": {
"bc-version": "28",
"technologies": [
"al"
],
"countries": [
"w1"
],
"application-area": [
"all"
],
"unknown": []
},
"requiredKnowledge": [
"microsoft/knowledge/performance/pair-findset-with-next-loop.md",
"microsoft/knowledge/performance/findset-true-applies-updlock-on-read.md",
"microsoft/knowledge/testing/use-library-codeunits-for-test-fixtures.md"
],
"optionalKnowledge": [
"microsoft/knowledge/performance/pass-var-record-to-preserve-partial-load-enumerator.md"
]
},
{
"id": "versioned-upgrade-plan-guidance",
"title": "Select guidance for a versioned data upgrade",
"expectedKind": "upgrade",
"expectedOutcome": "completed",
"expectedUnknown": [],
"requiresUnresolved": false,
"requiresMaterialUnresolved": false,
"development-plan": {
"kind": "upgrade",
"request": "Migrate existing customer tier text values to a new enum field in an app upgrade.",
"root-cause": "The new schema needs an explicit, rerunnable migration for existing tenant data.",
"affected-files": [
"src/Upgrade/CustomerTierUpgrade.Codeunit.al",
"test/Upgrade/CustomerTierUpgradeTests.Codeunit.al"
],
"proposed-changes": [
"Add a tagged upgrade step that copies existing values without validation triggers.",
"Add upgrade tests from multiple historical data versions."
],
"test-strategy": "Run upgrade tests from two prior data versions and verify a second invocation makes no further changes.",
"acceptance-criteria": [
"Existing values are preserved.",
"The migration is rerunnable.",
"Fresh installation does not execute upgrade migration."
]
},
"context": {
"bc-version": "28",
"technologies": [
"al"
],
"countries": [
"w1"
],
"application-area": [
"all"
],
"unknown": []
},
"requiredKnowledge": [
"microsoft/knowledge/upgrade/use-upgrade-tags-not-version-checks.md",
"microsoft/knowledge/upgrade/check-only-triggers-do-not-migrate-data.md",
"microsoft/knowledge/upgrade/install-code-does-not-run-on-version-upgrade.md"
],
"optionalKnowledge": [
"microsoft/knowledge/upgrade/datatransfer-skips-triggers-and-subscribers.md",
"microsoft/knowledge/upgrade/appversion-meaning-depends-on-execution-context.md"
]
},
{
"id": "no-additional-knowledge",
"title": "Honest empty enrichment does not prohibit ordinary work",
"expectedKind": "maintenance",
"expectedOutcome": "no-knowledge",
"expectedUnknown": [],
"requiresUnresolved": false,
"requiresMaterialUnresolved": false,
"development-plan": {
"kind": "maintenance",
"request": "Correct a spelling error in an existing internal explanatory comment in an AL procedure. Do not change the explanation, executable code, UI captions, schema, diagnostic text, configuration or behavior.",
"affected-files": ["src/Batch/EntryProcessor.Codeunit.al"],
"proposed-changes": ["Replace the misspelled word in the existing comment; add no new advice."],
"test-strategy": "Inspect the diff to confirm that only the comment spelling changes.",
"acceptance-criteria": ["Only the intended comment spelling changes; executable AL remains identical."]
},
"context": {
"bc-version": "28",
"technologies": ["al"],
"countries": ["w1"],
"application-area": ["all"],
"unknown": []
},
"requiredKnowledge": [],
"optionalKnowledge": []
},
{
"id": "unknown-material-version",
"title": "Materially unresolved version-sensitive guidance stays partial",
"expectedKind": "feature",
"expectedOutcome": "partial",
"expectedUnknown": ["bc-version"],
"requiresUnresolved": true,
"requiresMaterialUnresolved": true,
"development-plan": {
"kind": "feature",
"request": "Add an expensive Sum FlowField as the source of a usually-hidden page control. The proposed design relies on visibility suppressing calculation.",
"affected-files": ["src/Pages/EntryOverview.Page.al"],
"proposed-changes": ["Bind the page control directly to the FlowField and set Visible to a conditional expression."],
"test-strategy": "Verify aggregate queries are not executed while the control is hidden.",
"acceptance-criteria": ["Hidden controls do not cause expensive aggregate queries."],
"unknown": ["The deployment BC version and visible-only calculation feature state cannot be established from this fixture. Do not invent either."]
},
"context": {
"bc-version": "unknown",
"technologies": ["al"],
"countries": ["w1"],
"application-area": ["all"],
"unknown": ["bc-version"]
},
"requiredKnowledge": [
"microsoft/knowledge/performance/hidden-flowfields-still-calculate-before-bc26-opt-in.md"
],
"optionalKnowledge": []
},
{
"id": "partial-plan-decision",
"title": "Known platform context does not resolve an incomplete plan decision",
"expectedKind": "refactor",
"expectedOutcome": "partial",
"expectedUnknown": [],
"requiresUnresolved": true,
"requiresMaterialUnresolved": true,
"development-plan": {
"kind": "refactor",
"request": "Refactor a record read helper currently using FindFirst followed by Next. The caller contract does not establish whether to return one row or enumerate the entire filtered set.",
"affected-files": ["src/Queries/EntryReader.Codeunit.al"],
"proposed-changes": ["Choose the read method consistent with the intended cardinality after that decision is clarified."],
"test-strategy": "Add cardinality assertions after the caller contract is decided.",
"acceptance-criteria": ["The method and enumeration agree with the clarified caller contract."],
"unknown": ["Single-record versus multi-record caller intent remains materially unresolved."]
},
"context": {
"bc-version": "28",
"technologies": ["al"],
"countries": ["w1"],
"application-area": ["all"],
"unknown": []
},
"requiredKnowledge": [
"microsoft/knowledge/performance/pair-findset-with-next-loop.md"
],
"optionalKnowledge": []
}
]
}