Narrow AL development to read-only plan guidance

Retain shared knowledge enrichment and review guidance; defer standalone implementation and source-ingestion tracking. Add runner-owned baseline evidence, contract regressions, and explicit consumer/pilot boundaries.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This commit is contained in:
Jesper Schulz-Wedde 2026-09-07 11:47:42 +02:00
parent f6fca1d56d
commit 1deac52a53
26 changed files with 1486 additions and 11037 deletions

View file

@ -1,4 +1,4 @@
# AL review evaluation
# AL review and guidance evaluation
The evaluation is convention-driven. The harness discovers every `<layer>/skills/review/al-<domain>-review.md` leaf across the enabled `microsoft`, `community`, and `custom` layers. Duplicate domains resolve with `custom > community > microsoft` precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both `.bad.al` and `.good.al` companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.
@ -55,50 +55,116 @@ This credential-free check proves every selected leaf maps to a same-named knowl
The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.
## AL development evaluation
`development-fixtures.json` defines end-to-end development requests rather than
prewritten good/bad snippets. Each case declares its execution mode, the
Business Central capabilities it exercises, the knowledge that should shape
the implementation, acceptance criteria, and the real checks an external
runner must perform.
Validate the fixture and capability manifests and prepare opaque requests:
```powershell
pwsh ./tools/Test-DevelopmentFixtures.ps1 -Root . -PrepareDirectory ./.development-evaluation
```
Each request runs `al-development` in a fresh writable AL repository. Cases
may exercise feature, bug, refactor, upgrade, or maintenance mode.
The runner compiles the generated project, runs its tests, invokes the review
quality gate, and stores the resulting implementation report using the opaque
`caseId` from its request, for example `result-case-a1b2c3d4.json`. It keeps the
generated repository available at the wrapper's `workspaceRoot` so scoring can
verify reported changed paths. Score all results with:
```powershell
pwsh ./tools/Test-DevelopmentFixtures.ps1 -Root . -ResultsDirectory ./.development-evaluation
```
The initial fixtures cover setup-backed master data, document header/line
workflows, versioned API integrations, and surgical diagnosis and repair of a
batch-processing bug, plus a rerunnable data upgrade. The capability manifest
enforces a minimum fixture-backed coverage ratio; the broader roadmap lives in
`coverage/development-capabilities.json`.
### Read-only plan guidance
## Read-only plan guidance
`development-guidance-fixtures.json` evaluates the planning interface used by
specialized orchestrators. It supplies an existing development plan and expects
a referenced set of implementation constraints without any target-repository
changes.
existing workflows. It supplies an existing plan and expects referenced
constraints without target-repository changes. The initial-plan fixture is
anonymized and synthetic: a full document with metadata and a markdown body
covering root cause, proposed fix, affected files, tests, and acceptance
criteria. It is integration-shaped input, not private consumer content, a
continuation checkpoint, or proof that any production consumer is integrated.
### Credential-free contract and scorer coverage
CI validates the manifest, prepares opaque model requests, and runs
deterministic scorer regressions with controlled reports and temporary Git
repositories. These checks cover report shape, outcomes, reference paths, and
the evaluator's pre/post read-only comparison. They do **not** run an agent,
compile AL, run Business Central tests, or establish better code authoring.
Prepare requests in a runner-owned artifact directory outside every target
workspace:
```powershell
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . -PrepareDirectory ./.development-guidance-evaluation
$run = Join-Path ([IO.Path]::GetTempPath()) 'bcquality-guidance-run'
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . -PrepareDirectory $run
```
An external runner stores `result-<case-id>.json` beside the generated request
and retains the clean fixture repository at `workspaceRoot`. Score the result
with `-ResultsDirectory`; the scorer verifies knowledge recall and precision
and fails if the planning pass changed the repository.
Preparation is not a model run. A scorer can validate a citation's path and
required fields, but only external agent traces and expert evaluation can
establish that the article was opened and its normative constraints faithfully
applied. Expected knowledge recall/precision is fixture-specific, not a corpus
coverage or authoring capability percentage.
`no-knowledge` means no additional applicable BCQuality constraints, not unsafe
work. It requires empty `knowledge`. Partial evaluation, failed retrieval,
unknown context, and materially unresolved applicability must remain visible
and distinct; they cannot be counted as successful enrichment merely because
the JSON is parseable.
Run the deterministic regression suite without an agent or AL environment:
```powershell
pwsh ./tools/Test-DevelopmentGuidanceEvaluator.ps1
```
### Runner-owned read-only evidence
For an external guidance run, first provision a representative, standalone Git
repository for each manifest case. The runner supplies a JSON workspace map
whose keys are the manifest IDs (not the hashed model IDs) and whose values are
absolute workspace roots. It may pass `-WorkspaceMapPath` during preparation
to bind the generated requests to those roots. The model must not select its
own workspace for scoring.
Capture evidence **before** invoking the agent, with the manifest, workspace
map, and source checkout already finalized:
```powershell
# Runner-selected paths, all outside the targets and BCQuality checkout.
$map = Join-Path $evidenceDirectory 'workspace-map.json'
$baseline = Join-Path $evidenceDirectory 'baseline.json'
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
-CaptureBaseline -WorkspaceMapPath $map -BaselinePath $baseline
# Retain the printed SHA256 in runner-only state BEFORE agent invocation.
# After the external agent writes result-case-<hash>.json files:
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
-ResultsDirectory $resultsDirectory -BaselinePath $baseline `
-BaselineSha256 $preRunDigest
```
`$evidenceDirectory`, `$resultsDirectory`, and `$preRunDigest` are supplied by
the runner; the digest must not be recomputed from potentially modified evidence
after the agent runs. Protect the baseline, digest, evaluator, and invocation
from agent changes. Capture refuses to overwrite an existing baseline. Results
contain only `caseId` and `guidanceReport`; a legacy `workspaceRoot`, if present,
must agree with the independently captured binding and never overrides it.
Missing baselines or digests, malformed reports, and escaped reference paths
fail scoring.
The comparison checks target identity, Git HEAD, refs and index, filesystem
content and stable metadata, including tracked, untracked, ignored files and
empty directories. Committing edits or making an empty commit does not evade
the check. It also compares the actual knowledge checkout and manifest identity.
Targets must have internal Git storage; linked target worktrees, submodules,
sparse checkouts, links/junctions/reparse points, hard links, and alternate data
streams are unsupported and rejected rather than silently excluded. A linked
**knowledge** checkout is supported with its Git storage identity recorded.
Use quiescent, isolated repositories; concurrent changes also fail the gate.
This is before/after evidence, not an OS sandbox or a complete write monitor.
It cannot prove that no transient write was reverted, that articles were opened,
or that constraints are semantically faithful. Reports and generated artifacts
must stay outside all target workspaces and the knowledge checkout. The
regression suite creates and removes its own uniquely named fixture directory;
it does not run against or clean a caller's target.
### External agent/runtime pilot (follow-up)
Consumer uptake and a real before-authoring pilot are not implemented by these
fixtures. The consumer must normalize its normal initial plan, persist guidance
after state initialization, inject it into existing phases, re-enrich on
material changes, and run an independent final review. See
[the integration boundary](../agent-consumption.md#repository-specific-development-orchestrators).
Before claiming improved repairs, run an independent pinned baseline without
enrichment and a matched enriched run. Hold starting code, task, model, tools,
runtime, and gates constant; record actual immutable BCQuality checkout and
policy identities rather than trusting a configured ref. Use that same
recorded checkout for enrichment and final review. Retain external logs,
article-read traces, resulting diffs, compile/test outcomes, and independent
review evidence, including failures, no-knowledge, partial, and unresolved
results. No compile/run or authoring-quality claim follows from the
credential-free checks above.

View file

@ -1,237 +0,0 @@
{
"version": 1,
"skill": "microsoft/skills/development/al-development.md",
"maximumReviewRounds": 3,
"minimumKnowledgeRecall": 1.0,
"minimumKnowledgePrecision": 0.5,
"cases": [
{
"id": "setup-backed-master-data",
"title": "Build setup-backed loyalty member master data",
"capabilities": [
"setup-and-master-data"
],
"expectedKind": "feature",
"development-request": {
"kind": "feature",
"description": "Add a Loyalty Member feature to an existing Business Central AL app. Administrators configure the member number series on a singleton setup card. Users create members through list and card pages, the table assigns numbers, and blocked members cannot be selected by consuming records. Include least-privilege permission sets and automated tests.",
"acceptance-criteria": [
"The setup is a blank-key singleton surfaced by a Card page.",
"Member numbers use the current No. Series codeunit and support manual numbers according to setup.",
"Blocked validation occurs where a member is consumed, not only on the member table.",
"The feature includes assignable least-privilege permissions and automated tests."
]
},
"context": {
"technologies": [
"al"
],
"countries": [
"w1"
],
"application-area": [
"all"
]
},
"requiredKnowledge": [
"microsoft/knowledge/data-modeling/setup-table-is-a-singleton.md",
"microsoft/knowledge/data-modeling/master-table-no-from-number-series-in-oninsert.md",
"microsoft/knowledge/data-modeling/check-blocked-in-referencing-code-not-in-master.md",
"microsoft/knowledge/security/permission-set-avoid-wildcard-grants.md",
"microsoft/knowledge/testing/use-library-codeunits-for-test-fixtures.md"
],
"optionalKnowledge": [
"microsoft/knowledge/data-modeling/use-no-series-codeunit-not-noseriesmanagement.md",
"microsoft/knowledge/security/compose-permission-sets-with-included-sets.md",
"microsoft/knowledge/style/applicationarea-required-on-page-controls.md",
"microsoft/knowledge/style/tooltip-required-on-page-fields.md",
"microsoft/knowledge/ui/showmandatory-on-code-required-page-fields.md"
],
"requiredChecks": [
"compile",
"tests",
"review"
]
},
{
"id": "document-header-and-lines",
"title": "Build a document header and lines workflow",
"capabilities": [
"document-workflows"
],
"expectedKind": "feature",
"development-request": {
"kind": "feature",
"description": "Implement a Service Quote feature with a header, lines, document page, number series, posting and document dates, calculated totals, and tests. Line edits must refresh the total shown on the header. Structure initialization so API, UI, and test creation paths behave consistently.",
"acceptance-criteria": [
"The header assigns its number before InitRecord establishes document defaults.",
"The document page links lines correctly and refreshes parent totals after edits.",
"Tests exercise creation outside the UI as well as the document-page behavior.",
"The implementation contains no obsolete NoSeriesManagement dependency."
]
},
"context": {
"technologies": [
"al"
],
"countries": [
"w1"
],
"application-area": [
"service"
]
},
"requiredKnowledge": [
"microsoft/knowledge/data-modeling/initialize-document-defaults-in-initrecord.md",
"microsoft/knowledge/data-modeling/use-no-series-codeunit-not-noseriesmanagement.md",
"microsoft/knowledge/ui/updatepropagation-both-refreshes-main-page.md",
"microsoft/knowledge/testing/use-library-codeunits-for-test-fixtures.md"
],
"optionalKnowledge": [
"microsoft/knowledge/events/publish-thin-onbefore-onafter-integration-events.md",
"microsoft/knowledge/style/applicationarea-required-on-page-controls.md",
"microsoft/knowledge/style/tooltip-required-on-page-fields.md"
],
"requiredChecks": [
"compile",
"tests",
"review"
]
},
{
"id": "versioned-master-data-api",
"title": "Expose master data through a versioned API",
"capabilities": [
"api-integrations"
],
"expectedKind": "feature",
"development-request": {
"kind": "auto",
"plan": "Expose Loyalty Member master data through a Business Central API page. Use a stable v1.0 contract, address records by SystemId, support insert and update, use conventional entity naming, and include permissions and automated API-oriented tests.",
"acceptance-criteria": [
"The API declares all routing properties and a stable APIVersion.",
"ODataKeyFields uses SystemId and the exposed SystemId field is not editable.",
"EntityName is singular, EntitySetName is plural, and both are lower camel case.",
"The API is covered by least-privilege permissions and automated tests."
]
},
"context": {
"technologies": [
"al"
],
"countries": [
"w1"
],
"application-area": [
"all"
]
},
"requiredKnowledge": [
"microsoft/knowledge/web-services/set-required-api-page-properties.md",
"microsoft/knowledge/web-services/expose-systemid-as-the-api-key.md",
"microsoft/knowledge/style/api-page-delayedinsert-true.md",
"microsoft/knowledge/style/api-page-entity-naming-singular-plural.md",
"microsoft/knowledge/security/permission-set-avoid-wildcard-grants.md"
],
"optionalKnowledge": [
"microsoft/knowledge/style/api-page-camelcase-properties.md",
"microsoft/knowledge/web-services/version-apis-by-adding-not-mutating-published-versions.md"
],
"requiredChecks": [
"compile",
"tests",
"review"
]
},
{
"id": "fix-filtered-batch-processing",
"title": "Fix a batch routine that processes only one record",
"capabilities": [
"bug-diagnosis-and-fix"
],
"expectedKind": "bug",
"development-request": {
"kind": "bug",
"description": "Users report that an existing filtered batch routine updates only the first matching record. Reproduce the defect, identify why iteration stops, make the smallest safe correction, and add a regression test that selects multiple records and proves every selected record is processed.",
"acceptance-criteria": [
"The defect is reproduced or demonstrated by a failing regression test before the fix.",
"The root cause is corrected without widening the supplied record filters.",
"The routine processes every selected record with the appropriate update locking behavior.",
"A regression test covers more than one selected record."
]
},
"context": {
"technologies": [
"al"
],
"countries": [
"w1"
],
"application-area": [
"all"
]
},
"requiredKnowledge": [
"microsoft/knowledge/performance/pair-findset-with-next-loop.md",
"microsoft/knowledge/testing/use-library-codeunits-for-test-fixtures.md"
],
"optionalKnowledge": [
"microsoft/knowledge/performance/findset-true-applies-updlock-on-read.md",
"microsoft/knowledge/performance/pass-var-record-to-preserve-partial-load-enumerator.md"
],
"requiredChecks": [
"compile",
"tests",
"review"
]
},
{
"id": "versioned-data-upgrade",
"title": "Implement a rerunnable data upgrade",
"capabilities": [
"install-and-upgrade"
],
"expectedKind": "upgrade",
"development-request": {
"kind": "upgrade",
"description": "Upgrade an existing app from a text-based Customer Tier field to a new enum-backed Tier field. Preserve existing customer data, support tenants that skip intermediate app versions, keep fresh installation separate from migration, and add upgrade tests.",
"acceptance-criteria": [
"Existing tier values are migrated without running field validation triggers.",
"The migration is guarded by an upgrade tag and is safe when the upgrade runs again.",
"Check and validation triggers do not write data.",
"Fresh installation does not run version-upgrade migration.",
"Tests cover more than one historical source data version."
]
},
"context": {
"technologies": [
"al"
],
"countries": [
"w1"
],
"application-area": [
"all"
]
},
"requiredKnowledge": [
"microsoft/knowledge/upgrade/appversion-meaning-depends-on-execution-context.md",
"microsoft/knowledge/upgrade/install-and-upgrade-codeunits-have-no-order.md",
"microsoft/knowledge/upgrade/use-upgrade-tags-not-version-checks.md",
"microsoft/knowledge/upgrade/check-only-triggers-do-not-migrate-data.md",
"microsoft/knowledge/upgrade/install-code-does-not-run-on-version-upgrade.md"
],
"optionalKnowledge": [
"microsoft/knowledge/upgrade/datatransfer-skips-triggers-and-subscribers.md",
"microsoft/knowledge/upgrade/datatransfer-for-bulk-init.md",
"microsoft/knowledge/upgrade/register-upgrade-tags-with-subscribers.md",
"microsoft/knowledge/testing/transactionmodel-attribute-governs-test-transactions.md"
],
"requiredChecks": [
"compile",
"tests",
"review"
]
}
]
}

View file

@ -5,36 +5,18 @@
"minimumKnowledgePrecision": 0.67,
"cases": [
{
"id": "bcapps-filtered-batch-bug-plan",
"title": "Select guidance from a BCFIX-HANDOFF v1 payload",
"id": "synthetic-normal-initial-plan",
"title": "Anonymized normal initial-plan consumer boundary",
"evidenceType": "integration-shaped-synthetic",
"boundary": "A synthetic consumer produces metadata plus a markdown plan body, serializes the full document, and passes that document as the generic existing development-plan. This is not an external real pilot and contains no consumer workflow-state schema.",
"expectedKind": "bug",
"development-plan": {
"format": "BCFIX-HANDOFF",
"version": 1,
"issue": 4312,
"phase": "implement",
"status": "paused",
"baton": 2,
"rootCause": "The routine calls FindFirst and updates the current record without entering an enumerator loop, so only the first record in the supplied filtered set is modified.",
"harnessMap": {
"testCodeunit": "Update Selected Entries Tests",
"libraries": [
"Library - Random"
],
"pages": [],
"handlers": []
},
"iterationsUsed": 1,
"filesCommitted": [
"test/Batch/UpdateSelectedEntries.Codeunit.al"
],
"lastTestResult": "1 failing, 4 passing; the red test shows only the first of three selected records is updated.",
"deadEnds": [
"Changing the page selection did not help because the codeunit discarded the supplied enumerator."
],
"nextStep": "Replace the single-record read with an update-safe FindSet/Next loop that preserves the supplied filters, then rerun the red test."
},
"expectedOutcome": "completed",
"expectedUnknown": [],
"requiresUnresolved": false,
"requiresMaterialUnresolved": false,
"development-plan": "{\"metadata\":{\"kind\":\"bug\",\"request\":\"Update every entry in the supplied filtered record set while preserving the caller's selection.\",\"origin\":\"anonymized synthetic initial plan\"},\"body\":\"## Root cause and design\\nThe routine reads and updates only the first record rather than iterating the supplied filtered set. Preserve the supplied filters and update each selected row.\\n\\n## Proposed fix\\nUse an update-safe FindSet/Next loop with an explicit update on each selected entry.\\n\\n## Affected files\\n- src/Batch/UpdateSelectedEntries.Codeunit.al\\n- test/Batch/UpdateSelectedEntriesTests.Codeunit.al\\n\\n## Test strategy\\nUse existing AL test library codeunits to arrange three selected rows and an excluded row. Assert all selected rows are updated and the excluded row is unchanged. This is a proposed test, not a reported result.\\n\\n## Acceptance criteria\\n- Every selected entry is updated exactly once.\\n- The supplied filters remain effective.\\n- No excluded entry changes.\\n- Empty selections cause no changes.\\n\"}",
"context": {
"bc-version": "28",
"technologies": [
"al"
],
@ -43,7 +25,8 @@
],
"application-area": [
"all"
]
],
"unknown": []
},
"requiredKnowledge": [
"microsoft/knowledge/performance/pair-findset-with-next-loop.md",
@ -58,6 +41,10 @@
"id": "versioned-upgrade-plan-guidance",
"title": "Select guidance for a versioned data upgrade",
"expectedKind": "upgrade",
"expectedOutcome": "completed",
"expectedUnknown": [],
"requiresUnresolved": false,
"requiresMaterialUnresolved": false,
"development-plan": {
"kind": "upgrade",
"request": "Migrate existing customer tier text values to a new enum field in an app upgrade.",
@ -78,6 +65,7 @@
]
},
"context": {
"bc-version": "28",
"technologies": [
"al"
],
@ -86,7 +74,8 @@
],
"application-area": [
"all"
]
],
"unknown": []
},
"requiredKnowledge": [
"microsoft/knowledge/upgrade/use-upgrade-tags-not-version-checks.md",
@ -97,6 +86,90 @@
"microsoft/knowledge/upgrade/datatransfer-skips-triggers-and-subscribers.md",
"microsoft/knowledge/upgrade/appversion-meaning-depends-on-execution-context.md"
]
},
{
"id": "no-additional-knowledge",
"title": "Honest empty enrichment does not prohibit ordinary work",
"expectedKind": "maintenance",
"expectedOutcome": "no-knowledge",
"expectedUnknown": [],
"requiresUnresolved": false,
"requiresMaterialUnresolved": false,
"development-plan": {
"kind": "maintenance",
"request": "Correct a spelling error in an existing internal explanatory comment in an AL procedure. Do not change the explanation, executable code, UI captions, schema, diagnostic text, configuration or behavior.",
"affected-files": ["src/Batch/EntryProcessor.Codeunit.al"],
"proposed-changes": ["Replace the misspelled word in the existing comment; add no new advice."],
"test-strategy": "Inspect the diff to confirm that only the comment spelling changes.",
"acceptance-criteria": ["Only the intended comment spelling changes; executable AL remains identical."]
},
"context": {
"bc-version": "28",
"technologies": ["al"],
"countries": ["w1"],
"application-area": ["all"],
"unknown": []
},
"requiredKnowledge": [],
"optionalKnowledge": []
},
{
"id": "unknown-material-version",
"title": "Materially unresolved version-sensitive guidance stays partial",
"expectedKind": "feature",
"expectedOutcome": "partial",
"expectedUnknown": ["bc-version"],
"requiresUnresolved": true,
"requiresMaterialUnresolved": true,
"development-plan": {
"kind": "feature",
"request": "Add an expensive Sum FlowField as the source of a usually-hidden page control. The proposed design relies on visibility suppressing calculation.",
"affected-files": ["src/Pages/EntryOverview.Page.al"],
"proposed-changes": ["Bind the page control directly to the FlowField and set Visible to a conditional expression."],
"test-strategy": "Verify aggregate queries are not executed while the control is hidden.",
"acceptance-criteria": ["Hidden controls do not cause expensive aggregate queries."],
"unknown": ["The deployment BC version and visible-only calculation feature state cannot be established from this fixture. Do not invent either."]
},
"context": {
"bc-version": "unknown",
"technologies": ["al"],
"countries": ["w1"],
"application-area": ["all"],
"unknown": ["bc-version"]
},
"requiredKnowledge": [
"microsoft/knowledge/performance/hidden-flowfields-still-calculate-before-bc26-opt-in.md"
],
"optionalKnowledge": []
},
{
"id": "partial-plan-decision",
"title": "Known platform context does not resolve an incomplete plan decision",
"expectedKind": "refactor",
"expectedOutcome": "partial",
"expectedUnknown": [],
"requiresUnresolved": true,
"requiresMaterialUnresolved": true,
"development-plan": {
"kind": "refactor",
"request": "Refactor a record read helper currently using FindFirst followed by Next. The caller contract does not establish whether to return one row or enumerate the entire filtered set.",
"affected-files": ["src/Queries/EntryReader.Codeunit.al"],
"proposed-changes": ["Choose the read method consistent with the intended cardinality after that decision is clarified."],
"test-strategy": "Add cardinality assertions after the caller contract is decided.",
"acceptance-criteria": ["The method and enumeration agree with the clarified caller contract."],
"unknown": ["Single-record versus multi-record caller intent remains materially unresolved."]
},
"context": {
"bc-version": "28",
"technologies": ["al"],
"countries": ["w1"],
"application-area": ["all"],
"unknown": []
},
"requiredKnowledge": [
"microsoft/knowledge/performance/pair-findset-with-next-loop.md"
],
"optionalKnowledge": []
}
]
}