mirror of
https://github.com/microsoft/BCQuality.git
synced 2026-10-05 14:46:55 +01:00
Narrow development guidance to provisional contract
Keep plan enrichment internal and read-only pending consumer agreement and runtime pilot evidence. Move knowledge to its independent PR and remove the consumer-owned forensic evaluator. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 638b66d2-9f06-4f60-8781-808709e1485c
This commit is contained in:
parent
8f025ac679
commit
fa7eb750c6
40 changed files with 294 additions and 2053 deletions
|
|
@ -1,4 +1,4 @@
|
|||
# AL review and guidance evaluation
|
||||
# AL review evaluation
|
||||
|
||||
The evaluation is convention-driven. The harness discovers every `<layer>/skills/review/al-<domain>-review.md` leaf across the enabled `microsoft`, `community`, and `custom` layers. Duplicate domains resolve with `custom > community > microsoft` precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both `.bad.al` and `.good.al` companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.
|
||||
|
||||
|
|
@ -54,117 +54,3 @@ This credential-free check proves every selected leaf maps to a same-named knowl
|
|||
For a single combined stress-test result, use `-ResultsPath` instead.
|
||||
|
||||
The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.
|
||||
|
||||
## Read-only plan guidance
|
||||
|
||||
`development-guidance-fixtures.json` evaluates the planning interface used by
|
||||
existing workflows. It supplies an existing plan and expects referenced
|
||||
constraints without target-repository changes. The initial-plan fixture is
|
||||
anonymized and synthetic: a full document with metadata and a markdown body
|
||||
covering root cause, proposed fix, affected files, tests, and acceptance
|
||||
criteria. It is integration-shaped input, not private consumer content, a
|
||||
continuation checkpoint, or proof that any production consumer is integrated.
|
||||
|
||||
### Credential-free contract and scorer coverage
|
||||
|
||||
CI validates the manifest, prepares opaque model requests, and runs
|
||||
deterministic scorer regressions with controlled reports and temporary Git
|
||||
repositories. These checks cover report shape, outcomes, reference paths, and
|
||||
the evaluator's pre/post read-only comparison. They do **not** run an agent,
|
||||
compile AL, run Business Central tests, or establish better code authoring.
|
||||
|
||||
Prepare requests in a runner-owned artifact directory outside every target
|
||||
workspace:
|
||||
|
||||
```powershell
|
||||
$run = Join-Path ([IO.Path]::GetTempPath()) 'bcquality-guidance-run'
|
||||
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . -PrepareDirectory $run
|
||||
```
|
||||
|
||||
Preparation is not a model run. A scorer can validate a citation's path and
|
||||
required fields, but only external agent traces and expert evaluation can
|
||||
establish that the article was opened and its normative constraints faithfully
|
||||
applied. Expected knowledge recall/precision is fixture-specific, not a corpus
|
||||
coverage or authoring capability percentage.
|
||||
|
||||
`no-knowledge` means no additional applicable BCQuality constraints, not unsafe
|
||||
work. It requires empty `knowledge`. Partial evaluation, failed retrieval,
|
||||
unknown context, and materially unresolved applicability must remain visible
|
||||
and distinct; they cannot be counted as successful enrichment merely because
|
||||
the JSON is parseable.
|
||||
|
||||
Run the deterministic regression suite without an agent or AL environment:
|
||||
|
||||
```powershell
|
||||
pwsh ./tools/Test-DevelopmentGuidanceEvaluator.ps1
|
||||
```
|
||||
|
||||
### Runner-owned read-only evidence
|
||||
|
||||
For an external guidance run, first provision a representative, standalone Git
|
||||
repository for each manifest case. The runner supplies a JSON workspace map
|
||||
whose keys are the manifest IDs (not the hashed model IDs) and whose values are
|
||||
absolute workspace roots. It may pass `-WorkspaceMapPath` during preparation
|
||||
to bind the generated requests to those roots. The model must not select its
|
||||
own workspace for scoring.
|
||||
|
||||
Capture evidence **before** invoking the agent, with the manifest, workspace
|
||||
map, and source checkout already finalized:
|
||||
|
||||
```powershell
|
||||
# Runner-selected paths, all outside the targets and BCQuality checkout.
|
||||
$map = Join-Path $evidenceDirectory 'workspace-map.json'
|
||||
$baseline = Join-Path $evidenceDirectory 'baseline.json'
|
||||
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
|
||||
-CaptureBaseline -WorkspaceMapPath $map -BaselinePath $baseline
|
||||
|
||||
# Retain the printed SHA256 in runner-only state BEFORE agent invocation.
|
||||
# After the external agent writes result-case-<hash>.json files:
|
||||
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
|
||||
-ResultsDirectory $resultsDirectory -BaselinePath $baseline `
|
||||
-BaselineSha256 $preRunDigest
|
||||
```
|
||||
|
||||
`$evidenceDirectory`, `$resultsDirectory`, and `$preRunDigest` are supplied by
|
||||
the runner; the digest must not be recomputed from potentially modified evidence
|
||||
after the agent runs. Protect the baseline, digest, evaluator, and invocation
|
||||
from agent changes. Capture refuses to overwrite an existing baseline. Results
|
||||
contain only `caseId` and `guidanceReport`; a legacy `workspaceRoot`, if present,
|
||||
must agree with the independently captured binding and never overrides it.
|
||||
Missing baselines or digests, malformed reports, and escaped reference paths
|
||||
fail scoring.
|
||||
|
||||
The comparison checks target identity, Git HEAD, refs and index, filesystem
|
||||
content and stable metadata, including tracked, untracked, ignored files and
|
||||
empty directories. Committing edits or making an empty commit does not evade
|
||||
the check. It also compares the actual knowledge checkout and manifest identity.
|
||||
Targets must have internal Git storage; linked target worktrees, submodules,
|
||||
sparse checkouts, links/junctions/reparse points, hard links, and alternate data
|
||||
streams are unsupported and rejected rather than silently excluded. A linked
|
||||
**knowledge** checkout is supported with its Git storage identity recorded.
|
||||
Use quiescent, isolated repositories; concurrent changes also fail the gate.
|
||||
|
||||
This is before/after evidence, not an OS sandbox or a complete write monitor.
|
||||
It cannot prove that no transient write was reverted, that articles were opened,
|
||||
or that constraints are semantically faithful. Reports and generated artifacts
|
||||
must stay outside all target workspaces and the knowledge checkout. The
|
||||
regression suite creates and removes its own uniquely named fixture directory;
|
||||
it does not run against or clean a caller's target.
|
||||
|
||||
### External agent/runtime pilot (follow-up)
|
||||
|
||||
Consumer uptake and a real before-authoring pilot are not implemented by these
|
||||
fixtures. The consumer must normalize its normal initial plan, persist guidance
|
||||
after state initialization, inject it into existing phases, re-enrich on
|
||||
material changes, and run an independent final review. See
|
||||
[the integration boundary](../agent-consumption.md#repository-specific-development-orchestrators).
|
||||
|
||||
Before claiming improved repairs, run an independent pinned baseline without
|
||||
enrichment and a matched enriched run. Hold starting code, task, model, tools,
|
||||
runtime, and gates constant; record actual immutable BCQuality checkout and
|
||||
policy identities rather than trusting a configured ref. Use that same
|
||||
recorded checkout for enrichment and final review. Retain external logs,
|
||||
article-read traces, resulting diffs, compile/test outcomes, and independent
|
||||
review evidence, including failures, no-knowledge, partial, and unresolved
|
||||
results. No compile/run or authoring-quality claim follows from the
|
||||
credential-free checks above.
|
||||
|
|
|
|||
|
|
@ -1,175 +0,0 @@
|
|||
{
|
||||
"version": 1,
|
||||
"skill": "microsoft/skills/development/al-development-plan.md",
|
||||
"minimumKnowledgeRecall": 1.0,
|
||||
"minimumKnowledgePrecision": 0.67,
|
||||
"cases": [
|
||||
{
|
||||
"id": "synthetic-normal-initial-plan",
|
||||
"title": "Anonymized normal initial-plan consumer boundary",
|
||||
"evidenceType": "integration-shaped-synthetic",
|
||||
"boundary": "A synthetic consumer produces metadata plus a markdown plan body, serializes the full document, and passes that document as the generic existing development-plan. This is not an external real pilot and contains no consumer workflow-state schema.",
|
||||
"expectedKind": "bug",
|
||||
"expectedOutcome": "completed",
|
||||
"expectedUnknown": [],
|
||||
"requiresUnresolved": false,
|
||||
"requiresMaterialUnresolved": false,
|
||||
"development-plan": "{\"metadata\":{\"kind\":\"bug\",\"request\":\"Update every entry in the supplied filtered record set while preserving the caller's selection.\",\"origin\":\"anonymized synthetic initial plan\"},\"body\":\"## Root cause and design\\nThe routine reads and updates only the first record rather than iterating the supplied filtered set. Preserve the supplied filters and update each selected row.\\n\\n## Proposed fix\\nUse an update-safe FindSet/Next loop with an explicit update on each selected entry.\\n\\n## Affected files\\n- src/Batch/UpdateSelectedEntries.Codeunit.al\\n- test/Batch/UpdateSelectedEntriesTests.Codeunit.al\\n\\n## Test strategy\\nUse existing AL test library codeunits to arrange three selected rows and an excluded row. Assert all selected rows are updated and the excluded row is unchanged. This is a proposed test, not a reported result.\\n\\n## Acceptance criteria\\n- Every selected entry is updated exactly once.\\n- The supplied filters remain effective.\\n- No excluded entry changes.\\n- Empty selections cause no changes.\\n\"}",
|
||||
"context": {
|
||||
"bc-version": "28",
|
||||
"technologies": [
|
||||
"al"
|
||||
],
|
||||
"countries": [
|
||||
"w1"
|
||||
],
|
||||
"application-area": [
|
||||
"all"
|
||||
],
|
||||
"unknown": []
|
||||
},
|
||||
"requiredKnowledge": [
|
||||
"microsoft/knowledge/performance/pair-findset-with-next-loop.md",
|
||||
"microsoft/knowledge/performance/findset-true-applies-updlock-on-read.md",
|
||||
"microsoft/knowledge/testing/use-library-codeunits-for-test-fixtures.md"
|
||||
],
|
||||
"optionalKnowledge": [
|
||||
"microsoft/knowledge/performance/pass-var-record-to-preserve-partial-load-enumerator.md"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "versioned-upgrade-plan-guidance",
|
||||
"title": "Select guidance for a versioned data upgrade",
|
||||
"expectedKind": "upgrade",
|
||||
"expectedOutcome": "completed",
|
||||
"expectedUnknown": [],
|
||||
"requiresUnresolved": false,
|
||||
"requiresMaterialUnresolved": false,
|
||||
"development-plan": {
|
||||
"kind": "upgrade",
|
||||
"request": "Migrate existing customer tier text values to a new enum field in an app upgrade.",
|
||||
"root-cause": "The new schema needs an explicit, rerunnable migration for existing tenant data.",
|
||||
"affected-files": [
|
||||
"src/Upgrade/CustomerTierUpgrade.Codeunit.al",
|
||||
"test/Upgrade/CustomerTierUpgradeTests.Codeunit.al"
|
||||
],
|
||||
"proposed-changes": [
|
||||
"Add a tagged upgrade step that copies existing values without validation triggers.",
|
||||
"Add upgrade tests from multiple historical data versions."
|
||||
],
|
||||
"test-strategy": "Run upgrade tests from two prior data versions and verify a second invocation makes no further changes.",
|
||||
"acceptance-criteria": [
|
||||
"Existing values are preserved.",
|
||||
"The migration is rerunnable.",
|
||||
"Fresh installation does not execute upgrade migration."
|
||||
]
|
||||
},
|
||||
"context": {
|
||||
"bc-version": "28",
|
||||
"technologies": [
|
||||
"al"
|
||||
],
|
||||
"countries": [
|
||||
"w1"
|
||||
],
|
||||
"application-area": [
|
||||
"all"
|
||||
],
|
||||
"unknown": []
|
||||
},
|
||||
"requiredKnowledge": [
|
||||
"microsoft/knowledge/upgrade/use-upgrade-tags-not-version-checks.md",
|
||||
"microsoft/knowledge/upgrade/check-only-triggers-do-not-migrate-data.md",
|
||||
"microsoft/knowledge/upgrade/install-code-does-not-run-on-version-upgrade.md"
|
||||
],
|
||||
"optionalKnowledge": [
|
||||
"microsoft/knowledge/upgrade/datatransfer-skips-triggers-and-subscribers.md",
|
||||
"microsoft/knowledge/upgrade/appversion-meaning-depends-on-execution-context.md"
|
||||
]
|
||||
},
|
||||
{
|
||||
"id": "no-additional-knowledge",
|
||||
"title": "Honest empty enrichment does not prohibit ordinary work",
|
||||
"expectedKind": "maintenance",
|
||||
"expectedOutcome": "no-knowledge",
|
||||
"expectedUnknown": [],
|
||||
"requiresUnresolved": false,
|
||||
"requiresMaterialUnresolved": false,
|
||||
"development-plan": {
|
||||
"kind": "maintenance",
|
||||
"request": "Correct a spelling error in an existing internal explanatory comment in an AL procedure. Do not change the explanation, executable code, UI captions, schema, diagnostic text, configuration or behavior.",
|
||||
"affected-files": ["src/Batch/EntryProcessor.Codeunit.al"],
|
||||
"proposed-changes": ["Replace the misspelled word in the existing comment; add no new advice."],
|
||||
"test-strategy": "Inspect the diff to confirm that only the comment spelling changes.",
|
||||
"acceptance-criteria": ["Only the intended comment spelling changes; executable AL remains identical."]
|
||||
},
|
||||
"context": {
|
||||
"bc-version": "28",
|
||||
"technologies": ["al"],
|
||||
"countries": ["w1"],
|
||||
"application-area": ["all"],
|
||||
"unknown": []
|
||||
},
|
||||
"requiredKnowledge": [],
|
||||
"optionalKnowledge": []
|
||||
},
|
||||
{
|
||||
"id": "unknown-material-version",
|
||||
"title": "Materially unresolved version-sensitive guidance stays partial",
|
||||
"expectedKind": "feature",
|
||||
"expectedOutcome": "partial",
|
||||
"expectedUnknown": ["bc-version"],
|
||||
"requiresUnresolved": true,
|
||||
"requiresMaterialUnresolved": true,
|
||||
"development-plan": {
|
||||
"kind": "feature",
|
||||
"request": "Add an expensive Sum FlowField as the source of a usually-hidden page control. The proposed design relies on visibility suppressing calculation.",
|
||||
"affected-files": ["src/Pages/EntryOverview.Page.al"],
|
||||
"proposed-changes": ["Bind the page control directly to the FlowField and set Visible to a conditional expression."],
|
||||
"test-strategy": "Verify aggregate queries are not executed while the control is hidden.",
|
||||
"acceptance-criteria": ["Hidden controls do not cause expensive aggregate queries."],
|
||||
"unknown": ["The deployment BC version and visible-only calculation feature state cannot be established from this fixture. Do not invent either."]
|
||||
},
|
||||
"context": {
|
||||
"bc-version": "unknown",
|
||||
"technologies": ["al"],
|
||||
"countries": ["w1"],
|
||||
"application-area": ["all"],
|
||||
"unknown": ["bc-version"]
|
||||
},
|
||||
"requiredKnowledge": [
|
||||
"microsoft/knowledge/performance/hidden-flowfields-still-calculate-before-bc26-opt-in.md"
|
||||
],
|
||||
"optionalKnowledge": []
|
||||
},
|
||||
{
|
||||
"id": "partial-plan-decision",
|
||||
"title": "Known platform context does not resolve an incomplete plan decision",
|
||||
"expectedKind": "refactor",
|
||||
"expectedOutcome": "partial",
|
||||
"expectedUnknown": [],
|
||||
"requiresUnresolved": true,
|
||||
"requiresMaterialUnresolved": true,
|
||||
"development-plan": {
|
||||
"kind": "refactor",
|
||||
"request": "Refactor a record read helper currently using FindFirst followed by Next. The caller contract does not establish whether to return one row or enumerate the entire filtered set.",
|
||||
"affected-files": ["src/Queries/EntryReader.Codeunit.al"],
|
||||
"proposed-changes": ["Choose the read method consistent with the intended cardinality after that decision is clarified."],
|
||||
"test-strategy": "Add cardinality assertions after the caller contract is decided.",
|
||||
"acceptance-criteria": ["The method and enumeration agree with the clarified caller contract."],
|
||||
"unknown": ["Single-record versus multi-record caller intent remains materially unresolved."]
|
||||
},
|
||||
"context": {
|
||||
"bc-version": "28",
|
||||
"technologies": ["al"],
|
||||
"countries": ["w1"],
|
||||
"application-area": ["all"],
|
||||
"unknown": []
|
||||
},
|
||||
"requiredKnowledge": [
|
||||
"microsoft/knowledge/performance/pair-findset-with-next-loop.md"
|
||||
],
|
||||
"optionalKnowledge": []
|
||||
}
|
||||
]
|
||||
}
|
||||
Loading…
Add table
Add a link
Reference in a new issue