Keep plan enrichment internal and read-only pending consumer agreement and runtime pilot evidence. Move knowledge to its independent PR and remove the consumer-owned forensic evaluator. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 638b66d2-9f06-4f60-8781-808709e1485c |
||
|---|---|---|
| .. | ||
| README.md | ||
| review-fixtures.json | ||
AL review evaluation
The evaluation is convention-driven. The harness discovers every <layer>/skills/review/al-<domain>-review.md leaf across the enabled microsoft, community, and custom layers. Duplicate domains resolve with custom > community > microsoft precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both .bad.al and .good.al companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.
review-fixtures.json contains only global thresholds and optional exceptional overrides. An override may select a different article, add context when the generic convention cannot express a scenario, or use an articles array when one domain needs explicit regression coverage for several paired articles. Specify either article or articles, not both. The first selected article retains the stable <domain>-bad and <domain>-good manifest IDs; additional articles use slug-qualified IDs. Overrides should remain empty in the normal case.
Model-facing preparation hashes case IDs, neutralizes Good/Bad object-name tokens, and removes full-line sample comments so neither the article slug, domain, nor expected outcome reveals the answer.
Validate the corpus
pwsh ./tools/Test-ReviewFixtures.ps1 -Root .
This credential-free check proves every selected leaf maps to a same-named knowledge domain with at least one complete AL sample pair and that all configured overrides are valid.
Run a fast-model evaluation
-
Prepare neutral inputs:
pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -PrepareDirectory ./.evaluation-runThis is also the CI path. It derives all cases, builds the current index, requires the convention-selected article to rank naturally into the candidate cutoff, and prepares the neutral requests.
-
For a fast/small model, use one fresh invocation per
request-case-*.json. Each request embeds the exact leaf instructions, that domain's candidate index rows with authoritative paths, and one opaque case. The model opens only matching articles and copies finding IDs fromcandidateArticles[].path. Save each response with the matchingresult-case-*.jsonname in the same directory.
request-<domain>.json files provide optional leaf batches containing every selected case for that domain and identify the selected layer-owned skill path; save those as result-<domain>.json. A normal convention-selected domain has one bad/good pair, while an articles override contributes one pair per listed article. Directory scoring prefers result-case-*.json when present and otherwise falls back to result-*.json. review-request.json is an optional all-domains stress test for larger models. Neither batch form is the preferred fast-model profile.
-
Save only this result shape:
{ "cases": [ { "id": "case-a1b2c3d4", "findings": [ { "id": "microsoft/knowledge/appsource/permission-sets-cover-setup-and-usage-without-super.md" } ] } ] }Include every case. A clean control has an empty
findingsarray. -
Score all per-leaf results together:
pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -ResultsDirectory ./.evaluation-runFor a single combined stress-test result, use
-ResultsPathinstead.
The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.