The article was rewritten to say a reset is required only when the value can carry over, and its H1 was updated to match, but three artefacts still carried the old "always initialize to false" premise: - The slug still read `initialize-ishandled-to-false-before-publishing`, which contradicts the body. The slug is not cosmetic: Build-KnowledgeIndex.ps1 ranks candidates on keywords, frontmatter dimensions, domain, path and title, so a stale path pushes selection back toward the behaviour this change narrows. Renamed to `reset-ishandled-only-when-the-value-can-carry-over`, following the existing precedent for conditional slugs such as `unreleased-symbol-change-is-not-a-breaking-change`. - Keywords still listed `initialization` and `deterministic` and omitted `false-positive`, the tag this repository uses for suppression articles. Replaced with `carry-over` and `loop-iteration` and added `false-positive`. - The good sample demonstrated only the "prefer separate fresh locals" clause and contained no reset at all, so the article's headline case had no positive example. It was also asymmetric with the bad sample, which gained a loop procedure showing a local that carries `true` into the next iteration. Added the matching loop procedure to the good sample: a local declared outside the loop is reset at the top of each iteration. That case cannot be solved by introducing another local, because AL has no block scope, so it is the only shape that demonstrates the reset the article still requires. It also gives the engine the correct `suggested-code` shape for the loop finding; without it the one-click fix adapted from the good sample would propose splitting the variable rather than adding one line. Also renamed the sample codeunits from "IsHandled Init ..." to "IsHandled Carry Over ...", and updated the two references to the old slug: the events leaf skill cue and the events pin in evaluation/review-fixtures.json. validate_frontmatter.py reports 0 errors; Test-ReviewFixtures.ps1 passes with 32 cases across 16 leaf domains and resolves the events fixture to the renamed article. |
||
|---|---|---|
| .. | ||
| README.md | ||
| review-fixtures.json | ||
AL review evaluation
The evaluation is convention-driven. For every microsoft/skills/review/al-<domain>-review.md leaf, the harness finds microsoft/knowledge/<domain>/, selects the first article (by filename) with both .bad.al and .good.al companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.
review-fixtures.json contains only global thresholds and optional exceptional overrides. An override may select a different article or add context when the generic convention cannot express a scenario. It should remain empty in the normal case.
Model-facing preparation hashes case IDs, neutralizes Good/Bad object-name tokens, and removes full-line sample comments so neither the article slug, domain, nor expected outcome reveals the answer.
Validate the corpus
pwsh ./tools/Test-ReviewFixtures.ps1 -Root .
This credential-free check proves every registered leaf maps to a same-named knowledge domain with at least one complete AL sample pair and that all configured overrides are valid.
Run a fast-model evaluation
-
Prepare neutral inputs:
pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -PrepareDirectory ./.evaluation-runThis is also the CI path. It derives all cases, builds the current index, requires the convention-selected article to rank naturally into the candidate cutoff, and prepares the neutral requests.
-
For a fast/small model, use one fresh invocation per
request-case-*.json. Each request embeds the exact leaf instructions, that domain's candidate index rows with authoritative paths, and one opaque case. The model opens only matching articles and copies finding IDs fromcandidateArticles[].path. Save each response with the matchingresult-case-*.jsonname in the same directory.request-<domain>.jsonfiles provide optional two-case leaf batches; save those asresult-<domain>.json. Directory scoring prefersresult-case-*.jsonwhen present and otherwise falls back toresult-*.json.review-request.jsonis an optional all-domains stress test for larger models. Neither batch form is the preferred fast-model profile. -
Save only this result shape:
{ "cases": [ { "id": "case-a1b2c3d4", "findings": [ { "id": "microsoft/knowledge/appsource/object-affixes-prevent-collisions.md" } ] } ] }Include every case. A clean control has an empty
findingsarray. -
Score all per-leaf results together:
pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -ResultsDirectory ./.evaluation-runFor a single combined stress-test result, use
-ResultsPathinstead.
The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.