The narrowed UI-handler guidance still stated the execution rule without the qualifier the linked Microsoft reference uses. The article said every listed handler must execute at least once, and the testing leaf skill asked for `[HandlerFunctions(...)]` to match the invoked handlers exactly. The reference says every *nonoptional* listed handler must execute, and that send-notification and recall-notification handlers can be optional. As written, an agent could flag a deliberately unused optional notification handler. The discriminator is narrower than the handler type. Both `[SendNotificationHandler([HandlerIsOptional: Boolean])]` and `[RecallNotificationHandler([HandlerIsOptional: Boolean])]` take an explicit optionality argument, so `[SendNotificationHandler(true)]` is exempt while the same attribute written without the argument stays nonoptional like every other handler type. Keying the exemption on the argument rather than the type keeps it checkable from the diff and avoids the opposite false positive, where an agent stops flagging genuinely nonoptional notification handlers. Changes: - The article now states the nonoptional qualifier, explains that optionality is declared rather than inferred, and adds an explicit do-not-flag clause. That clause also forbids proposing removal, because the listed entry is what keeps the test passing on the runs where the notification does fire. - The testing leaf skill carries the same boundary in its `ui-handlers-in-tests` cue, and its mechanical-fix list no longer allows removing a listed optional notification handler as a one-click suggestion. - `SendNotificationHandler` and `RecallNotificationHandler` were missing from the skill's testing token list, so notification handlers were not reliably surfaced to the relevance step at all. Both are now listed. - The good sample gains a test that lists an unreached `[SendNotificationHandler(true)]`; the bad sample gains the mirror image, an unreached `[SendNotificationHandler]` with no optionality argument. The pair differs only by that argument, which is the point. - `evaluation/review-fixtures.json` pins the testing domain to `ui-handlers-in-tests` so the boundary is exercised: the good sample is the clean control at `minimumCleanRate` 1.0 and the bad sample is the expected finding. Keywords were retagged with `notification` and `optional-handler`. validate_frontmatter.py reports 0 errors; Test-ReviewFixtures.ps1 passes with 32 cases across 16 leaf domains and resolves the testing fixture to this article. |
||
|---|---|---|
| .. | ||
| README.md | ||
| review-fixtures.json | ||
AL review evaluation
The evaluation is convention-driven. For every microsoft/skills/review/al-<domain>-review.md leaf, the harness finds microsoft/knowledge/<domain>/, selects the first article (by filename) with both .bad.al and .good.al companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.
review-fixtures.json contains only global thresholds and optional exceptional overrides. An override may select a different article or add context when the generic convention cannot express a scenario. It should remain empty in the normal case.
Model-facing preparation hashes case IDs, neutralizes Good/Bad object-name tokens, and removes full-line sample comments so neither the article slug, domain, nor expected outcome reveals the answer.
Validate the corpus
pwsh ./tools/Test-ReviewFixtures.ps1 -Root .
This credential-free check proves every registered leaf maps to a same-named knowledge domain with at least one complete AL sample pair and that all configured overrides are valid.
Run a fast-model evaluation
-
Prepare neutral inputs:
pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -PrepareDirectory ./.evaluation-runThis is also the CI path. It derives all cases, builds the current index, requires the convention-selected article to rank naturally into the candidate cutoff, and prepares the neutral requests.
-
For a fast/small model, use one fresh invocation per
request-case-*.json. Each request embeds the exact leaf instructions, that domain's candidate index rows with authoritative paths, and one opaque case. The model opens only matching articles and copies finding IDs fromcandidateArticles[].path. Save each response with the matchingresult-case-*.jsonname in the same directory.request-<domain>.jsonfiles provide optional two-case leaf batches; save those asresult-<domain>.json. Directory scoring prefersresult-case-*.jsonwhen present and otherwise falls back toresult-*.json.review-request.jsonis an optional all-domains stress test for larger models. Neither batch form is the preferred fast-model profile. -
Save only this result shape:
{ "cases": [ { "id": "case-a1b2c3d4", "findings": [ { "id": "microsoft/knowledge/appsource/object-affixes-prevent-collisions.md" } ] } ] }Include every case. A clean control has an empty
findingsarray. -
Score all per-leaf results together:
pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -ResultsDirectory ./.evaluation-runFor a single combined stress-test result, use
-ResultsPathinstead.
The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.