bcquality/evaluation/README.md
Michael Dieringer 31206f6616 Fix four merge-critical blockers from Jesper's review; add 4 more patterns
Addresses microsoft/BCQuality#175 review feedback:
- Extend al-data-modeling-review's entry gate/relevance scope and token
  list to recognize document actions, Navigate subscribers, Report
  Selection registration, price-calculation/price-source extensibility,
  TransferFields posting-cascade mirroring, and barcode font-provider
  usage - previously excluded before any worklist cue could run.
- Fix document-print-and-email-actions-call-report-selections-directly:
  permit the legitimate stateless DocumentSendingProfile.TrySendToPrinter/
  TrySendToEMail path; rework the bad fixture to load a configured
  profile instead of demonstrating a trivial blank-record no-op.
- Fix extend-report-selection-usage-for-new-document-types: scope to the
  applicable single counterparty (ReportSelectionHandlerCZZ partitions
  strictly; only genuinely two-sided usages like Compensation need both),
  and add the page-facing usage-enum map/validate events alongside the
  filter-event subscription for full Document Layouts support.
- Fix a stale field-citation in custom-document-dispatch-must-not-bypass-
  report-selections (Custom Report Layout Code is field 7, not part of
  the 19-26 email-configuration range).
- Add deterministic positive/clean evaluation coverage (review-fixtures.json
  additionalArticles + Test-ReviewFixtures.ps1 support) so all 9 new
  good/bad pairs are actually exercised, not just present.
- Add 4 new patterns: activate-new-price-calculation-handler-via-
  onfindsupportedsetup, extend-price-source-type-must-sync-document-
  subset-enum, new-price-source-must-add-candidate-and-trigger-
  recalculation, report-barcodes-must-use-barcode-module-and-production-
  font-name.

All claims verified against live microsoft/BCApps source and Microsoft
Learn. Validators: frontmatter 0/0, review-fixtures 52 cases/17 domains
PASSED, knowledge-index 309 articles PASSED.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 06:33:04 +02:00

56 lines
3.5 KiB
Markdown

# AL review evaluation
The evaluation is convention-driven. The harness discovers every `<layer>/skills/review/al-<domain>-review.md` leaf across the enabled `microsoft`, `community`, and `custom` layers. Duplicate domains resolve with `custom > community > microsoft` precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both `.bad.al` and `.good.al` companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.
`review-fixtures.json` contains only global thresholds and optional exceptional overrides. An override may select a different article or add context when the generic convention cannot express a scenario. It should remain empty in the normal case. An override may also list `additionalArticles` — other same-domain slugs (each with a `.good.al`/`.bad.al` pair) that get their own deterministic positive/clean case pair alongside the convention-selected one. Use this when a single leaf's worklist covers several distinct, newly-added rules and each one needs its own proof of reachability rather than riding on whichever article the generic convention happens to select.
Model-facing preparation hashes case IDs, neutralizes `Good`/`Bad` object-name tokens, and removes full-line sample comments so neither the article slug, domain, nor expected outcome reveals the answer.
## Validate the corpus
```powershell
pwsh ./tools/Test-ReviewFixtures.ps1 -Root .
```
This credential-free check proves every selected leaf maps to a same-named knowledge domain with at least one complete AL sample pair and that all configured overrides are valid.
## Run a fast-model evaluation
1. Prepare neutral inputs:
```powershell
pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -PrepareDirectory ./.evaluation-run
```
This is also the CI path. It derives all cases, builds the current index, requires the convention-selected article to rank naturally into the candidate cutoff, and prepares the neutral requests.
2. For a fast/small model, use one fresh invocation per `request-case-*.json`. Each request embeds the exact leaf instructions, that domain's candidate index rows with authoritative paths, and one opaque case. The model opens only matching articles and copies finding IDs from `candidateArticles[].path`. Save each response with the matching `result-case-*.json` name in the same directory.
`request-<domain>.json` files provide optional two-case leaf batches and identify the selected layer-owned skill path; save those as `result-<domain>.json`. Directory scoring prefers `result-case-*.json` when present and otherwise falls back to `result-*.json`. `review-request.json` is an optional all-domains stress test for larger models. Neither batch form is the preferred fast-model profile.
3. Save only this result shape:
```json
{
"cases": [
{
"id": "case-a1b2c3d4",
"findings": [
{ "id": "microsoft/knowledge/appsource/object-affixes-prevent-collisions.md" }
]
}
]
}
```
Include every case. A clean control has an empty `findings` array.
4. Score all per-leaf results together:
```powershell
pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -ResultsDirectory ./.evaluation-run
```
For a single combined stress-test result, use `-ResultsPath` instead.
The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.