bcquality/evaluation/README.md
Michael Dieringer 31206f6616 Fix four merge-critical blockers from Jesper's review; add 4 more patterns
Addresses microsoft/BCQuality#175 review feedback:
- Extend al-data-modeling-review's entry gate/relevance scope and token
  list to recognize document actions, Navigate subscribers, Report
  Selection registration, price-calculation/price-source extensibility,
  TransferFields posting-cascade mirroring, and barcode font-provider
  usage - previously excluded before any worklist cue could run.
- Fix document-print-and-email-actions-call-report-selections-directly:
  permit the legitimate stateless DocumentSendingProfile.TrySendToPrinter/
  TrySendToEMail path; rework the bad fixture to load a configured
  profile instead of demonstrating a trivial blank-record no-op.
- Fix extend-report-selection-usage-for-new-document-types: scope to the
  applicable single counterparty (ReportSelectionHandlerCZZ partitions
  strictly; only genuinely two-sided usages like Compensation need both),
  and add the page-facing usage-enum map/validate events alongside the
  filter-event subscription for full Document Layouts support.
- Fix a stale field-citation in custom-document-dispatch-must-not-bypass-
  report-selections (Custom Report Layout Code is field 7, not part of
  the 19-26 email-configuration range).
- Add deterministic positive/clean evaluation coverage (review-fixtures.json
  additionalArticles + Test-ReviewFixtures.ps1 support) so all 9 new
  good/bad pairs are actually exercised, not just present.
- Add 4 new patterns: activate-new-price-calculation-handler-via-
  onfindsupportedsetup, extend-price-source-type-must-sync-document-
  subset-enum, new-price-source-must-add-candidate-and-trigger-
  recalculation, report-barcodes-must-use-barcode-module-and-production-
  font-name.

All claims verified against live microsoft/BCApps source and Microsoft
Learn. Validators: frontmatter 0/0, review-fixtures 52 cases/17 domains
PASSED, knowledge-index 309 articles PASSED.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-24 06:33:04 +02:00

3.5 KiB

AL review evaluation

The evaluation is convention-driven. The harness discovers every <layer>/skills/review/al-<domain>-review.md leaf across the enabled microsoft, community, and custom layers. Duplicate domains resolve with custom > community > microsoft precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both .bad.al and .good.al companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.

review-fixtures.json contains only global thresholds and optional exceptional overrides. An override may select a different article or add context when the generic convention cannot express a scenario. It should remain empty in the normal case. An override may also list additionalArticles — other same-domain slugs (each with a .good.al/.bad.al pair) that get their own deterministic positive/clean case pair alongside the convention-selected one. Use this when a single leaf's worklist covers several distinct, newly-added rules and each one needs its own proof of reachability rather than riding on whichever article the generic convention happens to select.

Model-facing preparation hashes case IDs, neutralizes Good/Bad object-name tokens, and removes full-line sample comments so neither the article slug, domain, nor expected outcome reveals the answer.

Validate the corpus

pwsh ./tools/Test-ReviewFixtures.ps1 -Root .

This credential-free check proves every selected leaf maps to a same-named knowledge domain with at least one complete AL sample pair and that all configured overrides are valid.

Run a fast-model evaluation

  1. Prepare neutral inputs:

    pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -PrepareDirectory ./.evaluation-run
    

    This is also the CI path. It derives all cases, builds the current index, requires the convention-selected article to rank naturally into the candidate cutoff, and prepares the neutral requests.

  2. For a fast/small model, use one fresh invocation per request-case-*.json. Each request embeds the exact leaf instructions, that domain's candidate index rows with authoritative paths, and one opaque case. The model opens only matching articles and copies finding IDs from candidateArticles[].path. Save each response with the matching result-case-*.json name in the same directory.

request-<domain>.json files provide optional two-case leaf batches and identify the selected layer-owned skill path; save those as result-<domain>.json. Directory scoring prefers result-case-*.json when present and otherwise falls back to result-*.json. review-request.json is an optional all-domains stress test for larger models. Neither batch form is the preferred fast-model profile.

  1. Save only this result shape:

    {
      "cases": [
        {
          "id": "case-a1b2c3d4",
          "findings": [
            { "id": "microsoft/knowledge/appsource/object-affixes-prevent-collisions.md" }
          ]
        }
      ]
    }
    

    Include every case. A clean control has an empty findings array.

  2. Score all per-leaf results together:

    pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -ResultsDirectory ./.evaluation-run
    

    For a single combined stress-test result, use -ResultsPath instead.

The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.