bcquality/evaluation/README.md
Jesper Schulz-Wedde 07e324ddbc
Some checks failed
Validate knowledge index / validate-index (push) Has been cancelled
Validate AL review fixtures / validate-review-fixtures (push) Has been cancelled
Validate skill index and report schemas / validate-contract (push) Has been cancelled
Validate frontmatter and structure / validate (push) Has been cancelled
Add source-verified Finance knowledge and review domain (#57)
* Add Finance posting domain-knowledge pilot

Adds three atomic domain-rule knowledge files under community/knowledge/finance/ as a pilot for Type-A (normative) Business Central domain knowledge: post through the posting engine, treat posted ledger entries as immutable, and treat the Dimension Set ID as the source of truth for dimensions. Includes a good/bad AL sample pair for the posting rule.

These encode BC-specific invariants that LLMs reliably get wrong, fitting the existing remedial/atomic knowledge grain with no schema or contract changes. Passes the repo frontmatter validator and is discovered by the knowledge index.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Clarify editable operational fields on posted ledger entries

Addresses review feedback from @JeremyVyska on PR #57: the immutability rule applies to financial content, not the whole entry. Reframes the Description around financial content and gives the operational-field exception (payment/application data, on-hold, applies-to, communication fields edited via CustEntry-Edit/VendEntry-Edit and the ledger entry pages) its own paragraph in Best Practice instead of understating it as a narrow set.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

* Expand Finance pilot into a source-verified review domain

Move Finance knowledge to the Microsoft-owned layer, add nine scoped rules with eighteen AL samples, and register bounded Finance review with complete paired evaluation coverage.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Fix Finance review applicability and ownership boundaries

Remove application-area gating and later VAT-field dependencies, align dynamic shared conventions, separate SCM ownership, and keep journal examples focused on the intended invariant.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Remove logo branding (#194)

Co-authored-by: Jesper Schulz-Wedde <jesper.schulzwedde@microsoft.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Add title and description to README

* Add foundational AL developer knowledge (#195)

Co-authored-by: Jesper Schulz-Wedde <jesper.schulzwedde@microsoft.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Clarify locale-safe DateFormula Evaluate inputs (#193)

* Clarify locale-safe DateFormula Evaluate inputs

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Normalize DateFormula article sections

Keep the analyzer-gap explanation in Description and its scoped probe evidence in References, without a novel Validation section. Normative guidance and fixtures are unchanged.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Jesper Schulz-Wedde <jesper.schulzwedde@microsoft.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Add SCM functional knowledge domain (#192)

* Add SCM functional knowledge domain

Introduce nine source-backed rules with original AL sample pairs, bounded SCM review routing, and complete positive/clean evaluation coverage.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

* Normalize SCM knowledge and review ownership

Align article and AL sample conventions, keep BC facts separate from review mechanics, and clarify reciprocal Finance ownership without bespoke shared test assertions.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Jesper Schulz-Wedde <jesper.schulzwedde@microsoft.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

---------

Co-authored-by: Jesper Schulz-Wedde <jesper.schulzwedde@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-21 10:20:41 +02:00

5.1 KiB

AL review evaluation

The evaluation is convention-driven. The harness discovers every <layer>/skills/review/al-<domain>-review.md leaf across the enabled microsoft, community, and custom layers. Duplicate domains resolve with custom > community > microsoft precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both .bad.al and .good.al companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.

review-fixtures.json contains only global thresholds and optional exceptional overrides. An override may select a different article, add context when the generic convention cannot express a scenario, or use an articles array when one domain needs explicit regression coverage for several paired articles. Specify either article or articles, not both. The first selected article retains the stable <domain>-bad and <domain>-good manifest IDs; additional articles use slug-qualified IDs. Overrides should remain empty in the normal case.

Model-facing preparation hashes case IDs, neutralizes Good/Bad object-name tokens, and removes full-line sample comments so neither the article slug, domain, nor expected outcome reveals the answer.

The SCM articles override deliberately selects every rule in the initial functional domain, producing nine positive cases and nine clean controls. Business context is executable: document/status TestField guards, calculated-revaluation fields, source-transfer base quantities, warehouse reconciliation steps, and additional-demand promising parameters survive neutralization. Do not move those preconditions into comments or generic "posting" helper names; removing them can turn a real defect into a valid alternative workflow. The clean pairs exercise the supported APIs selected by the same routing cues, not merely unrelated code that contains no SCM tokens.

The Finance override deliberately covers every paired Finance article, not only the first filename. Its shared context supplies the target version and localization but deliberately omits application area, as production callers often do. Finance applicability must come from its source-surface gate, not an artificial evaluation-only area hint. Scenario prerequisites live in executable AL: the document-balance cases check the template setting, and the VAT cases encode the imported net/VAT/gross totals and applicable VAT mode. Do not move these prerequisites into comments that preparation removes. Clean samples also retain supported operational edits, temporary ledger/set buffers, legitimate entry-number APIs, and reads of individual shortcut dimensions so these exceptions are exercised rather than blanket-excluded.

Validate the corpus

pwsh ./tools/Test-ReviewFixtures.ps1 -Root .

This credential-free check proves every selected leaf maps to a same-named knowledge domain with at least one complete AL sample pair and that all configured overrides are valid.

Run a fast-model evaluation

  1. Prepare neutral inputs:

    pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -PrepareDirectory ./.evaluation-run
    

    This is also the CI path. It derives all cases, builds the current index, requires the convention-selected article to rank naturally into the candidate cutoff, and prepares the neutral requests.

  2. For a fast/small model, use one fresh invocation per request-case-*.json. Each request embeds the exact leaf instructions, that domain's candidate index rows with authoritative paths, and one opaque case. The model opens only matching articles and copies finding IDs from candidateArticles[].path. Save each response with the matching result-case-*.json name in the same directory.

request-<domain>.json files provide optional leaf batches containing every selected case for that domain and identify the selected layer-owned skill path; save those as result-<domain>.json. A normal convention-selected domain has one bad/good pair, while an articles override contributes one pair per listed article. Directory scoring prefers result-case-*.json when present and otherwise falls back to result-*.json. review-request.json is an optional all-domains stress test for larger models. Neither batch form is the preferred fast-model profile.

  1. Save only this result shape:

    {
      "cases": [
        {
          "id": "case-a1b2c3d4",
          "findings": [
            { "id": "microsoft/knowledge/appsource/permission-sets-cover-setup-and-usage-without-super.md" }
          ]
        }
      ]
    }
    

    Include every case. A clean control has an empty findings array.

  2. Score all per-leaf results together:

    pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -ResultsDirectory ./.evaluation-run
    

    For a single combined stress-test result, use -ResultsPath instead.

The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.