bcquality/evaluation
2026-09-18 12:04:15 +02:00
..
development-guidance-fixtures.json Narrow AL development to read-only plan guidance 2026-09-07 11:47:42 +02:00
implementation-guidance-fixtures.json Add AL implementation guidance skill 2026-09-16 13:20:57 +02:00
README.md Merge 4ff0fe0160 into ab06998818 2026-09-18 12:04:15 +02:00
review-fixtures.json Add query filter semantics guidance (#186) 2026-09-15 12:51:15 +02:00

AL review and guidance evaluation

The evaluation is convention-driven. The harness discovers every <layer>/skills/review/al-<domain>-review.md leaf across the enabled microsoft, community, and custom layers. Duplicate domains resolve with custom > community > microsoft precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both .bad.al and .good.al companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.

review-fixtures.json contains only global thresholds and optional exceptional overrides. An override may select a different article, add context when the generic convention cannot express a scenario, or use an articles array when one domain needs explicit regression coverage for several paired articles. Specify either article or articles, not both. The first selected article retains the stable <domain>-bad and <domain>-good manifest IDs; additional articles use slug-qualified IDs. Overrides should remain empty in the normal case.

Model-facing preparation hashes case IDs, neutralizes Good/Bad object-name tokens, and removes full-line sample comments so neither the article slug, domain, nor expected outcome reveals the answer.

Validate the corpus

pwsh ./tools/Test-ReviewFixtures.ps1 -Root .

This credential-free check proves every selected leaf maps to a same-named knowledge domain with at least one complete AL sample pair and that all configured overrides are valid.

Run a fast-model evaluation

  1. Prepare neutral inputs:

    pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -PrepareDirectory ./.evaluation-run
    

    This is also the CI path. It derives all cases, builds the current index, requires the convention-selected article to rank naturally into the candidate cutoff, and prepares the neutral requests.

  2. For a fast/small model, use one fresh invocation per request-case-*.json. Each request embeds the exact leaf instructions, that domain's candidate index rows with authoritative paths, and one opaque case. The model opens only matching articles and copies finding IDs from candidateArticles[].path. Save each response with the matching result-case-*.json name in the same directory.

request-<domain>.json files provide optional leaf batches containing every selected case for that domain and identify the selected layer-owned skill path; save those as result-<domain>.json. A normal convention-selected domain has one bad/good pair, while an articles override contributes one pair per listed article. Directory scoring prefers result-case-*.json when present and otherwise falls back to result-*.json. review-request.json is an optional all-domains stress test for larger models. Neither batch form is the preferred fast-model profile.

  1. Save only this result shape:

    {
      "cases": [
        {
          "id": "case-a1b2c3d4",
          "findings": [
            { "id": "microsoft/knowledge/appsource/permission-sets-cover-setup-and-usage-without-super.md" }
          ]
        }
      ]
    }
    

    Include every case. A clean control has an empty findings array.

  2. Score all per-leaf results together:

    pwsh ./tools/Test-ReviewFixtures.ps1 -Root . -ResultsDirectory ./.evaluation-run
    

    For a single combined stress-test result, use -ResultsPath instead.

The committed gate requires full expected recall, the exact convention-derived article ID, and no findings on clean controls.

Read-only plan guidance

development-guidance-fixtures.json evaluates the planning interface used by existing workflows. It supplies an existing plan and expects referenced constraints without target-repository changes. The initial-plan fixture is anonymized and synthetic: a full document with metadata and a markdown body covering root cause, proposed fix, affected files, tests, and acceptance criteria. It is integration-shaped input, not private consumer content, a continuation checkpoint, or proof that any production consumer is integrated.

Credential-free contract and scorer coverage

CI validates the manifest, prepares opaque model requests, and runs deterministic scorer regressions with controlled reports and temporary Git repositories. These checks cover report shape, outcomes, reference paths, and the evaluator's pre/post read-only comparison. They do not run an agent, compile AL, run Business Central tests, or establish better code authoring.

Prepare requests in a runner-owned artifact directory outside every target workspace:

$run = Join-Path ([IO.Path]::GetTempPath()) 'bcquality-guidance-run'
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . -PrepareDirectory $run

Preparation is not a model run. A scorer can validate a citation's path and required fields, but only external agent traces and expert evaluation can establish that the article was opened and its normative constraints faithfully applied. Expected knowledge recall/precision is fixture-specific, not a corpus coverage or authoring capability percentage.

no-knowledge means no additional applicable BCQuality constraints, not unsafe work. It requires empty knowledge. Partial evaluation, failed retrieval, unknown context, and materially unresolved applicability must remain visible and distinct; they cannot be counted as successful enrichment merely because the JSON is parseable.

Run the deterministic regression suite without an agent or AL environment:

pwsh ./tools/Test-DevelopmentGuidanceEvaluator.ps1

Runner-owned read-only evidence

For an external guidance run, first provision a representative, standalone Git repository for each manifest case. The runner supplies a JSON workspace map whose keys are the manifest IDs (not the hashed model IDs) and whose values are absolute workspace roots. It may pass -WorkspaceMapPath during preparation to bind the generated requests to those roots. The model must not select its own workspace for scoring.

Capture evidence before invoking the agent, with the manifest, workspace map, and source checkout already finalized:

# Runner-selected paths, all outside the targets and BCQuality checkout.
$map = Join-Path $evidenceDirectory 'workspace-map.json'
$baseline = Join-Path $evidenceDirectory 'baseline.json'
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
    -CaptureBaseline -WorkspaceMapPath $map -BaselinePath $baseline

# Retain the printed SHA256 in runner-only state BEFORE agent invocation.
# After the external agent writes result-case-<hash>.json files:
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
    -ResultsDirectory $resultsDirectory -BaselinePath $baseline `
    -BaselineSha256 $preRunDigest

$evidenceDirectory, $resultsDirectory, and $preRunDigest are supplied by the runner; the digest must not be recomputed from potentially modified evidence after the agent runs. Protect the baseline, digest, evaluator, and invocation from agent changes. Capture refuses to overwrite an existing baseline. Results contain only caseId and guidanceReport; a legacy workspaceRoot, if present, must agree with the independently captured binding and never overrides it. Missing baselines or digests, malformed reports, and escaped reference paths fail scoring.

The comparison checks target identity, Git HEAD, refs and index, filesystem content and stable metadata, including tracked, untracked, ignored files and empty directories. Committing edits or making an empty commit does not evade the check. It also compares the actual knowledge checkout and manifest identity. Targets must have internal Git storage; linked target worktrees, submodules, sparse checkouts, links/junctions/reparse points, hard links, and alternate data streams are unsupported and rejected rather than silently excluded. A linked knowledge checkout is supported with its Git storage identity recorded. Use quiescent, isolated repositories; concurrent changes also fail the gate.

This is before/after evidence, not an OS sandbox or a complete write monitor. It cannot prove that no transient write was reverted, that articles were opened, or that constraints are semantically faithful. Reports and generated artifacts must stay outside all target workspaces and the knowledge checkout. The regression suite creates and removes its own uniquely named fixture directory; it does not run against or clean a caller's target.

External agent/runtime pilot (follow-up)

Consumer uptake and a real before-authoring pilot are not implemented by these fixtures. The consumer must normalize its normal initial plan, persist guidance after state initialization, inject it into existing phases, re-enrich on material changes, and run an independent final review. See the integration boundary.

Before claiming improved repairs, run an independent pinned baseline without enrichment and a matched enriched run. Hold starting code, task, model, tools, runtime, and gates constant; record actual immutable BCQuality checkout and policy identities rather than trusting a configured ref. Use that same recorded checkout for enrichment and final review. Retain external logs, article-read traces, resulting diffs, compile/test outcomes, and independent review evidence, including failures, no-knowledge, partial, and unresolved results. No compile/run or authoring-quality claim follows from the credential-free checks above.

Read-only implementation guidance

implementation-guidance-fixtures.json covers a distinct just-in-time consultation contract. Its production-shaped synthetic AL contexts exercise a new privacy constraint after the implementation surface expands beyond the plan, exact consumed-guidance omission, an honest no-additional-guidance control, missing current diff/decision context, and a public-interface checkpoint. They are neutral fixtures, not copies of knowledge samples or evidence of a production integration.

Validate and prepare the implementation manifest with the shared scorer:

$run = Join-Path ([IO.Path]::GetTempPath()) 'bcquality-implementation-guidance-run'
pwsh ./tools/Test-DevelopmentGuidanceFixtures.ps1 -Root . `
    -ManifestPath ./evaluation/implementation-guidance-fixtures.json `
    -PrepareDirectory $run
pwsh ./tools/Test-ImplementationGuidanceEvaluator.ps1

The same runner-owned baseline and workspace-map flow used for plan guidance proves that evaluated target repositories remain byte-for-byte and Git-identity stable across scoring. The implementation regression adds strict phase/decision/pin preservation and exact path + decision key + evidence fingerprint deduplication checks. This remains before/after evidence, not an OS sandbox; it cannot detect a transient reverted write.

BCQuality does not keep consumed state or run checkpoints automatically. The consumer chooses explicit checkpoints, persists consumed triples, recomputes evidence fingerprints, owns edits/build/tests/retries/delivery, and runs an independent final review. no-knowledge is only an additive retrieval outcome, never a functional-correctness or release-readiness claim.

A future experiment should compare four matched arms: baseline, plan-only, implementation-only, and combined. Hold task, starting code, model, tools, runtime, checkpoints, and gates constant. This is a recommended evaluation design, not a claimed result.