bcquality/microsoft/skills/review/al-testing-review.md
Michael Dieringer 0a8c9a8556
18 AL/BC patterns: style, data-modeling, web-services, appsource, breaking-changes, performance, testing (#156)
* Add 18 community AL/BC patterns across style, data-modeling, web-services, appsource, breaking-changes, performance, and testing

Contributed by CURABIS ApS, generalized from patterns observed across real AppSource/PTE development. Each article follows the knowledge file format (frontmatter, Description/Best Practice/Anti Pattern, sibling .good.al/.bad.al samples).

* Address Jesper Schulz-Wedde's review on PR #156

- Rename 3 articles so their .good.al/.bad.al companion stems match
  (do-not-change-primary-key, testfield-required-setup-field,
  al-identifiers-english), fixing the R14 orphan-sample errors.
- do-not-change-primary-key.good.al: include Flow in the new table's
  own primary key so it actually models the discriminating dimension.
- al-build-output-must-not-pollute-project-root.md: drop the
  unsubstantiated AL0197 causal claim and the non-existent
  al.outputPath setting; reframe as build-artifact hygiene sourced
  from ALTool --outfolder / al_build outputPath.
- prefer-email-module.md: Email Message is Codeunit 8904, not a table;
  distinguish it from the underlying Sent/Outbox/Draft storage.
- file-datatype-saas.md: File.Open/Create/Read/Write fails to compile
  against a Cloud-scoped project, it does not compile and silently
  fail at runtime.
- namespace-must-be-verified-from-source.md: narrow to "resolve from
  the referenced object's source or symbols," since source-file line
  one is not the only authoritative source (symbol packages, comments
  before the namespace line).
- test-data-must-be-random-and-complete.md: drop "assume an empty
  database" and "collision-free" absolutes; reframe around
  independence from unrelated business records and reserving explicit
  values for scenario-defining inputs.
- binary-choice-must-be-boolean.md: scope to genuine true/false
  semantics, not mechanical two-member-enum-to-boolean conversion.
- document-report-word-layout.md: scope down to a sourced Microsoft
  Learn recommendation instead of an unconditional performance
  guarantee; cite the three Learn pages.
- Wire the new articles into their review skills' candidate-selection
  signals (file-datatype-saas, prefer-email-module,
  namespace-must-be-verified-from-source, var-parameters-require-an-
  addressable-variable) so they can actually enter a worklist.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Fix dimension-management-wiring.md: ValidateShortcutDimCode and CreateDim
do not exist on the current DimensionManagement codeunit

Verified against microsoft/BCApps: the real master-table validation
procedure is ValidateDimValueCode (or ValidateShortcutDimValues when a
DimSetID is also needed), and the real document-side inheritance
procedure is GetDefaultDimID, not CreateDim. Caught from Jesper
Schulz-Wedde's review thread, which had been partially hidden by
GitHub's comment folding.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Address second round of Jesper Schulz-Wedde's review on PR #156

- dimension-management-wiring.md/.good.al: split into the two distinct
  models the article was conflating - master data (Default Dimension
  records via ValidateDimValueCode/SaveDefaultDim) vs. transactional/
  document data (a single Dimension Set ID assembled via AddDimSource +
  GetDefaultDimID, verified against BCApps' ExchRateAdjmtProcess.Codeunit.al).
  Added a compiling document-table example alongside the existing master
  table one.
- Deleted api-page-flowfields-must-be-calcfields (.md/.good.al/.bad.al):
  Microsoft's own FlowFields documentation states a FlowField used as a
  control's direct source expression is automatically calculated on any
  page - no API-page exception is documented, and none could be
  reproduced.
- prefer-email-module.bad.al/.md: Codeunit Mail has no Send/GetErrorDesc
  members; fixed to the real current 7-argument CreateMessage signature,
  and corrected the claim that the legacy path "still runs" - its base
  implementation no longer sends anything, only raises integration events.
- check-post-line-batch-pattern.md/.good.al: reframed from a universal
  invariant to the standard shape, naming the real Gen./Item/CA/Res./Job/
  Insurance/Mfg. Item/FA Jnl.-Check Line/-Post Line/-Post Batch codeunits
  it's based on. Added the missing Check Line companion codeunit so the
  good fixture is internally complete.
- test-data-must-be-random-and-complete.good.al: removed leftover
  "collision-free" wording contradicting the already-corrected article text.
- fixed-choice-set-must-use-enum-not-integer.md: removed the reintroduced
  state-count heuristic ("the line is the state count"), aligned with
  binary-choice-must-be-boolean.md's semantics-based distinction.
- namespace-must-be-verified-from-source.md: removed the false claim that
  the compiler and AL Language Server use different namespace-resolution
  rules.
- intrinsic-al-functions-must-use-modern-casing.md: removed the unverified
  claim that PascalCase is the VS Code formatter's default output.

Worklist completeness: added cues for the 8 rules in data-modeling,
testing, performance, and web-services that had none (Jesper's explicit
ask), plus the same gap in all 7 style rules from this PR (not explicitly
named this round, but the identical systemic issue) - 15 cues total across
al-data-modeling-review.md, al-testing-review.md, al-performance-review.md,
al-web-services-review.md, and al-style-review.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Fix remaining correctness issues from Jesper's 2026-09-15 re-review

- dimension-management-wiring: SaveDefaultDim's third argument is the
  shortcut dimension number (1-8), not the field's AL field ID; the
  fixture passed FieldNo(...) = 10. GetDefaultDimID's InheritFromDimSetID
  must be 0 when recomputing after the linking record changes, not the
  document's existing Dimension Set ID (which would retain the previous
  customer's leftover dimensions). Verified against
  DimensionManagement.Codeunit.al and BankDepositHeader.Table.al in the
  BCApps reference clone.
- check-post-line-batch-pattern: "Post Line writes exactly one line to
  the ledger" overclaimed - Gen. Jnl.-Post Line alone calls InsertGLEntry
  from a dozen call sites (balancing entry, VAT, currency rounding,
  deferrals) and can write several G/L Entries per journal line.
  Reworded to "posts exactly one journal line" and softened the
  "distinct, non-overlapping responsibilities" absolute.
- namespace-must-be-verified-from-source.bad.al: dropped the "resolves
  in a local build, fails in VS Code" comment (taught an inherent
  compiler/language-server disagreement that isn't real); reframed as
  stale/cached symbols, matching the prose fix already made.
- file-datatype-saas.good.al: replaced the deprecated 5-argument
  UploadIntoStream overload with the current 2-argument one, and
  actually staged through TempBlob as the article's own Best Practice
  instructs (the declared TempBlob variable was previously unused).
- test-data-must-be-random-and-complete: no longer treats a
  short-but-valid value as defective merely for being "underfilled" -
  AL field lengths are maxima, not minimums. Scoped to missing values
  or a scenario with an explicit length/format requirement (e.g. a
  truncation test). Updated the al-testing-review.md routing cue to
  match.
- stored-derived-fields-must-not-be-exposed-directly: stopped mandating
  source-field exposure as part of the core pattern: the good fixture
  exposed only one of the derived value's two inputs (Hours Used, not
  Budgeted Hours), making the claimed "so the consumer can verify it"
  impossible. Reframed as an optional, all-or-nothing addition and
  fixed the fixture to expose both inputs.

Rebased onto upstream/main to resolve conflicts in
al-breaking-changes-review.md, al-data-modeling-review.md,
al-performance-review.md, and al-style-review.md against merged PRs
#148 and #153; all sides' worklist tokens/cues retained.

* Fix remaining READ-convention sample links across this PR's 18 articles

The same plain-backtick "See sample: \`x.good.al\`." form fixed on
al-methods-limited-during-write-transactions (PR #161) turned up
repo-wide on 15 more of this PR's articles - Knowledge-Retrieval.ps1
requires the markdown-link form to associate a sample with its
article. All 16 fixed; the four local validators (frontmatter,
knowledge-index, knowledge-retrieval, review-fixtures, skill-index)
pass.

* Fix two merge-critical correctness issues from Jesper's 2026-09-22 review

- api-page-key-fields-must-be-editable-on-insert.good.al and
  stored-derived-fields-must-not-be-exposed-directly.good.al: both were
  writable API pages missing DelayedInsert = true, contradicting this
  repo's own api-page-delayedinsert-true rule - the canonical "good"
  samples were teaching code BCQuality itself flags.
- dimension-management-wiring.good.al: UpdateDimensionSetID exited
  early when Customer.Get failed, leaving the previous customer's
  shortcut dimension and Dimension Set ID in place - the same staleness
  bug the InheritFromDimSetID = 0 fix (from the prior review round) was
  meant to prevent, just triggered by a failed lookup instead of a
  successful one. Now clears the shortcut field and recomputes with an
  empty source list on a failed lookup too, so GetDefaultDimID
  correctly returns an empty Dimension Set ID instead of never running.

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-29 12:48:17 +02:00

14 KiB

kind id version title description inputs outputs bc-version technologies countries application-area
action-skill al-testing-review 1 AL testing review Performs an AL testing review against guidance from BCQuality.
pr-diff
file-path
folder-path
findings-report
all
al
w1
all

AL testing review

Reviews AL source changes against the testing knowledge domain in BCQuality and emits a findings report. This is a leaf action skill: it invokes no sub-skills. It is one of the skills composed by al-code-review.

An orchestrator invokes this skill with a pr-diff, file-path, or folder-path. Testing findings are narrow by design — they apply when the review scope contains test codeunits, test runners, test methods, handlers, assertions, or fixture construction. The skill returns not-applicable when none of those apply.

Source

Use READ's Bounded retrieval for review skills workflow with -Domain testing. Consume every catalog page across enabled layers before applying this leaf's Relevance and Worklist; preserve each exact catalog path and open complete bodies only for exact paths selected by the Worklist. If the helper or prepared index is unavailable or invalid, use READ's explicit path-discovery and bounded native-read fallback.

Relevance

Apply the frontmatter matching rules defined in READ (Frontmatter matching semantics) against the task context:

  • bc-version — the target BC version from the PR branch's app.json or the orchestrator-supplied version. If unavailable, the dimension is unknown.
  • technologies — [al].
  • countries — the countries declared in the consuming app's app.json. Default to the orchestrator's configured context; if absent, unknown.
  • application-area — the union of application areas declared by the changed objects. Pass the actual set; do not substitute [all]. If the area cannot be determined from the changes, the dimension is unknown.

Discard files that are not applicable. Retain conditionally applicable files (any dimension unknown) only when the orchestrator's configuration permits them; findings derived from those files MUST have confidence no higher than medium, AND the finding's message MUST name the dimension or dimensions that were unknown.

Worklist

Narrow the relevant files to the subset that applies to the changes under review. For each relevant file, compute overlap against:

  • The changed AL object names and types — especially codeunits with Subtype = Test, test runner codeunits with TestIsolation, test libraries, and codeunits that define UI handlers.
  • The changed methods and attributes, weighted toward [Test], [TransactionModel(...)], [TestPermissions(...)], [HandlerFunctions(...)], handler attributes, asserterror, ExpectedError, ExpectedErrorCode, fixture initialization, and test-library calls.
  • Tokens extracted from the diff that relate to testing (Subtype = Test, Subtype = TestRunner, TestIsolation, TestPermissions, Restrictive, NonRestrictive, Disabled, Permissions Mock, Library - Lower Permissions, TransactionModel, AutoRollback, AutoCommit, Commit, asserterror, ExpectedError, ExpectedErrorCode, HandlerFunctions, ConfirmHandler, MessageHandler, StrMenuHandler, ModalPageHandler, SendNotificationHandler, RecallNotificationHandler, Enqueue, Dequeue, AssertEmpty, Library Assert, LibraryVariableStorage, LibrarySales, LibraryPurchase, LibraryERM, LibraryInventory, LibraryRandom, Init, Insert).

A file enters the candidate worklist when its keywords intersect the extracted tokens or its topic (derived from the index entry's path, title, and description) matches a changed object type. Read an article's full file — its ## Best Practice / ## Anti Pattern bodies — only after it makes the worklist; candidate selection uses the index alone. When the diff contains no testing-related changes by any of the above signals, return outcome: "not-applicable" without evaluating files.

The following targeted checks cover every current testing article. Treat each as a candidate-selection cue: when the signal appears in changed code, add the named article to the worklist and evaluate it in Action.

  • A method in a Subtype = Test codeunit adds or changes [TransactionModel(...)], exercises code that calls Commit under AutoRollback, defaults broadly to AutoCommit, or chooses None for a writing test — transactionmodel-attribute-governs-test-transactions.
  • An AutoCommit test runs under a Subtype = TestRunner codeunit that omits TestIsolation or sets it to Disabled, leaving committed data between tests — testisolation-belongs-on-the-test-runner. Require runner/repository context; a standalone test file cannot prove which runner executes it.
  • A permission-sensitive test uses TestPermissions = Disabled, claims to test a restricted user without "Permissions Mock"/"Library - Lower Permissions", or declares [TestPermissions(...)] without applying that context — permission-tests-must-lower-the-execution-context.
  • Test fixture code manually calls Init/Insert, invents keys or prerequisite records, or bypasses available LibrarySales, LibraryPurchase, LibraryERM, LibraryInventory, LibraryRandom, or equivalent library codeunits — use-library-codeunits-for-test-fixtures.
  • asserterror is added or changed without a following Assert.ExpectedError, Assert.ExpectedErrorCode, or a purpose-built assertion such as ExpectedTestFieldError — asserterror-needs-expectederror-and-code.
  • A test path raises UI and [HandlerFunctions(...)] does not match the invoked handlers, or the test has no meaningful evidence of the UI result (for example, it treats a Boolean set before the action as proof of success) — ui-handlers-in-tests. A capture/reset/assert-after-RunModal pattern is valid. Enqueue/dequeue and AssertEmpty are required only when order, count, text, replies, or a scripted sequence is part of the contract. Only nonoptional handlers have to execute: a listed handler declared [SendNotificationHandler(true)] or [RecallNotificationHandler(true)] is optional by design, so do not treat it as unmatched when the run never raises the notification.
  • A test's [GIVEN]/setup looks up a hardcoded code/number/name assumed to already exist instead of creating it, leaves a mandatory field on a created record empty, uses a value that doesn't satisfy the scenario's own explicit length/format requirement (for example a truncation test whose value never exceeds the field), or a scenario-defining value (amount, quantity, percentage, date, threshold, rounding precision) is generated/randomized instead of an explicit chosen value — test-data-must-be-random-and-complete. Generating incidental fixture values (identifiers, names, descriptions) via the standard library codeunits is the compliant shape, not the signal to flag, and neither is a short-but-valid value in an otherwise-unremarkable field.

Once the candidate worklist is known, resolve layer-precedence conflicts per READ. Drop lower-precedence files whose normative guidance (## Best Practice or ## Anti Pattern) directly contradicts a higher-precedence candidate, and record each dropped file in suppressed with reason: "layer-precedence". Files that would have been candidates but are hidden because their layer is disabled in consumer configuration are recorded with reason: "configuration". Files that never became candidates are NOT recorded in suppressed.

When the post-conflict worklist is empty because no applicable testing knowledge exists, or because configuration suppressed every candidate, emit outcome: "no-knowledge". When the worklist is empty because no applicable testing knowledge matched the changes, emit outcome: "completed" with an empty findings array.

Action

For each worklist entry, evaluate the diff against the file's ## Best Practice and ## Anti Pattern sections. Emit findings as follows:

  • When the diff contains a clear match for an Anti Pattern, emit a finding with severity major or blocker, a message summarizing the anti-pattern, location pointing to the offending line or range, and a references entry pointing to the knowledge file. Use blocker only when the test can pass while verifying the wrong behavior or can leave committed data that contaminates later tests; otherwise the ceiling is major.
  • When the diff contains code that contradicts a Best Practice without being a full anti-pattern, emit minor with the same reference shape.
  • Applicability alone is not a finding. Emit info only for a concrete, non-actionable observation the article explicitly defines; otherwise emit nothing when no violation is present.

For ui-handlers-in-tests, use major when missing or incorrectly listed handlers make the test fail at runtime. Use minor when the test executes but lacks a meaningful semantic postcondition, including a pre-set Boolean used as proof. Do not escalate solely because a handler does not use queue storage or asserts inside the handler.

Set confidence to:

  • high when the detection is based on an unambiguous pattern match (attribute, handler declaration, assertion sequence, or fixture call).
  • medium when detection relies on heuristics or when any frontmatter dimension was unknown.
  • low when the finding is an advisory derived only from applicability.

After evaluating each worklist entry, also consider whether the diff exhibits a testing defect the agent recognises from its general AL knowledge that no knowledge file in the worklist covers. Such candidates are agent findings within this skill's domain — emit them with references: [], an id slug prefixed with agent:, confidence capped at medium, severity capped at minor (agent findings are advisory and non-gating), and a message that is self-contained (describing both the issue and a concrete recommendation, since there is no knowledge-file footer for the consumer to fall back on). Hold every candidate to the precision bar in skills/do.md (Agent findings): emit only a concrete, material testing defect a knowledgeable BC reviewer would agree is wrong — steelman it first and drop anything stylistic, speculative, dependent on code outside the diff, or merely a valid alternative; when in doubt, omit. The scope is strictly AL testing; defects outside this domain belong to other leaves and MUST NOT be emitted here. Before emitting, check the worklist for a knowledge file that matches the candidate — if one exists, upgrade the candidate to a knowledge-backed finding instead. See skills/do.md for the full contract.

For every emitted finding, decide whether the fix is mechanical. A fix is mechanical when it is small, local, and unambiguous from the diff context (for example: add the matching ExpectedError assertion after asserterror; add or remove a handler name in HandlerFunctions, except that a listed optional notification handler must never be proposed for removal; add LibraryVariableStorage.Clear or AssertEmpty when queue/LVS intentionally verifies interaction order, count, text, replies, or a scripted sequence; or replace hand-rolled fixture creation with an evident library call). For mechanical findings, emit findings[].suggested-code with the literal replacement for the source lines indicated by location. The payload must be a verbatim replacement — no diff markers, no fences, no commentary — that the consumer can render as a one-click suggestion. When a .good.al companion exists and the diff context matches the .bad.al shape, adapt the .good.al replacement into suggested-code.

Omit suggested-code only when the appropriate fix depends on context the skill cannot determine, when multiple defensible replacements exist, or when the fix spans non-contiguous code. If a finding is mechanical-looking but you omit suggested-code, set findings[].suggested-code-omission-reason to a short explanation. See skills/do.md for the full contract.

Outcome selection:

  • completed — the skill evaluated every worklist item.
  • no-knowledge — no applicable testing knowledge survived filtering.
  • not-applicable — the diff touches no test codeunit, runner, method, handler, assertion, or fixture surface.
  • partial — a budget was hit before the worklist was exhausted.
  • failed — an unrecoverable error occurred.

Output

Output conforms to the DO output contract. Every finding this skill emits MUST set findings[].domain to "Testing". A populated example:

{
  "skill": { "id": "al-testing-review", "version": 1 },
  "outcome": "completed",
  "summary": {
    "counts": { "blocker": 0, "major": 1, "minor": 0, "info": 0 },
    "coverage": { "worklist-size": 1, "items-evaluated": 1 }
  },
  "findings": [
    {
      "id": "microsoft/knowledge/testing/asserterror-needs-expectederror-and-code.md",
      "severity": "major",
      "message": "The negative test uses asserterror without checking the resulting message or error code, so any unrelated setup or permission error can make the test pass.",
      "location": {
        "file": "test/SalesPostingTests.Codeunit.al",
        "line": 42
      },
      "references": [
        { "path": "microsoft/knowledge/testing/asserterror-needs-expectederror-and-code.md" }
      ],
      "confidence": "high",
      "domain": "Testing",
      "suggested-code": "asserterror PostInvalidOrder();\nAssert.ExpectedError(ExpectedPostingErr);"
    }
  ],
  "suppressed": []
}

The empty-corpus case produces:

{
  "skill": { "id": "al-testing-review", "version": 1 },
  "outcome": "no-knowledge",
  "summary": {
    "counts": { "blocker": 0, "major": 0, "minor": 0, "info": 0 },
    "coverage": { "worklist-size": 0, "items-evaluated": 0 }
  },
  "findings": [],
  "suppressed": []
}