From 3eed0bba3848fbaa8ce89c38cf1975e0c2015e78 Mon Sep 17 00:00:00 2001 From: wenjiefan Date: Wed, 2 Sep 2026 14:55:49 +0200 Subject: [PATCH] experiment: shift the adjudication burden of proof toward rejection The first arbitration smoke run showed the gate changing nothing: comment volume stayed inside baseline noise, and on a clean fixture the two AppSource-affix findings the gate exists to stop survived untouched. The likely cause is the contract's own wording. 'When the evidence is genuinely ambiguous, keep the finding' let any finding whose precondition could not be positively disproved pass through, which is most of them - a reviewer cannot prove from a bare src/*.al diff that the app is not an AppSource submission. That framing asked the gate to disprove findings rather than to confirm them. This inverts it. Applicability must now be affirmatively evident in the diff or in observable repository context, project-wide assumptions (packaging model, deployment target, tenancy, release state) may not be invented to support a finding, and an unresolved doubt resolves to rejection instead of a pass. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: e33a7b96-4821-4797-9f72-96880f3c6412 --- microsoft/skills/review/al-code-review.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/microsoft/skills/review/al-code-review.md b/microsoft/skills/review/al-code-review.md index 72825cd..9ebb262 100644 --- a/microsoft/skills/review/al-code-review.md +++ b/microsoft/skills/review/al-code-review.md @@ -81,12 +81,12 @@ For each sub-skill in the worklist, executed one at a time per the discipline ab 3. If the sub-skill's `outcome` is `failed`, stop here for this sub-skill: its findings are not reliable per the DO contract and MUST NOT be copied into the super-skill's top-level `findings[]` or counted in `summary.counts`. 4. Otherwise, compare each entry from the sub-skill's `findings[]` with findings already rolled up. Two findings are duplicates when they point to the same file and overlapping line/range and prescribe materially the same correction, even when their knowledge-file IDs differ. Merge duplicates instead of appending both: keep the more specific domain owner, preserve that finding's optional `domain` field verbatim (including its absence), use its reference as `references[0]` and therefore as `id`, append the other references as supporting references, keep the highest severity and confidence justified by either report, and preserve one self-contained message. Article and leaf ownership notes decide specificity; do not choose by execution order. 5. **Adjudicate each non-duplicate finding before appending it.** The super-skill is the review's final quality gate. A leaf reports what its own domain knowledge matched, in isolation; only the super-skill sees that finding beside the full diff and the other leaves' reports. Re-read the cited lines and decide whether the finding survives. **Reject** it — do not append it, and record it in `adjudicated-out[]` as `{ id, from-sub-skill, location, reason }` — when any of the following holds: - - **Precondition not met.** The applicability condition behind the finding does not actually hold here: the construct the knowledge file targets is absent, or this file is outside the article's declared scope. + - **Applicability not established.** The burden sits with the finding: keep it only when the article's applicability condition is affirmatively evident in the diff or in repository context you can actually observe. Reject it when the construct the knowledge file targets is absent, when this file is outside the article's declared scope, or when the precondition is merely *assumed* rather than visible here. In particular, do not infer a project-wide context the reviewed changes do not themselves establish — packaging or distribution model, deployment target, tenancy, or release state — and then report a finding that only holds under that assumption. - **Unsupported by the code.** The finding asserts something the cited lines do not bear out — a symbol, property, call, or control-flow claim that is not there. Re-read the code before accepting; do not take the leaf's description on trust. - **Contradicted by knowledge.** A knowledge file loaded for this task explicitly permits what the leaf flagged (its `## Best Practice` or `## Anti Pattern` says the opposite). - **Not actionable.** The finding names no concrete, correct change the author could make at that location. - Reject on evidence, never on taste. A correct, in-scope, actionable finding stays even when it is low severity, repeats a theme already reported elsewhere in the diff, or is not one you would have raised yourself. When the evidence is genuinely ambiguous, keep the finding. + Reject on evidence, never on taste. A correct, in-scope, actionable finding stays even when it is low severity, repeats a theme already reported elsewhere in the diff, or is not one you would have raised yourself. But the burden of proof runs toward rejection: a finding earns its place by being affirmatively supported by the code in front of you. When you cannot establish from that code that the finding both applies here and is correct, reject it and record the reason — an unresolved doubt is a rejection, not a pass. 6. Append each surviving finding, setting `from-sub-skill` to the sub-skill's `skill.id` and preserving its optional `domain` field verbatim, including its absence. For non-citation findings (those whose `id` is a skill-defined slug rather than a reference path), prefix `id` with `:` to prevent collisions across sub-skills. Other finding fields are preserved. ### Agent self-review pass