Scope adjudication burden of proof to project-wide premises

The blanket burden-of-proof wording rejected any finding whose applicability was not affirmatively demonstrable in the diff. Findings about consequences that scale with data volume or codebase size cannot meet that bar by nature, so performance and style findings were rejected wholesale.

Narrow the rejection-on-doubt rule to the project-wide premises it was written for (packaging model, deployment target, tenancy, release state) and restore keep-on-ambiguity as the default elsewhere.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e33a7b96-4821-4797-9f72-96880f3c6412
This commit is contained in:
wenjiefan 2026-09-02 17:17:12 +02:00
parent 3eed0bba38
commit c2e15986ce

View file

@ -81,12 +81,12 @@ For each sub-skill in the worklist, executed one at a time per the discipline ab
3. If the sub-skill's `outcome` is `failed`, stop here for this sub-skill: its findings are not reliable per the DO contract and MUST NOT be copied into the super-skill's top-level `findings[]` or counted in `summary.counts`.
4. Otherwise, compare each entry from the sub-skill's `findings[]` with findings already rolled up. Two findings are duplicates when they point to the same file and overlapping line/range and prescribe materially the same correction, even when their knowledge-file IDs differ. Merge duplicates instead of appending both: keep the more specific domain owner, preserve that finding's optional `domain` field verbatim (including its absence), use its reference as `references[0]` and therefore as `id`, append the other references as supporting references, keep the highest severity and confidence justified by either report, and preserve one self-contained message. Article and leaf ownership notes decide specificity; do not choose by execution order.
5. **Adjudicate each non-duplicate finding before appending it.** The super-skill is the review's final quality gate. A leaf reports what its own domain knowledge matched, in isolation; only the super-skill sees that finding beside the full diff and the other leaves' reports. Re-read the cited lines and decide whether the finding survives. **Reject** it — do not append it, and record it in `adjudicated-out[]` as `{ id, from-sub-skill, location, reason }` — when any of the following holds:
- **Applicability not established.** The burden sits with the finding: keep it only when the article's applicability condition is affirmatively evident in the diff or in repository context you can actually observe. Reject it when the construct the knowledge file targets is absent, when this file is outside the article's declared scope, or when the precondition is merely *assumed* rather than visible here. In particular, do not infer a project-wide context the reviewed changes do not themselves establish — packaging or distribution model, deployment target, tenancy, or release state — and then report a finding that only holds under that assumption.
- **Applicability not established.** Reject when the construct the knowledge file targets is absent from the cited lines, or when this file is outside the article's declared scope. Reject as well when the finding rests on a project-wide premise the reviewed changes do not themselves establish — packaging or distribution model, deployment target, tenancy, or release state. Such a premise must be affirmatively evident in the diff or in repository context you can actually observe; it is never inferred, and a finding that holds only under an unestablished premise does not survive.
- **Unsupported by the code.** The finding asserts something the cited lines do not bear out — a symbol, property, call, or control-flow claim that is not there. Re-read the code before accepting; do not take the leaf's description on trust.
- **Contradicted by knowledge.** A knowledge file loaded for this task explicitly permits what the leaf flagged (its `## Best Practice` or `## Anti Pattern` says the opposite).
- **Not actionable.** The finding names no concrete, correct change the author could make at that location.
Reject on evidence, never on taste. A correct, in-scope, actionable finding stays even when it is low severity, repeats a theme already reported elsewhere in the diff, or is not one you would have raised yourself. But the burden of proof runs toward rejection: a finding earns its place by being affirmatively supported by the code in front of you. When you cannot establish from that code that the finding both applies here and is correct, reject it and record the reason — an unresolved doubt is a rejection, not a pass.
Reject on evidence, never on taste. A correct, in-scope, actionable finding stays even when it is low severity, repeats a theme already reported elsewhere in the diff, or is not one you would have raised yourself. Findings whose consequence scales with something the diff cannot show — performance under data volume, maintainability, style drift — are not rejected merely because that consequence is not demonstrable here; judge them on whether the pattern itself is really present in the cited lines. Genuine ambiguity therefore resolves toward keeping the finding, with one exception: where a finding depends on an unestablished project-wide premise above, an unresolved doubt is a rejection.
6. Append each surviving finding, setting `from-sub-skill` to the sub-skill's `skill.id` and preserving its optional `domain` field verbatim, including its absence. For non-citation findings (those whose `id` is a skill-defined slug rather than a reference path), prefix `id` with `<from-sub-skill>:` to prevent collisions across sub-skills. Other finding fields are preserved.
### Agent self-review pass