mirror of
https://github.com/microsoft/BCQuality.git
synced 2026-10-06 07:06:54 +01:00
Today the super-skill is a pass-through: Roll up step 5 appends every
non-duplicate leaf finding unconditionally, and the only rejection path is
step 3 (mechanical sub-skill failure). The self-review pass validates the
super-skill's own candidates against knowledge, and an explicit clause
exempts leaf findings from that check. No role in the pipeline can veto a
leaf.
That shows up in the BC-Bench code-review numbers: on the parallel-leaf
baseline the agent generates ~630 comments to match ~100 expected ones
(precision 0.15-0.16), and a single leaf (al-appsource) accounts for 27-29%
of everything emitted while the gold set expects zero findings from it.
This arm adds an adjudication gate at roll-up time. The super-skill re-reads
the cited code and rejects findings whose precondition does not hold, that
the code does not bear out, that loaded knowledge explicitly permits, or
that name no actionable change. Rejections are recorded in a new
audit-only 'adjudicated-out[]' field; the engine ignores unknown report
keys, so nothing downstream changes shape.
The gate is deliberately conservative ('when ambiguous, keep the finding')
so the measured result is a lower bound on what arbitration can buy.
Branched from
|
||
|---|---|---|
| .. | ||
| al-appsource-review.md | ||
| al-breaking-changes-review.md | ||
| al-code-review.md | ||
| al-data-modeling-review.md | ||
| al-error-handling-review.md | ||
| al-events-review.md | ||
| al-interfaces-review.md | ||
| al-performance-review.md | ||
| al-privacy-review.md | ||
| al-query-review.md | ||
| al-security-review.md | ||
| al-style-review.md | ||
| al-telemetry-review.md | ||
| al-testing-review.md | ||
| al-ui-review.md | ||
| al-upgrade-review.md | ||
| al-web-services-review.md | ||