Revert evaluation harness change for multiple articles per domain

The harness intentionally evaluates one paired article per domain. Keep it
as designed; how the retention pairs join privacy evaluation is left to the
maintainers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
Jeremy Vyska 2026-09-30 15:49:13 +02:00
parent 07a171386d
commit f9333addcd
3 changed files with 15 additions and 44 deletions

View file

@ -2,7 +2,7 @@
The evaluation is convention-driven. The harness discovers every `<layer>/skills/review/al-<domain>-review.md` leaf across the enabled `microsoft`, `community`, and `custom` layers. Duplicate domains resolve with `custom > community > microsoft` precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both `.bad.al` and `.good.al` companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.
`review-fixtures.json` contains only global thresholds and optional exceptional overrides. An override may select a different article, add context, or list `additionalArticles` whose sample pairs become extra positive and clean cases for that domain, when the generic convention cannot express a scenario. It should remain empty in the normal case.
`review-fixtures.json` contains only global thresholds and optional exceptional overrides. An override may select a different article or add context when the generic convention cannot express a scenario. It should remain empty in the normal case.
Model-facing preparation hashes case IDs, neutralizes `Good`/`Bad` object-name tokens, and removes full-line sample comments so neither the article slug, domain, nor expected outcome reveals the answer.