mirror of
https://github.com/microsoft/BCQuality.git
synced 2026-10-05 22:56:55 +01:00
Merge current main into development guidance
Reconcile the read-only guidance output with the machine-readable skill index, adopt linked sample references required by bounded retrieval, and update the guidance regression fixture for the retrieval helper dependency. Permit only the known endpoint-DLP metadata stream during read-only evidence capture. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 638b66d2-9f06-4f60-8781-808709e1485c
This commit is contained in:
commit
8f025ac679
127 changed files with 5251 additions and 136 deletions
|
|
@ -2,7 +2,7 @@
|
|||
|
||||
The evaluation is convention-driven. The harness discovers every `<layer>/skills/review/al-<domain>-review.md` leaf across the enabled `microsoft`, `community`, and `custom` layers. Duplicate domains resolve with `custom > community > microsoft` precedence. For each selected leaf, the harness finds paired knowledge across the same layers, applies the same precedence to duplicate article slugs, selects the first article (by filename) with both `.bad.al` and `.good.al` companions, and derives the expected positive and clean control automatically. Adding a conforming leaf requires no scoring-contract edit.
|
||||
|
||||
`review-fixtures.json` contains only global thresholds and optional exceptional overrides. An override may select a different article or add context when the generic convention cannot express a scenario. It should remain empty in the normal case.
|
||||
`review-fixtures.json` contains only global thresholds and optional exceptional overrides. An override may select a different `article`, add context when the generic convention cannot express a scenario, or use an `articles` array when one domain needs explicit regression coverage for several paired articles. Specify either `article` or `articles`, not both. The first selected article retains the stable `<domain>-bad` and `<domain>-good` manifest IDs; additional articles use slug-qualified IDs. Overrides should remain empty in the normal case.
|
||||
|
||||
Model-facing preparation hashes case IDs, neutralizes `Good`/`Bad` object-name tokens, and removes full-line sample comments so neither the article slug, domain, nor expected outcome reveals the answer.
|
||||
|
||||
|
|
@ -26,7 +26,7 @@ This credential-free check proves every selected leaf maps to a same-named knowl
|
|||
|
||||
2. For a fast/small model, use one fresh invocation per `request-case-*.json`. Each request embeds the exact leaf instructions, that domain's candidate index rows with authoritative paths, and one opaque case. The model opens only matching articles and copies finding IDs from `candidateArticles[].path`. Save each response with the matching `result-case-*.json` name in the same directory.
|
||||
|
||||
`request-<domain>.json` files provide optional two-case leaf batches and identify the selected layer-owned skill path; save those as `result-<domain>.json`. Directory scoring prefers `result-case-*.json` when present and otherwise falls back to `result-*.json`. `review-request.json` is an optional all-domains stress test for larger models. Neither batch form is the preferred fast-model profile.
|
||||
`request-<domain>.json` files provide optional leaf batches containing every selected case for that domain and identify the selected layer-owned skill path; save those as `result-<domain>.json`. A normal convention-selected domain has one bad/good pair, while an `articles` override contributes one pair per listed article. Directory scoring prefers `result-case-*.json` when present and otherwise falls back to `result-*.json`. `review-request.json` is an optional all-domains stress test for larger models. Neither batch form is the preferred fast-model profile.
|
||||
|
||||
3. Save only this result shape:
|
||||
|
||||
|
|
|
|||
|
|
@ -17,11 +17,40 @@
|
|||
"article": "set-defaultimplementation-on-enum"
|
||||
},
|
||||
"performance": {
|
||||
"article": "use-isempty-for-existence-check"
|
||||
"articles": [
|
||||
"use-isempty-for-existence-check",
|
||||
"job-queue-category-code-serializes-conflicting-jobs",
|
||||
"job-queue-external-effects-must-be-idempotent",
|
||||
"job-queue-handlers-must-not-require-ui",
|
||||
"job-queue-handlers-must-propagate-failures",
|
||||
"job-queue-on-hold-does-not-stop-running-work",
|
||||
"store-scheduled-task-id-to-avoid-duplicate-tasks"
|
||||
]
|
||||
},
|
||||
"privacy": {
|
||||
"article": "no-pii-in-telemetry-message-string"
|
||||
},
|
||||
"query": {
|
||||
"articles": [
|
||||
"dataitemtablefilter-cannot-be-overwritten-at-runtime",
|
||||
"reopening-query-resets-cursor-but-keeps-filters",
|
||||
"set-query-filters-before-open",
|
||||
"setfilter-overwrites-query-columnfilter"
|
||||
]
|
||||
},
|
||||
"reporting": {
|
||||
"articles": [
|
||||
"clear-report-variable-before-independent-runmodal",
|
||||
"currreport-break-ends-the-current-trigger",
|
||||
"currreport-quit-rolls-back-and-skips-onpostreport",
|
||||
"currreport-skip-does-not-stop-trigger-code",
|
||||
"report-output-in-a-loop-needs-one-client-download",
|
||||
"reportextension-dataitem-trigger-order-is-explicit",
|
||||
"reportextension-report-triggers-run-after-base-triggers",
|
||||
"settableview-cannot-broaden-dataitemtableview",
|
||||
"stop-when-runrequestpage-returns-empty-parameters"
|
||||
]
|
||||
},
|
||||
"style": {
|
||||
"article": "label-comment-explains-placeholders"
|
||||
},
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue