Complete AL review knowledge readiness (#108)
Some checks failed
Validate knowledge index / validate-index (push) Has been cancelled
Validate AL review fixtures / validate-review-fixtures (push) Has been cancelled
Validate frontmatter and structure / validate (push) Has been cancelled

* Complete AL review knowledge readiness

Fill telemetry and Query coverage, strengthen thin review domains, correct audited content defects, and add deterministic cheap-model evaluation and reference-integrity safeguards.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 9825b012-e653-496a-9310-c1f4b6f8ac27

* Generalize review fixture discovery

Derive smoke cases from the leaf, domain, and paired-sample conventions so new leaves require no scoring-contract changes. Keep only exceptional selection/context overrides and fail when retrieval metadata cannot rank the selected article.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 9825b012-e653-496a-9310-c1f4b6f8ac27

* Preserve published field IDs in sample

Keep the existing Email and Contact Email field IDs unchanged, clarify that the sample represents an independent baseline, and use a local breaking-change rule for the generic smoke evaluation.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 9825b012-e653-496a-9310-c1f4b6f8ac27

* Clarify published field identity rules

State explicitly that a published field keeps its ID, name, and type while a replacement is added as a separate field under an unused ID.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 9825b012-e653-496a-9310-c1f4b6f8ac27

* Align field obsoletion sample baselines

Use Email field ID 3 as the shared baseline so the bad example demonstrates a same-ID rename while the good example retains the original field and adds a separate replacement.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 9825b012-e653-496a-9310-c1f4b6f8ac27

---------

Co-authored-by: Jesper Schulz-Wedde <jesper.schulzwedde@microsoft.com>
This commit is contained in:
Jesper Schulz-Wedde 2026-07-15 10:55:25 +02:00 committed by GitHub
parent ae04938c03
commit 186d8a1314
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
105 changed files with 2229 additions and 212 deletions

View file

@ -63,11 +63,19 @@ plugin-root environment variable, prefer it.
`microsoft/skills/review/al-code-review.md`. For each dispatched skill, read the
file and execute its Source → Relevance → Worklist → Action steps, reading
`PLUGIN_ROOT/skills/read.md` and `PLUGIN_ROOT/skills/do.md` on demand.
When `al-code-review` composes its leaves and the host supports child contexts or
separate model calls, run each leaf in an isolated context and roll up the returned
JSON. Pass each call the exact index rows for that leaf's domain so references can
be copied verbatim. This is the preferred execution profile for fast/small models;
do not force one generation to retain all domain knowledge at once.
4. **Emit findings.** Produce the rolled-up findings report in the DO output contract,
including each review finding's producer-supplied `domain` label (`outcome`,
`findings`, `references`, `confidence`, `suppressed`). Do not invent a different
shape; downstream consumers parse the DO contract without skill-specific logic.
Apply DO's reference-integrity gate before returning: every knowledge-backed path
must exist in the installed tree, must have been opened in full, and must be copied
verbatim. Never synthesize a plausible article slug.
If Entry returns `no-match` or `failed`, return the dispatch record unchanged so the
caller can log the reason.

View file

@ -171,6 +171,8 @@ Consumers that render output MAY treat agent findings differently from knowledge
**`findings[].message`** — human-readable explanation of the finding. Single short paragraph. No markdown formatting assumptions.
**Applicability is not a finding.** Loading an article into the worklist only means its rule must be evaluated. If the changed code does not violate the article's normative guidance, emit nothing for that article. An `info` finding still requires a concrete observation defined by the article; skills MUST NOT use `info` to list guidance that merely happened to be relevant.
**`findings[].location`** — optional. When present:
- `file` MUST be a repo-relative path using forward slashes.
@ -186,6 +188,15 @@ Findings without a `location` are permitted (for example, repository-wide observ
The first reference is the **primary** reference: the knowledge file the finding most directly cites. Additional references provide supporting context and are not ranked. `references` MAY be empty only for **agent findings** (see the `findings[].id` section above for the full encoding); any other finding MUST have at least one reference.
**Reference-integrity gate (mandatory).** A knowledge-backed finding may cite only a path copied verbatim from the current knowledge index or from a file discovered by the index fallback, and the skill must have opened that exact file in full before citing it. Never construct a plausible slug or infer a path from a topic name. Immediately before emitting the JSON document:
1. Verify every non-empty `references[].path` exists in the live checkout and was opened during this skill run.
2. Verify every citation-based `findings[].id` exactly equals `references[0].path`.
3. Remove any candidate that cannot satisfy both checks; it is not a knowledge-backed finding. Do not convert it into an agent finding merely to preserve it.
4. If reference integrity cannot be checked reliably, return `outcome: "failed"` rather than emitting fabricated or unverified citations.
This gate applies independently to every leaf result and again to a super-skill's rolled-up result.
**`findings[].confidence`** — the skill's confidence that the finding is a true positive, given the evidence it evaluated. Not applicability confidence, not severity confidence. Values: `high`, `medium`, `low`.
**`findings[].from-sub-skill`** — optional. Set only by super-skills. The `skill.id` of the sub-skill that produced the finding, or the literal string `"agent"` for an agent finding the super-skill produced from its own cross-cutting reasoning. Absent on findings emitted directly by a leaf skill — including agent findings the leaf emits within its own domain, which appear in the leaf's own report without this field.

View file

@ -85,7 +85,10 @@ Before opening a pull request:
- No fenced code blocks.
- File is under 100 lines.
- File covers one concern.
- Frontmatter `domain` exactly matches the containing domain folder.
- File is in the correct layer and domain folder.
- Name is kebab-case and descriptive.
- Every companion sample is referenced by filename from the article, and every referenced sample exists.
- Every review-leaf domain has at least one article with both `.good.al` and `.bad.al` companions; the evaluation harness derives positive and clean controls from that convention automatically.
Agents scaffolding new files SHOULD run this checklist programmatically before emitting the file.