Complete AL review knowledge readiness

Fill telemetry and Query coverage, strengthen thin review domains, correct audited content defects, and add deterministic cheap-model evaluation and reference-integrity safeguards.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 9825b012-e653-496a-9310-c1f4b6f8ac27
This commit is contained in:
Jesper Schulz-Wedde 2026-07-15 07:19:15 +02:00
parent 809af9708e
commit e81632b4be
103 changed files with 2350 additions and 210 deletions

View file

@ -63,11 +63,18 @@ plugin-root environment variable, prefer it.
`microsoft/skills/review/al-code-review.md`. For each dispatched skill, read the
file and execute its Source → Relevance → Worklist → Action steps, reading
`PLUGIN_ROOT/skills/read.md` and `PLUGIN_ROOT/skills/do.md` on demand.
When `al-code-review` composes its leaves and the host supports child contexts or
separate model calls, run each leaf in an isolated context and roll up the returned
JSON. Pass each call the exact index rows for that leaf's domain so references can
be copied verbatim. This is the preferred execution profile for fast/small models;
do not force one generation to retain all domain knowledge at once.
4. **Emit findings.** Produce the rolled-up findings report in the DO output contract
(`outcome`, `findings`, `references`, `confidence`, `suppressed`). Do not invent a
different shape; downstream consumers parse the DO contract without skill-specific
logic.
logic. Apply DO's reference-integrity gate before returning: every knowledge-backed
path must exist in the installed tree, must have been opened in full, and must be
copied verbatim. Never synthesize a plausible article slug.
If Entry returns `no-match` or `failed`, return the dispatch record unchanged so the
caller can log the reason.

View file

@ -170,6 +170,8 @@ Consumers that render output MAY treat agent findings differently from knowledge
**`findings[].message`** — human-readable explanation of the finding. Single short paragraph. No markdown formatting assumptions.
**Applicability is not a finding.** Loading an article into the worklist only means its rule must be evaluated. If the changed code does not violate the article's normative guidance, emit nothing for that article. An `info` finding still requires a concrete observation defined by the article; skills MUST NOT use `info` to list guidance that merely happened to be relevant.
**`findings[].location`** — optional. When present:
- `file` MUST be a repo-relative path using forward slashes.
@ -185,6 +187,15 @@ Findings without a `location` are permitted (for example, repository-wide observ
The first reference is the **primary** reference: the knowledge file the finding most directly cites. Additional references provide supporting context and are not ranked. `references` MAY be empty only for **agent findings** (see the `findings[].id` section above for the full encoding); any other finding MUST have at least one reference.
**Reference-integrity gate (mandatory).** A knowledge-backed finding may cite only a path copied verbatim from the current knowledge index or from a file discovered by the index fallback, and the skill must have opened that exact file in full before citing it. Never construct a plausible slug or infer a path from a topic name. Immediately before emitting the JSON document:
1. Verify every non-empty `references[].path` exists in the live checkout and was opened during this skill run.
2. Verify every citation-based `findings[].id` exactly equals `references[0].path`.
3. Remove any candidate that cannot satisfy both checks; it is not a knowledge-backed finding. Do not convert it into an agent finding merely to preserve it.
4. If reference integrity cannot be checked reliably, return `outcome: "failed"` rather than emitting fabricated or unverified citations.
This gate applies independently to every leaf result and again to a super-skill's rolled-up result.
**`findings[].confidence`** — the skill's confidence that the finding is a true positive, given the evidence it evaluated. Not applicability confidence, not severity confidence. Values: `high`, `medium`, `low`.
**`findings[].from-sub-skill`** — optional. Set only by super-skills. The `skill.id` of the sub-skill that produced the finding, or the literal string `"agent"` for an agent finding the super-skill produced from its own cross-cutting reasoning. Absent on findings emitted directly by a leaf skill — including agent findings the leaf emits within its own domain, which appear in the leaf's own report without this field.
@ -288,5 +299,3 @@ Conforms to the DO output contract.
## How orchestrators consume output
An orchestrator invokes an action skill with an input appropriate to the skill's declared `inputs`, receives the JSON output, and maps findings to its delivery surface (PR comments, build gates, IDE diagnostics). The orchestrator MUST NOT interpret skill-specific fields beyond the schema above. Skills that need richer semantics MUST encode them within the schema (for example, by adding structured `message` text) rather than extending the output shape.

View file

@ -85,7 +85,10 @@ Before opening a pull request:
- No fenced code blocks.
- File is under 100 lines.
- File covers one concern.
- Frontmatter `domain` exactly matches the containing domain folder.
- File is in the correct layer and domain folder.
- Name is kebab-case and descriptive.
- Every companion sample is referenced by filename from the article, and every referenced sample exists.
- A changed review domain has a positive and clean control in `evaluation/review-fixtures.json`.
Agents scaffolding new files SHOULD run this checklist programmatically before emitting the file.