Add article-level leaf isolation to the AL review super-skill

Experiment branch. Redefines the unit of isolation in the Action step
from the sub-skill (domain) to the individual knowledge article: after a
sub-skill narrows its slice via Source/Relevance/Worklist, it dispatches
one isolated child agent per surviving article instead of evaluating the
whole slice in one pass.

Sub-skill ordering is unchanged, so this isolates granularity as the
only variable against the serial per-domain arm. Adds an
[article <domain>/<slug>: findings=<M>] progress marker so the
orchestrator can verify article-level isolation actually occurred.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e33a7b96-4821-4797-9f72-96880f3c6412
This commit is contained in:
wenjiefan 2026-09-03 09:42:26 +02:00
parent 712dee9ec1
commit 9d911aa206

View file

@ -66,10 +66,12 @@ The worklist is the list of sub-skills judged relevant by the previous step. Eve
The Action step is a sequence of **discrete iterations**, not one combined generation. The contract requires the super-skill to invoke each sub-skill in turn and then perform a self-review pass. Concretely this means: The Action step is a sequence of **discrete iterations**, not one combined generation. The contract requires the super-skill to invoke each sub-skill in turn and then perform a self-review pass. Concretely this means:
- **Isolate leaf invocations when the host supports it.** For fast/small models, each sub-skill SHOULD run in a fresh model call or child context containing only the task input, READ/DO contracts, the leaf instructions, a domain-filtered slice of the current knowledge index, and articles that leaf worklists. Preserve each index row's exact `path`; the leaf must copy references from that slice. The coordinator then collects the resulting JSON. This is the preferred fast-model profile: it bounds context, prevents later leaves from being skipped as attention is exhausted, and removes any reason to synthesize article paths. - **Isolate leaf invocations when the host supports it.** For fast/small models, each sub-skill SHOULD run in a fresh model call or child context containing only the task input, READ/DO contracts, the leaf instructions, a domain-filtered slice of the current knowledge index, and articles that leaf worklists. Preserve each index row's exact `path`; the leaf must copy references from that slice. The coordinator then collects the resulting JSON. This is the preferred fast-model profile: it bounds context, prevents later leaves from being skipped as attention is exhausted, and removes any reason to synthesize article paths.
- **Isolate one pass per knowledge article.** Inside a sub-skill, do not evaluate its whole knowledge slice in one combined reasoning step. Once the sub-skill's Source → Relevance → Worklist steps have narrowed its slice to the articles that actually apply to the changed files, dispatch **one isolated child agent per article that survived that narrowing**. Each child receives only the task input, the READ/DO contracts, that single article (its index row with `path` preserved verbatim, plus the article body), and the changed-file list. It evaluates that one article against the diff and returns only its own findings. The sub-skill concatenates its children's findings into its findings-report and applies its normal duplicate merge. Dispatch a sub-skill's article children concurrently when the host supports it; when it does not, evaluate the articles one at a time in sequence. Either way, do not begin the next sub-skill until every article child of the current one has returned.
- **Emit an article progress marker.** Immediately after each article child returns, emit exactly one line: `[article <domain>/<slug>: findings=<M>]`. This marker is a hard contract — it is how the orchestrator verifies that article-level isolation actually happened.
- Treat each sub-skill in the worklist as its own pass: read the sub-skill's instructions, apply its Source → Relevance → Worklist → Action steps to the orchestrator-supplied inputs, and produce that sub-skill's complete findings-report before moving on. - Treat each sub-skill in the worklist as its own pass: read the sub-skill's instructions, apply its Source → Relevance → Worklist → Action steps to the orchestrator-supplied inputs, and produce that sub-skill's complete findings-report before moving on.
- Do not collapse multiple sub-skills into one shared reasoning step. Each sub-skill has a distinct knowledge subset and a distinct evaluation procedure; sharing one rolled-up scan dilutes per-skill attention and causes leaves to silently underreport (this has been observed in production: leaf skills returned empty `findings[]` while their standalone runs against the same diff produced multiple matches). - Do not collapse multiple sub-skills into one shared reasoning step. Each sub-skill has a distinct knowledge subset and a distinct evaluation procedure; sharing one rolled-up scan dilutes per-skill attention and causes leaves to silently underreport (this has been observed in production: leaf skills returned empty `findings[]` while their standalone runs against the same diff produced multiple matches).
- The agent self-review pass is its own final iteration. Begin it only after every sub-skill in the worklist has completed and its sub-result is recorded. - The agent self-review pass is its own final iteration. Begin it only after every sub-skill in the worklist has completed and its sub-result is recorded.
- Sub-skills are independent: re-walking the diff once per sub-skill is correct and expected. The output schema accommodates this — `sub-results` carries one entry per sub-skill, each a complete findings-report. - Sub-skills are independent: re-walking the diff once per sub-skill is correct and expected. Article children are independent for the same reason: re-walking the diff once per article is also correct and expected, and its cost is never a reason to merge articles into a shared pass. The output schema accommodates this — `sub-results` carries one entry per sub-skill, each a complete findings-report.
- When isolated calls are unavailable and the current model cannot finish every leaf within its budget, return `partial` with completed `sub-results` and name the first unevaluated sub-skill in `outcome-reason`. Never silently mark the remaining leaves clean. - When isolated calls are unavailable and the current model cannot finish every leaf within its budget, return `partial` with completed `sub-results` and name the first unevaluated sub-skill in `outcome-reason`. Never silently mark the remaining leaves clean.
### Roll up sub-skill findings ### Roll up sub-skill findings