From 5aaa58e8ee38848eca70df4287644d7c387201df Mon Sep 17 00:00:00 2001 From: Jesper Schulz-Wedde Date: Fri, 17 Apr 2026 13:06:03 +0200 Subject: [PATCH] Introduce super-skill composition; refactor al-code-review into super + two leaves DO contract (skills/do.md) - New 'sub-skills' optional frontmatter field on action skills: when present and non-empty, the skill is a super-skill that composes other action skills. - New 'Composition (super-skills)' section covering section interpretation, outcome rollup, summary aggregation, and suppression scope. - Output schema gains three optional fields: 'from-sub-skill' on each finding, top-level 'sub-results[]' carrying nested findings-reports, and top-level 'skipped-sub-skills[]'. - Super-skills MUST NOT filter sub-skills by task content; leaves own task-level applicability and signal via outcome. - Findings from a failed sub-skill MUST NOT flow into the parent's findings[] or counts, consistent with DO's rule that consumers ignore a failed skill's findings. Reports are still preserved in sub-results[] for traceability. - Rolled-up non-citation finding ids MUST be prefixed with the sub- skill id to prevent collisions across sub-skills. Citation-based ids are already unique via repo path and are not rewritten. - Outcome rollup rules updated: 'partial' covers S = {partial}, {partial, partial}, and {partial, failed}. Empty worklist rolls up to 'not-applicable' with outcome-reason. - Nested super-skills are not permitted in v1. Reference skills (microsoft/skills/) - al-code-review.md rewritten as the canonical super-skill: lists al-performance-review and al-security-review as sub-skills, orch- estrates invocation, aggregates output, and includes a worked rolled-up JSON example plus the empty-corpus rollup. - al-performance-review.md added as a leaf reference skill for the performance knowledge domain. - al-security-review.md added as a leaf reference skill for the security knowledge domain. - Both leaves retain the leaf-level rules validated in the prior pass: partial-context message requirement, worklist-scoped suppression, application-area semantics, and the platform-guarantee threshold for blocker severity. README updated to describe leaf vs super-skill and link all three reference skills. Two rubber-duck passes tightened the contract and caught schema violations in the worked examples before commit. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> --- README.md | 2 +- microsoft/skills/al-code-review.md | 209 ++++++++++++++++------ microsoft/skills/al-performance-review.md | 126 +++++++++++++ microsoft/skills/al-security-review.md | 130 ++++++++++++++ skills/do.md | 60 ++++++- 5 files changed, 470 insertions(+), 57 deletions(-) create mode 100644 microsoft/skills/al-performance-review.md create mode 100644 microsoft/skills/al-security-review.md diff --git a/README.md b/README.md index d659f60..f48e262 100644 --- a/README.md +++ b/README.md @@ -36,7 +36,7 @@ Skills define how agents consume knowledge. They come in two flavors: Schema + Use and New Knowledge are deliberately separate: one is the reader's contract, the other is the writer's guide. New Knowledge depends on Schema + Use but does not duplicate it. -- **Action skills** — concrete skills that follow the Action Skill template to do real work (review code, audit telemetry, etc.). Action skills live inside the layers that own them (`/microsoft/skills/`, `/community/skills/`). The canonical reference implementation is [`microsoft/skills/al-code-review.md`](microsoft/skills/al-code-review.md) — skill authors should use it as a starting point. +- **Action skills** — concrete skills that follow the Action Skill template to do real work (review code, audit telemetry, etc.). Action skills live inside the layers that own them (`/microsoft/skills/`, `/community/skills/`). An action skill is either a **leaf** that evaluates knowledge files directly, or a **super-skill** that composes other action skills (declared via `sub-skills` in frontmatter). The canonical reference is [`microsoft/skills/al-code-review.md`](microsoft/skills/al-code-review.md) (super-skill), which composes [`microsoft/skills/al-performance-review.md`](microsoft/skills/al-performance-review.md) and [`microsoft/skills/al-security-review.md`](microsoft/skills/al-security-review.md) (leaves). ### Agent bootstrapping diff --git a/microsoft/skills/al-code-review.md b/microsoft/skills/al-code-review.md index b965a5c..4f27782 100644 --- a/microsoft/skills/al-code-review.md +++ b/microsoft/skills/al-code-review.md @@ -3,83 +3,79 @@ kind: action-skill id: al-code-review version: 1 title: AL code review -description: Reviews AL source changes against performance, security, UX, telemetry, and testing guidance from BCQuality. +description: Reviews AL source changes by composing the AL review leaf skills (performance, security, ...). inputs: [pr-diff, file-path] outputs: [findings-report] bc-version: [26..28] technologies: [al] countries: [w1] application-area: [all] +sub-skills: + - microsoft/skills/al-performance-review.md + - microsoft/skills/al-security-review.md --- # AL code review -Reviews AL source changes against applicable BCQuality guidance and emits a findings report. This skill is the canonical reference implementation of the DO contract — skill authors should copy its structure. +Reviews AL source changes by composing the leaf AL review skills. This is the canonical reference implementation of a **super-skill** — skill authors writing composed reviews should copy its structure. -An orchestrator invokes this skill with either a `pr-diff` (the standard PR-review entry point) or a `file-path` (single-file review, typically from an IDE). The skill produces a single JSON document conforming to the DO output contract. +`al-code-review` does not evaluate knowledge files directly. It invokes each of its sub-skills against the same task input, collects their findings-reports, and returns a rolled-up findings-report. + +An orchestrator invokes this skill with either a `pr-diff` (the standard PR-review entry point) or a `file-path` (single-file review). The skill produces a single JSON document conforming to the DO output contract, extended with `sub-results` and — when applicable — `skipped-sub-skills`. ## Source -Collect all knowledge files under `*/knowledge/**/*.md`, across every enabled layer (`/microsoft/`, `/community/`, `/custom/`). Files in any domain subfolder are included; the skill does not enumerate domains. Relevance trims the result to the subset that applies. +The sub-skills invoked by this skill are those listed in frontmatter `sub-skills`: + +- `microsoft/skills/al-performance-review.md` +- `microsoft/skills/al-security-review.md` + +Additional leaf skills (for example, UX, telemetry, testing) are added by updating the `sub-skills` list. The skill does not discover sub-skills implicitly. ## Relevance -Apply the frontmatter matching rules defined in READ (*Frontmatter matching semantics*) against the task context: +A sub-skill is relevant when both of the following hold: -- `bc-version` — the target BC version from the PR branch's `app.json` or the orchestrator-supplied version. If unavailable, the dimension is `unknown` (see READ's partial-context rule). -- `technologies` — `[al]`. -- `countries` — the countries declared in the consuming app's `app.json`. Default to the orchestrator's configured context; if absent, `unknown`. -- `application-area` — the union of application areas declared by the changed objects. Pass the actual set (e.g., `[finance, jobs]` for a PR touching both); do not substitute `[all]`. If the area cannot be determined from the changes, the dimension is `unknown`. +- The orchestrator has supplied inputs that satisfy the sub-skill's declared `inputs`. +- The orchestrator has not disabled the sub-skill via configuration. -Discard files that are not applicable. Retain conditionally applicable files (any dimension `unknown`) only when the orchestrator's configuration permits them; findings derived from those files MUST have `confidence` no higher than `medium`, AND the finding's `message` MUST name the dimension or dimensions that were unknown so reviewers can judge the finding's applicability. +Per the DO contract, the super-skill MUST NOT filter sub-skills by task content. `al-code-review` does not inspect the PR diff to predict whether, for example, there is anything for `al-security-review` to find. Each leaf is responsible for its own task-level applicability decision; leaves signal non-applicability by returning `outcome: "not-applicable"` or `outcome: "no-knowledge"`. + +Sub-skills that fail either check are not invoked and are recorded in `skipped-sub-skills`: + +- `reason: "configuration"` when the orchestrator disabled the sub-skill. +- `reason: "not-applicable"` when the orchestrator's inputs do not satisfy the sub-skill's declared `inputs`. ## Worklist -Narrow the relevant files to the subset that applies to the changes under review. For each relevant file, compute overlap against: - -- The changed AL object names and types (tables, pages, codeunits, reports, queries, xmlports, enums, permission sets). -- The changed procedures, triggers, and fields. -- Tokens extracted from the diff (identifier names, referenced objects, keywords). - -A file enters the candidate worklist when its `keywords` intersect the extracted tokens or its topic (derived from filename and Description) matches a changed object type. - -Once the candidate worklist is known, resolve layer-precedence conflicts per READ: for any two candidates whose normative guidance (`## Best Practice` or `## Anti Pattern`) directly contradicts, keep the file from the higher-precedence layer and drop the other. Every dropped file MUST be recorded in the output `suppressed` array with `reason: "layer-precedence"`. Files that would have been candidates but are hidden because their layer is disabled in consumer configuration MUST be recorded with `reason: "configuration"`. Files that were never candidates (failed Relevance or did not match task signal) are NOT recorded in `suppressed`. - -When the post-conflict worklist is empty because no applicable knowledge exists in the repo, or because configuration suppressed every candidate, emit `outcome: "no-knowledge"`. When the worklist is empty because no applicable knowledge matched the changes, emit `outcome: "completed"` with an empty `findings` array. +The worklist is the list of sub-skills judged relevant by the previous step. Every sub-skill in the worklist will be invoked in the Action step. ## Action -For each worklist entry, evaluate the diff against the file's `## Best Practice` and `## Anti Pattern` sections. Emit findings as follows: +For each sub-skill in the worklist: -- When the diff contains a clear match for an Anti Pattern, emit a finding with severity `major` or `blocker`, the message summarizing the anti-pattern, `location` pointing to the offending line or range, and a `references` entry pointing to the knowledge file. Use `blocker` only when the knowledge file states the anti-pattern violates a platform-level guarantee; when the file does not make that claim, the ceiling is `major`. -- When the diff contains code that contradicts a Best Practice without being a full anti-pattern, emit `minor` with the same reference shape. -- When the skill cannot detect a violation but the file is clearly applicable to the change, emit `info` citing the file so the author is nudged to read it. Repository-wide observations MAY omit `location`. +1. Invoke the sub-skill with the orchestrator's inputs, passing only the subset each sub-skill declares in its `inputs`. +2. Capture the sub-skill's complete findings-report verbatim and append it to `sub-results`. +3. If the sub-skill's `outcome` is `failed`, stop here for this sub-skill: its findings are not reliable per the DO contract and MUST NOT be copied into the super-skill's top-level `findings[]` or counted in `summary.counts`. +4. Otherwise, append each entry from the sub-skill's `findings[]` to the super-skill's top-level `findings[]`, setting `from-sub-skill` to the sub-skill's `skill.id`. For non-citation findings (those whose `id` is a skill-defined slug rather than a reference path), prefix `id` with `:` to prevent collisions across sub-skills. Other finding fields are preserved. -Set `confidence` to: +Aggregate `summary.counts` and `summary.coverage` as the sums across invoked sub-skills whose `outcome` is not `failed`. -- `high` when the detection is based on an unambiguous pattern match (identifier, syntax, object type). -- `medium` when detection relies on heuristics (name similarity, scope inference) or when any frontmatter dimension was `unknown`. -- `low` when the finding is an advisory derived only from applicability (no detection signal). +`suppressed[]` at the super-skill level remains empty. Knowledge-file-level suppression is reported by each sub-skill within its own entry in `sub-results`. -The outcome selection: - -- `completed` — the skill evaluated every worklist item. Default when the skill finishes normally, including when the resulting `findings` array is empty. -- `no-knowledge` — no applicable knowledge survived Source, Relevance, configuration filtering, and conflict resolution. `findings` is empty. -- `not-applicable` — the task context lacks an AL dimension (no AL changes in the diff, or `technologies` filter rejected the task). -- `partial` — a time or token budget was hit before the worklist was exhausted. `summary.coverage` reflects the evaluated subset; `outcome-reason` explains the cause. -- `failed` — an unrecoverable error occurred. `outcome-reason` is required. +Derive `outcome` using the DO rollup rules. `outcome-reason` is populated for `partial` and `failed` and SHOULD summarize per-sub-skill state, for example: *"al-security-review failed (tool timeout); al-performance-review completed."* ## Output -Output conforms to the DO output contract. A populated example: +Output conforms to the DO output contract, extended with `sub-results` and `skipped-sub-skills`. A populated example — both leaves ran, each produced findings: ```json { "skill": { "id": "al-code-review", "version": 1 }, "outcome": "completed", "summary": { - "counts": { "blocker": 0, "major": 1, "minor": 1, "info": 1 }, - "coverage": { "worklist-size": 3, "items-evaluated": 3 } + "counts": { "blocker": 1, "major": 1, "minor": 1, "info": 1 }, + "coverage": { "worklist-size": 4, "items-evaluated": 4 } }, "findings": [ { @@ -94,7 +90,33 @@ Output conforms to the DO output contract. A populated example: "references": [ { "path": "microsoft/knowledge/performance/filter-before-find.md" } ], - "confidence": "high" + "confidence": "high", + "from-sub-skill": "al-performance-review" + }, + { + "id": "community/knowledge/performance/use-setloadfields.md", + "severity": "info", + "message": "Posting routine iterates ledger entries; consider whether SetLoadFields applies per the linked guidance.", + "references": [ + { "path": "community/knowledge/performance/use-setloadfields.md" } + ], + "confidence": "low", + "from-sub-skill": "al-performance-review" + }, + { + "id": "microsoft/knowledge/security/no-plaintext-secrets-in-telemetry.md", + "severity": "blocker", + "message": "A bearer token is passed to Session.LogMessage as part of the CustomDimensions payload. The referenced guidance documents this as a platform-level data-protection violation.", + "location": { + "file": "src/Integration/ApiClient.Codeunit.al", + "line": 85, + "range": { "start-line": 85, "end-line": 89 } + }, + "references": [ + { "path": "microsoft/knowledge/security/no-plaintext-secrets-in-telemetry.md" } + ], + "confidence": "high", + "from-sub-skill": "al-security-review" }, { "id": "microsoft/knowledge/security/avoid-implicit-commit.md", @@ -107,28 +129,89 @@ Output conforms to the DO output contract. A populated example: "references": [ { "path": "microsoft/knowledge/security/avoid-implicit-commit.md" } ], - "confidence": "medium" - }, - { - "id": "community/knowledge/telemetry/log-posting-failures.md", - "severity": "info", - "message": "Posting routine touched; consider whether failure paths emit telemetry per the linked guidance.", - "references": [ - { "path": "community/knowledge/telemetry/log-posting-failures.md" } - ], - "confidence": "low" + "confidence": "medium", + "from-sub-skill": "al-security-review" } ], - "suppressed": [ + "suppressed": [], + "sub-results": [ { - "reference": { "path": "community/knowledge/performance/filter-before-find.md" }, - "reason": "layer-precedence" + "skill": { "id": "al-performance-review", "version": 1 }, + "outcome": "completed", + "summary": { + "counts": { "blocker": 0, "major": 1, "minor": 0, "info": 1 }, + "coverage": { "worklist-size": 2, "items-evaluated": 2 } + }, + "findings": [ + { + "id": "microsoft/knowledge/performance/filter-before-find.md", + "severity": "major", + "message": "FindSet is called on a record variable without any prior SetRange/SetFilter. This forces a full-table scan.", + "location": { + "file": "src/Sales/PostingRoutines.Codeunit.al", + "line": 140, + "range": { "start-line": 140, "end-line": 144 } + }, + "references": [ + { "path": "microsoft/knowledge/performance/filter-before-find.md" } + ], + "confidence": "high" + }, + { + "id": "community/knowledge/performance/use-setloadfields.md", + "severity": "info", + "message": "Posting routine iterates ledger entries; consider whether SetLoadFields applies per the linked guidance.", + "references": [ + { "path": "community/knowledge/performance/use-setloadfields.md" } + ], + "confidence": "low" + } + ], + "suppressed": [] + }, + { + "skill": { "id": "al-security-review", "version": 1 }, + "outcome": "completed", + "summary": { + "counts": { "blocker": 1, "major": 0, "minor": 1, "info": 0 }, + "coverage": { "worklist-size": 2, "items-evaluated": 2 } + }, + "findings": [ + { + "id": "microsoft/knowledge/security/no-plaintext-secrets-in-telemetry.md", + "severity": "blocker", + "message": "A bearer token is passed to Session.LogMessage as part of the CustomDimensions payload. The referenced guidance documents this as a platform-level data-protection violation.", + "location": { + "file": "src/Integration/ApiClient.Codeunit.al", + "line": 85, + "range": { "start-line": 85, "end-line": 89 } + }, + "references": [ + { "path": "microsoft/knowledge/security/no-plaintext-secrets-in-telemetry.md" } + ], + "confidence": "high" + }, + { + "id": "microsoft/knowledge/security/avoid-implicit-commit.md", + "severity": "minor", + "message": "An explicit COMMIT inside a posting routine may leave the ledger in an inconsistent state if subsequent steps fail.", + "location": { + "file": "src/Sales/PostingRoutines.Codeunit.al", + "line": 201 + }, + "references": [ + { "path": "microsoft/knowledge/security/avoid-implicit-commit.md" } + ], + "confidence": "medium" + } + ], + "suppressed": [] } ] } ``` -The empty-corpus case — BCQuality's state until knowledge files land — produces: +The empty-corpus case — BCQuality's state until knowledge files land — rolls up to `no-knowledge`: ```json { @@ -139,6 +222,22 @@ The empty-corpus case — BCQuality's state until knowledge files land — produ "coverage": { "worklist-size": 0, "items-evaluated": 0 } }, "findings": [], - "suppressed": [] + "suppressed": [], + "sub-results": [ + { + "skill": { "id": "al-performance-review", "version": 1 }, + "outcome": "no-knowledge", + "summary": { "counts": { "blocker": 0, "major": 0, "minor": 0, "info": 0 }, "coverage": { "worklist-size": 0, "items-evaluated": 0 } }, + "findings": [], + "suppressed": [] + }, + { + "skill": { "id": "al-security-review", "version": 1 }, + "outcome": "no-knowledge", + "summary": { "counts": { "blocker": 0, "major": 0, "minor": 0, "info": 0 }, "coverage": { "worklist-size": 0, "items-evaluated": 0 } }, + "findings": [], + "suppressed": [] + } + ] } ``` diff --git a/microsoft/skills/al-performance-review.md b/microsoft/skills/al-performance-review.md new file mode 100644 index 0000000..4b52d21 --- /dev/null +++ b/microsoft/skills/al-performance-review.md @@ -0,0 +1,126 @@ +--- +kind: action-skill +id: al-performance-review +version: 1 +title: AL performance review +description: Reviews AL source changes against performance guidance from BCQuality. +inputs: [pr-diff, file-path] +outputs: [findings-report] +bc-version: [26..28] +technologies: [al] +countries: [w1] +application-area: [all] +--- + +# AL performance review + +Reviews AL source changes against the `performance` knowledge domain in BCQuality and emits a findings report. This is a leaf action skill: it invokes no sub-skills. It is one of the skills composed by `al-code-review`. + +An orchestrator invokes this skill with either a `pr-diff` (the standard PR-review entry point) or a `file-path` (single-file review). The skill produces a single JSON document conforming to the DO output contract. + +## Source + +Collect all knowledge files under `*/knowledge/performance/**/*.md`, across every enabled layer (`/microsoft/`, `/community/`, `/custom/`). Relevance trims the result to the subset that applies. + +## Relevance + +Apply the frontmatter matching rules defined in READ (*Frontmatter matching semantics*) against the task context: + +- `bc-version` — the target BC version from the PR branch's `app.json` or the orchestrator-supplied version. If unavailable, the dimension is `unknown`. +- `technologies` — `[al]`. +- `countries` — the countries declared in the consuming app's `app.json`. Default to the orchestrator's configured context; if absent, `unknown`. +- `application-area` — the union of application areas declared by the changed objects. Pass the actual set; do not substitute `[all]`. If the area cannot be determined from the changes, the dimension is `unknown`. + +Discard files that are not applicable. Retain conditionally applicable files (any dimension `unknown`) only when the orchestrator's configuration permits them; findings derived from those files MUST have `confidence` no higher than `medium`, AND the finding's `message` MUST name the dimension or dimensions that were unknown. + +## Worklist + +Narrow the relevant files to the subset that applies to the changes under review. For each relevant file, compute overlap against: + +- The changed AL object names and types — especially tables, pages with SourceTable bindings, reports, queries, and codeunits performing record iteration. +- The changed procedures and triggers, weighted toward those that perform loops, Find/FindSet/FindFirst calls, CalcFields, CalcSums, FlowField access, or cross-table navigation. +- Tokens extracted from the diff that relate to data access (SetRange, SetFilter, SetLoadFields, SetCurrentKey, FindSet, Repeat…Until, CalcFields, CalcSums). + +A file enters the candidate worklist when its `keywords` intersect the extracted tokens or its topic (derived from filename and Description) matches a changed object type. + +Once the candidate worklist is known, resolve layer-precedence conflicts per READ. Drop lower-precedence files whose normative guidance (`## Best Practice` or `## Anti Pattern`) directly contradicts a higher-precedence candidate, and record each dropped file in `suppressed` with `reason: "layer-precedence"`. Files that would have been candidates but are hidden because their layer is disabled in consumer configuration are recorded with `reason: "configuration"`. Files that never became candidates are NOT recorded in `suppressed`. + +When the post-conflict worklist is empty because no applicable performance knowledge exists, or because configuration suppressed every candidate, emit `outcome: "no-knowledge"`. When the worklist is empty because no applicable performance knowledge matched the changes, emit `outcome: "completed"` with an empty `findings` array. + +## Action + +For each worklist entry, evaluate the diff against the file's `## Best Practice` and `## Anti Pattern` sections. Emit findings as follows: + +- When the diff contains a clear match for an Anti Pattern, emit a finding with severity `major` or `blocker`, a message summarizing the anti-pattern, `location` pointing to the offending line or range, and a `references` entry pointing to the knowledge file. Use `blocker` only when the knowledge file states the anti-pattern violates a platform-level guarantee (for example, documented query timeouts or transaction size limits). When the file does not make such a claim, the ceiling is `major`. +- When the diff contains code that contradicts a Best Practice without being a full anti-pattern, emit `minor` with the same reference shape. +- When the skill cannot detect a violation but the file is clearly applicable to the change, emit `info` citing the file. Repository-wide observations MAY omit `location`. + +Set `confidence` to: + +- `high` when the detection is based on an unambiguous pattern match (identifier, syntax, object type). +- `medium` when detection relies on heuristics or when any frontmatter dimension was `unknown`. +- `low` when the finding is an advisory derived only from applicability. + +Outcome selection: + +- `completed` — the skill evaluated every worklist item; default when the skill finishes normally, including when the resulting `findings` array is empty. +- `no-knowledge` — no applicable performance knowledge survived Source, Relevance, configuration filtering, and conflict resolution. `findings` is empty. +- `not-applicable` — the task context lacks an AL dimension (no AL changes in the diff, or `technologies` filter rejected the task). +- `partial` — a time or token budget was hit before the worklist was exhausted. `summary.coverage` reflects the evaluated subset; `outcome-reason` explains the cause. +- `failed` — an unrecoverable error occurred. `outcome-reason` is required. + +## Output + +Output conforms to the DO output contract. A populated example: + +```json +{ + "skill": { "id": "al-performance-review", "version": 1 }, + "outcome": "completed", + "summary": { + "counts": { "blocker": 0, "major": 1, "minor": 0, "info": 1 }, + "coverage": { "worklist-size": 2, "items-evaluated": 2 } + }, + "findings": [ + { + "id": "microsoft/knowledge/performance/filter-before-find.md", + "severity": "major", + "message": "FindSet is called on a record variable without any prior SetRange/SetFilter. This forces a full-table scan.", + "location": { + "file": "src/Sales/PostingRoutines.Codeunit.al", + "line": 140, + "range": { "start-line": 140, "end-line": 144 } + }, + "references": [ + { "path": "microsoft/knowledge/performance/filter-before-find.md" } + ], + "confidence": "high" + }, + { + "id": "community/knowledge/performance/use-setloadfields.md", + "severity": "info", + "message": "Posting routine iterates ledger entries; consider whether SetLoadFields applies per the linked guidance.", + "references": [ + { "path": "community/knowledge/performance/use-setloadfields.md" } + ], + "confidence": "low" + } + ], + "suppressed": [] +} +``` + +The empty-corpus case — BCQuality's state until performance knowledge files land — produces: + +```json +{ + "skill": { "id": "al-performance-review", "version": 1 }, + "outcome": "no-knowledge", + "summary": { + "counts": { "blocker": 0, "major": 0, "minor": 0, "info": 0 }, + "coverage": { "worklist-size": 0, "items-evaluated": 0 } + }, + "findings": [], + "suppressed": [] +} +``` diff --git a/microsoft/skills/al-security-review.md b/microsoft/skills/al-security-review.md new file mode 100644 index 0000000..cecfbd5 --- /dev/null +++ b/microsoft/skills/al-security-review.md @@ -0,0 +1,130 @@ +--- +kind: action-skill +id: al-security-review +version: 1 +title: AL security review +description: Reviews AL source changes against security guidance from BCQuality. +inputs: [pr-diff, file-path] +outputs: [findings-report] +bc-version: [26..28] +technologies: [al] +countries: [w1] +application-area: [all] +--- + +# AL security review + +Reviews AL source changes against the `security` knowledge domain in BCQuality and emits a findings report. This is a leaf action skill: it invokes no sub-skills. It is one of the skills composed by `al-code-review`. + +An orchestrator invokes this skill with either a `pr-diff` (the standard PR-review entry point) or a `file-path` (single-file review). The skill produces a single JSON document conforming to the DO output contract. + +## Source + +Collect all knowledge files under `*/knowledge/security/**/*.md`, across every enabled layer (`/microsoft/`, `/community/`, `/custom/`). Relevance trims the result to the subset that applies. + +## Relevance + +Apply the frontmatter matching rules defined in READ (*Frontmatter matching semantics*) against the task context: + +- `bc-version` — the target BC version from the PR branch's `app.json` or the orchestrator-supplied version. If unavailable, the dimension is `unknown`. +- `technologies` — `[al]`. +- `countries` — the countries declared in the consuming app's `app.json`. Default to the orchestrator's configured context; if absent, `unknown`. +- `application-area` — the union of application areas declared by the changed objects. Pass the actual set; do not substitute `[all]`. If the area cannot be determined from the changes, the dimension is `unknown`. + +Discard files that are not applicable. Retain conditionally applicable files (any dimension `unknown`) only when the orchestrator's configuration permits them; findings derived from those files MUST have `confidence` no higher than `medium`, AND the finding's `message` MUST name the dimension or dimensions that were unknown. + +## Worklist + +Narrow the relevant files to the subset that applies to the changes under review. For each relevant file, compute overlap against: + +- The changed AL object names and types — especially permission sets, codeunits handling authentication or authorization, objects touching `Isolated Storage`, `OAuth2` flows, web service endpoints, and API pages. +- The changed procedures and triggers, weighted toward those that call `HttpClient`, write to telemetry, read or write secrets, manipulate record-level security, or bypass the permission model (for example, `Record.WritePermission`, direct table access from a non-owning app). +- Tokens extracted from the diff that relate to security concerns (`IsolatedStorage`, `OAuth2`, `Secret`, `Password`, `Token`, `HttpClient`, `Permission`, `Session`, `UserSecurityId`, `Commit`). + +A file enters the candidate worklist when its `keywords` intersect the extracted tokens or its topic (derived from filename and Description) matches a changed object type. + +Once the candidate worklist is known, resolve layer-precedence conflicts per READ. Drop lower-precedence files whose normative guidance (`## Best Practice` or `## Anti Pattern`) directly contradicts a higher-precedence candidate, and record each dropped file in `suppressed` with `reason: "layer-precedence"`. Files that would have been candidates but are hidden because their layer is disabled in consumer configuration are recorded with `reason: "configuration"`. Files that never became candidates are NOT recorded in `suppressed`. + +When the post-conflict worklist is empty because no applicable security knowledge exists, or because configuration suppressed every candidate, emit `outcome: "no-knowledge"`. When the worklist is empty because no applicable security knowledge matched the changes, emit `outcome: "completed"` with an empty `findings` array. + +## Action + +For each worklist entry, evaluate the diff against the file's `## Best Practice` and `## Anti Pattern` sections. Emit findings as follows: + +- When the diff contains a clear match for an Anti Pattern, emit a finding with severity `major` or `blocker`, a message summarizing the anti-pattern, `location` pointing to the offending line or range, and a `references` entry pointing to the knowledge file. Use `blocker` only when the knowledge file states the anti-pattern violates a platform-level guarantee (for example, documented secret-handling rules, permission-model invariants, or data-protection requirements). When the file does not make such a claim, the ceiling is `major`. +- When the diff contains code that contradicts a Best Practice without being a full anti-pattern, emit `minor` with the same reference shape. +- When the skill cannot detect a violation but the file is clearly applicable to the change, emit `info` citing the file. Repository-wide observations MAY omit `location`. + +Set `confidence` to: + +- `high` when the detection is based on an unambiguous pattern match (identifier, syntax, object type). +- `medium` when detection relies on heuristics or when any frontmatter dimension was `unknown`. +- `low` when the finding is an advisory derived only from applicability. + +Outcome selection: + +- `completed` — the skill evaluated every worklist item; default when the skill finishes normally, including when the resulting `findings` array is empty. +- `no-knowledge` — no applicable security knowledge survived Source, Relevance, configuration filtering, and conflict resolution. `findings` is empty. +- `not-applicable` — the task context lacks an AL dimension (no AL changes in the diff, or `technologies` filter rejected the task). +- `partial` — a time or token budget was hit before the worklist was exhausted. `summary.coverage` reflects the evaluated subset; `outcome-reason` explains the cause. +- `failed` — an unrecoverable error occurred. `outcome-reason` is required. + +## Output + +Output conforms to the DO output contract. A populated example: + +```json +{ + "skill": { "id": "al-security-review", "version": 1 }, + "outcome": "completed", + "summary": { + "counts": { "blocker": 1, "major": 0, "minor": 1, "info": 0 }, + "coverage": { "worklist-size": 2, "items-evaluated": 2 } + }, + "findings": [ + { + "id": "microsoft/knowledge/security/no-plaintext-secrets-in-telemetry.md", + "severity": "blocker", + "message": "A bearer token is passed to Session.LogMessage as part of the CustomDimensions payload. The referenced guidance documents this as a platform-level data-protection violation.", + "location": { + "file": "src/Integration/ApiClient.Codeunit.al", + "line": 85, + "range": { "start-line": 85, "end-line": 89 } + }, + "references": [ + { "path": "microsoft/knowledge/security/no-plaintext-secrets-in-telemetry.md" } + ], + "confidence": "high" + }, + { + "id": "microsoft/knowledge/security/avoid-implicit-commit.md", + "severity": "minor", + "message": "An explicit COMMIT inside a posting routine may leave the ledger in an inconsistent state if subsequent steps fail.", + "location": { + "file": "src/Sales/PostingRoutines.Codeunit.al", + "line": 201 + }, + "references": [ + { "path": "microsoft/knowledge/security/avoid-implicit-commit.md" } + ], + "confidence": "medium" + } + ], + "suppressed": [] +} +``` + +The empty-corpus case — BCQuality's state until security knowledge files land — produces: + +```json +{ + "skill": { "id": "al-security-review", "version": 1 }, + "outcome": "no-knowledge", + "summary": { + "counts": { "blocker": 0, "major": 0, "minor": 0, "info": 0 }, + "coverage": { "worklist-size": 0, "items-evaluated": 0 } + }, + "findings": [], + "suppressed": [] +} +``` diff --git a/skills/do.md b/skills/do.md index 6434b68..9ef4f6c 100644 --- a/skills/do.md +++ b/skills/do.md @@ -45,6 +45,8 @@ application-area: [all] `inputs` is a list of abstract input types the skill consumes. Standard values: `pr-diff`, `object-list`, `file-path`, `repository`, `telemetry-query`. `outputs` is always a single-element list naming the output kind; today only `findings-report` is defined. +`sub-skills` is an optional field. When present and non-empty, the skill is a **super-skill** that composes other action skills; see *Composition* below. Values are repo-relative paths to action-skill files. + ## Required sections Every action skill MUST contain these five sections, in order: @@ -91,7 +93,8 @@ Every action skill emits a single JSON document that conforms to this schema: "references": [ { "path": "string", "sha": "string" } ], - "confidence": "high | medium | low" + "confidence": "high | medium | low", + "from-sub-skill": "string" } ], "suppressed": [ @@ -99,6 +102,15 @@ Every action skill emits a single JSON document that conforms to this schema: "reference": { "path": "string", "sha": "string" }, "reason": "layer-precedence | configuration" } + ], + "sub-results": [ + { "...full nested findings-report..." : null } + ], + "skipped-sub-skills": [ + { + "skill": { "id": "string", "version": 1 }, + "reason": "configuration | not-applicable" + } ] } ``` @@ -119,6 +131,8 @@ An empty `findings` array with `outcome: completed` means the skill ran and foun **`findings[].id`** — a stable identifier for the rule or concern that produced the finding. For citation-based findings (any finding with a non-empty `references`), `id` MUST equal `references[0].path` — the primary knowledge file's repo-relative path. For skills that detect concerns without a direct citation, `id` is a skill-defined slug (kebab-case, stable across versions of the skill). The same `id` produced in two runs MUST refer to the same concern; consumers MAY deduplicate findings by `id`. +When a super-skill rolls up a non-citation finding from a sub-skill (an `id` that is a slug, not a path), the super-skill MUST prefix the `id` with `:` to avoid collisions across sub-skills (for example, a slug `missing-test` from `al-security-review` becomes `al-security-review:missing-test`). Citation-based findings are already globally unique through their repo-relative path and MUST NOT be rewritten. + **`findings[].severity`** — see the taxonomy below. **`findings[].message`** — human-readable explanation of the finding. Single short paragraph. No markdown formatting assumptions. @@ -140,11 +154,17 @@ The first reference is the **primary** reference: the knowledge file the finding **`findings[].confidence`** — the skill's confidence that the finding is a true positive, given the evidence it evaluated. Not applicability confidence, not severity confidence. Values: `high`, `medium`, `low`. +**`findings[].from-sub-skill`** — optional. Set only by super-skills. The `skill.id` of the sub-skill that produced the finding. Absent on findings produced directly by the emitting skill. + **`suppressed`** — MUST list every knowledge file that was discarded due to layer precedence or consumer configuration, whenever that file would otherwise have contributed to the worklist. Each entry contains: - `reference` — the suppressed file (same object shape as `findings[].references`). - `reason` — `layer-precedence` when another layer won under READ's precedence rules; `configuration` when the consumer disabled the file's layer. +**`sub-results`** — super-skills only. Array of complete findings-reports, one per sub-skill that was invoked (i.e., every sub-skill not listed in `skipped-sub-skills`). Each entry MUST itself conform to this output contract. Leaf skills MUST NOT emit `sub-results`. + +**`skipped-sub-skills`** — super-skills only. Array of sub-skills that were declared in frontmatter but not invoked. `reason` is `configuration` when the orchestrator disabled the sub-skill, or `not-applicable` when the super-skill's Relevance step ruled it out. + Severity taxonomy: - `blocker` — violates platform-level guarantees; the work cannot proceed as-is. @@ -152,6 +172,44 @@ Severity taxonomy: - `minor` — quality concern; worth flagging but not a gate. - `info` — observation or context; not actionable on its own. +## Composition (super-skills) + +A **super-skill** is an action skill whose frontmatter declares a non-empty `sub-skills: [...]`. A super-skill does not evaluate knowledge files directly; it invokes other action skills and composes their output. + +Composition is flat: a super-skill MAY list only leaf skills (skills without their own `sub-skills`). Nested super-skills are not permitted in v1. + +### Section interpretation for super-skills + +The five required sections still apply. Their meaning shifts from knowledge files to sub-skills: + +- `## Source` — names the sub-skills invoked (mirrors `sub-skills` in frontmatter). +- `## Relevance` — rules for deciding which sub-skills apply to the current task. A sub-skill is relevant when its declared `inputs` are satisfied by the orchestrator's provided inputs and the orchestrator has not disabled it via configuration. The super-skill MUST NOT filter sub-skills by task content (for example, by inspecting the diff or the file). Task-level applicability is the sub-skill's own responsibility; sub-skills signal non-applicability by returning `outcome: "not-applicable"` or `outcome: "no-knowledge"`. +- `## Worklist` — the final list of sub-skills to invoke; the rest go to `skipped-sub-skills`. +- `## Action` — invoke each worklisted sub-skill with the appropriate subset of inputs, collect its findings-report verbatim into `sub-results`, and copy its `findings[]` into the super-skill's top-level `findings[]` with `from-sub-skill` set. Findings from a sub-skill with `outcome: "failed"` MUST NOT be copied into the super-skill's top-level `findings[]` and MUST NOT contribute to the super-skill's `summary.counts` (their report is still preserved in `sub-results` for traceability, consistent with DO's rule that consumers ignore a failed skill's findings). +- `## Output` — the super-skill's output contract, including `sub-results` and, if any, `skipped-sub-skills`. + +### Outcome rollup + +A super-skill's `outcome` is derived from its sub-skills' outcomes. Let S be the multiset of sub-skill outcomes for sub-skills in the worklist (skipped sub-skills do not contribute): + +- `failed` — every element of S is `failed`. +- `partial` — S contains at least one `partial`, OR S contains at least one `failed` alongside at least one non-`failed` outcome. +- `not-applicable` — every element of S is `not-applicable`. +- `no-knowledge` — every element of S is `no-knowledge` or `not-applicable`, and at least one is `no-knowledge`. +- `completed` — otherwise (every element of S is `completed`, `no-knowledge`, or `not-applicable`, with at least one `completed`). + +When the worklist is empty (every sub-skill was skipped), `outcome` is `not-applicable`; `outcome-reason` SHOULD describe the skip reasons, for example *"all sub-skills disabled by configuration"* or *"no sub-skill accepted the supplied inputs"*. + +`outcome-reason` is required for `partial` and `failed` and SHOULD summarize per-sub-skill state. + +### Rolled-up summary + +`summary.counts` is the sum of sub-skill counts. `summary.coverage.worklist-size` and `items-evaluated` are the sums across invoked sub-skills. + +### Suppression scope + +A super-skill's top-level `suppressed[]` remains knowledge-file-only and is typically empty. Knowledge-file suppression is reported by the leaf sub-skill inside its own entry in `sub-results`. Sub-skills the super-skill chose not to invoke belong in `skipped-sub-skills`, never in `suppressed`. + ## Worked example A minimal action skill that cites applicable guidance for a changed AL file, without generating findings of its own: