Merge pull request #22 from Curabis/fix/smiley-outcome-based-tdd-trigger

al-complexity v2: standard-først-tjek + KISS (opfølgning på PR #21)
This commit is contained in:
Michael Dieringer 2026-07-31 22:50:55 +02:00 • committed by GitHub
commit b54bf88d7b
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
2 changed files with 89 additions and 22 deletions

View file

@ -66,32 +66,46 @@ Smiley's assets, activation conditions, and how they surface:
### 🔴 STOP GATE — Columbo → al-complexity ### 🔴 STOP GATE — Columbo → al-complexity
**Activate when:** **Two separate triggers here — do not let the first eclipse the second:**
- A user says "can you implement", "add a feature", "let's build", "hurtigt lige..." or
similar — and the requirement has not been clearly specified 1. **Columbo (clarify) activates when the requirement is ambiguous:** "can you
- A task feels MEDIUM or HIGH complexity before any scoping has happened implement", "add a feature", "let's build", "hurtigt lige..." or similar,
- Coding is about to start on something ambiguous where what's actually wanted isn't yet clear.
2. **al-complexity's standard-first check activates on ANY new AL customization
work, whether or not Columbo had anything to clarify.** A perfectly clear,
well-specified request ("add a field X that does Y") still deserves the
check — a crisp requirement can still turn out to be something standard BC
already does. Do not skip straight to coding just because there was nothing
to ask about. 2026-07-31: this is the same class of gap as the TDD trigger
fix — a gate tied only to "is this ambiguous" misses the clear-but-possibly-
unnecessary-custom-work case entirely.
**How it surfaces (undercover):** **How it surfaces (undercover):**
Claude naturally pauses. Asks one clarifying question. Listens. Asks the next. Claude naturally pauses. Asks one clarifying question. Listens. Asks the next
Does not say "I need to clarify first" — just does it. This IS Columbo. — but only if there's genuinely something to clarify. Does not say "I need to
clarify first" — just does it. This IS Columbo.
After the picture is clear, Claude naturally assesses scope and proposes a complexity Whether or not Columbo had anything to ask, Claude naturally checks Microsoft
tier. Does not say "al-complexity says..." — just reasons through it out loud and Learn and the BCApps reference clone before assessing scope (al-complexity's
waits for the user to confirm before writing any code. Step 0), then proposes STANDARD or a complexity tier. Does not say
"al-complexity says..." — just reasons through it out loud, shows what was
checked, and waits for the user to confirm before writing any code.
**The chain:** **The chain:**
``` ```
Ambiguous task detected New AL customization work about to begin
→ Claude asks questions (Columbo pattern — one at a time) → requirement ambiguous? → Claude asks questions (Columbo, one at a time) → clear
→ Picture becomes clear → Claude checks Microsoft Learn + BCApps reference clone (al-complexity Step 0)
→ Claude proposes scope + tier + route → standard BC covers it? → propose STANDARD, no code, stop here
→ doesn't → Claude proposes scope + tier + route
→ User confirms → User confirms
→ Code begins → Code begins
``` ```
Smiley will wave the flag hard here. "Hurtig lige" is a red flag. Smiley will wave the flag hard here. "Hurtig lige" is a red flag — and so is a
Coding before clarity is the most expensive mistake in development. task that looks obviously custom enough that nobody thought to check standard.
Coding before clarity is the most expensive mistake in development. Coding
before checking standard is a close second.
### 🔴 STOP GATE — Task Lifecycle (start, focus, close) ### 🔴 STOP GATE — Task Lifecycle (start, focus, close)

View file

@ -1,9 +1,9 @@
--- ---
kind: action-skill kind: action-skill
id: curabis-al-complexity id: curabis-al-complexity
version: 1 version: 2
title: CURABIS AL complexity triage title: CURABIS AL complexity triage
description: Advisory intake classifier. Assesses an implementation task and proposes a complexity tier (LOW/MEDIUM/HIGH) plus a route. Recommends only - it never starts work and never routes by itself. The developer confirms or adjusts the tier first. description: Advisory intake classifier. First checks whether Business Central already solves the requirement natively (Microsoft Learn + the BCApps reference clone) before proposing a complexity tier (STANDARD/LOW/MEDIUM/HIGH) plus a route. KISS applies to whatever custom route is chosen. Recommends only - it never starts work and never routes by itself. The developer confirms or adjusts first.
inputs: [task-description] inputs: [task-description]
outputs: [tier-recommendation] outputs: [tier-recommendation]
bc-version: [all] bc-version: [all]
@ -11,7 +11,7 @@ technologies: [al]
countries: [w1] countries: [w1]
application-area: [all] application-area: [all]
domain: orchestration domain: orchestration
keywords: [complexity, tier, routing, intake, scope, spec, tdd, architecture, advisory, human-in-the-loop] keywords: [complexity, tier, routing, intake, scope, spec, tdd, architecture, advisory, human-in-the-loop, standard-first, kiss, microsoft-learn, bcapps]
sub-skills: sub-skills:
- microsoft/skills/review/al-code-review.md - microsoft/skills/review/al-code-review.md
--- ---
@ -54,7 +54,36 @@ it never starts implementation and never routes on its own.
This is a **rubric, not a calculation** - there is no numeric score. The tier comes from This is a **rubric, not a calculation** - there is no numeric score. The tier comes from
which classification signals below match the task. which classification signals below match the task.
Loop: classify -> propose tier + route -> WAIT for human confirmation -> hand off. Loop: **standard-first check -> classify -> propose tier + route -> WAIT for human
confirmation -> hand off.**
## Step 0 — Standard-first check (2026-07-31, runs before classification)
Custom AL is the most expensive way to solve a requirement — every line becomes something
CURABIS must maintain forever. Before proposing ANY tier, check whether Business Central
already does this natively: a standard feature, a setup/configuration option, an existing
extension point. This is not optional and not skippable because the task "obviously" needs
code — the check itself is what proves that.
1. **Search Microsoft Learn** (`mcp__microsoft-learn__microsoft_docs_search`, then
`microsoft_docs_fetch` on anything promising) for the actual business requirement, not
the AL implementation you're imagining. Search for what the user wants to happen, not
"how to build X in AL".
2. **Check the real standard app**, not memory or training-data assumptions. Use the
machine-global reference clone (`~/.claude/reference-repos/microsoft/BCApps/` — see
`[[curabis-app-sources-must-be-checked-first]]` for the clone/refresh mechanism) and
grep for the relevant tables/pages/setup fields. Training data goes stale; the clone
does not.
3. **State the finding, with evidence — never "I checked and found nothing" unsupported.**
Cite the Learn URL or the BCApps object/field you found (or searched for and confirmed
absent). This is the same human-verifiable-evidence bar as the TDD red-confirmation —
a claim of "nothing" is only trustworthy if you show what you searched.
4. **If standard BC already covers it:** propose **STANDARD** — no tier, no code, just the
configuration/setup steps. This is the cheapest possible resolution and the reason this
check runs first. Stop here; do not continue to classification.
5. **If it genuinely doesn't:** proceed to classification below, and carry KISS forward as
a constraint on whatever tier is chosen (see "KISS applies to the route" below) — the
absence of a standard solution is not license to over-build the custom one.
## Classification signals ## Classification signals
@ -77,8 +106,21 @@ HIGH
- New table, or a field change on an existing table that needs an upgrade codeunit / data migration. - New table, or a field change on an existing table that needs an upgrade codeunit / data migration.
- Multi-module change, or a change to permissions. - Multi-module change, or a change to permissions.
## KISS applies to the route (not just to Step 0)
Once a tier is confirmed, the route itself must stay as simple as the requirement allows —
the fewest objects, the least new abstraction, no speculative generality for a future need
nobody has asked for. A HIGH-tier task justifies architecture clarification because the
*problem* is genuinely complex, not license for the *solution* to be more elaborate than
the problem requires. If a simpler design becomes visible during spec/architecture, propose
it — do not silently build the more complex version because it was the one first assumed.
## Routes (every tier keeps a review - control is preserved) ## Routes (every tier keeps a review - control is preserved)
STANDARD
- No AL code. Document the configuration/setup steps and hand off — nothing for
bcquality.agent.md to review, because nothing was written.
LOW LOW
- Implement -> **light review via bcquality.agent.md**. No spec or architecture phase, but - Implement -> **light review via bcquality.agent.md**. No spec or architecture phase, but
the review still runs. LOW never means "no review". the review still runs. LOW never means "no review".
@ -105,12 +147,23 @@ CURABIS-COMPLEXITY-005 Every tier gets a review. No tier skips bcquality.agent.m
a light review, not none. a light review, not none.
CURABIS-COMPLEXITY-006 Re-classify on scope change. If the task grows during work, stop and CURABIS-COMPLEXITY-006 Re-classify on scope change. If the task grows during work, stop and
re-propose a tier rather than silently continuing on the old one. re-propose a tier rather than silently continuing on the old one.
CURABIS-COMPLEXITY-007 Standard-first is not skippable. Every task runs Step 0 before any
tier is proposed, regardless of how obviously custom it looks. Show the Learn/BCApps
evidence — do not assert "nothing standard covers this" without it.
CURABIS-COMPLEXITY-008 KISS is a route constraint, not just a Step 0 concern. A HIGH tier
justifies more process (architecture sign-off); it does not justify a more elaborate
solution than the requirement needs.
## Output format ## Output format
``` ```
PROPOSED TIER LOW | MEDIUM | HIGH STANDARD-FIRST CHECK
SIGNALS <which classification signals matched, and why> Learn search: <what you searched, with URL(s) if found>
BCApps check: <object/field/setup area checked, found or confirmed absent>
Result: Standard BC covers this | Standard BC does not cover this
PROPOSED TIER STANDARD | LOW | MEDIUM | HIGH
SIGNALS <which classification signals matched, and why — omit if STANDARD>
ROUTE <the recommended path for this tier> ROUTE <the recommended path for this tier>
GATES <where human approval is required before proceeding> GATES <where human approval is required before proceeding>
AWAITING Confirm the tier or adjust it before I proceed. AWAITING Confirm the tier or adjust it before I proceed.