Merge pull request #66 from Curabis/rule/testcase-must-fail-before-implementation-sharpen-fgwt

[BCQuality] Sharpen TDD rule: 1-to-n FGWT scenarios + reject test-on-request alternative
This commit is contained in:
Michael Dieringer 2026-08-13 08:13:05 +02:00 • committed by GitHub
commit 06f417b37e
No known key found for this signature in database
GPG key ID: B5690EEEBB952194

View file

@ -1,64 +1,80 @@
--- ---
bc-version: [all] bc-version: [all]
domain: testing domain: testing
keywords: [tdd, red-green, test-first, testcase, verification, human-checkpoint, lifecycle] keywords: [tdd, red-green, test-first, testcase, verification, human-checkpoint, lifecycle]
technologies: [al] technologies: [al]
countries: [w1] countries: [w1]
application-area: [all] application-area: [all]
--- ---
# Test case must fail before implementation begins # Test case must fail before implementation begins
## Description ## Description
Every task starts with a test case, and the test case must **demonstrably Every task starts with a test case, and the test case must **demonstrably
fail (red)** — verified by the developer, not self-certified by the AI — fail (red)** — verified by the developer, not self-certified by the AI —
before implementation begins. The task may only be completed when the same before implementation begins. The task may only be completed when the same
test case **passes (green)**. test case **passes (green)**.
The lifecycle gate, in order: The lifecycle gate, in order:
1. Task started (branch + BC status — see `[[one-task-in-progress-at-a-time]]`) 1. Task started (branch + BC status — see `[[one-task-in-progress-at-a-time]]`)
2. Test case written for the requirement — including any fields, setup 2. Test case written for the requirement — including any fields, setup
objects, or test data structures the scenario needs that do not yet exist objects, or test data structures the scenario needs that do not yet exist
3. Test is run; **the developer verifies the red result** — an AI session 3. Test is run; **the developer verifies the red result** — an AI session
must never assert "the test fails" without a run the developer has seen must never assert "the test fails" without a run the developer has seen
4. Implementation begins 4. Implementation begins
5. Test is run again; **green is a precondition for finishing the task** — 5. Test is run again; **green is a precondition for finishing the task** —
no merge to the track branch, no BC `Done`, while the test case is red no merge to the track branch, no BC `Done`, while the test case is red
## Why ## Why
A test written after the code proves only that the code does what the code A test written after the code proves only that the code does what the code
does. A test that was red first proves two separate things: that the test does. A test that was red first proves two separate things: that the test
actually exercises the requirement (red = the gap is real), and later that actually exercises the requirement (red = the gap is real), and later that
the requirement is met (green = the gap is closed). The human verification the requirement is met (green = the gap is closed). The human verification
of red is the cheap insurance: thirty seconds of looking at a failing test of red is the cheap insurance: thirty seconds of looking at a failing test
catches the test that accidentally passes vacuously — the most dangerous catches the test that accidentally passes vacuously — the most dangerous
test in any suite. test in any suite.
## Anti Pattern A task normally needs **1 to n** FEATURE/SCENARIO/GIVEN/WHEN/THEN cases
([[test-feature-scenario-tags]]), not exactly one — a requirement with
// Implementation written first, test added afterwards to "cover" it. several distinct behaviors gets a scenario per behavior, each taken through
// Test passes on first run — it has never been observed red. its own red-then-green cycle.
// Nobody knows whether it tests the requirement or just the code.
## Explicitly Rejected Alternative
## Best Practice
Some external guidance (e.g. community "vibe coding" rule sets) recommends
// [GIVEN] a customer with a tier-price agreement the opposite: don't write tests proactively, only on explicit request,
// [WHEN] FindPrice is called for quantity 100 (test-one-when-per-test) to keep the AI focused on "just" the implementation. CURABIS rejects this
// [THEN] the tier price is returned, not the unit price for all behavior-changing work — the reason is verifiability, not process
// for its own sake: a red test is the only way to confirm the requirement was
// Run 1 (before implementation): FAILS — developer confirms red ✓ real and that the fix actually closes it, rather than trusting that the
// ... implementation ... implementation "looks right." This is a deliberate, considered position,
// Run 2: PASSES — task may now be completed ✓ not an oversight — do not silently adopt the opposite external guidance.
Test structure follows the existing testing rules: ## Anti Pattern
`[[test-one-when-per-test]]`, `[[test-setup-must-use-library-codeunit]]`,
`[[test-data-must-be-random-and-complete]]`, `[[test-feature-scenario-tags]]`. // Implementation written first, test added afterwards to "cover" it.
// Test passes on first run — it has never been observed red.
## Scope // Nobody knows whether it tests the requirement or just the code.
All CURABIS repositories — customer apps and AppSource apps alike. Applies to ## Best Practice
every task that changes behavior. Pure refactorings keep existing tests green
throughout; documentation/translation tasks are exempt. // [GIVEN] a customer with a tier-price agreement
// [WHEN] FindPrice is called for quantity 100 (test-one-when-per-test)
// [THEN] the tier price is returned, not the unit price
//
// Run 1 (before implementation): FAILS — developer confirms red ✓
// ... implementation ...
// Run 2: PASSES — task may now be completed ✓
Test structure follows the existing testing rules:
`[[test-one-when-per-test]]`, `[[test-setup-must-use-library-codeunit]]`,
`[[test-data-must-be-random-and-complete]]`, `[[test-feature-scenario-tags]]`.
## Scope
All CURABIS repositories — customer apps and AppSource apps alike. Applies to
every task that changes behavior. Pure refactorings keep existing tests green
throughout; documentation/translation tasks are exempt.