A fix written before the bug is understood is a guess wearing a commit message. The sequence is fixed: reproduce, locate, name the cause, then change code — and the same reproduction that proved the bug proves the fix.
When to use
- Any report shaped like "X is broken / failing / wrong / flaky"
- An error message, stack trace, or failing test with no known cause
- Before writing any code intended as a fix
Steps
- Reproduce first. Turn the report into a command, test, or click-path that fails on demand. Can't reproduce it? That is the task now — gather the exact input, environment, and version until it fails; nothing gets "fixed" unreproduced.
- Read the actual error. The full message and the deepest frame that is your code — not the wrapper that re-threw it.
- Instrument, don't stare. Add targeted logging or a debugger at the boundary where good data should become bad; confirm which side of the boundary is wrong.
- Bisect the space. Halve it — by commit (
git bisectwhen a good version is known), by layer (API vs client), by input (which half of the payload triggers it) — until one component owns the failure. - Name the root cause in one sentence — mechanism, not vibes: "the cache key omits the locale, so the first locale wins" beats "caching issue". If the sentence can't be written, keep diagnosing.
- Fix the cause, prove it, pin it. Rerun the exact reproduction and watch it pass; add the regression test that fails without the fix; remove the instrumentation.
Rules
- No fix without a reproduction; no "fixed" claim without rerunning it.
- Fix causes, not symptoms — a retry, sleep, or try/catch around the crash site needs a written reason why the cause itself is unreachable.
- One hypothesis at a time; a change that didn't test a hypothesis gets reverted, not accumulated.
- Found a second bug on the way? Note it separately — don't widen this fix.
- Flaky counts as broken: reproduce by running enough iterations to make the failure rate a number, then diagnose the nondeterminism (ordering, time, shared state).
Examples
Good: "Repro: POST /import with a 2MB CSV fails 1-in-3. Bisected to the queue layer;
root cause: visibility timeout shorter than parse time, so the job runs twice.
Fix raises the timeout and makes the parser idempotent; repro passes 30/30;
regression test added."
Bad: "Wrapped the import in a retry — seems to work now."