← library
skill🧰 Engineering practicev1 · updated 2026-09-10

Bug Diagnosis

Diagnoses bugs by reproducing first and naming the root cause before any code changes — reproduce, instrument, bisect, prove the fix against the reproduction, leave a regression test. Use whenever something is reported broken, failing, or behaving wrong, before attempting a fix.

Run it as a prompt

Paste this into any AI agent, or fetch it: curl -s https://uplift.page/api/v1/prompts/bug-diagnosis/raw

prompt.md
# Bug Diagnosis

A fix written before the bug is understood is a guess wearing a commit message. The
sequence is fixed: reproduce, locate, name the cause, then change code — and the same
reproduction that proved the bug proves the fix.

## When to use

- Any report shaped like "X is broken / failing / wrong / flaky"
- An error message, stack trace, or failing test with no known cause
- Before writing any code intended as a fix

## Steps

1. **Reproduce first.** Turn the report into a command, test, or click-path that fails
   on demand. Can't reproduce it? That is the task now — gather the exact input,
   environment, and version until it fails; nothing gets "fixed" unreproduced.
2. **Read the actual error.** The full message and the deepest frame that is your code
   — not the wrapper that re-threw it.
3. **Instrument, don't stare.** Add targeted logging or a debugger at the boundary
   where good data should become bad; confirm which side of the boundary is wrong.
4. **Bisect the space.** Halve it — by commit (`git bisect` when a good version is
   known), by layer (API vs client), by input (which half of the payload triggers it) —
   until one component owns the failure.
5. **Name the root cause in one sentence** — mechanism, not vibes: "the cache key omits
   the locale, so the first locale wins" beats "caching issue". If the sentence can't
   be written, keep diagnosing.
6. **Fix the cause, prove it, pin it.** Rerun the exact reproduction and watch it pass;
   add the regression test that fails without the fix; remove the instrumentation.

## Rules

1. No fix without a reproduction; no "fixed" claim without rerunning it.
2. Fix causes, not symptoms — a retry, sleep, or try/catch around the crash site needs
   a written reason why the cause itself is unreachable.
3. One hypothesis at a time; a change that didn't test a hypothesis gets reverted, not
   accumulated.
4. Found a second bug on the way? Note it separately — don't widen this fix.
5. Flaky counts as broken: reproduce by running enough iterations to make the failure
   rate a number, then diagnose the nondeterminism (ordering, time, shared state).

Install it as a skill

Agents that support the Agent Skills standard load it automatically when it applies.

download
curl -fsSL https://uplift.page/p/bug-diagnosis/SKILL.md --create-dirs -o .agents/skills/bug-diagnosis/SKILL.md
or from any MCP client
MCP server: https://uplift.page/mcp
Tool: pull_skill  {"slug": "bug-diagnosis"}

The full skill

A fix written before the bug is understood is a guess wearing a commit message. The sequence is fixed: reproduce, locate, name the cause, then change code — and the same reproduction that proved the bug proves the fix.

When to use

  • Any report shaped like "X is broken / failing / wrong / flaky"
  • An error message, stack trace, or failing test with no known cause
  • Before writing any code intended as a fix

Steps

  1. Reproduce first. Turn the report into a command, test, or click-path that fails on demand. Can't reproduce it? That is the task now — gather the exact input, environment, and version until it fails; nothing gets "fixed" unreproduced.
  2. Read the actual error. The full message and the deepest frame that is your code — not the wrapper that re-threw it.
  3. Instrument, don't stare. Add targeted logging or a debugger at the boundary where good data should become bad; confirm which side of the boundary is wrong.
  4. Bisect the space. Halve it — by commit (git bisect when a good version is known), by layer (API vs client), by input (which half of the payload triggers it) — until one component owns the failure.
  5. Name the root cause in one sentence — mechanism, not vibes: "the cache key omits the locale, so the first locale wins" beats "caching issue". If the sentence can't be written, keep diagnosing.
  6. Fix the cause, prove it, pin it. Rerun the exact reproduction and watch it pass; add the regression test that fails without the fix; remove the instrumentation.

Rules

  1. No fix without a reproduction; no "fixed" claim without rerunning it.
  2. Fix causes, not symptoms — a retry, sleep, or try/catch around the crash site needs a written reason why the cause itself is unreachable.
  3. One hypothesis at a time; a change that didn't test a hypothesis gets reverted, not accumulated.
  4. Found a second bug on the way? Note it separately — don't widen this fix.
  5. Flaky counts as broken: reproduce by running enough iterations to make the failure rate a number, then diagnose the nondeterminism (ordering, time, shared state).

Examples

Good: "Repro: POST /import with a 2MB CSV fails 1-in-3. Bisected to the queue layer;
       root cause: visibility timeout shorter than parse time, so the job runs twice.
       Fix raises the timeout and makes the parser idempotent; repro passes 30/30;
       regression test added."
Bad:  "Wrapped the import in a retry — seems to work now."

Served from the uplift.page library and refreshed within 5 minutes of every update.