"Done" is a claim, not a fact. Before a diff becomes a commit or PR, interrogate it the way a hostile senior reviewer would — and either answer each challenge from the code or fix the code so the challenge dies. Presenting work skips straight past the doubts only when the doubts have been run down first.
When to use
- A feature, fix, or refactor is about to be committed or turned into a PR
- The user asks "is this ready?", "review this", or "anything I missed?"
- Handing work to a human reviewer whose time the grill protects
The grill
Work the diff, not the description of the diff. For each area, ask the question and answer it by pointing at code — file and line — or fix it on the spot:
- Correctness — What input makes this produce the wrong answer? What state was assumed that nothing enforces? Does every branch terminate, and what happens on the path where the happy call fails?
- Edge cases — Empty, null, zero, negative, huge, duplicated, concurrent, retried. Which of these actually reach this code, and what do they do when they arrive?
- Security — Where does untrusted input enter, and where is it trusted? Any query, command, path, or URL built by concatenation? Any authorization decided by data the caller controls?
- Performance — Any query-per-row loop, unbounded fetch, or work re-done per render or per request that a cache, join, or batch removes? What size of input makes this slow?
- Blast radius — Who else calls what this changed? What breaks for them, and did a caller's assumption silently change without its test changing?
Rules
- Every challenge gets a verdict: answered (cite the code that handles it) or fixed (change made). "Probably fine" is neither.
- Fix what you find before presenting; the grill is not a list of caveats to ship with.
- Report the grill's outcome in one short block — challenges raised, fixed, answered — so the human reviewer starts where the machine stopped.
- Scale depth to blast radius: a copy change gets thirty seconds; auth, payments, data deletion, and migrations get the full grill every time.
- The grill never replaces tests — a challenge that found a bug ends with a test that would have caught it.
Examples
Good: "Grilled the diff: 6 challenges. Fixed 2 (empty-cart checkout threw; webhook
retried non-idempotently). Answered 4 (citations in the PR description)."
Bad: "Looks good, ready for review." — no challenge was raised, so none was survived.