The human's attention is the scarcest thing in the loop. Verify everything you can yourself first, then hand over only what needs a person, set up so each check takes minutes and the reply takes one line.
When to use
- A major push lands: feature finished, PR opened, branch merged, preview or prod deployed
- You are about to write "please test", "try it out", "let me know if it works"
- You need the human to act (sign in, approve, use a real device) or decide something
Steps
- Verify everything you can first. List the checks the change needs, starting from what could hurt (money, data, access, anything irreversible), then run each one you can reach, failure paths included (bad input, missing config, repeats): tests, build, the flow in a real browser at desktop and ~375px, real requests, webhook replays, migrations on a local copy. Fix what fails; if the user only asked for a review, keep fixes as a flagged change they can drop.
- Sort what is left. A check is human-only when you lack the access, the hardware, or the authority. Slow or tedious does not count. If setup would make it yours (seed data, test-mode keys, a mail catcher, a fake clock), do it. If access is the only blocker, ask for the access (a test key in the env file) and offer to run the check yourself.
- Prepare the ground (preview deployed, data seeded, build installed) so each human check starts from an exact link or screen.
- Hand over in one message: a one-line verdict; what you verified, with evidence; the human-only checks, riskiest first; decisions with your default; how to reply.
- Close the loop. Fix each reported failure, re-run your own checks, and re-ask only the failed items plus anything the fix touched.
Rules
- Never hand the human work you could do: running tests, reading logs, checking the console, pasting errors, or a generic "make sure it works".
- Every human check gets exact steps, the expected result, what failure looks like, and why it needs a person.
- Five checks at most, riskiest first, each marked as blocking or not.
- If nothing needs a person, say so. Never invent checks to look thorough.
- Name what you could not verify and why. Untested must never read as verified.
- Never ask for secrets in chat; say where they go (env file, secret store).
What only a human can check
The test: could you do it with the tools you have, or with setup that does not need the human's credentials? If yes, it is yours. What is left falls into a few groups:
| Needs a human because | Examples |
|---|---|
| Physical hardware | Real phone touch and haptics, camera, mic, GPS, Bluetooth, biometrics, push on a locked screen, printers, POS terminals |
| Their identity | Signing in with their Google, Apple or SSO account; codes sent to their phone; their bank card |
| Real inboxes and channels | Email landing in the inbox and not spam; rendering in Gmail or Outlook; SMS to a real phone; a post landing in their Slack |
| Live systems you cannot reach | Production data, live-mode payments, provider dashboards, partner APIs with no sandbox |
| Judgment | Taste, brand voice, copy tone; whether the business rules or numbers are right; whether it solves the actual problem |
| Authority | Merging, prod deploys and migrations, emailing real customers, spending money, publishing |
| Real assistive tech and real users | A full screen-reader pass with VoiceOver or NVDA, a first-time user trying it cold |
Commonly handed over when the agent could have done them: tests, builds, console and
network checks, mobile-width layout, API calls, log reading, local database queries,
signed webhook replays, email rendering (render the template and look at it), two-user
flows (two browser sessions), time-based behaviour (a fake clock), accessibility scans.
references/catalog.md breaks this down by kind of change.
Handoff format
<Verdict in one line: ready / ready after X / not ready because Y>
Verified by me
- ✅ <check> — <evidence: command and result, screenshot, observed text>
- 🔧 Fixed: <bug> — <where the fix is, committed or not>
- ⚠️ Not verified: <check> — <why you could not>
Needs you: <N> checks, ~<M> min, <which block shipping>
1. <Outcome to confirm> (~<min>, blocking) — needs you because <reason>
Do: <exact steps, starting from a link or screen>
Expect: <what success looks like>
Wrong if: <what failure looks like>
Decisions (my default in brackets)
- <question only they can answer> [<default>]
Reply with the numbers, e.g. "1 ✅ 2 ❌ email went to spam".
Keep it scannable: one line per verified item, details only where the human acts. When nothing needs a person, keep the verified list and end with "Nothing needs you before this ships."
Decisions and approvals
The same rule applies to questions. Ask only what only the human can answer: product intent, cost, risk tolerance, or a go-ahead for something irreversible or outward-facing. Batch the questions into one message, offer a default for each, and keep working on everything that does not depend on the answers. Having access to production does not make a prod deploy, a data migration, or a customer email yours to approve.
Timing
Hand over once per milestone, not as a drip of small requests. A single handoff after the push beats five interruptions during it. If the human is away, finish your own checks and leave the handoff ready, so they can act on it the moment they return.
Examples
Good (after a checkout feature push):
Ready after two checks that need your Stripe account and inbox.
Verified by me
- ✅ 52 tests pass, types clean, production build OK
- ✅ Checkout driven at 1440px and 375px; console clean; card-declined path shows the error
- ✅ Webhook: signed replay marks the order paid; bad signature returns 400; a repeat
delivery creates no second order
- ⚠️ Not verified: Stripe's own delivery to the preview URL (needs your Stripe account)
Needs you: 2 checks, ~6 min, both block launch
1. A test-mode purchase reaches you end to end (~4 min, blocking). Needs your Stripe
login and your inbox.
Do: open https://preview-123.example.app/pricing, buy Pro with card 4242 4242 4242 4242
Expect: success page, then a receipt in your inbox within a minute showing $19.00
Wrong if: stuck on "Processing", no email after 5 minutes, or the email is in spam
2. The receipt reads right to you (~2 min, blocking). Wording and brand are your call.
...
Reply with the numbers, e.g. "1 ✅ 2 ❌ email went to spam".
Bad: "Pushed! Please test the checkout flow and let me know if anything breaks."
It hands over the tests, the browser check, and the webhook check, with no steps.