← library
skill✅ Testing & QAv1 · updated 2026-10-09

Human in the Loop

Hands work to the human only after verifying everything the agent can itself, then gives a short, risk-ordered list of the checks only a person can do (real devices, their accounts and inboxes, live payments, production, taste and business judgment), each with exact steps, the expected result, and why it needs a human. Use after any major code push, finished feature, opened PR, merge, or deploy, and whenever about to ask the user to test, try, check, confirm, or "let me know if it works".

Run it as a prompt

Paste this into any AI agent, or fetch it: curl -s https://uplift.page/api/v1/prompts/human-in-the-loop/raw

prompt.md
# Human in the Loop

The human's attention is the scarcest thing in the loop. Verify everything you can
yourself first, then hand over only what needs a person, set up so each check
takes minutes and the reply takes one line.

## When to use

- A major push lands: feature finished, PR opened, branch merged, preview or prod deployed
- You are about to write "please test", "try it out", "let me know if it works"
- You need the human to act (sign in, approve, use a real device) or decide something

## Steps

1. **Verify everything you can first.** List the checks the change needs, starting from
   what could hurt (money, data, access, anything irreversible), then run each one you
   can reach, failure paths included (bad input, missing config, repeats): tests, build,
   the flow in a real browser at desktop and ~375px, real requests, webhook replays,
   migrations on a local copy. Fix what fails; if the user only asked for a review, keep
   fixes as a flagged change they can drop.
2. **Sort what is left.** A check is human-only when you lack the access, the hardware, or
   the authority. Slow or tedious does not count. If setup would make it yours (seed data,
   test-mode keys, a mail catcher, a fake clock), do it. If access is the only blocker,
   ask for the access (a test key in the env file) and offer to run the check yourself.
3. **Prepare the ground** (preview deployed, data seeded, build installed) so each human
   check starts from an exact link or screen.
4. **Hand over** in one message: a one-line verdict; what you verified, with evidence; the
   human-only checks, riskiest first; decisions with your default; how to reply.
5. **Close the loop.** Fix each reported failure, re-run your own checks, and re-ask
   only the failed items plus anything the fix touched.

## Rules

1. Never hand the human work you could do: running tests, reading logs, checking the
   console, pasting errors, or a generic "make sure it works".
2. Every human check gets exact steps, the expected result, what failure looks like, and
   why it needs a person.
3. Five checks at most, riskiest first, each marked as blocking or not.
4. If nothing needs a person, say so. Never invent checks to look thorough.
5. Name what you could not verify and why. Untested must never read as verified.
6. Never ask for secrets in chat; say where they go (env file, secret store).

Install it as a skill

Agents that support the Agent Skills standard load it automatically when it applies.

download
curl -fsSL https://uplift.page/p/human-in-the-loop/SKILL.md --create-dirs -o .agents/skills/human-in-the-loop/SKILL.md
or from any MCP client
MCP server: https://uplift.page/mcp
Tool: pull_skill  {"slug": "human-in-the-loop"}

The full skill

The human's attention is the scarcest thing in the loop. Verify everything you can yourself first, then hand over only what needs a person, set up so each check takes minutes and the reply takes one line.

When to use

  • A major push lands: feature finished, PR opened, branch merged, preview or prod deployed
  • You are about to write "please test", "try it out", "let me know if it works"
  • You need the human to act (sign in, approve, use a real device) or decide something

Steps

  1. Verify everything you can first. List the checks the change needs, starting from what could hurt (money, data, access, anything irreversible), then run each one you can reach, failure paths included (bad input, missing config, repeats): tests, build, the flow in a real browser at desktop and ~375px, real requests, webhook replays, migrations on a local copy. Fix what fails; if the user only asked for a review, keep fixes as a flagged change they can drop.
  2. Sort what is left. A check is human-only when you lack the access, the hardware, or the authority. Slow or tedious does not count. If setup would make it yours (seed data, test-mode keys, a mail catcher, a fake clock), do it. If access is the only blocker, ask for the access (a test key in the env file) and offer to run the check yourself.
  3. Prepare the ground (preview deployed, data seeded, build installed) so each human check starts from an exact link or screen.
  4. Hand over in one message: a one-line verdict; what you verified, with evidence; the human-only checks, riskiest first; decisions with your default; how to reply.
  5. Close the loop. Fix each reported failure, re-run your own checks, and re-ask only the failed items plus anything the fix touched.

Rules

  1. Never hand the human work you could do: running tests, reading logs, checking the console, pasting errors, or a generic "make sure it works".
  2. Every human check gets exact steps, the expected result, what failure looks like, and why it needs a person.
  3. Five checks at most, riskiest first, each marked as blocking or not.
  4. If nothing needs a person, say so. Never invent checks to look thorough.
  5. Name what you could not verify and why. Untested must never read as verified.
  6. Never ask for secrets in chat; say where they go (env file, secret store).

What only a human can check

The test: could you do it with the tools you have, or with setup that does not need the human's credentials? If yes, it is yours. What is left falls into a few groups:

Needs a human because Examples
Physical hardware Real phone touch and haptics, camera, mic, GPS, Bluetooth, biometrics, push on a locked screen, printers, POS terminals
Their identity Signing in with their Google, Apple or SSO account; codes sent to their phone; their bank card
Real inboxes and channels Email landing in the inbox and not spam; rendering in Gmail or Outlook; SMS to a real phone; a post landing in their Slack
Live systems you cannot reach Production data, live-mode payments, provider dashboards, partner APIs with no sandbox
Judgment Taste, brand voice, copy tone; whether the business rules or numbers are right; whether it solves the actual problem
Authority Merging, prod deploys and migrations, emailing real customers, spending money, publishing
Real assistive tech and real users A full screen-reader pass with VoiceOver or NVDA, a first-time user trying it cold

Commonly handed over when the agent could have done them: tests, builds, console and network checks, mobile-width layout, API calls, log reading, local database queries, signed webhook replays, email rendering (render the template and look at it), two-user flows (two browser sessions), time-based behaviour (a fake clock), accessibility scans. references/catalog.md breaks this down by kind of change.

Handoff format

<Verdict in one line: ready / ready after X / not ready because Y>

Verified by me
- ✅ <check> — <evidence: command and result, screenshot, observed text>
- 🔧 Fixed: <bug> — <where the fix is, committed or not>
- ⚠️ Not verified: <check> — <why you could not>

Needs you: <N> checks, ~<M> min, <which block shipping>
1. <Outcome to confirm> (~<min>, blocking) — needs you because <reason>
   Do: <exact steps, starting from a link or screen>
   Expect: <what success looks like>
   Wrong if: <what failure looks like>

Decisions (my default in brackets)
- <question only they can answer> [<default>]

Reply with the numbers, e.g. "1 ✅ 2 ❌ email went to spam".

Keep it scannable: one line per verified item, details only where the human acts. When nothing needs a person, keep the verified list and end with "Nothing needs you before this ships."

Decisions and approvals

The same rule applies to questions. Ask only what only the human can answer: product intent, cost, risk tolerance, or a go-ahead for something irreversible or outward-facing. Batch the questions into one message, offer a default for each, and keep working on everything that does not depend on the answers. Having access to production does not make a prod deploy, a data migration, or a customer email yours to approve.

Timing

Hand over once per milestone, not as a drip of small requests. A single handoff after the push beats five interruptions during it. If the human is away, finish your own checks and leave the handoff ready, so they can act on it the moment they return.

Examples

Good (after a checkout feature push):
  Ready after two checks that need your Stripe account and inbox.

  Verified by me
  - ✅ 52 tests pass, types clean, production build OK
  - ✅ Checkout driven at 1440px and 375px; console clean; card-declined path shows the error
  - ✅ Webhook: signed replay marks the order paid; bad signature returns 400; a repeat
       delivery creates no second order
  - ⚠️ Not verified: Stripe's own delivery to the preview URL (needs your Stripe account)

  Needs you: 2 checks, ~6 min, both block launch
  1. A test-mode purchase reaches you end to end (~4 min, blocking). Needs your Stripe
     login and your inbox.
     Do: open https://preview-123.example.app/pricing, buy Pro with card 4242 4242 4242 4242
     Expect: success page, then a receipt in your inbox within a minute showing $19.00
     Wrong if: stuck on "Processing", no email after 5 minutes, or the email is in spam
  2. The receipt reads right to you (~2 min, blocking). Wording and brand are your call.
     ...
  Reply with the numbers, e.g. "1 ✅ 2 ❌ email went to spam".

Bad: "Pushed! Please test the checkout flow and let me know if anything breaks."
  It hands over the tests, the browser check, and the webhook check, with no steps.

Served from the uplift.page library and refreshed within 5 minutes of every update.