When asked how long / how much / what's the effort for any piece of work, always give two numbers side by side, never just one:
- Traditional — what a competent human developer would take at normal velocity. The number a contractor or a Jira board would quote: weeks/days, sequential, with normal review and meetings overhead. The familiar anchor.
- Agent-driven — what it costs to build the same thing in an AI-agent operating model: the agent builds, a human reviews fast, and the agent works increasingly unsupervised and gets faster as the codebase and patterns become familiar. Expressed as build-days + ~tokens + ~cost.
The agent-driven number is normally 3–10× smaller in wall-clock. Don't bury it — it's the number the owner plans around. The traditional number is the reference, not the recommendation.
When to use
- Any "how long / how much / what's the effort / what's the timeline" question about building something.
- Sizing a feature, migration, refactor, or new subsystem before committing to it.
- Comparing scope options (give each its own row / tier).
The operating assumption (state it every time)
A tight feedback loop: the human reviews quickly and unblocks fast. The agent does the implementation and, over a project, needs less supervision per unit of work and moves faster as patterns and the codebase become familiar. Later phases and repeat-shaped work get cheaper than the first of their kind.
Price this compounding in: the 4th similar page/endpoint/adapter is far cheaper than the 1st. Apply a learning-curve discount to repeated work within a project (see calibration).
Rules
- Always produce BOTH the traditional and the agent-driven figure — never one alone.
- Lead with a table; keep token/cost figures clearly labeled approximate, with ranges not points.
- List human-only steps separately — they don't get the agent discount and gate the calendar.
- State assumptions and the top 2–3 risks; label each discoverable-from-data vs genuinely unknowable.
- If scope is ambiguous, give 2–3 tiers (minimal / recommended / full) instead of picking silently.
Output format
| Phase | Traditional | Agent build | ~Tokens | ~Cost | Notes |
|---|---|---|---|---|---|
| <Xd / Xw> | <X–Y M> | <risk / dependency> | |||
| Total |
Then: Recommendation (one line) · Human-only steps · Assumptions & risks · Tiers (if ambiguous).
Human-only steps (never fold into the agent number)
Gate the wall-clock no matter how fast the agent is:
- Credentials, API keys, OAuth apps, domain/DNS, billing, account approvals.
- Real-money / real-world test transactions (test orders, live payment, physical fulfillment).
- External review latency: design sign-off, legal, a partner's API access, app-store review.
- Irreducible human review of each batch (fast, but nonzero — scale with diff size).
Where these dominate, say so: "build is ~2 days; the gate is the vendor test order + sign-off."
Calibration (anchors — adjust per project, label approximate)
- Build-day = one focused agent session on a coherent slice (≈ one nontrivial feature / page / endpoint / adapter, implemented + self-tested). Several run per calendar day with a responsive reviewer; wall-clock is gated by review, not by the agent.
- Tokens: substantial feature slice ≈ 0.5–3M end to end; large migration / new subsystem ≈ 3–10M+; small targeted change ≈ 0.1–0.5M.
- Cost: derive from the active model's posted per-Mtoken price × token estimate; state the model and that pricing moves — don't hardcode a rate that will age. Give a range.
- Learning-curve discount: first-of-a-kind = full; each subsequent similar item ≈ 0.4–0.7× the previous, flooring ~0.3×. "Build 10 similar pages" is NOT 10× one page.
- Traditional ↔ agent ratio: greenfield/additive web work 4–10× compression; gnarly debugging, ambiguous requirements, or heavy human gates 1.5–3× (the human loop, not the typing, is the bottleneck). When the ratio is low, say why.
Examples
Good: "Build ~3–4 agent build-days (~5–8M tokens, ~$X–Y) over ~1 week; traditional ~4 weeks.
Gate: one real test order + design sign-off — that, not the build, sets the date."
Bad: "About 4 weeks." — one number, no agent-driven figure, no human gates called out.