Build an agent that does real work through tools - with schemas tight enough that the model can't call them wrong, and evals that prove changes help.
Recommended stack
- TypeScript + an AI SDK - orchestration. Typed tool definitions, streaming, and provider portability in one layer.
- MCP (Model Context Protocol) - tool transport. Tools become reusable servers any MCP-capable agent can mount.
- zod / JSON Schema - tool contracts. The schema is the prompt: precise params + descriptions beat paragraphs of instructions.
- A trace + eval harness - quality. A dozen scripted tasks with pass/fail assertions catch regressions before users do.
Build steps
- Write the agent's job description first: inputs, the 3โ7 tools it may use, and what 'done' means.
- Define each tool with a strict schema, a one-sentence description that says when (not just what), and structured errors the model can react to.
- Keep the loop simple: model โ tool call โ result โ model; resist graph frameworks until a linear loop demonstrably fails.
- Log every run as a trace (messages, tool calls, latencies); build 10โ20 eval tasks from real failures.
- Gate prompt/tool changes on the eval suite, not vibes.
Watch out for
- Ten overlapping tools where three orthogonal ones would do - ambiguity is where agents fail.
- Swallowing tool errors; return them structured so the model can retry intelligently.
- No kill switch or budget cap on loops.
Definition of done
- Eval suite โฅ 90% on scripted tasks
- Every production run has an inspectable trace
- A malformed tool result degrades to a useful answer, not a crash