โ† library
prompt๐Ÿค– AI agentsv2 ยท updated 2026-06-12

AI agent with tools

A tool-using agent with typed tools, MCP, and an eval harness.

Run it as a prompt

Paste this into any AI agent, or fetch it: curl -s https://uplift.page/api/v1/prompts/ai-agent-tooling/raw

prompt.md
# AI agent with tools

Build an agent that does real work through tools - with schemas tight enough that the model can't call them wrong, and evals that prove changes help.

## Recommended stack

- **TypeScript + an AI SDK** - orchestration. Typed tool definitions, streaming, and provider portability in one layer.
- **MCP (Model Context Protocol)** - tool transport. Tools become reusable servers any MCP-capable agent can mount.
- **zod / JSON Schema** - tool contracts. The schema is the prompt: precise params + descriptions beat paragraphs of instructions.
- **A trace + eval harness** - quality. A dozen scripted tasks with pass/fail assertions catch regressions before users do.

## Build steps

1. Write the agent's job description first: inputs, the 3โ€“7 tools it may use, and what 'done' means.
2. Define each tool with a strict schema, a one-sentence description that says when (not just what), and structured errors the model can react to.
3. Keep the loop simple: model โ†’ tool call โ†’ result โ†’ model; resist graph frameworks until a linear loop demonstrably fails.
4. Log every run as a trace (messages, tool calls, latencies); build 10โ€“20 eval tasks from real failures.
5. Gate prompt/tool changes on the eval suite, not vibes.

## Watch out for

- Ten overlapping tools where three orthogonal ones would do - ambiguity is where agents fail.
- Swallowing tool errors; return them structured so the model can retry intelligently.
- No kill switch or budget cap on loops.

## Definition of done

- Eval suite โ‰ฅ 90% on scripted tasks
- Every production run has an inspectable trace
- A malformed tool result degrades to a useful answer, not a crash

The full prompt

Build an agent that does real work through tools - with schemas tight enough that the model can't call them wrong, and evals that prove changes help.

Recommended stack

  • TypeScript + an AI SDK - orchestration. Typed tool definitions, streaming, and provider portability in one layer.
  • MCP (Model Context Protocol) - tool transport. Tools become reusable servers any MCP-capable agent can mount.
  • zod / JSON Schema - tool contracts. The schema is the prompt: precise params + descriptions beat paragraphs of instructions.
  • A trace + eval harness - quality. A dozen scripted tasks with pass/fail assertions catch regressions before users do.

Build steps

  1. Write the agent's job description first: inputs, the 3โ€“7 tools it may use, and what 'done' means.
  2. Define each tool with a strict schema, a one-sentence description that says when (not just what), and structured errors the model can react to.
  3. Keep the loop simple: model โ†’ tool call โ†’ result โ†’ model; resist graph frameworks until a linear loop demonstrably fails.
  4. Log every run as a trace (messages, tool calls, latencies); build 10โ€“20 eval tasks from real failures.
  5. Gate prompt/tool changes on the eval suite, not vibes.

Watch out for

  • Ten overlapping tools where three orthogonal ones would do - ambiguity is where agents fail.
  • Swallowing tool errors; return them structured so the model can retry intelligently.
  • No kill switch or budget cap on loops.

Definition of done

  • Eval suite โ‰ฅ 90% on scripted tasks
  • Every production run has an inspectable trace
  • A malformed tool result degrades to a useful answer, not a crash

Served from the uplift.page library and refreshed within 5 minutes of every update.