← library
prompt💬 Chatbots & RAGv2 · updated 2026-06-12

RAG chatbot over your docs

Retrieval-augmented chat with pgvector, hybrid search, and citations.

Run it as a prompt

Paste this into any AI agent, or fetch it: curl -s https://uplift.page/api/v1/prompts/rag-chatbot/raw

prompt.md
# RAG chatbot over your docs

A chatbot that answers from your documents with citations, says 'I don't know' honestly, and streams responses.

## Recommended stack

- **Postgres + pgvector** - vector store. Your data is already relational; HNSW indexes make one database do both jobs.
- **An embedding model (e.g. gte-small / text-embedding-3)** - retrieval. Small embeddings are cheap and good enough when you chunk well.
- **An AI SDK with streaming** - generation. Streamed tokens + cited chunks is the product.
- **Next.js** - app. Chat UI, ingestion routes, and auth in one deploy.

## Build steps

1. Chunk by structure (headings/paragraphs), 200–500 tokens with small overlap; store source URL + heading with each chunk.
2. Embed chunks into a vector(N) column; add an HNSW index and a tsvector column for hybrid search.
3. Retrieve top-k from both vector and full-text, merge with reciprocal rank fusion, then (optionally) rerank.
4. Prompt the model to answer only from provided chunks and to cite chunk ids; render citations as links.
5. Build a 20-question eval set with known answers; track answer + citation accuracy on every change.

## Watch out for

- Chunking by fixed character count across heading boundaries - retrieval quality dies here.
- Skipping the 'no answer found' path; hallucinations live there.
- Re-embedding the whole corpus on every edit instead of hashing content.

## Definition of done

- Answers cite real sources a click away
- Questions outside the corpus get an honest refusal
- P95 first-token latency < 1.5s

The full prompt

A chatbot that answers from your documents with citations, says 'I don't know' honestly, and streams responses.

Recommended stack

  • Postgres + pgvector - vector store. Your data is already relational; HNSW indexes make one database do both jobs.
  • An embedding model (e.g. gte-small / text-embedding-3) - retrieval. Small embeddings are cheap and good enough when you chunk well.
  • An AI SDK with streaming - generation. Streamed tokens + cited chunks is the product.
  • Next.js - app. Chat UI, ingestion routes, and auth in one deploy.

Build steps

  1. Chunk by structure (headings/paragraphs), 200–500 tokens with small overlap; store source URL + heading with each chunk.
  2. Embed chunks into a vector(N) column; add an HNSW index and a tsvector column for hybrid search.
  3. Retrieve top-k from both vector and full-text, merge with reciprocal rank fusion, then (optionally) rerank.
  4. Prompt the model to answer only from provided chunks and to cite chunk ids; render citations as links.
  5. Build a 20-question eval set with known answers; track answer + citation accuracy on every change.

Watch out for

  • Chunking by fixed character count across heading boundaries - retrieval quality dies here.
  • Skipping the 'no answer found' path; hallucinations live there.
  • Re-embedding the whole corpus on every edit instead of hashing content.

Definition of done

  • Answers cite real sources a click away
  • Questions outside the corpus get an honest refusal
  • P95 first-token latency < 1.5s

Served from the uplift.page library and refreshed within 5 minutes of every update.