A chatbot that answers from your documents with citations, says 'I don't know' honestly, and streams responses.
Recommended stack
- Postgres + pgvector - vector store. Your data is already relational; HNSW indexes make one database do both jobs.
- An embedding model (e.g. gte-small / text-embedding-3) - retrieval. Small embeddings are cheap and good enough when you chunk well.
- An AI SDK with streaming - generation. Streamed tokens + cited chunks is the product.
- Next.js - app. Chat UI, ingestion routes, and auth in one deploy.
Build steps
- Chunk by structure (headings/paragraphs), 200–500 tokens with small overlap; store source URL + heading with each chunk.
- Embed chunks into a vector(N) column; add an HNSW index and a tsvector column for hybrid search.
- Retrieve top-k from both vector and full-text, merge with reciprocal rank fusion, then (optionally) rerank.
- Prompt the model to answer only from provided chunks and to cite chunk ids; render citations as links.
- Build a 20-question eval set with known answers; track answer + citation accuracy on every change.
Watch out for
- Chunking by fixed character count across heading boundaries - retrieval quality dies here.
- Skipping the 'no answer found' path; hallucinations live there.
- Re-embedding the whole corpus on every edit instead of hashing content.
Definition of done
- Answers cite real sources a click away
- Questions outside the corpus get an honest refusal
- P95 first-token latency < 1.5s