Our services
If you can think it, we can make it brainsoft.
If you can think it, we can make it brainsoft.
Written By: BrainSoft In AI & Research
A client asked us to wire an LLM to their internal docs last year. About 200 markdown files, a few PDFs, a wiki export. Not a huge corpus. The instinct is to reach for a managed vector database and a framework, but for a knowledge base this size you can build the whole retrieval loop yourself in an afternoon and understand every part of it. That matters when answers come back wrong and you have to figure out why.
You will chunk the documents, embed the chunks with a hosted embedding model, store the vectors in a plain Postgres table or a local index, retrieve the top matches for a question, and pass them into a prompt with a strict instruction to answer only from the supplied text. Then you test it against questions you already know the answers to.
Before any code, look at what you actually have. Convert everything to plain text and count it. For 200 files you are probably looking at under a million tokens total, which changes the design: you can afford to re-embed the whole corpus on every deploy, and you do not need incremental indexing machinery. You also do not need a separate vector service. Postgres with the pgvector extension handles a corpus this size without thinking about it.
The part people skip is cleanup. Strip navigation boilerplate from the wiki export, drop the table of contents pages, and keep the document title attached to every chunk. That title is often the strongest signal for retrieval, more than the body text.
Fixed-size character splitting is easy and usually wrong. It cuts sentences in half and separates a heading from the paragraph it describes. Split on structure instead: markdown headings first, then paragraphs, then merge small pieces up to a target of roughly 500 to 800 tokens with a little overlap.
def chunk(doc, target=600, overlap=80):
parts = split_on_headings(doc.text)
chunks, buf = [], ""
for p in parts:
if len(buf) + len(p) > target and buf:
chunks.append(doc.title + "\n" + buf)
buf = buf[-overlap:]
buf += "\n" + p
if buf:
chunks.append(doc.title + "\n" + buf)
return chunks
Store each chunk with its source path and heading so you can cite it later. Citations are not decoration. When a user can click through to the paragraph that produced the answer, they stop treating the system as an oracle and start using it as a search tool.
Embed the chunks once, store the vectors, and at query time embed the question and take the top 5 to 8 by cosine similarity. That is the whole retrieval step. Add a similarity floor: if the best match is below it, return "I could not find this in the documentation" instead of sending weak context to the model. Most bad answers come from confident generation over irrelevant chunks, not from the model being weak.
The prompt should be short and blunt. Give the model the chunks with their sources, then say: answer using only the text below, quote the source file for each claim, and say you do not know if the answer is not present. Keep temperature low. If you need multi-hop reasoning across documents, that is a different project, and for a small knowledge base it is usually unnecessary.
Then evaluate. Write 30 questions you know the answers to, run them, and read the outputs yourself. Track two things: did retrieval put the right chunk in the top 5, and did the answer match the source. Fixing retrieval failures is cheap. Fixing generation failures by tweaking the prompt rarely works. If you want help scoping this against your own docs, get in touch.
Until one of those is true, the hand-rolled version is easier to debug, cheaper to run, and lives in the same repo as the rest of your application. We build these alongside the surrounding web and backend work as part of our services, and the honest advice for a small corpus is usually to start small.
No. Postgres with pgvector, or even an in-memory index rebuilt at startup, handles a few hundred to a few thousand documents fine. A dedicated vector service adds an operational dependency you do not need until your corpus or query volume grows.
Start with five and look at the results. More context is not automatically better; irrelevant chunks give the model material to hallucinate from. A similarity threshold that returns nothing is often more useful than a low-confidence answer.
Build a set of questions with known answers from your own docs, run them after every change, and check both retrieval and the final text. Reading 30 outputs by hand takes an hour and catches more problems than any automated metric.