On building with AI
RAG knows what you said. It doesn't know what's true.
Point an AI assistant at your company's documents and something useful happens almost immediately. It reads your old emails, your support tickets, your internal docs, and it starts answering questions you never explicitly programmed it to answer. This is RAG (retrieval-augmented generation), and the reason it spread so fast is that it works on day one with material you already have, no authoring required.
Then a customer asks what you charge.
customer What does a project like this cost?
retrieved re: Proposal, Acme
…happy to move ahead at $2,000 for the…
answer Our rate is $2,000.Was $2,000 your rate? A discount for a long-standing client? A one-off for a smaller scope? A number from before you raised prices?
The assistant has no way to tell. That email is a record of something a person once said in a particular situation. It is not a statement of what you charge. But RAG doesn't distinguish between those two things, because to a vector index they look the same: text that scored well against the query. Retrieval ranks by similarity, and similarity has no opinion about authority.
So it quotes $2,000. And now that is your price, whether you like it or not.
This is not a RAG freshness problem
The first instinct is to treat this as staleness, which makes it sound like a tuning problem. Re-index more often. Add dates to the chunks. Filter to the last ninety days.
None of that fixes it. Suppose the email were from yesterday. It would still be the wrong source, because the assistant is reverse-engineering a fact out of a conversation. Sometimes that inference is correct. It is never authoritative, and you have no way of knowing which case you got.
The problem isn't that the document is old. It's that the document is the wrong kind of thing to answer that question with.
Data and directives
Almost everything an AI needs to know about your business falls into one of two categories, and they behave nothing alike.
Data is what accumulated. Tickets, emails, transcripts, past projects, published reference material. Nobody authored it as a statement of policy; it piled up as a side effect of doing business. It's large, it's inherited, and it tells you what happened.
Directives are what someone decided. Your rate. Your refund threshold. What the assistant is allowed to promise. Which client wants what. These are small and written on purpose. They don't describe the world, they bind it.
If a piece of data is wrong, that's an error. If a directive is wrong, that's a violation.
That's the test worth carrying around. A misremembered support ticket is embarrassing. A broken pricing rule costs money, and a broken compliance rule costs more than that. The two kinds cost very different amounts when they fail, which is why they need different tools.
| Data | Directives |
|---|---|
| Records | Rules |
| History | Decisions |
| Published reference | House positions |
| Large, inherited | Small, authored |
| True once | Binding now |
| Read to inform | Read to comply |
RAG is built for the left-hand column. It is very good there, and nothing below is an argument against using it. The argument is narrower: the right-hand column needs somewhere else to live, and most teams discover this only after a chatbot says something expensive.
Four limitations of RAG that better embeddings won't fix
Not "does badly." These are limits of the chunk-embed-retrieve pipeline itself. Rerankers, larger context windows and better embedding models improve RAG's recall without touching any of them, because none of the four is a recall problem.
It cannot represent asserted relationships
Embeddings encode similarity. They have no way to encode "this policy supersedes that one," "this step depends on that config," or "changing this breaks that."
A restaurant chain tightens its allergen policy. The old version and the new version are semantically near-identical, so similarity search ranks them side by side and may hand over the dead one. Meanwhile a runbook step and the configuration values it depends on share almost no vocabulary, so the two things that must travel together sit far apart in the index. Directed, author-asserted edges are a different kind of fact from closeness, and no improvement in embedding quality produces them.
It has no concept of absence
A RAG pipeline always returns its top-k. It hands back the least-bad match with exactly the same confidence it hands back a perfect one, which is the quiet mechanism behind a good share of chatbot hallucinations.
Ask an assistant about parental leave at a company that has never written a parental leave policy. It will return the sick leave policy, and it will tell an employee something false without hesitating. The correct answer was "there is no policy on file, ask HR." An assistant that knows it's missing context can ask a human. An assistant handed a plausible near-miss walks confidently into the error.
Retrieved chunks have no identity
Chunks are derived artifacts. RAG regenerates them every time you re-index, which means they can't carry version history, a record of where they came from, or an audit trail.
"What did this rule say in March, and who changed it?" is a question an insurance dispute or a failed audit will eventually ask you. Against a vector index it's unanswerable in principle, not merely unimplemented. The March chunks don't exist anymore.
Nothing can write back to it
RAG is read-only by design. The index sits downstream of a corpus, so an agent that discovers something during a task (a workaround, a new failure mode, a client's changed deadline) has nowhere to put it. You'd have to edit a source document and re-index, at which point the thing being written to is a knowledge base and retrieval is just a view over it.
Without a write path, every agent starts from scratch. The same edge case gets rediscovered four times because the fix lives in one closed ticket nobody re-indexed.
The alternative isn't better RAG
Every limitation above can be solved. Add stable IDs, a graph layer, versioning and a write path, and all four go away, and what you've built is no longer RAG. This is the gap Symbol is built for. Knowledge is authored as typed capsules with stable addresses rather than chunked out of documents. An assistant resolves @pricing/standard-rate and gets the whole rule, identically, every time. References between capsules are edges you assert, so changing your rate shows you the proposal template and two live projects that point at it. Capsules carry version history. Agents can write to them mid-task.
The same question, against directives instead of data:
customer What does a project like this cost?
resolved @pricing/standard-rate
Standard rate: $2,500/day.
Discounts require partner approval.
answer Our standard rate is $2,500 a day.Nothing was inferred. Nothing was ranked. The assistant read a statement someone wrote on purpose.
Running RAG and a knowledge base together
This is not a replacement architecture. In a production assistant RAG and a directive layer sit side by side, and the arrangement is simpler than it sounds.
A small set of directives loads before the conversation starts: pricing, tone, escalation rules, things the assistant must never claim. These can't be left to retrieval, because retrieval only fires when something in the question points toward the answer, and constraints don't announce themselves. Nothing about "draft a reply to this customer" would trigger a search for "never phrase returns as guaranteed."
Beyond that, both sources are tools the model chooses between. RAG for history and background. Directive lookup for policy and context. The system prompt carries one rule that matters more than the rest:
Directives are authoritative. Documents are evidence. If they conflict, directives win. And if no directive covers a policy question, say so rather than inferring one.
That second clause is the one teams skip, and it's the one that prevents the $2,000 email. Without it the model treats a well-ranked chunk as a fact.
Routing follows naturally. "What do you charge?" never touches RAG. "Has anyone reported this bug before?" never touches directives. "Can I refund an order from March?" uses both: the policy governs, the order history supplies the facts of the case. Where a question sits near the boundary, send it to the directive side and let the assistant say it needs to check. Erring toward a pause beats erring toward a number you can't honour.
One thing worth building early: log the conflicts. Every time RAG turns up something your directives contradict, you've learned either that a document is stale or that a rule is missing. That log is a to-do list generated by real demand instead of guesswork.
So is RAG enough?
For search over documents you inherited, yes. Nothing else does that job as cheaply. For anything a customer will act on, no. RAG tells your AI what your company said. It was never designed to tell it what's true, and that gap doesn't close with better embeddings, because the two are different kinds of knowledge with different costs when they fail.
Your data is for search. Your directives are for compliance. You need both. Just don't ask one to do the other's job.
Give your assistant directives, not guesses
Symbol is an AI-native knowledge base built entirely on the Model Context Protocol, so any MCP-capable assistant can read from it, and write back to it.
Get started with Symbol