Automation and AI

Retrieval-Augmented Generation (RAG)

Also called RAG, grounded generation

An AI setup where the tool searches your own documents first and answers from what it finds, not from memory.

Quick facts: Retrieval-Augmented Generation (RAG)

Category
Automation and AI
Also called
RAG, grounded generation
Level
Advanced
Affects
Chatbot accuracy, support workload, internal knowledge access, customer trust
Where to see it
Custom assistant builders, chatbot platforms with a knowledge base, automation tools such as n8n
In this article4
  1. How retrieval-augmented generation works
  2. Why RAG matters
  3. Where RAG goes wrong
  4. How to act on it

How retrieval-augmented generation works

An ordinary AI assistant answers from patterns learned during training. It has never seen your price list, your policies or last quarter’s proposals, so anything it says about them is guesswork dressed as fact. Retrieval-augmented generation puts a lookup step in front of the answer.

Your documents are split into passages and stored in a way that allows searching by meaning rather than by exact wording, so a question about “how long delivery takes” can find a paragraph headed “dispatch times”. When someone asks a question, the system retrieves the passages that look most relevant, hands them to the model along with the question, and instructs it to answer from that material. The model still writes the sentences; what has changed is where the facts came from.

Because the retrieved passages are known, a well-built system can also show its sources, letting a reader check the answer against the document it came from.

Why RAG matters

It is the difference between an assistant that sounds knowledgeable about your business and one that actually is. For a support chatbot, an internal helpdesk or a sales assistant, the useful answers are the specific ones: your refund terms, your service coverage, the clause in the contract you actually use.

It is also easier to keep current than the alternatives. Update a document and the next answer reflects the change, without retraining anything. And it narrows the space in which a model can invent, because the instruction is to answer from supplied text rather than from memory. That does not make invented detail impossible, but it makes it checkable, which is the property that matters when the answer goes to a customer.

Where RAG goes wrong

Most failures are retrieval failures, not model failures. If the right passage is never found, the model answers from whatever it was handed and sounds equally confident doing it. Contradictory documents cause the same problem: an old policy and a new one both retrieved, with no signal about which one wins.

Chunking causes quieter damage. Split a document badly and a passage arrives without the heading, the condition or the exception that gave it meaning — the answer is then technically present in the source and still wrong in effect. Access is a third risk: if everything is indexed together, a system can happily quote an internal note or a salary figure to whoever asks the right question. And a system that cites nothing gives a reader no way to catch any of this.

How to act on it

Start with the documents, not the technology. Decide which sources are authoritative, remove the outdated versions, and date what remains — a retrieval system built on a folder nobody has tidied will confidently repeat things you stopped doing years ago.

Then test with the questions people genuinely ask, including awkward ones, and check every answer against the source passage rather than judging how convincing it sounds. Require citations so answers stay auditable, keep an escalation path to a person for anything the system cannot support, and set a schedule for refreshing the content. If you are considering this for support or sales, scope it as a documentation project first and a custom AI assistant second.

Do and do not

Do

  • Tidy and date the source documents before building
  • Require answers to cite the passage they came from
  • Test with the awkward questions customers actually ask

Do not

  • Index internal or confidential files alongside public ones
  • Leave outdated versions of a policy in the source folder
  • Assume a confident answer means retrieval succeeded

Questions people ask about this

How is RAG different from fine-tuning a model?

Fine-tuning adjusts the model itself so it writes in a particular way or handles a particular kind of task. Retrieval-augmented generation leaves the model alone and supplies facts at the moment of the question. If your problem is that the assistant does not know your information, retrieval solves it; if the problem is style or format, that is a different fix.

What kind of documents work best in a RAG system?

Clear, current, well-headed text: policies, product details, service descriptions, procedures, past answers you were happy with. Scanned images without text, spreadsheets full of unlabelled figures and documents that contradict each other all cause poor answers. Tidying and dating the source material usually improves results more than any technical change.

Can RAG guarantee the AI will not invent answers?

No. It narrows the model to supplied passages, which greatly reduces invention, but the model can still misread a passage, blend two sources, or answer confidently when nothing relevant was retrieved. Requiring citations, testing with real questions, and giving people a route to a human are what make the remaining errors visible.

Related terms

Found this useful?

Share it, or ask an AI to summarise it

Back to the glossary

Knowing the term is the easy part

Applying it to your own site and budget is the work. Book a call and I will tell you what actually applies to you.