Back to case studies
Didactic case

Knowledge base for AI: organize documents before RAG

A didactic case study about preparing documents, permissions, source quality and risk before using retrieval-augmented generation.

Context

A company wants an AI assistant to answer internal or customer questions. Contracts, PDFs, spreadsheets and manuals exist, but the content is scattered, duplicated and not always current.

Observed symptoms

  • Documents have duplicated and conflicting versions.
  • Important rules are mixed with old drafts.
  • Confidential files are not clearly classified.
  • Users expect precise answers without verifiable sources.
  • The company wants to start with the tool before organizing knowledge.

Investigation path

01

Separate useful documents from noise

Not every file should enter the knowledge base. Start with reliable documents and clear owners.

02

Classify permissions

A customer assistant should not access the same content as an internal operations assistant.

03

Create test questions

Define real questions and expected answers before scaling the AI layer.

04

Require source and uncertainty

Answers should point to sources and admit when there is not enough evidence.

Technical decisions

  • Start with a small pilot using trustworthy documents.
  • Remove outdated versions before indexing content.
  • Separate public, internal and confidential material.
  • Measure quality through real questions, not impressive demos.

Expected outcome

AI becomes a layer on top of organized knowledge, not a magical fix for confused documents. That reduces risk and increases real usefulness.

FAQ

How to compare this case with your reality.

Should all documents go into the AI base?

No. Start with current, reliable and permission-safe documents.

What is the most important quality rule?

The assistant should show sources and admit uncertainty when the answer is not supported.

WhatsApp(12) 98855-9188