AI Tools

AI Knowledge Base: Make Company Knowledge Findable

Kollege weist eine neue Mitarbeiterin am Laptop ein – Sinnbild für eine KI-Wissensdatenbank, die Firmenwissen durchsuchbar macht

Images: created using AI

Short answer: An AI knowledge base answers questions about your internal documents in plain language and cites the source. The technology behind it is usually RAG: the model searches your files first, then answers. Success depends on curating the content, not on the model.

Every company above roughly 20 people has the same problem: the knowledge exists, but nobody can find it. It sits in a quote from 2023, in a PDF manual, in an email reply to a customer, in the head of the colleague who is on holiday. The instinct to solve this with a chatbot is right. Most projects fail not on the technology but on the question of which documents are allowed in at all.

How it works technically

The method is called retrieval augmented generation, or RAG. Microsoft describes it in its Azure AI Search documentation as an approach that grounds a language model’s answers in your own proprietary content. Google puts the benefit more bluntly in its documentation on grounding model responses: it reduces hallucinations and anchors answers to your data sources.

The sequence has four steps:

  1. Chunking. Your documents are cut into passages of a few hundred words. How intelligently they are cut decides answer quality later – a chunk that ends mid-table produces useless hits.
  2. Embedding. Each chunk gets a mathematical representation of its meaning, so the system finds material even when the question uses different words from the document.
  3. Retrieval. For each question, the most relevant chunks are pulled out, usually through a mix of semantic and classic keyword search.
  4. Generation. The language model receives the question plus the retrieved chunks and writes an answer, with a citation.

That last point is the real quality marker. A knowledge base without citations is worthless, because nobody can verify the answer. Insist that every vendor cites the document and ideally the page.

Four routes compared

Route Typical cost Time to deploy Strength Weakness
Better classic search in your file storage often included in the existing system days no new risks, immediate effect only finds what you can name
Chat feature in your existing office suite per-user licence uplift, in the tens of euros 1 to 3 weeks inherits permissions from the existing system vendor-bound, little control
Custom RAG on selected sources one-off setup plus running cost, usually mid four figures a year 4 to 10 weeks you control content, permissions and answer behaviour needs permanent ownership and curation
Fine-tuning a model on your data high, usually five figures months picks up your style and terminology knowledge ages with the training run, no citations
As of: August 2026. Cost ranges from mid-market project experience; not vendor statements.

For most mid-sized companies the third route is the right one, and the fourth almost never is. Fine-tuning solves a style problem, not a knowledge problem: whatever the model learned ages with every new price list. If you are weighing approaches in general, see choosing the right AI stack.

A worked example: 40 knowledge workers, 20 minutes a day

How much time internal information search really costs is surprisingly poorly evidenced. The most-quoted figure comes from a McKinsey Global Institute study published in 2012, which put nearly 20 percent of the working week for so-called interaction workers into searching for internal information and tracking down colleagues. The number is old and comes from a market report rather than a current survey – treat it as an order of magnitude, not as evidence about your company. Measuring in-house is better.

Take a conservative 20 minutes per person per day in a company with 40 knowledge workers and an internal rate of EUR 55 an hour:

  • Search effort today: 40 × 20 minutes = 800 minutes a day, or 13.3 hours. Across 215 working days that is 2,867 hours a year – around EUR 157,700.
  • Realistic reduction: 30 percent, not 80. That is around EUR 47,300 a year.
  • Cost: EUR 15,000 one-off for setup, a permissions concept and training. Running: EUR 350 a month for operation and model usage plus half a maintenance day a month, together roughly EUR 6,840 a year.

The investment therefore pays back inside the first year. The honest part is in the small print: the 30 percent only materialises if the content is curated. Point the system at an untidied file share and the reduction is zero – and you have paid for a tool that gives your people wrong answers in a confident tone.

The part everyone underestimates: selection and permissions

The most common wrong assumption is “we will just throw the whole drive in”. That creates two problems at once.

First, stale content. If the 2022 price list and the 2026 one sit in the same pot, the system answers with whatever matches the question better, not with whatever is current. Every source needs a validity date and an owner.

Second, permissions. If the knowledge base can see everything but its users cannot, it becomes a bypass around your access model. An apprentice asks about salary bands and gets them. Settle before the first import whether the system inherits permissions from your storage or needs rules of its own.

A third point concerns security. Germany’s Federal Office for Information Security notes in its security warning on indirect prompt injection that language models are vulnerable to this class of attack: an attacker can plant instructions inside the sources the model later processes. For an internal knowledge base that means documents from untrusted origins – inbound email attachments, uploaded customer files – do not belong in the index unchecked. Its publication on generative AI models gives the wider risk picture.

A realistic rollout plan

  1. Collect questions, not documents. For two weeks, everyone notes which questions they had to ask a colleague. That list is your requirements catalogue and your later test set.
  2. Pick one area. Not the whole company. One area with clearly bounded documents and an interested team: service, engineering, sales.
  3. Curate the sources. Fifty good documents beat five thousand unchecked ones. Each document gets a validity date and an owner.
  4. Test against the question list. Does the system answer 30 real questions correctly and with the right source? Anything below 80 percent is not production-ready.
  5. Anchor maintenance. A fixed monthly slot, a named owner, a list of questions answered badly. Skip this and quality decays within six months.

We build the connections into your existing systems as part of our automation services, and we select and operate the models under AI tools. If you would rather hand over day-to-day operation entirely, our managed service covers curation, monitoring and updates.

Start small, stay measurable
We build your knowledge base on one area with curated sources, test it against 30 real questions from your daily work and take over curation and operation on request.

See our AI tools services Hand over operation

Why projects fail

Nobody owns it. A knowledge base is not a project, it is an operation. With no named person removing stale documents and reworking bad answers, the result after six months is worse than the file search you had before.

Expectations are too high. Promising that the system answers everything guarantees disappointment. Say from the start which questions it covers – and that it should honestly decline everything else.

Nobody uses it. A tool that is not where the work happens goes unused. The search belongs inside the application your people already spend the day in, not on a separate site they have to remember.

Frequently asked questions

Do we need a locally hosted model for this?

Usually not. What matters is where your documents sit and whether the provider uses them for training. A European-hosted service with a processing agreement and training switched off meets most companies’ requirements. For particularly sensitive holdings a local model can make sense, though the operational burden is considerable. Our GDPR practice guide covers the data protection review.

How many documents make this worthwhile?

Fewer than most people think. From roughly 200 relevant documents, classic search becomes unwieldy. More important than volume is repetition: if five questions get asked several times a week, the build already pays.

What happens when the system says something wrong?

It will – which is why citations matter. Make it a rule: no answer without evidence, and for customer communication a human checks before anything leaves the building. For questions with legal or financial weight, the system should point to the responsible person instead of answering.

Can we include email?

Technically yes, advisable rarely. Email archives contain personal data, drafts, contradictory statements and confidential material. If at all, use a clearly bounded mailbox such as the service inbox – never staff members’ personal mailboxes.

How do we measure the benefit?

Through three numbers: questions per week, the share of answers with a usable source, and the number of follow-up questions that still go to colleagues. The third is the most honest.

Next step

We deliberately start these projects small: one area, curated sources, 30 real test questions. After four weeks you know whether it holds. Begin with our AI check or write to us through the contact form. If you want to lay the groundwork in the team first, our ChatGPT rollout guide helps.

Share
Newsletter

One real-world process, once a month

We only write when we have something substantial: a process we built ourselves, what it cost, what it saves — and where it got stuck. If nothing usable comes together in a given month, you get no email that month. You can unsubscribe from any email with one click.

  • Around once a month
  • Unsubscribe with one click
  • No sharing with third parties

Back to top

Get started

Ready to really put AI to work?

Tell us about your project in a few minutes — you’ll get an honest initial assessment, with no sales pressure.

  • Free & no obligation
  • Reply usually within 1 business day
  • GDPR-compliant