AI engineeringData
What makes enterprise knowledge search trustworthy
Answering questions from your own documents is easy to demonstrate and hard to trust. Trust comes from five properties you can design, test and keep testing: permissions, provenance, freshness, honest gaps and measured quality.
- By
- FromNine editorial team
- Published
- Reading time
- 6 min read
Key takeaways
- Enforce access rights at retrieval time, per user, from the source systems. Never rely on the model to withhold a passage it has already been given.
- Every statement in an answer should point to a passage a person can open and check in seconds, with its date and status visible.
- Measure retrieval and answer quality against real questions before launch, and keep measuring, because content and models change underneath you.
Contents
The demonstration problem
Most knowledge assistants follow the same pattern, known as retrieval-augmented generation (opens external site): find the passages that match a question, then have a language model write an answer grounded in those passages. On a curated folder of fifty documents this works convincingly within a week.
Production looks different. Twenty years of shared drives, wikis, tickets, contracts and policies, many of them duplicated, outdated or confidential. In that setting the assistant fails in two ways. It gives a wrong answer and people stop using it. Or, worse, it gives a wrong answer with confidence and people act on it.
Trustworthy knowledge search is therefore less about the model and more about five properties of the system around it. Each can be designed and, importantly, tested.
Permissions: retrieve only what the user may read
The rule is simple: the assistant must never show someone a passage they could not open in the source system. The enforcement point is what matters. If a passage reaches the model, assume it can appear in the answer. Permissions must therefore be applied before generation, at retrieval time, for the user who is asking.
OWASP lists both sensitive information disclosure (opens external site) and vector and embedding weaknesses (opens external site), including leakage across permission boundaries, among the main risks of language model applications. In practice that leads to four requirements.
- Store the access rights of each source document with every chunk in the index, taken from the source system rather than redefined by hand.
- Filter on those rights at query time, using the identity of the signed-in user, before ranking and generation.
- Synchronise permission changes at least as quickly as content changes, and agree how fast a revoked right must take effect.
- Test with real roles. "Can a trainee find the salary scales?" belongs in the test suite, not in a post-incident review.
Provenance: citations a person can check
Each statement in an answer should point to the passage it came from: document title, version or date, owner, and a link that opens at the right place. Show the passage itself, not only the document name, so the reader can see in seconds whether the citation supports the sentence.
Design the answer format so that unsupported text stands out. If no retrieved passage supports a sentence, the system should leave it out or mark it. Then check this automatically: compare each cited passage with the claim it is meant to support, and sample the results for human review. Citations that exist but do not support the claim are a common and quiet failure.
Freshness and authority
Organisational knowledge has versions. The old travel policy and the new one both match the question "can I book business class?", and they give different answers. Index the metadata that tells them apart: effective date, status (draft, approved, superseded) and owner. Rank current, approved documents first, exclude superseded ones unless the user asks for history, and show the date in every citation.
Decide per topic which source is authoritative. The HR policy comes from the HR portal, not from an email that quotes last year's version of it. Set a re-indexing rhythm per source, from near real time for case data to daily for a wiki, and monitor it: a stale index produces answers that look right and are not.
Honest gaps: "I could not find this"
Often the most valuable answer is an honest gap. Language models can produce fluent statements that are not supported by anything, which NIST calls confabulation (opens external site) in its generative AI profile. Grounding in retrieved passages reduces this; it does not remove it. OWASP treats the downstream effect, people acting on false output, as a risk in its own right under misinformation (opens external site).
Set a relevance threshold below which the system says it found no reliable source, and suggest where to ask instead. For questions about policy, rights or obligations, do not let the model fill gaps from general knowledge. Then track the no-answer rate by topic. In the first months, that list is often the most useful output of the whole project: it shows content owners exactly where documentation is missing.
The architecture behind the five properties
None of this requires an exotic stack. It requires that permissions, metadata and checks are first-class parts of the pipeline rather than features added after the demonstration.
Figure 1
Illustrative workflowRead the diagram as text
Documents are taken from the source systems together with their access rights and metadata such as date, status and owner, and indexed.
A question arrives with the identity of the signed-in user. A permission filter removes every passage that user may not read, before anything reaches the model.
The model drafts an answer with citations to the remaining passages. A grounding check compares each statement with its cited passage. The user receives either a cited answer or an honest statement that no reliable source was found.
Personal data deserves explicit attention. An index of documents is still a processing of the personal data in them, so the GDPR principles of purpose limitation and data minimisation (opens external site) (Article 5) apply. Leave out sources that the use case does not need, make sure deletions in a source system reach the index, and include the index when you answer a request for access.
Measured quality, before and after launch
You cannot assess trustworthiness by trying a few questions in a meeting. Agree quality criteria with the people who own the content and measure each layer separately, so you know where to fix a problem.
| Layer | Question | How to measure |
|---|---|---|
| Retrieval | Did the right passages come back? | Recall and precision against a labelled set of real questions |
| Grounding | Is every statement supported by its citation? | Automated claim-to-passage checks, plus sampled human review |
| Answer | Is the answer correct, complete and useful? | Expert review of a sample against agreed criteria |
| Access | Did anyone see something they should not? | A permission test suite per role, run on every index or connector change |
| Use | Do people rely on it, and where does it fail? | Feedback, no-answer rate by topic, follow-up questions, escalations |
Build the test set from real questions: search logs, helpdesk tickets, questions new colleagues ask. A few hundred questions with answers agreed by content owners is a practical start. Run the full set on every change, whether a new model, a new chunking strategy or a new source. Quality drifts silently when content changes, and only a fixed test set makes that drift visible.
What trustworthy looks like to a user
For the people using it, all of this comes down to a few observable behaviours. The assistant answers only from documents they are allowed to read. Every claim has a citation they can open. They can see how current a source is. They are told plainly when there is no reliable answer. And they know someone is measuring the quality, because they can report a wrong answer and see it fixed.
AI output can still be wrong, and some questions deserve a person rather than a search. Designing for that openly is not a weakness of the system. It is what makes people willing to rely on it for the rest.
Sources
- arXiv. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., 2020) (accessed )
- OWASP GenAI Security Project. LLM02:2025 Sensitive Information Disclosure (accessed )
- OWASP GenAI Security Project. LLM08:2025 Vector and Embedding Weaknesses (accessed )
- OWASP GenAI Security Project. LLM09:2025 Misinformation (accessed )
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1) (accessed )
- EUR-Lex. Regulation (EU) 2016/679 (General Data Protection Regulation), Article 5 (accessed )
Continue reading
- ServiceAI Engineering & GenAIRetrieval, model strategy, evaluation and LLMOps for generative AI that holds up in production.
- SolutionEnterprise knowledgeAnswers from your own documents, trimmed to what each person may see, with a source for every claim.
- InsightWhere AI agents need human approvalAgents can prepare almost anything. A practical way to decide which actions they may complete alone and which need a person to say yes.
- InsightThe EU AI Act for public bodies: what to prepare by December 2027As amended in 2026, the AI Act applies its high-risk rules from December 2027, but much of it already applies today. A dated preparation plan for public deployers.
Making your own knowledge searchable?
We build knowledge assistants on your existing sources, with permissions, citations and a test set agreed with your content owners from the first sprint.