DEEN
AI & Governance

RAG chatbots for banks: sources and answer quality

How RAG chatbots retrieve documents and support answers, with checks for permissions, knowledge maintenance and answer quality.

First version: Updated: Reading time: about 9 minutes

Symbolic image of a RAG chatbot in banking: stylised chat window with an answer bubble and a citation marker, glowing lines connect the citation marker to a highlighted passage in a stack of regulatory documents, a magnifying lens above, dark blue background with emerald accents

A RAG chatbot combines document search with an answer composed by a language model. It provides a way to access institutional knowledge, but also introduces checks. The right passage must be retrieved, the answer must represent it accurately, and the requester’s access rights must be respected.

How a RAG chatbot works

Retrieval augmented generation (RAG) combines document search with a language model. The system retrieves passages for a user’s question and supplies them as context for the model’s answer. Unlike search alone, it generates new text that must remain open to domain review. Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

  1. 01Ingestion and indexing

    Documents are prepared, split into sections and stored in a searchable index, typically with vector representations. Metadata such as source, version and confidentiality level belongs in from the start.

  2. 02Retrieval

    For every user question, the best-matching passages are retrieved from the index. The quality of this step decides whether the answer builds on the right passages.

  3. 03Augmentation and generation

    The retrieved passages are handed to the language model together with the question; the model formulates the answer from this context and cites the passages it used as sources.

For banks, the pattern is attractive for three reasons: freshness, because new documents become available immediately after indexing, without retraining the model; grounding, because answers are tied to concrete passages; and source attribution, because every answer can state which document and which section it came from.

Identify typical errors and test for them

RAG supplies a language model with selected source passages for its answer. This can reduce unsupported answers, but does not prevent errors. The actual application needs separate checks for relevant retrieval, supported claims and correct refusals. DSK guidance on Retrieval Augmented Generation, version 1.0 Correctness is not Faithfulness in RAG Attributions

Even with perfect retrieval, residual risk remains: models can over-interpret retrieved passages or fill gaps between two passages with plausible-sounding fabrications. The reverse also holds: if the decisive document is never retrieved, the result is a systematically wrong answer that still carries a citation. And the much-advertised freshness is a process question, not a property of the technology: it only materialises when re-ingestion and versioning of the knowledge base are organisationally governed, with clear responsibilities for which document versions enter the index and when.

Assess what citations support

A citation initially identifies the document referenced by an answer. Whether that passage supports the claim, is current and is accessible to the person asking requires separate checks. A citation that looks correct does not establish those properties. Correctness is not Faithfulness in RAG Attributions DSK guidance on Retrieval Augmented Generation, version 1.0

  1. 01Aim for supported answers

    Design for answers supported by the supplied context and explicitly test unsupported claims. An instruction alone does not guarantee that behaviour.

  2. 02Check supporting passages

    Map claims to passages that actually support them and include conflicting sources in evaluation.

  3. 03Detect discrepancies

    Compare automated checks with domain-reviewed examples; record false alarms and missed errors.

  4. 04Account for consequences

    Define suitable approval or escalation for answers with significant consequences.

Disclose AI interaction

Article 50(1) of the AI Act generally requires people to be informed of direct interaction with an AI system unless this is obvious. The information must fit the use context. For a customer chatbot, disclosure should be visible at the start. AI Act consolidated as at 27 July 2026

The AI-interaction disclosure duty has applied since 2 August 2026. Article 50 contains further rules for generated content, with their own scope and exceptions. Whether a chatbot is high-risk AI depends on its purpose and legal classification; RAG alone does not answer that question. AI Act consolidated as at 27 July 2026

Data protection in the actual process

Hosting location alone does not determine data protection compliance. The legal basis, access rights, deletion arrangements and actual roles of the parties need assessment. A processing agreement is required where the particular relationship constitutes processing on behalf of a controller. DSK guidance on Retrieval Augmented Generation, version 1.0

The German Data Protection Conference’s RAG guidance, version 1.0 of October 2025, describes checks for processing personal data. Data flows must be traceable through external components. This includes whether prompts or answers are retained or used for further purposes. DSK guidance on Retrieval Augmented Generation, version 1.0

Access control before generation

A chatbot should only use knowledge that the person asking the question is permitted to access. The permission logic must take effect in the retrieval step: only passages the asking person is authorised to see, by role, confidentiality level and tenant separation, are retrieved; only then does the model formulate the answer. Whoever instead filters the finished output has already created the problem: the confidential passages were then part of the model's answer context, and rephrasings, summaries or clever follow-up questions threaten data leakage through the answer.

On top of this comes an integrity risk that classical applications do not know in this form: manipulated documents in the index can steer model behaviour. Prompt injection via the knowledge base and data poisoning are the patterns: a prepared document contains instructions or targeted misinformation that the model adopts while generating. The data integrity of the index is therefore a security question for the ISMS, the institution's information security management system: who may add documents to the knowledge base, how changes are reviewed, and how manipulation is detected.

Classification under DORA and MaRisk

A RAG chatbot belongs in the existing risk management framework as an ICT asset. On 18 December 2025, BaFin published guidance on ICT risks in the use of AI: under the DORA regulation, AI systems are to be inventoried as ICT assets, run through a controlled lifecycle and enter third-party risk management as soon as external services are involved. If the chatbot connects external model APIs or external hosting, this falls under ICT third-party risk under DORA Chapter V.

Ultimate human responsibility remains with the institution: a compliance copilot is a research accelerator, not a decision-maker. Answers that decisions follow from need a responsible human in the process. How institutions organise model governance and risk management in general is covered in our articles on MaRisk and the DORA regulation.

Measure answer quality reproducibly

An evaluation tool can support comparison runs. Release still requires a documented question set and answers assessed by domain specialists. The tool, model version and knowledge-base state should be recorded for each run. DSK guidance on Retrieval Augmented Generation, version 1.0

Reviewers should be able to trace an answer to its retrieved passages, model version and document versions. The logging design determines the necessary content, authorised access and deletion periods. It should combine traceability with data minimisation; keeping every question and answer indefinitely is not a quality measure.

Combine RAG and curated knowledge

A RAG knowledge base can use approved wiki content, original documents or a controlled combination. Editorial maintenance and retrieval are separate tasks. The architecture replaces neither access rules nor source review. DSK guidance on Retrieval Augmented Generation, version 1.0 Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

RAG knowledge base and curated knowledge wiki compared
DimensionRAG knowledge baseCurated knowledge wiki
TaskRetrieve evidence for generationEdit and connect knowledge
SourceApproved documents or curated contentReviewed documents and expert contributions
UpdatesControlled ingestion into the search indexEditorial maintenance with version and ownership
Error checksAssess retrieval and generated claims separatelyReview claims and source mapping editorially

A curated wiki can serve as a source for a RAG system. What matters is which content is approved and how its current version reaches the search index. An additional generation step requires its own checks even when its sources are maintained editorially.

Checks before deployment

The following criteria support technical and domain preparation. Release requires demonstrated behaviour in the intended use; an architecture label is insufficient.

  1. 01Grounding with abstention as the default

    The objective is answers grounded in permitted context. Unanswerable questions and disallowed instructions belong in the test set.

  2. 02Mandatory citations per statement

    Map source passages to claims and assess whether they support them.

  3. 03Permission filtering in retrieval, not in the output

    Role, confidentiality level and tenant separation filter the retrievable passages before generation.

  4. 04Audit trail per interaction

    Log the information needed for traceability, with defined access rights and retention periods.

  5. 05Escalation path to humans

    High-consequence answers go to a responsible human instead of being answered automatically.

  6. 06Re-ingest and versioning process for the knowledge base

    A governed process determines which document versions enter the index and when; only that makes the promised freshness real.

  7. 07Continuous measurement

    Assess suitable quality measures against domain-reviewed questions and rerun them after relevant changes.

The US standards institute NIST lists confidently delivered false statements of generative AI as a distinct risk category, confabulation, and rates them as particularly dangerous precisely in financial and legal contexts (NIST AI 600-1, 2024).

  • EU AI Act for banks: risk classes, obligations and deadlines, including the transparency obligations under Art. 50.
  • AI agents in banking: when the chatbot does not just answer but acts, from approval thresholds to audit trails.
  • The DORA regulation at a glance: ICT risk management, incident reporting and the third-party regime under Chapter V.
  • MaRisk (BaFin): module system and model governance in the German supervisory framework.
FAQ

Frequently asked questions about RAG chatbots in banking

Can a RAG chatbot in a bank be GDPR-compliant?

Hosting location alone does not determine data protection compliance. The legal basis, access rights, deletion arrangements and actual roles of the parties need assessment. A processing agreement is required where the particular relationship constitutes processing on behalf of a controller.

Must our chatbot disclose that it is an AI?

Since 2 August 2026, direct AI interaction generally requires disclosure unless the AI nature is obvious. Article 50(1) of the AI Act sets out the applicable scope.

Is an internal GRC copilot high-risk AI?

Internal search across GRC documents is not high-risk AI simply because a bank uses it. The intended purpose and Article 6 with Annexes I and III determine the classification. Use for creditworthiness or certain employment decisions therefore requires a separate assessment.

Does RAG prevent hallucinations?

RAG supplies a language model with selected source passages for its answer. This can reduce unsupported answers, but does not prevent errors. The actual application needs separate checks for relevant retrieval, supported claims and correct refusals.

May customer data go into the knowledge base?

Only where the purpose, legal basis and permissions support that processing. First assess whether anonymised or less data would suffice. Customer information needs suitable access separation and deletion rules; external services require assessment of the actual recipients and data flows.

An offer from T-NEX GmbH

Discuss the project with T-NEX

A maintained answer library and generative RAG are distinct development stages. The offering explains source maintenance and answer boundaries; evaluation tests your own question set before wider deployment.

Management: Andreas Unruh and Christoph Gembruch.

Published by T-NEX GmbH.