Generative AI

RAG Is Easy to Demo and Hard to Operate

Production retrieval systems succeed when content ownership, evaluation and failure handling are designed before the chatbot interface.

BELFORT AI Engineering30 August 20266 min read
Generative AI knowledge retrieval system

Retrieval-augmented generation can look convincing with a clean document set. Production introduces stale policies, conflicting sources, access rules and questions that have no reliable answer.

Treat knowledge as a product

Assign owners to source collections, define freshness targets and test retrieval separately from generation.

A trustworthy assistant knows when its evidence is insufficient.

BELFORT AI Engineering

Start with the decision, not the technology

RAG Is Easy to Demo and Hard to Operate becomes useful only when it is attached to a specific operating decision: which evidence a retrieval system may use to answer a business question. The team should identify who makes that decision today, what information they use, how exceptions are handled and what a better outcome would look like. This framing prevents a technically impressive capability from being deployed without a practical purpose, accountable owner or credible definition of value.

Establish a defensible baseline

Before changing the workflow, measure its current cost, speed, quality and variation. The baseline should cover normal work as well as difficult cases, seasonal pressure and manual rework. For this topic, the most useful starting evidence includes retrieval relevance, source freshness, citation and answer-quality tests. A baseline is not bureaucracy: it is what allows leadership to distinguish real improvement from novelty, shifted work or a temporary demonstration effect.

Design the complete operating system

The model is only one component. Production design must include source systems, data contracts, interfaces, policy rules, user experience, access control, monitoring and a safe fallback. Each handoff needs a defined owner and service expectation. Teams should know what happens when data is late, a dependency is unavailable or confidence is low. That full chain determines whether the capability can support daily operations.

Treat data quality as an ongoing responsibility

Relevant data needs documented meaning, lineage, freshness and permitted use. Validation should run at ingestion and again near the point of decision because technically valid records can still be stale or unsuitable. Owners need alerts they can act on, plus a process for correcting defects and communicating schema changes. Historical training data must never be assumed to represent future users, conditions or operating policy without review.

Make risk concrete and testable

A useful risk review describes affected people, plausible harm, detection time and reversibility. The central failure mode here is confident answers based on missing, conflicting, unauthorized or obsolete material. Controls should be proportional to that consequence and independent from the component they supervise. High-impact actions need narrower permissions, stronger evidence and direct human checkpoints; low-impact assistance can move faster when users can inspect, correct and safely reject the output.

Build human judgment into the workflow

Human oversight is effective only when reviewers have time, context and authority. A confirmation button does not create meaningful control if staff cannot understand the recommendation or pause the process. Interfaces should show the evidence, uncertainty and applicable policy in plain language. Corrections, overrides and rejected outputs should become structured feedback, while employees need training on both appropriate use and known limitations.

Evaluate under realistic conditions

Evaluation should represent the actual population, language, workload and failure costs of the intended use. Test ordinary cases, rare but consequential cases and deliberately degraded conditions. Report results by relevant segment rather than hiding weaknesses in a single average. Compare against the existing process and a simple alternative. Independent review is particularly valuable before increasing autonomy or connecting the system to consequential actions.

Operate for change, incidents and recovery

After launch, inputs, user behavior, vendors and policies will change. The operating plan therefore needs versioned releases, change approval, observability, incident classification, rollback and retirement criteria. A named team must be able to answer which version made a decision and why. Monitoring should lead to predefined action: investigate, limit, revert or stop. Visibility without authority and response procedures is not operational control.

Measure value and system health together

A balanced scorecard for Generative AI should combine business outcomes with adoption, quality, reliability and risk. Relevant measures include grounded-answer rate, retrieval recall, abstention quality and content freshness. Every metric needs a definition, source, owner and response threshold. Cost should include integration, review, monitoring, incidents and future change—not only model access. Leaders can then compare realized value with the continuing cost of operating the capability responsibly.

Scale through reusable evidence

The first release should be deliberately bounded and designed to teach the organization. Capture assumptions, test sets, decisions, exceptions and user feedback as reusable assets. Expand scope only when evidence remains stable and teams can recover from failure confidently. Reusable data contracts, evaluation suites, control patterns and operating playbooks make the next use case faster without weakening accountability. Scale should mean repeating a trusted method, not multiplying disconnected pilots.