Load a few documents into the RAG Document Toolkit demo and ask a question, and the demo has to make a decision before the AI ever sees a word: which parts of your files should travel with the question? Send too little and the answer misses things. Send too much and you pay for tokens the question never needed.
The demo puts that decision in your hands with one slider, three numbers — and, since this page was first published, an AI document filter. This page explains what each control does and — more usefully — why it works the way it does.
The slider: send everything, or send what matters
The first decision is the big one, and it is governed by the threshold slider.
From the demo screen
Below the threshold (default: 20,000 tokens), the demo sends every chunk of every loaded file with every question. That sounds wasteful. It isn't — for two reasons:
- Nothing can be missed. There is no retrieval step, so there is no retrieval mistake. Every answer sees every page.
- Repeat questions are nearly free. Modern models cache the unchanged part of a prompt; cached input is billed at roughly a tenth of the normal rate. Since your file collection is the same on every question, the second question onward pays about 90% less for it. In a working session, sending a small collection whole is usually cheaper than retrieval — and better.
Above the threshold, sending everything stops being reasonable, and the demo turns to narrowing: first by document, then — if still necessary — by chunk. The context now changes with every question, so the cache discount no longer applies to it — but the set being sent is a small fraction of the collection, and the total cost is far lower.
The slider is where those two regimes meet. Move it and watch the economics flip.
The document filter: choose the files before touching the chunks
Most questions asked of a document collection concern only a few of its documents. A question about a lease does not need the production report; a question about the Q2 numbers does not need the wedding banquet order. Chunk-level retrieval eventually discovers this — but a whole stage earlier, the demo can often decide it outright, at the level where the decision is cheapest: the document.
From the demo screen
A quick, low-cost AI check reads each document's index entry and selects the files relevant to the question.
Questions about the collection itself ("what documents are loaded?") are answered from the index directly — no chunks needed.
Caps each document's contribution to the index. In practice the index typically runs well under the cap.
Here is how the filter works when the "AI document filter" box is checked:
Every file contributes a compact index entry. As the Toolkit converts a document, it collects the document's most characteristic phrases — heading text, bold and italic text, list items, table column headers, revision and comment authors, and person, organization, and place names — into one index line per document. The whole index typically runs a few percent of the collection's tokens (in our testing, around 5%); the Index size slider caps what each document may contribute.
A quick, low-cost AI check reads the index and the question, scores every document for relevance, and returns the few that matter. Three special verdicts round out the behavior: if the check concludes that all documents are relevant, or none clearly are, the demo simply proceeds with normal collection-wide retrieval — the filter never makes things worse than not having it. And if the question is about the collection itself, the check returns INDEX, and the demo answers directly from the index document — the behavior behind the nested checkbox.
Then comes the payoff — the fast path. If the selected documents' total tokens fit within the question's retrieval-and-enrichment budget, the demo skips retrieval and enrichment entirely and sends those documents whole, every chunk in order. Complete context, zero retrieval risk, for a handful of files instead of the whole collection. In our testing, more than half of all questions resolve this way. When the selection is too large for the budget, retrieval and enrichment run as described below — but confined to the selected files, with the entire budget spent where it counts.
The filter, in other words, is a third regime between the slider's two: send everything below the threshold, send whole selected files when the filter can narrow far enough, and send retrieved-and-enriched chunks only when nothing simpler suffices. Each step down the ladder is taken only when forced.
Above the threshold, questions have shapes
When retrieval does run — filter off, filter inconclusive, or a selection too big for the budget — not all questions want the same retrieval. “What is the warranty period?” lives in one or two places — the right answer is to follow the best matches wherever they are, even if all of them come from a single file. “Compare the payment terms across these contracts” is different: it is an aggregation, and an aggregate answer that silently skipped three of your seven files is worse than no answer.
So before retrieving, the demo runs a quick, low-cost AI check that classifies each question as pointed (local) or broad (global), and retrieves accordingly:
- Pointed questions get one retrieval across the whole collection (or the filtered set). The best-scoring chunks win, wherever they live.
- Broad questions get two passes: first a collection-wide retrieval to learn where the hits fall, then a separate retrieval into each file — so that every file contributes to the answer. That per-file guarantee is the difference between an honest summary of your collection and a confident summary of a third of it.
After either path, enrichment runs for every file that contributed chunks.
The three numbers
From the demo screen
How much of the collection each question retrieves. Larger = better coverage, higher cost.
A soft limit on added context. Nearby pages are dropped first when it's tight — split tables and revisions are always kept whole, so it can run slightly over. Zero turns enrichment off.
For broad questions, this share of retrieval is spread across every file, sized by file — the rest follows the best matches. Pointed questions always follow the matches.
1. Retrieval size — % of collection chunks (default 15%)
How much of the collection each question retrieves. It is a percentage rather than a fixed count for a practical reason: you add, remove, and replace files, and a percentage doesn't need retuning when you do. Fifteen percent of a small collection is a handful of chunks; fifteen percent of a large one scales up with it. (In code, the demo also applies a sensible floor and ceiling so tiny collections aren't starved and huge ones don't overspend.)
2. Enrichment budget — % of collection tokens (default 10%, zero = off)
Retrieval finds chunks; enrichment makes them whole. A retrieved chunk may hold half a table, a fragment of a tracked change, or a paragraph that only makes sense with its neighbors. Using the structural metadata every Toolkit chunk carries — table IDs, revision IDs, comment IDs, page numbers — the demo pulls in the missing pieces: the rest of the table, the other fragments of the revision, the surrounding page.
The budget caps how much enrichment may add, as a percentage of the collection's token count — larger collections tend to contain larger tables and wider-spread revisions, so the allowance scales with them. It is a soft limit: when the budget is tight, nearby-page additions are dropped first, but split tables and revisions are always kept whole, so the actual figure can run slightly over. That is deliberate. A partial table quietly misleads the AI; the Toolkit refuses to send one.
This same budget is what the filter's fast path measures against: selected documents that fit inside it are sent whole.
3. Coverage share — % of retrieval, broad questions only (default 25%)
For broad questions, this share of the retrieval is spread across every file, sized by each file's token count; the remaining share follows the hits, in proportion to where they landed. It is the knob behind the every-file guarantee: raise it and aggregate answers grow more evenly representative; lower it and they concentrate harder on the files with the strongest matches. Pointed questions ignore this number entirely — they always follow the matches, 100%.
Reading the meter
Every answer in the demo is accompanied by its actual numbers: which mode fired, the total tokens sent, and the retrieval/enrichment breakdown — for example:
Filter: 3 of 18 documents — 4,120 tokens, sent whole
Two things make those numbers worth a glance:
- They are exact, not estimates — every Toolkit chunk carries its precise token count in metadata, so the demo adds integers rather than guessing.
- They obey a one-line worst case you can plan with:
max context per question ≈ (retrieval % + enrichment %) × collection tokens
Set 15% and 10% on a 100,000-token collection and no question will cost meaningfully more than ~25,000 input tokens — a bound you can compute on a napkin before writing a line of code. With the document filter on, typical questions land far below that bound, because most of them resolve on the fast path with a few whole documents.
From demo knobs to production settings
These controls are not just demo conveniences; they are the calibration your own application needs. Because they are ratios of the collection rather than absolute numbers, the values you settle on while experimenting here transfer directly to your integration — across collections of different sizes, growing archives, and files that come and go. Watch the per-answer meter as you adjust them, find the balance of coverage, fidelity, and cost that fits your documents, and carry the same numbers into your code.
That is the demo's real purpose: not to hide the machinery, but to hand you the dials and show you the gauges.