Putting Your Documents to Work: RAG Document Toolkit Use Cases for Customer-Facing Apps and the Enterprise Interior
Most of what an organization knows lives in documents. Discharge instructions, equipment manuals, loan agreements, board packets, engineering change orders, benefits handbooks — written by people, for people, and stored as DOCX, RTF, and HTML. Large language models can now answer questions from that material, but only if the documents reach the model in a form it can use. That is the job of RAG Document Toolkit for .NET.
This post walks through where our customers are applying the toolkit — first in applications their own customers use, then inside the company for financial, technical, managerial, and HR information — and explains how the toolkit keeps those solutions accurate and inexpensive to run.
What the toolkit does, in one paragraph
RAG Document Toolkit is two .NET libraries. Dan converts DOCX, RTF, HTML, text, and Markdown files into AI-friendly Markdown chunks sized in tokens, and attaches metadata to every chunk: document id and title, page number, section, heading text, whether the chunk holds a table, and the revision and comment authors involved. Ven works on those chunks after retrieval: it completes tables that were split across chunks, gathers all of a reviewer's revisions or comments, adds neighboring pages or a whole section when the question needs context, and builds a compact per-collection index that an inexpensive model can use to route a question to the right documents. Your application keeps its own vector database, its own model provider, and its own security. The toolkit supplies the document intelligence.
Customer-facing applications
Healthcare (about half of our enterprise customers)
Healthcare organizations generate enormous volumes of narrative documents that patients, clinicians, and staff need to query without reading end to end.
- Patient portals. A patient asks "Can I take ibuprofen with the medication I was prescribed?" and the answer is drawn from their own discharge summary and the pharmacy's medication guide, with a citation to the document, section, and page. Because Dan preserves document structure, the model sees the Medications section as a section and the dosage table as a complete table, not a fragment.
- Clinical protocol lookup. Nurses and physicians query care pathways, formulary guides, and infection-control procedures. Sections and headings in the chunk metadata let the application restrict retrieval to the protocol in force, and the AI document filter routes a question about a named condition to the two or three documents that cover it.
- Payer and plan documents. Members ask what a plan covers; the answer comes from the summary of benefits and the evidence-of-coverage document, both table-heavy. Ven's table completion guarantees the model receives the whole benefits table, including the rows that landed in the next chunk.
- Clinical research and regulatory files. Protocol amendments and investigator brochures go through many tracked revisions. Ven can collect every change by a given author, so a query like "what did the medical monitor change in section 6" is answerable.
Privacy matters here more than anywhere. The toolkit runs inside your own .NET application. Documents are converted on your servers, chunks go into the database you already control, and you choose which model endpoint receives which text. Nothing about the toolkit requires sending documents to a third-party document service.
Manufacturing (about a fifth)
- Field service and technician support. Service manuals, torque specifications, wiring tables, and maintenance schedules are the classic split-table problem: a 40-row specification table rarely fits one chunk. Ven's table completion means a technician asking for the correct bearing clearance gets the full table, and a citation to the page a supervisor can verify.
- Distributor and customer portals. Product catalogs, safety data sheets, and installation guides answer "which model supports 480 V three-phase input" without the customer opening a PDF viewer. Heading metadata lets the application filter by product family before retrieval.
- Quality and compliance. Certificates, inspection procedures, and standard operating procedures are usually structured by section. Section-aware retrieval keeps the answer inside the procedure that applies.
Finance, legal, and others
- Financial services. Prospectuses, account agreements, fee schedules, and policy documents are dense with tables and defined terms. Client-facing assistants can answer fee and coverage questions from the governing document itself, with the page cited.
- Legal. Contract portfolios, matter files, and negotiated drafts live in Word with tracked changes and comments. Ven's revision and comment methods let a reviewer ask what opposing counsel changed in the indemnification clause and see only those revisions, rather than the whole redline.
- Customer support in any industry. Product documentation and knowledge bases become a support assistant that quotes the manual rather than improvising.
Inside the corporation: faster access to what the company already knows
The same toolkit serves employees, where the value is measured in hours not spent hunting through shared drives and email attachments.
Financial information. Budgets, monthly close packages, audit reports, and board decks are table-first documents. An analyst asks "what was the capital expenditure variance for the Austin plant in Q2" and the model receives the complete variance table, not the half that happened to match the query. Dan records which chunks hold tables and which do not, so the application can also route table-heavy questions differently from narrative ones.
Technical information. Specifications, design documents, runbooks, and engineering change notices are heavily sectioned and heavily revised. Section metadata lets an engineer scope a question to the current design, and revision tracking answers "what changed between revision C and D of the interface spec" from the tracked changes themselves.
Managerial information. Meeting minutes, project status reports, and policy documents accumulate decisions that nobody can find six months later. A manager asks "what did we decide about the vendor contract renewal in March" and receives the paragraph, the document, and the page. Comment-author metadata surfaces who raised an objection and where.
Human resources. Employee handbooks, benefits summaries, leave policies, and onboarding guides generate the same questions every week. An HR assistant answers from the policy in force and cites it, which is exactly what employees want and what HR wants to stand behind. Since your application controls which documents enter the collection and who can query it, role-based access is a decision you make in your own code, not a feature you hope a hosted service enforces.
Why this is efficient and cost-effective
A RAG system has two recurring costs: the tokens sent to the model on every query, and the engineering time spent making answers trustworthy. The toolkit is built to reduce both.
Structure-preserving chunks reduce retries. Plain-text chunking loses headings, breaks tables, and drops page boundaries. The result is answers that are wrong in ways users notice, followed by re-prompting and manual verification. Dan's Markdown chunks keep headings, table rows, and page information intact and self-describing, so the first answer is more often the right one.
Enrichment is targeted and budgeted. Ven does not blindly pad the context. It completes only what was cut (a table, a set of revisions) and adds pages or sections only when your application asks, all within a token budget you set. Every expansion method reports how many tokens it added, so you can tune the cost-versus-context trade-off per query type.
The AI document filter sends whole documents when that is cheaper. Ven builds a per-collection index — a compact summary of each document's title, headings, and key terms, typically a few percent of the collection's tokens. A low-cost model scores the documents against the question, and when a handful clearly apply, the application sends those documents whole and skips retrieval and enrichment entirely. In our demo collection of roughly 167,000 tokens, a targeted question such as "who treated this patient" selects two files and sends about 4,000 tokens to the answering model. Summarize-one-document and compare-two-documents questions behave the same way. Questions that span the collection fall back to standard retrieval automatically.
Small collections can skip retrieval altogether. Below a threshold you choose, the demo sends the entire collection on each query and relies on the model provider's cached-input discount. For a department's policy set or a single product's documentation, that is often the cheapest and most accurate option available.
Citations lower the cost of trust. Every answer can point to document, section, and page. In healthcare, finance, and legal work, an answer that can be verified in ten seconds is worth far more than one that must be checked from scratch.
Licensing that scales with your deployment, not your document count. RAG Document Toolkit is licensed per developer for desktop applications and per server for hosted ones, with unlimited-server and enterprise options. There are no per-document, per-query, or per-seat fees from Sub Systems. One license call covers both Dan and Ven, and renewals are discounted.
Integration measured in an afternoon. The toolkit ships as NuGet packages for .NET 9.0 and up, with a Windows Forms demo and its source. The demo shows the complete flow — import, convert, store, retrieve, enrich, filter, answer — in a handful of steps, and its retrieval, enrichment, and filter settings double as calibration guidance for your own application.
What the toolkit does not do
It is worth being clear. RAG Document Toolkit is not a chatbot and not a hosted search service. It does not store your documents, call a model on your behalf, or decide who may see what. It converts and enriches your documents so the application you build — on your database, with your model provider, under your access rules — answers well and cheaply. PDF and Excel input are planned for a future release; today the toolkit reads DOCX, RTF, HTML, text, and Markdown.
Getting started
Download the evaluation version, open the demo with a folder of your own documents, and ask it the questions your customers or employees ask you. The single-file help covers the license, a six-step code example, and the complete Dan and Ven API reference.
Sub Systems, Inc. has built document components for .NET and Windows developers for more than 36 years. Questions about RAG Document Toolkit, enterprise licensing, or your particular document set: info@subsystems.com or 512-733-2525.