Engineering Notes · RAG & Document Structure

The Day Vector Search Missed Every Comment in the Document

We build document conversion components at Sub Systems — our AI Markdown DLL converts DOCX, RTF, and HTML into AI-friendly Markdown, preserving the structures that most converters throw away: tables, tracked changes, and reviewer comments. To exercise the library, we maintain a C# document Q&A demo: load a document, chunk it, embed it into a vector store, and let an LLM answer questions from the retrieved chunks. A standard RAG pipeline, in other words.

Recently, that pipeline failed in a way worth writing about — because the failure was not a bug. It was similarity search doing exactly what similarity search does.

The failure

We loaded a large document that carried reviewer comments and asked a natural question:

“What did the reviewers comment on?”

The answer came back thin. The model discussed a couple of passages and missed most of what the reviewers had actually said. When we inspected the retrieved chunks, the cause was plain: retrieval had returned chunks whose body text happened to discuss comments and remarks — and had skipped most of the chunks that actually carried reviewer comments.

The word “comment” appeared in ordinary prose elsewhere in the document. Those chunks embedded strongly against the question. The chunks holding real comment tags did not.

Why this is not a bug

Vector retrieval works by embedding your question and returning the chunks whose embeddings sit closest to it. That mechanism answers one question extremely well: “what text resembles this question?” It cannot, even in principle, answer a different question: “give me all of X.”

Reviewer comments are the worst case for this mechanism:

Raising the retrieval count does not fix this. Top-40 instead of top-20 returns a larger sample with the same bias. A question that demands completeness — all the comments, all the changes a reviewer made — cannot be answered by a mechanism whose job is ranked resemblance.

Completeness is a structural property. And structure is exactly what our converter preserves.

The fix: three lines, no heroics

Every chunk our DLL emits carries a metadata dictionary alongside the text — page numbers, table IDs, revision IDs, comment IDs, author rollups. Our Vector Enrichment Library (Ven) reads that metadata to expand a retrieved chunk selection with the chunks that structure demands: the rest of a table that was split across chunk boundaries, the remaining fragments of a tracked change — or every chunk in the document that carries comment content.

The demo now does this:

// Ask the cheap questions first: does this document even have
// comments or revisions? (Answered from metadata — no model call.)
if (ven.VesHasComments())
    QuestionIsAboutComments = await IsUserQueryAboutDocComments(newUserQuestion);
if (ven.VesHasRevisions())
    QuestionIsAboutRevisions = await IsUserQueryAboutDocRevisions(newUserQuestion);
// If the user is asking about the collaboration layer,
// include it structurally — don't hope similarity finds it.
if (QuestionIsAboutComments) ven.VesAddCommentChunks(exp, "");
if (QuestionIsAboutRevisions) ven.VesAddRevisionChunks(exp, "");

Three things are happening here, and each carries a lesson bigger than this bug:

1. Metadata gates the model calls. VesHasComments() is answered from chunk metadata in microseconds. If the document has no comments — true for the overwhelming majority of documents our customers process — the intent classifier never runs. The rare feature costs nothing when it is absent.

2. A small model classifies intent. When the document does carry comments, a one-word classification by a budget-tier model decides whether the user's question is about them. This costs a few tokens and runs concurrently with our other routing.

3. Structure delivers completeness. VesAddCommentChunks does not search for comment chunks — it knows them, from the comment IDs in the metadata, and includes every one. The result is complete by construction. The model then answers from a context that provably contains everything the reviewers said.

After the change, the same question on the same document returned every comment, attributed to its author, in document order.

The general principle

Retrieval-augmented generation has two find-things mechanisms, and they are not interchangeable:

MechanismAnswersGuarantee
Similarity search“What resembles this question?”Ranked relevance
Structural metadata“What are all the X?”Completeness

Most RAG pipelines only have the first, because most converters emit bare text and give the pipeline nothing structural to act on. Then questions like “list every change Mary made” or “what did reviewers push back on?” fail quietly — the pipeline returns a plausible-looking partial answer, and nobody knows what was missed. In a legal or compliance setting, a change summary that silently omits changes is worse than no answer at all.

The fix is not a smarter embedding model. It is conversion that preserves structure as queryable metadata, and a retrieval layer willing to use it. Similarity search finds what is similar. Only structure can find what is exhaustive — and the questions that matter most in reviewed documents are exhaustive questions.

That is the design philosophy behind our AI Markdown DLL and the Vector Enrichment Library: everything the document knows about itself — its tables, its revisions, its comments, its authors — travels with the chunks, once as text the model can read, and once as metadata the machines can query. The day similarity search missed every comment in the document was the day that second channel paid for itself.