Blog Retrieval

Retrieval that knows when to say “I don’t know”

A RAG system that always answers is a liability. The useful ones retrieve carefully, cite what they used, and decline when the evidence isn’t there.

Fig. 0Many sources, one grounded answer

Retrieval-augmented generation makes a simple promise: answer questions from your own documents instead of the model’s memory. In practice, most RAG systems fail the same way. When retrieval comes back with the wrong passages, the model writes a fluent, confident answer anyway.

The fix isn’t a better prompt. It’s treating retrieval as a system you measure and tune, and treating “I don’t know” as a feature you design for.

Most RAG failures are retrieval failures

When an answer is wrong, look at what the model was given before you look at what it wrote. Very often the right passage wasn’t retrieved at all, or it was retrieved but ranked below noise, or it was split mid-sentence when the documents were chunked. The model did its best with bad material.

That’s good news, because retrieval is easier to measure than generation. You can build a labeled set of questions with the passages that answer them, and score retrieval on its own, before a model writes a word.

Chunk for meaning, not for size

Fixed-size chunks are the default in most tutorials and the cause of many quiet failures. A 500-token window doesn’t know that a table, a clause or a numbered procedure is one unit. Chunk along the document’s own structure instead:

  • split on headings and sections, and keep each chunk’s heading path with it (“Refunds › International orders › Timelines”);
  • keep tables, lists and code blocks whole;
  • store the metadata you’ll want to filter on later: document type, owner, effective date, region and access level.

Access level deserves its own mention. Filter by the user’s permissions at retrieval time, so a passage the user isn’t allowed to see never reaches the model in the first place.

Hybrid retrieval, fused

Embedding search is good at meaning and bad at exact terms: product codes, error messages, clause numbers, names. Keyword search (BM25) is the opposite. Run both and merge the results.

Reciprocal rank fusion is a simple, robust way to merge. It uses only each document’s position in each list, so it doesn’t matter that keyword scores and vector similarities aren’t on the same scale.

fusion.py
from collections import defaultdict


def reciprocal_rank_fusion(rankings: list[list[str]], k: int = 60) -> list[tuple[str, float]]:
    """Merge ranked lists of document ids, each ordered best first.

    A document's fused score is the sum of 1 / (k + rank) over every list it
    appears in. k damps the advantage of the very top ranks; 60 is the value
    from the original paper and a sensible default.
    """
    scores: dict[str, float] = defaultdict(float)
    for ranking in rankings:
        for rank, doc_id in enumerate(ranking, start=1):
            scores[doc_id] += 1.0 / (k + rank)
    return sorted(scores.items(), key=lambda item: item[1], reverse=True)


# Keyword and vector search each return their top 100; keep 50 for reranking.
fused = reciprocal_rank_fusion([keyword_ids, vector_ids])
candidates = [doc_id for doc_id, _ in fused[:50]]

Rerank, then cut

Fusion gives you a good candidate list. A cross-encoder reranker then reads each candidate together with the question and scores how well it answers it. Rerankers are slower than first-stage search, which is why they only see a short list, but they’re much better at telling a passage that mentions the right words from one that actually answers the question.

Then cut, hard. Pass the model the few passages that clear a relevance threshold, not a fixed top ten. More context isn’t free: irrelevant passages dilute the good ones and give the model material to misuse.

Make abstaining a feature

If no passage clears the threshold, the system shouldn’t fall back on the model’s general knowledge. It should say it couldn’t find the answer in the documents, and offer a next step: rephrase, narrow the question, or route it to the person who owns the topic.

A system that says “I don’t know” when it doesn’t is more useful than one that’s right 95% of the time and never tells you which 5% to doubt.

Abstaining needs its own evaluation, because it can fail in both directions. Measure how often the system declines questions it should have answered, and how often it answers questions the documents can’t support. Tune the threshold on that trade-off, slice by slice.

Cite everything

Every claim in an answer should point to the passage it came from, as a link the user can open. Citations do two jobs. They let users check the answer, which builds trust faster than any accuracy statistic. And they give you a cheap automated check: confirm that each cited passage supports the sentence that cites it, and flag answers where it doesn’t.

Measure retrieval on its own

Keep a labeled set of questions with the passages that answer them, and track retrieval separately from answer quality:

MetricThe question it answers
Recall@kIs a correct passage anywhere in the top k?
MRRHow high does the first correct passage rank?
Context precisionWhat share of the passages sent to the model were relevant?
Abstention accuracyWhen the documents can’t answer, does the system decline?

When answer quality drops, these numbers tell you whether to fix retrieval or generation, which saves days of guessing. The same set doubles as the regression suite for every change to chunking, embeddings or ranking; more on that in Evals before features.