Proofline pushed me to stop treating every wrong answer as a model problem. If the right evidence never reaches the model, a better prompt can still produce a beautifully written wrong answer. The debugging path has to begin earlier.
THE SIGNAL
When a RAG application gives a bad answer, teams often start by changing the prompt or replacing the model.
Sometimes the model never had the right evidence.
Anthropic’s Contextual Retrieval experiments make the point unusually clearly. In its test setup, the relevant-document top-20 retrieval failure rate started at 5.7%. Contextual embeddings reduced it to 3.7%; adding contextual BM25 reduced it to 2.9%; adding reranking reduced it further to 1.9%.
Those are Anthropic’s experimental results—not a universal benchmark.
But they illustrate an important debugging principle:
RAG quality has a pipeline.
THE QUESTION
When the user says “the AI was wrong,” where did the error actually happen?
A simplified failure chain is:
Source → parsing → chunking → indexing → retrieval → filtering → reranking → context → generation → citation → policy
If a policy document is stale, the model can faithfully repeat stale information.
If the correct API error code never appears in the retrieved context, a stronger prompt cannot recover it reliably.
If the right document is retrieved but the answer invents an unsupported claim, that is a generation/faithfulness problem.
These are different product failures and require different fixes.
THE OBVIOUS ANSWER
“Use a better model.”
Sometimes yes.
But Microsoft’s RAG evaluation documentation explicitly separates retrieval/process evaluation from response-level measures such as groundedness and relevance.
That separation is essential.
I would never want a single “RAG accuracy” number.
THE TENSION
A production support system needs multiple eval layers:
| Layer | What I would measure |
|---|---|
| Corpus | freshness, version correctness, ownership |
| Retrieval | recall@K / NDCG / exact-code retrieval |
| Reranking | most-applicable evidence reaches top context |
| Generation | faithfulness, completeness, unsupported claims |
| Citation | cited source actually supports claim |
| Safety | PII, prompt injection, scope/permission |
| Product | resolution, repeat contact, escalation quality |
OWASP also warns that RAG does not eliminate prompt-injection risk. Retrieved content itself can become an attack surface, which means “grounded” is not automatically “safe.”

MY PRODUCT TAKE
For high-stakes support, I would explicitly separate:
Can I find relevant evidence?
from
Can I answer from that evidence?
from
Am I allowed to answer this question?
That third question matters when a user asks for account-specific state that public documentation cannot provide.
A good refusal can be a successful product outcome.
WHAT I WOULD TEST
I would build an eval set with:
- exact API identifier;
- ambiguous natural-language question;
- stale document;
- conflicting policies;
- account-specific request;
- PII;
- unsupported factual question;
- malicious instruction inside retrieved text;
- answer where citation looks relevant but does not entail the claim.
Then I would classify failures by pipeline stage before making any model change.
WHAT WOULD CHANGE MY MIND
If the corpus is small enough to fit reliably into context, a complex RAG pipeline may be unnecessary. Anthropic itself notes that, for sufficiently small knowledge bases, passing the corpus directly can be simpler.
The point is not to use RAG.
The point is to use the simplest evidence architecture that produces trustworthy answers.
Secondary research
- Anthropic, Contextual Retrieval — https://www.anthropic.com/engineering/contextual-retrieval
- Microsoft Foundry, RAG evaluators — https://learn.microsoft.com/en-in/azure/ai-foundry/concepts/evaluation-evaluators/rag-evaluators?view=foundry
- OWASP LLM01 Prompt Injection — https://genai.owasp.org/llmrisk/llm01-prompt-injection/