Section 4
Failure modes
Ways a retrieval benchmark silently produces the wrong number, what each one corrupts, and how this project detects or prevents it.
5 of these were hit in the course of building this project and are marked as such. The rest are designed out, and the entry says how. The distinction matters: a list that presents prevented risks and observed bugs identically overstates what was actually found.
None of them announce themselves. Each one produces output that parses, counts that look plausible, and a run that completes successfully.
ANN post-filtering under Row-Level Security
Observed hereAn approximate nearest-neighbour index returns its ef_search closest vectors, and only then does the access policy discard the ones this role may not see. Nothing refills the discarded slots.
Why it is invisible
The query succeeds. It returns fewer rows, or none, and an empty result is indistinguishable from 'no evidence exists'.
What it corrupts
- recall_at_10
- recall_at_5
- ndcg_at_10
- mrr_at_10
- abstention_correct
How this project handles it
Measured against exact search under the identical policy. Fixed by partitioning per tenant so the index holds only permitted rows; iterative scan recovers part of it but not all.
Evidence
Measured on 36,000 chunks across 40 tenants: recall fell from 0.905 unrestricted to 0.023 for a single-tenant role, with 39 of 40 queries returning nothing while exact search returned a full ten. Raising ef_search from 40 to 800 changed nothing.
Element identity reuse during parsing
Observed hereSome parsers create element objects on demand and free them. Identities collected during one traversal are reused by unrelated elements in the next, so an identity-based set matches the wrong nodes.
Why it is invisible
Parsing succeeds and produces a plausible number, just a wrong one.
What it corrupts
- Corpus page count
How this project handles it
Elements are tested directly during a single traversal rather than pre-collecting identities, and page counts are re-derived independently and asserted.
Evidence
One filing recorded 535 pages against 162 actual page breaks — a fabricated figure reaching a published field through a bug rather than a decision.
MAX_TOKENS truncation parsed as a record
Observed hereA model that reasons before answering charges those tokens against the output budget. The budget runs out mid-JSON and the response arrives truncated but syntactically parseable up to the cut.
Why it is invisible
The call returns 200 with content. Only finishReason says anything, and nothing checks it.
What it corrupts
- faithfulness
- citation_accuracy
- citation_coverage
- answer_correctness
How this project handles it
finishReason == MAX_TOKENS is raised as an error and retried; thinking level is pinned so the budget is spent on output.
Evidence
Caught by a test asserting the client retries rather than returns: the truncated payload was a half-written question object that would have been saved as a real golden-set record.
The planner flip
Observed hereThe same query can execute as an exact sequential scan or as an approximate index scan depending on table statistics. The planner has no model of ef_search, so its row estimate for the index scan is fiction.
Why it is invisible
Recall moves upward when the plan flips to exact, so it never looks like a bug. A run simply scores higher on one machine than another.
What it corrupts
- recall_at_10
- ndcg_at_10
- mrr_at_10
- latency_p95_s
How this project handles it
Pin the plan and assert the executed plan shape in the harness; record it in the results JSON and fail the run on mismatch. Not yet built.
Evidence
Observed while measuring RLS behaviour: the same query planned as a sequential scan on one table and an HNSW index scan on another, with no error in between.
Truncated identifier collision
Observed hereA document id derived by slugging and truncating a filename is lossy. Two distinct source files can map to the same id, and the second silently overwrites the first.
Why it is invisible
Nothing raises. The only symptom was two counts on different lines of a log disagreeing.
What it corrupts
- Corpus document count
- every metric measured over the corpus
How this project handles it
Identifiers now carry a hash of the full source path, and registration refuses to overwrite an existing id belonging to a different document, recording a collision instead.
Evidence
26 real contracts were dropped this way. Because the resume check found the first file's entry intact, the later files were never extracted to disk either.
Superuser bypasses Row-Level Security
Prevented by designA PostgreSQL superuser is exempt from row-level security. Connect as postgres and every policy silently does nothing.
Why it is invisible
No error, no warning. Policies simply do not apply, so every access-control test passes for the wrong reason.
What it corrupts
- Every published security result
How this project handles it
The application role is NOSUPERUSER NOBYPASSRLS, verified rolsuper=f and rolbypassrls=f, and the connection string is fixed in configuration.
Evidence
Verified directly against the running database before any security measurement was taken.
Table owner bypasses Row-Level Security
Prevented by designA table's owner is also exempt unless FORCE ROW LEVEL SECURITY is set. relforcerowsecurity defaults to false.
Why it is invisible
Same silence as the superuser case, different cause. Migrations and ops queries running as the owner see everything.
What it corrupts
- Every published security result
How this project handles it
ALTER TABLE ... FORCE ROW LEVEL SECURITY on every table carrying a policy.
Evidence
Confirmed while building the RLS harness: running the probe as the table owner returned recall 0.95 with no restriction applied, because the owner is exempt.
The judge sees the drafter's labels
Prevented by designIf the model grading relevance can see which passages the query was written from, it agrees with them. The agreement rate then measures conformity rather than correctness.
Why it is invisible
The number goes up. A high agreement rate looks like a well-constructed set.
What it corrupts
- recall_at_10
- ndcg_at_10
- context_precision
- faithfulness
- answer_correctness
How this project handles it
The judging prompt interpolates only the question and the passage text. Source membership is recorded in the output but never sent to the judge.
Evidence
Verified by inspecting the compiled prompt template: the only interpolated fields are the relevance statement, the question, and the passages.
Unpinned model reference
Prevented by designA floating alias such as a -latest tag resolves to whatever the provider currently serves. Re-running months later silently uses a different judge.
Why it is invisible
Everything succeeds. The numbers simply move, and nothing in the output says why.
What it corrupts
- faithfulness
- answer_correctness
- answer_relevance
- cost_per_query_usd
How this project handles it
Every run records the resolved dated snapshot, not the requested name, in the run config and on every record the run produces.
Evidence
Checked against the provider's model list: a -latest alias reports a floating label as its version, while a dated model reports the snapshot it is pinned to.