Skip to content
Contents

Section 8

About

Zeroth is an open reconstruction of a production confidential-document retrieval platform, rebuilt over public documents so the architecture can be inspected, measured and argued with.

1 · Project overview

The original system was built for an employer over a private corpus and cannot leave. This is a from-scratch rebuild of the same architecture over public documents, together with a public evaluation board that measures it. Every number published here was measured on the corpus described in the methodology and applies only to it.

That framing is technical rather than merely ethical. Retrieval metrics are properties of a corpus-and-query-set pair, not of an architecture, so numbers measured here cannot validate or stand in for numbers measured on a different corpus.

2 · Problem

Retrieval systems are usually reported as a single quality number, and that number hides almost everything that determines whether the system works. It hides whether the retriever surfaced the passage, whether the generator used it, whether the system declined when there was nothing to say, and what access control did to any of it.

The project's objective is narrower and harder than a leaderboard score: measure a full pipeline end to end, publish the confidence intervals, publish the failure modes, and make every figure traceable to committed data so a stranger can re-run it.

3 · How it works

Documents are fetched from three public sources, parsed with page and section provenance preserved, chunked two ways, assigned to tenants and deduplicated. Chunks are embedded and indexed. A query runs through lexical and dense retrieval in parallel, the two ranked lists are fused by rank, a cross-encoder reorders the shortlist, and a language model answers under a schema constraint using only the retrieved passages. Citations are resolved, quotes are verified, and the system abstains when the evidence does not support an answer.

Query flows through lexical and dense retrieval, fused by reciprocal rank fusion, then reranking, constrained generation, citation and quote verification, and an abstention gate. Only the query stage is implemented; the rest arrive in Phase 2.QueryBM25lexicalP2DensepgvectorP2RRFfusionP2Rerankcross-encP2GenerateconstrainedP2Verifycite+quoteP2AbstaingateP2solid = implemented · dashed = planned, with delivering phase

Point at a stage for what it does. Solid = implemented; dashed = planned, with the phase that delivers it.

Figure 1. The pipeline under measurement. Solid stages are implemented; dashed stages are planned, marked with the phase that delivers them.

4 · System architecture

The platform runs locally in Docker. Only the site is publicly hosted, and it is a static export that renders committed JSON and never queries the platform. That split is what keeps hosting free and the public attack surface at zero.

Layer, responsibility and technology
LayerResponsibilityBuilt with
AcquisitionFetch three public sources, rate limited and resumablePython standard library only — urllib, no third-party HTTP client
ParsingExtract text with page and section provenancelxml for filing HTML; hand-written parsers for plain text
ChunkingTwo strategies behind one interfacetransformers tokenizer (bge-small vocabulary)
Embedding384-dimension vectors for every chunkBAAI/bge-small-en-v1.5 on PyTorch, CUDA
Lexical indexBM25 over the chunk corpusHand-rolled, array-backed postings, standard library
Vector indexApproximate nearest-neighbour searchPostgreSQL 16 with pgvector 0.8.6, HNSW
Access controlTenant isolation inside the queryPostgreSQL row-level security, non-superuser role
RerankingReorder the shortlistBAAI/bge-reranker-base cross-encoder
GenerationSchema-constrained answersvLLM 0.27.1, xgrammar constrained decoding
Golden setQuery drafting and relevance judgingGemini, pinned to dated snapshots
SiteStatic export, no runtimeNext.js App Router, TypeScript, Tailwind CSS v4
FiguresDiagrams and chartsHand-written inline SVG — no chart library

5 · The corpus

662 documents, 24,155 pages and 51,310 chunks across 47 tenants, from annual filings, the contract set and s. Corpus id edgar-cuad-rfc-v1.

Raw documents are not committed — the manifest is. It records source, identifier, URL, checksum, page count, licence and tenant per document, which is what lets someone reproduce the corpus without redistributing gigabytes.

6 · Key technical decisions

  • Metrics implemented explicitly, without an evaluation framework. The scoring logic is the credibility of the project, so it is written to be read.
  • Access control enforced in the database, not in application code. A forgotten filter in one query path is a data leak; a row-level security policy applies to every path.
  • Access-control effects reported separately from headline numbers. Approximate search under a policy loses recall, and folding that into a headline figure would misattribute it to retrieval quality.
  • Interactive demos replay committed measurements. The site has no backend, so every control moves over data captured by running the real pipeline offline. Nothing is simulated.
  • No fabricated data anywhere. Empty states say what has not happened yet. An agreement rate is withheld rather than published when the sample cannot support it.

7 · Statement of origin

Zeroth is an open reconstruction of a production confidential-document platform. The original was built for an employer over a private corpus and is not public. This is a from-scratch rebuild of the same architecture over public documents. Every number published here was measured on the public corpus described in the methodology, and applies only to it.

8 · Glossary

Every abbreviation the paper uses, written out. In the prose each one also carries its full form on hover or focus; this is the list for reading straight through.

Family tags

ABS
Abstention
CST
Cost
GRD
Grounding
PRF
Performance
RET
Retrieval

Terms

ANN
Approximate Nearest NeighbourTrades exactness for speed: the index may miss a true neighbour, which is why post-filtering under access control can silently cost recall.
BAAI
Beijing Academy of Artificial IntelligencePublisher of the bge embedding and reranker models used here.
BCP
Best Current PracticeAn RFC sub-series carrying operational practice rather than a wire protocol.
CC BY
Creative Commons AttributionA licence permitting reuse with attribution.
CI
Confidence intervalHere always a bootstrapped 95% interval over 1,000 resamples of the query set.
CUAD
Contract Understanding Atticus DatasetA set of commercial contracts with expert clause annotations, released by The Atticus Project.
CUDA
Compute Unified Device ArchitectureNVIDIA's GPU computing platform.
DCG
Discounted Cumulative GainSums graded relevance down the ranking, discounting each position logarithmically.
EDGAR
Electronic Data Gathering, Analysis, and RetrievalThe SEC's public filing system, and the source of the 10-K filings in the corpus.
HNSW
Hierarchical Navigable Small WorldThe graph index used for approximate nearest-neighbour search over the embedding vectors.
IDCG
Ideal Discounted Cumulative GainThe DCG of the best possible ordering, used as the normaliser.
IETF
Internet Engineering Task Force
RAG
Retrieval-Augmented GenerationGeneration conditioned on documents fetched at query time rather than on model weights alone.
RFC
Request for CommentsThe document series in which internet standards are published.
SEC
US Securities and Exchange Commission
TLS
Transport Layer Security

9 · Author and links

Built by Anant Sharma. AI Engineer who builds production Python systems that think in steps — agentic workflows, retrieval pipelines, and orchestration middleware where every output is gated by automated evals before it ships. A year and more turning generative AI research into shipped infrastructure. The project carries the platform, the harness, the corpus manifest and this site.

ABS Abstention

Abstention (correct)

What it measures

Proportion of genuinely unanswerable queries the system correctly declines.

Formula

Abst = |{ q in U : system declined }| / |U|
U
the unanswerable subset of the query set
declined
the abstention gate fired rather than an answer being produced

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · abstention_correct()

Arrives in Phase 4.

What can go wrong

Related — Abstention family

faithfulness · recall_at_10

GRD Grounding

Answer correctness

What it measures

Agreement with the reference answer, judged against a published rubric.

Formula

correctness = judge(answer, reference) in {0, 1}
reference
the golden set answer, model-drafted and partially human-verified

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · answer_correctness()

Arrives in Phase 4.

What can go wrong

Related — Grounding family

answer_relevance · faithfulness

GRD Grounding

Answer relevance

What it measures

Whether the answer addresses the question that was actually asked.

Formula

relevance = judge(answer, question) in {0, 1}
judge
LLM against a published rubric; the model reference is pinned per run

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · answer_relevance()

Arrives in Phase 4.

What can go wrong

Related — Grounding family

answer_correctness

GRD Grounding

Citation accuracy

What it measures

Proportion of citations that resolve to a chunk actually supporting the cited claim.

Formula

A = |{ citations resolving to a supporting chunk }| / |citations|
resolve
the cited chunk id exists in the index for this run
supporting
string containment first; LLM judge only where containment fails

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

platform/generation/verify.py · citation_accuracy()

Arrives in Phase 2.

What can go wrong

Related — Grounding family

citation_coverage · faithfulness

GRD Grounding

Citation coverage

What it measures

Proportion of factual claims that carry a citation at all.

Formula

C = |{ claims with >= 1 citation }| / |factual claims|
factual claim
an assertion that could be checked against a source

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

platform/generation/verify.py · citation_coverage()

Arrives in Phase 2.

What can go wrong

Related — Grounding family

citation_accuracy · faithfulness

RET Retrieval

Context precision

What it measures

Proportion of retrieved chunks that are actually relevant.

Formula

P(q) = |{ c in top_k(q) : grade(c) >= 2 }| / k
k
retrieval depth, recorded per run in the config

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · context_precision()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_10 · faithfulness

CST Cost

Cost per query

What it measures

Token counts multiplied by the rates in the run config. Local models cost zero, stated openly.

Formula

cost = (in_tokens * rate_in + out_tokens * rate_out) / 1e6
rate
from configs/pricing.yaml, recorded per run
local models
0.00 by definition; hardware cost is not amortised into this number

US dollars per query. Zero for fully local runs, which is stated rather than hidden.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

Reported as a point estimate. This is a measured quantity of the run rather than a sample statistic over queries.

How this project computes it

harness/eval/cost.py · cost_per_query()

Arrives in Phase 4.

What can go wrong

Related — Cost family

latency_p95_s

GRD Grounding

Faithfulness

What it measures

Proportion of generated claims entailed by the retrieved chunks.

Formula

F(a) = |{ claims in a entailed by context }| / |claims in a|
claims
atomic assertions extracted from the answer
entailed
judged by an LLM against a published rubric, not string match

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · faithfulness()

Arrives in Phase 4.

What can go wrong

Related — Grounding family

citation_accuracy · citation_coverage · answer_correctness

PRF Performance

p95 latency

What it measures

95th percentile end-to-end time, retrieval through verification.

Formula

p95 = quantile(latencies, 0.95) over 3 repeats per query
end-to-end
retrieval, rerank, generation, citation resolution, verification
repeats
three per query; the distribution is over all of them

Wall clock seconds. Lower is better. Machine-dependent, so it is only comparable within a run set.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

Reported as a point estimate. This is a measured quantity of the run rather than a sample statistic over queries.

How this project computes it

harness/eval/runner.py · latency_percentiles()

Arrives in Phase 4.

What can go wrong

Related — Performance family

cost_per_query

RET Retrieval

MRR@10

What it measures

Mean reciprocal rank of the first relevant chunk. Rewards getting one right answer high.

Formula

MRR@k = (1/|Q|) * sum_q 1 / rank_first_relevant(q)
rank_first_relevant(q)
1-indexed rank of the first chunk graded >= 2, or 0 contribution if none

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · mrr_at_k()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_10 · ndcg_at_10

RET Retrieval

NDCG@10

What it measures

Discounted cumulative gain over graded relevance, normalised against the ideal ranking.

Formula

DCG@k = sum_i (2^grade(c_i) - 1) / log2(i + 1)
NDCG@k = DCG@k / IDCG@k
c_i
the chunk at rank i
IDCG@k
DCG of the best possible ordering
grade
0-3, so a 3 counts 7x a 1

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · ndcg_at_k()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_10 · mrr_at_10

RET Retrieval

Recall@10

What it measures

Proportion of queries where at least one chunk graded 2 or higher appears in the top 10.

Formula

R@k = |{ q in Q : max grade(c) >= 2 for c in top_k(q) }| / |Q|
grade(c)
graded relevance 0-3 from the golden set
top_k(q)
the k chunks the retriever ranked highest for q
Q
the answerable query set

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · recall_at_k()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_5 · ndcg_at_10 · mrr_at_10 · context_precision

RET Retrieval

Recall@5

What it measures

The same measure at rank 5. Harder, and more sensitive to reranking.

Formula

R@5 = |{ q in Q : max grade(c) >= 2 for c in top_5(q) }| / |Q|
top_5(q)
the five chunks ranked highest for q

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · recall_at_k()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_10 · ndcg_at_10