Skip to content
Contents

Section 2

Methodology

How each number is produced, what it depends on, and where it can go wrong.

Metrics are implemented explicitly rather than through an evaluation framework. The scoring logic is the credibility of the project, so it is written to be read. Start with or .

2.1 Metrics

Grouped by family. The three-letter tag before each name is the family; the colour reinforces it but never carries it alone.

Retrieval

  • Proportion of retrieved chunks that are actually relevant.
  • Mean reciprocal rank of the first relevant chunk. Rewards getting one right answer high.
  • Discounted cumulative gain over graded relevance, normalised against the ideal ranking.
  • Proportion of queries where at least one chunk graded 2 or higher appears in the top 10.
  • The same measure at rank 5. Harder, and more sensitive to reranking.

Grounding

  • Agreement with the reference answer, judged against a published rubric.
  • Whether the answer addresses the question that was actually asked.
  • Proportion of citations that resolve to a chunk actually supporting the cited claim.
  • Proportion of factual claims that carry a citation at all.
  • Proportion of generated claims entailed by the retrieved chunks.

Abstention

  • Proportion of genuinely unanswerable queries the system correctly declines.

Performance

  • 95th percentile end-to-end time, retrieval through verification.

Cost

  • Token counts multiplied by the rates in the run config. Local models cost zero, stated openly.
Retrieval quality feeds grounding and abstention. A retrieval failure therefore moves faithfulness, citation accuracy and abstention numbers without any of those components being at fault.RetrievalRETGroundingGRDAbstentionABSPublished answer qualityA retrieval failure moves grounding and abstention numbers even when neither component is at fault.That is why access-control effects are reported separately, not folded into headline results.

Point at a node for the metrics it carries and what moves them.

Figure 8. Retrieval quality feeds grounding and abstention, so a retrieval failure moves numbers in components that are not themselves at fault.

2.2 Confidence intervals

Every quality metric carries a bootstrapped 95% confidence interval over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks, and rightly.

2.3 One factor at a time

Each published variant changes exactly one factor from the baseline, so every difference is attributable to a single change rather than to a bundle of them.

2.4 Chunking

Two strategies behind one interface, so the comparison is a run variable rather than a rewrite. Both preserve page and section provenance, which is what lets a citation resolve to a location a reader can check.

The same document chunked two ways. Fixed windows with overlap cut across section boundaries; section-aware packing never crosses one, so a chunk can be shorter but never straddles two clauses.document, with section boundariesfixed-512, 15% overlapwindows overlap and cut across boundaries — a clause can be splitsection-awarenever crosses a boundary — chunks vary in length, clauses stay whole
Figure 6. The same document under both strategies. Fixed windows overlap and cut across boundaries; section-aware packing never crosses one.

2.5 How the golden set is built

The query set is model-drafted and partially human-verified. It is not hand-labelled and is not described as such. A stratified quarter of it — 25% of each category independently, not a random draw — is graded by hand, and the agreement rate between the model judgments and those human grades is published.

Queries are drafted by one model, graded by a second model that never sees which passages the query was written from, then a stratified quarter of the set is graded by a human. The agreement rate compares the second model against the human.Draftqueries + sourcesJudgegrades 0–3Human samplestratified 25%Agreementpublishedthe judge never sees the drafter's labelsonly the question and the passage text cross this lineIf the judge could see them it would agree with them, and the agreement ratewould measure conformity rather than correctness.
Figure 5. The judge never sees which passages a query was drafted from. Only the question and the passage text cross that line.

2.6 Known limitations

Stated here rather than discovered by readers. Each is added as it is established by measurement.

  • The golden set is model-drafted and only partially human-verified. The agreement rate is published beside it.
  • Access control applies differently to the two retrieval paths. Lexical search filters before ranking and loses no recall. Vector search filters after approximate selection, so a restrictive role can lose candidates exact search would have returned. Measured, and reported separately from headline numbers.
  • CUAD tenants are contract type, not counterparty. The 510 contracts come from 463 distinct filers, so per-counterparty tenants would hold about 1.1 documents each and isolation could not be tested meaningfully. Contract type is bounded, deterministic from the source, and semantically coherent, so contracts sharing a tenant genuinely resemble each other. Tenants below a chunk floor are folded into a semantic sibling for the same reason, and every document records both its final tenant and the unmerged original.
  • Two document shapes from one publisher, not three independent sources. CUAD contracts are themselves drawn from EDGAR, so they are a different shape from the same publisher. Deduplication against the filing set is mandatory rather than optional.

ABS Abstention

Abstention (correct)

What it measures

Proportion of genuinely unanswerable queries the system correctly declines.

Formula

Abst = |{ q in U : system declined }| / |U|
U
the unanswerable subset of the query set
declined
the abstention gate fired rather than an answer being produced

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · abstention_correct()

Arrives in Phase 4.

What can go wrong

Related — Abstention family

faithfulness · recall_at_10

GRD Grounding

Answer correctness

What it measures

Agreement with the reference answer, judged against a published rubric.

Formula

correctness = judge(answer, reference) in {0, 1}
reference
the golden set answer, model-drafted and partially human-verified

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · answer_correctness()

Arrives in Phase 4.

What can go wrong

Related — Grounding family

answer_relevance · faithfulness

GRD Grounding

Answer relevance

What it measures

Whether the answer addresses the question that was actually asked.

Formula

relevance = judge(answer, question) in {0, 1}
judge
LLM against a published rubric; the model reference is pinned per run

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · answer_relevance()

Arrives in Phase 4.

What can go wrong

Related — Grounding family

answer_correctness

GRD Grounding

Citation accuracy

What it measures

Proportion of citations that resolve to a chunk actually supporting the cited claim.

Formula

A = |{ citations resolving to a supporting chunk }| / |citations|
resolve
the cited chunk id exists in the index for this run
supporting
string containment first; LLM judge only where containment fails

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

platform/generation/verify.py · citation_accuracy()

Arrives in Phase 2.

What can go wrong

Related — Grounding family

citation_coverage · faithfulness

GRD Grounding

Citation coverage

What it measures

Proportion of factual claims that carry a citation at all.

Formula

C = |{ claims with >= 1 citation }| / |factual claims|
factual claim
an assertion that could be checked against a source

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

platform/generation/verify.py · citation_coverage()

Arrives in Phase 2.

What can go wrong

Related — Grounding family

citation_accuracy · faithfulness

RET Retrieval

Context precision

What it measures

Proportion of retrieved chunks that are actually relevant.

Formula

P(q) = |{ c in top_k(q) : grade(c) >= 2 }| / k
k
retrieval depth, recorded per run in the config

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · context_precision()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_10 · faithfulness

CST Cost

Cost per query

What it measures

Token counts multiplied by the rates in the run config. Local models cost zero, stated openly.

Formula

cost = (in_tokens * rate_in + out_tokens * rate_out) / 1e6
rate
from configs/pricing.yaml, recorded per run
local models
0.00 by definition; hardware cost is not amortised into this number

US dollars per query. Zero for fully local runs, which is stated rather than hidden.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

Reported as a point estimate. This is a measured quantity of the run rather than a sample statistic over queries.

How this project computes it

harness/eval/cost.py · cost_per_query()

Arrives in Phase 4.

What can go wrong

Related — Cost family

latency_p95_s

GRD Grounding

Faithfulness

What it measures

Proportion of generated claims entailed by the retrieved chunks.

Formula

F(a) = |{ claims in a entailed by context }| / |claims in a|
claims
atomic assertions extracted from the answer
entailed
judged by an LLM against a published rubric, not string match

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · faithfulness()

Arrives in Phase 4.

What can go wrong

Related — Grounding family

citation_accuracy · citation_coverage · answer_correctness

PRF Performance

p95 latency

What it measures

95th percentile end-to-end time, retrieval through verification.

Formula

p95 = quantile(latencies, 0.95) over 3 repeats per query
end-to-end
retrieval, rerank, generation, citation resolution, verification
repeats
three per query; the distribution is over all of them

Wall clock seconds. Lower is better. Machine-dependent, so it is only comparable within a run set.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

Reported as a point estimate. This is a measured quantity of the run rather than a sample statistic over queries.

How this project computes it

harness/eval/runner.py · latency_percentiles()

Arrives in Phase 4.

What can go wrong

Related — Performance family

cost_per_query

RET Retrieval

MRR@10

What it measures

Mean reciprocal rank of the first relevant chunk. Rewards getting one right answer high.

Formula

MRR@k = (1/|Q|) * sum_q 1 / rank_first_relevant(q)
rank_first_relevant(q)
1-indexed rank of the first chunk graded >= 2, or 0 contribution if none

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · mrr_at_k()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_10 · ndcg_at_10

RET Retrieval

NDCG@10

What it measures

Discounted cumulative gain over graded relevance, normalised against the ideal ranking.

Formula

DCG@k = sum_i (2^grade(c_i) - 1) / log2(i + 1)
NDCG@k = DCG@k / IDCG@k
c_i
the chunk at rank i
IDCG@k
DCG of the best possible ordering
grade
0-3, so a 3 counts 7x a 1

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · ndcg_at_k()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_10 · mrr_at_10

RET Retrieval

Recall@10

What it measures

Proportion of queries where at least one chunk graded 2 or higher appears in the top 10.

Formula

R@k = |{ q in Q : max grade(c) >= 2 for c in top_k(q) }| / |Q|
grade(c)
graded relevance 0-3 from the golden set
top_k(q)
the k chunks the retriever ranked highest for q
Q
the answerable query set

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · recall_at_k()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_5 · ndcg_at_10 · mrr_at_10 · context_precision

RET Retrieval

Recall@5

What it measures

The same measure at rank 5. Harder, and more sensitive to reranking.

Formula

R@5 = |{ q in Q : max grade(c) >= 2 for c in top_5(q) }| / |Q|
top_5(q)
the five chunks ranked highest for q

Mathematical range. No typical value is stated: it would be a number with no run behind it.

Worked example

From the golden set, cross-document, query cross-document-000 8 human-verified judgments.

What was the total operating income for the North America segment in fiscal 2023, and which XBRL financial taxonomy members are associated with cash flow hedging using foreign exchange contracts during the same fiscal period?

  • grade 3edgar-0000829224-000082922424000057::fixed-512::00015
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00108
  • grade 3edgar-0000320187-000032018725000047::fixed-512::00183
  • grade 3edgar-0000320187-000032018724000044::fixed-512::00185
  • grade 3edgar-0000796343-000079634324000006::fixed-512::00150
  • grade 3edgar-0000796343-000079634325000004::fixed-512::00152

Confidence interval

95% interval by bootstrap resampling over 1,000 resamples of the query set. A point estimate over a few hundred queries without an interval is the first thing a reviewer attacks.

How this project computes it

harness/eval/scorers.py · recall_at_k()

Arrives in Phase 4.

What can go wrong

Related — Retrieval family

recall_at_10 · ndcg_at_10