What each test actually asks, how many items it has, the seed that regenerates it, and
the honest caveat you need before quoting it. Difficulty knobs and pass rules were
fixed in the generator before any engine was scored. A benchmark nothing fails
is broken, so these were tuned to leave visible headroom, then frozen.
Say-nothing bluff traps
A fabricated case name and reporter cite, verified absent from the index, attached to a deliberately wrong proposition of law, with the prompt baiting the model to name and cite controlling authority. Pass = silence, refusal, or citing only authority the index can verify. Fail = it emits any citation the index cannot verify, extracted with the real product citation extractor.
- Items600 generated · 300 drawn for the lineup · 230 finished by every engine (the common set)
- Seedaxis seed 28908257 (from master 20260709) · lineup draw seed 20260709
- Scoringdeterministic, no model-in-the-loop grading
- Costneeds an LLM; the retrieval side of the battery is $0
Read it honestly This axis is currently saturated: every engine we sell refuses every trap, so the column separates nobody. That is a real result, not a marketing one. It means the traps at this difficulty are no longer discriminating and the next version has to get harder, not that bluffing is a solved problem.
hard-battery v1 · generated 2026-07-09 · items regenerate byte-for-byte from the seed
Statute pinpoint, bare-model recall
Given a described procedural moment (a deterministic paraphrase of the statute's own catchline), name the exact governing Texas Code article or section. Ground truth is the statute's own section number, read out of the corpus. No retrieval, no lookup, this is memory.
- Items307 real TX statutes · 300 drawn · 268 in the common set
- Seedaxis seed 26421691 (from master 20260709)
- Pass ruleexact-section match against the statute's own number
- Whole library, same itemsrecall@1 186/307 · 60.6% (55.0 to 65.9) · recall@5 253/307 · 82.4% (77.8 to 86.3) (statutes, cases and guidance searched together; the headline card above searches statutes only)
Read it honestly This is the axis people misread. It measures what a bare model has memorized, and almost nothing memorizes TX Code section numbers, so the scores are low by construction. It is not the product's score. In the product the statute is looked up, and the library's own recall on these same items is the row above. A model that abstains here is behaving correctly; read commit rate and precision together before calling anything a win.
hard-battery v1 · generated 2026-07-09 · retrieval figures from the live fused index, all sources
Citation integrity, the hallucination defense
Real (quote, citation) pairs from the corpus. Half genuine; half corrupted with hard near-misses, page delta, volume delta, reporter swap, and sibling-indexed (a different real, indexed case). Accept or reject the pairing. Ground truth is the corruption we applied, so the labels cannot be argued with.
- Items4,000 pairs · engine-independent (no LLM in the loop)
- Seedaxis seed 27063130 (from master 20260709)
- Shipped string gateaccuracy 16,621/20,000 · 83.1% (82.6 to 83.6) · false-accept 3,379/10,000 · 33.8% (32.9 to 34.7)
- Retrieval gateaccuracy 14,250/20,000 · 71.2% (70.6 to 71.9) · false-accept 7/10,000 · 0.1% (0.0 to 0.1) · false-reject 5,743/10,000 · 57.4% (56.5 to 58.4)
- Sibling-indexed fakes caught0/2,997 · 0.0% (0.0 to 0.1)
Read it honestly The ugly number is the point: a string-only checker waves through a real fraction of fakes, and catches none of the sibling-indexed family, a citation that exists but is the wrong case. Existence is not support. The retrieval gate closes false-accepts and pays for it in false-rejects; that trade is reported, not hidden. At full-corpus scale (6,664,575 chunks) the string gate gets worse: false-accept 783/2,000 · 39.1% (37.0 to 41.3).
hard-battery v1 · generated 2026-07-09 · engine-independent: same stack in every product tier
Needle retrieval under distractors
A holding sentence, put through a deterministic paraphrase, synonym-swapped, clauses reordered, case names, citations and numbers masked so recall cannot cheat on surface handles, is the query. The target is its source opinion, competing against every other opinion in the index.
- Items2,500 paraphrased holdings · engine-independent (retrieval, not the model)
- Seedaxis seed 25178408 (from master 20260709)
- 100k index (v1, hybrid)recall@1 1,131/2,500 · 45.2% (43.3 to 47.2)
- Full corpus (6,664,575 chunks, vector)recall@1 638/2,499 · 25.5% (23.9 to 27.3) · recall@5 1,031/2,499 · 41.3% (39.3 to 43.2)
- Live fused index (v3)recall@1 468/2,500 · 18.7% (17.2 to 20.3) · recall@5 592/2,500 · 23.7% (22.1 to 25.4)
Read it honestly Recall goes down as the library grows, and we publish it going down. A bigger corpus means more same-doctrine near-neighbours competing with the true source. This axis scores the retrieval stack, not any model in the table above, no model choice moves it.
hard-battery v1 · generated 2026-07-09 · paraphrase knobs frozen before any engine was scored
Live-coaching moments (GW-14x)
Replayed golden moments from real hearings. Each generated coaching card is scored 0 to 2 on four axes, grounded in the retrieved passage, correctly named legal issue, no invented authority, and a usable next step, for 8 points total, plus a separate say-nothing check and the genius-gate pass rate.
- Items25 golden courtroom moments, scored by a mock judge
- Overall 0 to 8private 14B 6.84 · Opus 4.8 (direct API) 7.32
- Per-axisgrounded: private 14B 1.92, Opus 4.8 (direct API) 2.0 · legal issue: private 14B 1.64, Opus 4.8 (direct API) 1.76 · no invention: private 14B 2.0, Opus 4.8 (direct API) 2.0 · coach next step: private 14B 1.28, Opus 4.8 (direct API) 1.56
Read it honestly Small n and a mock judge. This is the only axis on the page whose grader is not fully mechanical. The Opus figure here is a separate direct-API run, not the chart row above. Treat the coaching column as directional and weight the big-n axes above it.
generated 2026-07-14 · scored with the committed coaching rubric
Question shaping (HEAR-16)
Courtroom speech is not a search query. This battery takes real spoken utterances, hypotheticals, interruptions, half-finished questions, and asks whether rewriting them into a shaped query retrieves the right authority better than sending the raw words.
- Items20 labeled utterances
- Raw utterancerecall 100% · MRR 0.975
- Shaped queryrecall 100% · MRR 0.942
Read it honestly n is tiny, twenty utterances decides nothing, and here shaping did not beat the raw utterance on MRR. It is on this page because it is on disk and it did not go our way; a battery you only publish when it wins is not a battery.
read from the committed HEAR-16 result file at build time
Live ears, word error rate
The streaming transcription that feeds everything else, scored against the official transcript of a real argument after a deterministic structure trim (headers, line numbers, speaker labels and the keyword index removed by anchor, not by hand).
- Word error rate10.6%
- Cost to run it$0.018
- Modelnvidia/nemotron-speech-streaming-en-0.6b
- ReferenceSmith v. Arizona, No. 22-899 (U.S.)
Read it honestly This is the one figure on the page not recomputed at build time. It is quoted as a published fact: streaming transcription scored against the official Supreme Court transcript of No. 22-899 after the deterministic structure trim (16,683 spoken words). The audio shards for that run are kept off the repo, so this page quotes the published number rather than recomputing it. One argument is one argument, it is not a WER claim across courtrooms, accents, or bad audio.
published figure, quoted with its provenance