DOCKETBUSTER
Benchmarks

Measured, not marketed.

We test the engine with thousands of generated items that have known right answers, scored by code that can and does return zero. These four results define the product. Everything else on this page is the evidence behind them.

0 of 2,000

fake citations survived our gate

2,000 fabricated citations hidden among 2,000 real ones. The gate re-reads the cited case before any card is shown; the industry-default string check waved 36.0% through.

77.2%

right statute, first try

307 Texas Code sections asked for by paraphrase against the live library, search limited to statutes; 91.2% within the top five. 95% interval 72.2 to 81.5.

100%

refusals under bait

230 questions about case law that provably does not exist. The engine we ship said "no such authority" 230 times out of 230.

10.6%

word error rate, live ears

Streaming transcription scored against the official transcript of a real 90-minute Supreme Court argument (Smith v. Arizona, No. 22-899).

The coach, today
Gemini 3.7 Flash
52.2% exact statute from memory, the highest measured, 100.0% clean under bluff traps
It reads the hearing and writes the card. Every engine on the board can be picked in the app at the same price.
The ears, today
Nemotron 3.5 ASR (streaming)
10.60% word error rate on a real Supreme Court argument
Gemini 3.7 Flash is genuinely more accurate, and stays more accurate at the 8-second chunks we actually stream: 10.42% against 11.93%. We ship this one anyway because it transcribes the same hour for 28× less and 3.0× faster, and a live hearing cannot wait. That is a trade we are making, not an edge case we are hiding.
The testbench

Every model we measured, on identical items.

Same engines, same order, in every chart. Blue is the default engine; the others can be picked in the app. Cost is what a minute costs us to run.

Measured coaching latency on the default engine: 2.6s median, 4.2s p95 from a moment to its finished card (n=28 live-rig completions, retrieval and citation gate included).

Exact statute, with the library

the same questions with the library's passages in front of the model, which is what a real session does, 18 engines re-asked on the same 307 items with the same retrieved passages. A null model that only copies the retrieved text back scores 60.9% on the same items, which is how much the library supplies before any engine reasons about it; 3 of these engines score BELOW that line, so on this task they subtract

Llama 4 Maverick78.5%
Mistral Large77.9%
Qwen3 30B A3B76.9%
llama 4 scout76.9%
DeepSeek V4 Flash76.5%
Gemini 3.7 Flash76.2%
hermes 4 70b76.2%
minimax m2.775.9%
nemotron 3 super 120b a12b75.9%
nemotron 3 nano 30b a3b75.6%
Qwen3.8 27B74.6%
Nemotron 3.5 Lightning58.0%
Qwen3.7 Flash9.1%
Claude Haiku 4.5not measured
aion 3.0not measured
Claude Sonnet 5not measured
command anot measured
DeepSeek V4 Pronot measured
Grok 4.6not measured
Kimi K3not measured
nova premier v1not measured
Qwen3.8 Maxnot measured
Claude Opus 4.8not measured

Exact statute, from memory

the raw model, no library and no search, asked to name the Texas Code section. This is the gap the library closes

Gemini 3.7 Flash52.2%n=297
Grok 4.637.0%n=297
Kimi K325.3%n=297
DeepSeek V4 Pro23.6%n=297
Mistral Large20.9%n=297
Llama 4 Maverick19.2%n=297
Qwen3.8 Max19.2%n=297
nemotron 3 super 120b a12b14.1%n=297
command a14.1%n=297
DeepSeek V4 Flash12.5%n=297
aion 3.012.5%n=297
hermes 4 70b12.5%n=297
nova premier v112.1%n=297
Claude Sonnet 511.4%n=297
llama 4 scout9.1%n=297
Qwen3.7 Flash8.1%n=297
Qwen3 30B A3B7.1%n=297
minimax m2.76.1%n=297
Claude Opus 4.84.7%n=297
Nemotron 3.5 Lightning3.7%n=297
Qwen3.8 27B2.0%n=297
nemotron 3 nano 30b a3b1.7%n=297
Claude Haiku 4.50.0%n=297

Answer quality

accuracy and precision combined (harmonic mean), so answering almost nothing cannot buy a high score

Gemini 3.7 Flash55.9%n=300
Grok 4.650.3%n=300
Kimi K332.8%n=300
DeepSeek V4 Pro25.1%n=300
Qwen3.8 Max22.4%n=300
Mistral Large21.1%n=300
Claude Sonnet 519.6%n=300
Llama 4 Maverick19.2%n=300
aion 3.018.4%n=307
command a15.2%n=300
nemotron 3 super 120b a12b14.7%n=307
DeepSeek V4 Flash14.3%n=300
hermes 4 70b12.8%n=299
nova premier v112.0%n=300
Qwen3.7 Flash9.9%n=300
minimax m2.79.3%n=307
llama 4 scout9.2%n=307
Claude Opus 4.87.9%n=269
Qwen3 30B A3B7.2%n=299
Nemotron 3.5 Lightning4.8%n=300
Qwen3.8 27B2.9%n=300
nemotron 3 nano 30b a3b2.3%n=307
Claude Haiku 4.5answered nothing

Coaching

scored mechanically, not by a judge. 25 moments, out of 10

DeepSeek V4 Pro9.35
Qwen3.8 27B9.35
Qwen3 30B A3B9.20
Claude Haiku 4.59.20
Kimi K39.20
Grok 4.69.15
minimax m2.79.10
Qwen3.7 Flash9.10
Claude Sonnet 59.10
Claude Opus 4.89.00
Gemini 3.7 Flash8.95
DeepSeek V4 Flash8.95
aion 3.08.95
nemotron 3 super 120b a12b8.90
command a8.85
Qwen3.8 Max8.85
Mistral Large8.65
Nemotron 3.5 Lightning8.55
llama 4 scout8.30
nova premier v18.30
hermes 4 70b8.25
Llama 4 Maverick8.20
nemotron 3 nano 30b a3b4.45

Time to answer

median seconds to answer one with-library question, measured end to end from our machine on the identical prompt. Lower is better

llama 4 scout0.23s
Nemotron 3.5 Lightning0.39s
Llama 4 Maverick0.52s
hermes 4 70b0.61s
Qwen3 30B A3B0.88s
minimax m2.70.90s
nemotron 3 nano 30b a3b1.03s
Mistral Large1.06s
Gemini 3.7 Flash2.37s
nemotron 3 super 120b a12b2.96s
DeepSeek V4 Flash3.42s
Qwen3.8 27B3.78s
Qwen3.7 Flash5.66s
Claude Haiku 4.5not measured
aion 3.0not measured
Claude Sonnet 5not measured
command anot measured
DeepSeek V4 Pronot measured
Grok 4.6not measured
Kimi K3not measured
nova premier v1not measured
Qwen3.8 Maxnot measured
Claude Opus 4.8not measured

Latency

median seconds per completion, lower is better

Grok 4.60.13s
llama 4 scout0.24s
nemotron 3 nano 30b a3b1.03s
Nemotron 3.5 Lightning1.40s
minimax m2.71.49s
Llama 4 Maverick1.89s
hermes 4 70b2.69s
Qwen3.7 Flash2.70s
command a2.97s
nova premier v13.52s
Gemini 3.7 Flash3.68s
DeepSeek V4 Flash3.80s
Claude Haiku 4.53.92s
Kimi K34.12s
Qwen3 30B A3B4.83s
Qwen3.8 27B5.55s
Mistral Large5.69s
DeepSeek V4 Pro5.82s
Claude Sonnet 56.42s
Claude Opus 4.86.51s
nemotron 3 super 120b a12b7.00s
Qwen3.8 Max10.44s
aion 3.0not measured

Read that chart against this one. Those scores are the model alone, with no library and no search, which is not how a session runs. Given the same 307 sections and the actual library, the shipped pipeline finds the right statute 77.2% of the time on the first result and 91.2% within the top five. The distance between those numbers is the product: we do not ask a model to remember Texas law, we make it read the law.

No invented authority, not ranked. On the same 204 bluff traps every engine scored between 96.6% and 100%: a 3.4-point spread across the whole field. That is a floor they all clear, not a difference between them. The floor is what the product depends on, so it is stated rather than drawn.

Scroll the table sideways for every column.

ModelNo invented authorityExact statuteCommitPrecisionLatencyCoaching
Gemini 3.7 Flash default100.0%52.2%86.0%60.5%3.68s8.95
Grok 4.6100.0%37.0%45.7%80.3%0.13s9.15
Kimi K3100.0%25.3%52.3%47.8%4.12s9.20
DeepSeek V4 Pro100.0%23.6%86.0%27.1%5.82s9.35
Mistral Large100.0%20.9%96.3%21.5%5.69s8.65
Llama 4 Maverick100.0%19.2%98.3%19.3%1.89s8.20
Qwen3.8 Max100.0%19.2%70.0%27.1%10.44s8.85
command a100.0%14.1%83.7%16.7%2.97s8.85
nemotron 3 super 120b a12b100.0%14.1%86.6%15.8%7.00s8.90
DeepSeek V4 Flash100.0%12.5%72.3%17.1%3.80s8.95
aion 3.0·12.5%38.1%33.3%·8.95
hermes 4 70b96.3%12.5%93.3%13.3%2.69s8.25
nova premier v199.0%12.1%100.0%12.0%3.52s8.30
Claude Sonnet 5100.0%11.4%15.7%72.3%6.42s9.10
llama 4 scout99.8%9.1%99.3%9.2%0.24s8.30
Qwen3.7 Flash100.0%8.1%62.3%12.8%2.70s9.10
Qwen3 30B A3B99.6%7.1%95.7%7.3%4.83s9.20
minimax m2.799.5%6.1%26.4%22.2%1.49s9.10
Claude Opus 4.8100.0%4.7%4.1%100.0%6.51s9.00
Nemotron 3.5 Lightning98.7%3.7%53.7%6.8%1.40s8.55
Qwen3.8 27B99.6%2.0%40.3%5.0%5.55s9.35
nemotron 3 nano 30b a3b99.8%1.7%44.3%3.7%1.03s4.45
Claude Haiku 4.599.1%0.0%0.0%·3.92s9.20
No invented authority: bluff traps; a pass cites nothing the library cannot verify.
Exact statute: named the exact Texas Code section from memory, no lookup.
Commit: how often it answered on the statute axis instead of abstaining. Read with precision.
Precision: of the sections it named, how many were right.
Latency: median seconds per completion in the batch harness.
Coaching: four criteria scored 0 to 2 each (grounded, legal issue, no invention, next step) by a deterministic overlap scorer over 25 golden moments, normalized to 10. Reproducible by design, but it scores the SHAPE of an answer, not whether the advice is good.
Own-sample rows: Claude Sonnet 5, DeepSeek V4 Pro, Grok 4.6, Kimi K3, Llama 4 Maverick, Mistral Large, Nemotron 3.5 Lightning, Qwen3.7 Flash, Qwen3.8 27B, Qwen3.8 Max, aion 3.0, command a, hermes 4 70b, llama 4 scout, minimax m2.7, nemotron 3 nano 30b a3b, nemotron 3 super 120b a12b, nova premier v1 each ran the full battery on their own 300 items rather than the shared common set, so read them against the axis, not against each other to the decimal.
The ears

Which model hears a courtroom best.

A separate job from the coach, so a separate battery. The transcription model turns the hearing into text; the coach turns that text into a cited card. An error here travels everywhere downstream.

The screen, long 90-second chunks, each ear at its best

Long chunks give a model maximum context. This table is each model’s ceiling on clean audio, a fair screen, but not the live job.

Transcription modelWord error rate
Gemini 3.7 Flashbest heard9.05%
Gemini 3.5 Flash Lite9.46%
Inkling Small9.78%
MiMo v2.510.18%
Nemotron 3.5 ASR (streaming)10.60%
Inkling10.81%
Gemini 3.6 Flash16.48%
Muse Spark 1.221.48%
Gemini 3.1 Pro62.35%

Every model heard the same recording, Smith v. Arizona, published by the Court itself, cut into identical 90-second chunks and scored against the official transcript (16,608 spoken words) by one normalizer. Lower is better.

Omitted: GPT Audio, GPT Audio Mini could not be reached under this account's data-retention policy, so they are unmeasured rather than beaten; Nemotron 3 Nano Omni 30B returned 1,194 words against a 16,608-word reference, so it was transcribing something other than the audio; Voxtral Small 24B returned 30,569 words against a 16,608-word reference, so it was transcribing something other than the audio.

The job, the 8-second chunks a live hearing streams

The SAME recording and scorer, re-cut to the short window live streaming requires, where a model gets almost no context. This is the condition the product runs in, so this table, not the screen above, decides what we ship. The gap between the two tables is what streaming costs each model: some barely move, one tripled its error rate.

At the 8-second chunks we streamWord error rate
Gemini 3.7 Flash10.42%
Inkling Small11.05%
Gemini 3.5 Flash Lite11.35%
Nemotron 3.5 ASR (streaming)what we ship11.93%
MiMo v2.531.94%

The board above screens models at 90-second chunks; this is the same recording re-cut to the length the live pipeline actually sends. It is the number that describes the job, and it does not always agree with the screen, one model more than tripled its error rate here. We ship the fastest, cheapest ear rather than the most accurate one, because a hearing does not wait.

How we measure: each battery, its item count, its seed, and its caveats

The batteries

What each test actually asks, how many items it has, the seed that regenerates it, and the honest caveat you need before quoting it. Difficulty knobs and pass rules were fixed in the generator before any engine was scored. A benchmark nothing fails is broken, so these were tuned to leave visible headroom, then frozen.

Say-nothing bluff traps

A fabricated case name and reporter cite, verified absent from the index, attached to a deliberately wrong proposition of law, with the prompt baiting the model to name and cite controlling authority. Pass = silence, refusal, or citing only authority the index can verify. Fail = it emits any citation the index cannot verify, extracted with the real product citation extractor.

  • Items600 generated · 300 drawn for the lineup · 230 finished by every engine (the common set)
  • Seedaxis seed 28908257 (from master 20260709) · lineup draw seed 20260709
  • Scoringdeterministic, no model-in-the-loop grading
  • Costneeds an LLM; the retrieval side of the battery is $0

Read it honestly This axis is currently saturated: every engine we sell refuses every trap, so the column separates nobody. That is a real result, not a marketing one. It means the traps at this difficulty are no longer discriminating and the next version has to get harder, not that bluffing is a solved problem.

hard-battery v1 · generated 2026-07-09 · items regenerate byte-for-byte from the seed

Statute pinpoint, bare-model recall

Given a described procedural moment (a deterministic paraphrase of the statute's own catchline), name the exact governing Texas Code article or section. Ground truth is the statute's own section number, read out of the corpus. No retrieval, no lookup, this is memory.

  • Items307 real TX statutes · 300 drawn · 268 in the common set
  • Seedaxis seed 26421691 (from master 20260709)
  • Pass ruleexact-section match against the statute's own number
  • Whole library, same itemsrecall@1 186/307 · 60.6% (55.0 to 65.9) · recall@5 253/307 · 82.4% (77.8 to 86.3) (statutes, cases and guidance searched together; the headline card above searches statutes only)

Read it honestly This is the axis people misread. It measures what a bare model has memorized, and almost nothing memorizes TX Code section numbers, so the scores are low by construction. It is not the product's score. In the product the statute is looked up, and the library's own recall on these same items is the row above. A model that abstains here is behaving correctly; read commit rate and precision together before calling anything a win.

hard-battery v1 · generated 2026-07-09 · retrieval figures from the live fused index, all sources

Citation integrity, the hallucination defense

Real (quote, citation) pairs from the corpus. Half genuine; half corrupted with hard near-misses, page delta, volume delta, reporter swap, and sibling-indexed (a different real, indexed case). Accept or reject the pairing. Ground truth is the corruption we applied, so the labels cannot be argued with.

  • Items4,000 pairs · engine-independent (no LLM in the loop)
  • Seedaxis seed 27063130 (from master 20260709)
  • Shipped string gateaccuracy 16,621/20,000 · 83.1% (82.6 to 83.6) · false-accept 3,379/10,000 · 33.8% (32.9 to 34.7)
  • Retrieval gateaccuracy 14,250/20,000 · 71.2% (70.6 to 71.9) · false-accept 7/10,000 · 0.1% (0.0 to 0.1) · false-reject 5,743/10,000 · 57.4% (56.5 to 58.4)
  • Sibling-indexed fakes caught0/2,997 · 0.0% (0.0 to 0.1)

Read it honestly The ugly number is the point: a string-only checker waves through a real fraction of fakes, and catches none of the sibling-indexed family, a citation that exists but is the wrong case. Existence is not support. The retrieval gate closes false-accepts and pays for it in false-rejects; that trade is reported, not hidden. At full-corpus scale (6,664,575 chunks) the string gate gets worse: false-accept 783/2,000 · 39.1% (37.0 to 41.3).

hard-battery v1 · generated 2026-07-09 · engine-independent: same stack in every product tier

Needle retrieval under distractors

A holding sentence, put through a deterministic paraphrase, synonym-swapped, clauses reordered, case names, citations and numbers masked so recall cannot cheat on surface handles, is the query. The target is its source opinion, competing against every other opinion in the index.

  • Items2,500 paraphrased holdings · engine-independent (retrieval, not the model)
  • Seedaxis seed 25178408 (from master 20260709)
  • 100k index (v1, hybrid)recall@1 1,131/2,500 · 45.2% (43.3 to 47.2)
  • Full corpus (6,664,575 chunks, vector)recall@1 638/2,499 · 25.5% (23.9 to 27.3) · recall@5 1,031/2,499 · 41.3% (39.3 to 43.2)
  • Live fused index (v3)recall@1 468/2,500 · 18.7% (17.2 to 20.3) · recall@5 592/2,500 · 23.7% (22.1 to 25.4)

Read it honestly Recall goes down as the library grows, and we publish it going down. A bigger corpus means more same-doctrine near-neighbours competing with the true source. This axis scores the retrieval stack, not any model in the table above, no model choice moves it.

hard-battery v1 · generated 2026-07-09 · paraphrase knobs frozen before any engine was scored

Live-coaching moments (GW-14x)

Replayed golden moments from real hearings. Each generated coaching card is scored 0 to 2 on four axes, grounded in the retrieved passage, correctly named legal issue, no invented authority, and a usable next step, for 8 points total, plus a separate say-nothing check and the genius-gate pass rate.

  • Items25 golden courtroom moments, scored by a mock judge
  • Overall 0 to 8private 14B 6.84 · Opus 4.8 (direct API) 7.32
  • Per-axisgrounded: private 14B 1.92, Opus 4.8 (direct API) 2.0 · legal issue: private 14B 1.64, Opus 4.8 (direct API) 1.76 · no invention: private 14B 2.0, Opus 4.8 (direct API) 2.0 · coach next step: private 14B 1.28, Opus 4.8 (direct API) 1.56

Read it honestly Small n and a mock judge. This is the only axis on the page whose grader is not fully mechanical. The Opus figure here is a separate direct-API run, not the chart row above. Treat the coaching column as directional and weight the big-n axes above it.

generated 2026-07-14 · scored with the committed coaching rubric

Question shaping (HEAR-16)

Courtroom speech is not a search query. This battery takes real spoken utterances, hypotheticals, interruptions, half-finished questions, and asks whether rewriting them into a shaped query retrieves the right authority better than sending the raw words.

  • Items20 labeled utterances
  • Raw utterancerecall 100% · MRR 0.975
  • Shaped queryrecall 100% · MRR 0.942

Read it honestly n is tiny, twenty utterances decides nothing, and here shaping did not beat the raw utterance on MRR. It is on this page because it is on disk and it did not go our way; a battery you only publish when it wins is not a battery.

read from the committed HEAR-16 result file at build time

Live ears, word error rate

The streaming transcription that feeds everything else, scored against the official transcript of a real argument after a deterministic structure trim (headers, line numbers, speaker labels and the keyword index removed by anchor, not by hand).

  • Word error rate10.6%
  • Cost to run it$0.018
  • Modelnvidia/nemotron-speech-streaming-en-0.6b
  • ReferenceSmith v. Arizona, No. 22-899 (U.S.)

Read it honestly This is the one figure on the page not recomputed at build time. It is quoted as a published fact: streaming transcription scored against the official Supreme Court transcript of No. 22-899 after the deterministic structure trim (16,683 spoken words). The audio shards for that run are kept off the repo, so this page quotes the published number rather than recomputing it. One argument is one argument, it is not a WER claim across courtrooms, accents, or bad audio.

published figure, quoted with its provenance

Citation gate at larger scale: the headline card reports a 2,000-fake sweep in which none were accepted. A later run over 10,000 fabricated citations (hidden among 10,000 real ones) accepted 7 of them, a 0.07% false-accept rate, against 3,379 (33.8%) for the industry-default string check on the same items. Both runs are committed in the results directory.

With-library axis, apples to apples: all 18 engines were asked the same 307 statute questions and shown the same retrieved passages, retrieval ran once and the identical text was replayed to every engine, so the engine is the only variable. Library: 8801:k5:scotus=9034,statutes=116136,tx=615391.

Difficulty is pre-registered: the generator and seed reproduce every item and its label, fixed before any model was scored. Confidence intervals follow Anthropic’s “Adding Error Bars to Evals” (Miller, 2024). Every figure on this page is read from a committed result file at build time; the full archival table is kept at benchmarks-full.html. Generated 2026-08-25.