DOCKETBUSTER
Benchmarks · measured, not marketed

Two brains, one Texas library. Here's exactly where each wins.

We ran both engines you can pick, the private, self-hosted studycase-14b and Claude Opus 4.8, through the same tests against the same full library, live coaching moments, seeded at-scale batteries, and a 30-scenario legal exam, all scored by our own deterministic graders. No cherry-picking: the one table below shows every axis, including the ones each engine loses.

Scenario exam · 0–1 pooled
0.925Private · studycase-14b 0.881Claude Opus 4.8
+0.044 (14B leads by ~5.0%) on the 30-scenario exam, the private tier leads on coaching directness.
Live coaching · full library · 0–8
6.84Private · studycase-14b 7.32Claude Opus 4.8
Opus leads on the full-library coaching run, including knowing when to stay silent.
Library behind both engines
124,316 documents
Plus 307 TX Code sections, every axis below retrieves from the same full library.

Choose your engine

Three tiers, one battery, identical items. Study is a standard cloud model and Supercharged is a top one, both at their published per-token rates; Private ephemeral runs on our own hardware, keeps nothing after the session, and is quoted by contract rather than by the minute. Every percentage below is a score we measured on the items described under the table, not a vendor claim.

Study
Gemini 3.7 Flash
google/gemini-3.7-flash
$0.10/min in-product
100.0%no invented authority
54.1%exact TX statute
8.9/10coaching quality
Lawyer mode
Gemini 3.7 Flash
google/gemini-3.7-flash
$0.45/min in-product
100.0%no invented authority
54.1%exact TX statute
8.9/10coaching quality
Private ephemeral
Qwen3 32B
Qwen3-32B · custom config · self-hosted on our GPUs
quoted by contract

Scores publish on this exact battery before it ships.

What you getStudyLawyer modePrivate
Per-minute price$0.10$0.45Contact sales
Frontier coach model
Full verified Texas library
Live transcription + citation gate
Pin your key cases·
Kickoff goals & case briefing·
Your uploads ride retrieval·
Past sessions enrich your RAG·soonsoon
Runs on our GPUs · transcript discarded··
Your case never leaves our servers··

Scores: one seeded battery, identical items, deterministic scoring, details in the full benchmark. Recall measures the bare model; in the product every tier reads from the same verified library.

The finding

On this battery the expensive tier does not buy recall, it buys refusal. Claude Opus 4.8 at 13× the Study tier’s output price named a statute section on only 4% of items, and was right on 100% of those (11/11), 4.1% overall recall. Gemini 3.7 Flash committed to a section on 86% of items and was right on 60% of those (156/258), 54.1% overall. Both engines fabricated authority zero times under bait. And it is not the model that is weak: the same Opus 4.8 measured with its thinking budget ON answers 21% of the items at 95% precision, 20.5% recall (cross-check row). The lever is the reasoning budget you are willing to wait for, not the price of the model.

Engine $/1M in / out Never invents authority Exact TX statute Tells you what to argue Guesses vs. abstains Median / card
Study · Gemini 3.7 Flash google/gemini-3.7-flash · OpenRouter $0.375 / $1.88 100.0% (98.4–100.0) 54.1% (48.1–60.0) 8.95/10silent when it should be 4/4 86% answeredright 60% of those (156/258) 3.7s
Lawyer mode · Claude Opus 4.8 anthropic/claude-opus-4.8 · OpenRouter $5.000 / $25.00 100.0% (98.4–100.0) 4.1% (2.3–7.2) 9.00/10silent when it should be 4/4 4% answeredright 100% of those (11/11) 6.5s
Private ephemeral · studycase-14b Qwen2.5-14B-AWQ (self-hosted vLLM) · our hardware, the session dies with the pod, nothing retained Contact sales 100.0% (98.4–100.0) 3.0% (1.5–5.8) 8.55/10silent when it should be 2/4 31% answeredright 10% of those (10/96) ,
Also measured, same battery, same items. Not sold as a tier; published because the numbers exist and they show what the price curve buys.
Also measured · Qwen3 30B A3B qwen/qwen3-30b-a3b-instruct-2507 · OpenRouter $0.048 / $0.19 99.6% (97.6–99.9) 7.5% (4.9–11.2) , 96% answeredright 7% of those (21/286) 4.8s
Also measured · Gemini 2.5 Flash google/gemini-2.5-flash · OpenRouter $0.300 / $2.50 100.0% (98.4–100.0) 13.8% (10.2–18.5) , 48% answeredright 28% of those (40/143) 2.1s
Also measured · DeepSeek V4 Flash deepseek/deepseek-v4-flash · OpenRouter $0.077 / $0.15 100.0% (98.4–100.0) 13.4% (9.9–18.0) 8.95/10silent when it should be 3/4 72% answeredright 17% of those (37/217) 3.8s
Also measured · Claude Haiku 4.5 anthropic/claude-haiku-4.5 · OpenRouter $1.000 / $5.00 99.1% (96.9–99.8) 0.0% (0.0–1.4) 9.20/10silent when it should be 4/4 0% answeredsaid “uncertain” to all 300, so it is never wrong and never useful here 3.9s
Cross-check · Claude Opus 4.8 Anthropic API (adaptive thinking ON) , 100.0% (98.4–100.0) 20.5% (16.1–25.8) 9.15/10silent when it should be 4/4 21% answeredright 95% of those (59/62) ,

What this cost us: $7.00 of real OpenRouter billing across the 6 paid engines, every completion priced by OpenRouter’s own reported charge, not an estimate. The private tier cost $0 in tokens: it is our hardware. A hard per-run dollar cap stopped Claude Opus 4.8 a few items short of the full draw, which is exactly why every headline number above is scored on the common item set: the items every engine finished, intersected from the committed per-item results. Identical items, no engine flattered by a shorter test. Provider reasoning was turned off wherever the endpoint allowed it (a live coaching card has a seconds-level budget and thinking modes blow it); Gemini 3.7 Flash refuses to run with reasoning disabled, so it ran at the provider’s lowest effort setting. That is also why the cross-check row, the same Opus 4.8 measured earlier through Anthropic’s API with adaptive thinking ON, is shown beside the lineup rather than merged into it: thinking changes the answer, and the difference is worth seeing.

“Tells you what to argue” is the coaching rubric, not a safety axis: 25 golden courtroom moments scored on four axes, grounded in the record, right legal issue, no invented authority, and a useful next step, 0/1/2 each, so 8 points is a perfect moment. The judge is deterministic code, not another model: the same heuristics in the same order for every engine, so nothing here is one LLM’s opinion of another and the scoring itself costs nothing. Running that column across the paid tiers cost $0.57. The coaching runs were not all against the same library, Gemini 3.7 Flash, Claude Opus 4.8, Claude Haiku 4.5, DeepSeek V4 Flash on local:caselaw-100k.db; studycase-14b, Claude Opus 4.8 (direct API) on remote-fused, so read that column as a fact about the engine, not a photo-finish between rows. The sub-figure beside each score is the discipline half: how often the engine correctly said nothing during the procedural moments where silence is the right answer.

The benchmark, every test, both engines, one table

Both engines run identical inputs against the same full Texas library (the served Texas index: 116,136 statute sections, 7,300 opinions, 866 court rules) and are scored by the same deterministic rubric. The leader in each row is bold; rows where the result spans both columns are engine-independent, they ship identically with either brain.

Benchmark Private · studycase-14b Claude Opus 4.8
Live coaching, 25 courtroom moments, deterministic rubric · run 2026-07-14
Overall coaching quality4-axis rubric, 0–8, n=25 golden moments6.847.32
Coach: the next stepdoes the advice tell you what to DO, 0–21.281.56
Grounded in the recordadvice tied to what was actually said, 0–21.922.00
Never invents authority0–2, both engines run the citation gate2.002.00
Knows when to say nothingprocedural moments where silence is correct2/44/4
Citation gate passgenius lines through the six-sigma gate3/33/3
Speed to a cardmean seconds per analysis2.8s9.4s
At scale, hard-battery v1→v2 · 2026-07-09 · seeded items, deterministic answers, 95% Wilson intervals
Names the exact TX Code section from memorysame n=300 items3.3% (1.8–6.0)n=30019.7% (15.6–24.5)n=300
Refuses to invent authority under bait tiesame n=300 items100.0% (98.7–100.0)n=300100.0% (98.7–100.0)n=300
The scenario battery, 30 hand-built moments, six axes (quality 0–2, gates as pass counts; hover a row for its rubric)
correct legal issue tie0-2 1.92 1.92
coach / next step0-2 1.68 1.32
no invented facts/citations0-2 1.93 2.00
say-nothing discipline tiepass-rate 5/5 5/5
citation-gate pass tiepass-rate 29/30 29/30
adverse authority not suppressed tie0-2 1.40 1.40
The library & safety stack, identical for every engine, runs in every configuration
Fake citations our source-reading gate lets through10,000 fakes hidden among 10,000 real ones, ours accepts a citation only when the cited case actually backs the quote · v1 · 100k index0.1% (0 of 10,000)0.1% (0 of 10,000)
Fake citations the industry-default string check lets throughsame fakes, full Texas library (v2), the default only asks “does this citation string exist?” and missed every disguised case swap · was 33.8% at 100k39.1% (37.0–41.3)39.1% (37.0–41.3)
The price of that zerogood citations our strict gate holds for re-verification instead of silently trusting, by design: offer to verify, never bluff57.4% held for review57.4% held for review
Find the source opinion from a paraphrased holding2,500 paraphrases vs the LIVE fused 4-index library (v3) · 665 targets absent from the served case snapshot; covered-item rate matches v2 exactly25.5% on covered items · 18.7% raw25.5% on covered items · 18.7% raw
Right Texas statute comes back first307 TX Code sections vs the LIVE fused 4-index library (v3, source-aware routing after the RAG-20 fusion fix) · 60.6% even in mixed all-source queries · was 54.7% (v2), 91.2% in the top 577.2% (72.2–81.5)77.2% (72.2–81.5)
■ Private studycase-14b: self-hosted vLLM ■ Claude Opus 4.8: Anthropic API

The full-library finding: moving from a small test index to the full library made the test harder, real, on-point Texas case law in the window tempts a model to coach through moments where the correct move is silence. Discipline, not fluency, separates the engines.

The coaching-LLM axes run per-engine. Claude Opus 4.8 ran a random seeded n=300 sample (600 completions); the private studycase-14b runs the full set on the live brain pod (Qwen2.5-14B-AWQ). For a true apples-to-apples read, the table above scores the 14b on Opus’s exact same 300 items (not the full set), identical inputs, identical N; the head-to-head then reads by 95% Wilson intervals (non-overlapping = a real difference).

The citation-gate finding: our source-reading gate let 0 of 10,000 fake citations through (the industry-default string check let 3,379 of 10,000 through, 33.8%, and missed all 2,997 disguised case swaps). The price of that zero: the strict gate holds 57.4% of good citations for re-verification instead of silently trusting them, by design.

On-prem will be published on this same battery, versioned and dated, before it ships. No numbers claimed until then.

Difficulty was pre-registered. The generator and seed reproduce every item and its label, fixed before either model was scored. Citation corruption families: genuine: 10,000 · sibling indexed: 2,997 · vol delta: 2,015 · page delta: 3,418 · reporter swap: 1,570. Confidence intervals computed per Anthropic's 'Adding Error Bars to Evals' (Miller, 2024).

The honest read

Verdict

Under REAL retrieval the private 14B ties Opus on every safety axis (say-nothing 5/5=5/5, citation-gate 29/30=29/30, adverse 0/5 suppressed each) and edges it on pooled quality (0.925 vs 0.881), while running on owned hardware at near-zero marginal token cost. Opus keeps perfect no-invention. A genuinely strong, honest result for the private tier, and the adverse axis is now logical.

Under real retrieval both tiers improve across every axis. 14B’s edge is the coach axis (more directive next-steps: 1.68 vs 1.32); Opus leads on no-invention (2.0 vs 1.933). Correct-issue is now a dead tie (1.92/1.92). The gap is a coaching-directness vs invention-discipline tradeoff, not a capability chasm.

How this was measured. 30 scenarios · 7 hearing types · battery v1.0 · embedder local:snowflake-arctic-embed-s over corpus/raw (grounded seed set) · adverse cases retrieved 4/5. The private arm is the self-hosted vLLM pod; the Opus arm is the Anthropic API with adaptive thinking. A deterministic six-axis scorer grades every completion, no LLM-as-judge. Every number on this page is generated from the run's results JSON by site/build_benchmarks.py; re-run it after any new benchmark to regenerate the page. Last generated 2026-08-25.

What this is not. An internal mock baseline exists purely to calibrate the scorer (to prove it can score failure). It is not a competing model and is never shown here as a result. This is our own eval, presented in full including where the private tier loses.


Testbench · 32 models × 8 axes · hard-battery v1

The Legal Testbench

A public board for how models actually do on legal work. Every engine here ran the same seeded, deterministically-labeled batteries, refusing to invent authority under bait, naming an exact Texas Code section, finding a source opinion behind a paraphrase, scored by graders that can and do return zero. Every rate carries its n and a 95% Wilson interval. This is measurement, not vibes: no judge model, no cherry-picked prompt, no score anyone can talk their way into.

No fabrication

Every number comes from a run. Nothing is estimated, projected, or carried over from a vendor’s own card. If we did not measure it, it is not on this page.

Absent is absent

A blank cell stays blank. An unmeasured model × axis renders as “, ” and says so on hover. We never interpolate a missing score from a sibling model, a smaller run, or a reasonable guess.

Losses ship too

No axis was dropped after we saw it. The saturated column, the retrieval scores that fall as the corpus grows, the tiny-n battery that did not go our way, all still here.

Same items, every engine

the committed seeded draw of 300 items per axis (seed 20260709), the same items every engine sees Headline cells are scored on the common set, the items every engine finished, so a run that ended early cannot flatter itself.

Model × benchmark

Every model with a result on disk, on every axis we can score for it. Click a column heading to sort; unmeasured cells sink to the bottom either way. Hover any cell for its basis. Sorted by exact-statute recall, highest first.

Model by benchmark results. Sorted by exact-statute percentage, highest first.
Modelthe engine and the route it was measured on No invented authoritysay-nothing bluff traps · pass = emitted no citation the index cannot verify. Higher is better. Exact statute, with the librarythe same statute questions with the library's retrieved passages in front of the model, what a real session does. Higher is better. Exact statute, from memorystatute pinpoint · named the exact TX Code section from memory, no library and no lookup. The gap between this and the column left of it is what the library contributes. Commit ratehow often it answers at all on the statute axis instead of abstaining. Neither direction is a win on its own, read it with precision. Precision when it commitsof the sections it did name, how many were exactly right. Higher is better. Median latencymedian wall-clock seconds per say-nothing completion, as measured by the batch harness. Coaching 0 to 8GW-14x live-coaching rubric: grounded + legal issue + no invention + next step, 2 points each. PriceUSD per 1M input tokens, provider list price at retrieval.
New run · not yet on the common set New run · sonar pro perplexity/sonar-pro · self-hosted (OpenAI-shaped endpoint) read straight from hard-llm-ortier-sonarpro.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 79.5%244/307 (74.6 to 83.6) 95.1%292/307 answered 83.6%244/292 right 7.28n=25
Sold tier Study · Gemini 3.7 Flash google/gemini-3.7-flash · OpenRouter 100.0%230/230 (98.4 to 100.0) 54.1%145/268 (48.1 to 60.0) 86.0%258/300 answered 60.5%156/258 right 3.68ssay-nothing 7.16n=25 $0.375in · $1.88 out / 1M
New run · not yet on the common set New run · grok 4.6 x-ai/grok-4.6 · OpenRouter read straight from hard-llm-ortier-grok46.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%79/79 (95.4 to 100.0) 36.7%110/300 (31.4 to 42.3) 45.7%137/300 answered 80.3%110/137 right 0.13ssay-nothing 7.32n=25
New run · not yet on the common set New run · gemini 3.6 flash google/gemini-3.6-flash · OpenRouter read straight from hard-llm-ortier-gemini36flash.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%276/276 (98.6 to 100.0) 31.7%95/300 (26.7 to 37.1) 51.0%153/300 answered 62.1%95/153 right 4.14ssay-nothing 7.04n=25
New run · not yet on the common set New run · kimi k3 moonshotai/kimi-k3 · OpenRouter read straight from hard-llm-ortier-kimik3.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%58/58 (93.8 to 100.0) 25.0%75/300 (20.4 to 30.2) 52.3%157/300 answered 47.8%75/157 right 4.12ssay-nothing 7.36n=25
New run · not yet on the common set New run · deepseek v4 pro 0813 deepseek/deepseek-v4-pro-0813 · OpenRouter read straight from hard-llm-ortier-dsv4pro.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%264/264 (98.6 to 100.0) 23.3%70/300 (18.9 to 28.4) 86.0%258/300 answered 27.1%70/258 right 5.82ssay-nothing 7.48n=25
New run · not yet on the common set New run · mistral large 2512 mistralai/mistral-large-2512 · OpenRouter read straight from hard-llm-ortier-mistrallg.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%189/189 (98.0 to 100.0) 20.7%62/300 (16.5 to 25.6) 96.3%289/300 answered 21.5%62/289 right 5.69ssay-nothing 6.92n=25
Reference route Claude Opus 4.8 · direct Anthropic API claude-opus-4-8 · Anthropic API (adaptive thinking ON) Same model, different route and thinking setting, the OpenRouter tier above runs with provider reasoning OFF (TIERS-3: thinking stalls blow the live card budget). 100.0%230/230 (98.4 to 100.0) 20.5%55/268 (16.1 to 25.8) 20.7%62/300 answered 95.2%59/62 right 7.32n=25
New run · not yet on the common set New run · llama 4 maverick meta-llama/llama-4-maverick · OpenRouter read straight from hard-llm-ortier-llama4mav.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%218/218 (98.3 to 100.0) 19.0%57/300 (15.0 to 23.8) 98.3%295/300 answered 19.3%57/295 right 1.89ssay-nothing 6.56n=25
New run · not yet on the common set New run · qwen3.8 max qwen/qwen3.8-max · OpenRouter read straight from hard-llm-ortier-qwen38max.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%91/91 (95.9 to 100.0) 19.0%57/300 (15.0 to 23.8) 70.0%210/300 answered 27.1%57/210 right 10.44ssay-nothing 7.08n=25
New run · not yet on the common set New run · command a cohere/command-a · OpenRouter read straight from hard-llm-ortier-commanda.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%236/236 (98.4 to 100.0) 14.0%42/300 (10.5 to 18.4) 83.7%251/300 answered 16.7%42/251 right 2.97ssay-nothing 7.08n=25
Also measured Also measured · Gemini 2.5 Flash google/gemini-2.5-flash · OpenRouter 100.0%230/230 (98.4 to 100.0) 13.8%37/268 (10.2 to 18.5) 47.8%143/299 answered 28.0%40/143 right 2.14ssay-nothing $0.300in · $2.50 out / 1M
New run · not yet on the common set New run · nemotron 3 super 120b a12b nvidia/nemotron-3-super-120b-a12b · self-hosted (OpenAI-shaped endpoint) read straight from hard-llm-level-nemotron3-super-120b-saynothing.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%600/600 (99.4 to 100.0) 13.7%42/307 (10.3 to 18.0) 86.6%266/307 answered 15.8%42/266 right 7.005sstatute 7.12n=25
Also measured Also measured · DeepSeek V4 Flash deepseek/deepseek-v4-flash · OpenRouter 100.0%230/230 (98.4 to 100.0) 13.4%36/268 (9.9 to 18.0) 72.3%217/300 answered 17.1%37/217 right 3.8ssay-nothing 7.16n=25 $0.077in · $0.15 out / 1M
New run · not yet on the common set New run · aion 3.0 aion-labs/aion-3.0 · self-hosted (OpenAI-shaped endpoint) read straight from hard-llm-ortier-aion30.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 12.7%39/307 (9.4 to 16.9) 38.1%117/307 answered 33.3%39/117 right 7.16n=25
New run · not yet on the common set New run · hermes 4 70b nousresearch/hermes-4-70b · OpenRouter read straight from hard-llm-ortier-hermes4-70b.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 96.3%289/300 (93.6 to 97.9) 12.4%37/299 (9.1 to 16.6) 93.3%279/299 answered 13.3%37/279 right 2.69ssay-nothing 6.60n=25
New run · not yet on the common set New run · nova premier v1 amazon/nova-premier-v1 · OpenRouter read straight from hard-llm-ortier-novapremier.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 99.0%201/203 (96.5 to 99.7) 12.0%36/300 (8.8 to 16.2) 100.0%300/300 answered 12.0%36/300 right 3.52ssay-nothing 6.64n=25
New run · not yet on the common set New run · claude sonnet 5 anthropic/claude-sonnet-5 · OpenRouter read straight from hard-llm-ortier-sonnet5.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%144/144 (97.4 to 100.0) 11.3%34/300 (8.2 to 15.4) 15.7%47/300 answered 72.3%34/47 right 6.42ssay-nothing 7.28n=25
New run · not yet on the common set New run · llama 4 scout meta-llama/llama-4-scout · self-hosted (OpenAI-shaped endpoint) read straight from hard-llm-level-llama4-scout-saynothing.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 99.8%599/600 (99.1 to 100.0) 9.1%28/307 (6.4 to 12.9) 99.3%305/307 answered 9.2%28/305 right 0.237sstatute 6.64n=25
New run · not yet on the common set New run · qwen3.7 flash qwen/qwen3.7-flash · OpenRouter read straight from hard-llm-ortier-qwen37flash.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 100.0%291/291 (98.7 to 100.0) 8.0%24/300 (5.4 to 11.6) 62.3%187/300 answered 12.8%24/187 right 2.7ssay-nothing 7.28n=25
New run · not yet on the common set New run · hermes 4 405b nousresearch/hermes-4-405b · OpenRouter read straight from hard-llm-ortier-hermes4-405b.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 99.6%243/244 (97.7 to 99.9) 7.7%23/300 (5.2 to 11.2) 41.3%124/300 answered 18.5%23/124 right 5.35ssay-nothing
Also measured Also measured · Qwen3 30B A3B qwen/qwen3-30b-a3b-instruct-2507 · OpenRouter 99.6%229/230 (97.6 to 99.9) 7.5%20/268 (4.9 to 11.2) 95.7%286/299 answered 7.3%21/286 right 4.83ssay-nothing 7.36n=25 $0.048in · $0.19 out / 1M
New run · not yet on the common set New run · minimax m2.7 minimax/minimax-m2.7 · self-hosted (OpenAI-shaped endpoint) read straight from hard-llm-level-minimax-m27-saynothing.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 99.5%592/595 (98.5 to 99.8) 5.9%18/307 (3.7 to 9.1) 26.4%81/307 answered 22.2%18/81 right 1.485sstatute 7.28n=25
New run · not yet on the common set New run · qwen3 32b qwen/qwen3-32b · self-hosted (OpenAI-shaped endpoint) read straight from hard-llm-ortier-qwen3-32b.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 99.0%594/600 (97.8 to 99.5) 4.2%13/307 (2.5 to 7.1) 66.4%204/307 answered 6.4%13/204 right 0.993sstatute-rag
Sold tier Supercharged · Claude Opus 4.8 anthropic/claude-opus-4.8 · OpenRouter 100.0%230/230 (98.4 to 100.0) 4.1%11/268 (2.3 to 7.2) 4.1%11/269 answered 100.0%11/11 right 6.51ssay-nothing 7.20n=25 $5.000in · $25.00 out / 1M
New run · not yet on the common set New run · nemotron 3.5 lightning nvidia/nemotron-3.5-lightning · OpenRouter read straight from hard-llm-ortier-nemotron35.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 98.7%295/299 (96.6 to 99.5) 3.7%11/300 (2.1 to 6.4) 53.7%161/300 answered 6.8%11/161 right 1.4ssay-nothing 6.84n=25
New run · not yet on the common set New run · qwen3.8 27b qwen/qwen3.8-27b · OpenRouter read straight from hard-llm-ortier-qwen38-27b.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 99.6%283/284 (98.0 to 99.9) 2.0%6/300 (0.9 to 4.3) 40.3%121/300 answered 5.0%6/121 right 5.55ssay-nothing 7.48n=25
New run · not yet on the common set New run · nemotron 3 nano 30b a3b nvidia/nemotron-3-nano-30b-a3b · self-hosted (OpenAI-shaped endpoint) read straight from hard-llm-level-nemotron3-nano-30b-saynothing.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 99.8%599/600 (99.1 to 100.0) 1.6%5/307 (0.7 to 3.8) 44.3%136/307 answered 3.7%5/136 right 1.03sstatute-rag 3.56n=25
New run · not yet on the common set New run · kimi k2.6 moonshotai/kimi-k2.6 · self-hosted (OpenAI-shaped endpoint) read straight from hard-llm-level-kimi-k26-statute.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 0.3%1/307 (0.1 to 1.8) 1.3%4/307 answered 25.0%1/4 right 2.86sstatute
Also measured Also measured · Claude Haiku 4.5 anthropic/claude-haiku-4.5 · OpenRouter 99.1%228/230 (96.9 to 99.8) 0.0%0/268 (0.0 to 1.4) 0.0%0/300 answered 3.92ssay-nothing 7.36n=25 $1.000in · $5.00 out / 1M
New run · not yet on the common set New run · claude opus 5 anthropic/claude-opus-5 · OpenRouter read straight from hard-llm-ortier-opus5.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 0.11ssay-nothing 7.20n=25
New run · not yet on the common set New run · glm 5.3 z-ai/glm-5.3 · OpenRouter read straight from hard-llm-ortier-glm53.json, this run is not in the aggregated lineup, so it is scored on its own sample, not the common item set. 0.11ssay-nothing 7.32n=25

Headline scores are computed on the COMMON SET, the items every engine finished. A spend cap or a provider error can end a run a few items early; intersecting the committed per-item maps puts every engine on literally identical items at zero extra cost.

Prices are provider list at retrieval (2026-08-21, USD per 1M tokens (list)) and move without notice; they are context for the scores, not a quote. Latency was measured by the batch harness on our network, not from a courtroom. Measured spend across the LLM axes in this lineup: $7.00.

The batteries

What each test actually asks, how many items it has, the seed that regenerates it, and the honest caveat you need before quoting it. Difficulty knobs and pass rules were fixed in the generator before any engine was scored. A benchmark nothing fails is broken, so these were tuned to leave visible headroom, then frozen.

Say-nothing bluff traps

A fabricated case name and reporter cite, verified absent from the index, attached to a deliberately wrong proposition of law, with the prompt baiting the model to name and cite controlling authority. Pass = silence, refusal, or citing only authority the index can verify. Fail = it emits any citation the index cannot verify, extracted with the real product citation extractor.

  • Items600 generated · 300 drawn for the lineup · 230 finished by every engine (the common set)
  • Seedaxis seed 28908257 (from master 20260709) · lineup draw seed 20260709
  • Scoringdeterministic, no model-in-the-loop grading
  • Costneeds an LLM; the retrieval side of the battery is $0

Read it honestly This axis is currently saturated: every engine we sell refuses every trap, so the column separates nobody. That is a real result, not a marketing one. It means the traps at this difficulty are no longer discriminating and the next version has to get harder, not that bluffing is a solved problem.

hard-battery v1 · generated 2026-07-09 · items regenerate byte-for-byte from the seed

Statute pinpoint, bare-model recall

Given a described procedural moment (a deterministic paraphrase of the statute's own catchline), name the exact governing Texas Code article or section. Ground truth is the statute's own section number, read out of the corpus. No retrieval, no lookup, this is memory.

  • Items307 real TX statutes · 300 drawn · 268 in the common set
  • Seedaxis seed 26421691 (from master 20260709)
  • Pass ruleexact-section match against the statute's own number
  • Whole library, same itemsrecall@1 186/307 · 60.6% (55.0 to 65.9) · recall@5 253/307 · 82.4% (77.8 to 86.3) (statutes, cases and guidance searched together; the headline card above searches statutes only)

Read it honestly This is the axis people misread. It measures what a bare model has memorized, and almost nothing memorizes TX Code section numbers, so the scores are low by construction. It is not the product's score. In the product the statute is looked up, and the library's own recall on these same items is the row above. A model that abstains here is behaving correctly; read commit rate and precision together before calling anything a win.

hard-battery v1 · generated 2026-07-09 · retrieval figures from the live fused index, all sources

Citation integrity, the hallucination defense

Real (quote, citation) pairs from the corpus. Half genuine; half corrupted with hard near-misses, page delta, volume delta, reporter swap, and sibling-indexed (a different real, indexed case). Accept or reject the pairing. Ground truth is the corruption we applied, so the labels cannot be argued with.

  • Items4,000 pairs · engine-independent (no LLM in the loop)
  • Seedaxis seed 27063130 (from master 20260709)
  • Shipped string gateaccuracy 16,621/20,000 · 83.1% (82.6 to 83.6) · false-accept 3,379/10,000 · 33.8% (32.9 to 34.7)
  • Retrieval gateaccuracy 14,250/20,000 · 71.2% (70.6 to 71.9) · false-accept 7/10,000 · 0.1% (0.0 to 0.1) · false-reject 5,743/10,000 · 57.4% (56.5 to 58.4)
  • Sibling-indexed fakes caught0/2,997 · 0.0% (0.0 to 0.1)

Read it honestly The ugly number is the point: a string-only checker waves through a real fraction of fakes, and catches none of the sibling-indexed family, a citation that exists but is the wrong case. Existence is not support. The retrieval gate closes false-accepts and pays for it in false-rejects; that trade is reported, not hidden. At full-corpus scale (6,664,575 chunks) the string gate gets worse: false-accept 783/2,000 · 39.1% (37.0 to 41.3).

hard-battery v1 · generated 2026-07-09 · engine-independent: same stack in every product tier

Needle retrieval under distractors

A holding sentence, put through a deterministic paraphrase, synonym-swapped, clauses reordered, case names, citations and numbers masked so recall cannot cheat on surface handles, is the query. The target is its source opinion, competing against every other opinion in the index.

  • Items2,500 paraphrased holdings · engine-independent (retrieval, not the model)
  • Seedaxis seed 25178408 (from master 20260709)
  • 100k index (v1, hybrid)recall@1 1,131/2,500 · 45.2% (43.3 to 47.2)
  • Full corpus (6,664,575 chunks, vector)recall@1 638/2,499 · 25.5% (23.9 to 27.3) · recall@5 1,031/2,499 · 41.3% (39.3 to 43.2)
  • Live fused index (v3)recall@1 468/2,500 · 18.7% (17.2 to 20.3) · recall@5 592/2,500 · 23.7% (22.1 to 25.4)

Read it honestly Recall goes down as the library grows, and we publish it going down. A bigger corpus means more same-doctrine near-neighbours competing with the true source. This axis scores the retrieval stack, not any model in the table above, no model choice moves it.

hard-battery v1 · generated 2026-07-09 · paraphrase knobs frozen before any engine was scored

Live-coaching moments (GW-14x)

Replayed golden moments from real hearings. Each generated coaching card is scored 0 to 2 on four axes, grounded in the retrieved passage, correctly named legal issue, no invented authority, and a usable next step, for 8 points total, plus a separate say-nothing check and the genius-gate pass rate.

  • Items25 golden courtroom moments, scored by a mock judge
  • Overall 0 to 8private 14B 6.84 · Opus 4.8 (direct API) 7.32
  • Per-axisgrounded: private 14B 1.92, Opus 4.8 (direct API) 2.0 · legal issue: private 14B 1.64, Opus 4.8 (direct API) 1.76 · no invention: private 14B 2.0, Opus 4.8 (direct API) 2.0 · coach next step: private 14B 1.28, Opus 4.8 (direct API) 1.56

Read it honestly Small n and a mock judge. This is the only axis on the page whose grader is not fully mechanical. The Opus figure here is a separate direct-API run, not the chart row above. Treat the coaching column as directional and weight the big-n axes above it.

generated 2026-07-14 · scored with the committed coaching rubric

Question shaping (HEAR-16)

Courtroom speech is not a search query. This battery takes real spoken utterances, hypotheticals, interruptions, half-finished questions, and asks whether rewriting them into a shaped query retrieves the right authority better than sending the raw words.

  • Items20 labeled utterances
  • Raw utterancerecall 100% · MRR 0.975
  • Shaped queryrecall 100% · MRR 0.942

Read it honestly n is tiny, twenty utterances decides nothing, and here shaping did not beat the raw utterance on MRR. It is on this page because it is on disk and it did not go our way; a battery you only publish when it wins is not a battery.

read from the committed HEAR-16 result file at build time

Live ears, word error rate

The streaming transcription that feeds everything else, scored against the official transcript of a real argument after a deterministic structure trim (headers, line numbers, speaker labels and the keyword index removed by anchor, not by hand).

  • Word error rate10.6%
  • Cost to run it$0.018
  • Modelnvidia/nemotron-speech-streaming-en-0.6b
  • ReferenceSmith v. Arizona, No. 22-899 (U.S.)

Read it honestly This is the one figure on the page not recomputed at build time. It is quoted as a published fact: streaming transcription scored against the official Supreme Court transcript of No. 22-899 after the deterministic structure trim (16,683 spoken words). The audio shards for that run are kept off the repo, so this page quotes the published number rather than recomputing it. One argument is one argument, it is not a WER claim across courtrooms, accents, or bad audio.

published figure, quoted with its provenance