Method, tests, and scoring
How the gauntlet works
ModelEval is a workload-driven Crucible for choosing local models for real agent work. Saved run evidence supports independent review and rescoring without rerunning inference.
Comparison integrity
Cards group the same model, runtime, quantization, thinking configuration, and physical hardware across complete runs. MTP and DFlash variants share one card. The top-level score and speed use the recommended configuration from the newest test generation; the detail view keeps every configuration and run, and test totals include all generations. Hardware profiles remain separate. Historical runs that omitted their context field are labeled with the confirmed 65,536-token set point; raw evidence remains unchanged.
Scores
Quality scores come from the stewarded Crucible rubric. Primary-agent and sub-agent scores use their named test families. The overall-performance index blends 70% quality, 15% long-context reliability, and 15% logarithmic long-context throughput. Concurrency never enters intelligence or quality scoring.
Each card also shows its best use case. The system averages the measured category scores in each use-case group. The group with the highest mean wins. Exact ties use the published group order. The card shows the runner-up and score margin.
Gauntlet time
Each card shows the median time for one complete gauntlet run. The exporter removes model-load time or detected first-test cold-start excess.
Long context
Long-context throughput is the minimum valid sustained decode observation across supporting runs. Genuine slowdowns remain; only corrupt or impossible telemetry is excluded. Context proof runs separately report configured context, acknowledgement fidelity, semantic proof results, speed, and peak server RSS. The app also shows average and worst-run Long Context Fidelity so unstable behavior is not hidden by a mean.
Every card shows a reliable context label. A 120K pair that passes both runs is 120K reliable context. If a lower checkpoint passes twice, the card shows that tier. If a later checkpoint passes after an earlier miss, the later pass is not a contiguous reliable ceiling. If no semantic checkpoint passed, the card shows exactly Zero reliable context. A missing proof is not a claim that configured context equals zero. Infrastructure stops are not model failures.
Pi Agent scores
Pi Agent scores are a separate metric family. They do not change Crucible overall, primary-agent, sub-agent, context, speed, or best-use-case scores. Repeated identical conditions publish the arithmetic mean and the source scores. Different runtimes, quantization values, thinking modes, and reasoning efforts stay separate. Remote aliases are provider-managed aliases, not fixed-weight snapshots.
Evidence and generations
Test additions, removals, and material definition changes create a comparison boundary. Every card links to run-level scores, performance, dates, runtime facts, RAM provenance, and limitations. Supplemental-generation evidence is labeled and excluded from the current generation ranking. The row-level bundle and current catalog are available for independent recomputation.