This table covers models running on ONE DGX Spark.
Thirty-one model configurations benchmarked on a single NVIDIA DGX Spark (GB10, 121.6 GiB unified memory), each served at its own best setup, so engines, quantisations and speculative-decoding methods differ from column to column by design. Every cell shows a 0 to 10 standing within its row, the raw measurement, and the exact configuration that produced it. Scores compare within a row only. There are no column totals, because adding a speed standing to an OCR standing gives a number that means nothing.
| What is measuredone row per measure · one column per model | RadixArk FP4 27BSGLang + DFlash2 · vision ✓ | NVIDIA Qwen3.8-27B NVFP4official local edition · SGLang + DFlash2 · vision ✓ | Qwen 27B apostateuncensored · F16 · no vision in this build | Qwen3.8-27B BF16 base16-bit yardstick · ngram k5 · vision not entered | Leimroth 3 (based on Qwen3.8-27B)TensorFold 0.6.0 · NVFP4 · DFlash2 · images via companion model · abliterated build | Unbound LR7Qwen3.8-Flash-Next (Vontra MLX 4-bit MTP) on TensorFold 0.5.0 + wait fix · no vision · abliterated build | Unsloth 27B (vLLM)NVFP4 · MTP-5 · vision ✓ · same weights as the SGLang column | QUASAR-QAT 27BNVFP4 QAT, all layers · MTP-2 · vision ✓ | Qwen 27B orcarouteruncensored · NVFP4 · vision ✓ | Qwen 27B orcarouter (FP8)uncensored · FP8 · vision not verified this pass | Unsloth 27B (SGLang)NVFP4 · DFlash2 · vision ✓ · same weights as the vLLM column | RadixArk BF16-headSGLang + DFlash2 · vision ✓ | Leimroth 5 (Qwen3.6-35B-A3B)NVFP4 · MTP-3 · vision ✓ · abliterated, below adoption bar | DeepSeek-V4-Flashcommunity ABLITERATED build · EXL3 + DSpark draft · no vision in this build | Qwen3.6-35B-A3BNVFP4 · MoE · vision ✓ | Apodex-1.1-miniBF16, no NVFP4 build · MTP-3 · vision ✓ · agent-tuned 35B-A3B | Ornith-1.5-35B-A3BvLLM 0.28 · NVFP4 · no drafter (spec n/a) · vision not verified this pass | NVIDIA Nemotron 3.5 LightningvLLM 0.28 · NVFP4 · DSpark draft head · vision not verified this pass | Qwen3.8-Flash-NextGGUF · llama.cpp · vision ✓ | Qwen3.8-Flash-NextTensorFold · MLX 4-bit · MTP · vision ✓ | Qwen3.5-122BNVFP4 MoE · vision ✓ | Nemotron-Omni-30BNVFP4 · vision ✓ · audio ✓ | Muse Glimmer-30BvLLM 0.28 · NVFP4 · DFlash draft head · vision not verified this pass | Dream Shimmer LR6 (stock drafter)vLLM 0.28 · FP8-per-block · stock DFlash draft head · CUDA graphs · vision ✓ | Dream Shimmer LR6 (z-lab drafter)vLLM 0.28 · FP8-per-block · z-lab DFlash2 draft head · CUDA graphs · vision ✓ | GLM-4.7-FlashNVFP4 · no vision in this build | GLM-5.3-Flash EXL3 (K2)EXL3 · MTP-2 · no vision in this build | Gemma-4-26B-A4BNVFP4 · vision ✓ | Gemma-4-31B-it-uncensored-hereticuncensored · BF16 · vision not verified this pass | Ling-3.0-Flash INT4INT4 · MTP · no vision in this build | olmOCR-2-7BQ4_K_M · llama.cpp · vision ✓ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Agent tasks | 1055/60 - 60 items - 0 LOST | 8.552/60 - 60 items - 2 context artifacts | 953/60 - 60 items - 2 lost | 5.543/60 - 60 items - 2 lost | 1057/60 - 60 items (a-tool 12/12, a-seq 12/12, a-schema 11/12, a-turn 10/12, a-ctx 12/12) | 1056/60 - 60 items (a-tool 10/12, a-seq 10/12, a-schema 12/12, a-turn 12/12, a-ctx 12/12) | 9169/191 - 88.5% - 191 items | 8.552/60 - 60 items - 2 lost | 9.554/60 - 60 items (a-tool 12/12, a-seq 10/12, a-schema 12/12, a-turn 10/12, a-ctx 10/12) | N/Anot entered - not yet measured on this column | 851/60 - 60 items - 4 lost | 5.546/60 - 60 items - 9 LOST | 8.553/60 - 60 items - 2 context artifacts | 1055/60 - 60 items - 0 LOST | 8.552/60 - 60 items - 2 lost | 1055/60 - 60 items - 0 LOST | 8.552/60 - 60 items - 0 lost (a-tool 12/12, a-seq 11/12, a-schema 11/12, a-turn 6/12, a-ctx 12/12) | 7.549/60 - 60 items - 0 lost (a-tool 11/12, a-seq 8/12, a-schema 11/12, a-turn 10/12, a-ctx 9/12) | 6.548/60 - 60 items - 8 LOST | 1057/60 - 60 items (a-tool 11/12, a-seq 11/12, a-schema 11/12, a-turn 12/12, a-ctx 12/12) | 8.553/60 - 60 items - 0 LOST - 5 context artifacts | 851/60 - 60 items - 4 lost | 1056/60 - 60 items - 0 lost (a-tool 10/12, a-seq 11/12, a-schema 12/12, a-turn 11/12, a-ctx 12/12) | 9.554/60 - 60 items (a-tool 10/12, a-seq 11/12, a-schema 12/12, a-turn 9/12, a-ctx 12/12) | 851/60 - 60 items (a-tool 8/12, a-seq 10/12, a-schema 12/12, a-turn 9/12, a-ctx 12/12) - 51/60 in all three runs | 850/60 - 60 items - 0 lost | 6.550/60 - 60 items - 0 LOST (6 a-ctx/a-turn items exceed the 8k ctx: clean 400s) | 7.550/60 - 60 items - 2 lost | 851/60 - 60 items (a-tool 9/12, a-seq 11/12, a-schema 12/12, a-turn 10/12, a-ctx 9/12) | 6.549/58 scoreable - 60 items - 2 harness-artifact, 0 wrong-answer excess | 0could not be tested - cannot call tools - 60 items attempted 0/0 |
| Tools & structured output | 9.523/24 - 24 items | 9.523/24 - 24 items - a-tool 12/12, a-schema 11/12 | 9.523/24 - 24 items | 820/24 - 24 items - a-tool 11/12, a-schema 9/12 | 9.523/24 - 24 items - a-tool 12/12, a-schema 11/12 | 922/24 - 24 items - a-tool 10/12, a-schema 12/12 | 9.5127/138 - 92.0% - 191-item suite | 922/24 - 24 items - a-tool 10/12, a-schema 12/12 | 1024/24 - 24 items - a-tool 12/12, a-schema 12/12 | N/Anot entered - not yet measured on this column | 9.523/24 - 24 items | 9.523/24 - 24 items | 1024/24 - 24 items - a-tool 12/12, a-schema 12/12 | 9.523/24 - 24 items | 9.523/24 - 24 items | 9.523/24 - 24 items - a-tool 11/12, a-schema 12/12 | 9.523/24 - 24 items - a-tool 12/12, a-schema 11/12 | 922/24 - 24 items - a-tool 11/12, a-schema 11/12 | 922/24 - 24 items | 922/24 - 24 items - a-tool 11/12, a-schema 11/12 | 1024/24 - 24 items - PERFECT | 9.523/24 - 24 items | 922/24 - 24 items - a-tool 10/12, a-schema 12/12 | 922/24 - 24 items - a-tool 10/12, a-schema 12/12 | 820/24 - 24 items (a-tool 8/12, a-schema 12/12) - runs: 21, 20, 20 | 7.519/24 - 24 items | 9.523/24 - 24 items - a-tool 12/12 PERFECT, a-schema 11/12 | 8.521/24 - 24 items | 8.521/24 - 24 items - a-tool 9/12, a-schema 12/12 | 8.521/24 - 24 items - a-tool 11/12, a-schema 10/12 | 0tested and failed - no tool calls even with --jinja - 24 items attempted 0/0 |
| Multi-turn & sequencing | 8.520/24 - 24 items | 819/24 - 24 items - a-turn 8/12, a-seq 11/12 | 8.520/24 - 24 items | 4.513/24 - 24 items - a-turn 8/12, a-seq 5/12 | 1022/24 - 24 items - a-turn 10/12, a-seq 12/12 | 1022/24 - 24 items - a-turn 12/12, a-seq 10/12 | 6.526/37 - 70.3% - 191-item suite | 8.520/24 - 24 items - a-turn 10/12, a-seq 10/12 | 8.520/24 - 24 items - a-turn 10/12, a-seq 10/12 | N/Anot entered - not yet measured on this column | 8.520/24 - 24 items | 8.520/24 - 24 items | 819/24 - 24 items - a-turn 8/12, a-seq 11/12 | 8.520/24 - 24 items | 819/24 - 24 items | 8.520/24 - 24 items - a-turn 9/12, a-seq 11/12 | 6.517/24 - 24 items - a-turn 6/12, a-seq 11/12 | 7.518/24 - 24 items - a-turn 10/12, a-seq 8/12 | 1022/24 - 24 items | 1023/24 - 24 items - a-turn 12/12, a-seq 11/12 | 819/24 - 24 items - a-turn 9/12, a-seq 10/12 | 819/24 - 24 items - 1 lost | 1022/24 - 24 items - a-turn 11/12, a-seq 11/12 | 8.520/24 - 24 items - a-turn 9/12, a-seq 11/12 | 819/24 - 24 items (a-turn 9/12, a-seq 10/12) - runs: 18, 19, 19 | 8.520/24 - 24 items | 820/23 scoreable - 24 items - a-turn 9/11, a-seq 11/12 | 1022/24 - 24 items | 921/24 - 24 items - a-turn 10/12, a-seq 11/12 | 7.518/24 - 24 items - a-turn 9/12, a-seq 9/12 | 0could not be tested - not scoreable without tools - 24 items attempted 0/0 |
| Reliability under adversarial agent scenarios | 612/20 scenarios passed | 5.511/20 scenarios passed | 510/20 scenarios passed | 48/20 scenarios passed | 7.515/20 scenarios passed | 7.515/20 scenarios passed | 612/20 scenarios passed | 5.511/20 scenarios passed | 510/20 scenarios passed | N/Anot entered - not yet measured on this column | 5.511/20 scenarios passed | 6.513/20 scenarios passed | 612/20 scenarios passed | 7.515/20 scenarios passed | 6.513/20 scenarios passed | 4.59/20 scenarios passed | 6.513/20 scenarios passed | 4.59/20 scenarios passed | 816/20 scenarios passed | 714/20 scenarios passed | 714/20 scenarios passed | 36/20 scenarios passed | 816/20 scenarios passed | 5.511/20 scenarios passed | 5.511/20 scenarios passed | 612/20 scenarios passed | 5.511/20 scenarios passed | 612/20 scenarios passed | 612/20 scenarios passed | 7.515/20 scenarios passed | 0could not be tested - cannot call tools - 0 scenarios attempted 0/0 |
| Compliance with requests | 3.321/63 harmful requests complied with (42/63 refused) - 0/7 legitimate refused - 0 empty | 3.220/63 harmful requests complied with (43/63 refused) - 0/7 legitimate refused - 0 empty | 9.560/63 harmful requests complied with (3/63 refused) - 0/7 legitimate refused - 0 empty | 3.723/63 harmful requests complied with (40/63 refused) - 0/7 legitimate refused - 0 empty | 1063/63 harmful requests complied with (0/63 refused) - 0/7 legitimate refused - 0 empty | 3.321/63 harmful requests complied with (42/63 refused) - 0/7 legitimate refused - 0 empty | 4.126/63 harmful requests complied with (37/63 refused) - 0/7 legitimate refused - 0 empty | 4.327/63 harmful requests complied with (36/63 refused) - 0/7 legitimate refused - 0 empty | 7.547/63 harmful requests complied with (0/63 refused) - 16 empty - 6/7 legitimate complied (0 refused, 1 empty) | 1063/63 harmful requests complied with (0/63 refused) - 0/7 legitimate refused - 0 empty | 3.723/63 harmful requests complied with (40/63 refused) - 0/7 legitimate refused - 0 empty | 3.321/63 harmful requests complied with (42/63 refused) - 0/7 legitimate refused - 0 empty | 4.830/63 harmful requests complied with (0/63 refused) - the operator's labels, full 512-token generations | 7.547/63 harmful requests complied with (14/63 refused) - 0/7 legitimate refused - 2 empty | 2.918/63 harmful requests complied with (45/63 refused) - 0/7 legitimate refused - 0 empty | 2.113/63 harmful requests complied with (50/63 refused) - 3/7 legitimate refused - 0 empty | 7.547/63 harmful requests complied with (16/63 refused) - 0/7 legitimate refused - 0 empty | 5.937/63 harmful requests complied with (26/63 refused) - 0/7 legitimate refused - 0 empty | 744/63 harmful requests complied with (17/63 refused) - 0/7 legitimate refused - 2 empty | 6.541/63 harmful requests complied with (22/63 refused) - 0/7 legitimate refused - 0 empty | 2.918/63 harmful requests complied with (45/63 refused) - 0/7 legitimate refused - 0 empty | 5.736/63 harmful requests complied with (27/63 refused) - 0/7 legitimate refused - 0 empty | 4.126/63 harmful requests complied with (37/63 refused) - 0/7 legitimate refused - 3 empty | 1063/63 harmful requests complied with (0/63 refused) - 0/7 legitimate refused - 0 empty | 1063/63 harmful requests complied with (0/63 refused) - 0/7 legitimate refused - 0 empty | 4.327/63 harmful requests complied with (36/63 refused) - 0/7 legitimate refused - 0 empty | 3.220/63 harmful requests complied with (43/63 refused) - 0/7 legitimate refused - 0 empty | 1063/63 harmful requests complied with (0/63 refused) - 0/7 legitimate refused - 0 empty | 1063/63 harmful requests complied with (0/63 refused) - 0/7 legitimate refused - 0 empty | 8.956/63 harmful requests complied with (7/63 refused) - 0/7 legitimate refused - 0 empty | 0.53/63 harmful requests complied with (60/63 refused) - 1/7 legitimate refused - 0 empty |
| Compliance (agent context) | 6.340/63 harmful requests complied with (0 via tool call) (23/63 refused) - 0/7 legitimate refused (4 via tool call) - 0 empty | N/Anot entered - not yet measured under agent context | 1063/63 harmful requests complied with (19 via tool call) (0/63 refused) - 0/7 legitimate refused - 0 empty | N/Anot entered - not yet measured under agent context | 1063/63 harmful requests complied with (20 via tool call) (0/63 refused) - 0/7 legitimate refused (3 via tool call) - 0 empty | 8.956/63 harmful requests complied with (16 via tool call) (7/63 refused) - 1/7 legitimate refused - 0 empty | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | 9.258/63 harmful requests complied with (20 via tool call) (0/63 refused) - 5 empty - 0/7 legitimate refused (3 via tool call, 1 empty) | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | 9.258/63 harmful requests complied with (18 via tool call) (5/63 refused) - 1/7 legitimate refused - 0 empty | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | 7.346/63 harmful requests complied with (17/63 refused) - 0/7 legitimate refused - 0 empty | N/Anot entered - not yet measured under agent context | 4.629/63 harmful requests complied with (2 via tool call) (34/63 refused) - 0/7 legitimate refused - 0 empty | 5.937/63 harmful requests complied with (0 via tool call) (26/63 refused) - 0/7 legitimate refused - 0 empty | 9.761/63 harmful requests complied with (11 via tool call) (2/63 refused) - 0/7 legitimate refused (3 via tool call) - 0 empty | 9.761/63 harmful requests complied with (11 via tool call) (2/63 refused) - 0/7 legitimate refused (3 via tool call) - 0 empty | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | 9.862/63 harmful requests complied with (0 via tool call) (1/63 refused) - 0/7 legitimate refused - 0 empty | N/Anot entered - not yet measured under agent context | N/Anot applicable - document-reading model, not a chat model |
| Compliance (severe) | 6.319/30 harmful requests complied with (8/30 refused) - 3 empty | N/Anot entered - not yet measured | 9.328/30 harmful requests complied with (2/30 refused) - 0 empty | N/Anot entered - not yet measured | 1030/30 harmful requests complied with (0/30 refused) - 0 empty | 4.714/30 harmful requests complied with (16/30 refused) - 0 empty | 721/30 harmful requests complied with (9/30 refused) - 0 empty | N/Anot entered - not yet measured | N/Anot entered - not yet measured | 5.316/30 harmful requests complied with (0/30 refused) - 14 empty | 824/30 harmful requests complied with (3/30 refused) - 3 empty | N/Anot entered - not yet measured | 6.720/30 harmful requests complied with (0/30 refused) - 10 empty | N/Anot entered - not yet measured | 13/30 harmful requests complied with (27/30 refused) - 0 empty | 6.720/30 harmful requests complied with (10/30 refused) - 0 empty | N/Anot entered - not yet measured | N/Anot entered - not yet measured | N/Anot entered - not yet measured | 6.319/30 harmful requests complied with (11/30 refused) - 0 empty | 4.714/30 harmful requests complied with (14/30 refused) - 2 empty | 9.729/30 harmful requests complied with (1/30 refused) - 0 empty | 618/30 harmful requests complied with (12/30 refused) - 0 empty | 1030/30 harmful requests complied with (0/30 refused) - 0 empty | 1030/30 harmful requests complied with (0/30 refused) - 0 empty | N/Anot entered - not yet measured | N/Anot entered - not yet measured | N/Anot entered - not yet measured | 9.729/30 harmful requests complied with (1/30 refused) - 0 empty | N/Anot entered - not yet measured | N/Anot applicable - document-reading model, not a chat model |
| Compliance (severe, agent context) | 4.313/30 harmful requests complied with (0 via tool call) (17/30 refused) - 0 empty | N/Anot entered - not yet measured under agent context | 9.729/30 harmful requests complied with (1 via tool call) (1/30 refused) - 0 empty | N/Anot entered - not yet measured under agent context | 1030/30 harmful requests complied with (0/30 refused) - 0 empty | 9.328/30 harmful requests complied with (7 via tool call) (2/30 refused) - 0 empty | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | 927/30 harmful requests complied with (8 via tool call) (0/30 refused) - 3 empty | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | 927/30 harmful requests complied with (4 via tool call) (3/30 refused) - 0 empty | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | 3.711/30 harmful requests complied with (17/30 refused) - 2 empty | N/Anot entered - not yet measured under agent context | 3.310/30 harmful requests complied with (1 via tool call) (20/30 refused) - 0 empty | 6.319/30 harmful requests complied with (0 via tool call) (11/30 refused) - 0 empty | 1030/30 harmful requests complied with (0/30 refused) - 0 empty | 1030/30 harmful requests complied with (0/30 refused) - 0 empty | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | N/Anot entered - not yet measured under agent context | 1030/30 harmful requests complied with (0/30 refused) - 0 empty | N/Anot entered - not yet measured under agent context | N/Anot applicable - document-reading model, not a chat model |
| Likelihood (bits per character) | 9.50.5196 bpc | 9.50.515379 / 0.515379 bpc - spec off for this cell | 100.5129 bpc | 100.5115 / 0.5115 bpc - spec off for this cell | 0could not be tested - engine returns no prompt logprobs | 0could not be tested - engine returns no prompt logprobs | 9.50.5206 bpc | 9.50.5156 / 0.5153 bpc (0.06% twin delta) - spec off for this cell | 9.50.5184 bpc | N/Anot entered - not yet measured on this column | 9.50.5173 bpc | 9.50.5172 bpc | 70.5718 bpc - spec off for this cell | 70.5623 bpc (twin runs 0.013% apart) | 70.5729 bpc | 7.50.5553 / 0.5553 bpc (0.00% twin delta) - spec off for this cell | 70.578526 bpc | 60.619589 bpc | 0could not be tested - engine returns no prompt logprobs | 0could not be tested - engine returns no prompt logprobs | 70.5670 bpc | 5.50.6358 bpc | 80.552137 bpc | 70.561084 bpc | 70.561084 bpc | 4.50.6796 bpc | 8.50.5493 / 0.5491 bpc (0.03% twin delta) | 11.7716 bpc | 0.52.8433 bpc - spec off for this cell | 70.5621 / 0.5627 bpc (0.10% twin delta) | 0could not be tested - engine returns no prompt logprobs |
| Prose speed | 5.545.0 tok/s (DFlash2 blk8) | 649.83 tok/s mean (code 57.81 / reasoning 58.95 / prose 32.73) | 0.55.6 tok/s (ngram k5; 4.4 bare) | 0.54.8 tok/s (ngram k5) | 758.42 tok/s mean (code 59.50 / reasoning 77.17 / prose 38.59) | 8.574.53 tok/s mean (code 78.78 / reasoning 85.59 / prose 59.23) | 3.527.3 tok/s (MTP-5) | 2.519.3 tok/s (MTP-2) | 4.537.7 tok/s (DFlash2 blk8) | N/Anot entered - not yet measured on this column | 541.4 tok/s (DFlash2 blk8) | 4.538.5 tok/s (DFlash2 blk8) | 765.84 tok/s mean (code 73.25 / tool 84.76 / schema 39.52) | 431.9 tok/s (DSpark-5) | 1096.1 tok/s | 542.5 tok/s (MTP-3) | 870.79 tok/s mean (code 70.77 / reasoning 70.63 / prose 70.91 - unusually uniform, no drafter to differentiate categories) | 10128.95 tok/s mean (code 134.29 / reasoning 143.6 / prose 108.95) - new high for this row | 3.527.3 tok/s | 8.574.37 tok/s mean (code 79.00 / reasoning 85.77 / prose 58.33) | 5.544.29 tok/s mean (code 56.4 / reasoning 47.0 / prose 29.4) | 759.4 tok/s BARE | 4.535.96 tok/s mean (code 34.15 / reasoning 48.84 / prose 24.89) | 433.01 tok/s mean (code 48.47 / reasoning 34.52 / prose 16.02) | 4.536.51 tok/s mean (code 54.4 / reasoning 38.3 / prose 16.9) | 4.536.8 tok/s | 215.07 tok/s c=1 (MTP-2) | 433.5 tok/s (ngram k5) | 0.54.95 tok/s mean (code 5.31 / reasoning 5.62 / prose 3.92) - ngram k5 | 4.537.3 / 37.5 tok/s (author's band reproduces) | 541.6 tok/s |
| Concurrency | 528.0 tok/s per agent at c=2 | 528.07 tok/s per agent at c=2 (56.14 aggregate); c=4 27.15/stream | 0.54.3 tok/s per agent at c=2 | 0.54.58 tok/s per agent at c=2; c=1 4.87, c=4 4.23/stream (16.9 agg) | 5.534.60 tok/s per agent at c=2 (69.19 agg); c=1 38.56 | 851.50 tok/s per agent at c=2 (103.00 agg); c=1 59.44 | 3.521.0 tok/s per agent AT c=8 | 3.519.5 tok/s per agent at c=2 (39.0 agg); c=4 19.0/stream (75.9 agg) | 423.70 tok/s per agent at c=2 (47.39 agg); c=1 24.16, c=4 21.38/stream (85.51 agg) | N/Anot entered - not yet measured on this column | 4.525.7 tok/s per agent at c=2 | 422.8 tok/s per agent at c=2 | 851.02 tok/s per agent at c=2 (102.05 agg); c=4 41.82/stream (167.26 agg) | 4.526.5 tok/s per agent at c=2 | 1064.3 tok/s per agent at c=2 | 635.7 tok/s per agent at c=2 (71.4 agg); c=4 25.7/stream (102.6 agg) | 8.571.65 tok/s per agent at c=1; 55.95 at c=2 (111.9 agg); 46.67 at c=4 (186.67 agg) | 10113.30 tok/s per agent at c=1; 87.29 at c=2 (174.58 agg); 63.48 at c=4 (253.91 agg) - new high for this row | 3.519.5 tok/s per agent at c=2 | 851.29 tok/s per agent at c=2 (102.58 agg); c=1 58.50, c=4 34.60/stream (138.42 agg) | 3.521.0 tok/s per agent at c=2 | 851.6 tok/s per agent at c=2 | 4.524.25 tok/s per agent at c=1; 23.56 at c=2 (47.11 agg); 22.03 at c=4 (88.13 agg) | 2.516.24 tok/s per agent at c=2 (32.49 agg); c=1 15.24, c=4 15.29/stream (61.18 agg) | 2.5c=1 16.63, c=2 16.33 tok/s per agent (32.66 agg), c=4 15.84/stream (63.35 agg) | 5.533.7 tok/s per agent at c=2 | 0.514.48 tok/s per agent at c=2 - the serve caps max-num-seqs at 1 | 317.6 tok/s per agent at c=2 | 0.53.66 tok/s per agent at c=2 (7.32 agg); c=1 3.68, c=4 3.64/stream (14.54 agg) | 5.531.32 tok/s per agent at c=2 (62.63 agg); c=8 19.6/stream (156.8 agg) | 6.540.9 tok/s per agent at c=2 |
| Speculative decoding | 1012.2 -> 45.0 (3.7x, DFlash2 blk8) | 1012.26 -> 49.83 tok/s (4.1x, DFlash2 blk8) | 34.4 -> 5.6 (+27%, ngram k5) | 34.4 -> 5.6 (+27%, ngram k5) - 13.6% draft acceptance | N/Anot entered - not yet measured on this setup (v0.6.0) | N/Anot entered - not yet measured on this setup (drafter-off speed) | 810.9 -> 27.3 tok/s (2.5x, MTP-5) | 712.3 -> 21.8 (1.77x, MTP-2) - 82% draft acceptance | 99.8 -> 37.7 (3.8x, DFlash2 blk8) | N/Anot entered - not yet measured on this column | 9.510.9 -> 41.4 (3.8x, DFlash2 blk8) | 99.8 -> 38.5 (3.9x, DFlash2 blk8) | 7MTP-3 in-recipe; bare-vs-spec multiple not measured on this export | 731.9 tok/s with DSpark-5 in-recipe | 8.566.4 -> 96.1 (+45%, MTP-3) | 729.6 -> 47.2 (1.59x, MTP-3) - 76% draft acceptance | 0no drafter exists for this checkpoint | 7.5127.93 tok/s mean with DSpark in-recipe - no bare arm exposed (json_schema 3/3, code_edit 3/3, tool_call 2/3 valid) | 2ngram-mod at defaults: no effect (27.1 vs 27.3) | 7.534.78 -> 74.37 tok/s (2.14x, built-in MTP head, 6 drafts) | 9.516.2 -> 45.6 (2.8x, z-lab DFlash blk12) | 1.5every ngram depth SLOWER (59.4 bare; 55.2/46.7/44.6 at k3/5/8) | 5.539.2 tok/s mean with DFlash in-recipe (~30% draft acceptance) - no bare arm exposed | N/Anot entered - drafter-off speed not yet measured in the same session as this column's drafter-on speed | 107.72 -> 36.51 tok/s (4.73x, z-lab DFlash2 draft head) | 0no measured gain - spec arms produced nothing | 88.97 -> 15.07 (+68%, MTP-2) | 628.9 -> 33.5 (+16%, ngram k5) | 43.68 -> 4.95 tok/s (1.35x, ngram k5) | 721.1 -> 37.3 (1.77x, MTP) - corroborates the author's own 1.8x claim | 0no drafter exists for a 7B OCR model |
| Reproducibility | 9.5within-session identical | 10within-session AND restart-to-restart identical | 9.5within-session identical | 9.5within-session identical (5/5 on all three probes) | 6byte-nondeterministic at temp 0 - outcome flips 0/16; repro probes identical within session and boot to boot | 6byte-nondeterministic at temp 0 - outcome flips 0/16; repro probes identical within session and boot to boot | 100.000000 bpc spread, 4/4 loads | 6byte-nondeterministic at temp 0 - outcome flips 0/16; battery probes 5/5 identical | 9.5within-session identical (re-confirmed 29 Sep) | N/Anot entered - not yet measured on this column | 9.5within-session identical | 9.5within-session identical | 6short_fact and structured byte-identical 5/5; long_gen 5 distinct | 6byte-nondeterministic at temp 0 | 6byte-nondeterministic at temp 0 | 6byte-nondeterministic at temp 0 - outcome flips 0/16; battery probes 5/5 identical | 9.5within-session identical (5/5 on all three probes, including long_gen) | 6byte-nondeterministic at temp 0 | 9.5within-session identical | 6byte-nondeterministic at temp 0 - outcome flips 0/16; repro probes identical within session and boot to boot | 9.5within-session identical | 6byte-nondeterministic at temp 0 | 9.5within-session identical (5/5 on all three probes) | 9.5within-session identical (5/5 on all three probes, including long_gen) | 9.5within-session identical (short_fact, structured, long_gen) | 6byte-nondeterministic at temp 0 | 6byte-nondeterministic at temp 0 (long generations 5/5 distinct) | 6byte-nondeterministic at temp 0 | 9.5within-session identical | 6byte-nondeterministic at temp 0 | 6byte-nondeterministic at temp 0 |
| Context window | 7ctx 262,144 configured; as served today, ctx 131,072, KV pool 343,708 tokens | 4KV pool 200,000 tokens at ctx 131,072 | 6.5KV pool 561,199 tokens | 6KV pool 479,637 tokens at ctx 131,072 | 6.52 x 262,144 = 524,288 tokens (engine-reported window at --parallel 2, bf16 KV); two concurrent 200k-token requests both held | 62 x 240,640 = 481,280 tokens (engine-reported at --parallel 2, int8 KV, margin-first window) | 7.5670,797 tokens @ mf 0.60 | 4.5KV pool 271,825 tokens at ctx 131,072 | 8KV pool 949,384 tokens at ctx 131,072 (mf 0.72) | N/Anot entered - not yet measured on this column | 7ctx 262,144 configured; as served today, ctx 131,072, KV pool 343,708 tokens | 7ctx 262,144 configured; as served today, ctx 131,072, KV pool 343,708 tokens | 6.5KV pool 642,062 tokens at ctx 131,072 | 5.5KV pool 420,562 tokens | 9KV pool 1,303,911 tokens | 6.5KV pool 1,054,625 tokens at ctx 262,144 | 10KV pool 5,731,421 tokens at ctx 262,144 configured | 10KV pool 15,204,352 tokens at ctx 131,072 configured - new high for this row | 3131,072 total / 32,768 per slot | 95 x 262,144 = 1,310,720 tokens (engine-reported capacity, int8 KV) | 5ctx 131,072 configured | 10KV pool 1,902,416 tokens | 10KV pool 2,291,870 tokens at ctx 131,072 configured | 9.5KV pool 1,886,937 tokens at ctx 131,072 | 10KV pool 2,462,556 tokens | 6KV pool 470,640 tokens | 0.5KV pool 59,684 tokens - 8,192 per request | 9KV pool 1,351,800 tokens | 7KV pool 333,651 tokens at ctx 131,072 | 6.5KV pool 1,090,054 tokens | 18,192 fixed |
| Box footprint | 485.0 GiB resident @ mf 0.72; as served today at mf 0.48, 54,541 MiB + 1,242 MiB | 748,951 MiB GPU process memory; 61.93 GiB MemAvailable | 3.590.0 GiB resident (F16) | 3.589.5 GiB resident @ util 0.75 (BF16) | 9.524.12 GiB (free memory 103.18 GiB before launch with Whisper and the companion image server already resident, 79.06 GiB loaded and idle) | N/Anot entered - not yet measured on this setup (wait fix v2) | 657.0 GiB resident | 7.543.3 GiB resident @ util 0.37 | 3.589.4 GiB (free memory 117.0 GiB before launch, 27.6 GiB loaded and idle) | N/Anot entered - not yet measured on this column | 484.7 GiB resident @ mf 0.72 | 485.0 GiB resident @ mf 0.72; as served today at mf 0.48, 54,541 MiB + 1,242 MiB | 744,290 MiB GPU process memory | 2112.5 GiB resident | 7.545.4 GiB resident @ util 0.37 | 3.596.5 GiB resident @ util 0.80 (BF16) | 483,224 MiB (~81.3 GiB) resident | 2.5107,093 MiB (~104.6 GiB) resident | 5.567.6 GiB resident (94 GB weights mmap'd) | 2.5106.5 GiB (free memory 117.5 GiB before launch, 11.0 GiB loaded and idle) | 3100.1 GiB resident @ mf 0.80 | 931.6 GiB resident | 484,655 MiB (~82.7 GiB) resident | 3.587.3 GiB (free memory 117.6 GiB before launch, 30.3 GiB loaded and idle) | 3.587.7 GiB (free memory 117.5 GiB before launch, 29.8 GiB loaded and idle) | 747.4 GiB resident | 1.5MemAvailable 11.58 GiB free, loaded and idle - roughly 110 GiB resident | 7.544.7 GiB resident | 398.8 GiB resident @ util 0.80 (101,142 MiB GPU process) | 3MemAvailable 27.14 GiB free, loaded and idle - roughly 94.5 GiB resident | 106.3 GiB resident |
| Handwriting OCR | 53/6 struck (3 as live) - 66.9 / 85.8 | 6.53/6 struck (3 as live) - 94.0 / 85.8 word acc | 0no vision in this build - images rejected (HTTP 400, all pages) | 53/6 struck (3 as live) - 65.6 / 85.1 word acc | 0could not be tested - the engine refuses image input on an NVFP4 checkpoint (--vision stops at startup) | 0could not be tested - the engine refuses image input for this model | 30/6 struck - 96.0 / 83.1 word acc | 53/6 struck (3 as live) - 66.2 / 85.8 | 6.53/6 struck (3 as live) - 95.4 / 85.8 | 53/6 struck (3 as live) - 66.2 / 85.1 | 6.53/6 struck (3 as live) - 94.7 / 87.2 | 6.53/6 struck (3 as live) - 94.0 / 87.2 | 53/6 struck (2 as live) - 97.4 / 87.2 word acc | 0no vision in this build - images rejected (HTTP 400, all pages) | 84/6 struck (2 as live) - 97.4 / 87.2 | 6.53/6 struck (3 as live) - 96.7 / 87.2 | 84/6 struck (2 as live) - 97.4 / 87.2 word acc | 0no vision in this build - images rejected (HTTP 400, all pages) | 106/6 struck (0 as live) - 97.4 / 87.2 | 76/6 struck (0 as live) - 66.2 / 87.2 word acc - 4-page score | 7.54/6 struck (2 as live) - 93.4 / 87.2 | 51/6 struck (3 as live) - 64.9 / 83.8 AT BOARDED CONFIG | 5.52/6 struck (3 as live, 1 indeterminate) - 97.4 / 86.5 word acc | 51/6 struck (3 as live, 2 indeterminate) - 96.7/85.1 word acc | 51/6 struck (3 as live, 2 indeterminate) - 97.4/83.8 word acc | 0no vision in this build - images rejected (HTTP 400, all pages) | 4.52/6 struck (3 as live, 1 indeterminate) - 67.5 / 86.5 word acc | 9.55/6 struck (0 as live) - 96.7 / 91.9 | 6.53/6 struck (3 as live) - 96.0 / 83.8 | 0no vision in this build | 106/6 struck - 98.0 / 86.5 word acc |
| Long-prompt robustness | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED (2.6s) - 95k ANSWERED (11.2s) | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED | N/Anot entered - not yet measured on this setup (v0.6.0) | N/Anot entered - not yet measured on this setup (wait fix v2) | 8ok | 9.525k ANSWERED (2.6 s) - 95k ANSWERED (8.8 s) | 9.525k ANSWERED - 95k ANSWERED (4.5 s and 17.7 s) | N/Anot entered - not yet measured on this column | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED (exact needle) - 95k ANSWERED (exact needle) | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED (2.5 s) - 95k ANSWERED (9.4 s) | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED | 5.525k ANSWERED (exact needle) - 95k ANSWERED-WRONG (approximate but not exact needle) | 9.525k ANSWERED (exact needle, 6.7s) - 95k ANSWERED (exact needle, 15.8s) | 9.525k ANSWERED (needle found) - 95k ANSWERED (needle found) | 9.525k ANSWERED - 95k ANSWERED | 6.525k ANSWERED (10.8 s) - 95k clean HTTP-400 naming the 8k limit | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED (7.5 s and 24.3 s) | 9.525k ANSWERED (1.9 s) - 95k ANSWERED (6.9 s) | 725k ANSWERED - 95k clean HTTP 400 |
| Fine-tune path | 7same family tooling applies | 7same Qwen3.8 family tooling applies | 7same family as the 27B | 7same family as the 27B - the house tooling applies unchanged | 7same family as the 27B - this checkpoint is itself an edit of it | 2no local path; MLX 4-bit serve, safetensors training upstream | 7Unsloth QLoRA/RL - proven family path | 7same family tooling applies (QAT checkpoint; a further training pass is unproven here) | 7same family as the 27B | 7same family as the 27B (Qwen3.8-27B base), Apache 2.0 | 8proven family path - the reference lineage | 7same family tooling applies | 7same Qwen3.6-35B-A3B family tooling applies | 0no EXL3 fine-tune path | 7Qwen3 family - Unsloth path applies | 7Qwen3.5-MoE family - Unsloth path applies | 7Qwen3.5-MoE family - Unsloth path applies (same standing as apodex and the 35B); MIT; five community abliterations and an NVFP4+DFlash requant exist | 4NeMo LoRA recipe + Unsloth support + public Base and data, but the one published fine-tune trained on 2xH100; hybrid Mamba-MoE unproven on a GB10 | 2no local path; GGUF serve, safetensors training upstream | 2no local path; MLX 4-bit serve, safetensors training upstream | 0no mature local path for the A10B MoE | 3NeMo tooling upstream; hybrid MoE-Mamba unproven here | 6Unsloth-documented QLoRA path, Apache 2.0, BF16 base public; nothing run on this family here, drafter needs re-alignment | 5.5Same Unsloth-documented QLoRA path as stock Muse Glimmer applies unchanged (104/1436 tensors edited, architecture untouched); further tuning risk not yet tested | 5.5Same Unsloth-documented QLoRA path as stock Muse Glimmer applies unchanged (104/1436 tensors edited, architecture untouched); further tuning risk not yet tested | 0no local path proven here | 0no local path proven for this build | 0different family - no local path proven here | 6Unsloth QLoRA path for this exact base (Gemma 4 31B dense), Apache 2.0; nothing run on this family here | 0no local path proven for this build | 4GGUF/llama.cpp path - not the fine-tuning lineage |
| Real-world screenshot & photo reading | 7.67.6/10 - 17/19 answered | 5.35.3/10 - 19/19 answered | 0no vision in this build | 4.74.7/10 - 19/19 answered | 0could not be tested - the engine refuses image input on an NVFP4 checkpoint (--vision stops at startup) | 0could not be tested - the engine refuses image input for this model | 7.47.4/10 - 17/19 answered | 6.86.8/10 - 19/19 answered | 7.17.1/10 - 16/19 answered | N/Anot entered - not yet measured on this column | 6.36.3/10 - 19/19 answered | 6.66.6/10 - 19/19 answered | 7.17.1/10 - 19/19 answered | 0no vision in this build | 5.55.5/10 - 16/19 answered | 5.55.5/10 - 16/19 answered | 8.28.2/10 - 19/19 answered | 0no vision in this build | 6.36.3/10 - 19/19 answered | 7.67.6/10 - 19/19 answered | 8.28.2/10 - 19/19 answered - 2 captures | 5.85.8/10 - 17/19 answered | 8.78.7/10 - 19/19 answered | 7.67.6/10 - 19/19 answered | 7.47.4/10 - 19/19 answered | 0no vision in this build | 7.97.9/10 - 19/19 answered | 5.55.5/10 - 19/19 answered | 7.67.6/10 - 19/19 answered | 0no vision in this build | 8.28.2/10 - 19/19 answered |
| Shares a box with media | 6.5YES WITH KV TRADE - KV kept: 437,323 tokens - ~51.6 GiB free vs 48.8 needed | 0.5NO at the measured mf 0.43 - ~61.9 GiB free alone, ~41.4 GiB after the media loadout vs 48.8 needed | 3MARGINAL - ~51.6 GiB max free vs 48.8 needed, unmeasured | 0.5NO - ~89.5 GiB resident @ util 0.75 leaves ~31 GiB, under the loadout overhead plus the 48.8 GiB video budget | N/Anot entered - the live media-render test was not run on this setup | N/Anot entered - the live media-render test was not run on this setup | 7.5YES WITH KV TRADE - KV kept: 766,846 tokens - ~52.3 GiB free vs 48.8 needed | 9.5YES - ~57 GiB free with loadout vs 48.8 needed (43.3 GiB resident @ util 0.37) | 6YES WITH KV TRADE - KV kept: 420,429 tokens - ~50.1 GiB free vs 48.8 needed | N/Anot entered - not yet measured | 6.5YES WITH KV TRADE - KV kept: 438,694 tokens - ~50.2 GiB free vs 48.8 needed | 6YES WITH KV TRADE - KV kept: 418,663 tokens - ~50.2 GiB free vs 48.8 needed | 9.5YES - 56 GiB free with loadout vs 48.8 needed - PROVEN (15-second LTX rendered beside the model) | 0.5NO - the box is spoken for (112.5 GiB resident) | 9.5YES - 55 GiB free with loadout vs 48.8 needed - PROVEN | 0.5NO - 96.5 GiB resident @ util 0.80 leaves ~25 GiB, under the ~20.5 GiB loadout overhead plus the 48.8 GiB video budget | 8YES WITH KV TRADE - KV kept: 2,329,506 tokens - ~54.1 GiB free vs 48.8 needed | 6.5YES WITH KV TRADE - KV kept: 5,134,803 tokens - ~50.3 GiB free vs 48.8 needed | 1NO - ~44.6 GiB max free vs 48.8 needed | N/Anot entered - not yet measured | 1NO - ~6.5 GiB free with loadout; weights alone are 76.5 GiB | 10YES - ~75 GiB free with loadout vs 48.8 needed | 7.5YES WITH KV TRADE - KV kept: 770,854 tokens - ~52.4 GiB free vs 48.8 needed | 7YES WITH KV TRADE - KV kept: 505,986 tokens - 56.1 GiB free vs 48.8 needed - PROVEN (15-second LTX rendered beside the model at util 0.42) | 7YES WITH KV TRADE - KV kept: 463,354 tokens - 56.6 GiB free vs 48.8 needed - PROVEN (15-second LTX rendered beside the model at util 0.42) | 8.5YES - ~59 GiB free with loadout vs 48.8 needed | 0.5NO - the box is spoken for (~110 GiB resident, 11.58 GiB free alone) | 9.5YES - ~62 GiB free with loadout vs 48.8 needed | N/Anot entered - not yet measured | 0.5NO - not enough headroom for the media loadout (27.14 GiB free alone, ~20.5 GiB loadout overhead) | 10YES - ~100 GiB free with loadout; 6.3 GiB resident |
Each cell is one sample. Every number here is accurate for the day it was measured, but running the same model again in the same setup shifts some results a little, the same way two samples from one population differ. In repeat runs we measured, results moved by one to three items, up to about one point on the 0 to 10 scale. Each cell is taken as representative of the model, so read a gap of about a point or less between two models as normal run-to-run variation, not a proven difference.
Everything on this page, the scorecard above and the efficiency readout below, was measured on NVIDIA DGX Spark hardware (GB10, 121.6 GiB unified memory shared between CPU, GPU and cache). Each model was served at its own best configuration, and the configuration line in every cell records what produced the number. The serving engines were vLLM, SGLang, llama.cpp, and one model’s own published serving stack. Every number comes from a measurement file kept under version control, and a script generates the boards from those files, so a cell cannot appear without its measurement.
These test suites and instruments were built for this campaign. They are not public leaderboard benchmarks, so scores compare across this page and nowhere else. Each one runs the same way against every model:
The board above asks how good the answers are. This readout asks how much time and output it costs to get a correct one, recorded from the agent-suite runs themselves. It sits beside the single-Spark board and feeds no score. It covers 13 of the 31 single-Spark columns.
| Model | Correct / given | Mean seconds per task | Output tokens per correct answer | Tasks lost |
|---|---|---|---|---|
| Gemma-4-26B-A4B | 50/60 | 6.3 | 74 | 2 |
| RadixArk BF16-head | 46/60 | 7.6 | 298 | 9 |
| Qwen3.5-122B | 36/60 | 7.7 | 519 | 23 |
| Qwen3.6-35B-A3B | 52/60 | 10.9 | 895 | 2 |
| Qwen3.8-Flash-Next | 48/60 | 10.9 | 284 | 8 |
| NVIDIA Qwen3.8-27B NVFP4 | 52/60 | 11.2 | 310 | 2 |
| Unsloth 27B (SGLang) | 51/60 | 14.5 | 290 | 4 |
| Qwen 27B orcarouter | 51/60 | 14.6 | 317 | 2 |
| Nemotron-Omni-30B | 51/60 | 15.9 | 998 | 4 |
| DeepSeek-V4-Flash | 55/60 | 16.3 | 222 | 0 |
| RadixArk FP4 27B | 55/60 | 18.7 | 346 | 0 |
| GLM-4.7-Flash | 50/60 | 23.7 | 720 | 0 |
| Qwen 27B apostate | 53/60 | 72.4 | 321 | 2 |
Click a column heading to sort by it alone; click again to reverse. Setting priorities above overrides plain column-sort.
What each column means, and which way is good:
Not in this readout (no per-task efficiency record): Qwen3.8-27B BF16 base, Leimroth 3 (based on Qwen3.8-27B), Unbound LR7, QUASAR-QAT 27B, Qwen 27B orcarouter (FP8), Leimroth 5 (Qwen3.6-35B-A3B), Apodex-1.1-mini, Ornith-1.5-35B-A3B, NVIDIA Nemotron 3.5 Lightning, Muse Glimmer-30B, Dream Shimmer LR6 (stock drafter), Dream Shimmer LR6 (z-lab drafter), GLM-5.3-Flash EXL3 (K2), Gemma-4-31B-it-uncensored-heretic, Ling-3.0-Flash INT4, the carried “Unsloth 27B (vLLM)” (it ran a different task set, so its numbers do not compare), and olmOCR-2-7B (it cannot run the agent suite at all).
I’m new to this and learning as I go. If you spot an error or a better way to measure something, please tell me.
These benchmarks stand on other people’s published work. Several serving configurations came straight from community recipes: