Tensor parallelism splits one model’s weights across both Sparks. That admits models too big for one box and changes the economics of everything else. Same rows, same instruments and same rules as the single-Spark board, and the same controls: click rows in priority order, remove models, read the config line on every cell.
Stability. Every configuration on this board serves and holds under sustained load, with no open cross-box failures. GLM-4.7 is the one that needed special treatment. It only serves when given a bigger share of the box’s memory than any other configuration here is allowed, and it pays for that room in context. It holds the least conversation memory of any configuration measured, and its cells show the trade.
| What is measuredone row per measure · one column per model | Qwen3.8-Flash-NextTP=2 · SGLang | GLM-5.3-Flash (RedHat)TP=2 · re-attributed checkpoint | Qwen3.5-122BTP=2 | DeepSeek-V4-FlashTP=2 · 1M context | GLM-5.3-FlashTP=2 | Qwen3-235BTP=2 · A22B | GLM-4.7 (full)TP=2 · NVFP4 · util 0.88 · context-thin | GLM-5.3-Flash EXL3 (kit c707598e)TP=2 · MiaAI kit, newer pin · DFlash k7 · speed arms only |
|---|---|---|---|---|---|---|---|---|
| Agent tasks | 8.554/60 - 60 items - ZERO artifacts (a-tool 10/12, a-seq 11/12, a-schema 11/12, a-turn 10/12, a-ctx 12/12 PERFECT) | 1056/60 - 60 items - 0 LOST, ZERO artifacts | 954/60 then 52/60 - 60 items - 0 LOST (a few items exceed the window: clean 400s) | 8.553/60 - 60 items - 0 LOST - ZERO artifacts | 1056/60 - 60 items - 0 LOST, ZERO artifacts (graphs arm, modelopt; EXL3 arm 55/60, 0 lost) | 7.551/60 then 49/60 - 60 items - 0 LOST (8 a-ctx items exceed the 32k ctx: clean 400s, same items both reps) | 748/60 - 60 items - 0 LOST (5 a-ctx exceed the 64k window, 2 a-turn hit the token cap) | N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column |
| Tools & structured output | 821/24 - 24 items - a-tool 10/12, a-schema 11/12 | 9.522/24 items | 1024/24 - 24 items - PERFECT (a-tool 12/12, a-schema 12/12) | 922/24 - 24 items - a-tool 12/12, a-schema 10/12 | 1023/24 items (graphs arm: a-tool 12/12 + a-schema 11/12; EXL3 arm 12/12) | 9.523/24 - 24 items - identical both reps (a-tool 12/12, a-schema 11/12) | 922/24 - 24 items - a-tool 12/12, a-schema 10/12 | N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column |
| Multi-turn & sequencing | 8.521/24 - 24 items - a-turn 10/12, a-seq 11/12 | 921/24 items | 8.520/24 - 24 items - a-seq 11/12, a-turn 9/12 | 819/24 - 24 items - a-turn 10/12, a-seq 9/12 | 921/24 items (graphs arm: a-turn 11/12 + a-seq 10/12) - EXL3 arm 20/24 | 1024/24 then 22/24 - 24 items - a-turn 12/12 both reps | 819/22 scoreable - 24 items - a-turn 8/10, a-seq 11/12 | N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column |
| Reliability under adversarial agent scenarios | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only |
| Compliance with requests | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only |
| Compliance (agent context) | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only |
| Compliance (severe) | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only |
| Compliance (severe, agent context) | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only | N/Anot entered - single-Spark programme only |
| Likelihood (bits per character) | 9.50.5180 / 0.5174 bpc (0.12% twin delta) | 90.5229 bpc | 70.5667 bpc - measured ON the two-box serve | 7.50.5555 / 0.5564 bpc twins (0.16% apart) | 80.537321 / 0.537219 bpc twins, 0.02% delta (graphs serve, modelopt) - EXL3 arm 0.5251 | 6.50.5852 / 0.5865 bpc twins (0.23% apart) | 6.50.5832 / 0.5800 bpc twins - 0.55% apart, the WIDEST spread measured | N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column |
| Prose speed | 4.540.93 tok/s c=1 (ttft 0.182 s) - +62% over the superseded config's 25.2 | 4.514.54 tok/s c=1 bare, 28.37 with MTP-4 (1.95x) | 2.520.28 tok/s bare eager on the fixed build (+23% over one box bare) | 646.35 tok/s (DSpark-5, ttft 0.175 s) | 3.523.47 tok/s c=1 (graphs serve, modelopt, MTP-4) - EXL3 arm 36.4 (DFlash2 k7 in-recipe) | 215.97 / 15.85 tok/s c=1 twins (fixed serve; TTFT 0.17 s) | 217.16 tok/s (ttft 0.162 s) | 3.524.4 tok/s c=1 (DFlash k7; mean decode 37.5: code 46.6, reasoning 41.6) |
| Concurrency | 633.89 tok/s per agent at c=2 (67.77 agg); c=8 24.30/stream (194.43 agg) | 2.514.22 tok/s per agent at c=2 (bare graphs serve) | 3.520.47 tok/s per agent at c=2 - FASTER per stream than solo | 5.532.73 tok/s per agent at c=2 - absolute numbers at last | 3c=1 23.47 / c=2 18.13 / c=4 13.56 tok/s per agent (graphs serve, modelopt) | 2.513.66 / 13.94 tok/s per agent at c=2 (twins, zero errors; c=4 holds 12.1/11.8) | 212.76 tok/s per agent at c=2 | N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column |
| Speculative decoding | 6MTP-3 baked into the recipe; no bare arm exposed | 714.54 -> 28.37 (1.95x, MTP-4 on the graphs serve) | 1.5no spec arm at TP=2 - z-lab drafter not ported | 7DSpark draft in-recipe; no bare arm exposed | 5MTP-4 active throughout this arm (no bare-vs-MTP A/B run on the corrected checkpoint) - EXL3 kit: bare 13.2 -> DFlash k7 37.5 (2.8x), MTP-2 26.7 | 1no MTP head in the checkpoint - 145,703 tensors, zero mtp/nextn | 0no speculative arm at this configuration | 913.2 -> 37.5 (2.8x, DFlash k7) - 84% draft acceptance; MTP-2 26.7 at 96% |
| Reproducibility | 6byte-nondeterministic at temp 0 (long generations 5/5 distinct) | 6byte-nondeterministic at temp 0 - outcome flips 1/16 | 6byte-nondeterministic at temp 0 (long generations diverge) | 6byte flips 16/16, outcome flips 0/16 - and pinning is measured shut | 6.5byte-nondeterministic at temp 0 - outcome flips 0/16 (graphs serve, modelopt) | 6byte-nondeterministic at temp 0 (long generations 5/5 distinct) | 6byte-nondeterministic at temp 0 (long generations 5/5 distinct) | N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column |
| Context window | 10KV pool 2,377,708 tokens - 1M-token YaRN window | 6.5KV pool 635,500 tokens | 10KV pool 3,455,361 tokens (fixed build; 3,546,542 pre-fix) | 10KV pool 2,461,175 tokens - ctx 1,048,576 per request | 5.5KV pool 507,041 tokens at ctx 262,144 (graphs serve, modelopt) | 1.5max_seq_len 32,768 (engine dump) - KV pool 345,440 / 334,896 tokens (twins) | 1KV pool 83,456 tokens - ctx 65,536 per request | 5.5KV pool 516,726 tokens at ctx 262,144 (DFlash arm; 1,087,746 with MTP-2, 1,426,897 bare) |
| Box footprint | 1.5head MemAvailable 5.89 GiB loaded, worker 10.38 GiB | 196.7 GiB resident on EACH of two boxes | 182.9 GiB resident on EACH of two boxes | 1100.9 GiB resident on EACH of two boxes | 199.11 GiB resident on EACH of two boxes (graphs serve, modelopt) | 185.4 / 84.9 GiB resident (twins) on EACH of two boxes | 1104.9 GiB resident on EACH of two boxes | N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column |
| Handwriting OCR | 6.55/6 struck (0 as live, 1 indeterminate) - 67.5 / 86.5 word acc | N/Anot entered - see the other GLM-5.3-Flash column | 7.54/6 struck (2 as live) - 94.7 / 87.2 word acc | 1images rejected (HTTP 400, all pages) | 0.5vision path PROVEN stable - transcription UNSCOREABLE at this config | 1images N/A - LANGUAGE-ONLY checkpoint | 1images N/A - text-only export | N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column |
| Long-prompt robustness | 9.525k ANSWERED - 95k ANSWERED (3.3 s and 8.5 s) | 9.525k ANSWERED - 95k ANSWERED | 9.525k ANSWERED - 95k ANSWERED (2.1 s and 5.7 s on the fixed build) | 9.525k ANSWERED - 95k ANSWERED (3.8 s and 10.5 s) | 9.525k ANSWERED 7.0s - 95k ANSWERED 13.6s (graphs serve, modelopt) | 9.525k ANSWERED - 95k ANSWERED (fixed-serve twins: 2.8/2.7 s and 9.9/9.7 s) | 9.525k ANSWERED - 95k ANSWERED (3.9 s and 16.0 s) | N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column |
| Real-world screenshot & photo reading | 7.17.1/10 - 19/19 answered | N/Anot entered - not measured | N/Anot entered - not measured | N/Anot entered - no vision in this build | N/Anot entered - not measured | N/Anot entered - no vision in this build | N/Anot entered - no vision in this build | N/Anot entered - not measured |
Each cell is one sample. Every number here is accurate for the day it was measured, but running the same model again in the same setup shifts some results a little, the same way two samples from one population differ. In repeat runs we measured, results moved by one to three items, up to about one point on the 0 to 10 scale. Each cell is taken as representative of the model, so read a gap of about a point or less between two models as normal run-to-run variation, not a proven difference.
Single vs dual-Spark context, in real terms. The default single-Spark model (RadixArk FP4 27B) holds 262,144 tokens, about 650 pages, which is longer than most novels. Qwen3.8-Flash-Next's dual-Spark configuration holds 2,377,708 tokens, about 9 times that, or 5,900 pages, the size of a multi-volume reference work or an entire codebase. Most single documents fit inside the single-Spark number. Dual-Spark is for the rare cases that don't: a huge archive, a full book series, a big monorepo.
These benchmarks stand on other people’s published work. Several serving configurations came straight from community recipes: