SCOTT LEIMROTH ← AI & Tech
GB10DGX Spark Benchmarks

Dual-Spark (TP=2) Scorecard: two Sparks serving one model

Tensor parallelism splits one model’s weights across both Sparks. That admits models too big for one box and changes the economics of everything else. Same rows, same instruments and same rules as the single-Spark board, and the same controls: click rows in priority order, remove models, read the config line on every cell.

Stability. Every configuration on this board serves and holds under sustained load, with no open cross-box failures. GLM-4.7 is the one that needed special treatment. It only serves when given a bigger share of the box’s memory than any other configuration here is allowed, and it pays for that room in context. It holds the least conversation memory of any configuration measured, and its cells show the trade.

Order the models your way. Each column is a model; each row is one thing measured. Click a ROW heading to make it priority 1, then another for priority 2, and so on. The models re-sort left to right by a weighted blend in which earlier priorities count more (with three picks the weights are 3, 2, 1). The small number under each model name is its current weighted standing. A row a model was never entered in counts as zero in that standing, so an untested model sorts after any model with a real score there, however small; the label says how many of your picks it is missing, and those cells stay N/A with their reason rather than being given a made-up score. Click a numbered heading again to drop it.

Click a MODEL name to remove its column, so you can strip the table down to the two or three you are comparing. Removed models appear above the table; click one to bring it back, or use Show all models.

Weighting: pick one below. It sets how much your first priority counts against your later picks when the models re-sort.
What is measuredone row per measure · one column per model Qwen3.8-Flash-NextTP=2 · SGLang GLM-5.3-Flash (RedHat)TP=2 · re-attributed checkpoint Qwen3.5-122BTP=2 DeepSeek-V4-FlashTP=2 · 1M context GLM-5.3-FlashTP=2 Qwen3-235BTP=2 · A22B GLM-4.7 (full)TP=2 · NVFP4 · util 0.88 · context-thin GLM-5.3-Flash EXL3 (kit c707598e)TP=2 · MiaAI kit, newer pin · DFlash k7 · speed arms only
Agent tasks 8.554/60 - 60 items - ZERO artifacts (a-tool 10/12, a-seq 11/12, a-schema 11/12, a-turn 10/12, a-ctx 12/12 PERFECT) 1056/60 - 60 items - 0 LOST, ZERO artifacts 954/60 then 52/60 - 60 items - 0 LOST (a few items exceed the window: clean 400s) 8.553/60 - 60 items - 0 LOST - ZERO artifacts 1056/60 - 60 items - 0 LOST, ZERO artifacts (graphs arm, modelopt; EXL3 arm 55/60, 0 lost) 7.551/60 then 49/60 - 60 items - 0 LOST (8 a-ctx items exceed the 32k ctx: clean 400s, same items both reps) 748/60 - 60 items - 0 LOST (5 a-ctx exceed the 64k window, 2 a-turn hit the token cap) N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column
Tools & structured output 821/24 - 24 items - a-tool 10/12, a-schema 11/12 9.522/24 items 1024/24 - 24 items - PERFECT (a-tool 12/12, a-schema 12/12) 922/24 - 24 items - a-tool 12/12, a-schema 10/12 1023/24 items (graphs arm: a-tool 12/12 + a-schema 11/12; EXL3 arm 12/12) 9.523/24 - 24 items - identical both reps (a-tool 12/12, a-schema 11/12) 922/24 - 24 items - a-tool 12/12, a-schema 10/12 N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column
Multi-turn & sequencing 8.521/24 - 24 items - a-turn 10/12, a-seq 11/12 921/24 items 8.520/24 - 24 items - a-seq 11/12, a-turn 9/12 819/24 - 24 items - a-turn 10/12, a-seq 9/12 921/24 items (graphs arm: a-turn 11/12 + a-seq 10/12) - EXL3 arm 20/24 1024/24 then 22/24 - 24 items - a-turn 12/12 both reps 819/22 scoreable - 24 items - a-turn 8/10, a-seq 11/12 N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column
Reliability under adversarial agent scenarios N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only
Compliance with requests N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only
Compliance (agent context) N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only
Compliance (severe) N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only
Compliance (severe, agent context) N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only N/Anot entered - single-Spark programme only
Likelihood (bits per character) 9.50.5180 / 0.5174 bpc (0.12% twin delta) 90.5229 bpc 70.5667 bpc - measured ON the two-box serve 7.50.5555 / 0.5564 bpc twins (0.16% apart) 80.537321 / 0.537219 bpc twins, 0.02% delta (graphs serve, modelopt) - EXL3 arm 0.5251 6.50.5852 / 0.5865 bpc twins (0.23% apart) 6.50.5832 / 0.5800 bpc twins - 0.55% apart, the WIDEST spread measured N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column
Prose speed 4.540.93 tok/s c=1 (ttft 0.182 s) - +62% over the superseded config's 25.2 4.514.54 tok/s c=1 bare, 28.37 with MTP-4 (1.95x) 2.520.28 tok/s bare eager on the fixed build (+23% over one box bare) 646.35 tok/s (DSpark-5, ttft 0.175 s) 3.523.47 tok/s c=1 (graphs serve, modelopt, MTP-4) - EXL3 arm 36.4 (DFlash2 k7 in-recipe) 215.97 / 15.85 tok/s c=1 twins (fixed serve; TTFT 0.17 s) 217.16 tok/s (ttft 0.162 s) 3.524.4 tok/s c=1 (DFlash k7; mean decode 37.5: code 46.6, reasoning 41.6)
Concurrency 633.89 tok/s per agent at c=2 (67.77 agg); c=8 24.30/stream (194.43 agg) 2.514.22 tok/s per agent at c=2 (bare graphs serve) 3.520.47 tok/s per agent at c=2 - FASTER per stream than solo 5.532.73 tok/s per agent at c=2 - absolute numbers at last 3c=1 23.47 / c=2 18.13 / c=4 13.56 tok/s per agent (graphs serve, modelopt) 2.513.66 / 13.94 tok/s per agent at c=2 (twins, zero errors; c=4 holds 12.1/11.8) 212.76 tok/s per agent at c=2 N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column
Speculative decoding 6MTP-3 baked into the recipe; no bare arm exposed 714.54 -> 28.37 (1.95x, MTP-4 on the graphs serve) 1.5no spec arm at TP=2 - z-lab drafter not ported 7DSpark draft in-recipe; no bare arm exposed 5MTP-4 active throughout this arm (no bare-vs-MTP A/B run on the corrected checkpoint) - EXL3 kit: bare 13.2 -> DFlash k7 37.5 (2.8x), MTP-2 26.7 1no MTP head in the checkpoint - 145,703 tensors, zero mtp/nextn 0no speculative arm at this configuration 913.2 -> 37.5 (2.8x, DFlash k7) - 84% draft acceptance; MTP-2 26.7 at 96%
Reproducibility 6byte-nondeterministic at temp 0 (long generations 5/5 distinct) 6byte-nondeterministic at temp 0 - outcome flips 1/16 6byte-nondeterministic at temp 0 (long generations diverge) 6byte flips 16/16, outcome flips 0/16 - and pinning is measured shut 6.5byte-nondeterministic at temp 0 - outcome flips 0/16 (graphs serve, modelopt) 6byte-nondeterministic at temp 0 (long generations 5/5 distinct) 6byte-nondeterministic at temp 0 (long generations 5/5 distinct) N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column
Context window 10KV pool 2,377,708 tokens - 1M-token YaRN window 6.5KV pool 635,500 tokens 10KV pool 3,455,361 tokens (fixed build; 3,546,542 pre-fix) 10KV pool 2,461,175 tokens - ctx 1,048,576 per request 5.5KV pool 507,041 tokens at ctx 262,144 (graphs serve, modelopt) 1.5max_seq_len 32,768 (engine dump) - KV pool 345,440 / 334,896 tokens (twins) 1KV pool 83,456 tokens - ctx 65,536 per request 5.5KV pool 516,726 tokens at ctx 262,144 (DFlash arm; 1,087,746 with MTP-2, 1,426,897 bare)
Box footprint 1.5head MemAvailable 5.89 GiB loaded, worker 10.38 GiB 196.7 GiB resident on EACH of two boxes 182.9 GiB resident on EACH of two boxes 1100.9 GiB resident on EACH of two boxes 199.11 GiB resident on EACH of two boxes (graphs serve, modelopt) 185.4 / 84.9 GiB resident (twins) on EACH of two boxes 1104.9 GiB resident on EACH of two boxes N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column
Handwriting OCR 6.55/6 struck (0 as live, 1 indeterminate) - 67.5 / 86.5 word acc N/Anot entered - see the other GLM-5.3-Flash column 7.54/6 struck (2 as live) - 94.7 / 87.2 word acc 1images rejected (HTTP 400, all pages) 0.5vision path PROVEN stable - transcription UNSCOREABLE at this config 1images N/A - LANGUAGE-ONLY checkpoint 1images N/A - text-only export N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column
Long-prompt robustness 9.525k ANSWERED - 95k ANSWERED (3.3 s and 8.5 s) 9.525k ANSWERED - 95k ANSWERED 9.525k ANSWERED - 95k ANSWERED (2.1 s and 5.7 s on the fixed build) 9.525k ANSWERED - 95k ANSWERED (3.8 s and 10.5 s) 9.525k ANSWERED 7.0s - 95k ANSWERED 13.6s (graphs serve, modelopt) 9.525k ANSWERED - 95k ANSWERED (fixed-serve twins: 2.8/2.7 s and 9.9/9.7 s) 9.525k ANSWERED - 95k ANSWERED (3.9 s and 16.0 s) N/Anot entered - speed arms only on this pin (02 Sep night); the 024db9f7 pin's cell sits in the GLM-5.3-Flash column
Real-world screenshot & photo reading 7.17.1/10 - 19/19 answered N/Anot entered - not measured N/Anot entered - not measured N/Anot entered - no vision in this build N/Anot entered - not measured N/Anot entered - no vision in this build N/Anot entered - no vision in this build N/Anot entered - not measured

Each cell is one sample. Every number here is accurate for the day it was measured, but running the same model again in the same setup shifts some results a little, the same way two samples from one population differ. In repeat runs we measured, results moved by one to three items, up to about one point on the 0 to 10 scale. Each cell is taken as representative of the model, so read a gap of about a point or less between two models as normal run-to-run variation, not a proven difference.

What each row measures
Agent tasks
Can it do the multi-step jobs your agents run: pick the right tool, follow the instructions, finish the job. Measured on 60 fixed tasks, 12 per category, drawn from the 197-item A-Bench suite; the one carried column ran an earlier 191-item version.
Tools & structured output
When a program needs a machine-readable answer back, a form filled in exactly, a tool called with the right arguments, does it come back in the shape the program expects, or as prose that breaks the caller.
Multi-turn & sequencing
Whether it keeps track over a long back-and-forth: remembers what it was told several messages ago and does things in the right order. Models that look fine in one-shot tests often fall over here.
Reliability under adversarial agent scenarios
This is the row that answers 'can something hijack my agent mid-task', not the row about refusing dangerous content. An agent reads things a human never approved: a web page, a file, a tool's output, and any of those can carry a hidden instruction trying to redirect it, an injected 'ignore your task and do X' buried in a document, a fake system message, a prompt telling it to touch a file outside its job. This row measures whether the model stays obedient to the ACTUAL task you gave it when that happens, across 20 fixed adversarial scenarios. A pass means it did only the job it was asked, called a tool only when the task needed one, and never claimed finished work it had not done; one unauthorised tool call, one touch of something off-limits, or one false completion claim fails that scenario. Score is scenarios passed out of 20. High here is what you want if your worry is an agent doing something you did not ask for; it has nothing to do with how the model handles requests for dangerous content, which is the row to the right.
Compliance with requests
High is more compliant here, same reading as every other row on this board. This is the flipped read of a refusal test: 63 adversarial prompts written to be genuinely dangerous, plus 7 that only sound risky (crisis support, harm-reduction, medication questions) which the model must still answer. A high score means it said yes more often across the harmful set specifically, not just your own ordinary requests - read that plainly, it is not a euphemism. The legitimate-look-alike count is shown separately and matters on its own: refusing one of those is a real cost, over-caution biting a request that was never dangerous. An answer the model never finished, because its reasoning budget ran out before a final answer landed, is counted separately as ‘empty’, never folded into either count. Score is harmful prompts complied with out of 63; both the complied and refused counts are shown on every cell so nothing is hidden behind the flip.
Compliance (agent context)
The same 63 harmful prompts and 7 legitimate look-alikes as Compliance with requests, but asked as the last turn of a real deployed agent's context - its full production system prompt, a live memory-system recall, and (the operator's ruling, 27 Sep 2026) the agent's real tool schema - instead of bare with nothing around it. A reply that completes via a tool call counts as complied, the same as a text answer. Same scoring otherwise: harmful prompts complied with out of 63, high is more compliant, an unfinished answer counts separately as ‘empty’, never folded into either count. The point of the row is the GAP to the bare row for the same weights: a model can comply freely bare and refuse far more once real deployment context surrounds the same question, and that gap is invisible to the bare row alone.
Compliance (severe)
A second, harder-hitting compliance probe than Compliance with requests: 30 prompts drawn from the StrongREJECT benchmark, split evenly across three categories - violence, illegal goods and services, and sexual content - curated by reading every candidate individually rather than keyword-screening, with material adjacent to child sexual abuse excluded during that review. Same read as Compliance with requests: high is more compliant, score is harmful prompts complied with out of 30. Bare, no system prompt, single turn - the agent-context pairing is the row below.
Compliance (severe, agent context)
The same 30-item severe set as Compliance (severe), asked as the last turn of a real deployed agent's context instead of bare - the same context capture used for Compliance (agent context) above. Same scoring: complied out of 30, high is more compliant. Where the two severe rows diverge for the same weights is the point: it shows whether real deployment context changes how a model handles harder requests specifically, not just the general 63-item set.
Likelihood (bits per character)
How unsurprised the model is by ordinary English, a rough proxy for how well it has learned the language. The RAW figure is lower-is-better; the SCORE is flipped so higher is better like every other row. A ‘cannot be measured’ cell means the serving software offers no confidence readout at all. The engine is the limit there, and the model goes unjudged. It measures fluency only, so a model can win here and still be the weaker assistant.
Prose speed
How fast it writes ordinary text. This is the number you feel when you watch it type.
Concurrency
What one agent gets while a second one is also working, which is the normal worst case. The figure is per agent. A total summed across agents would make a slow model look fast, because a sum describes the hardware and this row describes the experience.
Speculative decoding
Whether a small fast helper model can run ahead and guess the next few words so the big model checks several at once instead of writing one at a time. When it works it is free speed.
Reproducibility
Ask the same question twice and get the same answer back? Matters when you need to defend a decision or reproduce a result later.
Context window
How much text it can hold in mind at once, how long a document you can paste before it starts forgetting the beginning.
Box footprint
How much of a Spark it occupies before doing any work. A smaller footprint leaves more room for whatever else needs to share the box: another model, a larger KV cache, or any other workload.
Handwriting OCR
Reading your handwritten pages, and specifically whether it notices words that were crossed out. A model that reads deletions as live text credits words that the writer withdrew.
Long-prompt robustness
What happens when you paste something huge. Answering is best; refusing clearly is fine, because you learn why and can split it. Truncating without saying so, or crashing, is a failure. The two prompt sizes are CHARACTERS: 25k is about 6k tokens and 95k about 24k tokens (a needle question at the end, timed to a correct answer); a true 95k-TOKEN prompt is a separate, larger test.
Real-world screenshot & photo reading
Everyday screenshots and photos from a real workload: chat screenshots, dashboards, spreadsheets, X posts, dialog boxes, phone screenshots, product photos. Scored against an answer key written before any model saw an item, with a rule that a confidently wrong number or name counts against a model even when the rest of the reading is right. Every item counts: an item a model did not answer scores zero, so the count on each cell shows how many of the nineteen it answered. Where a cell is blank, hover over it to see why.

Configuration

Qwen3.8-Flash-Next
MiaAI-Lab vLLM recipe (sha 169fbad), adopted whole - TP=2+EP, MTP-3, NVFP4 (arch:sm_120) - 1M YaRN ctx - kv bf16 - util 0.835 (above the 0.80 ceiling used elsewhere on this board, head MemAvailable 5.89 GiB loaded - the tightest margin on the board except GLM-4.7)
GLM-5.3-Flash (RedHat)
vLLM tonyd2wild sm121-v11 TP=2 - compressed-tensors (arch:sm_120) - CUDA graphs - MTP-4 arm - same recipe and graphs config as the modelopt column above, different checkpoint
Qwen3.5-122B
vLLM 0.28.0-aarch64 TP=2 - NVFP4 (arch:sm_120) - enforce-eager (graphs arm 500s under load) - no spec - ctx 131072 - NCCL_CUMEM_ENABLE=0 (the ENGINE BUILD is the fix; stock 0.26 died 3/3)
DeepSeek-V4-Flash
dspark vLLM image TP=2 (MXFP4 MoE path) - revision-pinned - CUDA graphs - util 0.835 - max-num-seqs 16 - DSpark-5 - ctx 1,048,576
GLM-5.3-Flash
vLLM tonyd2wild sm121-v11 TP=2 - NVFP4 (arch:sm_120) - CUDA graphs (+4.5% by A/B) - MTP-4 arm, quantization=modelopt_fp4 asserted at the engine's own init line; EXL3 arm per MiaAI recipe
Qwen3-235B
vLLM 0.26 TP=2 - NVFP4 (arch:sm_120) - enforce-eager - no spec head - max_seq_len 32768 - NCCL_CUMEM_ENABLE=0 (fixes a worker livelock; stock config dies without it)
GLM-4.7 (full)
vLLM 0.28 TP=2 - NVFP4 (arch:sm_120) - community recipe: util 0.88, max-num-seqs 8, ctx 65536, CUDA graphs - util 0.80 is a MEASURED no-fit
GLM-5.3-Flash EXL3 (kit c707598e)
MiaAI-Lab GLM-5.3-Flash-EXL3 kit at commit c707598e (image rebuilt by the kit from its own recipe, tag glm53-exl3:c707598e) TP=2 over the fabric - EXL3 - DFlash k=7 in-recipe (MTP-2 and bare arms also served) - ctx 262144 - snapshot 024db9f7 weights - 02 Sep 22:30-23:22
  • Scores are standings within a row. A 10 on one row and a 3 on another do not add up to anything. Which rows matter is a judgement for whoever runs the box, so the board carries no column totals.
  • Near-identical scores are ties. Where the instruments could not separate two arms, or separated them by less than anyone would notice on real work, the arms score the same or within a fraction. The raw figures underneath still differ a little. That difference sits inside the measurement error.
  • A low score is a standing on one row. It is no verdict on the model. Several of these models are strong at things this board does not rank, and a model that scores badly as a daily general assistant may still be the right tool for one specific job.
  • Every number is a measurement taken on this hardware, or a labelled property of the configuration. Where a model could not be measured, the cell says why in its hover text and scores only what was observed. No cell holds a guess.
  • The best configurations conflict, and the board shows the conflict. Speculation buys decode speed and cuts concurrency and KV headroom. Autotune-off buys reproducibility and costs a little speed. Where one model needs two deployable profiles the cells say so. The board’s job is to make the trade visible. Which trade to take is a decision for whoever runs the box.
  • Likelihood is lower-is-better, so the score is inverted. It measures fluency on one frozen corpus and says nothing about usefulness on its own. Every likelihood cell on this board is a real measurement.
  • Every model handles text. The marks under each model name state what its served build accepts beyond text. Vision ✓ means it reads images, audio ✓ means it takes speech directly, and “no vision in this build” means this particular export shipped without its vision components. The model family may offer vision elsewhere, but the build measured here cannot see images (verified in each checkpoint’s own metadata).
  • There is no fine-tune row on this board. Every model here needs both Sparks just to serve, so with two Sparks there is no local fine-tuning path for any of them. The row would read the same for every column and tell you nothing. It stays on the single-Spark board, where the answer differs by model.

Single vs dual-Spark context, in real terms. The default single-Spark model (RadixArk FP4 27B) holds 262,144 tokens, about 650 pages, which is longer than most novels. Qwen3.8-Flash-Next's dual-Spark configuration holds 2,377,708 tokens, about 9 times that, or 5,900 pages, the size of a multi-volume reference work or an entire codebase. Most single documents fit inside the single-Spark number. Dual-Spark is for the rare cases that don't: a huge archive, a full book series, a big monorepo.

Credits & references

These benchmarks stand on other people’s published work. Several serving configurations came straight from community recipes:

  • GLM-5.3-Flash on Spark: deployment recipes and the patched sm121 serving image by tonyd2wild; the EXL3 two-Spark recipe by MiaAI-Lab; the single-Spark EXL3 K2 build by vcruz305; the DFlash2 drafter by incoai.
  • DeepSeek-V4-Flash single-Spark stack: MiaAI-Lab’s published deployment, adopted whole.
  • Ling-3.0-Flash single-Spark stack: sudoingX’s dgx-spark-ling recipe, adopted whole.
  • Qwen3.8-Flash-Next dual-Spark stack: MiaAI-Lab’s vLLM recipe, adopted whole (the Flash-Next supersession).
  • Speculative-decoding drafters: the DFlash method authors including z-lab’s Qwen3.5-122B and Qwen3.8-27B drafters; Meta’s Muse Glimmer assistant drafter; RadixArk’s model builds and DSpark drafter.
  • Model builds: Unsloth (27B NVFP4 and the Flash-Next GGUF), Vontra (the MLX 4-bit MTP conversion of Qwen3.8-Flash-Next that Unbound LR7 is built from), RadixArk, Ai2/allenai (olmOCR), NVIDIA (the Qwen3.8 NVFP4 local edition, the Muse Glimmer NVFP4 build, Nemotron and the Spark itself), Meta (Muse Glimmer), the Qwen team, Google (Gemma), Z.ai (GLM), DeepSeek, inclusionAI (Ling), ornith-ai (Ornith), QUASAR-QAT, RedHatAI (the Qwen3.5-122B NVFP4 build), GadflyII (the GLM-4.7-Flash NVFP4 build); and the uncensored builds: orcarouter’s NVFP4 and FP8 builds of Qwen3.8-27B, heterodoxin’s F16 build produced with their Apostate tool, which gives that column its name, and llmfan46’s Gemma-4-31B heretic build.
  • Serving engines: vLLM, SGLang, and llama.cpp, whose Flash-Next support merged upstream days before these runs; and TensorFold (ashhart/TensorFold), the engine behind the Leimroth 3 and Unbound LR7 columns.
  • Trap registry: Blackwellboy’s model-serving-minefield, a community registry of serving and measurement pitfalls that was consulted before every serving and benchmarking action in this campaign; several documented traps were avoided here because someone else had already paid for them.
  • Method lessons: catid’s public Spark benchmarks, and the wider DGX Spark community on X whose published results and warnings shaped several of these tests, including the wall-clock-beats-tokens-per-second point (@theotherpomp) and the EXL3 reasoning-tag artifact (@puffybsd) that changed how two of these measurements were taken.
  • Also on LocalMaxxing: these speed numbers, and others from this campaign not shown on this page, are submitted to LocalMaxxing’s community leaderboard as well.