Thirty-nine model configurations · two DGX Sparks · every number from a measurement file
Thirty-nine model configurations, each tuned to its own best engine, quantisation and speculative-decoding setup, then measured on DGX Spark hardware, one box and two. Every number traces back to a raw measurement file. When a config underperforms or the server dies, that goes on the board as a result.
Recipe swap
+62%
Qwen3.8-Flash-Next on two Sparks: 25.2 to 40.9 tok/s, single stream, same checkpoint. See the row →
Caught before publish
2.6% bpc
The 2.6% gap was two different quantizations. Reading the engine's own resolved config settled it. See the row →
The unstated trade
5.9GiB free
Head-node memory left with the model loaded and idle, measured on the box itself. See the row →
Two models, one number
67.5% tie
Their transcripts differ by several words each. The shared score is a floor in the ground truth. See the row →
How this is measured
Every model here runs on whatever combination of serving engine, quantisation and speculative-decoding setup gets the most out of it, so the engines and flags differ from column to column by design. Scores are standings within a row. A 10 is the best result measured on that one test, and the same model can score 3 on the next row. No number comes from a vendor spec sheet. Each one traces back to a raw measurement file, and a model that failed to serve stays on the board with its failure recorded.
Board one
Local models benchmarked on a single NVIDIA DGX Spark (GB10, 121.6 GiB unified memory). Click any row heading to reorder the models by what matters to you, or remove a model to compare just the ones you care about. Open Single-Spark →
Board two
Tensor parallelism splits one model across both Sparks. That admits models too large for one box and changes the economics of everything else measured. Same rows, same instruments, same rules. Open Dual-Spark →