SCOTT LEIMROTH ← AI & Tech
GB10DGX Spark Benchmarks

Thirty-nine model configurations · two DGX Sparks · every number from a measurement file

Which local model earns its configuration.

Thirty-nine model configurations, each tuned to its own best engine, quantisation and speculative-decoding setup, then measured on DGX Spark hardware, one box and two. Every number traces back to a raw measurement file. When a config underperforms or the server dies, that goes on the board as a result.

Recipe swap

+62%

Adopting a community vLLM recipe whole beat the original build

Qwen3.8-Flash-Next on two Sparks: 25.2 to 40.9 tok/s, single stream, same checkpoint. See the row →

Caught before publish

2.6% bpc

A directory name mislabeled a checkpoint. The measurement caught it.

The 2.6% gap was two different quantizations. Reading the engine's own resolved config settled it. See the row →

The unstated trade

5.9GiB free

A build framed as efficient turned out to have the tightest memory margin

Head-node memory left with the model loaded and idle, measured on the box itself. See the row →

Two models, one number

67.5% tie

Two unrelated models landed on the same OCR score for different reasons

Their transcripts differ by several words each. The shared score is a floor in the ground truth. See the row →

How this is measured

Each model at its own best

Every model here runs on whatever combination of serving engine, quantisation and speculative-decoding setup gets the most out of it, so the engines and flags differ from column to column by design. Scores are standings within a row. A 10 is the best result measured on that one test, and the same model can score 3 on the next row. No number comes from a vendor spec sheet. Each one traces back to a raw measurement file, and a model that failed to serve stays on the board with its failure recorded.

Board one

Single-Spark: thirty-one configurations, one box each

Local models benchmarked on a single NVIDIA DGX Spark (GB10, 121.6 GiB unified memory). Click any row heading to reorder the models by what matters to you, or remove a model to compare just the ones you care about. Open Single-Spark →

Board two

Dual-Spark (TP=2): eight configs, two boxes serving one model

Tensor parallelism splits one model across both Sparks. That admits models too large for one box and changes the economics of everything else measured. Same rows, same instruments, same rules. Open Dual-Spark →