How close are our estimates?

We set every estimate against a real measurement whenever we have one. This page shows each pair, how far off the estimate was, and whether we have enough pairs to adjust the estimates.

Status

All of our calibration so far comes from one machine: Apple MacBook Pro 16-inch (2021, M1 Max). Other chips and other makers may behave differently, so we keep the speed ranges wide and treat every speed as rough.

  • Dense models: 1 measured pair, 3 needed. Not enough measurements yet, so the default efficiency of 0.50 stays.
  • Mixture-of-experts models: 2 measured pairs, 3 needed. Not enough measurements yet, so the default efficiency of 0.30 stays.

The speed range runs from 0.70 to 1.30 times the middle estimate.

Measured against estimated

Speed is writing speed in tokens per second. Estimated is the middle of the estimate range, made with the default efficiency. A Spot row is a hand measurement; its method is in the last column. The scorecard row is the mean of the test jobs that ran.
LaptopMemory (GB)ModelKindSourceMeasuredEstimatedEstimate wasMethod
Apple MacBook Pro 16-inch (2021, M1 Max) 64 gemma3:27b Dense Spot 12 11.8 2% too low spot measurement, 3 runs after warm-up, ollama chat, short prompt
Apple MacBook Pro 16-inch (2021, M1 Max) Scorecard 64 qwen3:30b Mixture of experts Scorecard 53.7 58.4 9% too high scorecard: mean write speed over the tasks that ran
Apple MacBook Pro 16-inch (2021, M1 Max) 64 qwen3:30b Mixture of experts Spot 70.1 58.4 17% too low spot measurement, 3 runs after warm-up, ollama chat, short prompt

The numbers in use

  • Efficiency: 0.5 of memory bandwidth for dense models, 0.3 for mixture-of-experts models.
  • Speed range: 0.70 to 1.30 times the middle estimate.
  • Memory share the GPU may use on unified memory: 0.7.

The formulas are on How the estimates work. The full dataset is in finder.json, licensed CC BY 4.0.