Apple MacBook Pro 16-inch (2026, M5 Max 40-core GPU): which AI models can it run?

Estimated from the specifications below. Choose the memory size you have or plan to buy, then read the table for that size.

Specifications

Brand
Apple
Year
2026
Chip
Apple M5 Max
Processor
18-core CPU (6 super, 12 performance)
Graphics
40-core GPU
Memory type
unified memory
Memory options
128 GB
Memory bandwidth
614 GB/s

Sources: https://support.apple.com/en-us/126319, https://www.apple.com/newsroom/2026/03/apple-introduces-macbook-pro-with-all-new-m5-pro-and-m5-max/

Apple lists 614GB/s for the 40-core GPU variant and up to 128GB for M5 Max. Smaller options are not listed here because the Apple pages conflict.

With 128 GB of memory

About 89.6 GB of it can hold a model.

ModelFitNeeds (GB)Writing speed (tokens per second)SpeedQualityBasis
Llama 3.2 3B
llama3.2:3b
Fits 5 107 to 200 Fast not measured yet Estimated
Qwen3 4B
qwen3:4b · thinks first
Fits 6 86 to 160 Fast not measured yet Estimated
Gemma 3 4B
gemma3:4b
Fits 6.7 65 to 121 Fast not measured yet Estimated
Mistral 7B v0.3
mistral:7b
Fits 7.8 49 to 91 Fast not measured yet Estimated
Llama 3.1 8B
llama3.1:8b
Fits 8.3 44 to 81 Fast not measured yet Estimated
Qwen3 8B
qwen3:8b · thinks first
Fits 8.9 41 to 77 Fast not measured yet Estimated
Gemma 3 12B
gemma3:12b
Fits 15.9 27 to 49 Fast not measured yet Estimated
DeepSeek-R1 Distill Qwen 14B
deepseek-r1:14b · thinks first
Fits 13.7 24 to 44 Fast not measured yet Estimated
Phi-4 14B
phi4:14b
Fits 13.9 24 to 44 Fast not measured yet Estimated
Qwen3 14B
qwen3:14b · thinks first
Fits 13.4 23 to 43 Fast not measured yet Estimated
gpt-oss 20B
gpt-oss:20b · thinks first
Fits 16.5 53 to 99 Fast not measured yet Estimated
Gemma 3 27B
gemma3:27b
Fits 27.2 13 to 23 Fast not measured yet Estimated
Qwen3 30B (A3B)
qwen3:30b · thinks first
Fits 22.6 63 to 116 Fast 9.3/10 Estimated
Qwen3 32B
qwen3:32b · thinks first
Fits 26.3 11 to 20 Fast not measured yet Estimated
Llama 3.1 70B
llama3.1:70b
Fits 51.5 5 to 9.3 Unclear could be Slow or Moderate not measured yet Estimated
gpt-oss 120B
gpt-oss:120b · thinks first
Fits 70.5 46 to 85 Fast not measured yet Estimated

Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.

Speed is the speed of writing an answer, shown as a range. Reading speed is not estimated. How the estimates work · How close they are · Back to the finder