ASUS ProArt P16 (H7606): which AI models can it run?
Estimated from the specifications below. Choose the memory size you have or plan to buy, then read the table for that size.
Specifications
- Brand
- ASUS
- Chip
- AMD Ryzen AI 9 HX 370
- Processor
- AMD Ryzen AI 9 HX 370 (12 cores)
- Graphics
- NVIDIA GeForce RTX 5090 Laptop GPU
- Memory type
- LPDDR5X onboard
- Memory options
- 32 GB, 64 GB
- Graphics memory
- 24 GB
- Memory bandwidth
- 896 GB/s
Sources: https://www.asus.com/me-en/laptops/for-creators/proart/proart-p16-h7606/techspec/, https://www.nvidia.com/en-us/geforce/laptops/50-series/
GPU options on the spec page include RTX 5090 (24GB GDDR7), RTX 5070 (8GB) and RTX 4070 or 4060. Memory options 32GB and 64GB are listed for all configurations. GPU bandwidth from NVIDIA's RTX 50 laptop page.
With 32 GB of memory
About 24 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 157 to 291 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 125 to 233 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 95 to 176 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Fits | 7.8 | 71 to 132 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b |
Fits | 8.3 | 64 to 119 | Fast | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first |
Fits | 8.9 | 60 to 112 | Fast | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b |
Fits | 15.9 | 39 to 72 | Fast | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first |
Fits | 13.7 | 35 to 65 | Fast | not measured yet | Estimated |
| Phi-4 14B phi4:14b |
Fits | 13.9 | 34 to 64 | Fast | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first |
Fits | 13.4 | 34 to 63 | Fast | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first |
Fits | 16.5 | 78 to 145 | Fast | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 27.2 | 2.8 to 5.1 | Unclear could be Slow or Moderate | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b · thinks first |
Tight | 22.6 | 92 to 170 | Fast | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 26.3 | 2.4 to 4.4 | Slow | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 51.5 | 1.1 to 2 | Slow | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b |
Won't fit | 70.5 | not known | n/a | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
With 64 GB of memory
About 24 GB of it can hold a model.
| Model | Fit | Needs (GB) | Writing speed (tokens per second) | Speed | Quality | Basis |
|---|---|---|---|---|---|---|
| Llama 3.2 3B llama3.2:3b |
Fits | 5 | 157 to 291 | Fast | not measured yet | Estimated |
| Qwen3 4B qwen3:4b · thinks first |
Fits | 6 | 125 to 233 | Fast | not measured yet | Estimated |
| Gemma 3 4B gemma3:4b |
Fits | 6.7 | 95 to 176 | Fast | not measured yet | Estimated |
| Mistral 7B v0.3 mistral:7b |
Fits | 7.8 | 71 to 132 | Fast | not measured yet | Estimated |
| Llama 3.1 8B llama3.1:8b |
Fits | 8.3 | 64 to 119 | Fast | not measured yet | Estimated |
| Qwen3 8B qwen3:8b · thinks first |
Fits | 8.9 | 60 to 112 | Fast | not measured yet | Estimated |
| Gemma 3 12B gemma3:12b |
Fits | 15.9 | 39 to 72 | Fast | not measured yet | Estimated |
| DeepSeek-R1 Distill Qwen 14B deepseek-r1:14b · thinks first |
Fits | 13.7 | 35 to 65 | Fast | not measured yet | Estimated |
| Phi-4 14B phi4:14b |
Fits | 13.9 | 34 to 64 | Fast | not measured yet | Estimated |
| Qwen3 14B qwen3:14b · thinks first |
Fits | 13.4 | 34 to 63 | Fast | not measured yet | Estimated |
| gpt-oss 20B gpt-oss:20b · thinks first |
Fits | 16.5 | 78 to 145 | Fast | not measured yet | Estimated |
| Gemma 3 27B gemma3:27b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 27.2 | 2.8 to 5.1 | Unclear could be Slow or Moderate | not measured yet | Estimated |
| Qwen3 30B (A3B) qwen3:30b · thinks first |
Tight | 22.6 | 92 to 170 | Fast | 9.3/10 | Estimated |
| Qwen3 32B qwen3:32b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 26.3 | 2.4 to 4.4 | Slow | not measured yet | Estimated |
| Llama 3.1 70B llama3.1:70b Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 51.5 | 1.1 to 2 | Slow | not measured yet | Estimated |
| gpt-oss 120B gpt-oss:120b · thinks first Part of the model runs from system memory, which is much slower. |
Partial offload (slow) | 70.5 | 10 to 18 | Unclear could be Moderate or Fast | not measured yet | Estimated |
Models marked "thinks first" can write out their reasoning before they answer, so a job takes longer than the speed alone suggests.
Speed is the speed of writing an answer, shown as a range. Reading speed is not estimated. How the estimates work · How close they are · Back to the finder