AI system category

Experimental / Advanced

Mac mini plus external GPU paths for special CUDA compute cases. For AI compute only — not gaming or macOS graphics acceleration.

CategoryExperimental / AdvancedOptionsfor comparisonPricingreviewed

Important warning

External GPUs on Apple Silicon Macs are for AI compute workflows only. They do not accelerate macOS graphics, gaming, displays, Final Cut, or Blender viewport rendering.

Apple's official eGPU support is Intel-Mac-only. Apple Silicon support depends on TinyGPU/tinygrad-style AI compute drivers.

Experimental

Mac + eGPU

Configuration: Mac mini M4 Pro + Radeon AI PRO R9700 eGPU Experiment

Mac: Mac mini M4 Pro 24GB / 512GB (24 GB unified)

eGPU: Sonnet Breakaway Box 850T5 Thunderbolt 5 eGPU

GPU: ASUS Turbo Radeon AI PRO R9700 32GB (32 GB VRAM, RDNA 4)

Mainly for 7B/8B models

Safe starting point for chat and coding assistants; larger models need more VRAM or Apple unified memory.

Supported:

  • One explicitly qualified TinyGPU or tinygrad compute experiment using the external Radeon AI PRO R9700; normal Mac workloads continue to use the M4 Pro and MLX

Not supported:

  • Official Apple Silicon eGPU support
  • macOS graphics/display/gaming acceleration
  • Final Cut acceleration
  • native CUDA
  • general ROCm support on macOS
  • or production use without exact-stack qualification

UNQUALIFIED EXPERIMENTAL PATH: Apple does not support external GPUs on Apple Silicon, and Sonnet supports GPU cards in this enclosure on Windows rather than M-series Macs. Public component prices are planning references only; no hardware is quoted until the exact Mac, macOS, Sonnet GPU-850T5 enclosure, Radeon AI PRO R9700, TinyGPU/tinygrad version, PCIe mapping, compiler path, and target workload are reproduced successfully.

Mac: €1,948

Enclosure: €806

GPU: €1,733

The exact ASUS 90YV0MN0-M0NM00 card (266.7 x 111.1 x 40mm, 300W) is inside the Sonnet GPU-850T5 enclosure's published 345 x 155 x 70mm, triple-slot envelope. Physical and power fit do not establish Apple Silicon software support.

Experimental

Mac mini M4 Pro 48GB + Radeon AI PRO R9700 eGPU Experiment

Mac: Mac mini M4 Pro 48GB / 1TB (48 GB unified)

eGPU: Sonnet Breakaway Box 850T5 Thunderbolt 5 eGPU

GPU: ASUS Turbo Radeon AI PRO R9700 32GB (32 GB VRAM, RDNA 4)

Mainly for 7B/8B models

Safe starting point for chat and coding assistants; larger models need more VRAM or Apple unified memory.

Supported:

  • The same explicitly qualified TinyGPU or tinygrad compute experiment as the main eGPU build
  • with more host-side unified memory and storage for normal MLX
  • development
  • and dataset work

Not supported:

  • Official Apple Silicon eGPU support
  • macOS graphics/display/gaming acceleration
  • Final Cut acceleration
  • native CUDA
  • general ROCm support on macOS
  • or production use without exact-stack qualification

UNQUALIFIED EXPERIMENTAL PATH: the larger Mac configuration does not make Apple Silicon eGPU support official. No hardware is quoted until the exact 48GB Mac, macOS, Sonnet GPU-850T5 enclosure, Radeon AI PRO R9700, TinyGPU/tinygrad version, PCIe mapping, compiler path, and target workload are reproduced successfully.

Mac: €2,753

Enclosure: €806

GPU: €1,733

The external 32GB GPU memory and the Mac's 48GB unified memory are separate runtime pools. The exact ASUS card remains inside the Sonnet enclosure's published physical and power limits; that does not establish macOS software support.

Local AI examples

Entry local models onlyNot ideal for 70B+ models

Recommended model

Gemma 4 12B

A current multimodal everyday model that makes good use of a 16GB GPU.

It is the strongest comfortable general-purpose starting point for most 16GB builds.

Partial GPU offload only

Very slow

ollama run gemma4:12b-it-qat

Mac + eGPU runtime support is experimental; treat this as a validation target, not a normal recommendation. Quantized GPU memory fit looks reasonable, but longer context can still add pressure.

  • 24GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Qwen3.5 9B

Private chat, coding help, document summaries, and image questions

Partial GPU offload only

A strong everyday model for 12GB-class GPUs and a practical coding pick when speed matters.

Expected experience: Very slow

Mac + eGPU runtime support is experimental; treat this as a validation target, not a normal recommendation. Quantized GPU memory fit looks reasonable, but longer context can still add pressure.

  • 24GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Qwen3.6 27B

Serious coding help, complex reasoning, and longer document analysis

Not realistic here

A strong dense model for 24GB to 32GB GPUs and higher-memory Apple systems.

Expected experience: Not recommended

Needs at least 48GB system RAM; this machine reports 24GB.

  • 24GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Qwen3.6 35B-A3B

Coding agents, repository analysis, and complex local assistant workflows

Not realistic here

A capable MoE model that gives 32GB and 48GB workstations a meaningfully heavier coding target.

Expected experience: Not recommended

Needs at least 64GB system RAM; this machine reports 24GB.

  • 24GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache
Expandable technical details

Assumptions

  • GPU VRAM assumption: 32GB from ASUS Turbo Radeon AI PRO R9700 32GB.
  • System RAM: 24GB.
  • Mac + eGPU fit is experimental and depends on driver/runtime support, not just VRAM.
  • Ratings include model weights, estimated KV cache, runtime overhead, and safety margin for one local model running at a time. Treat them as fit guidance, not a speed guarantee.
  • Mac + eGPU examples are experimental. Individual setup pages must be checked before treating a model as practical.
Gemma 4 12B technical details

Family: Google Gemma 4

Parameters: 11.95B

Structure: 11.95B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Quantization: Q4_0 QAT

Approx. Q4 weights: 7.2 GB

Default estimate: 11.5 GB @ 8,192 tokens

Weights / KV / runtime / margin: 7.2 GB / 1.5 GB / 1.2 GB / 1.5 GB

CPU/RAM fallback: Not recommended

VRAM: 12 GB minimum / 16 GB recommended

RAM: 24 GB minimum / 32 GB recommended

Current run mode: Partial GPU offload only

Expected experience: Very slow

Full GPU offload: Only when the memory estimate and context fit

Context warning: Its large context window still requires substantial KV-cache headroom.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K7.2 GB1 GB1.2 GB1.5 GB11 GB
8K7.2 GB1.5 GB1.2 GB1.5 GB11.5 GB
16K7.2 GB2.5 GB1.2 GB1.5 GB12.5 GB
32K7.2 GB5 GB1.2 GB2 GB15.5 GB

Swipe the table sideways to inspect every memory estimate.

Qwen3.5 9B technical details

Family: Qwen3.5

Parameters: 9B

Structure: 9B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 6.6 GB

Default estimate: 11 GB @ 8,192 tokens

Weights / KV / runtime / margin: 6.6 GB / 1.5 GB / 1.2 GB / 1.5 GB

CPU/RAM fallback: Not recommended

VRAM: 8 GB minimum / 12 GB recommended

RAM: 16 GB minimum / 32 GB recommended

Current run mode: Partial GPU offload only

Expected experience: Very slow

Full GPU offload: Only when the memory estimate and context fit

Context warning: Start around 8K context even though the model supports much more.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K6.6 GB1 GB1.2 GB1.5 GB10.5 GB
8K6.6 GB1.5 GB1.2 GB1.5 GB11 GB
16K6.6 GB2.5 GB1.2 GB1.5 GB12 GB
32K6.6 GB5 GB1.2 GB2 GB15 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Qwen3.6 27B technical details

Family: Qwen3.6

Parameters: 27B

Structure: 27B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 17 GB

Default estimate: 23 GB @ 8,192 tokens

Weights / KV / runtime / margin: 17 GB / 2 GB / 1.5 GB / 2.5 GB

CPU/RAM fallback: Not recommended

VRAM: 24 GB minimum / 32 GB recommended

RAM: 48 GB minimum / 64 GB recommended

Current run mode: Not realistic here

Expected experience: Not recommended

Full GPU offload: Only when the memory estimate and context fit

Context warning: Repository-scale or very long document context can push a 24GB card beyond a comfortable fit.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K17 GB1 GB1.5 GB2.5 GB22 GB
8K17 GB2 GB1.5 GB2.5 GB23 GB
16K17 GB4 GB1.5 GB3 GB25.5 GB
32K17 GB8 GB1.5 GB3.5 GB30 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Qwen3.6 35B-A3B technical details

Family: Qwen3.6

Parameters: 35B

Structure: 35B total / 3B active MoE

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 24 GB

Default estimate: 31.5 GB @ 8,192 tokens

Weights / KV / runtime / margin: 24 GB / 2.5 GB / 2 GB / 3 GB

CPU/RAM fallback: Not recommended

VRAM: 32 GB minimum / 48 GB recommended

RAM: 64 GB minimum / 64 GB recommended

Current run mode: Not realistic here

Expected experience: Not recommended

Full GPU offload: Only when the memory estimate and context fit

Context warning: Long repository context can use the remaining headroom quickly, especially on 32GB GPUs.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K24 GB1.5 GB2 GB3 GB30.5 GB
8K24 GB2.5 GB2 GB3 GB31.5 GB
16K24 GB5 GB2 GB3.5 GB34.5 GB
32K24 GB10 GB2 GB4 GB40 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Local AI performance is approximate. Results depend on quantization, context length, backend, drivers, and whether the model plus KV cache fits in VRAM or Apple unified memory.