Mac + eGPU specialist setup

Mac mini M4 Pro 48GB + Radeon AI PRO R9700 eGPU Experiment

The external 32GB GPU memory and the Mac's 48GB unified memory are separate runtime pools. The exact ASUS card remains inside the Sonnet enclosure's published physical and power limits; that does not establish macOS software support.

experimental eGPU path

Is this build right for me?

This is a specialist Mac + external-GPU compute path, not a normal Mac upgrade. The 32 GB GPU has AI potential, but Apple Silicon has no official GPU-card eGPU support, so the exact hardware, software environment, and workload must be reproduced before it is relied on.

AI strength score

The native Mac path sets this score; external-GPU compute remains a separate validation-only option.

Value

Weak local-AI capability for the current estimated price.

What this build can handle

Comfortable

  • Smaller native MLX or Ollama models that fit the Mac's shared memory.
  • Planning and reproducing the exact eGPU software stack before relying on it.

Possible with limits

  • 32 GB external-GPU inference after the exact driver and software environment are validated.
  • A monitored experimental workload after it is reproduced on this exact hardware.

Choose this build if you have a specific external-GPU research workload and will validate it before buying.

Estimated parts price history

The chart combines the current estimated baseline with verified market prices. New checks automatically replace estimated points for the matching period.

Latest estimated parts total

€5,292

Final price has service fees included.

Component total

Time range

Tap a point to see its price. Swipe the chart sideways to inspect every date.

The chart slider reports the component total by date.

Estimated parts price historyThe line shows the estimated total for all components using both baseline estimates and available checked market prices. Values for individual dates remain available through the chart slider and data table.€4,762€5,001€5,239€5,478€5,7167 May18 Jun2 Aug

Component total by date
DateComponent totalCheckedEstimatedMac systemeGPU enclosureGPU
7 May€5,29203€2,753€806€1,733
10 May€5,29203€2,753€806€1,733
13 May€5,29203€2,753€806€1,733
16 May€5,29203€2,753€806€1,733
19 May€5,29203€2,753€806€1,733
22 May€5,29203€2,753€806€1,733
25 May€5,29203€2,753€806€1,733
28 May€5,29203€2,753€806€1,733
31 May€5,29203€2,753€806€1,733
3 Jun€5,29203€2,753€806€1,733
6 Jun€5,29203€2,753€806€1,733
9 Jun€5,29203€2,753€806€1,733
12 Jun€5,29203€2,753€806€1,733
15 Jun€5,29203€2,753€806€1,733
18 Jun€5,29203€2,753€806€1,733
21 Jun€5,29203€2,753€806€1,733
24 Jun€5,29203€2,753€806€1,733
27 Jun€5,29203€2,753€806€1,733
30 Jun€5,29203€2,753€806€1,733
3 Jul€5,29203€2,753€806€1,733
6 Jul€5,29203€2,753€806€1,733
9 Jul€5,29203€2,753€806€1,733
12 Jul€5,29203€2,753€806€1,733
15 Jul€5,29203€2,753€806€1,733
18 Jul€5,29203€2,753€806€1,733
21 Jul€5,29203€2,753€806€1,733
24 Jul€5,29203€2,753€806€1,733
27 Jul€5,29203€2,753€806€1,733
30 Jul€5,29230€2,753€806€1,733
2 Aug€5,29203€2,753€806€1,733

Core Configuration

Mac System

Mac mini M4 Pro 48GB / 1TB

Chip

Apple M4 Pro

Unified Memory

48 GB

Storage

1000 GB

GPU

ASUS Turbo Radeon AI PRO R9700 32GB

VRAM

32 GB

Performance & Power

eGPU Enclosure

Sonnet Breakaway Box 850T5 Thunderbolt 5 eGPU

Architecture

RDNA 4

Risk level

Experimental

Workload fit

Mainly for 7B/8B models

Mainly for 7B/8B models

Safe starting point for chat and coding assistants; larger models need more VRAM or Apple unified memory.

Mac unified memory48 GBNative Mac path
External GPU VRAM32 GBQualified eGPU path

Qualification protocol

PRE-QUOTE
  1. 01Define one target workload
  2. 02Confirm enclosure and GPU path
  3. 03Reproduce driver and software stack
  4. 04Make the go/no-go decision after testing

Supported Workloads

  • The same explicitly qualified TinyGPU or tinygrad compute experiment as the main eGPU build
  • with more host-side unified memory and storage for normal MLX
  • development
  • and dataset work

Not Supported

  • Official Apple Silicon eGPU support
  • macOS graphics/display/gaming acceleration
  • Final Cut acceleration
  • native CUDA
  • general ROCm support on macOS
  • or production use without exact-stack qualification

Buyer Warning

UNQUALIFIED EXPERIMENTAL PATH: the larger Mac configuration does not make Apple Silicon eGPU support official. No hardware is quoted until the exact 48GB Mac, macOS, Sonnet GPU-850T5 enclosure, Radeon AI PRO R9700, TinyGPU/tinygrad version, PCIe mapping, compiler path, and target workload are reproduced successfully.

Component Pricing Breakdown

Component rows show planning prices for the Mac, eGPU enclosure, and GPU. The separate service fee uses the shared 20–10% scale across all published complete PC and Mac/eGPU builds; the written quote confirms final price and compatibility.

ComponentProductEstimated market price
Mac SystemMac mini M4 Pro 48GB / 1TB€2,753
eGPU EnclosureSonnet Breakaway Box 850T5 Thunderbolt 5 eGPU€806
GPUASUS Turbo Radeon AI PRO R9700 32GB€1,733
Estimated parts subtotal€5,292
Service, setup, and validation fee (17.02%)€901
Estimated total€6,193

Local AI examples

Examples for Mac mini M4 Pro 48GB + Radeon AI PRO R9700 eGPU Experiment, based on GPU VRAM or Apple unified memory plus RAM headroom. System RAM is not treated as VRAM.

Entry local models onlyNot ideal for 70B+ models

Recommended model

Qwen3.6 27B

A strong dense model for 24GB to 32GB GPUs and higher-memory Apple systems.

It is the practical quality target when the machine has enough memory for a larger dense model.

Partial GPU offload only

Very slow

ollama run qwen3.6:27b

Mac + eGPU runtime support is experimental; treat this as a validation target, not a normal recommendation. Quantized GPU memory fit looks reasonable, but longer context can still add pressure.

  • 48GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Gemma 4 12B

Private assistant chat, document understanding, and image or audio analysis

Partial GPU offload only

A current multimodal everyday model that makes good use of a 16GB GPU.

Expected experience: Very slow

Mac + eGPU runtime support is experimental; treat this as a validation target, not a normal recommendation. Likely good memory headroom for this quantized model at normal context sizes.

  • 48GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Qwen3.6 35B-A3B

Coding agents, repository analysis, and complex local assistant workflows

Not realistic here

A capable MoE model that gives 32GB and 48GB workstations a meaningfully heavier coding target.

Expected experience: Not recommended

Needs at least 64GB system RAM; this machine reports 48GB.

  • 48GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Llama 3.3 70B Instruct

High-end local chat experiments on 48GB GPUs or 96GB+ Apple systems

Not realistic here

A clear upper-limit example for checking whether a machine can attempt a 70B-class model.

Expected experience: Not recommended

Needs at least 64GB system RAM; this machine reports 48GB.

  • 48GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Qwen3.5 9B

Private chat, coding help, document summaries, and image questions

Partial GPU offload only

A strong everyday model for 12GB-class GPUs and a practical coding pick when speed matters.

Expected experience: Very slow

Mac + eGPU runtime support is experimental; treat this as a validation target, not a normal recommendation. Likely good memory headroom for this quantized model at normal context sizes.

  • 48GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache
Expandable technical details

Assumptions

  • GPU VRAM assumption: 32GB from ASUS Turbo Radeon AI PRO R9700 32GB.
  • System RAM: 48GB.
  • Mac + eGPU fit is experimental and depends on driver/runtime support, not just VRAM.
  • Ratings include model weights, estimated KV cache, runtime overhead, and safety margin for one local model running at a time. Treat them as fit guidance, not a speed guarantee.
  • This setup depends on experimental Mac + external GPU software support. Treat compatibility ratings as a starting point for the pre-quote consultation.
Qwen3.6 27B technical details

Family: Qwen3.6

Parameters: 27B

Structure: 27B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 17 GB

Default estimate: 23 GB @ 8,192 tokens

Weights / KV / runtime / margin: 17 GB / 2 GB / 1.5 GB / 2.5 GB

CPU/RAM fallback: Not recommended

VRAM: 24 GB minimum / 32 GB recommended

RAM: 48 GB minimum / 64 GB recommended

Current run mode: Partial GPU offload only

Expected experience: Very slow

Full GPU offload: Only when the memory estimate and context fit

Context warning: Repository-scale or very long document context can push a 24GB card beyond a comfortable fit.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K17 GB1 GB1.5 GB2.5 GB22 GB
8K17 GB2 GB1.5 GB2.5 GB23 GB
16K17 GB4 GB1.5 GB3 GB25.5 GB
32K17 GB8 GB1.5 GB3.5 GB30 GB

Research sources

Researched: 2026-07-30

Gemma 4 12B technical details

Family: Google Gemma 4

Parameters: 11.95B

Structure: 11.95B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Quantization: Q4_0 QAT

Approx. Q4 weights: 7.2 GB

Default estimate: 11.5 GB @ 8,192 tokens

Weights / KV / runtime / margin: 7.2 GB / 1.5 GB / 1.2 GB / 1.5 GB

CPU/RAM fallback: Not recommended

VRAM: 12 GB minimum / 16 GB recommended

RAM: 24 GB minimum / 32 GB recommended

Current run mode: Partial GPU offload only

Expected experience: Very slow

Full GPU offload: Only when the memory estimate and context fit

Context warning: Its large context window still requires substantial KV-cache headroom.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K7.2 GB1 GB1.2 GB1.5 GB11 GB
8K7.2 GB1.5 GB1.2 GB1.5 GB11.5 GB
16K7.2 GB2.5 GB1.2 GB1.5 GB12.5 GB
32K7.2 GB5 GB1.2 GB2 GB15.5 GB
Qwen3.6 35B-A3B technical details

Family: Qwen3.6

Parameters: 35B

Structure: 35B total / 3B active MoE

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 24 GB

Default estimate: 31.5 GB @ 8,192 tokens

Weights / KV / runtime / margin: 24 GB / 2.5 GB / 2 GB / 3 GB

CPU/RAM fallback: Not recommended

VRAM: 32 GB minimum / 48 GB recommended

RAM: 64 GB minimum / 64 GB recommended

Current run mode: Not realistic here

Expected experience: Not recommended

Full GPU offload: Only when the memory estimate and context fit

Context warning: Long repository context can use the remaining headroom quickly, especially on 32GB GPUs.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K24 GB1.5 GB2 GB3 GB30.5 GB
8K24 GB2.5 GB2 GB3 GB31.5 GB
16K24 GB5 GB2 GB3.5 GB34.5 GB
32K24 GB10 GB2 GB4 GB40 GB

Research sources

Researched: 2026-07-30

Llama 3.3 70B Instruct technical details

Family: Meta Llama 3.3

Parameters: 70.60B

Structure: 70.6B dense

License: Llama 3.3 Community License; allowed with terms; gated access. For guidance only; review the model licence before commercial use.

Native context: 131,072 tokens

Quantization: Q4_K_M

Approx. Q4 weights: 42.5 GB

Default estimate: 51 GB @ 4,096 tokens

Weights / KV / runtime / margin: 42.5 GB / 3 GB / 1.5 GB / 4 GB

CPU/RAM fallback: Not recommended

VRAM: 48 GB minimum / 64 GB recommended

RAM: 64 GB minimum / 96 GB recommended

Current run mode: Not realistic here

Expected experience: Not recommended

Full GPU offload: Often limited

Context warning: The Q4 weights alone use about 42.5GB before KV cache, runtime overhead, and desktop headroom.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K42.5 GB3 GB1.5 GB4 GB51 GB
8K42.5 GB6 GB1.5 GB4 GB54 GB
16K42.5 GB12 GB1.5 GB4.5 GB60.5 GB
32K42.5 GB24 GB1.5 GB5.5 GB73.5 GB

Research sources

Researched: 2026-07-30

Qwen3.5 9B technical details

Family: Qwen3.5

Parameters: 9B

Structure: 9B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 6.6 GB

Default estimate: 11 GB @ 8,192 tokens

Weights / KV / runtime / margin: 6.6 GB / 1.5 GB / 1.2 GB / 1.5 GB

CPU/RAM fallback: Not recommended

VRAM: 8 GB minimum / 12 GB recommended

RAM: 16 GB minimum / 32 GB recommended

Current run mode: Partial GPU offload only

Expected experience: Very slow

Full GPU offload: Only when the memory estimate and context fit

Context warning: Start around 8K context even though the model supports much more.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K6.6 GB1 GB1.2 GB1.5 GB10.5 GB
8K6.6 GB1.5 GB1.2 GB1.5 GB11 GB
16K6.6 GB2.5 GB1.2 GB1.5 GB12 GB
32K6.6 GB5 GB1.2 GB2 GB15 GB

Research sources

Researched: 2026-07-30

Local AI performance is approximate. Results depend on quantization, context length, backend, drivers, and whether the model plus KV cache fits in VRAM or Apple unified memory.

AI terms in plain language

VRAM

Memory on the graphics card; usually the main limit for local AI model size.

System RAM vs GPU memory

RAM helps apps and data work; GPU memory usually decides which model size can run quickly.

CUDA

NVIDIA’s software layer for many AI tools. Macs and AMD GPUs do not run CUDA workflows the same way.

Apple unified memory

Apple Silicon memory shared by CPU, GPU, macOS, and apps. Useful for Mac AI, but not the same as NVIDIA VRAM.

7B / 8B / 14B / 70B

Approximate model size in billions of parameters. Larger numbers usually need more memory and may run slower.

Quantized models

Compressed models, such as Q4, that use less memory with possible quality or speed tradeoffs.

Context length

How much text the model can keep in mind at once. Longer context uses more memory.

Inference

Running an existing model for chat, coding help, summaries, or document workflows.

Fine-tuning vs adapter tuning

Fine-tuning adapts a model; adapter tuning, such as LoRA/QLoRA, is a lighter way to steer an existing model with examples.

Dual-GPU limitations

Two GPUs do not automatically combine VRAM into one large pool. Software must explicitly support multiple GPUs.

eGPU limitations

External GPUs need enclosure, driver, and runtime validation, and usually do not mean macOS graphics or gaming acceleration.

Planning estimate / written quote

The displayed estimate helps with budgeting. A written quote contains the exact price, confirmed availability, and any proposed substitutions.

Request a quote

Planning estimate

About €6,193

The external 32GB GPU memory and the Mac's 48GB unified memory are separate runtime pools. The exact ASUS card remains inside the Sonnet enclosure's published physical and power limits; that does not establish macOS software support.

Mac + eGPU setups start with a consultation. Component prices are planning estimates; the written quote confirms price, GPU and enclosure availability, software compatibility, and suitability for your workload.