Model Customization

Model Tuning PC

PC builds for customizing smaller AI models and testing fine-tuning workflows without implying full model training from scratch.

CategoryModel CustomizationOptionsfor comparisonPricingreviewed

Best for

  • LoRA and QLoRA adapter training
  • Longer sustained workloads with stable cooling
  • More system RAM for datasets and tooling

Not ideal for

  • Overbuilt if you only want local chat
  • Full model training still needs much larger hardware
  • Costs more than inference-first systems

AI fit is a rough estimate; model/runtime/quantization/context affects results.

Model Tuning PC

Configuration: Radeon AI PRO 32GB Model Tuning Workstation

A flexible workstation for adapting smaller AI models with your own examples, datasets, and experiments.

GPU: ASUS Turbo Radeon AI PRO R9700 32GB

CPU: AMD Ryzen 7 9700X

RAM: 64 GB | Storage: 4000 GB

Recommended model class: Adapter-based model tuning, not full model training from scratch.

30B-class Q4 only with caveats

Fit assumes quantization, moderate context, and runtime validation; 70B is not a normal target without 48GB+ VRAM or large unified memory.

About €4,497

Planning estimate; exact total in the written quote.

Radeon AI PRO 32GB Storage Alternative

This keeps the main build's 32GB Radeon AI PRO and CPU performance, but gives the quote process a second exact 64GB EXPO kit and a second trackable 4TB TLC NVMe drive. The exact Linux/ROCm stack and target workload are reproduced before quote.

GPU: ASUS Turbo Radeon AI PRO R9700 32GB

CPU: AMD Ryzen 7 9700X

RAM: 64 GB | Storage: 4000 GB

Recommended model class: 7B-14B LoRA/QLoRA after exact ROCm workload validation

30B-class Q4 only with caveats

Fit assumes quantization, moderate context, and runtime validation; 70B is not a normal target without 48GB+ VRAM or large unified memory.

About €4,597

Planning estimate; exact total in the written quote.

Local AI examples

Good fit for private chatGood fit for coding helpGood fit for document summariesNot ideal for 70B+ models

Recommended model

Qwen3.6 27B

A strong dense model for 24GB to 32GB GPUs and higher-memory Apple systems.

It is the practical quality target when the machine has enough memory for a larger dense model.

Comfortable GPU fit

Fast

ollama run qwen3.6:27b

Likely good memory headroom for this quantized model at normal context sizes.

  • 64GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Qwen3.6 35B-A3B

Coding agents, repository analysis, and complex local assistant workflows

Practical quantized GPU fit

A capable MoE model that gives 32GB and 48GB workstations a meaningfully heavier coding target.

Expected experience: Usable

Practical only with Q4-style quantization and moderate context; larger context can require offload or a smaller model.

  • 64GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Qwen3.5 9B

Private chat, coding help, document summaries, and image questions

Comfortable GPU fit

A strong everyday model for 12GB-class GPUs and a practical coding pick when speed matters.

Expected experience: Fast

Likely good memory headroom for this quantized model at normal context sizes.

  • 64GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache

Llama 3.3 70B Instruct

High-end local chat experiments on 48GB GPUs or 96GB+ Apple systems

Not realistic here

A clear upper-limit example for checking whether a machine can attempt a 70B-class model.

Expected experience: Not recommended

This model needs roughly 51GB GPU memory at Q4_K_M with moderate context; this machine has about 32GB VRAM.

  • 64GB system/unified memory available
  • 32GB effective accelerator memory for model weights and cache
Expandable technical details

Assumptions

  • GPU VRAM assumption: 32GB from ASUS Turbo Radeon AI PRO R9700 32GB.
  • System RAM: 64GB.
  • Ratings include model weights, estimated KV cache, runtime overhead, and safety margin for one local model running at a time. Treat them as fit guidance, not a speed guarantee.
  • Profile page uses a representative listed build. Open a build detail page for exact component-level fit.
Qwen3.6 27B technical details

Family: Qwen3.6

Parameters: 27B

Structure: 27B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 17 GB

Default estimate: 23 GB @ 8,192 tokens

Weights / KV / runtime / margin: 17 GB / 2 GB / 1.5 GB / 2.5 GB

CPU/RAM fallback: Not recommended

VRAM: 24 GB minimum / 32 GB recommended

RAM: 48 GB minimum / 64 GB recommended

Current run mode: Comfortable GPU fit

Expected experience: Fast

Full GPU offload: Only when the memory estimate and context fit

Context warning: Repository-scale or very long document context can push a 24GB card beyond a comfortable fit.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K17 GB1 GB1.5 GB2.5 GB22 GB
8K17 GB2 GB1.5 GB2.5 GB23 GB
16K17 GB4 GB1.5 GB3 GB25.5 GB
32K17 GB8 GB1.5 GB3.5 GB30 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Qwen3.6 35B-A3B technical details

Family: Qwen3.6

Parameters: 35B

Structure: 35B total / 3B active MoE

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 24 GB

Default estimate: 31.5 GB @ 8,192 tokens

Weights / KV / runtime / margin: 24 GB / 2.5 GB / 2 GB / 3 GB

CPU/RAM fallback: Not recommended

VRAM: 32 GB minimum / 48 GB recommended

RAM: 64 GB minimum / 64 GB recommended

Current run mode: Practical quantized GPU fit

Expected experience: Usable

Full GPU offload: Only when the memory estimate and context fit

Context warning: Long repository context can use the remaining headroom quickly, especially on 32GB GPUs.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K24 GB1.5 GB2 GB3 GB30.5 GB
8K24 GB2.5 GB2 GB3 GB31.5 GB
16K24 GB5 GB2 GB3.5 GB34.5 GB
32K24 GB10 GB2 GB4 GB40 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Qwen3.5 9B technical details

Family: Qwen3.5

Parameters: 9B

Structure: 9B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 6.6 GB

Default estimate: 11 GB @ 8,192 tokens

Weights / KV / runtime / margin: 6.6 GB / 1.5 GB / 1.2 GB / 1.5 GB

CPU/RAM fallback: Not recommended

VRAM: 8 GB minimum / 12 GB recommended

RAM: 16 GB minimum / 32 GB recommended

Current run mode: Comfortable GPU fit

Expected experience: Fast

Full GPU offload: Only when the memory estimate and context fit

Context warning: Start around 8K context even though the model supports much more.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K6.6 GB1 GB1.2 GB1.5 GB10.5 GB
8K6.6 GB1.5 GB1.2 GB1.5 GB11 GB
16K6.6 GB2.5 GB1.2 GB1.5 GB12 GB
32K6.6 GB5 GB1.2 GB2 GB15 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Llama 3.3 70B Instruct technical details

Family: Meta Llama 3.3

Parameters: 70.60B

Structure: 70.6B dense

License: Llama 3.3 Community License; allowed with terms; gated access. For guidance only; review the model licence before commercial use.

Native context: 131,072 tokens

Quantization: Q4_K_M

Approx. Q4 weights: 42.5 GB

Default estimate: 51 GB @ 4,096 tokens

Weights / KV / runtime / margin: 42.5 GB / 3 GB / 1.5 GB / 4 GB

CPU/RAM fallback: Not recommended

VRAM: 48 GB minimum / 64 GB recommended

RAM: 64 GB minimum / 96 GB recommended

Current run mode: Not realistic here

Expected experience: Not recommended

Full GPU offload: Often limited

Context warning: The Q4 weights alone use about 42.5GB before KV cache, runtime overhead, and desktop headroom.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K42.5 GB3 GB1.5 GB4 GB51 GB
8K42.5 GB6 GB1.5 GB4 GB54 GB
16K42.5 GB12 GB1.5 GB4.5 GB60.5 GB
32K42.5 GB24 GB1.5 GB5.5 GB73.5 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Local AI performance is approximate. Results depend on quantization, context length, backend, drivers, and whether the model plus KV cache fits in VRAM or Apple unified memory.