Local AI workstation

LLMLab Forge Ultra

A maximum-capacity single-GPU workstation with 96 GB ECC VRAM and 256 GB eight-channel ECC system memory.

Is this build right for me?

AI strength score

Headroom for large local models and demanding AI work.

Value

Weak local-AI capability for the current estimated price.

What this build can handle

Comfortable

  • Smaller models in large batches and many 70B-class quantized inference workloads after exact qualification.

Possible with limits

  • Selected larger or longer-context workloads when runtime, memory overhead, speed, and serving load are validated.

Not recommended

  • Assuming one GPU trains foundation models from scratch or that every multi-user production workload is automatically qualified.

Price estimate history

Uses market prices where available and reference estimates for the rest.

Reference estimate

€33,150

System estimate

Time range

Tap a point to see its price. Swipe the chart sideways to inspect every date.

The chart slider reports the estimated system total by date.

Price estimate historyThe line shows the estimated system total over time. Values for individual dates remain available through the chart slider and data table.€29,821€31,316€32,812€34,307€35,8021 Jul15 Aug29 Sept

System estimate by date
DateSystem estimateCPUGPURAMStorageMotherboardPSUCaseCooler
1 Jul€33,135€4,400€13,900€8,109€1,200€1,359€459€164€152
6 Jul€33,135€4,400€13,900€8,109€1,200€1,359€459€164€152
11 Jul€33,135€4,400€13,900€8,109€1,200€1,359€459€164€152
16 Jul€33,135€4,400€13,900€8,109€1,200€1,359€459€164€152
21 Jul€33,135€4,400€13,900€8,109€1,200€1,359€459€164€152
26 Jul€33,135€4,400€13,900€8,109€1,200€1,359€459€164€152
31 Jul€33,135€4,400€13,900€8,109€1,200€1,359€459€164€152
5 Aug€33,135€4,400€13,900€8,109€1,200€1,359€459€164€152
10 Aug€33,135€4,400€13,900€8,109€1,200€1,359€459€164€152
15 Aug€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152
20 Aug€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152
25 Aug€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152
30 Aug€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152
4 Sept€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152
9 Sept€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152
14 Sept€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152
19 Sept€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152
24 Sept€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152
29 Sept€33,150€4,400€13,900€8,109€1,200€1,359€459€178€152

Core Configuration

CPU

AMD Ryzen Threadripper PRO 9975WX

GPU

PNY NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96GB

VRAM

96 GB

RAM

256 GB

Storage

8000 GB

Recommended model class

70B-class quantized inference with long-context headroom after exact workload validation

Performance & Power

Throughput

Runtime qualification required; no fixed throughput promise

System power

~950 W

Recommended PSU

1600 W

Cooling

Full-IHS sTR5 workstation air cooling with a full-tower airflow path

Component Pricing Breakdown

Component rows show Estonian market or reference prices. Service, assembly, and configuration are shown separately below.

ComponentProductEstimated market price
CPUAMD Ryzen Threadripper PRO 9975WX€4,620
GPUPNY NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96GB€14,595
RAMKingston FURY Renegade Pro 256GB (8x32GB) DDR5-5600 CL28 ECC RDIMM€8,515
StorageWD Black SN850X 8TB€1,260
MotherboardASUS Pro WS WRX90E-SAGE SE€1,427
PSUbe quiet! Dark Power Pro 13 1600W€482
CasePhanteks Enthoo Pro 2 Tempered Glass Black€172
CoolerNoctua NH-U14S TR5-SP6€160
Estimated component total€31,230
Service / assembly / configuration fee€3,437
Estimated build total€34,667

Local AI examples

Good fit for private chatGood fit for coding helpGood fit for document summaries70B-class models need strict caveats

Recommended model

Qwen3.6 35B-A3B

A capable MoE model that gives 32GB and 48GB workstations a meaningfully heavier coding target.

Comfortable GPU fit

Usable

ollama run qwen3.6:35b

Likely good memory headroom for this quantized model at normal context sizes.

  • 256GB system/unified memory available
  • 96GB effective accelerator memory for model weights and cache

Qwen3.6 27B

Serious coding help, complex reasoning, and longer document analysis

Comfortable GPU fit

A strong dense model for 24GB to 32GB GPUs and higher-memory Apple systems.

Expected experience: Fast

Likely good memory headroom for this quantized model at normal context sizes.

  • 256GB system/unified memory available

Qwen3.5 9B

Private chat, coding help, document summaries, and image questions

Comfortable GPU fit

A strong everyday model for 12GB-class GPUs and a practical coding pick when speed matters.

Expected experience: Fast

Likely good memory headroom for this quantized model at normal context sizes.

  • 256GB system/unified memory available

Llama 3.3 70B Instruct

High-end local chat experiments on 48GB GPUs or 96GB+ Apple systems

Comfortable GPU fit

A clear upper-limit example for checking whether a machine can attempt a 70B-class model.

Expected experience: Usable

Likely good memory headroom for this quantized model at normal context sizes.

  • 256GB system/unified memory available
Expandable technical details

Assumptions

  • GPU VRAM assumption: 96GB from PNY NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96GB.
  • System RAM: 256GB.
  • Ratings include model weights, estimated KV cache, runtime overhead, and safety margin for one local model running at a time. Treat them as fit guidance, not a speed guarantee.
Qwen3.6 35B-A3B technical details

Family: Qwen3.6

Parameters: 35B

Structure: 35B total / 3B active MoE

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 24 GB

Default estimate: 31.5 GB @ 8,192 tokens

Weights / KV / runtime / margin: 24 GB / 2.5 GB / 2 GB / 3 GB

CPU/RAM fallback: Not recommended

VRAM: 32 GB minimum / 48 GB recommended

RAM: 64 GB minimum / 64 GB recommended

Current run mode: Comfortable GPU fit

Expected experience: Usable

Full GPU offload: Only when the memory estimate and context fit

Context warning: Long repository context can use the remaining headroom quickly, especially on 32GB GPUs.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K24 GB1.5 GB2 GB3 GB30.5 GB
8K24 GB2.5 GB2 GB3 GB31.5 GB
16K24 GB5 GB2 GB3.5 GB34.5 GB
32K24 GB10 GB2 GB4 GB40 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Qwen3.6 27B technical details

Family: Qwen3.6

Parameters: 27B

Structure: 27B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 17 GB

Default estimate: 23 GB @ 8,192 tokens

Weights / KV / runtime / margin: 17 GB / 2 GB / 1.5 GB / 2.5 GB

CPU/RAM fallback: Not recommended

VRAM: 24 GB minimum / 32 GB recommended

RAM: 48 GB minimum / 64 GB recommended

Current run mode: Comfortable GPU fit

Expected experience: Fast

Full GPU offload: Only when the memory estimate and context fit

Context warning: Repository-scale or very long document context can push a 24GB card beyond a comfortable fit.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K17 GB1 GB1.5 GB2.5 GB22 GB
8K17 GB2 GB1.5 GB2.5 GB23 GB
16K17 GB4 GB1.5 GB3 GB25.5 GB
32K17 GB8 GB1.5 GB3.5 GB30 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Qwen3.5 9B technical details

Family: Qwen3.5

Parameters: 9B

Structure: 9B dense

License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.

Native context: 262,144 tokens

Extended context: 1,010,000 via YaRN; not a default fit assumption

Quantization: Q4_K_M

Approx. Q4 weights: 6.6 GB

Default estimate: 11 GB @ 8,192 tokens

Weights / KV / runtime / margin: 6.6 GB / 1.5 GB / 1.2 GB / 1.5 GB

CPU/RAM fallback: Not recommended

VRAM: 8 GB minimum / 12 GB recommended

RAM: 16 GB minimum / 32 GB recommended

Current run mode: Comfortable GPU fit

Expected experience: Fast

Full GPU offload: Only when the memory estimate and context fit

Context warning: Start around 8K context even though the model supports much more.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K6.6 GB1 GB1.2 GB1.5 GB10.5 GB
8K6.6 GB1.5 GB1.2 GB1.5 GB11 GB
16K6.6 GB2.5 GB1.2 GB1.5 GB12 GB
32K6.6 GB5 GB1.2 GB2 GB15 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Llama 3.3 70B Instruct technical details

Family: Meta Llama 3.3

Parameters: 70.60B

Structure: 70.6B dense

License: Llama 3.3 Community License; allowed with terms; gated access. For guidance only; review the model licence before commercial use.

Native context: 131,072 tokens

Quantization: Q4_K_M

Approx. Q4 weights: 42.5 GB

Default estimate: 51 GB @ 4,096 tokens

Weights / KV / runtime / margin: 42.5 GB / 3 GB / 1.5 GB / 4 GB

CPU/RAM fallback: Not recommended

VRAM: 48 GB minimum / 64 GB recommended

RAM: 64 GB minimum / 96 GB recommended

Current run mode: Comfortable GPU fit

Expected experience: Usable

Full GPU offload: Often limited

Context warning: The Q4 weights alone use about 42.5GB before KV cache, runtime overhead, and desktop headroom.

ContextWeightsKVRuntimeMarginEstimated GPU memory
4K42.5 GB3 GB1.5 GB4 GB51 GB
8K42.5 GB6 GB1.5 GB4 GB54 GB
16K42.5 GB12 GB1.5 GB4.5 GB60.5 GB
32K42.5 GB24 GB1.5 GB5.5 GB73.5 GB

Swipe the table sideways to inspect every memory estimate.

Research sources

Researched: 2026-07-30

Local AI performance is approximate. Results depend on quantization, context length, backend, drivers, and whether the model plus KV cache fits in VRAM or Apple unified memory.

AI terms in plain language

GPU memory and RAM

GPU memory usually limits model size; RAM supports apps, data, and model offload.

CUDA

NVIDIA’s software layer for many AI tools. Macs and AMD GPUs do not run CUDA workflows the same way.

Apple unified memory

Memory shared by the CPU, GPU, macOS, and apps; it is not the same as NVIDIA VRAM.

7B / 8B / 14B / 70B

Approximate model size in billions of parameters; larger models usually need more memory.

Quantized models

Lower-precision models, such as Q4, that use less memory with possible quality or speed tradeoffs.

Context length

How much text the model can keep in mind at once. Longer context uses more memory.

Inference

Running an existing model for chat, coding help, summaries, or document workflows.

Fine-tuning vs adapter tuning

LoRA and QLoRA train small adapters and need fewer resources than full fine-tuning.

Throughput (tokens/s)

Model output speed. Compare the same model, quantization, context, runtime, and user count; prompt processing is a separate measurement.

Request a quote

€34,667

Estimated total. Final price confirmed before ordering.

Request a quote