Frequently Asked Questions

Plain answers for choosing a local AI computer, understanding the tradeoffs, and ordering without surprises.

Start here

The shortest path if you are comparing the current main LLMLab builds.

What does local AI mean?

Local AI means the model runs on your own computer instead of only in a cloud service. Local chat is the simple version: you talk to a local AI model through an app such as Ollama, LM Studio, or Open WebUI.

Can these PCs run ChatGPT?

Not ChatGPT itself. ChatGPT and Claude are hosted services. LLMLab builds run local/open models that can feel similar for chat, coding, and document work, but results depend on the exact model and settings.

Why run AI locally instead of using ChatGPT or Claude?

Local AI can improve privacy, offline use, and cost control for repeated work. Cloud tools can still be stronger for many tasks. Many buyers use both.

What model sizes can I run?

Smaller 7B/8B models are the easiest local target. 13B/14B models benefit from at least 16 GB of VRAM or enough Apple unified memory. 30B/34B-class work usually needs more headroom. 70B-class models need strict validation and are not a normal expectation for a 16 GB GPU.

Is payment taken when I request a quote?

No. A quote request only collects your contact details and requirements. We then email a written quote with the exact total, availability, configuration, and next payment step.

Why do some products require a written quote?

Component prices and stock move quickly. Verified-price products can use Stripe checkout; other configurations receive a written quote with the exact total and confirmed availability first.

Can you help me choose if I do not understand the specs?

Yes. Start with what you want to do, not the GPU name. If the choice is unclear, request a quote and describe the apps, models, games, or documents you care about.

Technical terms in plain language

Short definitions for words you will see on build pages and order forms.

LLM / AI model

An AI model is the file or set of weights that does the work. An LLM is a text model for chat, writing, coding, summarizing, and similar tasks.

GPU, VRAM, and RAM

The GPU does the heavy AI work. VRAM is memory on the graphics card. RAM is the computer's general memory for apps, data, and offload.

CUDA

NVIDIA's compute platform used by many AI tools. It is a major reason NVIDIA PCs are often the simplest path for local AI.

Inference

Using an already-trained model to answer, write, code, summarize, or generate output. Most local chat is inference.

Quantization

A smaller, compressed version of a model. It saves memory and makes local use practical, but can affect quality or speed.

Context length

How much text the model can consider at once. Longer context uses more memory, so the same model can fit or fail depending on context.

RAG

Retrieval augmented generation: the app searches your documents and gives relevant snippets to the model. It is not permanent training.

Fine-tuning, LoRA, and QLoRA

Fine-tuning adapts a model to examples. LoRA and QLoRA are lighter adapter methods; they are not the same as training a new foundation model from scratch.

Unified memory and eGPU

Apple unified memory is shared by the CPU and GPU, but it is not the same as dedicated GPU memory. An eGPU is an external GPU. Our Experimental Mac + eGPU build is a quote-only research setup—not a supported Apple Silicon graphics upgrade.

Written quote, component total, assembly and setup

A quote request does not collect payment. Hardware and service are shown separately, and the written quote confirms the exact total, availability, and work included.

Choosing a build

What is the difference between Local AI PC and LLMLab Forge?

Local AI PC is the €2,500–€3,500 hardware tier for private chat, coding, and everyday CUDA inference with 16 GB VRAM. LLMLab Forge is the €5,000–€6,000 business tier with a 32 GB ECC professional GPU, professional drivers, and more workload headroom.

How do LLMLab Forge, LLMLab Forge Pro, and LLMLab Forge Ultra differ?

The three business tiers step from 32 GB ECC GPU memory at €5,000–€6,000, to 72 GB at €10,000–€12,000, to a full 96 GB RTX PRO 6000 and 256 GB ECC system memory at €30,000–€35,000. These are hardware budgets before service, and exact workload fit is checked before quote.

Should I choose a PC or a Mac for local AI?

Choose a PC if CUDA tools, NVIDIA VRAM, upgradeability, or broad AI software compatibility matter most. Choose Local AI Mac if you prefer macOS, quiet operation, MLX/Metal workflows, and simpler local chat with Apple unified memory limits.

Can I request changes to a build?

Yes. Build pages are starting points. When requesting a quote, you can ask for more storage, RAM, quieter cooling, a different case, or a specific software setup.

Models and performance

Why do model size, quantization, and context length matter?

Model size affects memory and capability. Quantization reduces memory use. Context length controls how much text the model remembers at once. A model that fits with short context may fail or slow down with long context.

Are model recommendations guaranteed?

No. They are conservative planning guidance. Actual performance depends on model version, quantization, context, software, drivers, thermals, and settings.

How does LLMLab decide if a build is suitable?

The check starts with memory fit, then considers GPU support, system RAM, storage, cooling, power, software maturity, and whether the workflow needs quote review.

Gaming and creator work

Can the AI + Gaming PC also play games well?

Yes, that is the point of the build. It is meant to be a strong gaming and creator PC that also handles useful local AI work, with performance still depending on game, resolution, settings, and drivers.

Will AI work reduce gaming performance if both run at once?

Usually yes. If a model is using GPU and VRAM, games have less headroom. Schedule long inference, image, or tuning jobs when you are not gaming.

Is LLMLab Forge Pro useful for rendering or creator workloads?

Yes, often. High-VRAM NVIDIA workstation builds can help with rendering, 3D, video, and creator workflows, but exact benefit depends on the application and renderer.

Mac and eGPU setups

Can Macs run local AI?

Yes. Local AI Mac systems can run local AI through Apple-friendly tools such as MLX, Metal, Core ML, Ollama, and LM Studio, with limits set by unified memory and software support.

Does Mac run CUDA workloads?

No. Macs do not run CUDA workloads on Apple Silicon. If your workflow requires CUDA, a NVIDIA PC is usually the practical path.

What is unified memory?

Unified memory is Apple memory shared by CPU and GPU. It can be useful for local AI, but it is not identical to NVIDIA VRAM and should not be compared one-to-one.

Is the Experimental Mac + eGPU build a normal recommendation?

No. It is a clearly labeled experimental build using a Sonnet GPU-850T5 enclosure with two independently checkable Estonian retail prices. Apple Silicon does not officially support a normal eGPU graphics path, so we only quote hardware after reproducing the exact workload and rechecking current stock.

Orders, pricing, delivery, and support

What does quote-only mean?

Quote-only means the request does not collect payment or card details. We first check pricing, availability, compatibility, order size, and workload fit, then send a written quote.

Do you support upgrades later?

Yes, where the platform allows it. GPUs, RAM, storage, cooling, and power supplies are easier to plan on desktop PCs than on Macs.

Do you install the AI software?

A practical starter setup can be included when agreed in the order or quote. Software ecosystems change quickly, so later model or tool updates may need follow-up work.

What support is included?

Support focuses on the agreed build, handover notes, and initial software path. Ongoing model changes, custom datasets, and workflow engineering may need separate follow-up.

Decision guide

Still deciding?

The recommendation guide loads when you reach it.