LLM / AI model
An AI model is the file or set of weights that does the work. An LLM is a text model for chat, writing, coding, summarizing, and similar tasks.
Plain answers for choosing a local AI computer, understanding the tradeoffs, and ordering without surprises.
The shortest path if you are comparing the current main LLMLab builds.
Local AI means the model runs on your own computer instead of only in a cloud service. Local chat is the simple version: you talk to a local AI model through an app such as Ollama, LM Studio, or Open WebUI.
Not ChatGPT itself. ChatGPT and Claude are hosted services. LLMLab builds run local/open models that can feel similar for chat, coding, and document work, but results depend on the exact model and settings.
Local AI can improve privacy, offline use, and cost control for repeated work. Cloud tools can still be stronger for many tasks. Many buyers use both.
Smaller 7B/8B models are the easiest local target. 13B/14B models benefit from at least 16 GB of VRAM or enough Apple unified memory. 30B/34B-class work usually needs more headroom. 70B-class models need strict validation and are not a normal expectation for a 16 GB GPU.
No. A quote request only collects your contact details and requirements. We then email a written quote with the exact total, availability, configuration, and next payment step.
Component prices and stock move quickly. Verified-price products can use Stripe checkout; other configurations receive a written quote with the exact total and confirmed availability first.
Yes. Start with what you want to do, not the GPU name. If the choice is unclear, request a quote and describe the apps, models, games, or documents you care about.
Short definitions for words you will see on build pages and order forms.
An AI model is the file or set of weights that does the work. An LLM is a text model for chat, writing, coding, summarizing, and similar tasks.
The GPU does the heavy AI work. VRAM is memory on the graphics card. RAM is the computer's general memory for apps, data, and offload.
NVIDIA's compute platform used by many AI tools. It is a major reason NVIDIA PCs are often the simplest path for local AI.
Using an already-trained model to answer, write, code, summarize, or generate output. Most local chat is inference.
A smaller, compressed version of a model. It saves memory and makes local use practical, but can affect quality or speed.
How much text the model can consider at once. Longer context uses more memory, so the same model can fit or fail depending on context.
Retrieval augmented generation: the app searches your documents and gives relevant snippets to the model. It is not permanent training.
Fine-tuning adapts a model to examples. LoRA and QLoRA are lighter adapter methods; they are not the same as training a new foundation model from scratch.
Apple unified memory is shared by the CPU and GPU, but it is not the same as dedicated GPU memory. An eGPU is an external GPU. Our Experimental Mac + eGPU build is a quote-only research setup—not a supported Apple Silicon graphics upgrade.
A quote request does not collect payment. Hardware and service are shown separately, and the written quote confirms the exact total, availability, and work included.
Local AI PC is the €2,500–€3,500 hardware tier for private chat, coding, and everyday CUDA inference with 16 GB VRAM. LLMLab Forge is the €5,000–€6,000 business tier with a 32 GB ECC professional GPU, professional drivers, and more workload headroom.
The three business tiers step from 32 GB ECC GPU memory at €5,000–€6,000, to 72 GB at €10,000–€12,000, to a full 96 GB RTX PRO 6000 and 256 GB ECC system memory at €30,000–€35,000. These are hardware budgets before service, and exact workload fit is checked before quote.
Choose a PC if CUDA tools, NVIDIA VRAM, upgradeability, or broad AI software compatibility matter most. Choose Local AI Mac if you prefer macOS, quiet operation, MLX/Metal workflows, and simpler local chat with Apple unified memory limits.
Yes. Build pages are starting points. When requesting a quote, you can ask for more storage, RAM, quieter cooling, a different case, or a specific software setup.
Model size affects memory and capability. Quantization reduces memory use. Context length controls how much text the model remembers at once. A model that fits with short context may fail or slow down with long context.
No. They are conservative planning guidance. Actual performance depends on model version, quantization, context, software, drivers, thermals, and settings.
The check starts with memory fit, then considers GPU support, system RAM, storage, cooling, power, software maturity, and whether the workflow needs quote review.
Yes, that is the point of the build. It is meant to be a strong gaming and creator PC that also handles useful local AI work, with performance still depending on game, resolution, settings, and drivers.
Usually yes. If a model is using GPU and VRAM, games have less headroom. Schedule long inference, image, or tuning jobs when you are not gaming.
Yes, often. High-VRAM NVIDIA workstation builds can help with rendering, 3D, video, and creator workflows, but exact benefit depends on the application and renderer.
Yes. Local AI Mac systems can run local AI through Apple-friendly tools such as MLX, Metal, Core ML, Ollama, and LM Studio, with limits set by unified memory and software support.
No. Macs do not run CUDA workloads on Apple Silicon. If your workflow requires CUDA, a NVIDIA PC is usually the practical path.
Unified memory is Apple memory shared by CPU and GPU. It can be useful for local AI, but it is not identical to NVIDIA VRAM and should not be compared one-to-one.
No. It is a clearly labeled experimental build using a Sonnet GPU-850T5 enclosure with two independently checkable Estonian retail prices. Apple Silicon does not officially support a normal eGPU graphics path, so we only quote hardware after reproducing the exact workload and rechecking current stock.
Quote-only means the request does not collect payment or card details. We first check pricing, availability, compatibility, order size, and workload fit, then send a written quote.
Yes, where the platform allows it. GPUs, RAM, storage, cooling, and power supplies are easier to plan on desktop PCs than on Macs.
A practical starter setup can be included when agreed in the order or quote. Software ecosystems change quickly, so later model or tool updates may need follow-up work.
Support focuses on the agreed build, handover notes, and initial software path. Ongoing model changes, custom datasets, and workflow engineering may need separate follow-up.
Decision guide
The recommendation guide loads when you reach it.