LLM / AI model
An AI model is the file or set of weights that does the work. An LLM is a text model for chat, writing, coding, summarizing, and similar tasks.
Plain answers for choosing a local AI computer, understanding the tradeoffs, and ordering without surprises.
The shortest path if you are comparing the current main LLMLab builds.
Local AI means the model runs on your own computer instead of only in a cloud service. Local chat is the simple version: you talk to a local AI model through an app such as Ollama, LM Studio, or Open WebUI.
Not ChatGPT itself. ChatGPT and Claude are hosted services. LLMLab builds run local/open models that can feel similar for chat, coding, and document work, but results depend on the exact model and settings.
Smaller 7B/8B models are the easiest local target. 13B/14B models benefit from at least 16 GB of VRAM or enough Apple unified memory. 30B/34B-class work usually needs more headroom. 70B-class models need strict validation and are not a normal expectation for a 16 GB GPU.
No. A quote request only collects your contact details and requirements. We then email a written quote with the exact total, availability, configuration, and next payment step.
Component prices and stock move quickly. Verified-price products can use Stripe checkout; other configurations receive a written quote with the exact total and confirmed availability first.
Yes. Start with what you want to do, not the GPU name. If the choice is unclear, request a quote and describe the apps, models, games, or documents you care about.
Short definitions for words you will see on build pages and order forms.
An AI model is the file or set of weights that does the work. An LLM is a text model for chat, writing, coding, summarizing, and similar tasks.
The GPU does the heavy AI work. VRAM is memory on the graphics card. RAM is the computer's general memory for apps, data, and offload.
NVIDIA's compute platform used by many AI tools. It is a major reason NVIDIA PCs are often the simplest path for local AI.
Using an already-trained model to answer, write, code, summarize, or generate output. Most local chat is inference.
A smaller, compressed version of a model. It saves memory and makes local use practical, but can affect quality or speed.
How much text the model can consider at once. Longer context uses more memory, so the same model can fit or fail depending on context.
Retrieval augmented generation: the app searches your documents and gives relevant snippets to the model. It is not permanent training.
Fine-tuning adapts a model to examples. LoRA and QLoRA are lighter adapter methods; this is what Model Tuning PC is for, not training a new foundation model from scratch.
Apple unified memory is shared by the CPU and GPU, but it is not the same as dedicated GPU memory. An eGPU is an external GPU. Our Mac + eGPU build is an experimental, quote-only research setup—not a supported Apple Silicon graphics upgrade.
A quote request does not collect payment. The component total is the estimated cost of the parts; the separately listed service fee covers building, configuring, testing, and handing over the computer. For complete PC builds, the shared service fee scales from 20% on the lowest-priced complete build to 10% on the highest-priced complete build.
Local Chat PC is for everyday private chat, coding assistants, and small RAG projects. Model Tuning PC adds a 32 GB Radeon AI PRO GPU for Linux and ROCm-based Python workflows, datasets, checkpoints, and adapter-tuning experiments; the exact software stack is checked before quote.
AI + Gaming PC combines games, creator apps, and local AI in one capable computer. Pro AI Workstation is for larger projects, heavier creative or rendering work, more memory headroom, and workloads that need individual compatibility checks.
Choose a PC if CUDA tools, NVIDIA VRAM, upgradeability, or broad AI software compatibility matter most. Choose AI-Ready Mac if you prefer macOS, quiet operation, MLX/Metal workflows, and simpler local chat with Apple unified memory limits.
Yes. Build pages are starting points. When requesting a quote, you can ask for more storage, RAM, quieter cooling, a different case, or a specific software setup.
Local AI can improve privacy, offline use, and cost control for repeated work. Cloud tools can still be stronger for many tasks. Many buyers use both.
Local chat is a ChatGPT-like interface connected to a model running on your own computer. The answer quality depends on the model, prompt, quantization, context length, and software.
Inference means using a model after it has already been trained. Chatting, coding help, summarizing, and document Q&A are inference workloads.
Model size affects memory and capability. Quantization reduces memory use. Context length controls how much text the model remembers at once. A model that fits with short context may fail or slow down with long context.
No. They are conservative planning guidance. Actual performance depends on model version, quantization, context, software, drivers, thermals, and settings.
The check starts with memory fit, then considers GPU support, system RAM, storage, cooling, power, software maturity, and whether the workflow needs quote review.
Yes, that is the point of the build. It is meant to be a strong gaming and creator PC that also handles useful local AI work, with performance still depending on game, resolution, settings, and drivers.
Yes, but not at full load at the same time. Heavy AI jobs and games share the GPU, VRAM, power, and cooling.
Usually yes. If a model is using GPU and VRAM, games have less headroom. Schedule long inference, image, or tuning jobs when you are not gaming.
Yes, often. High-VRAM NVIDIA workstation builds can help with rendering, 3D, video, and creator workflows, but exact benefit depends on the application and renderer.
Yes. AI-Ready Mac systems can run local AI through Apple-friendly tools such as MLX, Metal, Core ML, Ollama, and LM Studio, with limits set by unified memory and software support.
No. Macs do not run CUDA workloads on Apple Silicon. If your workflow requires CUDA, a NVIDIA PC is usually the practical path.
Unified memory is Apple memory shared by CPU and GPU. It can be useful for local AI, but it is not identical to NVIDIA VRAM and should not be compared one-to-one.
No. It is a clearly labeled experimental build using a Sonnet GPU-850T5 enclosure with two independently checkable Estonian retail prices. Apple Silicon does not officially support a normal eGPU graphics path, so we only quote hardware after reproducing the exact workload and rechecking current stock.
Quote-only means the request does not collect payment or card details. We first check pricing, availability, compatibility, order size, and workload fit, then send a written quote.
You receive a confirmation email and no payment is taken. We review the request, then reply with questions or a written quote containing the exact total, availability, and proposed adjustments.
Yes, where the platform allows it. GPUs, RAM, storage, cooling, and power supplies are easier to plan on desktop PCs than on Macs.
A practical starter setup can be included when agreed in the order or quote. Software ecosystems change quickly, so later model or tool updates may need follow-up work.
Support focuses on the agreed build, handover notes, and initial software path. Ongoing model changes, custom datasets, and workflow engineering may need separate follow-up.
Decision guide
The recommendation guide loads when you reach it.