VRAM
Memory on the graphics card; usually the main limit for local AI model size.
Component overview
ASUS graphics card
Brand
ASUS
Verified manufacturer part number
90YV0LX0-M0NA00
Release
2025 Q1
VRAM
16GB GDDR7
Architecture
Blackwell
Display Power
360W
Connector Standard
12V-2x6
Minimum PSU
850W
Dual GPU Capable
No
Memory Bus
256-bit
Bandwidth
960 GB/s
CUDA Cores
10752
Tensor Cores
336
RT Cores
84
Base / Boost Clock
2295 / 2685 MHz
TDP
360W
PCIe Generation
PCIe 5.0
Slot Width
3-slot
Length
304mm
Power Connectors
1x 16-pin 12V-2x6
Recommended PSU
850W
AI Score
96
Source
https://www.asus.com/motherboards-components/graphics-cards/prime/prime-rtx5080-o16g/techspec/
Inference Notes
Exact ASUS PRIME-RTX5080-O16G. Multiple Estonian retailer offers were verified; checkout still requires fresh exact-SKU observations.
13B/14B Q4 at practical context
Everyday local LLM fit assumes quantization and moderate context; 30B is not a normal target without more VRAM.
This build handles 12B-class models better than larger dense models; choose a 24GB+ VRAM build for a practical 27B target.
VRAM
Memory on the graphics card; usually the main limit for local AI model size.
System RAM vs GPU memory
RAM helps apps and data work; GPU memory usually decides which model size can run quickly.
CUDA
NVIDIA’s software layer for many AI tools. Macs and AMD GPUs do not run CUDA workflows the same way.
Apple unified memory
Apple Silicon memory shared by CPU, GPU, macOS, and apps. Useful for Mac AI, but not the same as NVIDIA VRAM.
7B / 8B / 14B / 70B
Approximate model size in billions of parameters. Larger numbers usually need more memory and may run slower.
Quantized models
Compressed models, such as Q4, that use less memory with possible quality or speed tradeoffs.
Context length
How much text the model can keep in mind at once. Longer context uses more memory.
Inference
Running an existing model for chat, coding help, summaries, or document workflows.
Fine-tuning vs adapter tuning
Fine-tuning adapts a model; adapter tuning, such as LoRA/QLoRA, is a lighter way to steer an existing model with examples.
Dual-GPU limitations
Two GPUs do not automatically combine VRAM into one large pool. Software must explicitly support multiple GPUs.
eGPU limitations
External GPUs need enclosure, driver, and runtime validation, and usually do not mean macOS graphics or gaming acceleration.
Planning estimate / written quote
The displayed estimate helps with budgeting. A written quote contains the exact price, confirmed availability, and any proposed substitutions.
Estonian market planning estimate. A written quote includes the exact price and confirmed availability.
Swipe the chart sideways to inspect every date.
Planning estimate: €1,758
The displayed price is a planning estimate. A written quote provides the exact price and availability.
Online payment is unavailable for this product; request a quote instead.
Use the quote request on this page. It does not collect payment or card details; we email the exact price and availability.
What happens after your quote request
Support and questions continue through the order or quote email thread.
Request a quote
The request does not collect payment or card details. We send the exact price, availability, and any proposed substitutions in a written quote.
Examples for ASUS Prime GeForce RTX 5080 OC Edition 16GB GDDR7, based on GPU VRAM or Apple unified memory plus RAM headroom. System RAM is not treated as VRAM.
Recommended model
A current multimodal everyday model that makes good use of a 16GB GPU.
It is the strongest comfortable general-purpose starting point for most 16GB builds.
Fast
Likely good memory headroom for this quantized model at normal context sizes.
Private chat, coding help, document summaries, and image questions
A strong everyday model for 12GB-class GPUs and a practical coding pick when speed matters.
Expected experience: Fast
Likely good memory headroom for this quantized model at normal context sizes.
High-end local chat experiments on 48GB GPUs or 96GB+ Apple systems
A clear upper-limit example for checking whether a machine can attempt a 70B-class model.
Expected experience: Not recommended
Needs at least 64GB system RAM; this machine reports 32GB.
Fast private chat, short summaries, and first local-AI experiments
A current compact model for learning local AI, multilingual chat, and short document work.
Expected experience: Fast
Likely good memory headroom for this quantized model at normal context sizes.
Assumptions
Family: Google Gemma 4
Parameters: 11.95B
Structure: 11.95B dense
License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.
Native context: 262,144 tokens
Quantization: Q4_0 QAT
Approx. Q4 weights: 7.2 GB
Default estimate: 11.5 GB @ 8,192 tokens
Weights / KV / runtime / margin: 7.2 GB / 1.5 GB / 1.2 GB / 1.5 GB
CPU/RAM fallback: Not recommended
VRAM: 12 GB minimum / 16 GB recommended
RAM: 24 GB minimum / 32 GB recommended
Current run mode: Comfortable GPU fit
Expected experience: Fast
Full GPU offload: Only when the memory estimate and context fit
Context warning: Its large context window still requires substantial KV-cache headroom.
| Context | Weights | KV | Runtime | Margin | Estimated GPU memory |
|---|---|---|---|---|---|
| 4K | 7.2 GB | 1 GB | 1.2 GB | 1.5 GB | 11 GB |
| 8K | 7.2 GB | 1.5 GB | 1.2 GB | 1.5 GB | 11.5 GB |
| 16K | 7.2 GB | 2.5 GB | 1.2 GB | 1.5 GB | 12.5 GB |
| 32K | 7.2 GB | 5 GB | 1.2 GB | 2 GB | 15.5 GB |
Research sources
Researched: 2026-07-30
Family: Qwen3.5
Parameters: 9B
Structure: 9B dense
License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.
Native context: 262,144 tokens
Extended context: 1,010,000 via YaRN; not a default fit assumption
Quantization: Q4_K_M
Approx. Q4 weights: 6.6 GB
Default estimate: 11 GB @ 8,192 tokens
Weights / KV / runtime / margin: 6.6 GB / 1.5 GB / 1.2 GB / 1.5 GB
CPU/RAM fallback: Not recommended
VRAM: 8 GB minimum / 12 GB recommended
RAM: 16 GB minimum / 32 GB recommended
Current run mode: Comfortable GPU fit
Expected experience: Fast
Full GPU offload: Only when the memory estimate and context fit
Context warning: Start around 8K context even though the model supports much more.
| Context | Weights | KV | Runtime | Margin | Estimated GPU memory |
|---|---|---|---|---|---|
| 4K | 6.6 GB | 1 GB | 1.2 GB | 1.5 GB | 10.5 GB |
| 8K | 6.6 GB | 1.5 GB | 1.2 GB | 1.5 GB | 11 GB |
| 16K | 6.6 GB | 2.5 GB | 1.2 GB | 1.5 GB | 12 GB |
| 32K | 6.6 GB | 5 GB | 1.2 GB | 2 GB | 15 GB |
Family: Meta Llama 3.3
Parameters: 70.60B
Structure: 70.6B dense
License: Llama 3.3 Community License; allowed with terms; gated access. For guidance only; review the model licence before commercial use.
Native context: 131,072 tokens
Quantization: Q4_K_M
Approx. Q4 weights: 42.5 GB
Default estimate: 51 GB @ 4,096 tokens
Weights / KV / runtime / margin: 42.5 GB / 3 GB / 1.5 GB / 4 GB
CPU/RAM fallback: Not recommended
VRAM: 48 GB minimum / 64 GB recommended
RAM: 64 GB minimum / 96 GB recommended
Current run mode: Not realistic here
Expected experience: Not recommended
Full GPU offload: Often limited
Context warning: The Q4 weights alone use about 42.5GB before KV cache, runtime overhead, and desktop headroom.
| Context | Weights | KV | Runtime | Margin | Estimated GPU memory |
|---|---|---|---|---|---|
| 4K | 42.5 GB | 3 GB | 1.5 GB | 4 GB | 51 GB |
| 8K | 42.5 GB | 6 GB | 1.5 GB | 4 GB | 54 GB |
| 16K | 42.5 GB | 12 GB | 1.5 GB | 4.5 GB | 60.5 GB |
| 32K | 42.5 GB | 24 GB | 1.5 GB | 5.5 GB | 73.5 GB |
Family: Qwen3.5
Parameters: 4B
Structure: 4B dense
License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.
Native context: 262,144 tokens
Extended context: 1,010,000 via YaRN; not a default fit assumption
Quantization: Q4_K_M
Approx. Q4 weights: 3.4 GB
Default estimate: 6 GB @ 8,192 tokens
Weights / KV / runtime / margin: 3.4 GB / 0.5 GB / 0.8 GB / 1 GB
CPU/RAM fallback: Small-model fallback only
VRAM: 4 GB minimum / 8 GB recommended
RAM: 8 GB minimum / 16 GB recommended
Current run mode: Comfortable GPU fit
Expected experience: Fast
Full GPU offload: Only when the memory estimate and context fit
Context warning: The advertised long context is not a practical default on low-memory machines.
| Context | Weights | KV | Runtime | Margin | Estimated GPU memory |
|---|---|---|---|---|---|
| 4K | 3.4 GB | 0.5 GB | 0.8 GB | 1 GB | 6 GB |
| 8K | 3.4 GB | 0.5 GB | 0.8 GB | 1 GB | 6 GB |
| 16K | 3.4 GB | 1 GB | 0.8 GB | 1 GB | 6.5 GB |
| 32K | 3.4 GB | 2 GB | 0.8 GB | 1 GB | 7.5 GB |
Local AI performance is approximate. Results depend on quantization, context length, backend, drivers, and whether the model plus KV cache fits in VRAM or Apple unified memory.
Trust and process
Exact item and availability
We confirm the exact model, current price, and availability before fulfillment. Submitting a quote request does not collect payment.
Compatibility scope
We review known fit information and flag evident size, power, or interface constraints. Final compatibility depends on the rest of your system and the details you provide.
Handover in Estonia
Pickup or local delivery method and timing are agreed after availability is confirmed.
Returns, warranty, and support
Return and warranty handling depends on order state, the component, manufacturer, and retailer. Questions continue through the order or quote email thread.
Trust details
Contact and support
Questions continue through the order or quote email thread. Replying to the confirmation is the fastest path.
Warranty
Warranty handling depends on the component, manufacturer, and retailer; the practical path is confirmed case by case.
Exact item
The manufacturer, model, availability, and final amount are confirmed in checkout or in writing before fulfillment. A replacement is never made silently.
Compatibility limits
Catalog specifications and known constraints are buyer guidance. Confirming fit with an existing system requires the full parts list plus dimensions, power, interfaces, and software requirements.
Delivery and returns
Pickup or local delivery is agreed after availability is confirmed. Returns and cancellations depend on order state and the applicable terms.
Pricing basis
The page labels whether pricing uses recent market evidence, a planning estimate, or a written quote. A component price does not imply whole-system assembly, software setup, or completed-system testing.