GPU memory and RAM
GPU memory usually limits model size; RAM supports apps, data, and model offload.
Component overview
ASUS graphics card
Brand
ASUS
Verified manufacturer part number
90YV0MH2-M0NA00
Release
2025 Q1
VRAM
16GB GDDR7
Architecture
Blackwell
Display Power
180W
Connector Standard
PCIe 8-pin
Minimum PSU
550W
Dual GPU Capable
No
Memory Bus
128-bit
Bandwidth
448 GB/s
CUDA Cores
4608
Tensor Cores
144
RT Cores
36
Base / Boost Clock
2407 / 2647 MHz
TDP
180W
PCIe Generation
PCIe 5.0
Slot Width
3-slot
Length
304mm
Power Connectors
1x 8-pin
Recommended PSU
550W
Inference Notes
Exact ASUS PRIME-RTX5060TI-O16G. Reference price is backed by reviewed exact-MPN offers; checkout still requires current accepted observations.
13B/14B Q4 at practical context
Everyday local LLM fit assumes quantization and moderate context; 30B is not a normal target without more VRAM.
This build handles 12B-class models better than larger dense models; choose a 24GB+ VRAM build for a practical 27B target.
GPU memory and RAM
GPU memory usually limits model size; RAM supports apps, data, and model offload.
CUDA
NVIDIA’s software layer for many AI tools. Macs and AMD GPUs do not run CUDA workflows the same way.
Apple unified memory
Memory shared by the CPU, GPU, macOS, and apps; it is not the same as NVIDIA VRAM.
7B / 8B / 14B / 70B
Approximate model size in billions of parameters; larger models usually need more memory.
Quantized models
Lower-precision models, such as Q4, that use less memory with possible quality or speed tradeoffs.
Context length
How much text the model can keep in mind at once. Longer context uses more memory.
Inference
Running an existing model for chat, coding help, summaries, or document workflows.
Fine-tuning vs adapter tuning
LoRA and QLoRA train small adapters and need fewer resources than full fine-tuning.
Throughput (tokens/s)
Model output speed. Compare the same model, quantization, context, runtime, and user count; prompt processing is a separate measurement.
Uses market prices where available and reference estimates otherwise.
Swipe the chart sideways to inspect every date.
€903
Estimated total. Final price confirmed before ordering.
Online payment is unavailable for this product; request a quote instead.
Use the quote request on this page. It does not collect payment or card details; we email the exact price and availability.
Request a quote
What happens after your quote request
Support and questions continue through the order or quote email thread.
Recommended model
A current multimodal everyday model that makes good use of a 16GB GPU.
Fast
Likely good memory headroom for this quantized model at normal context sizes.
Private chat, coding help, document summaries, and image questions
A strong everyday model for 12GB-class GPUs and a practical coding pick when speed matters.
Expected experience: Fast
Likely good memory headroom for this quantized model at normal context sizes.
High-end local chat experiments on 48GB GPUs or 96GB+ Apple systems
A clear upper-limit example for checking whether a machine can attempt a 70B-class model.
Expected experience: Not recommended
Needs at least 64GB system RAM; this machine reports 32GB.
Assumptions
Family: Google Gemma 4
Parameters: 11.95B
Structure: 11.95B dense
License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.
Native context: 262,144 tokens
Quantization: Q4_0 QAT
Approx. Q4 weights: 7.2 GB
Default estimate: 11.5 GB @ 8,192 tokens
Weights / KV / runtime / margin: 7.2 GB / 1.5 GB / 1.2 GB / 1.5 GB
CPU/RAM fallback: Not recommended
VRAM: 12 GB minimum / 16 GB recommended
RAM: 24 GB minimum / 32 GB recommended
Current run mode: Comfortable GPU fit
Expected experience: Fast
Full GPU offload: Only when the memory estimate and context fit
Context warning: Its large context window still requires substantial KV-cache headroom.
| Context | Weights | KV | Runtime | Margin | Estimated GPU memory |
|---|---|---|---|---|---|
| 4K | 7.2 GB | 1 GB | 1.2 GB | 1.5 GB | 11 GB |
| 8K | 7.2 GB | 1.5 GB | 1.2 GB | 1.5 GB | 11.5 GB |
| 16K | 7.2 GB | 2.5 GB | 1.2 GB | 1.5 GB | 12.5 GB |
| 32K | 7.2 GB | 5 GB | 1.2 GB | 2 GB | 15.5 GB |
Swipe the table sideways to inspect every memory estimate.
Research sources
Researched: 2026-07-30
Family: Qwen3.5
Parameters: 9B
Structure: 9B dense
License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.
Native context: 262,144 tokens
Extended context: 1,010,000 via YaRN; not a default fit assumption
Quantization: Q4_K_M
Approx. Q4 weights: 6.6 GB
Default estimate: 11 GB @ 8,192 tokens
Weights / KV / runtime / margin: 6.6 GB / 1.5 GB / 1.2 GB / 1.5 GB
CPU/RAM fallback: Not recommended
VRAM: 8 GB minimum / 12 GB recommended
RAM: 16 GB minimum / 32 GB recommended
Current run mode: Comfortable GPU fit
Expected experience: Fast
Full GPU offload: Only when the memory estimate and context fit
Context warning: Start around 8K context even though the model supports much more.
| Context | Weights | KV | Runtime | Margin | Estimated GPU memory |
|---|---|---|---|---|---|
| 4K | 6.6 GB | 1 GB | 1.2 GB | 1.5 GB | 10.5 GB |
| 8K | 6.6 GB | 1.5 GB | 1.2 GB | 1.5 GB | 11 GB |
| 16K | 6.6 GB | 2.5 GB | 1.2 GB | 1.5 GB | 12 GB |
| 32K | 6.6 GB | 5 GB | 1.2 GB | 2 GB | 15 GB |
Swipe the table sideways to inspect every memory estimate.
Family: Meta Llama 3.3
Parameters: 70.60B
Structure: 70.6B dense
License: Llama 3.3 Community License; allowed with terms; gated access. For guidance only; review the model licence before commercial use.
Native context: 131,072 tokens
Quantization: Q4_K_M
Approx. Q4 weights: 42.5 GB
Default estimate: 51 GB @ 4,096 tokens
Weights / KV / runtime / margin: 42.5 GB / 3 GB / 1.5 GB / 4 GB
CPU/RAM fallback: Not recommended
VRAM: 48 GB minimum / 64 GB recommended
RAM: 64 GB minimum / 96 GB recommended
Current run mode: Not realistic here
Expected experience: Not recommended
Full GPU offload: Often limited
Context warning: The Q4 weights alone use about 42.5GB before KV cache, runtime overhead, and desktop headroom.
| Context | Weights | KV | Runtime | Margin | Estimated GPU memory |
|---|---|---|---|---|---|
| 4K | 42.5 GB | 3 GB | 1.5 GB | 4 GB | 51 GB |
| 8K | 42.5 GB | 6 GB | 1.5 GB | 4 GB | 54 GB |
| 16K | 42.5 GB | 12 GB | 1.5 GB | 4.5 GB | 60.5 GB |
| 32K | 42.5 GB | 24 GB | 1.5 GB | 5.5 GB | 73.5 GB |
Swipe the table sideways to inspect every memory estimate.
Local AI performance is approximate. Results depend on quantization, context length, backend, drivers, and whether the model plus KV cache fits in VRAM or Apple unified memory.
Trust and process
Exact item and availability
We confirm the model, price, and availability before fulfillment. A quote request does not collect payment.
Compatibility scope
We check known size, power, and interface constraints. Final compatibility depends on the rest of your system.
Handover in Estonia
Pickup or local delivery method and timing are agreed after availability is confirmed.
Returns, warranty, and support
Handling depends on order state and the component, manufacturer, and retailer terms. Questions continue by email.
Trust details
Contact and support
Replying to the order or quote confirmation is the fastest path.
Exact item and fit
We confirm the model, availability, and price. Existing-system compatibility depends on the complete parts list.
Delivery and returns
Handover is agreed after availability is confirmed. Returns depend on order state and applicable terms.
Pricing basis
The page distinguishes market data, estimates, and written quotes. A component price does not include whole-system service.