Family: Qwen3.6
Parameters: 35B
Structure: 35B total / 3B active MoE
License: Apache License 2.0; allowed. For guidance only; review the model licence before commercial use.
Native context: 262,144 tokens
Extended context: 1,010,000 via YaRN; not a default fit assumption
Quantization: Q4_K_M
Approx. Q4 weights: 24 GB
Default estimate: 31.5 GB @ 8,192 tokens
Weights / KV / runtime / margin: 24 GB / 2.5 GB / 2 GB / 3 GB
CPU/RAM fallback: Not recommended
VRAM: 32 GB minimum / 48 GB recommended
RAM: 64 GB minimum / 64 GB recommended
Current run mode: Comfortable GPU fit
Expected experience: Usable
Full GPU offload: Only when the memory estimate and context fit
Context warning: Long repository context can use the remaining headroom quickly, especially on 32GB GPUs.