Running Qwen 125B Models on Consumer GPUs like RTX 4090 at 100 Tokens/s
- Authors

- Name
- Nino
- Occupation
- Senior Tech Editor
Running frontier-class large language models locally has traditionally required multi-GPU server clusters or expensive enterprise accelerators like the NVIDIA H100 or A100. However, recent advancements in Mixture-of-Experts (MoE) architectures, low-bit quantization algorithms, and speculative decoding techniques have shifted the paradigm. Developers can now run 100B+ parameter models, such as Qwen series MoE models, on single consumer-grade graphics cards like the NVIDIA GeForce RTX 4090 at real-time speeds exceeding 100 tokens per second (T/s).
Achieving this performance level requires a thorough understanding of hardware bottlenecks, memory bandwidth throughput, tensor parallelism, and hybrid CPU-GPU memory offloading framework designs. For high-availability enterprise applications or developers who want instant, unconstrained scale without hardware investment, platforms like n1n.ai provide unified access to top-tier LLM endpoints. However, for edge deployments and local experimentation, pushing local hardware to its theoretical limits remains an invaluable engineering achievement.
The Memory Bandwidth Wall: Why 125B Models Struggle on Consumer GPUs
To understand how 100 T/s is possible on an RTX 4090, we must first address the foundational hardware limitation: memory bandwidth.
An NVIDIA GeForce RTX 4090 features 24 GB of GDDR6X VRAM with a maximum memory bandwidth of approximately 1,008 GB/s. In contrast, standard dense models with 125 billion parameters stored in FP16 precision require 250 GB of VRAM just to hold the weights in memory. Even at 4-bit quantization (INT4), a dense 125B model occupies roughly 65 GB to 70 GB of memory—far exceeding the 24 GB capacity of an RTX 4090.
Furthermore, LLM token generation (autoregressive decoding) is strictly memory-bandwidth bound. During each token step, every weight parameter in the model must be loaded from memory into the compute cores. The theoretical maximum generation speed (in tokens/sec) for a dense model can be expressed as:
If a dense 125B model could somehow be compressed to 24 GB to fit entirely within VRAM, the maximum theoretical generation speed on an RTX 4090 would be:
Real-world overheads (KV cache allocation, PCIe transfer overheads, kernel launch latency) drop this theoretical limit to under 25-30 T/s. To achieve 100 T/s, we must break free from dense architecture constraints and standard autoregressive generation loops.
Architectural Pillars of the 100 T/s Benchmark
Achieving 100 T/s on consumer hardware relies on four key architectural techniques operating in tandem:
+-------------------------------------------------------------------------+
| 100 T/s Inference Engine |
+-------------------------------------------------------------------------+
| 1. MoE Active Parameter Sparsity (Reduce total parameters loaded/step) |
| 2. KTransformers CPU-GPU Offloading (Offload cold weights to DDR5 RAM) |
| 3. Flash-Decoding & KV Cache FP8 Quantization (Minimize memory footprint)|
| 4. Speculative Decoding with Draft Engine (Execute multiple tokens/step)|
+-------------------------------------------------------------------------+
1. Mixture-of-Experts (MoE) Active Sparsity
Modern 125B models such as Qwen MoE architectures do not activate all 125 billion parameters for every token. Instead, they use top- routing to pass inputs to a tiny fraction of expert networks. For instance, a 125B parameter MoE model might only activate 14B to 22B parameters per token.
Because only the active parameters are loaded per generation step, the memory bandwidth required per token is dramatically reduced. If active parameters account for 16 GB of memory transfer per token, the raw theoretical bandwidth ceiling on an RTX 4090 jumps from 42 T/s to over 63 T/s.
2. KTransformers & Heterogeneous Memory Hierarchies
To fit the inactive expert weights without exceeding 24 GB VRAM, state-of-the-art runtimes like KTransformers (built on top of llama.cpp and vLLM architecture principles) partition the model dynamically:
- GPU VRAM (24 GB): Holds the attention layers, shared routing layers, and top active expert parameters.
- System DDR5 RAM (64 GB - 128 GB): Holds inactive expert weights. Utilizing high-speed PCIe 4.0/5.0 lanes with NUMA-aware CPU memory streaming, system memory handles the non-critical path of weight swapping.
- CPU Acceleration: Modern CPUs with AVX-512 or AMX instruction sets compute offloaded expert layers in parallel with GPU operations.
3. Speculative Decoding with Small Draft Models
Speculative decoding is the single most critical factor for pushing past 60 T/s to exceed 100 T/s.
Instead of evaluating the 125B target model token-by-token, a small, highly optimized draft model (e.g., Qwen-0.5B or Qwen-1.5B running at FP16 entirely in VRAM) generates a sequence of candidate tokens at ultra-high speeds ( T/s).
The 125B target model then evaluates all candidate tokens in a single forward pass using matrix-matrix multiplications () rather than matrix-vector multiplications (). Because matrix-matrix ops are compute-bound rather than memory-bandwidth bound, validating 5 tokens takes nearly the same time as generating 1 token.
If the acceptance rate of draft tokens is high (typically with fine-tuned draft models), the effective generation throughput multiplies by a factor of to .
Hardware Prerequisites & System Configuration
To replicate these benchmark results on a developer workstation, ensure your setup meets the following specs:
| Component | Minimum Specification | Recommended Specification |
|---|---|---|
| GPU | NVIDIA RTX 4090 (24GB VRAM) | NVIDIA RTX 4090 (24GB VRAM) |
| System RAM | 64 GB DDR5-5600 MHz | 128 GB DDR5-6400 MHz (Dual Channel) |
| CPU | Intel Core i9-13900K / Ryzen 9 7950X | Intel Xeon / AMD EPYC with AVX-512 |
| PCIe Interface | PCIe 4.0 x16 (31.5 GB/s bandwidth) | PCIe 5.0 x16 (63.0 GB/s bandwidth) |
| Storage | NVMe M.2 SSD (Read speed > 5000 MB/s) | PCIe Gen4/Gen5 NVMe Enterprise SSD |
Step-by-Step Implementation Guide
The following section details setting up an optimized execution environment using KTransformers and speculative inference hooks.
Step 1: Install Required Dependencies
Ensure CUDA 12.2+ and PyTorch 2.3+ are installed in your Python environment.
# Create isolated environment
conda create -n qwen-flash python=3.10 -y
conda activate qwen-flash
# Install PyTorch with CUDA 12.1 build
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121
# Install FlashAttention-2 for high-performance context processing
pip install flash-attn --no-build-isolation
# Install KTransformers runtime framework
pip install ktransformers
Step 2: Configure Quantized Model Weights
Download the GGUF/EXL2 quantized representation of the 125B model. We recommend an IQ3_M or Q4_K_M quant format for optimal quality-to-memory ratios.
from huggingface_hub import snapshot_download
# Download target model (Qwen 125B MoE Quantized)
model_path = snapshot_download(
repo_id="Qwen/Qwen1.5-MoE-A2.7B-GGUF