AI Tutorials
Running Qwen3.8-Flash-Next 125B on Three RTX 3090s at 80 Tokens per Second
Discover how expert usage sorting, dynamic tier-based quantization, and llama.cpp CUDA kernel patching enable running the 125B Qwen3.8-Flash-Next MoE model on three consumer RTX 3090 GPUs at 80 tokens/sec.
Read more →