Qwen3.8-Flash Now Run Locally via Unsloth GGUF
Qwen3.8-Flash, a 125B MoE model, is now available for local deployment as a GGUF quantization from Unsloth, running on roughly 75 GB of RAM. The Qwen3.8-Flash-Next variant is designed for CPU RAM and unified memory, delivering speeds close to running from VRAM. According to Unsloth’s benchmarks, the model surpasses Claude Opus 4.6 Max in several tests.
Related: Qwen Releases Qwen3.8-Max with 2.4 Trillion Parameters for Autonomous Coding, Qwen3.8-27B Legal Agent Outperforms Opus 4.8