← News
News

ISTA-DASLab publishes GSQ-RCO quantization family: 32GB RAM runs Qwen3.8-Flash-Next 180B

ISTA-DASLab publishes GSQ-RCO GGUF quants on HuggingFace — Q2_0 / IQ2_XS / IQ3_XXS / IQ3_S / Coder IQ1_M; weights 32–83GB, spanning 32GB-RAM workstations to 128GB memory servers.

Source: HuggingFace ISTA-DASLab

What it is

ISTA-DASLab has published the GSQ-RCO GGUF quantization family for Qwen3.8-Flash-Next on HuggingFace, covering everything from consumer workstations to memory-rich servers:

Quant Weight size Min RAM Recommended scenario
Coder (IQ1_M) ~32 GB 48 GB Code-specialized, speed first
IQ2_XS ~45 GB 64 GB Recommended — speed/quality balance
IQ3_XXS ~76 GB 96 GB Quality-sensitive
IQ3_S ~83 GB 128 GB Close to full-model quality
Q2_0 ~66 GB 96 GB Q2_K style, CPU-friendly

Why GSQ-RCO instead of vanilla GGUF

GSQ is ISTA-DASLab's own Grouped Scaling Quantization algorithm; combined with RCO (Row-Column Ordering), it works especially well on MoE models with extreme activation skew (Qwen3.8-Flash-Next's top-10 routing is a textbook case). In short: at the IQ2_XS tier, GSQ-RCO has higher quality and less speed loss than vanilla GGUF IQ2_XS.

Pairing with Strata

GSQ-RCO GGUF is fully compatible with Strata 0.1.38+ — Strata consumes GGUF directly; GSQ-RCO needs no extra patch. Download GSQ-RCO IQ2_XS (45GB) and you're good to go with Strata:

strata serve --model Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS.gguf \
  --gpu-layers 12 --ctx 4096 --batch 1

Pick-by-scenario guidance

  • First try: IQ2_XS (45GB) — sweet spot for 64GB RAM + 12GB GPU
  • Code-heavy workload: Coder IQ1_M (32GB) — starts at 48GB RAM, actually fastest (more aggressive sparse routing)
  • Quality priority: IQ3_XXS (76GB) — 96GB RAM, holds its own against fp16
  • Maximum quality: IQ3_S (83GB) — 128GB RAM, essentially bfloat16-grade

Local deployment angle

Download straight from HuggingFace; the repo includes SHA256 checksums. Recommended path: grab IQ2_XS first to validate the pipeline, then swap quant tiers as needed.