[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f352gpr7kjfuu5":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":15,"bodyZh":16,"bodyEn":17,"sourceName":18,"sourceUrl":19,"updatedAt":20},"gsq-rco-quants","NEWS","2026-10-05T04:48:00.000Z","ISTA-DASLab 发布 GSQ-RCO 量化家族：32GB RAM 也能跑 Qwen3.8-Flash-Next 180B","ISTA-DASLab publishes GSQ-RCO quantization family: 32GB RAM runs Qwen3.8-Flash-Next 180B","ISTA-DASLab 在 HuggingFace 发布 GSQ-RCO GGUF 量化家族：Q2_0 \u002F IQ2_XS \u002F IQ3_XXS \u002F IQ3_S \u002F Coder IQ1_M，权重从 32GB 到 83GB，覆盖从 32GB RAM 工作站到 128GB 内存服务器的不同档位。","ISTA-DASLab publishes GSQ-RCO GGUF quants on HuggingFace — Q2_0 \u002F IQ2_XS \u002F IQ3_XXS \u002F IQ3_S \u002F Coder IQ1_M; weights 32–83GB, spanning 32GB-RAM workstations to 128GB memory servers.",[12],{"id":13,"name":14},11,"Qwen3.8-Flash-Next",[],"### 是什么\n\n[ISTA-DASLab](https:\u002F\u002Fhuggingface.co\u002FISTA-DASLab) 在 HuggingFace 上线了 Qwen3.8-Flash-Next 的 **GSQ-RCO GGUF** 量化家族，覆盖从消费工作站到内存服务器的完整档位：\n\n| 量化档 | 权重大小 | 最低 RAM | 推荐场景 |\n|---|---|---|---|\n| Coder (IQ1_M) | ~32 GB | 48 GB | 代码专用，速度优先 |\n| IQ2_XS | ~45 GB | 64 GB | **推荐档**——速度\u002F质量平衡 |\n| IQ3_XXS | ~76 GB | 96 GB | 质量敏感 |\n| IQ3_S | ~83 GB | 128 GB | 接近全模型质量 |\n| Q2_0 | ~66 GB | 96 GB | Q2_K 风格，CPU 友好 |\n\n### 为什么发 GSQ-RCO 而不是普通 GGUF\n\nGSQ 是 ISTA-DASLab 自研的**分组缩放量化**算法，结合 RCO（Row-Column Ordering，行列重排）后，对**激活分布极端不均**的 MoE 模型（Qwen3.8-Flash-Next 走 top-10 路由就是典型）效果尤其好。简单说：在 IQ2_XS 这个量级下，GSQ-RCO 比 vanilla GGUF IQ2_XS 质量更高、速度损失更小。\n\n### 与 Strata 的搭配\n\nGSQ-RCO GGUF 跟 **Strata 0.1.38+ 完全兼容**——Strata 直接消费 GGUF 文件，GSQ-RCO 不需要额外 patch。下载 GSQ-RCO IQ2_XS（45GB）后即可用 Strata 跑起来：\n\n```\nstrata serve --model Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS.gguf \\\n  --gpu-layers 12 --ctx 4096 --batch 1\n```\n\n### 选档建议\n\n- **第一次试**：IQ2_XS（45GB）——64GB RAM + 12GB 显卡的甜点\n- **写代码为主**：Coder IQ1_M（32GB）——48GB RAM 起步，速度反而最快（因为 sparse 路由更激进）\n- **质量优先**：IQ3_XXS（76GB）——96GB RAM，能跟 fp16 抗衡\n- **极致质量**：IQ3_S（83GB）——128GB RAM，基本等同于 bfloat16\n\n### 本地部署角度\n\n下载链接直接走 HuggingFace，仓库里有 SHA256 校验。建议先下 IQ2_XS 跑通流程，再按需换档。","### What it is\n\n[ISTA-DASLab](https:\u002F\u002Fhuggingface.co\u002FISTA-DASLab) has published the **GSQ-RCO GGUF** quantization family for Qwen3.8-Flash-Next on HuggingFace, covering everything from consumer workstations to memory-rich servers:\n\n| Quant | Weight size | Min RAM | Recommended scenario |\n|---|---|---|---|\n| Coder (IQ1_M) | ~32 GB | 48 GB | Code-specialized, speed first |\n| IQ2_XS | ~45 GB | 64 GB | **Recommended** — speed\u002Fquality balance |\n| IQ3_XXS | ~76 GB | 96 GB | Quality-sensitive |\n| IQ3_S | ~83 GB | 128 GB | Close to full-model quality |\n| Q2_0 | ~66 GB | 96 GB | Q2_K style, CPU-friendly |\n\n### Why GSQ-RCO instead of vanilla GGUF\n\nGSQ is ISTA-DASLab's own **Grouped Scaling Quantization** algorithm; combined with RCO (Row-Column Ordering), it works especially well on **MoE models with extreme activation skew** (Qwen3.8-Flash-Next's top-10 routing is a textbook case). In short: at the IQ2_XS tier, GSQ-RCO has higher quality and less speed loss than vanilla GGUF IQ2_XS.\n\n### Pairing with Strata\n\nGSQ-RCO GGUF is **fully compatible with Strata 0.1.38+** — Strata consumes GGUF directly; GSQ-RCO needs no extra patch. Download GSQ-RCO IQ2_XS (45GB) and you're good to go with Strata:\n\n```\nstrata serve --model Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS.gguf \\\n  --gpu-layers 12 --ctx 4096 --batch 1\n```\n\n### Pick-by-scenario guidance\n\n- **First try**: IQ2_XS (45GB) — sweet spot for 64GB RAM + 12GB GPU\n- **Code-heavy workload**: Coder IQ1_M (32GB) — starts at 48GB RAM, actually fastest (more aggressive sparse routing)\n- **Quality priority**: IQ3_XXS (76GB) — 96GB RAM, holds its own against fp16\n- **Maximum quality**: IQ3_S (83GB) — 128GB RAM, essentially bfloat16-grade\n\n### Local deployment angle\n\nDownload straight from HuggingFace; the repo includes SHA256 checksums. Recommended path: grab IQ2_XS first to validate the pipeline, then swap quant tiers as needed.","HuggingFace ISTA-DASLab","https:\u002F\u002Fhuggingface.co\u002FISTA-DASLab\u002FQwen3.8-Flash-Next-GSQ-RCO-GGUF","2026-10-05T02:23:19.806Z"]