【帖子标题】:Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀 【帖子标题】:Qwen3.8-Flash-Next。一旦权重发布,这种架构可能会对本地部署出奇地友好。 👀
【帖子正文】: Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate: 【帖子正文】: Qwen3.8-Flash-Next(约125B-A6B + 51B n-gram)内存估算:
Ideal 4-bit quant ≈ 82 GB 理想的4位量化 ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables) (58 GB主权重 + 24 GB n-gram表)
Real-world quants likely land in the 80–90 GB range. 实际量化后的显存占用很可能在80–90 GB范围内。
The big n-gram table is sparsely accessed → excellent candidate for system RAM offload. 庞大的n-gram表访问频率较低 → 非常适合分流至系统内存。
This architecture could be surprisingly local-friendly once the weights drop. 一旦权重发布,这种架构可能会对本地部署出奇地友好。