【文章标题】:Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at 12 tok/s
【文章标题】:技术分享:在48GB内存Mac上以12 token/秒速度运行104GB版Qwen3.8-Flash-Next
【文章正文】: Run Qwen3.8-Flash-Next on a Mac that cannot hold it. The model is 104 GB at 4-bit; slotstream streams it from SSD and runs it in whatever memory you give it, down to an 8.1 GB planned floor. One Swift binary with the commonly used Ollama and OpenAI chat/generate endpoints. 在内存不足的Mac上运行Qwen3.8-Flash-Next。这个4位量化的模型有104GB;slotstream工具会从SSD流式加载,并能在你分配的任何内存中运行(最低支持到8.1GB)。该Swift二进制程序内置了常用的Ollama和OpenAI聊天/生成接口。
| on a 48 GB Mac | |
|---|---|
| Warm decode | ~12 tok/s |
| Cold start to first token | ~3 s |
| Peak memory | 32 GB (auto-sized; you can cap it) |
| Weights on disk | 104 GB |
| 48GB内存Mac实测 | |
|---|---|
| 热解码速度 | ~12 token/秒 |
| 冷启动到首token | ~3秒 |
| 峰值内存占用 | 32GB(自动调节,可设上限) |
| 磁盘权重文件 | 104GB |
Disk is the gate that bites first. You need ~110 GB free, so a 512 GB Mac is the realistic minimum however much memory it has. The weights are a one-time 104 GB download: well under an hour on a fast connection, several hours on a slow one (table below). 磁盘空间是首要门槛。需要约110GB空闲空间,因此512GB存储的Mac是实际最低要求(无论内存大小)。权重文件只需一次性下载104GB:快速网络下不到1小时,慢速网络需数小时(见下表)。
| memory | expect |
|---|---|
| 8 GB | below the 8.1 GB floor; doctor warns that it will page |
| 16 GB | ~5 tok/s estimated |
| 24 GB | ~8 tok/s estimated |
| 32 GB | ~10 tok/s estimated |
| 48 GB and up | ~12 tok/s — and auto stops at 33 GB here, so the rest of the machine stays yours |
| 内存 | 预期表现 |
|---|---|
| 8GB | 低于8.1GB下限,工具会警告将使用分页 |
| 16GB | 约5 token/秒 |
| 24GB | 约8 token/秒 |
| 32GB | 约10 token/秒 |
| 48GB及以上 | 约12 token/秒(自动限制在33GB内存占用,其余内存仍可供系统使用) |
Only the 48 GB row is measured on real hardware; the rest come from the same measured curve, and smaller Macs also have slower SSDs. Run slotstream doctor to see what your machine would get, and whether you have the disk for the weights, before downloading anything. 仅48GB行是实体机实测数据,其余数据来自同一测量曲线(且小内存Mac通常配备较慢SSD)。下载前可运行slotstream doctor命令检测设备性能和磁盘空间。
curl -fsSL https://raw.githubusercontent.com/carloslfu/slotstream/main/install.sh | sh
Installs a prebuilt binary to /.slotstream/bin and puts it on your PATH.
通过命令行安装预编译二进制文件到/.slotstream/bin目录并添加至PATH环境变量:
curl -fsSL https://raw.githubusercontent.com/carloslfu/slotstream/main/install.sh | sh
Needs Apple Silicon and macOS 14+. Re-run the same line to upgrade; uninstall with rm -rf ~/.slotstream. 要求Apple Silicon芯片和macOS 14+系统。升级只需重复执行同一命令,卸载使用rm -rf ~/.slotstream。
Releases are built by CI from the tagged commit with signed provenance, so you can check an asset yourself rather than trusting the download: gh attestation verify slotstream-arm64.tar.gz —repo carloslfu/slotstream 所有发布版本均由CI系统从标签提交构建并附带签名凭证,支持自主验证: gh attestation verify slotstream-arm64.tar.gz —repo carloslfu/slotstream
Or build it yourself — Command Line Tools are enough, no Xcode needed: git clone https://github.com/carloslfu/slotstream && cd slotstream make build 也可自行编译(仅需命令行工具,无需Xcode): git clone https://github.com/carloslfu/slotstream && cd slotstream make build
The binary is small; the weights are not. 103.8 GB across 24 files, one time. 二进制文件很小,但权重文件达103.8GB(含24个文件,一次性下载)。
serve and run offer the download on first run, and slotstream pull does it on its own: slotstream serve 首次运行serve/run命令时会触发下载,也可单独使用slotstream pull命令: slotstream serve
Either way it prints the size, the destination and your free disk and waits for a yes before transferring anything, and it refuses outright if the disk cannot hold it. 两种方式都会显示文件大小、目标路径和剩余空间,需确认后才开始传输,若磁盘空间不足会直接拒绝。
Hugging Face is the bottleneck, not your link. Past four connections it plateaus: 4, 8, 16 and 32 all landed in the same 36 to 57 MB/s band, and so did hf_xet, Hugging Face’s own fastest client, while the same link did 134 MB/s to an ordinary host. So past roughly 400 Mbps, more bandwidth buys nothing: Hugging Face是下载瓶颈而非网络带宽。超过4个连接后速度稳定在36-57MB/s区间(即使使用HF官方最快客户端hf_xet),而相同网络对其他主机可达134MB/s。因此带宽超过400Mbps后不再提升速度:
| your connection | wait |
|---|---|
| 400 Mbps or faster | 30–50 min — Hugging Face’s day, not your link |
| 200 Mbps | ~1 h 10 |
| 100 Mbps | ~2 h 20 |
| 50 Mbps | ~4 h 40 |
| 25 Mbps | ~9 h |
| 网络带宽 | 预计耗时 |
|---|---|
| 400Mbps以上 | 30-50分钟(受限于HF服务器) |
| 200Mbps | 约1小时10分 |
| 100Mbps | 约2小时20分 |
| 50Mbps | 约4小时40分 |
| 25Mbps | 约9小时 |
A real install here took 35 min; the top row is wide because Hugging Face’s own throughput moved between sessions. The rows below it are arithmetic over 103.8 GB at your full rated speed, so treat them as best cases. 实测安装耗时35分钟。首行区间较大是因为HF服务器吞吐波动,其余行是按103.8GB全速下载的理论值(视为最佳情况)。
Interrupting is safe: it resumes at the exact byte it stopped on, and all 24 files are checked against sha256 hashes compiled into the binary, so a truncated, same-size, or corrupted download cannot reach the engine. 支持断点续传(精准到字节),且所有24个文件都会与二进制内编译的sha256哈希校验,确保残缺/等大/损坏的下载文件不会被加载。
pull —verify re-hashes an existing copy in under 10 s — 7.7 s here, hashed in parallel. pull —verify可在10秒内完成现有文件的哈希复核(实测7.7秒,并行计算)。
serve listens on port 11434 and implements the chat/generate subset used by Ollama clients and OpenAI SDKs: serve命令监听11434端口,实现Ollama客户端和OpenAI SDK常用的聊天/生成子集:
curl localhost:11434/api/chat -d ’{ “model”: “qwen3.8-flash-next:4bit”, “messages”: [{“role”: “user”, “content”: “hello”}] }’
OLLAMA_HOST=http://localhost:11434 ollama run qwen3.8-flash-next:4bit
Open WebUI, the Ollama CLI, and the OpenAI SDKs are tested for this subset. Streaming, CORS, and the usual sampling options (temperature, top_p, top_k, min_p, presence_penalty, seed, num_predict, stop) are all supported. Unsupported semantics such as tools, images, JSON-schema output, logprobs, and alternate model names return a clear 400 instead of being silently ignored. 已测试兼容Open WebUI、Ollama CLI和OpenAI SDK。支持流式传输、CORS和常规采样参数(温度/top_p/top_k/min_p/存在惩罚/随机种子/预测数/停止符)。对工具调用、图像、JSON模式输出、对数概率和替代模型名等未实现功能会明确返回400错误而非静默忽略。
Follow-up turns in a conversation only prefill what is new, so time to first token stays flat as a chat grows — measured over eight turns, 6.0 s instead of climbing to 25.8 s. One consequence worth knowing: reusing that state is not bit-identical to recomputing it, so a reply can occasionally differ where two tokens were nearly tied. —no-prefix-cache turns it off if you need exact reproducibility. 对话后续轮次仅预填充新增内容,使得首token响应时间不随对话增长而增加(实测8轮对话保持6.0秒,而非升至25.8秒)。需注意:状态复用与重新计算存在位级差异,当两个token概率相近时可能产生不同回复。如需完全确定性可启用—no-prefix-cache关闭该优化。
Prompt plus completion is capped at 32,768 tokens (—max-context). Long prompts are the slow axis: prefill runs at roughly 50 tok/s on a 16 GB Mac and 125 on a 48 GB one, so an 8,000-token prompt waits somewhere between about a minute and about three before its first token. A per-user lock enforces one model process at a time. 提示词+补全上限32,768 token(—max-context)。长提示词是性能瓶颈:16GB Mac预填充约50 token/秒,48GB机型约125 token/秒,因此8,000 token的提示词需要等待1-3分钟才出首token。单用户锁机制确保同一时间只运行一个模型进程。
With no flags slotstream sizes itself to your machine and tells you what it chose. This is a 48 GB Mac — it reads 52 GB because everything here counts in decimal GB, while Apple markets the same memory as 48: 无参数运行时slotstream会自动适配设备配置并告知选择。这是在48GB内存Mac上的示例(显示52GB是因为采用十进制GB计算,而苹果宣传使用的是48GB二进制标准): slot