【文章标题】:My local model setup on an M4 Pro Mac Mini
【文章标题】:我的M4 Pro Mac Mini本地模型配置方案
【文章正文】:
I run a local LLM server on my M4 Pro Mac mini with 48 GB of RAM. It handles everything from my Hermes agent backend to quick chat queries on my phone. The whole thing takes about 30 minutes to set up.
我在配备48GB内存的M4 Pro Mac mini上运行本地LLM服务器。它处理从Hermes智能体后端到手机快捷聊天查询的所有任务,整套配置仅需30分钟即可完成。
Here is the stack:
技术栈如下:
- Qwen3.6-35B-A3B-OptiQ-4bit: my main model for anything that needs reasoning or depth
- Qwen3.6-35B-A3B-OptiQ-4bit:用于需要深度推理任务的主模型
- Gemma-4-E4B-it-OptiQ-4bit: lightweight model for simple chats, formatting, and other routine tasks
- Gemma-4-E4B-it-OptiQ-4bit:处理简单对话、文本格式化等日常任务的轻量模型
- oMLX: the inference server
- oMLX:推理服务器
- Tailscale: tailnet connecting the Mac mini, my iPhone, and my MacBook
- Tailscale:连接Mac mini、iPhone和MacBook的专属网络
Hermes runs as the agent backend on the Mac mini, with my MacBook running the desktop client and my phone running Telegram. For non-Hermes usage I use Apollo on iOS for quick chats (reads like Claude, good for throwaway questions), Pi as my coding agent (I already wrote about that setup), and Raycast AI on my Mac for random things.
Hermes作为智能体后端运行在Mac mini上,MacBook运行桌面客户端,手机通过Telegram接入。非Hermes场景下:iOS使用Apollo进行快捷对话(类Claude体验,适合临时问题)、Pi作为编程助手(此前撰文介绍过配置)、Mac端通过Raycast AI处理杂项任务。
Why bother?
为何大费周章?
The main reason to run local: cloud APIs are rented land. They can change their pricing, hit your usage limits, or swap the model being served behind the scenes whenever they feel like it. I was regularly maxing out two $200/month subscriptions and it felt like I was getting different things from them at different points. Sometimes a model was fine, sometimes it degraded with no notice.
本地化核心原因:云API如同租赁土地。服务商可随时调整定价、触发用量限制或暗中切换模型。我曾每月耗尽两个200美元订阅额度,却常遭遇服务波动——模型时而正常时而无故降级。
Data privacy is another issue. You do not know what these companies do with your data once they have it. They might limit how it gets used, they might sell it, they might expose it. Either way, it creates an operational security risk. If you work with sensitive code, client data, or proprietary workflows, sending it to a third-party API is a decision you make once and cannot undo.
数据隐私是另一关键。企业获取数据后的处理方式无从知晓——可能限制使用、转售或泄露。对于敏感代码、客户数据或专有工作流,第三方API传输如同不可逆的单向阀,必然带来运营安全风险。
Then there is AI sovereignty. I have been watching how the US government has limited the rollout of various models. That can happen at any point from any government, for any reason, and you have no control over it. If your workflow depends on a cloud model that gets restricted, you have to stop or scramble. The only way to avoid that is to own your compute.
AI主权同样重要。美国政府已多次限制模型部署,任何政府都可能在任意时刻以任何理由实施管制。若工作流依赖的云端模型突遭封禁,只能被迫中断或仓促调整。唯有掌握本地算力才能规避风险。
Other practical advantages:
其他实际优势:
-
Cost predictability. APIs are variable. Your usage spikes and your bill follows. With local hardware, the cost is the hardware purchase plus electricity. Flat. After that, every inference is free.
-
成本可控:云API费用随用量波动,本地硬件仅需一次性购置成本加固定电费,后续推理完全免费
-
Latency. No network roundtrip means faster responses for everyday tasks. The M4 Pro’s media engine handles inference at speeds that feel instant for most prompts.
-
低延迟:省去网络往返,日常任务响应更快。M4 Pro媒体引擎的推理速度对多数提示词而言近乎瞬时
-
Offline capability. No internet, still works. For agent workflows that run in the background, this matters more than it sounds.
-
离线能力:无网络仍可运行,对后台智能体工作流至关重要
-
No rate limits. API providers throttle you when you hit usage thresholds. Your own machine does not care how much you run.
-
无速率限制:摆脱API提供商的用量阈值节流,本地硬件可无限调用
How I actually use it
实际使用场景
The Mac mini is always on. It sits on my desk and I barely notice it except when I need it.
Mac mini保持常开状态,静置桌面几乎无感,仅在需要时唤醒。
Hermes runs on the Mac mini as well, using a local model on the same machine. I access my agent through Telegram (on my phone) and the Hermes desktop app on my MacBook. The Hermes desktop app acts as a ‘shell’ and connects to a Hermes backend on another device (in this case the Mac mini). This means I share a backend, conversation history, and skillset across all my devices.
Hermes后端与本地模型共同运行于Mac mini,通过手机Telegram和MacBook桌面客户端接入。桌面应用作为”外壳”连接至Mac mini的后端,实现跨设备共享后台服务、对话历史和技能集。
Then there is everything else:
其他应用场景:
-
Apollo on iOS for quick throwaway chats. I want something that reads like Claude but does not require an API key or a subscription. Connect Apollo to http://[mac-mini-tailnet-url]/v1 and you are done. Good for “rewrite this paragraph” or “what does this error mean” type questions.
-
iOS端Apollo处理临时对话:无需API密钥或订阅的类Claude体验,直连Mac mini的Tailscale地址即可处理”重写段落”或”错误解读”类请求
-
Raycast also on my Mac for random things I don’t want to install anything for.
-
Mac端Raycast处理零散需求:免安装轻量解决方案
-
Pi for coding. Already wrote about that setup.
-
Pi编程助手:此前已详述配置方案
The point is not to replace API-based models. It is to handle the 80% of requests that do not need GPT-5 or Claude Opus. And when I do need those, they are already available. Local just covers more of my day-to-day for free.
核心并非取代API模型,而是处理80%无需GPT-5或Claude Opus的请求。当需要顶级模型时仍可随时调用,本地方案免费覆盖了更多日常需求。
The model breakdown
模型解析
Running a large model locally comes down to one thing: how much RAM it actually needs in memory. Most people look at the parameter count and get the wrong idea, because the difference between dense and mixture-of-experts (MoE) models matters a lot on consumer hardware.
本地运行大模型的关键在于实际内存占用。多数人误解参数数量,实则稠密模型与专家混合模型(MoE)在消费级硬件上差异显著。
Here is how to read the identifier:
模型标识解读:
Qwen3.6-35B-A3B-OptiQ-4bit
- Qwen3.6 : model family and version
- Qwen3.6:模型系列及版本
- 35B : total parameters across all experts
- 35B:全体专家参数总量
- A3B : active parameters per token (3 billion, not 35)
- A3B:单token激活参数(30亿而非350亿)
- OptiQ-4bit : mixed-precision quantization (4-bit mostly, 8-bit on sensitive layers)
- OptiQ-4bit:混合精度量化(主体4-bit,敏感层8-bit)
gemma-4-e4b-it-4bit
- gemma-4 : Google’s Gemma 4 family
- gemma-4:谷歌Gemma 4系列
- e4b : encoding size, roughly 4 billion parameters total
- e4b:编码规模,约40亿参数总量
- it : instruction-tuned
- it:指令微调版
- 4bit : uniform 4-bit quantization
- 4bit:统一4-bit量化
The key difference is the A3B part. A dense 27B model has 27 billion parameters loaded in RAM at all times, for every single token. An MoE model like the Qwen3.6-35B-A3B has 35 billion total parameters spread across 256 experts, but only about 3 billion are actually activated per token. The other 32 billion sit in RAM doing nothing.
关键差异在于A3B标识:稠密27B模型需常驻270亿参数,而Qwen3.6-35B-A3B这类MoE模型虽总参数量350亿(分布在256个专家中),但单token仅激活约30亿参数,其余320亿参数闲置内存。
On my 48GB Mac mini, the Qwen3.6-35B-A3B in 4-bit takes
在我的48GB内存Mac mini上,4-bit量化的Qwen3.6-35B-A3B模型占用