【帖子标题】:Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release
【帖子标题】:为Qwen 3.5、3.6及全新3.8版本修复的Jinja聊天模板
【帖子正文】: Qwen just released their first 3.8 model.
【帖子正文】: Qwen刚刚发布了他们的首个3.8模型。
The main addition in 3.8 is prompt-steered reasoning effort. You can tell the model how deeply to think by setting
3.8版本的主要新增功能是提示引导的推理深度控制。你可以通过设置以下参数来告诉模型思考的深度:
reasoning_effort
reasoning_effort
to
为
xhigh
xhigh
,
、
medium
medium
, or
或
low
low
.
。
However, the official template still has some serious problems:
然而,官方模板仍然存在一些严重问题:
You cannot disable thinking.
你无法禁用思考功能。
If you pass
如果你传入
enable_thinking=false
enable_thinking=false
, it 3.8 crashes with a hard exception.
,3.8版本会直接抛出硬异常而崩溃。
Chat history gets poisoned.
聊天历史记录会被污染。
In multi-turn chats, the official template injects blank
在多轮对话中,官方模板会在真实思考内容之前注入空白的
thinking response
思考 回应
tags before real thoughts.
标签。
Tool calling crashes.
工具调用会崩溃。
If your client passes arguments as JSON strings (the standard OpenAI API format), the official template crashes.
如果你的客户端以JSON字符串形式(标准的OpenAI API格式)传递参数,官方模板就会崩溃。
Agent stalls.
智能体会停滞。
The official template often drops mid-dialogue system messages and wedges multi-step tool loops.
官方模板经常丢弃对话中途的系统消息,并卡住多步骤工具循环。
I maintain a single, drop-in fixed Jinja template that works across all Qwen 3.5, 3.6, and 3.8 models:
我维护了一个可直接替换使用的修复版Jinja模板,适用于所有Qwen 3.5、3.6和3.8模型:
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
What this template does:
该模板的功能如下:
Full 3.8 reasoning effort support:
完整的3.8推理深度支持:
Steer reasoning depth with
通过
reasoning_effort
reasoning_effort
(
(
xhigh
xhigh
,
、
high
high
,
、
low
low
,
、
medium
medium
).
)来控制推理深度。
Restores the thinking toggle:
恢复思考开关功能:
Turn off reasoning whenever you want fast answers, either via kwargs or by typing
无论何时你想要快速回答,都可以通过kwargs参数或输入
<|think_off|>
<|think_off|>
in your prompt.
来关闭推理。
100% KV Cache hits:
100% KV缓存命中:
Keeps past thoughts intact by default so your prefix cache stays warm across turns.
默认保留过去的思考内容完整,使你的前缀缓存在多轮对话中保持热度。
llama.cpp support:
llama.cpp支持:
Native support for the new
原生支持新的
--reasoning-preserve
--reasoning-preserve
flag.
标志。
Universal tool parsing:
通用工具解析:
Handles both Python dicts and JSON strings. Works on llama.cpp, vLLM, LM Studio, and MLX.
同时支持Python字典和JSON字符串。可在llama.cpp、vLLM、LM Studio和MLX上运行。
Recommended llama-server launch command:
推荐的llama-server启动命令:
llama-server -m your_model.gguf —jinja —chat-template-file chat_template.jinja —reasoning-format deepseek
llama-server -m your_model.gguf —jinja —chat-template-file chat_template.jinja —reasoning-format deepseek
(The
(
--reasoning-format deepseek
--reasoning-format deepseek
flag separates thinking into the OpenAI
标志会将思考内容分离到OpenAI的
reasoning_content
reasoning_content
field so OpenCode, Claude Code, and other harnesses do not stall on raw tokens).
字段中,这样OpenCode、Claude Code以及其他工具框架就不会因为原始token而停滞)。
Note on hardware:
关于硬件的说明:
I cannot run a 2.4 trillion parameter model on my local rig. The template passes all 28 automated tests and tokenizer parity checks, but I would appreciate feedback from anyone testing it with Qwen 3.8.
我无法在本地设备上运行一个2.4万亿参数的模型。该模板通过了全部28项自动化测试和分词器一致性检查,但我希望能得到任何使用Qwen 3.8进行测试的人的反馈。
submitted by
由
/u/ex-arman68
/u/ex-arman68
[link]
[链接]
[comments]
[评论]