【帖子标题】:Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release

【帖子标题】:为Qwen 3.5、3.6及全新3.8版本修复的Jinja聊天模板

【帖子正文】: Qwen just released their first 3.8 model.

【帖子正文】: Qwen刚刚发布了他们的首个3.8模型。

The main addition in 3.8 is prompt-steered reasoning effort. You can tell the model how deeply to think by setting

3.8版本的主要新增功能是提示引导的推理深度控制。你可以通过设置以下参数来告诉模型思考的深度:

reasoning_effort

reasoning_effort

to

为

xhigh

xhigh

,

、

medium

medium

, or

或

low

low

.

。

However, the official template still has some serious problems:

然而,官方模板仍然存在一些严重问题:

You cannot disable thinking.

你无法禁用思考功能。

If you pass

如果你传入

enable_thinking=false

enable_thinking=false

, it 3.8 crashes with a hard exception.

,3.8版本会直接抛出硬异常而崩溃。

Chat history gets poisoned.

聊天历史记录会被污染。

In multi-turn chats, the official template injects blank

在多轮对话中,官方模板会在真实思考内容之前注入空白的

thinking response

思考 回应

tags before real thoughts.

标签。

Tool calling crashes.

工具调用会崩溃。

If your client passes arguments as JSON strings (the standard OpenAI API format), the official template crashes.

如果你的客户端以JSON字符串形式(标准的OpenAI API格式)传递参数,官方模板就会崩溃。

Agent stalls.

智能体会停滞。

The official template often drops mid-dialogue system messages and wedges multi-step tool loops.

官方模板经常丢弃对话中途的系统消息,并卡住多步骤工具循环。

I maintain a single, drop-in fixed Jinja template that works across all Qwen 3.5, 3.6, and 3.8 models:

我维护了一个可直接替换使用的修复版Jinja模板,适用于所有Qwen 3.5、3.6和3.8模型:

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

What this template does:

该模板的功能如下:

Full 3.8 reasoning effort support:

完整的3.8推理深度支持:

Steer reasoning depth with

通过

reasoning_effort

reasoning_effort

(

(

xhigh

xhigh

,

、

high

high

,

、

low

low

,

、

medium

medium

).

)来控制推理深度。

Restores the thinking toggle:

恢复思考开关功能:

Turn off reasoning whenever you want fast answers, either via kwargs or by typing

无论何时你想要快速回答,都可以通过kwargs参数或输入

<|think_off|>

<|think_off|>

in your prompt.

来关闭推理。

100% KV Cache hits:

100% KV缓存命中:

Keeps past thoughts intact by default so your prefix cache stays warm across turns.

默认保留过去的思考内容完整,使你的前缀缓存在多轮对话中保持热度。

llama.cpp support:

llama.cpp支持:

Native support for the new

原生支持新的

--reasoning-preserve

--reasoning-preserve

flag.

标志。

Universal tool parsing:

通用工具解析:

Handles both Python dicts and JSON strings. Works on llama.cpp, vLLM, LM Studio, and MLX.

同时支持Python字典和JSON字符串。可在llama.cpp、vLLM、LM Studio和MLX上运行。

Recommended llama-server launch command:

推荐的llama-server启动命令:

llama-server -m your_model.gguf —jinja —chat-template-file chat_template.jinja —reasoning-format deepseek

llama-server -m your_model.gguf —jinja —chat-template-file chat_template.jinja —reasoning-format deepseek

(The

(

--reasoning-format deepseek

--reasoning-format deepseek

flag separates thinking into the OpenAI

标志会将思考内容分离到OpenAI的

reasoning_content

reasoning_content

field so OpenCode, Claude Code, and other harnesses do not stall on raw tokens).

字段中,这样OpenCode、Claude Code以及其他工具框架就不会因为原始token而停滞)。

Note on hardware:

关于硬件的说明:

I cannot run a 2.4 trillion parameter model on my local rig. The template passes all 28 automated tests and tokenizer parity checks, but I would appreciate feedback from anyone testing it with Qwen 3.8.

我无法在本地设备上运行一个2.4万亿参数的模型。该模板通过了全部28项自动化测试和分词器一致性检查,但我希望能得到任何使用Qwen 3.8进行测试的人的反馈。

submitted by

由

/u/ex-arman68

/u/ex-arman68

[link]

[链接]

[comments]

[评论]