【帖子标题】:8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics
【帖子标题】:8个未审查版Qwen 3.8 27B变体对比:1个基础模型,消耗167 GPU小时——Abliterlitics分析

【帖子正文】:
This comparison was requested by a few people, and certainly we were all eager to see the final results. The comparison had taken 11 days and the GPU was crunching numbers for ~167 hours.
本次对比应多位研究者要求开展,我们团队也对最终结果充满期待。整个测试耗时11天,GPU累计运算约167小时。

We’ve been comparing different abliterated models from huggingface to see if they really are what they claim to be. So far the results have been interesting.
我们系统对比了HuggingFace上多个经过”消除处理”的模型,验证其是否名副其实。目前发现的现象相当耐人寻味。

The pipeline includes a weight comparison, KL divergence measurement, 13 benchmarks and measuring refusals with the HarmBench 400 classic. Qwen 3.8 27b is a thinker, and with ourselves using xhigh we had to set our token budget much higher per request.
测试流程包含权重对比、KL散度测量、13项基准测试,以及使用经典HarmBench 400进行的拒绝率测定。Qwen 3.8 27b属于思考型模型,由于我们采用xhigh配置,必须为每个请求分配更高的token预算。

Lets check out how Qwen 3.8 27b stacks up comparing 8 variants.
现在让我们看看Qwen 3.8 27b的8个变体表现如何。

Full report:
完整报告:
abliterlitics.dev/models/qwen38-27b

The rankings
排名情况

LLM Judge HarmBench ASR (attack success rate), best to worst, with the one-line story:
基于LLM Judge HarmBench攻击成功率(ASR)的排名(从优到劣)及简评:

orcarouter
82.2%, the winner. Arditi-style single direction at layer 38, 131 matrices, and the only card where every claim checked out against the weights. Best copyright unlock in the set at 39%
冠军orcarouter:82.2%成功率。采用Arditi式单方向修改(第38层,131个矩阵),是唯一权重与声明完全吻合的变体。39%的版权绕过率为本组最佳

apostate
78.7%, best value. Their new KCRN method, 41 real edits, lowest KL measured at 0.0439, near-identity capabilities. Packaging quirks: text-only re-save with no vision and no MTP, stored FP16
性价比之王apostate:78.7%成功率。采用新型KCRN方法(41处实质编辑),0.0439的KL散度最低,性能近乎原版。注意:仅保留文本功能(无视觉/MTP),FP16存储格式

huihui
75.6%, the classic method, reliable. Clean unlock everywhere except copyright, where it sits at 3%
经典之选huihui:75.6%成功率。传统方法制作,除版权项仅3%解锁率外,其余场景解锁干净

ultra_heretic
70.5%, Heretic v2 with MPOA. Works, but the heaviest truthfulness drop outside obliteratus and 118 soft refusals
ultra_heretic:70.5%成功率。MPOA加持的Heretic v2版本,但存在最严重的真实性下降问题(仅次于obliteratus),出现118次软拒绝

coder3101
70.0%, vanilla Heretic. The card calls itself the weakest removal at 33 of 100 refusals. Measured: 5 explicit refusals in 400. The card undersells it
coder3101:70.0%成功率。原始Heretic版本,虽自称”最弱消除版(100次测试33次拒绝)“,实际400次测试仅5次明确拒绝,表现优于声明

blackfrost
68.5%, closed method. The weights say single direction, heaviest magnitude in the panel, 100% rank-1, which refutes the rank-k direction bank story. Also ships a jailbreak system prompt inside its chat template, meaning every single prompt you make will have a modified chat template injecting a jailbreak
blackfrost:68.5%成功率。封闭方法制作,权重显示单方向修改(幅度全组最大),100%秩1数据推翻其宣称的秩k方向库。注意:聊天模板内置越狱指令,所有提示都会自动注入越狱代码

obliteratus
63.9%,
avoid
. The most aggressive edit in the panel at 841 of 850 tensors, and it performs like it. 44.8% of responses never finish thinking, and it is the only variant that got meaningfully dumber
obliteratus:63.9%成功率——建议避坑。850个张量中修改841个(全组最激进),44.8%响应无法完成思考,是唯一出现明显智力下降的变体

trohrbaugh
57.5%, last of the variants because it still refuses. 122 explicit refusals, the most surviving alignment of any variant, and the cleanest capability profile in the comparison. This is the one I use at home and it’s been great for me.
trohrbaugh:57.5%成功率。因122次明确拒绝(保留最多对齐特性)垫底,但能力谱最纯净。个人自用推荐款

base
4.5%, a wall. Zero compliance on chem and bio, harassment, harmful content and copyright
基础版:4.5%成功率。化学/生物、骚扰、有害内容及版权项完全无法绕过

The highlights
核心发现

Surgical beats heavy, again, and this time it is not close. The top two spots went to the two smallest verified edits. The heaviest edit of all landed second-to-last. At 27B, editing everything mostly buys you a model that thinks in circles
精准修改再次完胜粗暴改造,本次差距尤为显著。前两名都是经核实的微调版本,而修改量最大的变体位列倒数第二。对于27B模型,过度编辑只会导致思维循环

The thinking loop story is the big new finding for this model. Qwen 3.8 thinks before answering, and on the aggressive arms up to 45% of HarmBench responses never close their think block before the 15,360-token budget dies. The judge reads the full trace, so compliance inside a loop still counts. But a model that only delivers the goods inside an unterminated monologue is not a usable model
思维循环是本模型的重要新发现。Qwen 3.8会先思考再应答,激进版本中45%的HarmBench响应在15360 token预算耗尽前无法结束思考。虽然裁判会读取完整轨迹(循环内合规仍有效),但只会输出未终止独白的模型不具备实用性

GSM8K loops are gone at this budget. The same arms that loop 40%+ on HarmBench finish their math reasoning fine, every arm within 1.2pp of base on answered-only. School math converges, adversarial deliberation does not
GSM8K数学题在此预算下无循环现象。那些在HarmBench上40%+循环的变体,在数学推理时都能正常完成,各版本与基础版的答题差距仅1.2个百分点。校园数学能收敛,对抗性思考则不然

Copyright is the new universal wall. Nobody exceeds 39%, five of nine sit at or below 3.2%. Chem and bio, historically the hardest category, is now the easiest unlock. The walls moved
版权项成为新壁垒。无一变体超过39%解锁率,5/9版本≤3.2%。而传统难点化学/生物类反而成为最易突破项,防御格局已变

Chat template forensics was needed for the first time. blackfrost ships a 1457-character jailbreak prompt inside its template. obliteratus ships thinking-off. ultra_heretic deletes the stock reasoning-effort prompt. We pinned the stock template for every arm, because the template is a stronger behavioural lever than most people assume
首次需要进行聊天模板取证:blackfrost模板内置1457字符越狱指令;obliteratus关闭思考功能;ultra_heretic删除原始推理提示。我们固定了所有变体的原始模板,因其对行为的影响远超常人预期

Card honesty check
卡片真实性核查

orcarouter verified 4 of 4 claims exactly, the model card is honest.
orcarouter四项声明全部核实,模型卡片完全诚实

trohrbaugh’s KL calibrated within 9% of our measurement. However the card mentions 0/100 refusals, with our measurement this model had the most refusals, yet also had preserved capabilities.
trohrbaugh的KL散度与我们测量值偏差9%。但卡片称0/100拒绝率,实测却是拒绝最多的版本(同时保留了最佳能力)

The rest diverge, and none of it is dishonesty, KL is non-deterministic and moves with CUDA version and hardware, or by what method used. Read KL as a within-comparison spread
其余存在偏差但非虚假宣传,因KL散度具有非确定性(受CUDA版本/硬件/计算方法影响)。建议将KL值视作对比区间参考

Obliteratus published honest changes to how their model was changed, yet we were unable to replicate the ‘0% refusals, 0% deflection’ claims. 44.8% of harmbench never finished thinking, and reasoning analysis found 142 deflections.
Obliteratus如实公布了修改方法,但我们无法复现其”0%拒绝/0%回避”声明。44.8%的HarmBench响应未完成思考,推理分析发现142次回避

blackfrost - mentions they have internal direction bank, suggesting that multiple directions are changed. Yet we only found one direction changed. Their refusal numbers are accurate however, even with ourselves not using their jailbreak chat template we got similar results. The modified chat template also is not disclosed
blackfrost宣称使用多方向修改库,但我们仅发现单方向修改。其拒绝率数据准确(即使不使用越狱模板结果相似),但未披露修改后的聊天模板

🔗 知识库双向关联