Abliterlitics Audited 8 'Uncensored' Qwen 3.8 27B Variants. The Heaviest Edit Lost
An 11-day audit of eight uncensored Qwen 3.8 27B fine-tunes found the smallest edits unlock most, and one model hides a jailbreak in its chat template.

Refusal-removal fine-tunes are usually measured by whoever uploaded them, which is why independent audits are rare. On September 6, the Abliterlitics project posted one covering eight abliterated variants of the Qwen 3.8 27B reasoning model, plus the base, over 11 days and roughly 167 GPU hours on a single RTX 5090. The thread drew roughly 500 points and more than 150 comments on r/LocalLLaMA; its headline finding upends the genre: the loudest, most aggressive edits lost.
What Happened
Every arm ran identically in dynamic FP8 through four axes: tensor-by-tensor weight comparison, KL divergence, a 13-task benchmark suite, and HarmBench with 400 harmful behaviours at xhigh effort with a 15,360-token thinking budget. An LLM judge read all 3,600 responses, thinking included, and sorted each into complied, deflected, refused, or broken. The full report is at abliterlitics.dev. Ranked by judge-assessed attack success rate, best to worst:
- orcarouter/Qwen3.8-27B-Uncensored — 82.2%: one Arditi-style direction edit at layer 38 (131 matrices); the only card whose claims all verified; best copyright unlock at 39%
- apostate (heterodoxin) — 78.7%: KCRN method, 41 real edits, lowest KL; text-only re-save without vision or MTP, stored FP16
- huihui-ai — 75.6%: the classic method, clean unlock except a copyright wall near 3%
- ultra_heretic (llmfan46) — 70.5%: Heretic v2 with MPOA; heaviest truthfulness drop outside obliteratus; 118 soft refusals
- coder3101 — 70.0%: vanilla Heretic; card claims 33 of 100 refusals, judge counted 5 in 400, "the card undersells it"
- blackfrost — 68.5%: closed method; weights show one edited direction, not the claimed "direction bank"; jailbreak prompt hidden in its chat template
- obliteratus — 63.9%, avoid: most aggressive edit, 841 of 850 tensors; 44.8% of responses never finish thinking
- trohrbaugh (heretic-ara) — 57.5%: refuses 122 of 400, cleanest capability profile; "the one I use at home"
- Qwen/Qwen3.8-27B (base) — 4.5%: "a wall"; zero compliance on chem/bio, harassment, harmful content, copyright
Why This Matters
Surgical beats heavy: the top two spots went to the smallest verified edits, the heaviest landed second-to-last. As the tester puts it, "at 27B, editing everything mostly buys you a model that thinks in circles." The thinking-loop finding is new here: aggressive arms leave up to 45% of HarmBench responses stuck in an unclosed think block at the 15,360-token budget, yet the same arms finish school-math problems within 1.2 points of base. "School math converges, adversarial deliberation does not."
Two walls moved. Copyright is the new universal wall, nobody exceeds 39% and five of nine sit at or below 3.2%, while chem and bio, historically the hardest category, is now the easiest unlock. Chat-template forensics also mattered for the first time: blackfrost ships a 1,457-character jailbreak system prompt inside its bundled template, injected into every conversation, a supply-chain risk for anyone self-hosting these weights.
What's Next
Model cards proved unreliable in both directions, from orcarouter's verified claims to obliteratus's unreplicated "zero refusals" framing. Abliterlitics open-sources its toolkit and is building template-aware measurement. Until then, the takeaway: check the chat template before self-hosting, and treat card claims as hypotheses.
For most users, the base model, an Apache-2.0 reasoning model that runs on consumer hardware, remains the more interesting story; we covered running it at 100K context last week.
