Qwen3.5-9B · served locally on a single RTX 3060, 12GB VRAM

Benchmarks

Qwen3.5-9B, three ways: our abliterated baseline, an independent reproducibility rerun of that same baseline, and our run of the official weights — scored and timed against each other, and against Qwen's own published numbers where one exists.

queue idle · updated 10:10pm PDT
Abliterated¹ (ours)

huihui-ai abliterated build, served serially through the desktop adapter → production Ollama, 2026‑08‑16–20.

Abliterated² (ours, rerun)

Same abliteration, independently‑verified genuine GGUF, served via the 16‑way raw‑llama.cpp bypass, 2026‑08‑23 — 7 of 7 tasks done.

Qwen 3.5:9b (ours)

Unsloth official Qwen3.5‑9B weights, same harness/prompting, run 2026‑08‑22–23 (two tasks via the raw‑llama.cpp bypass).

Qwen3.5‑9B (posted)

Numbers from Qwen's own model card. Not run by us — quoted as-is.

Score comparison

One table per benchmark, all scores as percentages. Δ is the change versus the row above: abliterated¹→abliterated² on the rerun row, abliterated²→Qwen 3.5:9b on the ours row. The posted row isn't chained in — see the readout cards below for why that comparison needs its own caveats.

Score comparison chart
gsm8k
80.5
82.5
85.7 ± 1.0
not reported
humaneval_instruct
37.8 ± 3.8
61.6 ± 3.8
47.6 ± 3.9
not reported
mbpp_instruct_fixed
39.4 ± 2.2
48.2 ± 2.2
48.0 ± 2.2
not reported
bbh_cot_zeroshot
46.3 ± 0.5
45.6 ± 0.5
65.8 ± 0.5
not reported
mmlu_pro
61.6 ± 0.4
62.0 ± 0.4
64.7 ± 0.4
82.5
ifeval
66.1
64.7
58.8
91.5
cnn_dailymail
not reported
7.9
10.0
not reported

gsm8kexact_match · 1,319 samples

grade-school math word problems

SourceScoreΔ
Abliterated¹80.5strict 80.9
Abliterated² (rerun)82.5strict 82.6▲ 2.0
Qwen 3.5:9b (ours)85.7 ± 1.0▲ 3.2
Qwen3.5‑9B (posted)not reported1

humaneval_instructpass@1 · 164 samples

Python function generation from docstrings

SourceScoreΔ
Abliterated¹37.8 ± 3.8
Abliterated² (rerun)61.6 ± 3.8▲ 23.8
Qwen 3.5:9b (ours)47.6 ± 3.9▼ 14.0
Qwen3.5‑9B (posted)not reported1

mbpp_instruct_fixedpass@1 · 500 samples

basic Python programming problems

SourceScoreΔ
Abliterated¹39.4 ± 2.2
Abliterated² (rerun)48.2 ± 2.2▲ 8.8
Qwen 3.5:9b (ours)48.0 ± 2.2▼ 0.2
Qwen3.5‑9B (posted)not reported1

bbh_cot_zeroshotexact_match, flexible-extract · 6,511 samples

diverse reasoning — logic, algorithms, language

SourceScoreΔ
Abliterated¹46.3 ± 0.5
Abliterated² (rerun)45.6 ± 0.5▼ 0.7
Qwen 3.5:9b (ours)65.8 ± 0.5▲ 20.2
Qwen3.5‑9B (posted)not reported1

mmlu_proexact_match · 12,032 samples

broad academic knowledge, 57 subjects

SourceScoreΔ
Abliterated¹61.6 ± 0.4
Abliterated² (rerun)62.0 ± 0.4▲ 0.5
Qwen 3.5:9b (ours)64.7 ± 0.4▲ 2.7
Qwen3.5‑9B (posted)82.52

ifevalloose‑instruction acc · 541 samples

instruction-following & formatting compliance

SourceScoreΔ
Abliterated¹66.1strict‑p 51.0 · strict‑i 63.8 · loose‑p 54.2
Abliterated² (rerun)64.7strict‑p 48.8 · strict‑i 61.9 · loose‑p 52.7▼ 1.4
Qwen 3.5:9b (ours)58.8strict‑p 44.2–44.43 · strict‑i 57.4–57.7 · loose‑p 45.7▼ 5.9
Qwen3.5‑9B (posted)91.52

cnn_dailymailrouge · 11,490 samples

news article summarization, single ROUGE score

SourceScoreΔ
Abliterated¹not reported4
Abliterated² (rerun)7.9
Qwen 3.5:9b (ours)10.0▲ 2.2
Qwen3.5‑9B (posted)not reported1
  1. gsm8k, humaneval, mbpp, bbh_cot_zeroshot, and cnn_dailymail don't appear on Qwen3.5's own model card at all — the blank isn't a missed lookup, no public number exists.
  2. mmlu_pro and ifeval are the only two rows that are genuine name‑for‑name matches with a posted number — and even then, exact prompt version and scoring protocol aren't guaranteed identical. Treat the gap's size as directional, not exact.
  3. ifeval's strict submetrics show small run‑to‑run nondeterminism on identical cached inputs (four replays: strict‑prompt 43.99–44.36, strict‑inst 57.43–57.67). Loose metrics were stable. The spread is far smaller than the gaps in this table, but don't read these two columns to more than one decimal place.
  4. cnn_dailymail has no abliterated¹ baseline — it was queued 2026‑08‑23, after the six‑task suite above (and its abliterated¹ baseline run, 2026‑08‑16–20) were already locked in. It's the only gap in this suite tied directly to production use (digest-produce's actual job is summarization), added after the fact rather than run alongside the others.

Qwen 3.5:9b beats abliterated, 4 of 7

gsm8k, bbh_cot_zeroshot, mmlu_pro, cnn_dailymail favor Qwen 3.5:9b by 2.2–20.2 points. bbh_cot_zeroshot's +20.2 is the largest gap in the comparison.

3 of 7 tasks reverse the pattern

humaneval_instruct, mbpp_instruct_fixed, ifeval all score lower for Qwen 3.5:9b than abliterated, by 0.2–14.0 points. Don't generalize "Qwen 3.5:9b always wins" from the other 4.

Two confounds vs. the posted column

Qwen's posted numbers are almost certainly thinking‑mode (the model's default); ours are forced non‑thinking. Ours also run Q4_K_M quantized, not full precision. The gap to "posted" reflects both of those on top of any real capability difference.

Runtime & throughput

Per‑item speed for each run, wall‑clock runtime alongside it for reference. Not a clean speed benchmark — see the callout below before reading these as apples‑to‑apples.

Three different serving conditions, not one axis: abliterated¹ ran serially (one request at a time) through the desktop adapter, spread across 2026‑08‑16–20 with routine contention from scheduled jobs on the same machine — its rates run slow and somewhat noisy (it has no row at all for cnn_dailymail, added after this baseline was already locked in — see footnote 4). Qwen 3.5:9b ran serially for gsm8k/ifeval/humaneval_instruct/mbpp_instruct_fixed/cnn_dailymail, but through the 16‑way raw‑llama.cpp bypass for mmlu_pro/bbh_cot_zeroshot — two different axes inside the same column. Abliterated² runs through the 16‑way bypass for five of those six tasks, one continuous session, minimal contention — cnn_dailymail is the exception, run serially like its qwen row instead, so that pair is still apples‑to‑apples despite abliterated²'s bypass elsewhere. Only compare rows where both sides used the same serving mode.
Runtime comparison chart

Bars are scaled per task (each task's own longest bar is 100% width) — runtimes span two orders of magnitude across tasks, so a single shared scale would flatten the short ones to nothing.

gsm8k
2h44m
36m10s
2h19m
humaneval_instruct
24m39s
3m18s
14m26s
mbpp_instruct_fixed
2h21m
13m09s
1h02m
bbh_cot_zeroshot
39h28m
2h05m
1h44m
mmlu_pro
52h48m
5h54m
3h51m
ifeval
7h43m
33m51s
2h49m
cnn_dailymail
not reported
5h41m
6h43m

gsm8kexact_match · 1,319 samples

grade-school math word problems

SourceSpeedRuntime
Abliterated¹ (serial)6.00s/it†raw 7.46s/it2h44m
Abliterated² (bypass)1.65s/itfrom 7.46s/it serial36m10s
Qwen 3.5:9b (ours)6.32s/itserial2h19m

humaneval_instructpass@1 · 164 samples

Python function generation from docstrings

SourceSpeedRuntime
Abliterated¹ (serial)7.50s/it†raw 9.02s/it24m39s
Abliterated² (bypass)1.20s/itfrom 9.02s/it serial3m18s
Qwen 3.5:9b (ours)5.28s/itserial14m26s

mbpp_instruct_fixedpass@1 · 500 samples

basic Python programming problems

SourceSpeedRuntime
Abliterated¹ (serial)12.00s/it†raw 16.97s/it2h21m
Abliterated² (bypass)1.58s/itfrom 16.97s/it serial13m09s
Qwen 3.5:9b (ours)7.42s/itserial1h02m

bbh_cot_zeroshotexact_match, flexible-extract · 6,511 samples

diverse reasoning — logic, algorithms, language

SourceSpeedRuntime
Abliterated¹ (serial)21.82s/it‡39h28m
Abliterated² (bypass)1.16s/itfrom 21.82s/it serial2h05m
Qwen 3.5:9b (ours)0.96s/it16‑way bypass1h44m

mmlu_proexact_match · 12,032 samples

broad academic knowledge, 57 subjects

SourceSpeedRuntime
Abliterated¹ (serial)15.79s/it‡52h48m
Abliterated² (bypass)1.77s/itfrom 15.79s/it serial5h54m
Qwen 3.5:9b (ours)1.15s/it16‑way bypass3h51m

ifevalloose‑instruction acc · 541 samples

instruction-following & formatting compliance

SourceSpeedRuntime
Abliterated¹ (serial)51.00s/it†raw 51.35s/it7h43m
Abliterated² (bypass)3.75s/itfrom 51.35s/it serial33m51s
Qwen 3.5:9b (ours)18.77s/itserial2h49m

cnn_dailymailrouge · 11,490 samples

news article summarization, single ROUGE score

SourceSpeedRuntime
Abliterated¹ (serial)not run
Abliterated² (bypass)1.78s/itserial5h41m
Qwen 3.5:9b (ours)2.11s/itserial6h43m

Where both abliterated¹ and abliterated² rows are filled in, the drop in per‑item rate is entirely the 16‑way bypass replacing serial production Ollama — same weights, same harness, same prompts. It is not evidence about the abliterated vs. official comparison itself.

Outlier‑corrected median, gsm8k/humaneval_instruct/mbpp_instruct_fixed/ifeval only: reconstructed from the desktop adapter's per‑request log for each task's exact time window, then IQR‑filtered to strip contention spikes (9.8–17.5% of that task's wall‑clock, depending on task) before taking the median of what's left. This is the cleaner read on per‑item cost with routine scheduled‑job contention removed.

bbh_cot_zeroshot and mmlu_pro's abliterated¹ rates could not get the same treatment. Reconstructing their windows the same way (each result file's own reported duration, worked backward from its finish time) shows the two runs started within two minutes of each other on 2026‑08‑17 and ran concurrently on claw‑01 for nearly their entire span — a genuine self‑contention bug between two of our own jobs, not just digest, predating the concurrency locks added 2026‑08‑22 to prevent exactly this. No per‑task log survived from that original run and the adapter log doesn't tag requests by task, so their rates can't be cleanly separated after the fact. Treat both raw numbers as inflated by more than ordinary contention.