QWEN3.6-27B · TOOL CALLING
Q4 matched Q8: 94 of 100 selected cases scored correct.
We tested four GGUF quantizations on the same 100 selected BFCL V4 cases. Q8_0, Q5_K_M, and Q4_K_M each scored 94/100. Q3_K_M scored 92/100.
After inspecting the first score, we corrected Qwen name normalization and rescored the unchanged raw outputs. No inference result or selected case changed.
The result
| Quantization | simple_python | parallel_multiple | Total |
|---|---|---|---|
| Q8_0 | 48/50 | 46/50 | 94/100 |
| Q5_K_M | 48/50 | 46/50 | 94/100 |
| Q4_K_M | 48/50 | 46/50 | 94/100 |
| Q3_K_M | 48/50 | 44/50 | 92/100 |
Q4_K_M matched Q8_0 in both tested categories
Q3_K_M lost 2 cases in parallel_multiple.
Source: corrected exploratory BFCL V4 pilot, published July 27, 2026. 50 selected cases per category and quantization.
Two selected non-live categories only; not a full BFCL leaderboard result.
What the rule said
Our pre-specified “never below Q4” slogan did not survive. Q3_K_M’s 4-point loss in parallel_multiple was not greater than the required 5-point cutoff. That is a failed rule, not evidence that Q3 and Q4 are equivalent.
Limit
This post-result-corrected exploratory pilot covers one model revision and two selected non-live categories. It does not establish quantization equivalence, non-inferiority, full-BFCL performance, or performance on other prompts, runtimes, or hardware.