QWEN3.6-27B · TOOL CALLING
Q4 matched Q8: 94 of 100 selected cases scored correct.
We tested four GGUF quantizations on the same 100 selected BFCL V4 cases. Q8_0, Q5_K_M, and Q4_K_M each scored 94/100. Q3_K_M scored 92/100.
After inspecting the first score, we corrected Qwen name normalization and rescored the unchanged raw outputs. No inference result or selected case changed.
The result
Swipe or use the arrow keys to compare all four columns →
| Quantization | simple_python | parallel_multiple | Total |
|---|---|---|---|
| Q8_0 | 48/50 | 46/50 | 94/100 |
| Q5_K_M | 48/50 | 46/50 | 94/100 |
| Q4_K_M | 48/50 | 46/50 | 94/100 |
| Q3_K_M | 48/50 | 44/50 | 92/100 |
Q4_K_M matched Q8_0 in both tested categories
Q3_K_M lost 2 cases in parallel_multiple.
Source: corrected exploratory BFCL V4 pilot, published July 27, 2026. 50 selected cases per category and quantization.
Two selected non-live categories only; not a full BFCL leaderboard result.
What the rule said
Our pre-specified “never below Q4” slogan did not survive. Q3_K_M’s 4-point loss in parallel_multiple was not greater than the required 5-point cutoff. That is a failed rule, not evidence that Q3 and Q4 are equivalent.
Limit
This post-result-corrected exploratory pilot covers one model revision and two selected non-live categories. It does not establish quantization equivalence, non-inferiority, full-BFCL performance, or performance on other prompts, runtimes, or hardware.