QWEN3.6-27B · TOOL CALLING

Q4 matched Q8: 94 of 100 selected cases scored correct.

We tested four GGUF quantizations on the same 100 selected BFCL V4 cases. Q8_0, Q5_K_M, and Q4_K_M each scored 94/100. Q3_K_M scored 92/100.

After inspecting the first score, we corrected Qwen name normalization and rescored the unchanged raw outputs. No inference result or selected case changed.

The result

Quantizationsimple_pythonparallel_multipleTotal
Q8_048/5046/5094/100
Q5_K_M48/5046/5094/100
Q4_K_M48/5046/5094/100
Q3_K_M48/5044/5092/100

Q4_K_M matched Q8_0 in both tested categories

Q3_K_M lost 2 cases in parallel_multiple.

Source: corrected exploratory BFCL V4 pilot, published July 27, 2026. 50 selected cases per category and quantization.

Two selected non-live categories only; not a full BFCL leaderboard result.

What the rule said

Our pre-specified “never below Q4” slogan did not survive. Q3_K_M’s 4-point loss in parallel_multiple was not greater than the required 5-point cutoff. That is a failed rule, not evidence that Q3 and Q4 are equivalent.

Limit

This post-result-corrected exploratory pilot covers one model revision and two selected non-live categories. It does not establish quantization equivalence, non-inferiority, full-BFCL performance, or performance on other prompts, runtimes, or hardware.

Download the datasetInspect all 400 scored rowsRead the commands