QWEN3.6-27B · TOOL CALLING

Q4 matched Q8: 94 of 100 selected cases scored correct.

We tested four GGUF quantizations on the same 100 selected BFCL V4 cases. Q8_0, Q5_K_M, and Q4_K_M each scored 94/100. Q3_K_M scored 92/100.

After inspecting the first score, we corrected Qwen name normalization and rescored the unchanged raw outputs. No inference result or selected case changed.

The result

Swipe or use the arrow keys to compare all four columns →

Quantizationsimple_pythonparallel_multipleTotal
Q8_048/5046/5094/100
Q5_K_M48/5046/5094/100
Q4_K_M48/5046/5094/100
Q3_K_M48/5044/5092/100

Q4_K_M matched Q8_0 in both tested categories

Q3_K_M lost 2 cases in parallel_multiple.

Source: corrected exploratory BFCL V4 pilot, published July 27, 2026. 50 selected cases per category and quantization.

Two selected non-live categories only; not a full BFCL leaderboard result.

What the rule said

Our pre-specified “never below Q4” slogan did not survive. Q3_K_M’s 4-point loss in parallel_multiple was not greater than the required 5-point cutoff. That is a failed rule, not evidence that Q3 and Q4 are equivalent.

Limit

This post-result-corrected exploratory pilot covers one model revision and two selected non-live categories. It does not establish quantization equivalence, non-inferiority, full-BFCL performance, or performance on other prompts, runtimes, or hardware.

Download the datasetInspect all 400 scored rowsRead the commands