1. Choose a label to flip
Select any context point (X, Y) for one label flip. 0 = outside; 1 = inside.
Same scale and X/Y range: −1.2 to +1.2; top/right are positive. Dots: context rows; gold circle: flipped row; faint dotted circle: synthetic label rule. Boundaries follow the recorded grid classes; no interpolation. Reach rings can extend beyond the map edge.
The full 16-flip finding
Across all 16 flips, one wrong label changed Kumo’s predicted class at a median of 13.5 of 169 grid points, up to 1.70 units from the flipped row. k-NN(3) changed a median of 8.5, never beyond 0.58 units. Logistic regression changed at most 1.
| Model | Changed classes (of 169) | Largest probability shift (percentage points) | Farthest class change (synthetic units) |
|---|---|---|---|
| Kumo | 13.5 / 14 | 22.3 / 29.3 | 1.32 / 1.70 |
| k-NN(3) | 8.5 / 11 | 33.3 / 33.3 | 0.58 / 0.58 |
| Logistic regression | 0 / 1 | 19.2 / 26.2 | 0.00 / 0.42 |
One tiny two-feature synthetic table, one seed, one checkpoint and fixed recipes. Kumo before-map retained from v1. No v1 inference rerun. Recipes differ in capacity and inductive bias; logistic regression’s linear boundary is mismatched to this circular-label task, and equidistant k-NN neighbors can depend on the fixed row order. Lower changed-class counts are not evidence of superior robustness. No accuracy, calibration or general robustness claim. Farthest reach is distance from the flipped row to a query whose class changes, in synthetic units. No changed classes: “none”; 0 in aggregate distance statistics.
2. Inspect any query
Updates are immediate, with no animation. All 16 label buttons and both query sliders work with a keyboard.
What the contrast shows
In this fixed grid, Kumo’s label flips changed classes farther from the flipped row than the classical comparators did at the median and maximum. k-NN had larger maximum probability shifts, while logistic regression usually kept the same classes. These are different behaviors, not a ranking of model quality. Every flip and unchanged query is retained.
Logistic regression uses a linear boundary on a circular-label task, so its low changed-class count is not a robustness win. Its probabilities can move while its predicted class stays outside. k-NN uses a local vote; Kumo uses its default preprocessing and pretrained weights. These recipes are not matched in capacity or inductive bias.
No Kumo weights were updated. The classical estimators were refitted for each condition. Fixed row order matters for equidistant k-NN neighbors. The model’s output probabilities are not validated confidence estimates.
Settings & provenance
All 16 flips and reach values
Each model lists changed-class count for all 16 flips in row order. Full probabilities and all three reach metrics are in the raw JSON.
Source and file hashes
Nearest comparator: Classifier comparison
The scikit-learn Classifier comparison gallery example contrasts classical decision boundaries on two-dimensional synthetic datasets. This lens adds an in-context tabular model, all 16 single-label interventions and inspectable saved probabilities. It is narrower: one table, fixed inputs and no accuracy comparison.
The NVIDIA release article and model card are architecture and benchmark sources. The community WebGPU Space by mrfakename is an additional source; its predictions are untested here.
Verification and execution instructions · Frozen v2 protocol · Pre-execution hashes · Class-order amendment history