AI & Technology / Background
Addition in words improves—with a larger inference budget
Simon Willison’s October 4 experiment, shared on X October 6 Beijing time, asks a local four-bit Qwen3.8-27B model to express integer sums only in English words. Correct formatting is distinct from correct arithmetic.
A non-reasoning sweep scored 1,195/5,070, or 23.57%, across digit-length pairs from one to thirteen. A separate medium-reasoning run scored 167/169, or 98.82%. Different samples prevent treating those percentages as a controlled comparison.
On the same frozen 169 questions, non-reasoning scored 45 and reasoning 167. However, reasoning strength, output allowance and execution order changed: the non-reasoning run allowed 128 output tokens; the reasoning run allowed longer completion. This compares configurations rather than isolating a causal reasoning effect.
Median latency rose from 1.50 to 27.62 seconds and median completion tokens from 14 to 313. The setup used DGX Spark and Qwen3.8-27B-Q4_K_M. Those costs and conditions bound the result; we have not rerun it or generalized it to other tasks.
Sources and further reading
Edited report · Sources and limitations in the text
Selected through a verified followed account: @simonw.