X PAPER / Full editionIssue 002 · Morning candidate中文

AI & Technology / Background

Addition in words improves—with a larger inference budget

Source author: Simon Willison · @simonw
Source posted:
Written and edited by

Simon Willison’s October 4 experiment, shared on X October 6 Beijing time, asks a local four-bit Qwen3.8-27B model to express integer sums only in English words. Correct formatting is distinct from correct arithmetic.

A non-reasoning sweep scored 1,195/5,070, or 23.57%, across digit-length pairs from one to thirteen. A separate medium-reasoning run scored 167/169, or 98.82%. Different samples prevent treating those percentages as a controlled comparison.

On the same frozen 169 questions, non-reasoning scored 45 and reasoning 167. However, reasoning strength, output allowance and execution order changed: the non-reasoning run allowed 128 output tokens; the reasoning run allowed longer completion. This compares configurations rather than isolating a causal reasoning effect.

Median latency rose from 1.50 to 27.62 seconds and median completion tokens from 14 to 313. The setup used DGX Spark and Qwen3.8-27B-Q4_K_M. Those costs and conditions bound the result; we have not rerun it or generalized it to other tasks.

Sources and further reading

Edited report · Sources and limitations in the text

Selected through a verified followed account: @simonw.