kapynResearch

Qwen3.8 27B addition in words

Qwen 3.8 27B tests how LLMs handle arithmetic word outputs. The experiment benchmarks addition accuracy across varying digit lengths when models must return the sum spelled out in words rather than digits. Developers can view detailed heatmap charts and raw logs to analyze failure modes in tokenization and numerical reasoning.

Simon Willison·Oct 4, 2026

Opening Kapyn…