Don’t Learn What You Can Compute: Arithmetic Residual Blocks for Exact Arithmetic in Transformers
Abstract
Large language models spend enormous parameter capacity approximating arithmetic, a deterministic function with an exact solution, and still fail reliably on multi-digit operations. We propose Arithmetic Residual Blocks (ARBs): frozen, differentiable modules inserted into the transformer forward pass that compute exact integer arithmetic using Residue Number System (RNS) encoding on unit circles. The model learns only the interface, a gated injection pathway and a low-rank adapter on the language model head, while the base model and all arithmetic computation remain frozen. The architecture embodies a design principle we call Don't Learn What You Can Compute: any deterministic function expressible as tensor operations can be embedded as a frozen residual block, and the model will learn to route through it because the gradient reward for exact answers dominates internal approximation. We validate this principle by inserting ARBs into a frozen SmolLM2-360M model, training only 1.7M parameters (0.47% of the base model). On exact-match evaluation across addition, subtraction, multiplication, and division with operands up to 3 digits, the augmented model achieves 99.9%+ accuracy across all operations, compared to 4.6% mean accuracy for the unmodified base model. A logit analysis of the frozen base model confirms that the injection pathway overrides the base model's prior uniformly across operations, and that the residual errors (0.047%) concentrate at later digit positions in autoregressive generation, not at any particular operation. These findings motivate the hypothesis that pre-training with ARBs, where no misaligned frozen representations exist, would yield ceiling performance across all operations.