Do Mathematical LLMs Survive Analog Inference?
Abstract
Analog in-memory computing (AIMC) offers a promising path toward energy-efficient LLM inference, but mathematical reasoning poses a particularly demanding deployment setting: small perturbations can propagate across dependent reasoning steps and alter both correctness and solution structure. We study Llama 3.1 8B and Phi-3 Mini under simulated analog inference with AIHWKit on two mathematical reasoning benchmarks: GSM8K and SVAMP. Across additive and PCM-inspired noise, we find structured degradation: reasoning-trace divergence and extra steps are more frequent than explicit arithmetic errors, while output-format reliability can degrade substantially even when final-answer accuracy remains relatively high. Transformer vulnerability is also strongly submodule- and noise-model-dependent. Vulnerability-guided hybrid execution can preserve task accuracy, while retaining LoRA adaptation digitally can improve end-to-end robustness under analog perturbations. These results motivate reasoning-aware evaluation of mathematical LLMs before AIMC deployment, since final-answer accuracy alone can mask important hardware-induced changes in generated solution behavior.