Availability Is Not Utilization: Trace-Based Evaluation of Tool Allocation in Foundation-Model Agents
Abstract
Autonomous foundation-model agents produce rich runtime traces as they deliberate, call tools, recover from failures, and attempt task completion. Final success rates compress these trajectories into a single outcome and can hide important behavioral differences. We introduce ThinkCompute, a brokered framework for trace-based behavioral interpretability in which neural deliberation, Python/SymPy, SMT, and Lean actions have explicit operational semantics and final correctness is independently verified. Across three model configurations, we find substantial differences in voluntary tool allocation, post-tool behavior, and the effect of integrating deterministic computation into the reasoning process. The same allocation policy degrades verified solving for one configuration, has little effect on another, and improves performance for the strongest configuration, producing a large cross-model policy interaction. Trace analysis separates selection, execution, uptake, repair, and verification, showing that models can behave differently even when tool execution rates are similar. These results show that agent behavior cannot be inferred from tool availability or final accuracy alone, and motivate evaluation methods that treat runtime trajectories as first-class empirical evidence.