agent-vitals: Section-Level Interpretability Under Harness Evolution
Abstract
Harness optimizers learn from an agent's outputs and task outcomes, but they need additional tooling to observe how the agent uses its context. Existing interpretability tools generally require access to probe the model during serving, while a harness optimizer may retain only logs. We introduce \texttt{agent-vitals}, a package for section-level interpretability of agentic harnesses around open-weight models. The harness declares and annotates its own sections, and after each benchmark Vitals replays the logged episodes through a reference copy of the model to produce sectional measures of where the agent reads, how concentrated that reading is, and how it changes between confident and uncertain steps. Across harnesses generated by an external optimizer agent, which we call the proposer, the measurements reveal systematic patterns in how and when agents read different parts of the scaffold. We also show that sectional measures make internal evidence available for optimization. In matched harness optimization runs, the proposer cited statistics generated by Vitals in every hypothesis and used them to justify its placement choices.