When Does Feedback Help? Auditing Learned Voltage-Clamp Experimental Design
Abstract
A learned scientific controller can outperform a misspecified planner without benefiting from feedback. We audit two synthetic voltage-clamp tasks with a training-selected fixed protocol, current masking, a Bayesian reference, and an excluded target. Five recurrent policies repeat one fixed protocol in all 1,200 off-grid rollouts; masking current changes no decision. In a crossed four-hypothesis task, a fixed diagnostic pulse is followed by one learned choice. Policies then use the first current, but tie a correctly specified Bayesian controller, and a fitted threshold reproduces the learned actions without being that controller. The excluded-target prediction gain is 4.6991 × 10^-6 nats per sample. Feedback use and protocol quality are separate tests. These cases establish neither biological discovery nor general RL superiority.