Can a Meta-Agent Choose the Safer Branch? A Controlled Audit of Pre-Commit Fork Selection
Abstract
Forking an agent is useful only if its supervisor can decide which branch to release. We study this missing pre-commit decision: given two candidate events from the same task, state, and execution prefix, can a meta-agent select the better-supported continuation—and does one more quarantined observation help? ForkAudit evaluates this question on AFTraj-Fork, a controlled set of task-unique strict-prefix pairs whose candidates share the task, state, and execution prefix but differ in outcome construction. Across three fixed open-weight auditors, one matched observation improves the same pairwise gate by a mean of 6.3 percentage points. This improvement is not clean evidence of semantic oversight: task-blind probes and auditors recover the reference label from construction cues without access to the task. Yet a no-shadow absolute scorer remains stronger, and selective dual-order decisions retain high conditional error risk. ForkAudit therefore treats the observation as an incremental benchmark gain, not a deployment-ready safety capability, and evaluates commit gates by attacking their interpretation and measuring risk–coverage rather than significance alone.