OrthogonalAgent: An Auditable Control Layer for Agentic Statistical Reasoning
meng tang
Abstract
Adaptive analysis can invalidate ordinary confidence intervals even when each candidate is familiar. We present \method{}, an auditable control layer for tool-mediated causal workflows that records data-use permissions and reconstructs inference, separating fixed-procedure from selection validity. Our primary controlled statistical evaluation uses scripted simulations. Separately, actual LLM proposers made real API tool calls in exploratory experiments with gpt-4o-mini, gpt-4o, gpt-5.6-luna, and gpt-6-luna. The formal Luna follow-up planned 1,020 original episodes: 1,005 reached a terminal status (579 completed, 388 abstained, 38 recorded as blocked) and 15 did not; 15 linked GPT-6 retries (10 completed, 3 abstained, 2 recorded as blocked) are post-hoc attempts, not added independent episodes. Earlier exploratory protocol traces had condition contamination and no verified final-report gate; the protocol specifies corrected conditions, but its typed-plan, feedback, and release-gate package does not isolate a verifier-only effect. In 500 frozen exact-null repetitions, the raw, pre-verifier false-positive rate for same-data maximum-$|t|$ search rose from 0.036 at one candidate to 0.166 at 13 candidates; the verifier rejected all 500 such runs. Bonferroni, selection splitting, and fold-local selection yielded rates of 0.018, 0.066, and 0.072, respectively (their Monte Carlo intervals overlap). Correction cannot repair invalid base inference; the handcrafted 72-record suite is a rule-aligned regression test. These results are bounded evidence, not a new theorem or general agent-safety claim.
Chat is not available.
Successful Page Load