ML Configuration Artifacts Should Be Closed Under Review
Abstract
As training runs become more expensive and failures more time-consuming to diagnose, configuration mistakes are no longer minor annoyances; they can waste substantial compute, delay iteration, and weaken the experimental record. Yet many configuration systems do not treat reviewability as a main design goal. Our position is that ML experiment configuration artifacts should be closed under review. The experiment specification must be recoverable from the artifact without executing host-language code or running a composition engine. In otherwords, a reviewer should be able to fully understand the experiment by reading the configuration artifact. Popular systems fail this property because meaning is distributed across helper code and runtime context, and we argue this failure is structural rather than incidental. We organize the configuration design space along two axes, authoring freedom and semantic locality, and identify recurring antipatterns that arise when locality is weak or the configuration surface becomes too expressive. Closure is achieved by high semantic locality combined with authoring surfaces bounded enough to exclude host-language execution. Constrained text-first wiring representations meet both conditions. We present pfig, a Python object-wiring DSL, as evidence that this design point is practical.