pAI/MSc: Automating ML-Theory Research with Humans on the Loop
Mahmoud Abdelmoneum ⋅ Pierfrancesco Beneventano ⋅ Tomaso Poggio
Abstract
Advances in LLM-based models have resulted in agents and agentic systems that can search the literature, write code, run experiments, and draft papers, yet difficult research still depends on judgments that current models make unreliably. These include choosing a worthwhile problem, recognizing when a proof attempt has failed, and deciding whether the evidence supports the paper's main claim. We present pAI/MSc, a human-on-the-loop multiagent system for automating machine-learning theory research. Given a research question, specialist agents develop the literature review, theory, experiments, and manuscript, while a human researcher monitors the work and intervenes at a small number of high-level checkpoints. The system also records the intermediate work supporting each claim so that unsupported conclusions can be revised before submission. We evaluate its planning and theoretical-reasoning components on existing benchmarks and study one complete research project on the Muon optimizer. On planning tasks, pAI/MSc scores higher than two of three single-agent baselines, although an 80-task ablation finds no average benefit from the multiagent planning group alone. In the Muon project, seven high-level human interventions guided a 31-day run that produced nearly $29{,}000$ recorded downstream actions and outputs. An early draft proposed an incorrect implicit-bias theorem. Later verification rejected that theorem and replaced it with a narrower result supported by a proof and numerical checks. This case illustrates both the promise and the current limit of human-on-the-loop research automation: agents can carry out substantial research work between human interventions, but expert judgment remains essential to the scientific direction and validity of the final paper.
Chat is not available.
Successful Page Load