Agentic Job Scheduling for Scientific Distributed Computing via Active Knowledge Extraction
Abstract
Large language model (LLM) agents have recently emerged as a promising paradigm for iterative optimization, yet most existing agent benchmarks provide only coarse environmental feedback, such as scalar rewards. In contrast, many real-world systems naturally generate rich execution traces that explain why decisions succeed or fail. We study this setting in distributed scheduling for High Energy Physics (HEP) computing infrastructures, where execution logs capture detailed information about resource utilization, queueing delays, data transfers, task failures, and system bottlenecks. We propose Active Knowledge-Extraction Agent (AKEA), an agentic scheduler that actively extracts knowledge from execution logs while jointly evolving scheduling policies and reusable knowledge-extraction skills. Experiments on realistic HEP scheduling workloads show that AKEA consistently outperforms representative heuristic, neural, and recent LLM-based schedulers. Our results suggest that rich environmental feedback should also be treated as a source of reasoning rather than merely an optimization signal.