Beyond the Trial-and-Error Loop: Hybrid Projection and Automated Tuning for Distributed Training
Abstract
Configuring distributed LLM training across hundreds of GPUs requires jointly optimizing tensor, pipeline, expert, context, and data parallelism alongside micro-batch size, recomputation, sharding, and overlap — a combinatorial space where the boundary between memory-legal and OOM can hinge on a single byte of per-element activation accounting. Today, teams navigate it through expensive multi-node trial and error, burning GPU-hours on crashes and suboptimal recipes. We present Primus Projection, which closes this loop with two coupled components. First, a hybrid projection engine that measures what it can and simulates what it cannot: GPU microbenchmarks on as few as a single device capture compute kernels and any in-scope communication; a sub-node downscale–measure–upscale workflow analytically restores out-of-scope collectives, FSDP overlap, and the full pipeline schedule at the multi-node target. A pure simulation mode runs entirely on CPU, enabling pre-silicon planning without any GPU. Second, a tuning agent treats the projection engine as a fast scoring oracle — seconds per candidate instead of tens of minutes — and uses it to drive an intelligent search over the configuration space. A deterministic seed planner sweeps single axes; an LLM investigation loop proposes cross-axis combinations the planner cannot reach, with an analytic memory pre-filter that rejects infeasible candidates before any GPU time is spent. Validated on Llama 3.1 (dense) and Mixtral 8×22B (MoE) across two GPU generations, all projections stay within ~10% of measured multi-node throughput. In a Mixtral 8×22B case study, the agent discovers a configuration delivering +27% throughput over the published 4-node BF16 reference at the same scale — in under 30 minutes of single-node exploration, with no hand-written configs and no full-cluster profiling pass.