Can Hybrid-Parallel Planning Support Alternating Model–Strategy Design? Dependency-Keyed Per-Layer Primitives for Re-Planning
Abstract
Production training of large models alternates between model edits and hybrid-parallel strategy planning. Fixed-configuration planners solve each revision from scratch, while unvalidated reuse across edits risks returning a different plan than full re-search would. We introduce a dependency-keyed per-layer primitive interface for re-planning: closed-form layer primitives expose communication, memory, and time to the solver, and the same templates expose the input dependencies needed for reuse-equivalent re-planning. CONCORD attaches dependency sets and value signatures to primitive evaluations, enabling dependency-keyed re-planning to reuse only signature-valid primitive-evaluation work while preserving the same logical candidate set and feasibility predicate as cache-disabled full search. On production MoE strategy-quality runs—including a 4096-device MoE-438B deployment and a 32k-token long-context point—CONCORD's single-call plan achieves 1.14–1.24× training step-time speedup over an expert-tuned Megatron-style baseline (1.70× on the long-context point). Across a 9-fixture sweep at 256–4096 cards, CONCORD's warm re-planning achieves median 2.74× planner wall-clock speedup over cache-disabled full search across 54 in-dispatch cells. Full-search equivalence holds throughout: candidate and feasibility counts match in every cell, the top-1 plan matches in 54/54 cells, and stale-hit mismatches (cached value disagreeing with cold recomputation) are 0/162 over paired trials.