Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR
Abstract
Multi-Level Intermediate Representation (MLIR) underlies modern ML compiler infrastructure including TensorFlow, JAX via StableHLO, PyTorch, Inductor, IREE, etc yet it appears only in trace amounts in code-LM pretraining corpora. MLIR is also extensible by design: new dialects ship per application domain, so maintaining a fine-tuned model per dialect does not scale. We ask whether inference-time priors derived mechanically from each dialect’s Operation Definition Specification (ODS) can substitute for gradient-based adaptation. We make two contributions. First, we release four natural-language-to-MLIR benchmarks across three dialects; MLIR-Spec-150, Linalg-Spec-30, StableHLO-Spec-30, and StableHLO-Held-Out-200, totaling 410 in-scope NL→MLIR pairs, plus a 25-program StableHLO-Out-Of-Grammar stress set and a hand-authored n=30 functional reference set (435 instances total). All artifacts ship under Apache-2.0 with Gebru datasheets and Croissant 1.0 metadata. Second, on top of these benchmarks we build a three-layer schema-derived constraint stack: a context-free grammar over op signatures (C1), type-domain splits from an ODS-extracted type lattice (C2), and an SSA-scope validator driving five-retry rejection sampling (C3). Porting the stack from arith+func+memref+linalg to StableHLO required no new constraint-layer code. Empirically, on dialects whose verifier semantics are dominated by structural constraints, schema-derived priors let SmolLM2-1.7B match or exceed 15B–34B code LMs at 8–25×the per-generation speed: on linalg, SmolLM2 reaches 80.0% verify-valid (three-seed mean, n=125, every seed 80.0%), beating CodeLlama-34B, Granite-Code-34B, and StarCoder2-15B by 21 to 44 pp with non-overlapping CIs, and surviving a same-family fp16 precision control. On arith+func and on the templated parametric StableHLO-Held-Out-200, where verifier semantics turn on attribute values rather than structure, the same baselines match or beat the SLM; we scope these explicitly as non-win cells. We release benchmarks, decoder, every per-prompt generation, and a reproducibility Docker image.