From Enterprise Pipelines to Agent Context: Distilling Compact Models for Schema Lineage
Abstract
Enterprise data pipelines encode operational knowledge in multilingual scripts, yet coding and data agents struggle to recover the long-range dependencies, transformations, and aggregations needed to interpret derived columns. We present a compact-model framework that converts production pipeline code into target-specific schema lineage for downstream agents. GPT-4.1 transforms expert annotations into reasoning traces and structured answers. Traces whose final answers score above 0.95 under the Schema Lineage Composite Evaluation (SLiCE) are used for supervised fine-tuning of Qwen2.5-Coder-7B, followed by group relative policy optimization with Graded SLiCE, a reward that retains hard validity constraints while providing component-level partial credit. The resulting 7B model reaches GPT-4o-level pipeline-macro SLiCE (0.5843 vs. 0.5764) under the same evaluator. Supplying its lineage to a text-to-SQL agent increases execution accuracy on lineage-dependent questions from 27.4% with matched metadata but no lineage to 75.0%. These results show that a compact, locally deployable task-specialized model can recover structured knowledge from enterprise code and make it useful to downstream agents.