ReaXpert: Evolutionary Prompt Optimization for Factor-Level Reaction Condition Reasoning
Abstract
Adapting a chemistry agent by fine-tuning is not always practical: experimental labels are costly, frontier-model weights may be unavailable, and parameter updates give an expert no direct record of what changed. ReaXpert instead specializes a frozen model with Genetic-Pareto (GEPA), an evolutionary prompt optimizer that diagnoses failures from natural-language execution traces and revises the instruction. To our knowledge, this is the first application of GEPA to reaction chemistry. The result is a readable adaptation artifact that a chemist can inspect and amend, produced for approximately USD 14 in model calls without weight access or accelerator hardware. The evolved prompts predict yield through six precedent-grounded chemistry factors. This structure is motivated by a measured failure mode: direct prediction recovers only 7.9% of true-low reactions. The named factors expose where that optimism arises and provide a surface for explicit calibration. Appending failure-attributed calibration rules raises true-low recall to 48.3% and narrows the per-class recall spread from 52.5 to 5.3 percentage points. The complete prompt pipeline comes within 0.4 percentage points of an open-weight low-rank adaptation (LoRA) fine-tune on exact accuracy, exceeds it on macro-averaged F1, and requires no weight training. The same factor reasoning also supplies a chemically informed warm start for condition optimization.