ConjectureLineage: Learning Mathematical Taste from Documented Follow-Up
Abstract
While mathematical taste is central to a mathematician's choice of research questions, it remains largely unaddressed in automated reasoning. We define it as the ability to identify questions that are "interesting" to the mathematical community because their investigation can produce useful mathematics, including methods and related results that leave the original question open. We operationalize that judgment through documented mathematical developments and train a detector of this evidence-based notion of taste. We introduce CONJECTURELINEAGE, an expert-reviewed, LLM-assisted dataset of 586 algebraic conjectures and questions with three-year follow-up evidence. Its score combines direct progress, related mathematical contributions, and a small research-attention term. A logistic readout trained on frozen GPT-Neo features reaches 0.7040 AUROC, versus 0.4914 for the untrained readout, and cuts the log-score error by 72.9%. This 1.3B frozen backbone is comparable to frontier zero-shot ranking on the same questions while predicting the reviewed score more accurately than all twelve reported language models. Our longer-term plan is to expand and improve the dataset, study whether small critics can teach larger models, and invite mathematicians to assess whether the resulting models pose more worthwhile questions.