Can We Trust What We Measure? Auditing LLM Validity-Awareness in Predictive Modeling
Jennifer Zhang ⋅ Amanda L Coston
Abstract
Practitioners increasingly use large language models (LLMs) to write the code behind predictive models, including in societally high-stakes settings such as child-welfare screening, care-management enrollment, and job-training assignment. A structural failure common to these settings, decision dependency, can make predictions systematically wrong even when every standard metric looks fine. Measuring whether an LLM recognizes and handles it is itself hard: held-out accuracy cannot detect the failure, since a naive model is an accurate estimate of the wrong quantity, and a model may appear to recognize it by recalling a published critique rather than reasoning about the deployment context in front of it. We build and validate a measurement pipeline that survives these traps — audited memorization control, and a cross-vendor LLM-as-judge agreeing with human labels at Cohen's $\kappa \geq 0.8$ in 19 of 20 vendor/domain/dimension cells — then use it for a five-dimension study across three domains, two model families, and a prompt hint ladder from explicit instruction to realistic practitioner framing. Verbal recognition of the failure is common but frequently fails to translate into correct code, and the confidence a model volunteers unasked carries almost no signal about whether it succeeded. Asked afterwards whether its output is the quantity a decision-maker needs, a model that got it wrong usually declines to rate it at all — and that refusal, not the rating that accompanies it, is the signal.
Chat is not available.
Successful Page Load