JSLBench: Where Do LLMs Fail in Statistical Programming?
Abstract
Statistical code generated by large language models (LLMs) can execute successfully while failing to implement the requested analysis. When such analyses inform consequential decisions, evaluation must expose the analytical choices a program makes, not only whether it runs. We introduce JSLBench, a benchmark and evaluation framework for JMP Scripting Language (JSL), a specialized statistical language underrepresented in existing benchmarks. JSLBench contains 10K natural-language-to-JSL tasks, each pairing a request with dataset context and a reference script previously executed in JMP; LLM-as-judge prompt validation and expert review of sampled tasks support task quality. On JSLBench-Diverse, a 500-task subset spanning 220 datasets, eight model and context configurations pass all five automatic checks, from response generation through JMP execution (end-to-end success), on 49.8% to 86.6% of tasks. Adding a JSL skill or prompt restrictions yields net gains in end-to-end success but also regressions on individual tasks. Rule-based categorization of the 1,228 end-to-end failures attributes most to unresolved names, argument faults, and object-model or column errors rather than syntax. Expert code inspection rates no program that passes every automatic check unsatisfactory, but notes partial fulfillment or code-quality concerns on 23 of 69, including variable-role choices that execution cannot check. JSLBench pairs traceable execution evidence with expert review of analysis requirements as validation, and provides a foundation for evaluating LLM-generated JSL and developing statistical agents.