Longtail-SWE: High-Quality SWE Tasks for Underrepresented Languages
Abstract
Challenging repository-level benchmarks remain limited for many programming language ecosystems. When candidate tasks are scarce, filtering out instances with ambiguous descriptions or defective tests can discard valuable engineering problems. We introduce Longtail-SWE, a benchmark of 237 issue-resolution tasks across 13 languages with limited coverage in existing repository-level benchmarks: Clojure, Dart, Elixir, Erlang, F#, Haskell, Julia, Kotlin, Lua, OCaml, R, Scala, and Swift. Its construction uses a solver-judge-repair pipeline: a solver attempts each candidate task, and a judge reviews the task together with the solver's trajectory and test results. Its findings guide revisions to tests and task descriptions. These revisions aim to correct tests that reject valid alternative implementations or accept incorrect solutions while retaining challenging tasks and preserving their intended requirements. We evaluate seven model configurations, including frontier models, and analyze task-solving performance, consistency across repeated attempts, and resource usage. The release includes all benchmark instances, executable environments, audit and repair tooling, and agent trajectories, together with 296 additional, easier tasks to support evaluation of smaller models.