LongBanana: An Expert-Verified Benchmark for Long-Context Multi-Reference Image Synthesis
Abstract
We introduce LongBanana, to our knowledge the \textbf{first expert-designed and expert-verified benchmark} for \emph{long-context multi-reference image synthesis} in the age of \textbf{agent-controlled image generation}. The task requires retrieving dispersed visual evidence, rejecting distractors, binding entities and relations, and composing one coherent final image from multiple references. LongBanana contains 90 carefully designed tasks organized by a 3-by-3 taxonomy; each task includes about 32 reference images, 20 constraints, and 50 evaluation items, yielding \textbf{4{,}670 expert-specified atomic evaluation items} in total. Its graph-grounded evaluation design instantiates each task with reference images and an entity-relation graph, then converts node, edge, and whole-scene obligations into checklist items for Content Fidelity, Relational Binding, and Global Constraint Satisfaction. Reviewed expert labels include a 15\% double-labeled subset with \textbf{89.4\% exact agreement} and \textbf{0.83 quadratic weighted kappa}. We further introduce LongBanana-Gen, an agentic controller framework for long-horizon image generation; experiments show that minimal agent harnesses improve over direct generation, while ReAct/Plan-Exec/Reflect yields controller-dependent gains by strengthening evidence grounding and repair. Together, LongBanana provides an auditable testbed for measuring dense visual evidence use and a shared target for developing stronger agent-controlled image-generation systems.