Comparing human-AI and AI-AI teams in multi-turn layout creation
Abstract
Co-creating layouts such as websites and slides is a useful testbed for evaluating multi-turn instruction following of LLM systems. Since an intended layout’s appearance cannot be well expressed in a single turn, the user must wait for the LLM to generate a layout then iteratively refine it with the LLM over multiple turns. We present a small-scale, exploratory (n=6) study of this multi-turn process to compare the differences of human-AI and AI-AI teams on the same layout- generation task. Given a reference wire-frame layout, a participant or a language model ‘speaker‘ issues natural-language instructions to a fixed ‘listener‘ LLM over multiple turns until the speaker is satisfied. Rathar than an LLM as the judge, we fit a scoring function over geometric layout features using curated human preference data (n=19). We find that LLMs tend to encode the entire layout in the first turn, while humans iterate over many turns on average, writing much shorter instructions. We also find that human-AI teams consistently scored lower than AI-AI teams. This gap is driven by a limiter of human participants in specifying precise sizes and positions (e.g. 42 pixels to the right). Humans perform almost at par with language models on other features like ordering the elements into nested structures. We then perform interventions on the speaker LLMs to make them more human-like, and find it difficult. This raises the challenge of using LLMs as user models to evaluate multi-turn co-creation tasks.